Decentralization distance correlation network construction algorithm based on data driving

Through a data-driven decentralized distance correlation network construction algorithm, the center of mass difference and random disturbance test are used to automatically screen lung cancer-specific network signals, and the parameters dependence and complexity problems of deep learning models in lung cancer diagnosis are solved, achieving the accuracy and biological explanatory improvement of early diagnosis of lung cancer.

CN120432009APending Publication Date: 2025-08-05ANSHAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510303226.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing deep learning network model relies on prior knowledge in the early diagnosis of lung cancer, resulting in low overfitting and marker specificity, making it difficult to accurately understand the dynamic changes in cancer development, and the network construction mechanism is complex and difficult to biologically explain, affecting the diagnostic effect.

Method used

The data-driven decentralized distance correlation network construction algorithm is used to construct t statistics and random disturbance tests through the centroid difference method, and the specific network signals at different stages of lung cancer are automatically determined, prospective warning markers for lung cancer are screened, biological networks are constructed and topological analysis is carried out.

Benefits of technology

Without manual parameters setting, it can effectively reflect the robustness of inter-molecular correlation, improve the accuracy of early diagnosis of lung cancer, discover key early warning markers, deeply understand the pathogenic mechanism of lung cancer, and improve clinical diagnosis effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120432009A_ABST
    Figure CN120432009A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological data analysis, in particular to a decentralized distance correlation network construction algorithm based on data driving, which comprises the following steps: inputting omics data of different stages of lung cancer occurrence and development, defining distance correlation of a pair of molecular characteristics fi and fj on a sample t, decentralization distance correlation is used for measuring the difference change of the intermolecular incidence relation in the lung cancer occurrence and development process, the centroid difference between the local decentralization distance correlation and the global decentralization distance correlation of the features fi and fj on the ck class of samples is calculated, random disturbance testing is used for measuring the robustness degree of the molecular features to the centroid difference of the (fi, fj), and the robustness degree of the molecular features to the centroid difference of the (fi, fj) is calculated. Constructing a biological network; the method does not need priori knowledge for manual setting of threshold parameters, is helpful for accurately understanding the occurrence and development process of the lung cancer, focuses on specific information of a network topology structure in the early stage of the lung cancer, screens more effective prospective early warning markers of the lung cancer, and improves the early diagnosis effect of the lung cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biological data analysis, and in particular to a data-driven decentralized distance correlation network construction algorithm. Background Art

[0002] Lung cancer, primarily consisting of adenocarcinoma and squamous cell carcinoma, poses a serious threat to human health. Over the past 20 years, clinical treatments such as surgery, radiotherapy, or a combination of these have been frequently used to reduce mortality in lung cancer patients. However, long-term survival remains far from ideal, with the overall 5-year survival rate for patients with advanced lung cancer less than 20%. Delays in diagnosis and treatment are often the primary factors contributing to the increased mortality and mortality of patients with advanced lung cancer. Because lung cancer often lacks obvious symptoms in its early stages, most patients present with symptoms at advanced stages, missing the optimal time for treatment. Therefore, prospective early warning biomarker screening and clinical early diagnosis of lung cancer have become hot topics. However, the development and progression of lung cancer involves a complex interplay of multiple factors, including genetic inheritance, environmental factors, and metabolic disorders, and its pathogenic mechanisms remain largely unclear. Furthermore, individual variability in biochemical reactions, such as gene regulation and metabolic phenotypes, can vary from patient to patient. This individual variability directly impacts the accuracy of data analysis and biomarker selection in clinical lung cancer research, ultimately impacting the effectiveness of early diagnosis and drug therapy.

[0003] In living organisms, disturbances in key metabolic activities such as non-essential amino acid metabolism, proline metabolism, and glucose metabolism can induce tumor cells to use metabolic reprogramming biochemical reactions to achieve uncontrolled proliferation needs. Changes in the concentration and type of molecules caused by disturbances in metabolic pathways can also lead to carcinogenesis in the body. For example, aerobic glycolysis with low energy production efficiency and its intermediates can promote the rapid growth, proliferation, and invasion of tumor cells. In addition, related metabolic mechanisms, such as metabolic stress activation in tumor-infiltrating immune cells, can affect the functional activity of immune cells and the anti-tumor immune response, thereby allowing tumor cells to evade the surveillance of the immune system. Therefore, a comprehensive analysis of abnormal changes in metabolic reaction activities during the development and progression of lung cancer will help to better understand the pathogenesis of lung cancer and improve clinical diagnosis, monitoring, and treatment.

[0004] In living organisms, molecules at different levels, such as genes and metabolites, interact with each other to ensure the normal functioning of various metabolic and physiological activities. During the development and progression of cancer, metabolic reactions related to cancer are synergistically carried out by genes and metabolites. For example, aberrant expression of genes regulating metabolic reactions can disrupt metabolic pathways and metabolites in cells that are closely related to tumor growth and transformation. In turn, metabolic abnormalities influence and alter epigenetic processes and interfere with the activity of biochemical reactions such as transcriptional regulators, leading to abnormal gene expression. Genomics and metabolomics both play important roles in abnormal cell differentiation, rapid invasive growth, and early metastasis in cancer. In oncology research, genomics focuses on exploring DNA sequencing variations associated with cancer development and progression, analyzing cancer-specific mutations and chromosomal rearrangements to classify or subtype cancers. Because gene mutations significantly increase the risk of cancer, genomics is widely used in cancer risk prediction, patient diagnosis, and precision medicine. Metabolomics, as the end product of gene expression, focuses on changes in endogenous small molecules to understand the underlying mechanisms of cancer and identify key biomarkers. Because metabolites are regulated by genes and are generally considered products of cellular biochemical reactions, metabolic abnormalities can be used to directly characterize changes in tumor cell function. The organic integration of genomics and metabolomics, leveraging biological information at different levels within the body, can provide a more comprehensive description of lung cancer, from genes to phenotypes. This will help better understand the development and progression of lung cancer and address the challenges of early clinical diagnosis and treatment.

[0005] Although deep learning network models based on models like Transformer and AlphaFold have achieved some success in clinical cancer research, their numerous layers and complex structures require extensive parameter configuration. For example, parameters such as the number of training cycles, attention weights, and loss functions must be set in advance based on prior knowledge. With the rapid development of high-throughput technologies, an increasing amount of massive medical big data is being used in early clinical cancer diagnosis research. However, different data often require different parameter settings. Improper parameter settings hinder accurate understanding of the dynamics of cancer development and progression, and can also lead to overfitting, resulting in low specificity of selected biomarkers and, consequently, ineffectiveness in improving early clinical cancer diagnosis. Furthermore, the complex network construction mechanisms within deep learning network models often hinder effective biological analysis and interpretation based on clinical knowledge. Furthermore, network model construction is affected by weakly robust intermolecular correlations. Adding or removing even a small number of samples, or even a single sample, can significantly impact the calculated intermolecular correlations, leading to suboptimal clinical application of the selected biomarkers. Consequently, the resulting results fail to effectively characterize the differences in biochemical activity between pre- and post-malignant states, hindering further research into cancer pathogenesis and early diagnosis. Summary of the Invention

[0006] To address the above issues, the present invention provides a data-driven decentralized distance correlation network construction algorithm. It uses the centroid difference method to construct a decentralized distance correlation t-statistic to systematically compare and analyze the changing patterns of intermolecular correlations at different stages of lung cancer, and adopts a random perturbation test method to measure the robustness of intermolecular correlations. It automatically determines the specific network signals of different stages of lung cancer in a data-driven manner, thereby helping to deeply understand the occurrence and development process of lung cancer, focusing on the specific information of the network topology structure in the early stages of lung cancer, and screening more effective prospective warning markers for lung cancer, in order to improve the early diagnosis of lung cancer.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] A data-driven decentralized distance correlation network construction algorithm includes the following steps:

[0009] S1, set F = {f1, f2, ..., f m} is defined as a molecular feature set, m represents the number of molecular features; X={x1,x2,...,x n} is defined as a sample set, n represents the number of samples; C={c1,c2,…,c z} is defined as the class label set, z represents the number of class labels;

[0010] S2. Input omics data of different stages of lung cancer development;

[0011] S3, a pair of molecular features f i and f j The distance correlation on sample t is defined as Dc ijt =||x it –x jt ||, where x it is the feature f i The expression value on sample t, x jt is the feature f j Expression value on sample t;

[0012] S4. Use decentralized distance correlation to measure the differential changes in molecular associations during the development and progression of lung cancer. The calculation method is as follows:

[0013]

[0014] Among them, Dc ipt is the molecular feature f i and f p Distance correlation on sample t, Dc qjt is the molecular feature f q and f j Distance correlation on sample t, Dc pqt is the molecular feature f p and f q distance correlation on sample t;

[0015] S5. Definition Represents feature f i and feature f j In c k The average value of the decentralized distance correlation of class samples, u ij Represents feature f i and feature f j The average value of the decentralized distance correlation over all samples, then the feature f i and feature f j In c k The centroid difference between the local decentralized distance correlation and the global decentralized distance correlation on the class samples is calculated as follows:

[0016]

[0017] Among them, σ ij is the feature f i and feature f j The within-class standard deviation of the decentralized distance correlation is calculated as follows:

[0018]

[0019] Where σ0 is σ ij median,e k is the balance coefficient;

[0020] S6. Use random perturbation test to measure molecular feature pairs (f i ,f j ) The robustness of the centroid difference, after num perturbation tests, is based on the p-value ijk The value measures the molecular characteristics of the pair (f i ,f j ) Robustness of centroid difference, p-value ijk The value is calculated as follows:

[0021]

[0022] When p-value ijk When it is less than 0.05, it is considered that the molecular feature pair (f i ,f j )’s centroid differences are robust;

[0023] Among them, count ijk Represents molecular feature pairs (f i ,f j ) The center of mass difference is the number of times it is unstable, and num is the number of perturbation tests;

[0024] S7. Building Biological Networks: Using G k =(V(G k ),E(G k ),W(G k )) indicates lung cancer k The weighted undirected biological network constructed in the stage, V(G k )=F is the node set, is the edge set, W(G k ) is the weight set of edges, w(f i ,f j ) is the biological network graph G k Middle node f i With node f j The weight of the edge; if w(f i ,f j )≥ε and p-value ijk <0.05, then the node f i and node f j Use edges of one color to connect them; if w(f i,f j )≤-ε and p-value ijk <0.05, then the node f i and node f j They are connected by edges of another color; where ε is the edge weight threshold.

[0025] Furthermore, the balance coefficient e k Set to When, e k ×σ ij for The estimated standard error of k For c k The number of class samples, therefore, for a pair of molecular features (f i ,f j ) The centroid difference values of different stages of lung cancer It conforms to the t distribution.

[0026] Furthermore, the molecular feature pair (f i ,f j ) The method for determining whether the centroid difference is unstable is: in each disturbance test, the feature f i and feature f j The decentralized distance values on all samples will be randomly reassigned and the centroid difference value after perturbation will be calculated If hour, or when hour, It means that the molecular feature pair (f i ,f j ) The center of mass difference is unstable.

[0027] Furthermore, it also includes a network topology analysis method. In the constructed biological network diagram, the star-shaped subgraph composed of the node with the largest degree in the biological network and the nodes directly connected to it is extracted as an important network signal, and the node with the largest degree is screened as a potential warning marker for lung cancer.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1) Due to It complies with the t distribution, that is, there is only one peak pv. When ε=pv, the node f i and node f jEdges only exist in the network graph of a certain stage of lung cancer, but not in the network graphs of other stages. Therefore, they can be used as specific information of lung cancer in that stage. Since the peak PV is determined by the data itself, the setting of the threshold ε is automatically determined based on data-driven methods, without the need for manual parameter setting based on prior knowledge.

[0030] 2) In systems biology, distance correlation can effectively reflect the closeness of pathway reaction activities between molecular features and can measure both linear and nonlinear relationships between molecules, which is of great biological significance. Therefore, the biological network constructed based on the distance correlation t statistic can be effectively analyzed and interpreted biologically using clinical knowledge, thereby making the potential biomarkers screened have practical clinical application value.

[0031] 3) Using random perturbation testing can effectively eliminate the influence of weak and robust intermolecular associations on data analysis, thereby more accurately reflecting the changes in related metabolic mechanisms during the development and progression of lung cancer, providing assistance for in-depth research on the pathogenic mechanism of lung cancer and providing important targets for the early diagnosis of lung cancer;

[0032] 4) In-depth analysis of the dynamic changes of specific networks in different stages of lung cancer from a temporal perspective will help further understand the occurrence and development mechanism of lung cancer. Focusing on the changing patterns of specific network signals in the early stages of lung cancer and conducting comprehensive comparison and analysis with specific network signals in other stages can help discover more effective prospective warning information for lung cancer, thereby improving the clinical early diagnosis of lung cancer.

[0033] 5) The design concept and method principles of the present invention have good universality and can also be applied to the pathogenic mechanism and early clinical diagnosis research of other cancers or complex diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is the gene network subgraph SG of the living organism in the healthy stage in the embodiment of the present invention. I topological structure.

[0035] Figure 2 This is the gene network subgraph SG of the living organism in stage I of lung cancer in the embodiment of the present invention. I topological structure.

[0036] Figure 3 This is the gene network subgraph SG of the living organism in stage II of lung cancer in the embodiment of the present invention. I topological structure.

[0037] Figure 4 This is the gene network subgraph SG of the living organism in stage III lung cancer in the embodiment of the present invention.I topological structure.

[0038] Figure 5 This is the gene network subgraph SG of the living organism in stage IV of lung cancer in the embodiment of the present invention. I topological structure.

[0039] Figure 6 It is the metabolic network subgraph SMN of the living organism in the healthy stage in the embodiment of the present invention LC topological structure.

[0040] Figure 7 It is the metabolic network subgraph SMN of the living organism in the benign tumor stage of lung cancer in the embodiment of the present invention. LC topological structure.

[0041] Figure 8 It is the metabolic network subgraph SMN of the living organism in the malignant tumor stage of lung cancer in the embodiment of the present invention. LC topological structure. DETAILED DESCRIPTION

[0042] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings:

[0043] The present invention proposes a data-driven decentralized distance correlation network construction algorithm, which takes lung cancer genomics data and lung cancer metabolomics data as research objects, discovers early diagnostic markers for lung cancer as research goals, and uses dynamic differential biological network construction and analysis as research methods to explore changes in metabolic activity mechanisms during the occurrence and development of lung cancer, thereby helping to promote clinical diagnostic research and application of lung cancer; in living organisms, the correlations between different omics molecules are complex and diverse, and systematic and comprehensive exploration of the differential changes in the correlations between different omics molecules during the occurrence and development of lung cancer will help to screen prospective warning signals that can effectively warn of the occurrence of lung cancer, thereby improving the clinical early diagnosis effect of lung cancer.

[0044] The present invention proposes a data-driven decentralized distance correlation network construction algorithm DCN. This algorithm uses decentralized distance correlation to construct a t-statistic and combines it with a random perturbation test to establish a specific network that can effectively and robustly reflect the different stages of canceration in the body. It also automatically screens specific network markers for early-stage lung cancer in a data-driven manner. The algorithm specifically includes the following steps:

[0045] S1, set F = {f1, f2, ..., f m} is defined as a molecular feature set, m represents the number of molecular features; X={x1,x2,...,x n} is defined as a sample set, n represents the number of samples; C={c1,c2,…,c z} is defined as the class label set, and z represents the number of class labels.

[0046] S2. Input omics data of different stages of lung cancer development and progression.

[0047] S3, a pair of molecular features f i and f j The distance correlation on sample t is expressed by Euclidean distance or other calculation formula, which is defined as Dc ijt =||x it –x jt ||, where x it is the feature f i The expression value on sample t, x jt is the feature f j The expression value on sample t.

[0048] S4. To reduce the impact of differences in molecular expression values on data analysis, decentralized distance correlation was used to measure the differential changes in molecular associations during the development and progression of lung cancer. The calculation method is as follows:

[0049]

[0050] Among them, Dc ipt is the molecular feature f i and f p Distance correlation on sample t, Dc qjt is the molecular feature f q and f j Distance correlation on sample t, Dc pqt is the molecular feature f p and f q The distance correlation on sample t.

[0051] S5. Definition Represents feature f i and feature f j In c k The average value of the decentralized distance correlation of class samples, u ij Represents feature f i and feature f j The average value of the decentralized distance correlation over all samples, then the feature f i and feature f j In c k The centroid difference between the local decentralized distance correlation and the global decentralized distance correlation on the class samples is calculated as follows:

[0052]

[0053] Among them, σ ij is the feature fi and feature f j The within-class standard deviation of the decentralized distance correlation is calculated as follows:

[0054]

[0055] Where σ0 is σ ij The median is used to further eliminate the impact of differences in the magnitude of molecular expression values on data analysis;

[0056] e k is the balance coefficient, when the balance coefficient e k Set to When, e k ×σ ij for The estimated standard error of k For c k Therefore, for a pair of molecular features (f i ,f j ) The centroid difference values of different stages of lung cancer It conforms to the t distribution.

[0057] S6. Use random perturbation test to measure molecular feature pairs (f i ,f j ) The robustness of the centroid difference; in each perturbation test, the feature f i and feature f j The decentralized distance values on all samples will be randomly reassigned and the centroid difference value after perturbation will be calculated If hour, or when hour, It means that the molecular feature pair (f i ,f j ) The centroid difference is unstable. After num perturbation tests, the p-value ijk The value measures the molecular characteristics of the pair (f i ,f j ) Robustness of centroid difference, p-value ijk The value is calculated as follows:

[0058]

[0059] Among them, count ijk Represents molecular feature pairs (f i ,f j ) The centroid difference is unstable, num is the number of disturbance tests; when p-value ijkWhen it is less than 0.05, it is considered that the molecular feature pair (f i ,f j ) has strong robustness to the centroid difference.

[0060] S7. Building Biological Networks: Using G k =(V(G k ),E(G k ),W(G k )) indicates lung cancer k The weighted undirected biological network constructed in the stage, V(G k )=F is the node set, is the edge set, W(G k ) is the weight set of edges, w(f i ,f j ) is the biological network graph G k Middle node f i With node f j The weight of the edge; if w(f i ,f j )≥ε and p-value ijk <0.05, then the node f i and node f j Use edges of one color to connect them; if w(f i ,f j )≤-ε and p-value ijk <0.05, then the node f i and node f j They are connected by edges of another color; where ε is the edge weight threshold.

[0061] The dynamic differential biological network constructed by the present invention can more comprehensively and accurately characterize the changes in the molecular association relationships under different physiological and pathological conditions of the body. Dynamic analysis of the topological structure of the constructed biological network helps to discover important network signals and potential biomarkers that can warn of the early occurrence of lung cancer, provide assistance for further understanding the pathogenic mechanism of lung cancer and early and accurate diagnosis of patients, and provide theoretical guidance for research in the fields of disease genomics data analysis, translational medicine, etc.

[0062] It also includes a network topology analysis method. In the dynamic differential biological network diagram constructed by the present invention, the node with the largest degree indicates that the molecular feature participates in multiple pathway reaction activities during the occurrence and development of lung cancer, and is at the central hub of the biochemical reaction, which is of great significance to the understanding of the body's lung cancer process and pathogenic mechanism; the star-shaped subgraph composed of the node with the largest degree and its connecting edges is extracted as an important network early warning signal for lung cancer. The change in its topological structure can intuitively reflect the change in the molecular association relationship in the early stage of lung cancer, thereby helping to promote the research on the occurrence mechanism and early diagnosis of lung cancer. The present invention screens the node with the largest degree as a potential early warning marker, in order to effectively improve the clinical early diagnosis effect of lung cancer.

[0063] The following examples are implemented under the premise of the technical solution of the present invention, and provide detailed implementation methods and specific operating processes, but the scope of protection of the present invention is not limited to the following examples. The methods used in the following examples are conventional methods unless otherwise specified.

[0064] [Example]

[0065] The present invention proposes a data-driven decentralized distance correlation network construction algorithm for screening early diagnostic markers for lung cancer based on genomic data and metabolomics data related to body metabolism.

[0066] (1) Collection of lung cancer genomic data related to body metabolism

[0067] The training set of lung cancer genomic data in this experiment was derived from the Cancer Genome Atlas database, including 108 healthy samples N in the control group and 1005 samples M in the model group. The model group samples consisted of 520 stage I lung cancer samples, 284 stage II lung cancer samples, 168 stage III lung cancer samples, and 33 stage IV lung cancer samples. Based on functional enrichment analysis, a total of 875 genes were closely related to the body's metabolic reaction activities. Therefore, these 875 genes were used for subsequent network construction and analysis. The validation set of lung cancer genomic data in this experiment was derived from the Gene Expression Omnibus database, consisting of data sets GSE33532, GSE27262, and GSE75037, to further verify the clinical early diagnostic effect of potential lung cancer markers screened based on the training set.

[0068] (2) Collection of lung cancer metabolomics data related to body metabolism

[0069] The training set of lung cancer metabolomics data in this experiment consists of 30 healthy samples HC from the control group, 6 benign lung cancer samples BC, and 30 malignant lung cancer samples LC. The training set contains a total of 158 metabolites. In addition, the validation set of lung cancer metabolomics data in this experiment consists of another 35 control group samples HC and 35 malignant lung cancer samples LC.

[0070] (3) Construction of dynamic differential biological networks based on lung cancer genomics data

[0071] In the lung cancer genomics data, the training set model includes healthy samples of the control group, samples of lung cancer stage I, samples of lung cancer stage II, samples of lung cancer stage III, and samples of lung cancer stage IV. i ,f j ) in each sample and p-value ijk Build a dynamic network and output 5 network diagrams, see Figure 1-5 .

[0072] (4) For the dynamic network of lung cancer genes, G I The biological network constructed based on stage I lung cancer samples can be used more effectively to discover the prospective warning gene information of lung cancer and extract G I The star-shaped subgraph SG consisting of the five nodes with the largest median degree and the nodes directly connected to them I As a prospective network warning signal for lung cancer, the five nodes with the largest degree are used as potential gene markers for early diagnosis of clinical lung cancer;

[0073] Figure 1-5 SG is the gene network subgraph of the organism at different stages of lung cancer I The topological structure of Figure 1 In the figure, the decentralized distance correlation of the selected genes in the healthy state is negatively correlated, that is, the color of the edges is green. Figure 2 In the figure, the decentralized distance correlation of the selected genes in the lung cancer stage I state is positively correlated, that is, the color of the edges is red, while in Figure 3 、 Figure 4 and Figure 5 That is, the selected genes show weak correlation in the status of lung cancer stage II, lung cancer stage III and lung cancer stage IV, that is, there is no edge. Therefore, the gene network subgraph SG I Changes in topological structure can effectively warn of the occurrence of lung cancer.

[0074] (5) Construction of dynamic differential biological networks based on lung cancer metabolomics data

[0075] In the lung cancer metabolomics data, the training set model contains three different types of samples, based on the molecular feature pairs (fi,f j ) in each sample and p-value ijk Build a dynamic network and output 3 network diagrams, see Figure 6-8 .

[0076] (6) For the dynamic metabolic network of lung cancer, MN LC The biological network constructed based on lung cancer malignant tumor samples can be used more effectively to discover clinical diagnostic markers for lung cancer and extract MN LC The star-shaped subgraph SMN consisting of the five nodes with the largest median degree and the nodes directly connected to them LC As a prospective metabolic network warning signal for lung cancer, the five nodes with the largest degree are used as potential diagnostic metabolic markers for clinical lung cancer;

[0077] Figure 6-8 SMN is the metabolic network subgraph for different stages of lung cancer LC The topological structure changes, Figure 6 In the healthy state, the number of edges in the network is small and mainly concentrated on the two metabolites of aspartic acid and xanthine. The other metabolites show weak correlations, so there are no connecting edges. Figure 7 In the benign state of lung cancer, the number of edges in the network is also small and mainly concentrated on the metabolite 3-hydroxypropanoic acid. Figure 8 In the case of lung cancer malignancy, the selected metabolites all showed strong correlations, that is, the number of edges increased suddenly, and the color of the edges was consistent with the Figure 7 In contrast, the metabolic network subgraph SMN LC The changes in the topological structure of the mitochondria can effectively characterize the different stages of lung cancer development.

[0078] (7) The area under the curve (AUC) was used to verify the clinical diagnostic effect of the potential gene markers and metabolic markers of lung cancer screened by the present invention. Tables 1, 2 and 3 respectively give the comparison results of the present invention and other methods on genomic and metabolomics data. The comparative experimental results show that the potential markers screened by the present invention exhibit good discrimination ability in both genomic and metabolomics training sets and validation sets.

[0079] Table 1 Comparison results of different methods on genomic data training set

[0080]

[0081]

[0082] Table 2 Comparison results of different methods on the genomics data validation set

[0083]

[0084] Table 3 Comparison results of different methods on metabolomics data training set and validation set

[0085]

[0086]

Claims

1. A data-driven decentralized distance correlation network construction algorithm, characterized by: The steps include: S1, set F = {f1, f2, ..., f m } is defined as a molecular feature set, m represents the number of molecular features; X={x1,x2,...,x n } is defined as a sample set, n represents the number of samples; C={c1,c2,…,c z } is defined as the class label set, z represents the number of class labels; S2. Input omics data of different stages of lung cancer development; S3, a pair of molecular features f i and f j The distance correlation on sample t is defined as Dc ijt =||x it –x jt ||, where x it is the feature f i The expression value on sample t, x jt is the feature f j Expression value on sample t; S4. Use decentralized distance correlation to measure the differential changes in molecular associations during the development and progression of lung cancer. The calculation method is as follows: Among them, Dc ipt is the molecular feature f i and f p Distance correlation on sample t, Dc qjt is the molecular feature f q and f j Distance correlation on sample t, Dc pqt is the molecular feature f p and f q distance correlation on sample t; S5. Definition Represents feature f i and feature f j In the c k The average value of the decentralized distance correlation of class samples, u ij Represents feature f i and feature f j The average value of the decentralized distance correlation over all samples, then the feature f i and feature f j In the c k The centroid difference between the local decentralized distance correlation and the global decentralized distance correlation on the class samples is calculated as follows: Among them, σ ij is the feature f i and feature f j The within-class standard deviation of the decentralized distance correlation is calculated as follows: Where σ0 is σ ij median,e k is the balance coefficient; S6. Use random perturbation test to measure molecular feature pairs (f i ,f j ) The robustness of the centroid difference, after num perturbation tests, is based on the p-value ijk The value measures the molecular characteristics of the pair (f i ,f j ) Robustness of centroid difference, p-value ijk The value is calculated as follows: When p-value ijk When it is less than 0.05, it is considered that the molecular feature pair (f i ,f j )’s centroid differences are robust; Among them, count ijk Represents molecular feature pairs (f i ,f j ) The center of mass difference is the number of times it is unstable, and num is the number of perturbation tests; S7. Building Biological Networks: Using G k =(V(G k ),E(G k ),W(G k )) indicates lung cancer k The weighted undirected biological network constructed in the stage, V(G k )=F is the node set, is the edge set, W(G k ) is the weight set of edges, w(f i ,f j ) is the biological network graph G k Middle node f i With node f j The weight of the edge; if w(f i ,f j )≥ε and p-value ijk <0.05, then the node f i and node f j Use edges of one color to connect them; if w(f i ,f j )≤-ε and p-value ijk <0.05, then the node f i and node f j They are connected by edges of another color; where ε is the edge weight threshold.

2. The data-driven decentralized distance correlation network construction algorithm according to claim 1 is characterized in that: The balance coefficient e k Set to When, e k ×σ ij for The estimated standard error of k For c k The number of class samples, therefore, for a pair of molecular features (f i ,f j ) The centroid difference values of different stages of lung cancer It conforms to the t distribution.

3. The data-driven decentralized distance correlation network construction algorithm according to claim 1 is characterized in that: The molecular feature pair (f i ,f j ) The method for determining whether the centroid difference is unstable is: in each disturbance test, the feature f i and feature f j The decentralized distance values on all samples will be randomly reassigned and the centroid difference value after perturbation will be calculated If hour, or when hour, It means that the molecular feature pair (f i ,f j ) The center of mass difference is unstable.

4. The data-driven decentralized distance correlation network construction algorithm according to claim 1, characterized in that: It also includes a network topology analysis method. In the constructed biological network diagram, the star-shaped subgraph composed of the node with the largest degree in the biological network and the nodes directly connected to it is extracted as an important network signal, and the node with the largest degree is screened as a potential warning marker for lung cancer.