Omics data analysis method combining molecular network and sample network

By constructing a multi-view molecular network and combining the topological characteristics evaluation of the sample network, key modules are identified and support vector machine classifiers are established, the problem of ignoring biological molecules in traditional methods is solved, and the analysis effect of omics data is improved.

CN120260673APending Publication Date: 2025-07-04DALIAN UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510430857.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing omics data analysis methods often ignore the interaction between biomolecules, resulting in the inability to adequately capture key biological information, and traditional feature selection methods are not effective in high-dimensional complex omics data.

Method used

A multi-perspective molecular network is constructed, and a fast-great community detection algorithm is used to divide the modules, combine the topological characteristic evaluation module of the sample network to identify key modules, and classifiers are built through support vector mechanisms to integrate multi-perspective information to improve the accuracy of feature selection.

Benefits of technology

Through multi-perspective analysis and module division, the feature subspace with strong distinction capabilities is identified, which improves the classification performance of omics data and provides more reliable data analysis methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260673A_ABST
    Figure CN120260673A_ABST
Patent Text Reader

Abstract

The invention provides an omics data analysis method combining a molecular network and a sample network, and belongs to the technical field of omics data analysis. According to the method, the interaction relationship between omics data features (molecules) is measured from multiple perspectives, a molecular network is constructed, and a network module is determined by adopting a rapid-greedy community detection algorithm; a topological characteristic evaluation module of the sample network is utilized to identify a key module with distinguishing capability; according to the method, feature subspaces containing rich biological information are determined by fusing key modules obtained from multiple perspectives, a high-performance SVM classifier is established for biomics data analysis, a practical and effective data analysis means is provided for research of genomics, metabonomics, proteomics and other omics data, and the method has high application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of omics data analysis, measures the synergistic effects between molecules (features) from multiple perspectives, constructs a molecular network, and uses the fast-greedy community detection algorithm to divide the network modules; evaluates the modules by using the topological characteristics of the sample network, and selects a feature subset with strong discrimination ability, which is an omics data analysis method that comprehensively uses molecular networks and sample networks. Background Art

[0002] With the development of high-throughput biotechnology, the accumulation of omics data (such as genomics, transcriptomics, proteomics, and metabolomics data) has increased explosively. These data provide rich data support and unprecedented opportunities for understanding the pathogenesis, subtype classification, and prognosis evaluation of diseases. Omics data has the characteristics of small sample size, high dimension, much noise, and highly complex information. Therefore, effective analysis and mining of omics data are of great significance for disease research, precision medicine, etc.

[0003] Feature selection is a commonly used omics data processing method, aiming to identify and extract key features from high-dimensional and complex omics data to reduce the data dimension and improve the interpretability and prediction performance of the model. Traditional machine learning feature selection methods include the Relief series, SVM-RFE, genetic algorithms, etc. However, these methods often only focus on the discrimination ability of a single biomolecule itself, ignoring the interactions between biomolecules, and may not be able to fully capture the key biological information contained in these interactions. Currently, omics data analysis methods based on network analysis have received increasing attention. By measuring the interaction relationships of biomolecules to construct a molecular network, it can more comprehensively reveal the correlations between molecules, and determine the key molecular modules and key network nodes related to the disease pathogenesis mechanism. However, the interactions between molecules in biological systems are complex and diverse. Examining the synergistic effects between molecules (features) from multiple dimensions helps to comprehensively understand the molecular interaction mechanism.

[0004] The present invention measures the interaction relationships between omics data features from multiple perspectives and constructs a molecular network; uses the fast-greedy community detection algorithm to divide the network modules to ensure that the correlations between features within the modules are close and the correlations between features between modules are sparse. The present invention constructs a sample network based on the determined feature subset within the module, evaluates the discrimination ability of the module by using the topological characteristics of the sample network, and identifies the key modules. Then, the key modules from multiple perspectives are fused to determine an important feature subspace containing rich biological information, and a high-performance classification model is established. Summary of the Invention

[0005] The object of the present invention is to establish an omics data analysis method combining a molecular network and a sample network, measure the complex interaction relationships between molecules from multiple perspectives, establish molecular networks based on different perspectives, determine sub-network markers under different perspectives, obtain an information subspace after multi-perspective information fusion, and establish a sample classification model. This technology is applicable to the analysis and research of omics data, comprehensively excavates important information in omics data from different angles, and can be used in fields such as omics data analysis and precision medicine. The core technology of this method is to propose a new omics data analysis method combining a molecular network and a sample network.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] An omics data analysis method combining a molecular network and a sample network, in which serum, urine, and tissue are used as samples, and genes, proteins, and metabolites are used as features according to omics data. The following steps are adopted:

[0008] (1) Construct a multi-perspective molecular network

[0009] Let S = {S1, S2, …, S n} be a sample set, F = {f1, f2, …, f m} be a feature (biomolecule) set, n be the number of samples, and m be the number of features. Y = (Y1, Y2, …, Y n ) is a label vector of n samples, where y i ∈ {0, 1} (i = 1, 2, …, n), 0 represents a normal sample, and 1 represents a disease sample.

[0010] First, based on the disease samples in S and the feature set F, the Spearman correlation coefficient, the maximum information coefficient, and the cosine similarity are respectively used to measure the correlation relationships between features, and a multi-perspective molecular network is established. The Spearman correlation coefficient is a rank correlation coefficient, which measures the rank correlation between two features and reflects the monotonicity of the correlation between features; while the maximum information coefficient is a concept based on information theory, which measures the relationship by finding the maximum mutual information between features and can capture various functional relationships between features, including linear and non-linear relationships; the cosine similarity starts from the perspective of the vector space, regards features as vectors, and calculates the cosine value of the angle between two vectors. Through these three measurement methods, the complex and diverse interaction relationships between features are comprehensively excavated.

[0011] For two features (molecules) f i and f j (1 ≤ i ≠ j ≤ m), let SCC(f i , f j ) represent the feature f i and fj The Spearman correlation coefficient between MIC(f i , f j ) represents the maximum information coefficient between feature f i and f j The cosine similarity between Cos(f i , f j ) represents the feature f i and f j The specific calculation formulas are shown in (1), (2), and (3):

[0012]

[0013] Among them, represents the difference in the order of feature f i and f j on the k-th sample, and n d represents the total number of disease samples; the larger the absolute value of SCC(f i , f j ) represents the stronger the Spearman correlation between feature f i and f j . MI ij represents the mutual information between feature f i and f j . div i and div j represent the number of grid partitions in the directions of f i and f j respectively. B is a variable whose value is taken as max(n α , 4), where α is default set to 0.6. The larger the value of MIC(f i , f j ) represents the stronger the correlation between feature f i and f j . Similarly, the larger the absolute value of Cos(f i , f j ) represents the stronger the correlation between feature f i and f j .

[0014] The vertex set of the molecular network established under each perspective is F. For each feature f i (1 ≤ i ≤ m), the top k features with the strongest correlation with it are selected to establish connections, and the weight of the edge is equal to the corresponding correlation value; there are no connections with the remaining features, and the weight is 0. Let the three molecular networks constructed based on disease samples be the Spearman correlation coefficient network the maximum information coefficient network and the cosine similarity network where Vm =F represents the node set, F is the feature set, and They represent the edge sets calculated based on the Spearman correlation coefficient, maximum information coefficient, and cosine similarity, respectively.

[0015] (II) Dividing modules

[0016] For the three molecular networks, the fast-greedy community detection algorithm was used to divide the modules. Modularity is an indicator to measure the strength of the network module structure. For example, assume that the module set M is divided into S ={m S1 ,m S2 ,…,m Sk},m Si Represents a module, which contains at least one feature (node), and each feature belongs to and only belongs to one module.

[0017] Modularity Q S The calculation formula is:

[0018] in, It's the network The number of edges in the middle; A is The adjacency matrix of the network Medium i and f j If there is an edge between them, then A ij =1; otherwise, A ij =0; k i is the feature (node) f i degree; and Respectively represent the features (nodes) f i and f j The module to which it belongs, when f i and f j In the same module, otherwise

[0019]

[0020] Initially, each feature (node) is regarded as an independent module; then, the modules are gradually merged through a greedy optimization strategy: the modular increment ΔQ of each pair of modules after merging is calculated, and a pair of modules that maximizes ΔQ and ΔQ ≥ 0 are selected for merging, and this step is repeated until there are no modules that make ΔQ ≥ 0 available for merging.

[0021] (III) Identifying key modules

[0022] Sample networks are established respectively based on the features contained in each module. The cosine similarity is used to measure the similarity between samples, and the edges of the sample network are established in the KNN manner. Let the sample network established based on the features contained in one of the modules be G s =(V s , E), where V s =S, E = {(s i , s j ) | s i , s j ∈V s , i ≠ j, s i ∈N(s j ) or s j ∈N(s i )}, where N(s i ), N(s j ) represent the sets of the neighboring nodes of the samples (nodes) s i and s j respectively. If the discriminative ability of the feature subspace is strong, then in this subspace, the edges between the same-class samples are dense, and the edges between different-class samples are sparse. Therefore, the difference between the number of edges between the same-class samples and the number of edges between different-class samples in the sample network established based on the module is used as the score of the module: Score = |E intra | - |E inter | (5)

[0023] Then, for each perspective, the modules are sorted according to the module scores from high to low, and T modules with the highest scores and the number of features contained not exceeding 20% of |F| are selected, where |F| represents the total number of features.

[0024] Suppose T1, T2, and T3 key modules are selected from the molecular networks based on three different perspectives respectively, and the feature subsets under the three perspectives are determined accordingly: and represent the modules sorted by score, and Features(·) represents the set of features contained in the module.

[0025] (IV) Establishing a classifier based on the feature subspace

[0026] The three feature subsets obtained respectively from the perspectives of considering the monotonic relationship, generalized correlation, and direction consistency between features are rich in different biological information. The union of the three feature subsets is taken to obtain a fused feature subspace to ensure the richness of information in the feature subspace. A support vector machine (SVM) is applied on the fused feature subspace to construct a classifier, and the data is classified to obtain the final result.

[0027] Support Vector Machine (SVM), as a commonly used machine learning classification method, aims to find the best separating hyperplane to distinguish samples of different categories. This hyperplane should have the largest margin with the "support vectors" to classify the training samples. The SVM method with a linear kernel function is selected in this invention.

[0028] To address the impact of small sample size and high dimensionality of omics data samples on the effectiveness of analysis methods, starting from the complex and diverse characteristics of intermolecular interaction relationships, this invention adopts multiple perspectives to comprehensively and systematically analyze the complex molecular relationships in genomics, metabolomics and other omics data, eliminating the limitations of a single perspective; combining molecular networks and sample networks, using the differences in the association strengths between similar samples and between dissimilar samples in the sample network to evaluate the quality of molecular network module partitioning, determining the characteristic subspace with strong discrimination ability, and providing a new technology for the analysis and processing of omics data. Experimental results on multiple different omics public datasets show that the data analysis method combining molecular networks and sample networks proposed in this invention is superior to other comparative data analysis methods in terms of classification performance. Through theoretical and experimental analysis, this invention can provide reliable data analysis means for the research of omics data such as metabolomics, genomics and proteomics, and has strong application value. Brief Description of the Drawings

[0029] Figure 1 It is the overall architecture diagram of this invention. Detailed Description of the Invention

[0030] The following takes the miRNA dataset of human gastric cancer as an example to detail this invention in combination with the technical solution and the drawings.

[0031] The omics data used in this embodiment is the public miRNA dataset of human gastric cancer. After effective analysis and processing of the samples by relevant biological analysis techniques, the dataset contains a total of 28 pairs of miRNA microarrays of gastric cancer tissues and matched normal mucosal tissues, with 817 features, fully meeting the characteristics of small sample size and high dimensionality of the omics dataset targeted by this invention.

[0032] 1. Construct a multi - perspective molecular network

[0033] There is an existing sample set S = {S1, S2, …, S 56}, feature set F = {f1, f2, …, f 817}, Y = (y1, y2, …, y 56 ) is the label vector of 56 samples, where y i ∈{0, 1} (i = 1, 2, …, 56), 0 represents normal samples, and 1 represents gastric cancer samples. For any two features f i and f j(1 ≤ i ≠ j ≤ 817), based on the gastric cancer samples in S, the Spearman correlation coefficients SCC(f i , f j ), the maximum information coefficient MIC(f i , f j ), and the cosine similarity Cos(f i , f j ) are calculated using formulas (1), (2), and (3) respectively; for each measurement method, for each feature f i (1 ≤ i ≤ 817), the three features with the strongest correlation with it are selected to establish connections, and the weight of the edge is the corresponding correlation value. Based on this, three molecular networks are constructed, namely the Spearman correlation coefficient network the maximum information coefficient network and the cosine similarity network where the node set V m = F represents miRNA features, and represent the edge sets calculated based on the Spearman correlation coefficient, the maximum information coefficient, and the cosine similarity respectively. The weights of the edges in the three networks are w S (f i , f j ) = SCC(f i , f j ), w M (f i , f j ) = MIC(f i , f j ) and w C (f i , f j ) = Cos(f i , f j ).

[0034] 2. Module Partition

[0035] The fast - greedy community detection algorithm is used to partition the modules of the molecular networks and respectively. Initially, each feature (node) is regarded as an independent module, and the modularity value is calculated using formula (4). Then, two modules are tried to be merged, and the modularity value after the merge is calculated using formula (4) and compared with the modularity value before the merge to obtain the difference. A pair of modules that can maximize the modularity increment ΔQ and ΔQ ≥ 0 is selected for merging, and this step is repeated until there are no modules available for merging that can make ΔQ ≥ 0.

[0036] 3. Identification of Key Modules

[0037] Sample networks are established respectively based on the features included in each module. The cosine similarity is used to measure the similarity between samples. For each sample, the seven samples with the strongest correlation with it are selected to establish connections, and the sample network is constructed. Then, formula (5) is used to calculate the difference between the number of connections between samples of the same type and the number of connections between samples of different types in the network as the score of the module. For all modules under the same perspective, the modules are sorted according to the module scores from high to low, and the number of features included in each module is counted.

[0038] For any perspective, assume that the upper limit of the number of selectable features is |F| * 20% (i.e., 163 features). Initialize the set of selected modules SM and the set of selected features SF as empty sets. Traverse the sorted modules. If the total number of features is still less than 163 after adding the features of the current module to SF, then add this module to SM and add its features to SF; otherwise, stop traversing. If no module is selected after traversing, then add the module with the highest score to SM and add its features to SF. Finally, based on the molecular networks of the three perspectives, important feature subsets Subset_SCC, Subset_MIC, and Subset_Cos are obtained respectively.

[0039] 4. Establish a classifier based on the feature subspace

[0040] Take the union of the three obtained feature subsets to get the fused feature subspace, and apply the support vector machine (SVM) method on the feature subspace to construct a classifier for data classification. The linear kernel function is selected for the SVM method.

[0041] The following table shows the comparison of the classification performance of the method of the present invention (MAS-SVM) and other common data analysis methods in bioinformatics data analysis (including SVM-RFE, DBN, and GRACES) on five bioinformatics public datasets. The experiment uses ten-fold ten-fold cross-validation and uses accuracy (ACC), sensitivity (SEN), and specificity (SPE) as performance evaluation indicators. The bold fonts in the table represent the optimal performance of each method on each dataset. W / T / L represent the number of times the comparison method wins, ties, and loses compared with MAS-SVM. The experimental results show that MAS-SVM has achieved better classification performance in terms of ACC, SEN, and SPE indicators, verifying the effectiveness of the method of the present invention.

[0042] Table 1 Comparison of the accuracy of MAS-SVM and other effective methods

[0043]

[0044] Table 2 Comparison of the sensitivity of MAS-SVM and other effective methods

[0045]

[0046] Table 3 Comparison of Specificity between MAS-SVM and Other Effective Methods

[0047]

Claims

1. An omics data analysis method combining a molecular network and a sample network, characterized in that The steps are as follows: (1) Construct a multi-perspective molecular network Let \(S = \{S_1, S_2, \ldots, S n \}\) be the sample set, \(F = \{f_1, f_2, \ldots, f m \}\) be the feature (biomolecule) set, \(n\) be the number of samples, and \(m\) be the number of features; \(Y=(Y_1, Y_2, \ldots, Y n )\) is the label vector of \(n\) samples, where \(y i \in\{0, 1\}(i = 1, 2, \ldots, n)\), \(0\) represents a normal sample, and \(1\) represents a disease sample; First, based on the disease samples and feature set F in S, the Spearman correlation coefficient, the maximum information coefficient, and the cosine similarity are used to measure the correlation relationships between features respectively, and a multi-perspective molecular network is established. The Spearman correlation coefficient is a rank correlation coefficient that measures the rank correlation between two features and reflects the monotonicity of the correlation between features. The maximum information coefficient is based on the concept of information theory and measures the relationship by finding the maximum mutual information between features, which can capture various functional relationships between features, including linear and non-linear relationships. The cosine similarity, from the perspective of vector space, regards features as vectors and calculates the cosine value of the angle between two vectors. Through these three measurement methods, the complex and diverse interaction relationships between features can be comprehensively mined. For two features (molecules) f i and f j (1 ≤ i ≠ j ≤ m), let SCC(f i , f j ) denote the Spearman correlation coefficient between features f i and f j , MIC(f i , f j ) denote the maximum information coefficient between features f i and f j , and Cos(f i , f j ) denote the cosine similarity between features f i and f j . The specific calculation formulas are shown in (1), (2), and (3) as follows: Among them, represents the difference in order between feature f on the k-th sample i and f j ; n d represents the total number of disease samples; the larger the absolute value of SCC(f i , f j ), the stronger the Spearman correlation between feature f i and f j ; MI ij represents the mutual information between feature f i and f j ; div i and div j represent the number of grid divisions in the directions of f i and f j respectively; B is a variable whose value is taken as max(n α , 4), where α is default set to 0.6; the larger the value of MIC(fi, f j ), the stronger the correlation between feature fi and f j ; similarly, the larger the absolute value of Cos(fi, f j ), the stronger the correlation between feature fi and f j . The vertex set of the molecular network established under each perspective is F. For each feature fi (1 ≤ i ≤ m), the top k features with the strongest correlation with it are selected to establish connections, and the weight of the edge is equal to the corresponding correlation value; there are no connections with the remaining features, and the weight is 0; let the three molecular networks constructed based on disease samples be the Spearman correlation coefficient network the maximum information coefficient network and the cosine similarity network where V m = F represents the node set, F is the feature set, and respectively represent the edge sets calculated based on the Spearman correlation coefficient, the maximum information coefficient, and the cosine similarity; (2) Divide modules For three molecular networks, the fast-greedy community detection algorithm is used to perform module partitioning respectively; modularity is an index to measure the strength of the network module structure; taking the network as an example, assume that the partitioned module set M S ={m S1 , m S2 , …, m Sk}, where m Si represents a module, which contains at least one feature (node), and each feature belongs to and only belongs to one module; the calculation formula of modularity Q S is as follows: Among them, is the number of edges in the network ; A is the adjacency matrix of . If there is an edge between fi and f in the network j , then A ij = 1; otherwise, A ij = 0; k i is the degree of the feature (node) fi; and represent the modules to which the features (nodes) fi and f j belong, respectively. When fi and f j are in the same module, otherwise Initially, each feature (node) is regarded as an independent module. Subsequently, the modules are gradually merged through a greedy optimization strategy: calculate the modularity increment ΔQ after the merger of each pair of modules, and select a pair of modules with the largest ΔQ and ΔQ≥0 for merger. Repeat this step until there are no modules with ΔQ≥0 available for merger. (3) Identify key modules Sample networks are established separately based on the features contained in each module. The cosine similarity is used to measure the similarity between samples, and the edges of the sample network are established in the KNN manner. Let the sample network established based on the features contained in one of the modules be G s =(V s , E), where V s =S, E = {(s i , s j ) | s i , s j ∈V s , i≠j, s i ∈N(s j ) or s j ∈N(s i )}, where N(s i ) and N(s j ) respectively represent the sets of the neighboring nodes of the samples (nodes) s i and s j ; if the feature subspace has strong discrimination ability, then in this subspace, the edges between samples of the same class are dense, and the edges between samples of different classes are sparse; therefore, the difference between the number of edges between samples of the same class and the number of edges between samples of different classes in the sample network established based on the module is used as the score of the module: Score=|E intra |-|E inter | (5) Then, for each perspective, the modules are sorted according to the module scores from high to low, and T modules with the highest scores and the number of features included not exceeding 20% of |F| are selected, where |F| represents the total number of features. Suppose that T1, T2, and T3 key modules are selected from the molecular networks based on three different perspectives, and the feature subsets under the three perspectives are thus determined: and represent the modules sorted by score, and Features(·) represents the set of features included in the module; (4) Establish a classifier based on the feature subspace Three feature subsets obtained from the perspectives of considering the monotonic relationship, the generalized correlation, and the direction consistency between features respectively are rich in different biological information. Take the union of the three feature subsets to obtain a fused feature subspace to ensure the information richness of the feature subspace. Apply a support vector machine (SVM) on the fused feature subspace to construct a classifier and classify the data to obtain the final result.