Processing method, processing system and storage medium for asthma omics data
Patent Information
- Application Number
- CN202310396153.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-04-12
AI Technical Summary
虽然该方法初始参数的随机初始化使得目标函数的解不唯一且对噪声敏感,但是可以通过对初始参数采用奇异值初始化并加入多种网络正则化约束来构建竞争性内源RNA网络的方式,得到解决并获得了良好的效果,但该算法仍然无法兼顾不同组学数据之间的非线性关系
[0007]为了克服现有技术存在的上述缺陷,本发明提供了一种哮喘病组学数据的处理方法、一种哮喘病组学数据的处理系统,以及一种计算机可读存储介质,能够基于数据驱动的临床信息规则提取方法,在融合样本先验信息的同时,整合多种具有非线性关联的组学数据、抵抗数据中的噪声,从而挖掘哮喘病相关的生物标志物,以为哮喘病的诊断和治疗靶点开发提供重要参考。
Smart Images

Figure CN118800324B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, and in particular to a method for processing asthma omics data, a system for processing asthma omics data, and a computer-readable storage medium. Background Technology
[0002] As a heterogeneous inflammatory disease, asthma is characterized by recurrent wheezing, chest tightness, and cough, which have serious negative impacts on human health. Therefore, it is particularly important to explore biomarkers for asthma and construct diagnostic models for it.
[0003] Existing diagnostic methods can explore biomarkers for asthma and construct diagnostic models by acquiring and integrating data from multiple omics datasets. For example, logistic regression models using asthma omics data have identified several single nucleotide polymorphisms associated with asthma-related transcription factors. However, logistic regression assumes a linear relationship between the original omics data, which is insufficient for omics datasets with non-linear relationships.
[0004] Furthermore, in existing technologies, nonnegative matrix factorization algorithms can effectively utilize prior information in omics data to improve the effectiveness of biomarker mining. Although the random initialization of the initial parameters in this method makes the solution of the objective function non-unique and sensitive to noise, this can be addressed and good results can be achieved by using singular value initialization of the initial parameters and adding various network regularization constraints to construct a competitive endogenous RNA network. However, this algorithm still cannot take into account the nonlinear relationships between different omics data.
[0005] To overcome the aforementioned shortcomings of existing technologies, there is an urgent need in this field for a processing technology for asthma omics data, based on a data-driven clinical information rule extraction method. This method integrates multiple omics data with nonlinear correlations while fusing prior information from samples and resisting noise in the data, thereby mining asthma-related biomarkers to provide a reference for the development of diagnostic and therapeutic targets for asthma. Summary of the Invention
[0006] The following provides a brief overview of one or more aspects to offer a basic understanding of them. This overview is not an exhaustive summary of all conceived aspects, nor is it intended to identify key or decisive elements of all aspects, nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed descriptions that follow.
[0007] To overcome the aforementioned deficiencies in existing technologies, this invention provides a method for processing asthma omics data, a system for processing asthma omics data, and a computer-readable storage medium. Based on a data-driven clinical information rule extraction method, this invention integrates multiple omics data with nonlinear correlations while fusing prior information from samples and resisting noise in the data, thereby uncovering asthma-related biomarkers and providing important references for the development of diagnostic and therapeutic targets for asthma.
[0008] Specifically, a method for processing asthma omics data according to a first aspect of the present invention includes the following steps: acquiring transcriptome data and DNA methylation data of asthma patients and control samples; performing differential expression analysis on the transcriptome data and the DNA methylation data to determine differentially expressed genes and differentially methylated sites in the asthma patients and control samples; reconstructing a gene expression matrix and a DNA methylation expression matrix by changing the distribution of the differentially expressed genes and the differentially methylated sites according to the group of the asthma patients and the control samples using a deep subspace reconstruction algorithm; performing joint deep semi-supervised nonnegative matrix decomposition on the gene expression matrix and the DNA methylation expression matrix to construct a cooperating module; constructing a diagnostic model based on the cooperating module using deep orthogonal canonical correlation analysis and machine learning algorithms; and determining asthma-related biomarkers using the diagnostic model.
[0009] Furthermore, in some embodiments of the present invention, the transcriptome data and DNA methylation data of the asthma patient are tabular data with no missing values. Each row of the tabular data corresponds to one asthma patient, and each column of the tabular data corresponds to one gene or one methylation site of the asthma patient.
[0010] Furthermore, in some embodiments of the present invention, the step of performing differential expression analysis on the transcriptome data and the DNA methylation data to determine the differentially expressed genes and differentially methylated sites of the asthma patients and the control samples includes: dividing the transcriptome data and DNA methylation data of each of the asthma patients and each of the control samples according to a preset ratio to construct a training set and a test set; and performing the differential expression analysis on the transcriptome data and DNA methylation data of each of the asthma patients and each of the control samples in the training set via the limma algorithm to determine the differentially expressed genes and differentially methylated sites of the asthma patients and the control samples.
[0011] Furthermore, in some embodiments of the present invention, the step of reconstructing a gene expression matrix and a DNA methylation expression matrix by altering the distribution of the differentially expressed genes and the differentially methylated sites according to the group of the asthma patient and the control sample via a deep subspace reconstruction algorithm includes: determining sample tags for each of the differentially expressed genes and each of the differentially methylated sites according to the group of the asthma patient and the control sample; and altering the distribution of the differentially expressed genes and the differentially methylated sites according to the sample tags to reconstruct the gene expression matrix and the DNA methylation expression matrix.
[0012] Furthermore, in some embodiments of the present invention, the step of changing the distribution of the differentially expressed genes and the differentially methylated sites according to the sample tags to reconstruct the gene expression matrix and the DNA methylation expression matrix includes: determining the initial expression matrix of each differentially expressed gene and each differentially methylated site; and inputting the initial expression matrix and the corresponding sample tags into a deep subspace reconstruction algorithm to make sample data with the same tags close together in a subspace and to make sample data with different tags far apart in the subspace, so as to reconstruct the gene expression matrix and the DNA methylation expression matrix.
[0013] Furthermore, in some embodiments of the present invention, the step of performing joint deep semi-supervised non-negative matrix factorization on the gene expression matrix and the DNA methylation expression matrix to construct a cooperating module includes: inputting the reconstructed gene expression matrix and the DNA methylation expression matrix into a joint deep semi-supervised non-negative matrix factorization algorithm to obtain multiple candidate cooperating modules, wherein each candidate cooperating module includes a different number of genes and methylation sites; and screening the multiple candidate cooperating modules based on at least one preset screening rule to obtain at least one key cooperating module.
[0014] Furthermore, in some embodiments of the present invention, the step of inputting the reconstructed gene expression matrix and the DNA methylation expression matrix into a joint deep semi-supervised nonnegative matrix factorization algorithm to obtain multiple candidate cooperating modules includes: decomposing the reconstructed gene expression matrix and the DNA methylation expression matrix into a common sample latent matrix and multiple feature latent matrices; and performing nonlinear feature association analysis by performing layer-by-layer dimensionality reduction on the multiple feature latent matrices and nonlinear transformation during the layer-by-layer activation function dimensionality reduction process to obtain the multiple candidate cooperating modules.
[0015] Furthermore, in some embodiments of the present invention, the step of constructing a diagnostic model based on the coordinating module via deep orthogonal canonical correlation analysis and machine learning algorithms includes: inputting the expression levels of at least one gene and at least one methylation site in the key coordinating module into a deep orthogonal canonical correlation analysis model to obtain the weights of each gene and each methylation site; taking the absolute value of each weight and sorting them according to their size to obtain a preset number of target genes and target methylation sites; and using a logistic regression algorithm to construct a diagnostic model for the target genes and the target methylation sites.
[0016] Furthermore, a system for processing asthma omics data according to a second aspect of the present invention includes a memory and a processor. The memory stores computer instructions. The processor is connected to the memory and configured to execute the computer instructions stored in the memory to implement the asthma omics data processing method as described in any one of the first aspects of the present invention.
[0017] Furthermore, according to a third aspect of the present invention, a computer-readable storage medium is provided thereon storing computer instructions. When the computer instructions are executed by a processor, the method for processing asthma genomics data as described in any one of the first aspects of the present invention is implemented. Attached Figure Description
[0018] The above-described features and advantages of the present invention will be better understood after reading the following detailed description of embodiments of the present disclosure in conjunction with the accompanying drawings. In the drawings, components are not necessarily drawn to scale, and components having similar related properties or features may have the same or similar reference numerals.
[0019] Figure 1 A flowchart illustrating a method for processing asthma genomics data according to some embodiments of the present invention is shown.
[0020] Figures 2A to 2D A schematic diagram of dimensionality-reduced visualization scatter plots of gene expression data and methylation data before and after deep subspace reconstruction provided according to some embodiments of the present invention is shown.
[0021] Figures 3A to 3D A schematic diagram of the ROC curve of a diagnostic model provided according to some embodiments of the present invention is shown. Detailed Implementation
[0022] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention is presented in conjunction with preferred embodiments, this does not mean that the features of the invention are limited to these embodiments. On the contrary, the purpose of describing the invention in conjunction with embodiments is to cover other options or modifications that may be derived based on the claims of the present invention. To provide a thorough understanding of the invention, many specific details will be included in the following description. The invention may also be implemented without using these details. Furthermore, to avoid confusion or obscuring the focus of the invention, some specific details will be omitted in the description.
[0023] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0024] Furthermore, the terms "upper," "lower," "left," "right," "top," "bottom," "horizontal," and "vertical" used in the following description should be understood as the orientations shown in the relevant paragraphs and accompanying drawings. These relative terms are for illustrative purposes only and do not imply that the described apparatus must be manufactured or operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0025] It is understood that although terms such as "first," "second," and "third" may be used herein to describe various components, regions, layers, and / or parts, these components, regions, layers, and / or parts should not be limited by these terms, and these terms are only used to distinguish different components, regions, layers, and / or parts. Therefore, the first components, regions, layers, and / or parts discussed below may be referred to as second components, regions, layers, and / or parts without departing from some embodiments of the present invention.
[0026] As described above, to overcome the aforementioned deficiencies in the existing technology, this invention provides a method for processing asthma omics data, a system for processing asthma omics data, and a computer-readable storage medium. Based on a data-driven clinical information rule extraction method, this invention integrates multiple omics data with non-linear correlations while fusing prior information from samples and resisting noise in the data, thereby mining asthma-related biomarkers and providing important references for the development of diagnostic and therapeutic targets for asthma.
[0027] In some non-limiting embodiments, the asthma omics data processing method provided in the first aspect of the present invention can be implemented based on the asthma omics data processing system provided in the second aspect of the present invention. Specifically, the asthma omics data processing system is configured with a memory and a processor. The memory includes, but is not limited to, the computer-readable storage medium provided in the third aspect of the present invention, on which computer instructions are stored. The processor is connected to the memory and configured to execute the computer instructions stored in the memory to implement the asthma omics data processing method provided in the first aspect of the present invention.
[0028] The following will illustrate the embodiments of the present invention with specific examples of methods for processing asthma omics data. Those skilled in the art will understand that these embodiments of asthma omics data processing methods are merely non-limiting implementations provided by the present invention, intended to clearly demonstrate the main concepts of the invention and provide specific solutions convenient for public implementation, rather than limiting all functions or operating methods of the asthma omics data processing system. Similarly, the asthma omics data processing system is also merely a non-limiting implementation of the present invention and does not limit the executing entity or execution order of the steps in these asthma omics data processing methods. Any solutions achieved by adding, subtracting, and / or replacing steps in the prior art based on the principles of the present invention should be included within the scope of protection of the present invention.
[0029] Please refer to the following first. Figure 1 , Figure 1 A flowchart illustrating a method for processing asthma genomics data according to some embodiments of the present invention is shown.
[0030] like Figure 1 As shown, the asthma omics data processing method includes the following steps S11: obtaining transcriptome data and DNA methylation data of asthma patients and their control samples.
[0031] Specifically, the transcriptomic and DNA methylation data for asthma patients and their control samples were presented as tabular data with no missing values. Here, each row of the tabular data represents a patient sample, and each column represents a characteristic of that patient.
[0032] Furthermore, technicians can download the gene expression profiling dataset GSE40732 and the DNA methylation expression profiling dataset GSE40576, which collect samples from the same asthma population, from the GEO database. The samples in both datasets are DNA and RNA from peripheral blood mononuclear cells of children aged 6–12 years in the city center, used to compare gene expression and methylation patterns in persistent atopic asthma and healthy controls. A total of 194 samples were included, including 97 healthy samples and 97 corresponding disease samples. The data platform for GSE40732 was GPL16025 (NimbleGen human expression array). The data platform for GSE40576 was GPL13534 (Illumina Human Methylation 450 BeadChip). All probe names in the datasets are labeled using the chip's GPL platform file.
[0033] In addition, the GSE40888 and GSE109446 datasets contain gene expression data and DNA methylation expression data for the external validation set, respectively. GSE40888 includes gene expression data from 14 allergic asthma samples and 14 control samples, and its data platform is GPL6244 (Affymetrix Human Gene 1.0ST Array). GSE109446 includes DNA methylation data from 29 asthma samples and 29 control samples, and its data platform is GPL13534 (Illumina Human Methylation 450 BeadChip).
[0034] Please continue to refer to this. Figure 1 After obtaining transcriptome data and DNA methylation data from asthma patients and their control samples, this asthma omics data processing system can first perform step S12: perform differential expression analysis on transcriptome data and DNA methylation data to identify differentially expressed genes and differentially methylated sites in asthma patients and control samples, and then perform step S13: obtain sample tags from asthma patients and their control samples.
[0035] Specifically, the processing system can divide the transcriptome data and DNA methylation data of asthma patients and their control samples into training and test sets in an 8:2 ratio. After differential expression analysis using the limma algorithm and setting the screening threshold p to less than 0.05, genes and methylation sites with significant differences in expression between the asthma group and the control group can be obtained.
[0036] Furthermore, the processing system can extract their expression levels and organize them into tabular data. Let the number of samples be n, and the processed gene expression matrix X1∈R n×p The processed DNA methylation matrix X2∈R n×qHere, p represents the number of differentially expressed genes, and q represents the number of differentially expressed methylation sites.
[0037] The entire dataset contains 97 asthma samples and 97 healthy samples. The processing system randomly selects 80% of the samples as the training set and the remaining 20% as the test set. 154 training samples (including 77 diseased samples and 77 healthy samples) are used for collaborative modules and biomarker discovery, while the remaining 40 test samples (including 20 diseased samples and 20 healthy samples) are used for case studies. Next, the processing system uses the R package limma algorithm to screen for differentially expressed genes and methylation sites between asthma and healthy samples, ultimately retaining 449 differentially expressed genes and 856 differentially methylation sites with p < 0.05.
[0038] Furthermore, the processing system can add labels to the above samples, labeling as 1 for asthma patient samples and 0 for control group samples. Thus, the diagnostic labels can be represented as l∈R n×1 express.
[0039] Please continue to refer to this. Figure 1 After identifying and tagging differentially expressed genes and differentially methylated sites in asthma patients and control samples, the asthma omics data processing method can perform step S14: via a deep subspace reconstruction algorithm, the distribution of the differentially expressed genes and differentially methylated sites is changed according to the group of the asthma patients and the control samples to reconstruct and obtain the gene expression matrix and DNA methylation expression matrix.
[0040] Please refer to the details. Figures 2A to 2D , Figures 2A to 2D A schematic diagram of dimensionality-reduced visualization scatter plots of gene expression data and methylation data before and after deep subspace reconstruction provided according to some embodiments of the present invention is shown.
[0041] like Figures 2A to 2D As shown, deep subspace reconstruction algorithms can alter the distribution of the original data. Existing techniques demonstrate that prior information can improve the correlation analysis performance of multi-omics fusion models and algorithms. However, most multi-omics fusion schemes directly apply various penalty terms to the original data, which contains technical and biological noise, leading to estimation bias during fusion. Therefore, considering the underlying multi-subspace structure of the data, deep subspace reconstruction is a first-of-its-kind approach that integrates subject and sample diagnostic information into the original data.
[0042] Specifically, deep subspace reconstruction utilizes the self-representational properties of data. It ensures that samples with the same diagnostic label will be collected in the same subspace, while samples with different diagnostic labels will not be collected in the same subspace.
[0043] Furthermore, deep subspace reconstruction first performs a nonlinear transformation on the original data using a multi-layer feedforward neural network, and then reconstructs the data embedded in the subspace of the network's output layer. Specifically, we input X1, X2, and l into the algorithm, perform a nonlinear transformation on the original data using a multi-layer feedforward neural network, and then reconstruct the data embedded in the subspace of the network's output layer.
[0044] This invention uses X1 = [x1, x2, ... x i ,…,x N The deep subspace reconstruction algorithm is illustrated using [example 1]. First, X1 groups data according to labels, such as [example 2]. in, This represents the first sample of class i.
[0045] set up And set θ to {W (m) ,b (m) Let {C, m = 1: M, i = 1: N} be the hyperparameters in the feedforward neural network, where m represents the number of layers in the current network. Then, the output of the i-th sample with the first class label in the m-th layer is defined as follows:
[0046]
[0047] in, and d represents the weight matrix and bias matrix of the m-th layer, respectively. m This represents the dimension of the m-th layer of the neural network. The output of the last layer of the network is... in, Let represent the i-th sample after nonlinear mapping through a multi-layer neural network. The expression for the coefficient matrix C obtained by reconstructing the i-th sample with the d-th class label is given below.
[0048]
[0049] Here, C represents the block structure, which reflects the similar structure of the reconstructed data by reconstructing the original data. Ultimately, the reconstructed data can be represented as follows.
[0050] f*X i ) = C i X i (i=1,2)3)
[0051] Where C1 and C2 represent the self-expression coefficient matrices of X1 and X2, respectively. f(X1) and f(X2) are the expression matrices of the reconstructed gene and DNA methylation, respectively.
[0052] Please continue to refer to this. Figure 1 After reconstructing and obtaining the gene expression matrix and DNA methylation expression matrix, the asthma omics data processing method can perform step S15: perform joint deep semi-supervised non-negative matrix decomposition processing on the gene expression matrix and the DNA methylation expression matrix to construct a collaborative module.
[0053] Specifically, the processing system can input gene expression and methylation data reconstructed from deep subspaces into a joint deep semi-supervised non-negative matrix factorization algorithm to construct a collaborative module. This algorithm decomposes the two types of omics data reconstructed from deep subspaces into a common sample latent matrix and multiple feature latent matrices. Nonlinear feature association analysis is achieved through layer-by-layer dimensionality reduction of the feature latent matrices and nonlinear transformations during this dimensionality reduction process using activation functions. The objective function is as follows:
[0054]
[0055] in, It is the sample latent matrix. It is the feature latent matrix of the first layer of the neural network. It is the feature latent matrix of the (n+1)th layer. It is the connection latent matrix. In joint deep semi-supervised nonnegative matrix factorization, it is necessary to satisfy k0 < k i <k i+1 <k n <min{n,p i} The nonlinear decomposition is achieved by the sigmoid activation function, the expression of which is as follows:
[0056]
[0057] Ultimately, the processing system can obtain U and H. 10 and H 20 The cleaned data matrices (such as the gene expression matrix f(X1) and the DNA methylation matrix f(X2)) may share a sample latent matrix U, which can be viewed as a common feature basis matrix of the samples, where each feature represents a cooperative module among a group of samples. 10 and H 20 The decomposed feature coefficient matrix represents the potential relationship between the feature base (such as hidden sample representation) and the original feature (such as gene or methylation site).
[0058] To determine key features from the basis matrix U, the z-score method is used for each feature coefficient vector in the coefficient matrix. It is defined as:
[0059]
[0060] Among them, h ij This refers to the characteristic coefficient, μ j This refers to the mean of the characteristic coefficients of characteristic j, σ. j This refers to the standard deviation of these feature coefficients. For each feature, if its z-score is greater than a threshold T, it is considered a key feature of a feature base. The complete set of key features for each feature base contains a collaborative module.
[0061] In summary, by employing a joint deep semi-supervised nonnegative matrix factorization algorithm, the processing system obtained 151 cooperating modules. Then, 78 modules without any key features were removed, leaving 9 modules, exceeding 2% of the total number of genes and methylation sites. Since modules with too few elements are detrimental to subsequent analysis, the system calculated the average number of elements in each module and retained 4 modules with a number of elements greater than the average. Here, the escape rate is defined as the percentage of elements in a module that have no overlap with any other module. Specifically, modules 12, 25, 31, and 75 had escape rates of 23.8%, 86.67%, 16.67%, and 100% at the gene expression level, and 79.31%, 14.29%, 81.48%, and 4.17% at the methylation level, respectively. The joint deep semi-supervised nonnegative matrix factorization effectively reconstructs the original data at the module level.
[0062] Please continue to refer to this. Figure 1 After the collaborative module is constructed, the processing method for this asthma omics data can further execute S16: construct a diagnostic model based on the collaborative module through deep orthogonal canonical correlation analysis and machine learning algorithms, thereby identifying asthma-related biomarkers.
[0063] Specifically, for each collaborative module, this invention develops a canonical correlation analysis algorithm based on orthogonal constraints by applying orthogonal constraints to the weight vectors of canonical correlation analysis. This reduces the impact of collinearity of key features on feature ranking and selection in constructing the diagnostic model. The input data for this method is the data reconstructed from a deep subspace corresponding to a collaborative module, and its objective function is as follows:
[0064]
[0065]
[0066] in, p is the gene expression matrix reconstructed from the deep subspace in the cooperating module k. [k] This represents the number of genes in module k. q is the methylation expression matrix reconstructed from the deep subspace in the co-operating module k.[k] The number of methylation sites in module k. and Let λu and λv represent the CCA weight vectors for genes and methylation, respectively. I is the identity matrix, and λ1 and λ2 are two hyperparameters used to control the sparsity intensity of u and v, respectively. The objective function above can be rewritten as:
[0067]
[0068]
[0069] This objective function can be optimized by iteratively updating u and v alternately using the Lagrange operator:
[0070] First, this invention can fix u, and then perform a rescue operation on v, setting the partial derivative of 10) with respect to u to 0, to obtain the following equation:
[0071]
[0072] Then, the iterative formula for u can be obtained by the present invention:
[0073]
[0074] Similarly, the iterative formula for v is as follows:
[0075]
[0076] Furthermore, through deep orthogonal canonical correlation analysis, we can obtain the feature weights in the key collaborative modules. The processing system can sort the features according to their weights based on their absolute values, and construct a diagnostic model using the top 10 genes and top 10 methylation sites in this module based on the logistic regression algorithm. The logistic regression algorithm can be performed using IBM SPSS Statistics 26.
[0077] Please refer to the reference. Figures 3A-3D And Tables 1-2. Figures 3A-3B The diagrams show the ROC curves for the internal and external test sets of gene expression data, respectively. Figures 3C-3D The diagrams show the ROC curves for the internal and external test sets of the DNA methylation expression data. Tables 1 and 2 present the top 10 genes and methylation sites, along with their weights.
[0078] Table 1. Top 10 genes by weight.
[0079]
[0080]
[0081] Table 2. Top 10 methylation sites by weight.
[0082] cg00456348 0.1496 cg00053393 0.1175 cg00313876 0.1058 cg00322319 0.0932 cg00483304 0.0774 cg00266865 0.0757 cg00240732 0.0640 cg00146676 0.0638 cg00610021 0.0533 cg00594129 0.0504
[0083] like Figures 3A-3D As shown, for genes and methylation sites in key collaborative modules, the DOCCA algorithm proposed in this invention can use CCA to assign feature weights, sort the absolute values of feature weights from DOCCA from high to low, and use the top 10 gene sites and methylation sites in the test set samples of the module to construct a diagnostic model.
[0084] Finally, the processing system can use these features to construct a diagnostic model for asthma based on the logistic regression algorithm, and then use this diagnostic model to identify asthma-related biomarkers.
[0085] In summary, this invention's deep association-based asthma omics data integration and analysis method first acquires transcriptomic and DNA methylation data from the same batch of samples. Then, it extracts differentially expressed genes and methylation sites using the limma algorithm, and subsequently reconstructs the two sets of data using a deep subspace reconstruction algorithm, thereby altering the data distribution. The reconstructed data is input into a joint deep semi-supervised nonnegative matrix factorization algorithm to obtain a synergistic module composed of gene loci and methylation sites. The weights of the two data features are obtained using a deep orthogonal canonical correlation analysis algorithm. The weights are then sorted from largest to smallest by taking their absolute values, and the top 10 gene loci and methylation sites are selected. Logistic regression is then used to construct and validate diagnostic models for asthma.
[0086] Although the methods described above are illustrated and depicted as a series of actions for the sake of simplicity, it should be understood and appreciated that these methods are not limited by the order of the actions, as some actions may occur in a different order and / or concurrently with other actions from the illustrations and descriptions herein or not illustrated and described herein but which may be understood by those skilled in the art, according to one or more embodiments.
[0087] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in a generalized manner in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention.
[0088] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing asthma genomics data, characterized in that, Includes the following steps: Obtain transcriptomic and DNA methylation data from asthma patients and control samples; Differential expression analysis was performed on the transcriptome data and the DNA methylation data to identify differentially expressed genes and differentially methylated sites in the asthma patients and the control samples. Using a deep subspace reconstruction algorithm, the distribution of differentially expressed genes and differentially methylated sites is altered according to the groups of the asthma patients and the control samples, in order to reconstruct and obtain the gene expression matrix and DNA methylation expression matrix; The gene expression matrix and the DNA methylation expression matrix are subjected to joint deep semi-supervised non-negative matrix decomposition to construct a collaborative module; A diagnostic model based on the collaborative module is constructed using deep orthogonal canonical correlation analysis and machine learning algorithms. The deep orthogonal canonical correlation analysis includes orthogonalizing the weight vectors of the canonical correlation analysis to reduce the impact of collinearity of key features on the feature ranking and selection for constructing the diagnostic model. The diagnostic model was used to identify asthma-related biomarkers.
2. The processing method as described in claim 1, characterized in that, The transcriptome data and DNA methylation data of the asthma patients are tabular data with no missing values, wherein each row of the tabular data corresponds to one asthma patient, and each column of the tabular data corresponds to one gene or one methylation site of the asthma patient.
3. The processing method as described in claim 1, characterized in that, The step of performing differential expression analysis on the transcriptome data and the DNA methylation data to determine the differentially expressed genes and differentially methylated sites in the asthma patient and the control samples includes: Transcriptome and DNA methylation data of each asthma patient and control sample were divided according to a predetermined ratio to construct training and testing sets; and The differential expression analysis was performed on the transcriptomic data and DNA methylation data of each asthma patient and each control sample in the training set using the limma algorithm to identify the differentially expressed genes and differentially methylated sites of the asthma patients and the control samples.
4. The processing method as described in claim 1, characterized in that, The step of reconstructing the gene expression matrix and DNA methylation expression matrix by altering the distribution of differentially expressed genes and differentially methylated sites according to the group of the asthma patients and the control samples using a deep subspace reconstruction algorithm includes: Based on the group classification of the asthma patients and the control samples, sample tags were determined for each differentially expressed gene and each differentially methylated site; and Based on the sample tags, the distribution of the differentially expressed genes and the differentially methylated sites is changed to reconstruct the gene expression matrix and the DNA methylation expression matrix.
5. The processing method as described in claim 4, characterized in that, The step of modifying the distribution of the differentially expressed genes and the differentially methylated sites according to the sample tags to reconstruct the gene expression matrix and the DNA methylation expression matrix includes: Determine the initial expression matrix for each differentially expressed gene and each differentially methylation site; and The initial expression matrix and the corresponding sample labels are paired and input into a deep subspace reconstruction algorithm, so that sample data with the same label are close together in a subspace, and sample data with different labels are far apart in the subspace, so as to reconstruct the gene expression matrix and the DNA methylation expression matrix.
6. The processing method as described in claim 1, characterized in that, The step of performing joint deep semi-supervised non-negative matrix decomposition on the gene expression matrix and the DNA methylation expression matrix to construct a cooperating module includes: The reconstructed gene expression matrix and DNA methylation expression matrix are input into a joint deep semi-supervised nonnegative matrix factorization algorithm to obtain multiple candidate cooperating modules, wherein each candidate cooperating module includes a varying number of genes and methylation sites; and The multiple candidate collaborative modules are filtered based on at least one preset filtering rule to obtain at least one key collaborative module.
7. The processing method as described in claim 6, characterized in that, The step of inputting the reconstructed gene expression matrix and the DNA methylation expression matrix into a joint deep semi-supervised nonnegative matrix factorization algorithm to obtain multiple candidate cooperating modules includes: The reconstructed gene expression matrix and DNA methylation expression matrix are decomposed into a common sample latent matrix and multiple feature latent matrices; and By performing layer-by-layer dimensionality reduction on the multiple latent feature matrices and nonlinear transformations during the layer-by-layer activation function dimensionality reduction process, nonlinear feature correlation analysis is conducted to obtain the multiple candidate collaborative modules.
8. The processing method as described in claim 6, characterized in that, The step of constructing a diagnostic model based on the collaborative module via deep orthogonal canonical correlation analysis and machine learning algorithms includes: The expression levels of at least one gene and at least one methylation site in the key collaborative module are input into a deep orthogonal canonical correlation analysis model to obtain the weights of each gene and each methylation site. The absolute values of each weight are taken and sorted by their magnitude to obtain a predetermined number of target genes and target methylation sites; and A diagnostic model for the target gene and the target methylation site is constructed using the logistic regression algorithm.
9. A system for processing asthma genomics data, characterized in that, include Memory, on which computer instructions are stored; and A processor, connected to the memory, and configured to execute computer instructions stored in the memory to implement the method for processing asthma omics data as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instructions are executed by the processor, the method for processing asthma genomics data as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Transcriptome and DNA methylation data correlation analysis method and system
CN112201302A
IncRNA-disease association prediction method based on high-order proximity and matrix completion algorithm
CN113160880A