A method of processing respiratory system disease data and a system thereof
By combining single-cell pathway scoring methods and genetic association data, and utilizing the genetically related pathway activity score gPAS, the problem of identifying respiratory disease-related cell subpopulations in existing technologies has been solved, achieving deeper disease risk identification and improved statistical efficiency.
Patent Information
- Application Number
- CN202211277916.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-19
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-10-19
AI Technical Summary
Existing technologies struggle to effectively identify cell subpopulations associated with respiratory diseases using single-cell RNA sequencing (scRNA-seq) data and genetic association data. Furthermore, existing methods suffer from complex parameter adjustments and neglect of internal heterogeneity when identifying cell types, resulting in limited statistical power.
A single-cell pathway-based scoring method was used, combined with scRNA-seq data and genetic association data. Through machine learning and multi-gene regression models, genes, cells and biological pathways related to respiratory diseases were inferred. Statistical analysis was performed using the genetically related pathway activity score gPAS to identify trait-related genes and cell subpopulations.
It significantly improves statistical power and biological interpretability, enabling in-depth analysis of life patterns from multiple dimensions, identification of disease-related early developmental events and key cell types, overcoming the limitations of existing technologies, and providing more accurate disease risk identification capabilities.
Smart Images

Figure CN116486911B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of gene sequencing, and more particularly to a respiratory disease data processing method and system. BACKGROUND
[0002] Respiratory disease is a common and frequently-occurring disease, and the main lesions are in the trachea, bronchus, lung and chest. The light lesions are mostly cough, chest pain, and affected breathing, and the severe lesions are dyspnea, hypoxia, and even respiratory failure and death. In the city, the mortality rate ranks third, while in the countryside, it ranks first. More attention should be paid to the fact that due to air pollution, smoking, population aging and other factors, the incidence and mortality of chronic obstructive pulmonary disease (including chronic bronchitis, emphysema, and pulmonary heart disease), bronchial asthma, lung cancer, pulmonary interstitial fibrosis, and pulmonary infection have increased.
[0003] Coronaviruses are a large family of viruses that can cause illnesses ranging from the common cold to more severe diseases such as Middle East Respiratory Syndrome (MERS) and Severe Acute Respiratory Syndrome (SARS). Common signs of infection with coronaviruses in humans include respiratory symptoms, fever, cough, shortness of breath, and difficulty breathing. In more severe cases, infection can lead to pneumonia, severe acute respiratory syndrome, kidney failure, and even death. Many of the symptoms caused by coronaviruses can be managed, so treatment should be based on the patient's clinical condition. In addition, supportive care for infected individuals can be very effective, and self-protection measures include maintaining basic hand and respiratory hygiene, adhering to safe eating habits, etc. Understanding the impact of host genetic components on the immune response to severe infection can help develop effective vaccines and treatments to control related respiratory disease pandemics. With the rapid development of sequencing technology, single-cell sequencing technology provides a more comprehensive opportunity to reveal the relevant mechanisms of related respiratory diseases.
[0004] Identifying key cell subpopulations associated with complex diseases or traits using single-cell RNA sequencing (scRNA-seq) technology is crucial for understanding the mechanisms of complex diseases. However, scRNA-seq data is not suitable for large-scale sequencing due to its high cost and low throughput, and most current single-cell-based studies have fewer than 20 samples, resulting in limited statistical power and an inability to accurately reveal risk subsets associated with diseases or traits in cell subpopulations. In addition, scRNA-seq data has high sparsity, technical noise, and variance instability at the gene level. Genetic association data, such as genome-wide association studies (GWAS), are widely used to study different complex diseases or traits. Associating scRNA-seq data with phenotypic-related genetic information from large-scale GWAS samples is considered a practical and effective method that can reveal the genetic molecular mechanisms of complex diseases or traits at single-cell resolution.
[0005] Methods of combining GWAS with scRNA-seq data to identify cell types associated with complex diseases include methods such as LDSC-SEG, MAGMA, RolyPoly, but the above methods require a large number of parameter adjustments in order to annotate cell types with known marker genes, and largely ignore the internal heterogeneity of each cell type. In addition, the prior art can identify genes with high expression levels, but the potential defect is that overemphasizing high expression genes can underestimate the functional role of genes with relatively low expression levels but important for revealing cell fate. SUMMARY
[0006] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application provides a method for processing respiratory system disease data and a system thereof; the method of the present application infers genes, cells, etc. associated with respiratory system diseases by combining scRNA-seq data and genetic association data based on the scoring method of single-cell pathways, and determines the potential relationship between genes, cells, cell subpopulations, biological pathways, etc. and respiratory system diseases from a deep level to mine the life laws hidden behind single-cell sequencing data.
[0007] The present application discloses a method for processing respiratory system disease data, comprising:
[0008] obtaining single-cell sequencing sequence data to be analyzed;
[0009] processing the single-cell sequencing sequence data to be analyzed by a machine learning method to obtain a PAS scoring matrix of cell pathways and a PAS of cell pathways;
[0010] obtaining genetic association data of respiratory diseases, processing the genetic association data to obtain pathway data with SNP annotations;
[0011] statistically analyzing and processing the PAS of the cell pathways and the pathway data with SNP annotations to obtain an estimation coefficient;
[0012] multiplying the estimation coefficient by the PAS and summing to obtain a genetic correlation pathway activity score gPAS of cells;
[0013] outputting the genetic correlation pathway activity score gPAS.
[0014] The step of statistically analyzing and processing the PAS of the cell pathways and the pathway data with SNP annotations to obtain an estimation coefficient comprises:
[0015] obtaining genetic effect values of all SNPs in single-pathway data based on the pathway data with SNP annotations;
[0016] based on the PAS and the genetic effect value, performing parameter estimation on a distribution of the genetic effect value to obtain an estimated coefficient;
[0017] Optionally, the genetic effect value is obtained according to a formula: wherein β represents a theoretical effect size vector of m SNPs, ε represents a random environmental error, R represents an LD matrix, X represents a standard genotype of the SNPs in the genetic association data sample, and Y represents a standard phenotype of the genetic association data sample. T
[0018] Optionally, the estimated coefficient is obtained according to a method comprising:
[0019]
[0020] wherein τ i,j represents an estimated coefficient of pathway i in cell j, τ0 represents an intercept term, σ 2 represents a variance of SNP effect size in the pathway, and represents a weighted PAS.
[0021] The genetic correlation pathway activity score gPAS (gPj) is obtained according to a formula:
[0022]
[0023] wherein the is an optimized estimated coefficient.
[0024] The method further comprises: performing correlation analysis and sorting on the genetic correlation pathway activity score gPAS and gene expression of each cell to screen N trait-related genes.
[0025] Optionally, the method of performing correlation analysis and sorting on the genetic correlation pathway activity score gPAS and gene expression of each cell comprises: determining the correlation between expression of a single gene and the gPAS through a Pearson correlation coefficient (PCC), and sorting the genes according to the correlation to obtain the N trait-related genes.
[0026] Optionally, the N trait-related genes are the top 1000 or bottom 1000 trait-related genes sorted in descending or ascending order of correlation degree.
[0027] Optionally, the N trait-related genes comprise one or more of the following: CALM3, PIK3R1, IL32, CD3E, B2M, PRS29, and GZMB.
[0028] The method further comprises: calculating a trait-related score TRS of each cell according to the N trait-related genes; and clustering according to the trait-related score TRS and a horizontal P value of a single cell to obtain trait-related cells related to different levels of severity of respiratory diseases.
[0029] Optionally, a cell scoring method is used to calculate a trait-related score TRS of the N trait genes.
[0030] Optionally, the different levels of severity of the respiratory diseases include mild, moderate and severe.
[0031] The method further comprises: obtaining a trait-related cell type or subpopulation based on a block bootstrap method.
[0032] The method further comprises: ranking the genetic correlation pathway activity score gPAS, and obtaining a trait-related pathway according to the ranking result, a P value of a pathway at a cell type level and the statistical importance value.
[0033] Detecting new Use of a product of a CD8+ T cell subpopulation in the preparation of a product for diagnosing respiratory diseases.
[0034] A respiratory system disease data processing device, the device comprising: a memory and a processor;
[0035] The memory is configured to store program instructions, and the processor is configured to invoke the program instructions, and when the program instructions are executed, the program instructions are configured to perform the respiratory system disease data processing method.
[0036] A respiratory system disease data processing system, comprising:
[0037] An acquisition unit configured to acquire single-cell sequencing sequence data to be analyzed;
[0038] A first processing unit configured to process the single-cell sequencing sequence data to be analyzed by using a machine learning method to obtain a PAS score matrix of a cell pathway and a PAS of a cell pathway;
[0039] A second processing unit configured to acquire genetic correlation data of respiratory diseases, and process the genetic correlation data to obtain pathway data with SNP annotations;
[0040] A third processing unit configured to statistically analyze and process the PAS of the cell pathway and the pathway data with SNP annotations to obtain an estimation coefficient;
[0041] The fourth processing unit is used to multiply the estimated coefficients by PAS and then sum them to obtain the cell's genetic pathway activity score gPAS;
[0042] The output unit is used to output the genetically related pathway activity score gPAS.
[0043] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for processing respiratory disease data.
[0044] This application has the following beneficial effects:
[0045] 1. This application innovatively discloses a method for processing respiratory disease data by combining single-cell sequencing data and genetic association data. This method can infer, from a deeper and more multi-dimensional perspective, genes, cells, cell subpopulations, and related biological pathways associated with respiratory diseases. Understanding the impact of these host genetic components on the immune response to severe infections helps in the development of effective vaccines and treatments to control disease pandemics, contributing to research on respiratory diseases. Based on a single-cell pathway scoring method, this method has a powerful ability to identify disease-risk cell types. It integrates the functional roles of different genes involved in the same biological pathways to obtain stable cell states, significantly increasing statistical power, biological interpretability, and result reproducibility. It overcomes the limitations of known cell type annotation and may discover key genes or pathways in new genetically related subpopulations and cell types. It has wide applications and strong practicality. For example, this method can identify new gene-driven genes that can be prioritized. CD8+ T cell subsets may play an important role in mediating the immune response in patients with severe respiratory diseases.
[0046] 2. This application innovatively discloses a scoring method based on single-cell pathways, employing a multi-gene regression model. It utilizes scRNA-seq data of pathway activity transformation and genetic association study data to reveal genes and cell subpopulations associated with traits. This effectively overcomes the current limitations in identifying genes and cell subpopulations associated with multi-gene risks in complex diseases, which are largely hindered by the small sample size and high sparsity of scRNA-seq data, resulting in limited statistical power and an inability to accurately reveal risk subsets related to diseases or traits within cell subpopulations. This method delves deeper into the life patterns hidden behind single-cell sequencing data, conducting in-depth analysis from multiple dimensions, including the relationship between population genetic mutations and diseases and single-cell sequencing gene abundance information, significantly improving the accuracy and depth of data analysis.
[0047] 3、The present application is based on large-scale simulation and real data, and combines scRNA-seq data and genetic correlation data by using the above scoring method, which can effectively overcome the problem that a large number of parameters need to be adjusted in the prior art for the convenience of annotating cell types by using known marker genes, and the internal heterogeneity in each cell type is largely ignored; there is no problem of underestimating the functional role of genes with relatively low expression levels but important for revealing cell fate due to excessive attention to highly expressed genes, which helps to identify early developmental events or progenitor cells related to diseases by aggregating the effects of genes with low average expression levels, such as key transcription factors related to cell development; at the same time, the sparsity and technical noise of scRNA-seq data can be effectively reduced, and good robustness and ability in identifying feature-related cell types and subgroups are shown. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained according to these drawings without creative labor for those skilled in the art.
[0049] Figure 1 is an analysis schematic flowchart of the processing method of respiratory system disease data provided by the embodiments of the present application;
[0050] Figure 2 is a processing device schematic diagram of respiratory system disease data provided by the embodiments of the present application;
[0051] Figure 3 is a processing system schematic flowchart of respiratory system disease data provided by the embodiments of the present application;
[0052] Figure 4 is an overview diagram of obtaining gPAS based on the scoring method of single cell pathway, and outputting TRS, trait-related genes, trait-related cells, trait-related cell types / subgroups, and trait-type tube pathway provided by the embodiments of the present application. DETAILED DESCRIPTION
[0053] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application.
[0054] In some of the flowcharts described in the specification and claims of the present application and in the above-mentioned drawings, a plurality of operations are included in the description of the flowcharts in a specific order, but it should be clearly understood that these operations can be executed or performed in parallel or in the order in which they appear in this text, and the serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these flowcharts can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this text are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do "first" and "second" represent different types.
[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0056] Figure 1 is a schematic flowchart of a respiratory disease data processing method provided by an embodiment of the present application. Specifically, the method comprises the following steps:
[0057] 101: obtaining single-cell sequencing sequence data to be analyzed;
[0058] In one embodiment, the single-cell sequencing data includes seven independent single-cell RNA-seq (scRNA-seq) or single-nucleus RNA-seq (snRNA-seq) data sets, covering 1.39 million cells from homo sapiens and mus musculus. For blood cells, two scRNA-seq data sets based on human BMMC (N=35,582 cells) and human PBMC (N=97,039 cells) were collected to reveal trait-related cell subpopulations or types. For immune / metabolic-related diseases / traits, scRNA-seq data sets from human cells (HCL, N=513,707 cells in 35 adult tissues) were used to construct a pseudo-bulk expression profile for each tissue and a priority risk tissue related to the disease / traits.
[0059] In one embodiment, to identify immune cell populations associated with severe respiratory disease, a large-scale PBMC scRNA-seq dataset (N = 469,453 cells) was collected, comprising 254 peripheral blood samples with varying degrees of respiratory disease severity (mild N = 109 samples, moderate N = 102 samples, severe N = 50 samples) and 16 healthy controls. Optionally, single-cell sequencing data to be analyzed included single-cell sequencing data from healthy controls and respiratory diseases of different severity levels.
[0060] 102: Machine learning methods were used to process the single-cell sequencing data to be analyzed, and the PAS score matrix and PAS of the cell pathway were obtained.
[0061] In one embodiment, the steps for processing the single-cell sequencing sequence data to be analyzed using machine learning methods to obtain the PAS score matrix of the cell pathway and the PAS of the cell pathway include:
[0062] Obtain pathway data for respiratory diseases;
[0063] The gene-cell matrix in single-cell sequencing data was standardized to obtain a standardized gene-cell matrix. Specifically, a variance stabilization transformation parameter with a scaling factor of 10,000 was used to standardize the sparse gene-cell matrix in scRNA-seq data, obtaining the standardized expression of a single gene in a single cell. The standardization formula is as follows: Among them, a g,j This indicates the original expression of gene g in cell j, e g,j This represents the normalized expression of gene g in cell j;
[0064] Based on pathway data of respiratory diseases, machine learning methods are used to convert the standardized gene-cell matrix into a pathway-cell matrix. The pathway-cell matrix is then used to obtain the cell-pathway PAS score matrix, which includes the pathway activity score (PAS) of a single cell in a single pathway.
[0065] In one embodiment, the pathway data is KEGG pathway data, with pathways from the KEGG database serving as the default gene set for evaluating PAS. The normalized gene-cell matrix is converted into a pathway-cell matrix using singular value decomposition (SVD). P... i This represents the set of genes in pathway i. For each pathway i, matrix A is selected from the normalized gene-cell matrix A. i , where matrix A i The columns are all N cells, and the rows are the pathway gene sets P. i China | Pi | gene, obtained according to SVD , where U represents an N x N orthogonal matrix, Σ represents a diagonal matrix with all zero except the main diagonal elements, and V T represents | P i | x | P i | orthogonal matrix; for right orthogonal matrix The t-th column vector v t represents the t-th principal component, reflecting the coordinated expression variability of genes in the pathway in single-cell data; since the first principal component PC1 represents the largest variance variation, the projection of the cell j feature in PC1 represents the PASs of the pathway i i,j ; for cell j, all expression variances in the pathway i are used to adjust the original PASs i,j ; for gene g in the pathway i, the minimum-maximum value scaling method is used to re-adjust the gene expression e g,j The adjusted gene expression is
[0066] In one embodiment, the pathway activity score PAS is optimized to obtain a weighted PAS;
[0067] The obtaining method of the weighted PAS includes:
[0068] wherein, represents the weighted PAS, represents the standardized expression of gene g in cell i after optimization, s i,j represents the pathway activity score PAS of cell j in pathway i;
[0069] In one embodiment, The obtaining method of the weighted PAS includes:
[0070]
[0071] wherein, represents the standardized expression of gene g in cell i, MAX(e g,j ) represents the maximum value of gene expression in pathway i, and MIN(e g,j ) represents the minimum value of gene expression in pathway i.
[0072] Optionally, the machine learning method includes a singular value decomposition SVD method; the singular value decomposition SVD method greatly improves the calculation efficiency of analyzing sparse matrices, and can obtain eigenvalues without calculating the variance matrix; the singular value decomposition method is used to convert the standardized gene-cell matrix into a low-dimensional pathway-cell matrix.
[0073] 103: obtaining genetic association data of respiratory diseases, processing the genetic association data to obtain pathway data with SNPs annotation;
[0074] In one embodiment, the genetic association data of respiratory diseases comprises genetic association data of severe respiratory diseases;
[0075] In one embodiment, the step of processing the genetic association data to obtain pathway data with SNPs annotation comprises:
[0076] Screening SNPs of single genes from the genetic association data, mapping the SNPs of single genes to corresponding pathways based on the pathway data of respiratory diseases to obtain pathway data with SNPs annotation;
[0077] Optionally, the step of obtaining SNPs of single genes comprises: after obtaining SNPs of genes in the genetic association data, respectively assigning SNPs gene pairs to obtain assignment results;
[0078] Respectively treating repeated genes of several single SNPs corresponding to multiple genes in the assignment results as independent SNP gene associations; retaining SNPs with minor allele frequency (MAF) greater than 0.1 in the assignment results; deleting SNPs on sex chromosomes; obtaining SNPs of single genes;
[0079] After the SNPs of single genes are summarized, they are all the SNPs of genes. Specifically, the genetic association data is GWAS data, and the SNPs in the GWAS summary statistics are assigned to related genes with 20 kb as the default parameter; the symbol g(k) is used to represent the gene g with SNP k; through the assignment of SNP gene pairs, there are several single SNPs corresponding to multiple genes; since the whole process needs to infer parameters from thousands of SNPs, the above SNPs corresponding to multiple genes have no effect on the inference process, so the above repeated genes need to be treated as independent SNP gene associations; SNPs with minor allele frequency (MAF) greater than 0.1 are retained, SNPs on sex chromosomes are deleted, and finally SNPs of related genes are obtained;
[0080] Based on the pathways in the KEGG database, the genes with associated SNPs are annotated to the pathways, and the set of SNPs in the pathway i is represented by Si= formula 2; the linkage disequilibrium (LD) of the SNPs extracted from the GWAS summary data is calculated by using the data of the third phase of the 1000 Genomes Project; the functional gene sets such as GO, Reactome, and MSigDB are provided as alternative options. In addition, the major histocompatibility complex region Chr6:25-35Mbp with extensive LD is deleted.
[0081] In one embodiment, the GWAS data has given phenotypes, and the phenotype annotation of the given phenotypes includes dichotomy, continuous dependent characteristics or internal phenotypes and central measurements.
[0082] 104: statistically analyze and process the PAS of the cellular pathways and the pathway data with SNP annotation, to obtain an estimated coefficient;
[0083] In one embodiment, the step of statistically analyzing and processing the PAS of the cellular pathways and the pathway data with SNP annotation to obtain an estimated coefficient includes:
[0084] Based on the pathway data with SNP annotation, the genetic effect value of all SNPs in a single pathway data is obtained;
[0085] Based on the polygenic regression model of the genetic correlation data of respiratory diseases, based on the PAS and the genetic effect value, the distribution of the genetic effect value is parameter estimated to obtain an estimated coefficient;
[0086] Optionally, the formula for obtaining the genetic effect value is: Wherein, β represents the theoretical effect size vector of m SNPs, ε represents a random environmental error, R represents an LD matrix, X T represents the standard genotype of SNPs in the genetic correlation data sample;
[0087] Optionally, the estimated coefficient is obtained by:
[0088]
[0089] Wherein, τ i,j represents the estimated coefficient of pathway i in cell j, the estimated coefficient reflects the influence of cell-specific PAS on the variance of GWAS effect size, that is, the influence of genetics on the reaction; τ0 represents an intercept term, σ 2 represents the variance of SNP effect size in the pathway, represents the weighted PAS.
[0090] In one embodiment, S iS represents the SNP set containing all SNPs in the positioning gene of each pathway i, the polygenic model assumes that the effect sizes of all SNPs of the prior pathway i follow a multivariate normal distribution, wherein σ 2 S represents the SNP set containing all SNPs in the positioning gene of each pathway i, the polygenic model assumes that the effect sizes of all SNPs of the prior pathway i follow a multivariate normal distribution, wherein σ i S represents the SNP set containing all SNPs in the positioning gene of each pathway i, the polygenic model assumes that the effect sizes of all SNPs of the prior pathway i follow a multivariate normal distribution, wherein σ i S represents the SNP set containing all SNPs in the positioning gene of each pathway i, the polygenic model assumes that the effect sizes of all SNPs of the prior pathway i follow a multivariate normal distribution, wherein σ
[0091] In an embodiment, based on the prior assumption, the distribution of genetic effect values is estimated by using the following formula: The estimated coefficients are optimized by using this formula;
[0092] In an embodiment, in order to optimize the estimated coefficients of each pathway in the polygenic regression model, the polygenic regression model is optimized by using the method-of-moments approach which can significantly improve the computational efficiency and the consistent convergence of the estimation; then, the observed and expected squared effects of the SNPs related to each pathway are fitted, and the expected value is estimated by the following formula: wherein Tr represents the matrix trace.
[0093] 105: multiply the estimated coefficients by PAS and sum up to obtain the genetic correlation pathway activity score gPAS of the cell;
[0094] In an embodiment, the genetic correlation pathway activity score gPAS (gPj) is the gPAS related to the respiratory system disease, and the formula is:
[0095]
[0096] wherein, is the optimized estimated coefficient.
[0097] In an embodiment, the method further comprises: performing correlation analysis and sorting on the genetic correlation pathway activity score gPAS and the gene expression of each cell, and screening N trait-related genes;
[0098] Optionally, the method of performing correlation analysis and sorting on the genetic correlation pathway activity score gPAS and the gene expression of each cell comprises: determining the correlation between the expression of a single gene and gPAS by using the Pearson correlation coefficient (PCC), and sorting the genes according to the correlation to obtain N trait-related genes related to the respiratory system disease; specifically, in order to maximize the efficacy, the expression of each gene g is inversely weighted by the gene-specific technical noise level, and the noise level is estimated by modeling the average variance relationship between genes in the scRNA-seq data;
[0099] Optionally, N trait-related genes are the top 1000 or bottom 1000 trait-related genes sorted according to the descending or ascending order of relevance; N is not limited to 1000, and N is a natural integer. N trait-related genes include one or more of the following: CALM3, PIK3R1, IL32, CD3E, B2M, PRS29, and GZMB;
[0100] In one embodiment, the method further includes: calculating a trait-related score (TRS) for each cell based on N trait-related genes; clustering cells based on the TRS and the level P-value of individual cells to obtain trait-related cells associated with different grades of respiratory disease severity; the trait-related genes are significantly enriched in the following trait-related cells, including hay bone marrow naive T16 cells. T16 cells), lung naïve CD8+ T cells (lung CD8+ T cells, liver NKT cells, and brain naive T cells. T cells); the formula for obtaining the trait-related score (TRS) is: TRS = average RE(GS) - average RE(CG); where average RE(GS) is the average relative expression value of N trait-related gene sets in a given cell, and average RE(CG) is the average relative expression value of the same number of control gene sets randomly selected from an existing gene bank; RE stands for relative expression; GS stands for gene set; and CG stands for control gene set.
[0101] Optionally, the severity levels of respiratory illnesses can be categorized as mild, moderate, and severe.
[0102] Optionally, the trait correlation score (TRS) of N genes can be calculated using the cell scoring method of the AddModuleScore function in Seurat.
[0103] In one embodiment, the method further includes: obtaining trait-related cell types or subpopulations associated with respiratory diseases based on the block bootstrap method, and determining whether the cell type of an individual cell is related. Trait-related cell types or subpopulations (associated with severe respiratory diseases) include one or more of the following: CD8+ T cells, megakaryocytes, CD16+ monocytes; Genes highly expressed in CD8+ T cells include memory effector marker genes (GZMK, AQP3, GZMA, PRF1, and GNLY) and exhaustive effector marker genes (LAG3, TIGIT, GZMA, GZMB, PRDM1, and IFNG); specifically, a set of cells is treated as a pseudo-bulk transcriptome profile and the gene expression across cells within a given cell type is averaged; for associated cell types, the standard error is estimated using the block bootstrap method and the t-statistic for each cell type corresponding P-value is calculated. Given that the goal of the block bootstrap method is to preserve the data structure when sampling from the empirical distribution, the gene set is divided into multiple biologically meaningful blocks using the pathways of KEGG database and the above-mentioned pathway-based blocks are replaced with sampling. Under the default parameters, 200 iterations are performed for each cell type association analysis, and the default parameters can be modified during actual execution.
[0104] In one embodiment, the method further comprises: ranking the genetic correlation pathway activity scores gPAS, obtaining the traits related pathways related to respiratory diseases according to the ranking results (selecting the top ranked pathways in the ranking results) and the P-value of the pathways at the cell type level; the traits related pathways include ribosome, T cell receptor signaling pathway, primary immunodeficiency, natural killer cell-mediated cytotoxicity and platelet activation.
[0105] Specifically, the gPAS is ranked based on the central limit theorem; the symbol C t represents the cell type t, and the pathway percentage rank of each cell j in the cell type t is calculated using the following formula: t wherein, represents the gPAS rank of the pathway i in the cell j, and M represents the total number of pathways; similarly, the statistical importance value T of each pathway i in the cell type t is calculated using the following formula: i t wherein, The hypothesis is: H0: T i t = 0 vs H1: T i t > 0; the P-value of each pathway i in the cell type t is:
[0106] In one embodiment, the statistical significance of a single cell is determined by calculating the rank distribution of trait-associated genes to further assess whether the cell is significantly associated with the trait of interest; specifically, the percentage rank of trait-associated genes in the cell is derived, wherein r g,j represents the expression rank of gene g in cell j, G represents the number of trait-associated genes designated; the gene percentage rank follows a normal distribution U(0, 1), and the statistical value T j of each cell is obtained under the null hypothesis that there is no correlation between the percentage ranks of genes
[0107] Based on a large number of cells in single-cell data, the distribution of T j is derived using the central limit theorem: where N is the total number of cells; the hypothesis for significance testing is: H0: T j = 0 vs H1: T j > 0; the P value of each cell j is: p j = Pr(T j ≤ t).
[0108] 106: output genetic correlation pathway activity score gPAS;
[0109] An application for detecting new CD8+ T cell subsets in the preparation of products for diagnosing respiratory diseases; new CD8+ T cell subsets are new discoveries related to respiratory diseases.
[0110] Figure 2 is The embodiment of the present application provides a kind of respiratory system disease data processing equipment schematic flow chart, equipment includes: memory and processor;Memory is used to store program instruction;Processor is used to call program instruction, when program instruction is executed, it is used to execute the processing method of respiratory system disease data described above.
[0111] Figure 3 is The embodiment of the present application provides a kind of respiratory system disease data processing system schematic flow chart, comprising:
[0112] The acquisition unit 301 is used to acquire single-cell sequencing sequence data to be analyzed;
[0113] The first processing unit 302 is used to process the single-cell sequencing sequence data to be analyzed by adopting the method of machine learning, to obtain the PAS score matrix of cell pathway and the PAS of cell pathway;
[0114] The second processing unit 303 is configured to obtain genetic correlation data of respiratory diseases, and process the genetic correlation data to obtain pathway data with SNP annotation.
[0115] The third processing unit 304 is configured to perform statistical analysis and processing on the PAS of the cell pathway and the pathway data with SNP annotation to obtain an estimated coefficient.
[0116] The fourth processing unit 305 is configured to multiply the estimated coefficient by the PAS and sum up to obtain a genetic correlation pathway activity score gPAS of the cell.
[0117] The output unit 306 is configured to output the genetic correlation pathway activity score gPAS.
[0118] A computer readable storage medium has a computer program stored thereon, and the computer program is executed by a processor to implement the processing method of the respiratory system disease data.
[0119] Figure 4 is The scoring method based on the single cell pathway provided in the embodiment of the present application obtains gPAS, and outputs TRS, trait-related genes, trait-related cells, trait-related cell types / subpopulations, and trait-related cell pathway overview diagrams;
[0120] wherein A represents converting a gene-cell matrix into a pathway-cell matrix by using a singular value decomposition method, PC1 represents the PAS of each pathway; B represents annotating SNPs in the GWAS data into the corresponding pathways; C represents a polygenic regression model; wherein the diagram at the top represents inferring the estimated coefficient in each pathway by using the polygenic regression model, and then calculating gPAS by using the estimated coefficient and the corresponding PAS, the diagram at the bottom represents a Pearson correlation model, which is used to associate the gPAS of each cell with all single cells to rank the trait-related genes; the top N trait-related genes (the default is the top 1,000) are obtained by using the AddModuleScore function in Seurat; D represents output, which includes four outputs: trait-related cells, trait-related cell types, trait-related pathways and trait-related genes.
[0121] The verification results of the verification embodiment show that assigning the inherent weight to the indication can moderately improve the performance of the method compared with the default setting.
[0122] Those skilled in the art can clearly understand the specific working processes of the system, device and unit described above for the convenience and brevity of description, and the corresponding processes in the foregoing method embodiments can be referred to, which will not be described herein.
[0123] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other manners. For example, the above-described device embodiments are merely illustrative, for example, the division of the units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the shown or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0124] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0125] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0126] Those skilled in the art can understand that all or part of the steps in the above embodiments can be completed by programs instructing relevant hardware, and the programs can be stored in a computer readable storage medium, which can include read only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0127] Those skilled in the art can understand that all or part of the steps in the above embodiments can be completed by programs instructing relevant hardware, and the programs can be stored in a computer readable storage medium, which can include read only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0128] The above has introduced in detail a computer device provided by the present application. For those skilled in the art, according to the idea of the embodiment of the present application, there will be changes in specific implementation manners and application ranges. In conclusion, the content of the present application should not be understood as a limitation.
Claims
1. A method for processing respiratory disease data, comprising: Obtain single-cell sequencing data to be analyzed; Machine learning methods were used to process the single-cell sequencing data to be analyzed, resulting in a PAS score matrix and PAS values for the cell pathway. Genetic association data of respiratory diseases are obtained, and the genetic association data are processed to obtain pathway data with SNP annotations. Statistical analysis was performed on the PAS data of the cellular pathway and the pathway data annotated with SNPs to obtain estimated coefficients; The steps to obtain the estimated coefficients include: obtaining the genetic effect values of all SNPs in a single pathway based on pathway data with SNP annotations; using a multi-gene regression model based on genetic association data of respiratory diseases, estimating the distribution of genetic effect values based on PAS and genetic effect values to obtain the estimated coefficients; the formula for obtaining the genetic effect values is: ;in, This represents the vector of theoretical effect sizes of m SNPs. Let X represent random environmental error, R represent the LD matrix, and X represent the random environmental error. T This represents the standard genotype of SNPs in a genetic association data sample; Multiply the estimated coefficients by PAS and then sum them to obtain the cell’s genetic pathway activity score gPAS; Output the activity score gPAS of the genetically related pathways.
2. The method for processing respiratory disease data according to claim 1, characterized in that, The methods for obtaining the estimated coefficients include: in, The estimated coefficient of pathway i in cell j represents the effect of cell-specific PAS on the size and variance of the GWAS effect, i.e. the effect of genetics on the response. Represents the intercept term. The variance representing the magnitude of the SNP effect in the pathway. This indicates weighted PAS.
3. The method for processing respiratory disease data according to claim 1, characterized in that, The formula for obtaining the genetic pathway activity score gPAS is as follows: Wherein, the gP j For gPAS, the These are the optimized estimated coefficients.
4. The method for processing respiratory disease data according to claim 1, characterized in that, The method further includes: performing correlation analysis and ranking of the genetic pathway activity score gPAS and the gene expression level of each cell to screen out N trait-related genes; specifically, the method includes: determining the correlation between the expression of a single gene and gPAS using the Pearson correlation coefficient, ranking the genes according to the correlation, and obtaining the N trait-related genes.
5. The method for processing respiratory disease data according to claim 4, characterized in that, The N trait-related genes are the top 1000 or bottom 1000 trait-related genes sorted according to the descending or ascending order of correlation.
6. The method for processing respiratory disease data according to claim 4, characterized in that, The N trait-related genes include one or more of the following: CALM3, PIK3R1, IL32, CD3E, B2M, PRS29, and GZMB.
7. The method for processing respiratory disease data according to claim 4, characterized in that, The method further includes: calculating the trait-related score (TRS) for each cell based on the N trait-related genes; and clustering the cells based on the TRS and the level P-value of individual cells to obtain trait-related cells associated with different levels of respiratory disease severity.
8. The method for processing respiratory disease data according to claim 7, characterized in that, The trait correlation score (TRS) of the N trait-related genes was calculated using a cell scoring method.
9. The method for processing respiratory disease data according to claim 7, characterized in that, The different severity levels of the respiratory diseases include mild, moderate, and severe.
10. The method for processing respiratory disease data according to claim 1, characterized in that, The method also includes obtaining trait-related cell types or subpopulations based on the block bootstrap method.
11. The method for processing respiratory disease data according to claim 1, characterized in that, The method further includes: ranking the activity scores gPAS of the genetically related pathways, and obtaining the trait-related pathways based on the ranking results and the P-values of the pathways at the cell type level.
12. A device for processing respiratory system disease data, the device comprising: Memory and processor; The memory is used to store program instructions; The processor is used to invoke program instructions, which, when executed, are used to perform the method for processing respiratory disease data according to any one of claims 1-11.
13. A system for processing respiratory disease data, comprising: The acquisition unit is used to acquire single-cell sequencing sequence data to be analyzed. The first processing unit is used to process the single-cell sequencing sequence data to be analyzed using machine learning methods to obtain the PAS score matrix of the cell pathway and the PAS of the cell pathway. The second processing unit is used to acquire genetic association data of respiratory diseases, process the genetic association data, and obtain pathway data with SNP annotations. The third processing unit is used to perform statistical analysis on the PAS of the cell pathway and the pathway data annotated with SNPs to obtain estimated coefficients. The steps to obtain the estimated coefficients include: obtaining the genetic effect values of all SNPs in a single pathway based on pathway data with SNP annotations; using a multi-gene regression model based on genetic association data of respiratory diseases, estimating the distribution of genetic effect values based on PAS and genetic effect values to obtain the estimated coefficients; the formula for obtaining the genetic effect values is: ;in, This represents the vector of theoretical effect sizes of m SNPs. Let X represent random environmental error, R represent the LD matrix, and X represent the random environmental error. T This represents the standard genotype of SNPs in a genetic association data sample; The fourth processing unit is used to multiply the estimated coefficients by PAS and then sum them to obtain the cell's genetic pathway activity score gPAS; The output unit is used to output the genetically related pathway activity score gPAS.
14. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for processing respiratory disease data according to any one of claims 1-11.