A method for screening reliability of single-cell expression pattern difference based on multiple group comparisons
Patent Information
- Application Number
- CN202311487781.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-08
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-02-08
AI Technical Summary
但由于单细胞组学庞大的数据量与细胞量、复杂的数据状况,现有技术难以直接、准确评估细胞表达模式的差异
本发明通过计算各细胞类型下,各组的特征表达谱与细胞类型表达谱的距离,对所有距离进行排名,基于距离排名计算各细胞类型的差异富集得分,可以比较各组在各细胞类型中的差异富集得分差异,获得相应的表达模式差异幅度评估结果,为科学选择后续研究方向提供依据。
Smart Images

Figure CN117457072B_ABST
Abstract
Description
[0001] Cross-references to related applications This application is a divisional application of Chinese patent application No. 2023100835857, filed on February 8, 2023, entitled "A method for evaluating the difference in single-cell expression patterns based on multi-group comparison". Technical Field
[0002] This invention relates to the field of bioinformatics analysis technology, and in particular to a method for evaluating differences in single-cell expression patterns based on multi-group comparisons. Background Technology
[0003] Single-cell transcriptomics, single-cell ATAC omics (high-throughput analysis of single-cell chromatin transposase accessibility), and single-cell epigenetics, among other single-cell level sequencing technologies, can obtain RNA and chromatin information for thousands or tens of thousands of genes within a single cell, comprehensively revealing gene expression differences between various cells. High-throughput single-cell sequencing platforms (such as those from 10X Genomics) utilize microfluidics, droplet encapsulation, and barcode tagging technologies to achieve high-throughput cell sorting and capture, capable of isolating and labeling hundreds or even tens of thousands of cells at once. After amplification and sequencing, information such as transcriptome or chromatin and site methylation can be obtained for each cell, offering advantages such as high cell throughput, low library construction cost, and high efficiency. Combined with the characteristic signal features (marker genes) of different cell types or cell type labeling algorithms (such as SingleR), this technology can be used to analyze the expression, chromatin, or methylation characteristics of different cell types, and further applied to research on biological development, disease progression, and immune changes.
[0004] The analysis of single-cell omics data typically involves the following steps: filtering low-quality cells, identifying cell types, preliminary analysis of the overall characteristics or expression patterns of each cell type, selecting cell types for in-depth analysis, and performing personalized analysis for the target cell type (such as pathway enrichment analysis, prediction of differentiation trajectories, transcription factor activity, cell communication, etc.). The final step, personalized analysis, can theoretically be performed using an unlimited number of existing software programs, thus usually requiring the most manpower, effort, and time. The preceding step of selecting cell types is the foundation of personalized analysis—selecting key and appropriate cell types allows subsequent personalized analysis to uncover important and meaningful information; conversely, inappropriate cell type selection may result in significant time and effort being wasted, only to uncover worthless or low-value information. Especially in single-cell datasets with multiple comparisons, in-depth research on each cell type increases the research cost exponentially compared to no comparison group or two-group comparisons.
[0005] Currently, researchers often choose subsequent research directions for single-cell data with multiple treatment groups (three or more groups) based on existing biological knowledge, rather than on the characteristics and differences of the data itself. This may lead to the overlooking of many potentially valuable research directions. The existence, magnitude, and significance of differences among different cell types across multiple groups all indicate the degree of influence of treatments or experimental conditions on the expression patterns of different cells. However, due to the massive amount of data and cell quantity, and the complex data conditions in single-cell omics, current technologies struggle to directly and accurately assess differences in cell expression patterns. Summary of the Invention
[0006] To address at least one of the technical problems mentioned in the background section, the present invention aims to provide a method for evaluating the differences in single-cell expression patterns based on multiple group comparisons. This method can provide precise indicators for assessing the magnitude and enrichment of differences in single-cell expression profiles under multiple group comparisons, thus providing a basis for the scientific selection of subsequent research directions.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for assessing single-cell expression pattern differences based on multi-group comparisons includes the following steps: S101, which combines single-cell transcriptome expression profiles from multiple groups; S102, grouping all cells in the merged single-cell transcriptome expression profile; S103, cell type identification for each cell population; S104, screen out cell types that coexist in two or more groups, and extract the co-expression profile of the cell type; S105, Calculate the differential enrichment score for the corresponding cell type based on the coexistence expression profile; S106, rank the cell types according to their differential enrichment scores from highest to lowest.
[0008] Furthermore, the method for calculating the difference enrichment score is as follows: S1051, based on the coexistence expression profile, the overall characteristic expression profile of this cell type was obtained; S1052, the total feature expression spectrum is divided into several sub-feature expression spectra according to the group; S1053, calculate the distance between each sub-feature expression spectrum and the corresponding total feature expression spectrum; S1054, repeat S1051 to S1053 for the remaining coexisting expression profiles to obtain the corresponding distances; S1055: Sort all distances from largest to smallest and assign them a ranking; S1056, Calculate the differential enrichment score for cell types based on ranking: ; ; ; in, The score represents the difference enrichment score. Represents ranking, Representing cell types of the same genus The rank set corresponding to the distance, index Represents cell type, subscript Representative group; Represents the total number of rankings; Representing the same cell type The number of groups; To enrich weights, .
[0009] Furthermore, .
[0010] Furthermore, the distance is either Euclidean distance or Manhattan distance.
[0011] Furthermore, the method for solving the total feature expression profile is as follows: for the cell population corresponding to the screened cell type, the average value of the gene expression level of each gene is taken as the feature value of the gene, and the feature values of several genes constitute the total feature expression profile.
[0012] Furthermore, in step S106, the cell type is also subjected to reliability screening: S1061, compare the ranking of all distances in this cell type with the ranking of all distances in other cell types, and calculate the p-value using the Wilcoxon rank-sum test; S1062, obtain the corresponding p-values for other cell types; S1063, delete cell types with p values higher than a preset threshold.
[0013] Furthermore, in S102, the cells are normalized and / or PCA dimensionality reduction is performed before the cells are clustered.
[0014] Furthermore, in S103, cell types are identified using cell type-specific marker genes, and cell populations that highly express marker genes are identified as the corresponding cell types; or, SingleR software is used to identify each cell type.
[0015] A computer storage medium storing a computer program, characterized in that, when executed by a processor, the program implements the single-cell expression pattern difference assessment method based on multiple group comparisons as described above.
[0016] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the single-cell expression pattern difference assessment method based on multiple group comparisons as described above.
[0017] Compared with the prior art, the beneficial effects of the present invention are: This invention calculates the distance between the characteristic expression profile of each group and the expression profile of each cell type under each cell type, ranks all distances, and calculates the differential enrichment score of each cell type based on the distance ranking. It can compare the difference in differential enrichment scores of each group in each cell type, obtain the corresponding expression pattern difference magnitude assessment results, and provide a basis for scientifically selecting subsequent research directions. Attached Figure Description
[0018] Figure 1 This is an overall flowchart of an embodiment of the present invention.
[0019] Figure 2 This is a flowchart of the difference enrichment scoring process according to an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Example 1: Please see Figure 1 This embodiment provides a method for assessing single-cell expression pattern differences based on multiple group comparisons, including the following steps: S101, a single-cell transcriptome expression profile combining multiple groups; as shown in Table 1 below: Table 1 Merged expression profile
[0022] In Table 1 above, Group1 to Group3 represent groups; Samp1 to Samp6 represent sample names; C1 to C22001 represent cells; Gene1 to Gene3 represent genes; and the numbers in the main text represent gene expression levels.
[0023] S102 uses Seurat software to cluster all cells in the merged single-cell transcriptome expression profile; it can also perform normalization and dimensionality reduction, with "LogNormalize" selected as the normalization method and "vst" used as the screening method for hypervariable genes.
[0024] S103, cell type identification was performed on each cell population. There are two common methods for identification: one uses cell type-specific marker genes (such as CD4 and CD8 genes in human T cells) to identify cell populations with high expression of these marker genes as the corresponding cell type; the other uses software with cell identification capabilities (such as SingleR and scCATCH) to determine the cell type of each cell population. The identification results are shown in Table 2 below: Table 2 Cell identification results
[0025] In Table 2 above, Type 1 to Type 4 are the identified cell types.
[0026] S104, count the number of groups in which each cell type exists, screen out cell types that exist in two or more groups at the same time, and extract the co-expression profile of the cell type.
[0027] If cell type Type1 exists in all three groups, then the number of groups in which cell type Type1 exists is 3.
[0028] Remove cell types with fewer than 2 subgroups, and denote the remaining n cell types as follows: to The number of corresponding groups is marked as to And extract its corresponding expression profile. Table 3 below shows the co-expression profiles corresponding to cell type Type 1. : Table 3 Expression profiles corresponding to cell type Type 1
[0029]
[0030] In Table 3 above, T1 represents the original cell type Type 1, which is present in all three groups. =3.
[0031] S105, Calculate the differential enrichment score for the corresponding cell type based on the coexistence expression profile to measure the degree of differential enrichment for that cell type.
[0032] Specifically, such as Figure 2 As shown, the method for calculating the difference enrichment score is as follows: S1051, based on coexistence expression profile The overall characteristic expression profile of this cell type was obtained. (The 0 in the subscript represents the overall characteristic expression profile of this cell type).
[0033] The method for solving the overall characteristic expression profile is as follows: For the cell population corresponding to the screened cell type, the average gene expression level of each gene is taken as the characteristic value of that gene, and the characteristic values of several genes constitute the overall characteristic expression profile. Table 4 below shows the cell type Type 1 (i.e., the corresponding co-existing expression profile). Total Feature Expression Spectrum : Table 4 Overall Characteristic Expression Profile of Cell Type 1
[0034] In Table 4 above, taking a value of 17.6 as an example, the value was calculated using the coexistence expression profile. The average gene expression level of Gene1 was obtained.
[0035] In other embodiments, the method for calculating the characteristic expression profile can be adjusted according to the data distribution—when the gene expression level is normally or uniformly distributed, it can be obtained by calculating the average gene expression level of each gene; if it is otherwise distributed, it can be obtained by the median of gene expression levels, or by calculating the average of the squares (or square roots or logarithms of gene expression levels, which can eliminate skewness in the distribution).
[0036] S1052, the total feature expression spectrum is divided into several sub-feature expression spectra according to the groups. Specifically, there are... Each group includes cell type T1. Calculate the values for group 1 to group 2. Sub-characteristic expression profiles of T1 cells were analyzed and labeled as follows. to ( to The first 1 in the subscript represents cell type T1, the second 1 and... (Representative group). Table 5 below shows the sub-expression profiles of the total characteristic expression profile of cell type T1, divided into several sub-expression profiles: Table 5. Sub-characteristic expression profiles of the total characteristic expression profile of cell type T1.
[0037] S1053, calculate the distance between each sub-feature expression spectrum and the corresponding total feature expression spectrum; Calculate the expression spectrum of these sub-features ( to Each cell type characteristic expression profile The distance. As shown in Table 5 above, firstly, the expression profile of cell type T1 is... Segmented by group Each expression profile is calculated using the same method as the overall characteristic expression profile in step S1051 (e.g., by calculating the average or median gene expression level for each gene, or by calculating the square, square root, or logarithm of the gene expression level). Sub-feature expression profiles of an expression profile to Then calculate the sub-feature expression spectrum separately. to With total feature expression spectrum European distance to (Manhattan distance can also be used instead), to Merge into a set .
[0038] S1054, for other cell types ( to By repeating S1051 to S1053 for the corresponding coexistence expression profiles and obtaining the corresponding distances, the set can be obtained. to . Set to Merge them to form a set B of all Euclidean distances, containing all distances for each cell type, i.e. to , to … to .
[0039] S1055: Sort all distances in descending order and assign them a ranking. Specifically, sort all distances in set B in descending order and assign them a corresponding ranking. ,Will to The corresponding ranking is defined as a set. to .
[0040] S1056, Calculate the differential enrichment score for cell types based on ranking: ; ; ; in, The score represents the difference enrichment score. Represents ranking, Representing cell types of the same genus The rank set corresponding to the distance, index Represents cell type, subscript Representative group; Represents the total number of rankings; Representing the same cell type The number of groups; To enrich weights, It can be adjusted as needed, but is usually set to 0.75.
[0041] S106. The cell types are scored according to differential enrichment. The cells are ranked from largest to smallest to assess the degree of differential enrichment in expression patterns across multiple groups, and to rank them according to their potential research value. The differential enrichment score is used to rank these groups. The larger the value, the higher the potential research value.
[0042] Example 2: Based on Example 1, Example 2 further performs reliability screening on the cell type in S106, specifically including the following steps: S1061, compare the ranking of all distances in this cell type with the ranking of all distances in other cell types, and calculate the p-value using the Wilcoxon rank-sum test; S1062, obtain the corresponding p-values for other cell types; S1063, delete cell types with p values higher than a preset threshold G, which is usually set to 0.05.
[0043] Example 3: This embodiment provides a computer storage medium storing a computer program, characterized in that, when the program is executed by a processor, it implements the single-cell expression pattern difference assessment method based on multiple group comparisons as described in Embodiment 1 or Embodiment 2.
[0044] Example 4: This embodiment provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The feature is that when the processor executes the computer program, it implements the single-cell expression pattern difference assessment method based on multiple group comparisons as described in Embodiment 1 or Embodiment 2.
[0045] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of equivalents of the claims be included within the present invention.
Claims
1. A reliable screening method for single-cell expression pattern differences based on multi-group comparisons, characterized in that, Includes the following steps: S101, which combines single-cell transcriptome expression profiles from multiple groups; S102, grouping all cells in the merged single-cell transcriptome expression profile; S103, cell type identification for each cell population; S104, screen out cell types that coexist in two or more groups, and extract the co-expression profile of the cell type; S105, Calculate the differential enrichment score for the corresponding cell type based on the coexistence expression profile; S106, rank the cell types according to their differential enrichment scores from highest to lowest; The cell types were subjected to reliability screening; the screening method is as follows: S1061, compare the ranking of all distances in this cell type with the ranking of all distances in other cell types, and calculate the p-value using the Wilcoxon rank-sum test; S1062, obtain the corresponding p-values for other cell types; S1063, delete cell types with p values higher than the preset threshold G; The difference enrichment score is calculated as follows: S1051, based on the co-existence expression profile of one of them, the overall characteristic expression profile of this cell type is obtained; S1052, the total feature expression spectrum is divided into several sub-feature expression spectra according to the group; S1053, calculate the distance between each sub-feature expression spectrum and the corresponding total feature expression spectrum; S1054, repeat S1051 to S1053 for the remaining coexisting expression profiles to obtain the corresponding distances; S1055: Sort all distances from largest to smallest and assign them a ranking; S1056, Calculate the differential enrichment score for cell types based on ranking: in, Representative difference enrichment score, Represents ranking, Representing cell types of the same genus The rank set corresponding to the distance, index Represents cell type, subscript Representative group; Represents the total number of rankings; Representing the same cell type The number of groups; To enrich weights, .
2. The method for reliable screening of single-cell expression pattern differences based on multi-group comparisons according to claim 1, characterized in that, 。 3. The method for reliable screening of single-cell expression pattern differences based on multi-group comparisons according to claim 1, characterized in that, The distance is either Euclidean distance or Manhattan distance.
4. The method for reliable screening of single-cell expression pattern differences based on multi-group comparisons according to claim 1, characterized in that, The method for solving the total feature expression profile is as follows: For the cell population corresponding to the screened cell type, the average value of the gene expression level of each gene is taken as the feature value of the gene, and the feature values of several genes constitute the total feature expression profile.
5. The method for reliable screening of single-cell expression pattern differences based on multiple group comparisons according to claim 1, characterized in that, In S102, the cells are normalized and / or PCA dimensionality reduction is performed before the cells are grouped.
6. The method for reliable screening of single-cell expression pattern differences based on multiple group comparisons according to claim 1, characterized in that, In step S103, cell types are identified using cell type-specific marker genes, and cell populations that highly express marker genes are identified as the corresponding cell types; or, SingleR software is used to identify each cell type.
7. A computer storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the single-cell expression pattern difference reliability screening method based on multiple group comparisons as described in any one of claims 1 to 6.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the single-cell expression pattern difference reliability screening method based on multiple group comparisons as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Analysis method based on 10X unicell transcriptome sequencing data
CN109979538A
Cell subset annotation method based on single cell transcriptome sequencing
CN112700820A