A method and system for calculating molecular markers of each subpopulation based on single cell omics
By optimizing molecular biomarker identification using the 'searchMarker' and 'filterMarker' methods based on single-cell omics, the problems of slow computation speed and low specificity in single-cell omics are solved, and efficient and accurate molecular biomarker identification is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHONGNAN HOSPITAL OF WUHAN UNIV
- Filing Date
- 2023-07-31
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies in single-cell omics are slow in computation, have low specificity, and are difficult to efficiently identify molecular markers of cell subpopulations, especially when computing large samples, they are inefficient and lack specificity for low samples.
A computational method based on single-cell omics is provided, including 'searchMarker' and 'filterMarker', which optimizes molecular marker identification and improves computational efficiency and specificity through maximum normalization, specificity scoring and visualization.
It significantly improves the computational speed and specificity of molecular markers in single-cell sequencing data, enhancing computational efficiency and the accuracy of molecular markers.
Smart Images

Figure CN116913381B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics technology, specifically relating to a novel, independent algorithm and system for identifying molecular markers of various cell subpopulations in the process of single-cell omics analysis. This algorithm is fast, accurate, and highly specific. Background Technology
[0002] Single-cell and single-nuclear RNA sequencing (scRNA-seq and snRNA-seq) reveals cell type heterogeneity, enabling the identification of different cell subpopulations through various methods. Annotation of cell subpopulations requires identifying marker genes with highly specific expression patterns for each subpopulation. These marker genes are used to distinguish between different subpopulations and facilitate the integration of single-cell sequencing data with subsequent experiments, such as fluorescence-activated cell sorting.
[0003] At the same time, as the number of cells generated in single-cell sequencing research continues to increase, and as technological advancements and sample throughput increase, researchers' demands for computing resources and time costs are also increasing, making analysis less efficient. This limits the flexibility of data analysis and poses challenges to the scalability and efficiency of effectively managing single-cell sequencing workflows.
[0004] On the other hand, thanks to the increased number of cells included in the study, researchers can iteratively re-cluster from cell subpopulations of interest to study biological processes within more precise cell types. However, after multiple re-clusterings, the differences between subpopulations gradually become small, thus requiring methods capable of identifying marker genes with higher specificity. Simultaneously, identifying subpopulations that exist only under specific conditions requires computational methods capable of identifying novel marker genes from small-scale subpopulations.
[0005] Seurat is widely considered one of the most important and powerful tools for single-cell data analysis. Marker gene identification in Seurat primarily relies on the "FindMarkers" or "FindAllMarkers" functions. This algorithm is based on the following logic: marker genes in a subpopulation are significantly upregulated compared to other subpopulations; therefore, averaging certain genes may weaken these characteristics. Ideally, marker genes should be present only in a single subpopulation; therefore, calculating genes for each subpopulation requires redundant computation and provides redundant information, especially when dealing with experiments with larger scales and more complex annotations. Thus, further screening methods may be needed after iterative re-clustering to find highly specific marker genes. Furthermore, identifying marker genes can be challenging for subpopulations with low cell counts.
[0006] Therefore, we are currently working on a rapid, accurate, and highly specific method for calculating molecular markers for each subgroup based on single-cell omics. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides a rapid, accurate, and highly specific method and system for calculating molecular markers for various subpopulations based on single-cell omics. This method can solve the problems of slow calculation speed for large samples and low specificity for small samples in the calculation of molecular markers during single-cell sequencing.
[0008] The method described in this invention comprises two independent methods, Method A and Method B, denoted as "searchMarker" and Method B as "filterMarker". "searchMarker" takes single-cell expression data and annotation data as input, efficiently and accurately providing molecular markers for each cell subpopulation. "filterMarker" further optimizes the molecular marker results calculated using the traditional Seurat scheme to obtain molecular markers with higher specificity and accuracy. Therefore, the computational method provided by this invention can efficiently search for molecular markers in single-cell sequencing data, and is expected to become a key technology in single-cell research.
[0009] The technical solution provided by this invention is as follows:
[0010] In a first aspect, the present invention provides a method for calculating molecular markers of various subgroups based on single-cell omics, including one or both of calculation method A and method B, wherein method A and method B are calculated independently;
[0011] Method A includes the following steps:
[0012] (1) Acquire single-cell omics data and accurately annotate the data by cell type;
[0013] (2) Calculate the average expression matrix using the annotation information;
[0014] (3) Perform maximum normalization calculation on the average expression matrix;
[0015] (4) Based on the cell type subgroup where the gene is most highly expressed, the cell type subgroup where the gene is located is divided into a high expression group and a low expression group;
[0016] (5) Filter low-expression genes, and for each gene, calculate the positive molecular index, negative molecular index and molecular index in the high expression group and low expression group respectively;
[0017] (6) Based on the number of high expression groups of genes, the basic expression level, and the molecular index, a comprehensive ranking was performed;
[0018] (7) Use heatmaps and bubble charts to visualize the analysis results, and provide all intermediate calculation indicators involved in the calculation process;
[0019] Method B's calculations are based on the output results of single-cell subset molecular marker calculations provided by the Seraut software, and include the following steps:
[0020] (1) Specificity score of the calculated output structure;
[0021] (2) Based on the cell type subpopulation where the gene is most highly expressed, the gene is assigned to each group;
[0022] (3) Sort according to gene specificity and eliminate redundant information;
[0023] (4) Visualize the results using heatmaps and bubble charts, and provide an optimized molecular marker matrix.
[0024] Furthermore, in step (1) of method A, the single-cell omics data is stored in the form of a Seurat S4 object, and the annotation of cell types is also stored in the S4 object.
[0025] Furthermore, in step (3) of method A, the method for calculating the maximum normalization is as follows:
[0026]
[0027] Where, x ij The average expression level of gene i in cell subpopulation type j; max x i This represents the maximum expression level of gene i in each cell subpopulation;
[0028] This represents the expression level after maximum standardization.
[0029] Furthermore, in step (4) of method A, the index S1, with a value ranging from 0 to 1, is used to perform binary classification on the subpopulations of each gene after maximum normalization. High expression is defined as the cell subpopulation with a value greater than S1 after maximum normalization, while low expression is defined as the other remaining cell subpopulations. The number of groups with values higher than S1 is denoted as n, and the number of groups with values lower than S1 is denoted as m.
[0030] Furthermore, in step (4) of method A, the cell subpopulations with a maximum normalized expression level of 1 for each gene are stored, and the calculation formula is as follows:
[0031]
[0032] Where c i The cell subpopulation representing the maximum normalized expression level of gene i is 1;
[0033] C represents all cell subpopulations;
[0034] This represents the maximum normalized expression level of gene i in cell subpopulation j;
[0035] Furthermore, in step (5) of method A, the gene expression ratio S2 at the end of the index is used, with a value range of 0-1, to sort the genes in each group according to their average expression level, and to filter the genes with the gene expression ratio S2 at the end of the index. Since the expression level and specificity are negatively correlated in the molecular marker identification calculation process, users can select an appropriate gene expression threshold according to their actual experimental situation to identify high-specificity molecular markers at that threshold.
[0036] The positive numerator index (PMI) and the negative numerator index (NMI) are calculated using the following methods, with the meanings of each variable being the same as described above:
[0037]
[0038]
[0039]
[0040]
[0041] in, This is represented as the average value of cell subpopulations higher than S1. x represents the average value of each cell subpopulation below S1. ij The average expression level of gene i in cell subpopulation type j. The expression level represents the maximum standardized value; PMI i With NMI i These represent the positive and negative molecule indices, respectively.
[0042] The molecular index (MI) is calculated using the following method:
[0043] MI i =PMI i -NMI i
[0044] Furthermore, in step (6) of method A, the sorting method is as follows: sort in ascending order according to the number of n, and sort in descending order according to the calculated value of the numerator index MI:
[0045]
[0046] Furthermore, in Method B, the specificity score T is calculated using the molecular marker matrix output by Seurta. i The specific calculation method is as follows:
[0047]
[0048] Where G1 represents the set of all cells expressing gene i, G clust The set represents the cell subpopulation with gene i as the molecular marker in the calculation results, and card() represents the number of elements in the set.
[0049] Furthermore, in step (3) of method B, the sorting method is as follows: sorting in descending order of specificity to obtain a molecular marker matrix arranged from high to low specificity; for duplicate genes, only T is retained. i The item that has the maximum value.
[0050] Secondly, the present invention provides a system of molecular markers for various subgroups based on single-cell omics, comprising: one or both of independent system A and system B;
[0051] System A includes a data format conversion module, which checks and unifies single-cell omics data and cell type annotations in various formats, and converts them into an average expression matrix to accelerate subsequent calculations.
[0052] The core computation module is used to calculate various indicators that measure the ability of marker genes, including: a maximum normalization module, a filtering algorithm module, a hierarchical mean module, and a ranking module. Among them, the maximum normalization module is used to extract features of gene expression levels in each cell subpopulation; the filtering algorithm module is used to filter low-expression genes according to user-specified indicators; the hierarchical mean module is used to calculate molecular markers; and the ranking module is used to rank genes according to their ability to serve as cell subpopulation markers based on multiple indicators.
[0053] Visualization module: Used to visualize the genes of each cell subpopulation;
[0054] System B includes:
[0055] Specificity calculation module: used to rank the ability of genes to be called subpopulation molecular markers based on the calculation output of single-cell subpopulation molecular markers provided by Seurat software;
[0056] The sorting module is used to sort genes according to their ability to serve as markers of cell subpopulations based on multiple indicators.
[0057] Visualization module: Used to visualize the genes of each cell subpopulation.
[0058] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method for calculating molecular markers of various subgroups based on single-cell omics as described in the first aspect.
[0059] Beneficial technical effects of the present invention:
[0060] a) This invention can significantly improve computational speed and efficiency compared to traditional computational schemes when the number of cells is large and the annotation information is complex, through the calculation of molecular markers.
[0061] b) The molecular markers screened by this invention have significantly improved specificity compared to traditional analytical methods. Attached Figure Description
[0062] Figure 1 This is a schematic diagram of the overall architecture and computing process;
[0063] Figure 2 This is a schematic diagram illustrating the calculation logic of the molecule index;
[0064] Figure 3 This is a diagram illustrating the table content in the output file;
[0065] Figure 4 Bubble graphs of molecular markers obtained under different S2 conditions; A, S2 = 0.0; B, S2 = 0.4; C, S2 = 0.8;
[0066] Figure 5 The bar chart shows the computation time of a relatively traditional molecular marker calculation scheme (Seurat-“FindAllMarkers”) tested on three different samples; A represents the computation time difference between the present invention (“searchMarker”) and the traditional scheme (Seurat) in human prefrontal cortex single-cell sequencing test data; B represents the computation time difference between the present invention (“searchMarker”) and the traditional scheme (Seurat) in human heart single-cell sequencing test data; C represents the computation time difference between the present invention (“searchMarker”) and the traditional scheme (Seurat) in mouse kidney single-cell sequencing test data. In the figure, starTracer represents the “searchMarker” of the present invention in the examples, and Seurat represents the method of the traditional scheme in the examples.
[0067] Figure 6 A bar chart illustrating the improvement in specificity compared to traditional molecular marker calculation schemes. The vertical axis represents the specificity of marker genes for each cell subpopulation calculated using this invention, further normalized to the corresponding Seurat-calculated cell subpopulation marker gene specificity value (fold(T)). iThe horizontal axis Vas represents the subgroups of each single-cell analysis, where "Seurat" represents the average value of the marker gene calculated by Seurat after normalization to itself.
[0068] Figure 7 T is a specificity measure. i A schematic diagram;
[0069] Figure 8 Stacked violin plot of the top 3 molecular markers for each cell subpopulation before and after filtering with “filterMarker”. Detailed Implementation
[0070] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0071] The following terms or definitions are provided merely to aid in understanding the invention. These definitions should not be construed as having specific meanings or implications.
[0072] It is less than the scope understood by those skilled in the art.
[0073] Unless otherwise defined below, all technical and scientific terms used in the specific embodiments of this invention are intended to have the same meaning as commonly understood by those skilled in the art. While it is believed that the following terms will be well understood by those skilled in the art, the following definitions are set forth to better explain the invention.
[0074] As used in this invention, the terms “comprising,” “including,” “having,” “containing,” or “involving” are inclusive or open-ended and do not exclude other unlisted elements or method steps. The term “consisting of” is considered a preferred embodiment of the term “comprising.” If a group is defined below as comprising at least a certain number of embodiments, this should also be understood to disclose a group that preferably consists only of those embodiments.
[0075] When referring to a singular noun, the indefinite or definite article used, such as "a" or "a kind of," "the," includes the plural form of the noun.
[0076] The terms "approximately" and "generally" in this invention refer to a range of accuracy that, as would be understood by those skilled in the art, still guarantees the technical effects of the discussed features. This term typically indicates a deviation from the indicated value of ±10%, preferably ±5%.
[0077] Furthermore, the terms first, second, third, (a), (b), (c), and similar terms used in the specification and claims are for distinguishing similar elements and are not necessary for the order of description or chronological sequence. It should be understood that such terms are interchangeable in appropriate contexts, and the embodiments described in this invention can be implemented in a different order than that described or illustrated in this invention.
[0078] The embodiments of the present invention will be described in detail below with reference to examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and should not be considered as limiting the scope of the invention. Unless otherwise specified in the examples, conventional conditions or conditions recommended by the manufacturer are followed. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased on the market.
[0079] Example 1
[0080] A computational method for molecular markers of various subpopulations based on single-cell omics, comprising the following steps:
[0081] 1. Acquire single-cell omics data and accurately annotate the data by cell type;
[0082] The specific steps are as follows:
[0083] 1.1 Obtain the mm39 genome for sequencing; the sequencing platform is Illuminal Novo-seq 6000.
[0084] 1.2 The data from the machine were processed using CellRanger and compared to the mouse mm39 genome.
[0085] 1.3 Use the Seurat package to read the output file of cellranger and perform doublet filtering and low-quality cell filtering.
[0086] 1.4. The samples were standardized using the SCTransform scheme, and the cells were clustered according to their gene expression patterns using the Shared Nearest Neighbor (SNN) clustering algorithm. The UMAP scheme was used for dimensionality reduction analysis and visualization.
[0087] 1.5. Use the Azimuth database to assist in cell type annotation and validate the annotation using common marker genes.
[0088] Further subgrouping of excitatory neurons was performed. Normalization was achieved using the SCTransform scheme, cell clustering based on gene expression patterns was performed using the Shared Nearest Neighbor (SNN) algorithm, and dimensionality reduction and visualization were conducted using the UMAP scheme.
[0089] 2. Using the processed single-cell experimental subjects as input, the system directly returns the matrix of marker genes for each subpopulation and the visualization results to the user; specifically, the calculation steps are as follows:
[0090] 2.1 Calculate the average representation matrix using annotation information:
[0091] The average expression matrix can be obtained by taking the data that has been standardized by UMI from the single-cell sparse expression matrix and then calculating the average expression level of each gene in cells belonging to the same cell subpopulation.
[0092] Specifically, it includes the following sub-steps:
[0093] 2.1.1 For users who input Seurat objects, this algorithm will call Seurat's "AverageExpression" function to calculate the average expression matrix.
[0094] 2.1.2 For users who directly input the average representation matrix, this algorithm does not perform any further processing.
[0095] 2.1.3 For users who input sparse matrices and annotation information, this algorithm will use an internal function to calculate the average representation matrix.
[0096] 3. Perform maximum normalization calculation on the average expression matrix;
[0097] The specific steps are as follows:
[0098] Using S4 class objects of excitatory neurons as input, the maximum normalization is calculated using the following formula:
[0099]
[0100] Where, x ij The average expression level of gene i in cell subpopulation type j; max x i This represents the maximum expression level of gene i in each cell subpopulation;
[0101] This represents the expression level after maximum standardization.
[0102] 4. Based on the cell type subpopulation where the gene is most highly expressed, the cell type subpopulation where the gene is located is divided into a high expression group and a low expression group;
[0103] The specific steps are as follows:
[0104] In the process of grouping cell subtypes with different expressions of any gene, S1 = 0.5 is set, and this index is used to binary classify the subpopulations of each gene after maximum normalization. High expression is defined as cell subpopulations with a value greater than 0.5 after maximum normalization, and low expression is defined as the remaining cell subpopulations. The number of groups with a value higher than 0.5 is denoted as n, and the number of groups with a value lower than 0.5 is denoted as m. The calculation logic is as follows: Figure 2 As shown.
[0105] Furthermore, in the high-expression group, cell subpopulations with a maximum normalized expression level of 1 for each gene were stored, and the specific calculation method was as follows:
[0106]
[0107] Where c i The cell subpopulation representing the maximum normalized expression level of gene i is 1;
[0108] C represents all cell subpopulations;
[0109] This represents the maximum normalized expression level of gene i in cell subpopulation j.
[0110] Each gene will be determined according to c i The cell subpopulations whose maximum normalized value is 1 are assigned to the corresponding groups. Specifically, each cell subpopulation will correspond to a gene set, representing marker genes that can serve as molecular markers for that subpopulation.
[0111] 5. Filter out low-expression genes. For each gene, calculate the positive molecular index, negative molecular index, and molecular index in both the high-expression and low-expression groups.
[0112] The specific steps are as follows:
[0113] Using the proportion of genes ending in the index S2, set to 0.3, genes in each group are sorted according to their average expression levels, and genes with a proportion of 30% ending in the index are filtered out. Since expression level and specificity are negatively correlated in the molecular marker identification calculation process, users can select an appropriate gene expression threshold based on their actual experimental conditions to identify high-specificity molecular markers below that threshold. This step can provide high-specificity molecular markers for each cell subpopulation in a very short time.
[0114] For each group, calculate the positive numerator index (PMI) and negative numerator index (NMI):
[0115]
[0116]
[0117]
[0118]
[0119] in, This is represented as the average value of cell subpopulations higher than S1. x represents the average value of each cell subpopulation below S1. ij The average expression level of gene i in cell subpopulation type j. This represents the expression level after maximum standardization. PMI i With NMI i These represent the positive and negative molecule indices, respectively.
[0120] The molecular index (MI) is calculated using the following formula:
[0121] MI i =PMI i -NMI i
[0122] 6. Rank the genes based on the number of overexpressed groups, baseline expression levels, and molecular indices.
[0123] The specific steps are as follows:
[0124] Genes are sorted in ascending order based on the number of genes with the same number of genes (n). For genes with the same number of genes (n), they are further sorted in descending order based on their MI values. After the calculation is complete, all intermediate variables from the calculation process will be output.
[0125] The formula is as follows:
[0126]
[0127] Sorting method examples are as follows Figure 3 As shown.
[0128] 7. Visualize the analysis results using heatmaps and bubble charts, and provide all intermediate calculation metrics involved in the calculation process.
[0129] The specific steps are as follows:
[0130] After obtaining the analysis results, heatmaps and bubble charts are used to visualize the results, while providing all intermediate calculation metrics involved in the calculation process.
[0131] refer to Figure 1 A system A based on molecular markers for various subpopulations in single-cell omics, comprising:
[0132] Data format conversion module: Used to check and unify annotations of single-cell omics data and cell types in various formats, and convert them into an average expression matrix to accelerate subsequent calculations.
[0133] The core computational module is used to calculate various indicators to measure the ability of marker genes, including: a maximum normalization module, a filtering algorithm module, a hierarchical mean module, and a ranking module. Among them, the maximum normalization module is used to extract features of gene expression levels in each cell subpopulation; the filtering algorithm module is used to filter low-expression genes according to user-specified indicators; the hierarchical mean module is used to calculate molecular markers; and the ranking module is used to rank genes according to their ability to serve as cell subpopulation markers based on multiple indicators.
[0134] Visualization module: Used to visualize the genes of each cell subpopulation.
[0135] Changes in molecular markers obtained under different S2 conditions, such as Figure 4 As shown. From Figure 4 As can be seen, with the continuous increase of S2, the expression level of molecular markers is improved, but the specificity gradually decreases.
[0136] refer to Figure 5 The figures show bar charts comparing the computation time of the present invention with that of the traditional molecular marker calculation scheme (Seurat-“FindAllMarkers”) tested on three different samples. A) Using the present invention in human prefrontal cortex single-cell sequencing data, the computation time was 10.79 seconds; the traditional scheme (Seurat) took 562.86 seconds, a fold increase of 52.15. B) Using the present invention in human heart single-cell sequencing data, the computation time was 0.76 minutes; the traditional scheme took 381.28 minutes, a fold increase of 503.2. C) Using the present invention in mouse kidney single-cell sequencing data, the computation time was 1.45 seconds; the traditional scheme took 44.28 seconds, a fold increase of 30.50. The results show that the present invention achieves a significant improvement in computation speed compared to the traditional scheme.
[0137] refer to Figure 6 The results showed that the specificity of molecular markers obtained using this analysis was improved in single-cell sequencing data of the human prefrontal cortex. The results indicated that the specificity of molecular markers obtained from all cell subpopulations was enhanced.
[0138] Example 2
[0139] For subgroup molecular markers that have already been obtained using the traditional Seurat software, this embodiment provides "filterMarker" as a downstream optimization strategy for the traditional approach to obtain subgroup molecular markers with higher specificity and accuracy.
[0140] The specific steps are as follows:
[0141] 1. Obtain the single-cell subpopulation molecular marker matrix output by the traditional analysis software Seurat.
[0142] 2. Input the marker matrix and the Seurat object into "filterMarker". After calculation, the optimized molecular marker matrices for each subgroup and the visualization results are returned directly. The specific calculation steps are as follows:
[0143] 2.1 When calculating the specificity score using the molecular marker matrix output by Seurat, the calculation method is as follows (refer to...). Figure 7 )for:
[0144]
[0145] Among them G i G represents the set of all cells that express gene i. clust This represents the set of cells in the cell subpopulation with gene i as the molecular marker in the calculation results; card() represents the number of elements in the set.
[0146] 2.2. Sort the markers in descending order using "filterMarker" to obtain a molecular marker matrix arranged from highest to lowest specificity. For duplicate genes, only T markers are retained. i The item that has the maximum value.
[0147] 3. Visualize the results using heatmaps and bubble charts, and provide an optimized molecular marker matrix.
[0148] refer to Figure 1 A system B based on molecular markers for various subgroups of single-cell omics, comprising:
[0149] Specificity calculation module: used to rank the ability of genes to be called subpopulation molecular markers based on the calculation output of single-cell subpopulation molecular markers provided by Seurat software;
[0150] The sorting module is used to sort genes according to their ability to serve as markers of cell subpopulations based on multiple indicators.
[0151] Visualization module: Used to visualize the genes of each cell subpopulation.
[0152] The test results are as follows Figure 8 As shown, from Figure 8As can be seen, after optimization using the "filterMaker" invention (right), the top three molecular markers (vertical axis) showed a significant improvement in specificity across various cell subpopulations (horizontal axis) compared to before optimization. This improvement was particularly pronounced for inhibitory interneurons (shown by the dashed box).
[0153] Another embodiment of the present invention also provides an electronic device, comprising:
[0154] One or more processors:
[0155] A memory for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to perform the method described in any of the above technical solutions.
[0156] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the invention.
Claims
1. A method for calculating molecular markers for various subpopulations based on single-cell omics, characterized in that, This includes method A or a combination of method A and method B; method A and method B are calculated independently. Method A includes the following steps: (1) Obtain single-cell omics data and accurately annotate the data by cell type; (2) Calculate the average expression matrix using the annotation information; (3) Perform maximum standardization calculation on the average expression matrix; (4) Based on the cell type subgroup where the gene is most highly expressed, the cell type subgroup where the gene is located is divided into a high expression group and a low expression group; Among them, the use of indicators The value ranges from 0 to 1. For each gene subpopulation after maximum standardization, a binary classification is performed, where the subpopulation is higher than... The number of groups is denoted as n lower than The number of groups is denoted as m ; (5) Filter low-expression genes. For each gene, calculate the positive molecular index, negative molecular index and molecular index in the high-expression group and the low-expression group respectively. Among them, the proportion of genes at the end of the indicator is used. The values range from 0 to 1. Genes in each group are ranked according to their average expression levels, and the proportion of genes at the end of the gene sequence is also considered. Filtering genes; The positive numerator index (PMI) and the negative numerator index (NMI) are calculated using the following methods: in, This indicates that gene i is higher than The average of the highest standardized expression levels in the group. This indicates that gene i is not higher than The average of the highest standardized expression levels in the group. The average expression level of gene i in cell subpopulation type j. This represents the expression level after maximum standardization. and These represent the positive and negative molecular indices, respectively. The molecular index (MI) is calculated using the following method: ; (6) Based on the number of high expression groups, basic expression levels, and molecular indices of genes, a comprehensive ranking is performed; (7) Use heatmaps and bubble charts to visualize the analysis results, and provide all intermediate calculation indicators involved in the calculation process; Method B's calculations are based on the output results of single-cell subset molecular marker calculations provided by the Seurat software, and include the following steps: (1) Specificity score of the calculation output; (2) Based on the cell type subpopulation where the gene is most highly expressed, the gene is assigned to each group; (3) Sort the data according to gene specificity and exclude redundant information; (4) Use heatmaps and bubble charts to visualize the results and provide an optimized molecular marker matrix.
2. The calculation method according to claim 1, characterized in that, In step (1) of method A, the single-cell omics data is stored as a Seurat S4 object, and the cell type annotation is also stored in the S4 object.
3. The method according to claim 1, characterized in that, In step (3) of method A, the method for calculating the maximum normalization is as follows: in, The average expression level of gene i in cell subpopulation type j; max x i This represents the maximum expression level of gene i in each cell subpopulation; This represents the expression level after maximum standardization.
4. The calculation method according to claim 1, characterized in that, In step (4) of method A, the cell subpopulations with a maximum normalized expression level of 1 for each gene are stored, and the calculation formula is as follows: in The cell subpopulation representing the maximum normalized expression level of gene i is 1; C Represents all cell subsets; This represents the maximum normalized expression level of gene i in cell subpopulation j.
5. The calculation method according to claim 1, characterized in that, In step (6) of method A, the sorting method is as follows: sort in ascending order according to the number of n, and sort in descending order according to the calculated value of the numerator index MI: 。 6. The calculation method according to claim 1, characterized in that, In method B, the specificity score is calculated using the molecular marker matrix output by Seurat. The specific calculation method is as follows: in The set of all cells that express gene i. The set represents the cell subpopulation with gene i as the molecular marker in the calculation results; card() represents the number of elements in the set.
7. A system based on molecular markers for various subpopulations in single-cell omics, characterized in that: For implementing the method of claim 1, the system comprises: a standalone system A or a combination of system A and system B; System A includes: Data format conversion module: Used to check and unify annotations of single-cell omics data and cell types in various formats, and convert them into an average expression matrix to accelerate subsequent calculations; The core computational module is used to calculate various indicators to measure the ability of marker genes, including: a maximum normalization module, a filtering algorithm module, a hierarchical mean module, and a ranking module. Among them, the maximum normalization module is used to extract features of gene expression levels in each cell subpopulation; the filtering algorithm module is used to filter low-expression genes according to user-specified indicators; the hierarchical mean module is used to calculate molecular markers; and the ranking module is used to rank genes according to their ability to serve as cell subpopulation markers based on multiple indicators. Visualization module: Used to visualize the genes of each cell subpopulation; System B includes: Specificity calculation module: used to rank the ability of genes to be called subpopulation molecular markers based on the calculation output of single-cell subpopulation molecular markers provided by Seurat software; The sorting module is used to sort genes according to their ability to serve as markers of cell subpopulations based on multiple indicators. Visualization module: Used to visualize the genes of each cell subpopulation.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the computational method for molecular markers of each subgroup based on single-cell omics as described in any one of claims 1 to 6.