Cell annotation method based on single cell space transcriptome and related device

Through a cell annotation method based on single-cell spatial transcriptome, multiple branch annotation models are used to perform branch annotation and comprehensive annotation prediction, the problems of low annotation efficiency and insufficient accuracy in the prior art are solved, and efficient and accurate cell type prediction is achieved.

CN120126546APending Publication Date: 2025-06-10BGI RES SOUTHWEST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311686928.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, the cell annotation efficiency of the single-cell spatial transcriptome is low, and it is difficult to obtain more accurate annotation results while ensuring the annotation efficiency.

Method used

A cell annotation method based on single-cell spatial transcriptome is proposed. By obtaining single-cell transcriptome data, spatial transcriptome data and multiple branch annotation models, these models are used for branch annotation and comprehensive annotation prediction, and comprehensive annotation results are generated for target single-cell spatial transcriptome.

Benefits of technology

While ensuring cell annotation efficiency, the accuracy of annotation results is significantly improved. By combining the annotation results of multiple branches, more reliable and accurate cell type predictions are generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126546A_ABST
    Figure CN120126546A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological information analysis, in particular to a cell annotation method based on a single cell space transcriptome and a related device. According to the cell annotation method based on the single-cell spatial transcriptome, single-cell transcriptome data, spatial transcriptome data and a plurality of branch annotation models need to be acquired firstly; converting the single cell transcriptome data into a single cell annotation file according to a first format constraint condition; and converting the spatial transcriptome data into a spatial annotation file according to a second format constraint condition. Inputting the single cell annotation file and the space annotation file into corresponding branch annotation models for branch annotation to obtain a branch annotation result output by each branch annotation model, and finally performing comprehensive annotation prediction based on a plurality of branch annotation results to obtain a comprehensive annotation result of the target single cell space transcriptome. According to the method, a more accurate annotation result can be obtained while good efficiency of cell annotation is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics analysis technology, and in particular, to a cell annotation method and related device based on single-cell spatial transcriptomics. Background Art

[0002] Single-Cell Spatial Transcriptomics is a genomics technology that can provide precise images of the gene expression and spatial distribution of cells in cell tissues. It is a high-dimensional cell analysis technology used to provide precise images of the gene expression and spatial distribution of cells in cell tissues.

[0003] There are various methods for annotating cells in single-cell spatial transcriptomics, and each type of cell has its suitable annotation method. In order to obtain more accurate annotation results, in related technologies, it is necessary to use multiple annotation methods to annotate the cells of the same single-cell spatial transcriptomics and then compare them. However, the efficiency of cell annotation in this way is relatively low. Therefore, how to obtain more accurate annotation results while ensuring good efficiency of cell annotation has become an urgent problem in the industry. Summary of the Invention

[0004] This application aims to at least solve one of the technical problems existing in the prior art. For this reason, this application proposes a cell annotation method and related device based on single-cell spatial transcriptomics, which can obtain more accurate annotation results while ensuring good efficiency of cell annotation.

[0005] The cell annotation method based on single-cell spatial transcriptomics according to the first aspect embodiment of this application includes:

[0006] Obtain single-cell transcriptome data, spatial transcriptome data, and multiple branch annotation models; wherein, the single-cell transcriptome data and the spatial transcriptome data are derived from the same target single-cell spatial transcriptome;

[0007] According to the first format constraint condition corresponding to each branch annotation model, convert the single-cell transcriptome data into a single-cell annotation file;

[0008] According to the second format constraint condition of each branch annotation model, convert the spatial transcriptome data into a spatial annotation file;

[0009] Input the single-cell annotation file and the spatial annotation file into the corresponding branch annotation model for branch annotation to obtain the branch annotation results output by each branch annotation model;

[0010] Based on the multiple branch annotation results, comprehensive annotation prediction is performed to obtain the comprehensive annotation result of the target single-cell spatial transcriptome.

[0011] In some embodiments of the present application, the converting the single-cell transcriptome data into a single-cell annotation file according to the first format constraint condition corresponding to each branch annotation model includes:

[0012] Extract annotation information from the single-cell transcriptome data to obtain a single-cell expression matrix and cell type information;

[0013] Generate the single-cell annotation file that meets the first format constraint condition according to the single-cell expression matrix and the cell type information.

[0014] In some embodiments of the present application, the converting the spatial transcriptome data into a spatial annotation file according to the second format constraint condition corresponding to each branch annotation model includes:

[0015] Extract annotation information from the spatial transcriptome data to obtain a spatial group expression matrix and spatial single-cell coordinate information;

[0016] Generate the spatial annotation file that meets the second format constraint condition according to the spatial group expression matrix and the spatial single-cell coordinate information.

[0017] In some embodiments of the present application, the performing comprehensive annotation prediction based on the multiple branch annotation results to obtain the comprehensive annotation result of the target single-cell spatial transcriptome includes:

[0018] Based on the branch annotation results, determine the type prediction result obtained by each branch annotation model for predicting the target single-cell spatial transcriptome;

[0019] Determine the target cell type according to the multiple type prediction results, and generate the comprehensive annotation result based on the target cell type.

[0020] In some embodiments of the present application, the number of the type prediction results is the first number;

[0021] The determining the target cell type according to the multiple type prediction results and generating the comprehensive annotation result based on the target cell type includes:

[0022] Group the first number of type prediction results to obtain the second number of prediction result groups; wherein, each prediction result group includes at least one of the type prediction results, and each prediction result group corresponds to a predicted cell type;

[0023] Determine a target group from multiple prediction result groups based on the number of the type prediction results included in the prediction result group;

[0024] Determine the predicted cell type corresponding to the target group as the target cell type;

[0025] Generate the comprehensive annotation result based on the target cell type.

[0026] In some embodiments of the present application, the determining a target group from multiple prediction result groups based on the number of the type prediction results included in the prediction result group includes:

[0027] Score each prediction result group based on the number of the type prediction results included in the prediction result group;

[0028] When the number of the type prediction results included in the prediction result group is more than two, the prediction result group gets a first score;

[0029] When the number of the type prediction results included in the prediction result group is less than two, the prediction result group gets a second score; wherein, the second score is less than the first score;

[0030] Determine the target group from the prediction result groups that get the first score.

[0031] In some embodiments of the present application, the determining a target group from multiple prediction result groups based on the number of the type prediction results included in the prediction result group includes:

[0032] Compare the numbers of the type prediction results included in multiple prediction result groups, and determine the prediction result group that includes the most type prediction results as the target group.

[0033] In some embodiments of the present application, each branch annotation model has a corresponding cell annotation sub-method built therein. After obtaining the comprehensive annotation result of the target single-cell spatial transcriptome through comprehensive annotation prediction based on multiple branch annotation results, it further includes:

[0034] Compare the similarities between multiple branch annotation results and the comprehensive annotation result to obtain an annotation comparison result;

[0035] When the annotation comparison result meets a preset annotation matching condition, configure an annotation applicable label for the cell annotation sub-method corresponding to the branch annotation result; wherein, the annotation applicable label is used to indicate that the cell annotation sub-method is applicable to annotating the cell type corresponding to the target single-cell spatial transcriptome.

[0036] In some embodiments of the present application, the branch annotation model is configured with adaptation information corresponding to the target single-cell spatial transcriptome;

[0037] The obtaining of the single-cell transcriptome data, the spatial transcriptome data, and multiple branch annotation models includes:

[0038] Obtaining the single-cell transcriptome data and the spatial transcriptome data;

[0039] Based on the adaptation information, multiple branch annotation models are selected from a preset alternative set of annotation models.

[0040] In some embodiments of the present application, after comprehensively annotating and predicting based on multiple branch annotation results to obtain a comprehensive annotation result for the target single-cell spatial transcriptome, it further includes:

[0041] Scoring multiple branch annotation results based on the comprehensive annotation result to obtain annotation adaptation scores corresponding one-to-one to the branch annotation results;

[0042] Updating the adaptation information for the corresponding branch annotation model according to the annotation adaptation scores.

[0043] In some embodiments of the present application, the scoring multiple branch annotation results based on the comprehensive annotation result to obtain annotation adaptation scores corresponding one-to-one to the branch annotation results includes:

[0044] When the branch annotation result is the same as the comprehensive annotation result, the branch annotation result obtains a third score;

[0045] When the branch annotation result is different from the comprehensive annotation result, the branch annotation result obtains a fourth score; where the fourth score is less than the third score.

[0046] According to the second aspect of the embodiments of the present application, a cell annotation device based on single-cell spatial transcriptome includes:

[0047] A data acquisition unit for acquiring single-cell transcriptome data, spatial transcriptome data, and multiple branch annotation models; wherein the single-cell transcriptome data and the spatial transcriptome data are derived from the same target single-cell spatial transcriptome;

[0048] A first conversion unit for converting the single-cell transcriptome data into a single-cell annotation file according to the first format constraint condition corresponding to each branch annotation model;

[0049] A second conversion unit for converting the spatial transcriptome data into a spatial annotation file according to the second format constraint condition corresponding to each branch annotation model;

[0050] A branch annotation unit, configured to input the single-cell annotation file and the spatial annotation file into the corresponding branch annotation model for branch annotation, so as to obtain the branch annotation results output by each branch annotation model;

[0051] An annotation integration unit, configured to perform comprehensive annotation prediction based on multiple branch annotation results, so as to obtain the comprehensive annotation result of the target single-cell spatial transcriptome.

[0052] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the cell annotation method based on single-cell spatial transcriptome according to any one of the embodiments in the first aspect of the present application.

[0053] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the storage medium stores a program, and when the program is executed by a processor, it implements the cell annotation method based on single-cell spatial transcriptome according to any one of the embodiments in the first aspect of the present application.

[0054] According to the cell annotation method and related device based on single-cell spatial transcriptome in the embodiments of the present application, it has at least the following beneficial effects:

[0055] According to the cell annotation method based on single-cell spatial transcriptome in the present application, it is necessary to first obtain single-cell transcriptome data, spatial transcriptome data, and multiple branch annotation models; among them, the single-cell transcriptome data and the spatial transcriptome data are derived from the same target single-cell spatial transcriptome; further, according to the first format constraint condition corresponding to each branch annotation model, the single-cell transcriptome data is converted into a single-cell annotation file; furthermore, according to the second format constraint condition of each branch annotation model, the spatial transcriptome data is converted into a spatial annotation file. Then, the single-cell annotation file and the spatial annotation file are input into the corresponding branch annotation model for branch annotation, so as to obtain the branch annotation results output by each branch annotation model, and finally, comprehensive annotation prediction is performed based on multiple branch annotation results to obtain the comprehensive annotation result of the target single-cell spatial transcriptome. Since the comprehensive annotation result is obtained by comprehensively summarizing and predicting numerous branch annotation results, in this way, the cell annotation method based on single-cell spatial transcriptome in the present application can obtain more accurate annotation results while ensuring good efficiency of cell annotation.

[0056] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0058] Figure 1 is a flowchart of a cell annotation method based on single-cell spatial transcriptomics provided by an embodiment of the present application;

[0059] Figure 2 is Figure 1 a flowchart of step S101 in

[0060] Figure 3 is Figure 1 a flowchart of step S102 in

[0061] Figure 4 is Figure 1 a flowchart of step S103 in

[0062] Figure 5 is Figure 1 a flowchart of step S105 in

[0063] Figure 6 is Figure 5 a flowchart of step S502 in

[0064] Figure 7 is Figure 6 a flowchart of step S602 in

[0065] Figure 8 is Figure 1 an optional flowchart after step S105 in

[0066] Figure 9 is Figure 8 a flowchart of step S801 in

[0067] Figure 10 is a schematic flowchart of a specific embodiment provided by the present application;

[0068] Figures 11(a) to 11(e) is a schematic diagram of experimental data of an embodiment of the present application;

[0069] Figure 12 is a schematic structural diagram of a cell annotation device based on single-cell spatial transcriptomics provided by an embodiment of the present application;

[0070] Figure 13 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0071] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.

[0072] In the description of the present application, the meaning of "a number of" is one or more, the meaning of "a plurality of" is two or more, and understandings such as "greater than", "less than", "exceeding", etc. do not include the recited number, and understandings such as "above", "below", "within", etc. include the recited number. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0073] In the description of the present application, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by "upper", "lower", "left", "right", "front", "rear", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present application.

[0074] In the description of this specification, the description with reference to terms such as "an embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0075] In the description of the present application, it should be noted that unless otherwise clearly defined, words such as "arrangement", "installation", "connection", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution. In addition, the identification of specific steps hereinafter does not represent a limitation on the step sequence and execution logic. The execution sequence and execution logic between each step should be understood and inferred with reference to the content described in the embodiment.

[0076] Single-Cell Spatial Transcriptomics is a genomics technology that can provide an accurate picture of the gene expression and spatial distribution of cells in cell tissues. It is a high-dimensional cell analysis technology used to provide an accurate picture of the gene expression and spatial distribution of cells in cell tissues. For single-cell spatial transcriptomics, specifically, the cells in cell tissues need to be segmented into single cells, and then each cell is labeled with a fluorescent dye to determine its position. Next, high-throughput sequencing technology is used to sequence the gene expression of each cell to obtain the gene expression profile of each cell. Finally, the gene expression profiles of each cell are combined with their position information to obtain an accurate picture of the gene expression and spatial distribution of cells in cell tissues.

[0077] High-throughput sequencing is a technology for sequencing nucleic acid molecules, which can simultaneously perform parallel sequence determination on a large number of nucleic acid molecules at one time. It should be noted that during the high-throughput sequencing process, nucleic acid molecules need to be ligated based on a solid surface, and complementary probes with fluorescent groups are ligated to the nucleic acid molecules, and then the base sequence is sequentially confirmed through fluorescence imaging.

[0078] There are various methods for annotating the cells of single-cell spatial transcriptomics, and each type of cell will have a suitable annotation method. In order to obtain more accurate annotation results, in related technologies, multiple annotation methods need to be used to annotate the cells of the same single-cell spatial transcriptomics and then compare them. However, the efficiency of such cell annotation is relatively low. Therefore, how to obtain more accurate annotation results while ensuring good efficiency of cell annotation has become an urgent problem to be solved in the industry.

[0079] This application aims to solve at least one of the technical problems existing in the prior art. For this purpose, this application proposes a cell annotation method and related device based on single-cell spatial transcriptomics, which can effectively improve the accuracy of base interpretation during the process of sequencing nucleic acid molecules.

[0080] The following will be further described with reference to the accompanying drawings.

[0081] This application aims to solve at least one of the technical problems existing in the prior art. For this purpose, this application proposes a cell annotation method and related device based on single-cell spatial transcriptomics, which can obtain more accurate annotation results while ensuring good efficiency of cell annotation.

[0082] Refer to Figure 1 , according to the cell annotation method based on single-cell spatial transcriptomics provided by the embodiments of this application, it may include, but is not limited to, the following steps S101 to step S105.

[0083] Step S101: Obtain single-cell transcriptome data, spatial transcriptome data, and multiple branch annotation models. Among them, the single-cell transcriptome data and the spatial transcriptome data are derived from the same target single-cell spatial transcriptome.

[0084] Step S102: Convert the single-cell transcriptome data into a single-cell annotation file according to the first format constraint condition corresponding to each branch annotation model.

[0085] Step S103: Convert the spatial transcriptome data into a spatial annotation file according to the second format constraint condition corresponding to each branch annotation model.

[0086] Step S104: Input the single-cell annotation file and the spatial annotation file into the corresponding branch annotation model for branch annotation to obtain the branch annotation results output by each branch annotation model.

[0087] Step S105: Perform comprehensive annotation prediction based on multiple branch annotation results to obtain the comprehensive annotation result of the target single-cell spatial transcriptome.

[0088] According to the method for cell annotation based on single-cell spatial transcriptome of the present application shown in Steps S101 to S105, it is necessary to first obtain single-cell transcriptome data, spatial transcriptome data, and multiple branch annotation models. Among them, the single-cell transcriptome data and the spatial transcriptome data are derived from the same target single-cell spatial transcriptome. Further, convert the single-cell transcriptome data into a single-cell annotation file according to the first format constraint condition corresponding to each branch annotation model. Still further, convert the spatial transcriptome data into a spatial annotation file according to the second format constraint condition corresponding to each branch annotation model. Then input the single-cell annotation file and the spatial annotation file into the corresponding branch annotation model for branch annotation to obtain the branch annotation results output by each branch annotation model. Finally, perform comprehensive annotation prediction based on multiple branch annotation results to obtain the comprehensive annotation result of the target single-cell spatial transcriptome. Since the comprehensive annotation result is obtained by comprehensively summarizing and predicting numerous branch annotation results, in this way, the method for cell annotation based on single-cell spatial transcriptome of the present application can obtain more accurate annotation results while ensuring good efficiency of cell annotation.

[0089] In some embodiments, in Step S101, obtain single-cell transcriptome data, spatial transcriptome data, and multiple branch annotation models. Among them, the single-cell transcriptome data and the spatial transcriptome data are derived from the same target single-cell spatial transcriptome.

[0090] It should be emphasized that the basic principle of single-cell spatial transcriptomics is as follows: cells in cell tissues are separated into individual cells, and then each cell is labeled with a fluorescent dye to determine its location. Next, high-throughput sequencing is used to sequence the gene expression of each cell to obtain the gene expression profile of each cell. Finally, the gene expression profile of each cell is combined with its location information to obtain an accurate image of the gene expression and spatial distribution of cells in the cell tissue.

[0091] Based on this, it should be noted that single-cell spatial transcriptomics can provide an accurate image of the gene expression and spatial distribution of cells in cell tissues, while the target single-cell spatial transcriptomics is an accurate image obtained by sequencing the cell tissue as the sequencing target, and the target single-cell spatial transcriptomics is used to describe the gene expression and spatial distribution of the sequencing target cells. It should be pointed out that single-cell transcriptome data is used to describe the gene expression of the sequencing target cells, and spatial transcriptome data is used to describe the spatial distribution of the sequencing target cells.

[0092] It should be pointed out that there are various optional types of cell annotation sub-methods, which can include but are not limited to the following:

[0093] First, use SingleR for cell annotation. SingleR is an R package for automatically annotating cell types in single-cell RNA-seq sequencing data. An R package is a collection of R language functions, example data, and pre-compiled code, including R programs, annotation documents, examples, test data, etc. The basic principle of SingleR is to identify cell types using the correlation between the gene expression profiles of known cell types and the gene expression profiles of individual cells.

[0094] Second, use RCDT for cell annotation. RCDT uses annotated scRNA-Seq data to create a cell type profile of the expected cell populations in the data, and then uses supervised learning methods to label the pixels of the spatial transcriptome with cell types. RCDT can relatively accurately detect the localization of cell types in simulated and real spatial transcriptome data. RCDT can also detect subtle transcriptomic differences, thereby mapping cell subtypes spatially.

[0095] Third, use SPOTlight for cell annotation. SPOTlight can integrate spatial transcriptomics with scRNA-seq data to infer the location of cell types and states in complex tissues. It is based on a seed non-negative matrix factorization regression, uses cell type marker genes and non-negative least squares initialization, and then deconvolves the spatial transcriptome data to capture the location. SPOTlight also has high prediction accuracy in low-depth sequencing or small-scale scRNA-seq reference datasets.

[0096] It should be emphasized that there are various methods for annotating cells in single-cell spatial transcriptomics, and each type of cell has a suitable annotation method. However, there is currently no consensus in the industry that "a specific annotation method is more suitable for annotating a specific type of cell." Therefore, relevant technologies use multiple annotation methods to annotate the cells of the same single-cell spatial transcriptome, and then further select the optimal solution based on the annotation results combined with the biological characteristics of the target cells to be sequenced. Therefore, the efficiency of cell annotation in this implementation order is relatively low.

[0097] In each branch annotation model of the embodiments of the present application, there is a type of cell annotation sub-method built in. The cell annotation sub-methods built in different branch annotation models are different. Each branch annotation model can annotate the sequencing data of the target single-cell spatial transcriptome based on different cell annotation sub-methods.

[0098] Refer to Figure 2 , in some embodiments of the present application, the branch annotation model is configured with adaptation information corresponding to the target single-cell spatial transcriptome. Step S101 may include, but is not limited to, the following steps S201 to step S202.

[0099] Step S201, obtain single-cell transcriptome data and spatial transcriptome data;

[0100] Step S202, based on the adaptation information, select multiple branch annotation models from a preset alternative set of annotation models.

[0101] For steps S201 to S202 of some embodiments, obtain single-cell transcriptome data and spatial transcriptome data; based on the adaptation information, select multiple branch annotation models from a preset alternative set of annotation models. It should be noted that since there are various methods for annotating cells in single-cell spatial transcriptomics, and each type of cell has a suitable annotation method. However, there is currently no consensus in the industry that "a specific annotation method is more suitable for annotating a specific type of cell." Based on this, the branch annotation model in the embodiments of the present application has adaptation information corresponding to the target single-cell spatial transcriptome; wherein, the adaptation information is used to reflect the adaptation degree between the branch annotation model and the target single-cell spatial transcriptome. The preset alternative set of annotation models includes multiple pre-set annotation models for alternative, and each annotation model has a type of cell annotation method built in.

[0102] In some embodiments, adaptation information for the branch annotation model can be configured based on past empirical data. For example, by simulating different reference quantity and quality data, it can be confirmed that SPOTlight has high prediction accuracy in low-depth sequencing or small-scale scRNA-seq reference datasets; the SPOTlight deconvolution of the mouse brain correctly maps the subtle neuronal cell states of the cortical layer and the specific structures of the hippocampus. On this basis, applying SPOTlight to the annotation of single-cell spatial transcriptomes of pancreatic cells helps to determine the spatial organization of relevant immune cell states. Therefore, if a branch annotation model has a method for cell annotation using SPOTlight built in, then it can be reflected in the adaptation information of the branch annotation model that there is a good degree of adaptation between the single-cell spatial transcriptome corresponding to the type of "animal pancreatic cells" and the branch annotation model.

[0103] It should be understood that since there are various methods for annotating cells in single-cell spatial transcriptomes, each type of cell will have a suitable annotation method, and the adaptation information can reflect the degree of adaptation between the branch annotation model and the target single-cell spatial transcriptome. In this way, based on the adaptation information, a branch annotation model that is more suitable for the target single-cell spatial transcriptome can be selected from a preset set of alternative annotation models.

[0104] In the embodiments of the present application shown in steps S201 to S202, single-cell transcriptome data and spatial transcriptome data are obtained; based on the adaptation information, multiple branch annotation models are selected from a preset set of alternative annotation models. Multiple branch annotation models that are more suitable for the target single-cell spatial transcriptome can be obtained based on the adaptation information, which helps to obtain more accurate annotation results.

[0105] In steps S102 to S103 of some embodiments, according to the first format constraint condition corresponding to each branch annotation model, the single-cell transcriptome data is converted into a single-cell annotation file; according to the second format constraint condition of each branch annotation model, the spatial transcriptome data is converted into a spatial annotation file.

[0106] It should be noted that since different cell annotation sub-methods are built into different branch annotation models, and different cell annotation sub-methods have different format requirements for input data. Based on this, for multiple branch annotation models, it is necessary to clarify the first format constraint condition and the second format constraint condition corresponding to each branch annotation model, and then use the first format constraint condition to convert the single-cell transcriptome data into a single-cell annotation file, and use the second format constraint condition to convert the spatial transcriptome data into a spatial annotation file. In this way, the problem of difficult format conversion when the input data formats are not unified can be solved, providing support for efficient cell annotation.

[0107] Referring to Figure 3 , in some embodiments of the present application, step S102 may include, but is not limited to, the following steps S301 to S302.

[0108] Step S301: Extract annotation information from the single-cell transcriptome data to obtain a single-cell expression matrix and cell type information;

[0109] Step S302: Generate a single-cell annotation file that meets the first format constraint condition according to the single-cell expression matrix and the cell type information.

[0110] For steps S301 to S302 of some embodiments, extract annotation information from the single-cell transcriptome data to obtain a single-cell expression matrix and cell type information; generate a single-cell annotation file that meets the first format constraint condition according to the single-cell expression matrix and the cell type information. It should be noted that the single-cell transcriptome data is used to describe the gene expression of the sequenced target cells. Specifically, by extracting annotation information from the single-cell transcriptome data, a single-cell expression matrix and cell type information can be obtained. Among them, the single-cell expression matrix is a matrix used to describe the gene expression level in a single cell. In gene expression research, usually a large number of cells are analyzed together as a whole. However, the gene expression differences between different cells may be very large, and this average analysis method cannot reveal the heterogeneity of individual cells. The generation of the single-cell expression matrix depends on single-cell transcriptome sequencing, where the mRNA of a single cell is transcribed into cDNA and then high-throughput sequencing is performed to obtain the expression level of each gene in each cell; the cell type information is descriptive information used to reflect the type characteristics of a single sequenced target cell.

[0111] It should be noted that after extracting annotation information from the single-cell transcriptome data to obtain a single-cell expression matrix and cell type information, a single-cell annotation file that meets the first format constraint condition can be generated according to the single-cell expression matrix and the cell type information. Among them, each branch annotation model corresponds to a first format constraint condition for converting the single-cell transcriptome data into a single-cell annotation file. Therefore, generating a single-cell annotation file needs to be carried out on the premise of meeting the first format constraint condition.

[0112] In some more specific embodiments, SingleR uses the SummarizedExperiment method to construct a single-cell annotation file, RCDT uses the Reference method to construct a single-cell annotation file, and Spotlight uses the SingleCellExperiment method to construct a single-cell annotation file. Each of the above cell annotation sub-methods has different requirements for the data structure of the input data, so the first format constraint conditions are different. Therefore, in order to perform branch annotation using different branch annotation models, it is necessary to first transform the single-cell expression matrix and cell type information based on the first format constraint conditions.

[0113] The R language is a mathematical programming language designed for mathematical researchers, mainly used for statistical analysis, graphing, and data mining. Machine learning is a multi-disciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance.

[0114] The R language provides two storage methods, one is the.Rds file and the other is the Rdata file. Among them: the.Rds file is a file for saving data sets, such as the iris data; the Rdata file is similar to a project file and will store all imported data sets and processed data.

[0115] In the above embodiments, SingleR, RCDT, and Spotlight are all annotation tools based on the R language. Therefore, the single-cell annotation file can be an RDS object, that is, the single-cell annotation file is a.Rds file and can be represented as sc_references.Rds.

[0116] The embodiment of the present application shown by steps S301 to S302 provides a specific embodiment of converting single-cell transcriptome data into a single-cell annotation file based on the first constraint condition, which solves the problems of inconsistent input data format that needs to be converted and difficult extraction, and facilitates the efficient cell annotation of the embodiment of the present application.

[0117] Refer to Figure 4 , in some embodiments of the present application, step S103 may include, but is not limited to, the following steps S401 to S402.

[0118] Step S401, extract annotation information from the spatial transcriptome data to obtain a spatial group expression matrix and spatial single-cell coordinate information;

[0119] Step S402: Generate a spatial annotation file that meets the second format constraint conditions based on the spatial group expression matrix and the spatial single-cell coordinate information.

[0120] In steps S401 to S402 of some embodiments, the annotation information of the spatial transcriptome data is extracted to obtain the spatial group expression matrix and the spatial single-cell coordinate information. Based on the spatial group expression matrix and the spatial single-cell coordinate information, a spatial annotation file that meets the second format constraint conditions is generated. It should be noted that the spatial transcriptome data is used to describe the spatial distribution of the sequencing target cells. Specifically, the spatial group expression matrix is a matrix used to describe the expression levels of genes in the spatial transcriptome; the spatial single-cell coordinate information is used to reflect the coordinate distribution of the sequencing target cells in space.

[0121] It should be noted that after extracting the annotation information of the spatial transcriptome data to obtain the spatial expression matrix and the spatial single-cell coordinate information, a spatial annotation file that meets the second format constraint conditions can be generated based on the spatial expression matrix and the spatial single-cell coordinate information. Among them, each branch annotation model corresponds to a method for converting the spatial transcriptome data into the second format constraint conditions of the spatial annotation file. Therefore, generating the spatial annotation file needs to be carried out on the premise of meeting the second format constraint conditions.

[0122] In some more specific embodiments, SingleR, RCDT, and Spotlight are all annotation tools based on R language. Therefore, based on the spatial expression matrix and the spatial single-cell coordinate information, seurat can be used to construct the spatial annotation file. It should be noted that seurat is a single-cell data analysis integration software package, and its functions not only include basic data analysis processes, such as quality control, cell screening, cell type identification, characteristic gene selection, differential expression analysis, data visualization, etc. It also includes some advanced functions, such as time-series single-cell data analysis, integration analysis of different omics single-cell data, etc.

[0123] The embodiments shown in steps S401 to S402 of the present application provide a specific embodiment of converting the spatial transcriptome data into a spatial annotation file based on the second constraint condition, solving the problems of inconsistent input data formats that need to be converted and difficult extraction, and facilitating the efficient cell annotation of the embodiments of the present application.

[0124] In step S104 of some embodiments, the single-cell annotation file and the spatial annotation file are input into the corresponding branch annotation model for branch annotation to obtain the branch annotation results output by each branch annotation model. It should be noted that after obtaining the corresponding single-cell annotation file and spatial annotation file for each branch model, the single-cell annotation file and the spatial annotation file can be further input into the corresponding branch annotation model, and the built-in cell annotation sub-method in the branch annotation model is used for branch annotation, so as to obtain the branch annotation results output by each branch annotation model. It should be understood that since each branch annotation result is obtained by branch annotation based on different cell annotation sub-methods, and different cell annotation sub-methods have different adaptabilities to specific types of cells. Therefore, the outputs of multiple annotation models can annotate the target single-cell spatial transcriptome from different annotation dimensions respectively to obtain the branch annotation results.

[0125] In step S105 of some embodiments, based on multiple branch annotation results, comprehensive annotation prediction is performed to obtain the comprehensive annotation result of the target single-cell spatial transcriptome. It should be noted that after obtaining the branch annotation results output by multiple annotation models, comprehensive annotation prediction can be performed based on the multiple branch annotation results to determine a more reliable and accurate comprehensive annotation result.

[0126] In some more specific embodiments, the comprehensive annotation results can be stored in various files, including but not limited to the following:

[0127] First, anno.Rds: It represents the RDS object of the spatial annotation file after annotation, stores the expression matrix of spatial single cells and the cell annotation results, and can be directly used for downstream analysis subsequently;

[0128] Second, anno.csv: It represents a csv file of the spatial single-cell coordinate information and the comprehensive annotation results, which is convenient for subsequent reading and plotting;

[0129] Third, celltype_maintype.pdf: It represents the global mapping diagram of the comprehensive annotation results, visually showing the spatial distribution of the annotation results on the sequencing chip;

[0130] Fourth, celltype_maintype_split.pdf: It represents the mapping diagram for each cell type in the comprehensive annotation results, visually showing the spatial distribution of each type of cell on the sequencing chip.

[0131] By obtaining the comprehensive annotation results through the above several methods, the branch annotation results of each branch annotation model can be integrated with a unified format, which is convenient for users to read and use. In addition, the output result file types and formats of each method and prediction method are the same, which is also convenient for reading and comparison, thus helping to obtain more accurate annotation results while ensuring good efficiency of cell annotation.

[0132] Referring to Figure 5 , in some embodiments of the present application, step S105 may include, but is not limited to, the following steps S501 to S502.

[0133] Step S501, based on the branch annotation results, determine the type prediction results obtained by each branch annotation model for predicting the target single-cell spatial transcriptome.

[0134] Step S502, determine the target cell type according to multiple type prediction results, and generate a comprehensive annotation result based on the target cell type.

[0135] For steps S501 to S502 of some embodiments, based on the branch annotation results, determine the type prediction results obtained by each branch annotation model for predicting the target single-cell spatial transcriptome, determine the target cell type according to multiple type prediction results, and generate a comprehensive annotation result based on the target cell type. It should be noted that the branch annotation model performs cell annotation on the target single-cell spatial transcriptome to obtain branch annotation results. Among them, the branch annotation results may include type prediction results for predicting the corresponding target cell type of the target single-cell spatial transcriptome; the branch annotation results may also include a global mapping diagram of the annotation of the target single-cell spatial transcriptome to visually display the distribution of the annotation results; the branch annotation results may further include coordinate information corresponding to specific annotation results for convenient subsequent reading and plotting. It should be understood that the information contained in the branch annotation results is diverse and is not limited to the above examples.

[0136] In the embodiments of the present application, in order to perform comprehensive annotation prediction based on multiple branch annotation results, it is necessary to determine the type prediction results obtained by each branch annotation model for predicting the target single-cell spatial transcriptome based on the branch annotation results. It should be pointed out that each type prediction result can reflect a predicted cell type, and the predicted cell type is used to reflect the cell type predicted by the corresponding branch annotation model for the target single-cell spatial transcriptome. Since each branch annotation result is obtained by performing branch annotation based on different cell annotation sub-methods, and different cell annotation sub-methods have different adaptabilities to specific types of cells, the predicted cell types reflected by different type prediction results may be the same or different. Therefore, in the embodiments of the present application, the target cell type is determined based on the predicted cell types reflected by multiple type prediction results, and then a comprehensive annotation result is generated.

[0137] Referring to Figure 6 , in some embodiments of the present application, the number of type prediction results is the first number. Step S502 may include, but is not limited to, the following steps S601 to S604.

[0138] Step S601: Group the first number of type prediction results to obtain the second number of prediction result groups; wherein each prediction result group includes at least one type prediction result, and each prediction result group corresponds to a predicted cell type;

[0139] Step S602: Determine the target group from the multiple prediction result groups based on the number of type prediction results included in the prediction result group;

[0140] Step S603: Determine the predicted cell type corresponding to the target group as the target cell type;

[0141] Step S604: Generate a comprehensive annotation result based on the target cell type.

[0142] In step S601 of some embodiments, the first number of type prediction results are grouped to obtain the second number of prediction result groups; wherein each prediction result group includes at least one type prediction result, and each prediction result group corresponds to a predicted cell type. It should be noted that the type prediction results and the number of branch annotation models are in one-to-one correspondence. For one branch annotation model, after inputting the single-cell annotation file and the spatial annotation file into it for branch annotation, a branch annotation result can be obtained, and then a type prediction result for reflecting the predicted cell type can be determined from the branch annotation result. Therefore, the first number of branch annotation models can correspondingly determine the first number of type prediction results. Based on this, the first number of type prediction results are grouped, and the type prediction results of the same predicted cell type will be grouped into the same group, so as to obtain the second number of prediction result groups. Wherein, each prediction result group corresponds to a predicted cell type, and each prediction result group includes at least one type prediction result.

[0143] In steps S602 to S604 of some embodiments, based on the number of type prediction results included in the prediction result groups, determine a target group from multiple prediction result groups; determine the predicted cell type corresponding to the target group as the target cell type; generate a comprehensive annotation result based on the target cell type. It should be noted that after obtaining the second number of prediction result groups, in order to obtain a more accurate and reliable comprehensive annotation result, a target group can be determined from multiple prediction result groups based on the number of type prediction results included in the prediction result groups. The reason is that since the various type prediction results of the same prediction result group reflect the same predicted cell type, the more the number of type prediction results in the same prediction result group, the more it means that the number of branch annotation models that predict the cell type corresponding to the target single-cell spatial transcriptome as this predicted cell type is more. Therefore, the predicted cell type corresponding to this prediction result group is considered to be more reliable as the accurate cell type of the target single-cell spatial transcriptome. It should be pointed out that after determining the target group from multiple prediction result groups based on the number of type prediction results included in the prediction result groups, the predicted cell type corresponding to the target group can be determined as the target cell type, and a comprehensive annotation result can be generated based on the target cell type. It should be pointed out that the comprehensive annotation result can include, but is not limited to, the target cell type. Therefore, the target cell type can be regarded as a component of the comprehensive annotation result.

[0144] In some embodiments of the present application, step S602 of determining a target group from multiple prediction result groups based on the number of type prediction results included in the prediction result groups can include, but is not limited to: comparing the numbers of type prediction results included in the multiple prediction result groups, and determining the prediction result group with the largest number of type prediction results among them as the target group. It should be emphasized that since the various type prediction results of the same prediction result group reflect the same predicted cell type, the more the number of type prediction results in the same prediction result group, the more it means that the number of branch annotation models that predict the cell type corresponding to the target single-cell spatial transcriptome as this predicted cell type is more. Therefore, the predicted cell type corresponding to this prediction result group is considered to be more reliable as the accurate cell type of the target single-cell spatial transcriptome. On this basis, the numbers of type prediction results included in the multiple prediction result groups can be compared, and the prediction result group with the largest number of type prediction results among them can be determined as the target group. In this way, the predicted cell type corresponding to the target group can be considered as the most accurate and reliable cell type among the predicted cell types corresponding to the multiple prediction result groups. Based on this, the predicted cell type of the target group can be determined as the target cell type.

[0145] Refer to Figure 7, in some embodiments of the present application, step S602 may determine a target group from multiple prediction result groups based on the number of type prediction results included in the prediction result group, and may include, but are not limited to, the following steps S701 to S704.

[0146] Step S701, score each prediction result group based on the number of type prediction results included in the prediction result group;

[0147] Step S702, when the number of type prediction results included in the prediction result group is more than two, the prediction result group gets a first score;

[0148] Step S703, when the number of type prediction results included in the prediction result group is less than two, the prediction result group gets a second score; wherein, the second score is less than the first score;

[0149] Step S704, determine the target group from the prediction result groups that get the first score.

[0150] In step S701 of some embodiments, each prediction result group is scored based on the number of type prediction results included in the prediction result group. It should be noted that since the various type prediction results of the same prediction result group reflect the same predicted cell type, the more the number of type prediction results in the same prediction result group, the more it means that the number of branch annotation models that predict the cell type corresponding to the target single-cell spatial transcriptome as this predicted cell type is more. Therefore, each prediction result group can be scored based on the number of type prediction results included in the prediction result group, where the higher the score, the more reliable and accurate the predicted cell type corresponding to the prediction result group is;

[0151] In step S702 of some embodiments, when the number of type prediction results included in the prediction result group is more than two, the prediction result group gets a first score. It should be noted that when the number of type prediction results included in the prediction result group is more than two, it means that more than two branch annotation models have obtained the same type prediction result. At this time, the first score is given to this prediction result group, and the first score reflects the accuracy of the predicted cell type corresponding to this prediction result group.

[0152] In step S703 of some embodiments, when the number of type prediction results included in the prediction result group is less than two, the prediction result group gets a second score; wherein, the second score is less than the first score. It should be noted that when the number of type prediction results included in the prediction result group is less than two, it means that less than two branch annotation models have obtained the same type prediction result. At this time, the second score is given to this prediction result group, and the second score reflects the accuracy of the predicted cell type corresponding to this prediction result group.

[0153] It should be understood that the higher the score, the more reliable and accurate the predicted cell type corresponding to the predicted result group. Therefore, a predicted result group containing more than two types of predicted results needs to have a higher score than a predicted result group containing less than two types of predicted results. Thus, the second score is less than the first score.

[0154] In step S704 of some embodiments, the target group is determined from the predicted result group that obtains the first score. It should be noted that after scoring each predicted result group, it is possible to select the predicted result group that obtains the first score from among the numerous predicted result groups, and then determine the predicted result group that obtains the first score as the target group. Subsequently, the predicted cell type corresponding to the target group is determined as the most reliable and accurate target cell type.

[0155] In the embodiments of the present application illustrated by steps S701 to S704, if more than two branch annotation models yield the same type of predicted result, the predicted cell type corresponding to this type of predicted result can be determined as the target cell type, facilitating the generation of a more accurate and reliable comprehensive annotation result, and contributing to obtaining a more accurate annotation result while ensuring good efficiency in cell annotation.

[0156] In some more specific embodiments of the present application, scores are given based on the number of type prediction results included in the predicted result group. The more type prediction results are included in the predicted result group, the higher the score obtained by the predicted result group, which means that the predicted cell type corresponding to the predicted result group is more reliable and accurate. Therefore, if there are multiple predicted result groups that obtain the first score, the predicted result group with a higher score can be selected from them, determined as the target group, and then the predicted cell type corresponding to the target group is determined as the most reliable and accurate target cell type.

[0157] It should be understood that there are various embodiments for determining the target group from multiple predicted result groups based on the number of type prediction results included in the predicted result group, which may include, but are not limited to, the specific embodiments cited above.

[0158] In the embodiments of the present application shown in steps S601 to S604, group the first number of type prediction results to obtain the second number of prediction result groups; where each prediction result group includes at least one type prediction result, and each prediction result group corresponds to a predicted cell type; further, based on the number of type prediction results included in the prediction result group, determine the target group from multiple prediction result groups; then determine the predicted cell type corresponding to the target group as the target cell type; finally, generate a comprehensive annotation result based on the target cell type. In this way, the target group can be determined based on the number of type prediction results included in the prediction result group, so that the predicted cell type corresponding to the target group is determined as the target cell type, and a comprehensive annotation result is generated based on the target cell type, which helps to obtain a more accurate annotation result while ensuring good efficiency of cell annotation.

[0159] In the embodiments of the present application shown in steps S501 to S502, based on the branch annotation results, determine the type prediction results obtained by each branch annotation model predicting the target single-cell spatial transcriptome, determine the target cell type according to multiple type prediction results, and generate a comprehensive annotation result based on the target cell type. Generating a comprehensive annotation result based on the target cell type reflected by the type prediction results helps to determine a more reliable and accurate comprehensive annotation result.

[0160] In some embodiments of the present application, each branch annotation model has a corresponding cell annotation sub-method built in. After step S105, the cell annotation method of the embodiments of the present application based on the single-cell spatial transcriptome may further include, but is not limited to, the following steps:

[0161] Perform a similarity comparison between multiple branch annotation results and the comprehensive annotation result to obtain an annotation comparison result;

[0162] When the annotation comparison result meets the preset annotation matching condition, configure an annotation applicable label for the cell annotation sub-method corresponding to the branch annotation result; where the annotation applicable label is used to identify that the cell annotation sub-method is applicable to annotating the cell type corresponding to the target single-cell spatial transcriptome.

[0163] It should be noted that each branch annotation model is built-in with a corresponding cell annotation sub-method. Different branch annotation models respectively use the cell annotation sub-methods to perform branch annotation on the cell types corresponding to the target single-cell spatial transcriptome, and obtain the corresponding branch annotation results. Then, based on multiple branch annotation results, comprehensive annotation prediction is performed to obtain the comprehensive annotation result of the target single-cell spatial transcriptome. Since the comprehensive annotation result is a relatively accurate and reliable annotation result for the target single-cell spatial transcriptome, multiple branch annotation results are compared with the comprehensive annotation result in terms of similarity, aiming to determine which of the multiple branch annotation results are relatively accurate and reliable. It should be pointed out that the annotation comparison result reflects the similarity degree between the branch annotation result and the comprehensive annotation result. When the annotation comparison result of a certain branch annotation result meets the preset annotation matching condition, it means that the branch annotation result is relatively accurate and reliable. On this basis, an annotation applicable label can be configured for the cell annotation sub-method corresponding to the branch annotation result. It should be pointed out that the annotation applicable label is used to identify that the cell annotation sub-method is applicable to annotating the cell types corresponding to the target single-cell spatial transcriptome.

[0164] In some more specific embodiments, if a cell annotation sub-method is configured with an annotation applicable label, it means that this cell annotation sub-method is applicable to annotating the cell types corresponding to the target single-cell spatial transcriptome. At this time, directly using this cell annotation sub-method to annotate the cell types corresponding to the target single-cell spatial transcriptome can obtain a more accurate annotation result while ensuring good efficiency of cell annotation.

[0165] Refer to Figure 8 , in some embodiments of the present application, after step S105, the cell annotation method based on single-cell spatial transcriptome of the embodiments of the present application may further include, but is not limited to, the following steps S801 to step S802.

[0166] Step S801, score multiple branch annotation results based on the comprehensive annotation result to obtain annotation adaptation scores corresponding one by one to the branch annotation results;

[0167] Step S802, update the adaptation information for the corresponding branch annotation model according to the annotation adaptation score.

[0168] It should be noted that in some embodiments, multiple branch annotation models need to be selected from a preset alternative set of annotation models based on the adaptation information. In the embodiments of the present application, with each cell annotation, the configuration information can be updated so that the adaptation information can more accurately reflect the adaptation degree between the branch annotation model and the target single-cell spatial transcriptome.

[0169] In steps S801 to S802 of some embodiments, the multiple branch annotation results are scored based on the comprehensive annotation result to obtain annotation adaptation scores corresponding one-to-one to the branch annotation results; and the adaptation information is updated for the corresponding branch annotation models according to the annotation adaptation scores. It should be noted that since the comprehensive annotation result is a relatively reliable and accurate annotation result for the target single-cell spatial transcriptome, the multiple branch annotation models can be scored based on the comprehensive annotation result to determine the annotation adaptation scores corresponding one-to-one to the branch annotation results. It should be clear that the annotation adaptation scores are used to update the adaptation information for the corresponding branch annotation models so that the adaptation degree between the branch annotation models and the target single-cell spatial transcriptome can be reflected more accurately. In this way, as the cell annotation method based on the single-cell spatial transcriptome in the embodiments of the present application performs cell annotation on various target single-cell spatial transcriptomes, the adaptation information of each branch annotation model will be updated, so as to obtain more accurate annotation results while ensuring good efficiency of cell annotation.

[0170] Referring to Figure 9 , in some embodiments of the present application, step S801 of scoring the multiple branch annotation results based on the comprehensive annotation result to obtain the annotation adaptation scores corresponding one-to-one to the branch annotation results may include, but is not limited to, the following steps S901 to S902.

[0171] In step S901, when the branch annotation result is the same as the comprehensive annotation result, the branch annotation result obtains a third score;

[0172] In step S902, when the branch annotation result is different from the comprehensive annotation result, the branch annotation result obtains a fourth score; wherein, the fourth score is less than the third score.

[0173] For steps S901 to S902 of some embodiments, when the branch annotation result is the same as the comprehensive annotation result, the branch annotation result obtains a third score; when the branch annotation result is different from the comprehensive annotation result, the branch annotation result obtains a fourth score; wherein, the fourth score is less than the third score. It should be noted that when the branch annotation result is the same as the comprehensive annotation result, it means that the branch annotation model can make a relatively reliable and accurate cell annotation for the target single-cell spatial transcriptome. At this time, the branch annotation result obtains a third score; when the branch annotation result is different from the comprehensive annotation result, it means that the cell annotation made by the branch annotation model for the target single-cell spatial transcriptome has poor reliability and low accuracy. At this time, the branch annotation result obtains a fourth score. It should be pointed out that the fourth score is less than the third score, which means that the branch annotation model that annotates different results from the comprehensive annotation result has poorer reliability and accuracy than the branch annotation model that annotates the same result as the comprehensive annotation result. On this basis, according to the annotation adaptation score, update the adaptation information for the corresponding branch annotation model, and the adaptation information of each branch annotation model will be updated, so as to obtain a more accurate annotation result while ensuring good efficiency of cell annotation.

[0174] Referring to Figure 10 In a relatively specific embodiment shown, first obtain and input single-cell transcriptome data and spatial transcriptome data. Among them, the single-cell transcriptome data includes a single-cell expression matrix and cell type information, and the spatial transcriptome data includes a spatial group expression matrix and spatial single-cell coordinate information.

[0175] It should be pointed out that there are various optional types of cell annotation sub-methods built in the branch annotation model, including but not limited to the following:

[0176] SingleR is built in the branch annotation model for cell annotation; among them, SingleR uses the SummarizedExperiment method to construct a single-cell annotation file.

[0177] RCDT is built in the branch annotation model for cell annotation; among them, RCDT uses the Reference method to construct a single-cell annotation file.

[0178] SPOTlight is built in the branch annotation model for cell annotation; among them, Spotlight uses the SingleCellExperiment method to construct a single-cell annotation file.

[0179] Therefore, the single-cell expression matrix and cell type information can be input into each branch annotation model. In the above embodiments, SingleR, RCDT, and Spotlight are all R-language-based annotation tools. Therefore, the single-cell annotation file can be an RDS object, that is, the single-cell annotation file is an.Rds file, which can be represented as sc_references.Rds.

[0180] It should be emphasized that in the above embodiments, SingleR, RCDT, and Spotlight are all R-language-based annotation tools. Therefore, an RDS object of the corresponding spatial annotation file needs to be generated according to the spatial transcriptome expression matrix and spatial single-cell coordinate information, which is represented as the spatial annotation file.Rds.

[0181] Furthermore, the single-cell annotation file sc_references.Rds and the spatial annotation file.Rds are input into the corresponding branch annotation model for branch annotation to obtain the branch annotation results. Among them, based on multiple branch annotation results for comprehensive annotation prediction, the comprehensive annotation result of the target single-cell spatial transcriptome can be obtained.

[0182] Furthermore, after obtaining the branch annotation results corresponding to each branch annotation model and the comprehensive annotation result, since the comprehensive annotation result is a relatively reliable and accurate annotation result, the branch annotation results corresponding to each branch annotation model can be compared with the comprehensive annotation result for similarity in turn. Among them, the more similar the branch annotation result is to the comprehensive annotation result, the higher the reliability and accuracy of the branch annotation result; the less similar the branch annotation result is to the comprehensive annotation result, the lower the reliability and accuracy of the branch annotation result. In this way, the configuration information of each branch annotation model can be updated based on the results of the similarity comparison.

[0183] In addition, the comprehensive annotation result can be stored in the following formats:

[0184] anno.Rds: Represents the RDS object of the spatial annotation file after annotation, stores the expression matrix of spatial single cells and the cell annotation results, and can be directly used for downstream analysis subsequently;

[0185] anno.csv: Represents a csv file of spatial single-cell coordinate information and comprehensive annotation results, which is convenient for subsequent reading and plotting;

[0186] celltype_maintype.pdf: Represents the global mapping diagram of the comprehensive annotation result, intuitively showing the spatial distribution of the annotation result on the sequencing chip;

[0187] celltype_maintype_split.pdf: It represents the mapping diagram for each cell type in the comprehensive annotation results, visually showing the spatial distribution of cells of each type on the sequencing chip.

[0188] In this way, the cell annotation method based on single-cell spatial transcriptomics of the present application can obtain more accurate annotation results while ensuring good efficiency of cell annotation.

[0189] In some more specific embodiments of the present application, a wing primordium sample was collected, and the corresponding target single-cell spatial transcriptome was obtained based on the sequencing of the wing primordium sample.

[0190] The annotation results of cell annotation of the target single-cell spatial transcriptome by SingleR are shown in Figure 11(a);

[0191] The annotation results of cell annotation of the target single-cell spatial transcriptome by RCDT are shown in Figure 11(b);

[0192] The annotation results of cell annotation of the target single-cell spatial transcriptome by SPOTlight are shown in Figure 11(c);

[0193] It should be noted that the cell annotation method based on single-cell spatial transcriptomics in the embodiments of the present application involves three types of branch annotation models; namely, the branch annotation model with SingleR built-in for cell annotation, the branch annotation model with RCDT built-in for cell annotation, and the branch annotation model with SPOTlight built-in for cell annotation.

[0194] It can be clearly seen that the similarity between Figure 11(a) and Figure 11(b) is relatively high, that is, the annotation results of cell annotation of the target single-cell spatial transcriptome by SingleR are relatively close to the annotation results of cell annotation of the target single-cell spatial transcriptome by RCDT. Then, when using the cell annotation method based on single-cell spatial transcriptomics of the present application to perform cell annotation on the target single-cell spatial transcriptome obtained by sequencing the wing primordium sample, the comprehensive annotation results will mainly be determined based on the branch annotation results obtained by the branch annotation model with SingleR built-in for cell annotation and the branch annotation results obtained by the branch annotation model with RCDT built-in for cell annotation.

[0195] Based on this, by using the cell annotation method based on single-cell spatial transcriptomics provided in the embodiments of the present application to perform cell annotation on the target single-cell spatial transcriptome obtained by sequencing the wing primordium sample, the comprehensive annotation results are shown in Figure 11(d).

[0196] Figure 11(e) shows the percentage of each cell type in the annotation results of each branch annotation model, as well as the percentage of each cell type in the annotation result obtained by the cell annotation method of the embodiment of the present application.

[0197] It can be clearly seen that the cell annotation method based on single-cell spatial transcriptomics provided by the embodiment of the present application is relatively close to the branch annotation results obtained by the branch annotation model for cell annotation using the built-in SingleR; and the cell annotation method based on single-cell spatial transcriptomics provided by the embodiment of the present application is also relatively close to the branch annotation results obtained by the branch annotation model for cell annotation using the built-in RCDT. Specifically:

[0198] The similarity score between the annotation result obtained by the cell annotation method of the embodiment of the present application and the branch annotation result obtained by the branch annotation model for cell annotation using the built-in SPOTlight: 955 / 15785 = 0.6153;

[0199] The similarity score between the annotation result obtained by the cell annotation method of the embodiment of the present application and the branch annotation result obtained by the branch annotation model for cell annotation using the built-in SingleR: 15369 / 15785 = 0.9736;

[0200] The similarity score between the annotation result obtained by the cell annotation method of the embodiment of the present application and the branch annotation result obtained by the branch annotation model for cell annotation using the built-in RCDT: 15443 / 15785 = 0.9783.

[0201] Among them, the similarity score of the branch annotation result obtained by the branch annotation model for cell annotation using the built-in RCDT is the highest.

[0202] Through Figures 11(a) to 11(e) The experimental data reflected by the embodiments in can be clearly seen that the cell annotation method based on single-cell spatial transcriptomics of the present application can obtain more accurate annotation results while ensuring good efficiency of cell annotation.

[0203] Referring to Figure 12 , according to some embodiments, the present application provides a cell annotation device 1200 based on single-cell spatial transcriptomics, including:

[0204] A data acquisition unit 1201, configured to acquire single-cell transcriptome data, spatial transcriptome data, and multiple branch annotation models; wherein, the single-cell transcriptome data and the spatial transcriptome data are derived from the same target single-cell spatial transcriptome;

[0205] A first conversion unit 1202, configured to convert the single-cell transcriptome data into a single-cell annotation file according to the first format constraint condition corresponding to each branch annotation model;

[0206] A second conversion unit 1203, configured to convert the spatial transcriptome data into a spatial annotation file according to the second format constraint conditions of each branch annotation model;

[0207] A branch annotation unit 1204, configured to input the single-cell annotation file and the spatial annotation file into the corresponding branch annotation model for branch annotation, so as to obtain the branch annotation results output by each branch annotation model;

[0208] An annotation integration unit 1205, configured to perform comprehensive annotation prediction based on multiple branch annotation results, so as to obtain the comprehensive annotation result of the target single-cell spatial transcriptome.

[0209] It can be seen that the content in the above embodiments of the cell annotation method based on single-cell spatial transcriptome is applicable to the embodiments of the cell annotation device based on single-cell spatial transcriptome. The functions specifically implemented by the embodiments of the cell annotation device based on single-cell spatial transcriptome are the same as those in the above embodiments of the cell annotation method based on single-cell spatial transcriptome, and the beneficial effects achieved are also the same as those in the above embodiments of the cell annotation method based on single-cell spatial transcriptome.

[0210] Referring to Figure 13 , Figure 13 FIG. schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0211] A processor 1301, which can be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0212] A memory 1302, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1302 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of the present specification through software or firmware, the relevant program codes are stored in the memory 1302 and are called by the processor 1301 to execute the cell annotation method based on single-cell spatial transcriptome of the embodiments of the present application;

[0213] An input / output interface 1303, configured to implement information input and output;

[0214] A communication interface 1304, which is used to implement communication and interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0215] A bus 1305, which transmits information between various components of the device (such as a processor 1301, a memory 1302, an input / output interface 1303, and a communication interface 1304);

[0216] Among them, the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304 achieve communication connections with each other inside the device through the bus 1305.

[0217] An embodiment of this application also provides a computer program product, which includes a computer program. The processor of the computer device reads and executes this computer program, so that the computer device executes the cell annotation method based on single-cell spatial transcriptomics described above.

[0218] It should be understood that in this disclosure, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this disclosure described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "comprise" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0219] It should be understood that in this disclosure, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0220] It should be understood that in the description of the embodiments of the present application, the meaning of "a plurality of (or multiple)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.

[0221] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0222] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0223] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0224] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM for short), random access memory (RAM for short), magnetic disks, or optical discs that can store program codes.

[0225] It should also be understood that the various implementation manners provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0226] The above is a specific description of the embodiments of the present disclosure. However, the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.

Claims

1. A method for cell annotation based on single-cell spatial transcriptomics, characterized in that, it includes: Obtaining single-cell transcriptome data, spatial transcriptome data, and multiple branch annotation models; wherein, the single-cell transcriptome data and the spatial transcriptome data are derived from the same target single-cell spatial transcriptome; Converting the single-cell transcriptome data into a single-cell annotation file according to the first format constraint condition corresponding to each branch annotation model; Converting the spatial transcriptome data into a spatial annotation file according to the second format constraint condition corresponding to each branch annotation model; Inputting the single-cell annotation file and the spatial annotation file into the corresponding branch annotation model for branch annotation to obtain the branch annotation results output by each branch annotation model; Performing comprehensive annotation prediction based on multiple branch annotation results to obtain a comprehensive annotation result for the target single-cell spatial transcriptome.

2. The method according to claim 1, characterized in that, The step of converting the single-cell transcriptome data into a single-cell annotation file according to the first format constraint condition corresponding to each branch annotation model includes: Extracting annotation information from the single-cell transcriptome data to obtain a single-cell expression matrix and cell type information; Generating the single-cell annotation file that meets the first format constraint condition according to the single-cell expression matrix and the cell type information.

3. The method according to claim 1, characterized in that, The step of converting the spatial transcriptome data into a spatial annotation file according to the second format constraint condition of each branch annotation model includes: Extracting annotation information from the spatial transcriptome data to obtain a spatial group expression matrix and spatial single-cell coordinate information; Generating the spatial annotation file that meets the second format constraint condition according to the spatial group expression matrix and the spatial single-cell coordinate information.

4. The method according to claim 1, characterized in that, The step of performing comprehensive annotation prediction based on multiple branch annotation results to obtain a comprehensive annotation result for the target single-cell spatial transcriptome includes: Based on the branch annotation results, determining the type prediction results obtained by each branch annotation model for predicting the target single-cell spatial transcriptome; Determining the target cell type according to multiple type prediction results, and generating the comprehensive annotation result based on the target cell type.

5. The method according to claim 4, characterized in that, The number of the type prediction results is the first number; The step of determining the target cell type according to multiple type prediction results and generating the comprehensive annotation result based on the target cell type includes: Grouping the first number of type prediction results to obtain the second number of prediction result groups; wherein, each prediction result group includes at least one type prediction result, and each prediction result group corresponds to a predicted cell type; Determining the target group from multiple prediction result groups based on the number of type prediction results included in the prediction result group; Determine the predicted cell type corresponding to the target group as the target cell type; Generate the comprehensive annotation result based on the target cell type.

6. The method according to claim 5, wherein, The determining the target group from multiple prediction result groups based on the number of the type prediction results included in the prediction result group includes: Scoring each prediction result group based on the number of the type prediction results included in the prediction result group; When the number of the type prediction results included in the prediction result group is more than two, the prediction result group gets a first score; When the number of the type prediction results included in the prediction result group is less than two, the prediction result group gets a second score; wherein, the second score is less than the first score; Determine the target group from the prediction result groups that get the first score.

7. The method according to claim 5, wherein, The determining the target group from multiple prediction result groups based on the number of the type prediction results included in the prediction result group includes: Compare the numbers of the type prediction results included in multiple prediction result groups, and determine the prediction result group with the largest number of the type prediction results as the target group.

8. The method according to claim 1, wherein, Each of the branch annotation models is built-in with a corresponding cell annotation sub-method. After comprehensively annotating and predicting based on multiple branch annotation results to obtain the comprehensive annotation result of the target single-cell spatial transcriptome, it further includes: Perform a similarity comparison between multiple branch annotation results and the comprehensive annotation result to obtain an annotation comparison result; When the annotation comparison result meets the preset annotation matching condition, configure an annotation applicable label for the cell annotation sub-method corresponding to the branch annotation result; wherein, the annotation applicable label is used to identify that the cell annotation sub-method is applicable to annotating the cell type corresponding to the target single-cell spatial transcriptome.

9. The method according to claim 1, wherein, The branch annotation model is configured with adaptation information corresponding to the target single-cell spatial transcriptome; The obtaining the single-cell transcriptome data, spatial transcriptome data, and multiple branch annotation models includes: Obtain the single-cell transcriptome data and the spatial transcriptome data; Based on the adaptation information, select multiple branch annotation models from a preset alternative set of annotation models.

10. The method according to claim 9, wherein, After comprehensively annotating and predicting based on multiple branch annotation results to obtain the comprehensive annotation result of the target single-cell spatial transcriptome, it further includes: Score multiple branch annotation results based on the comprehensive annotation result to obtain annotation adaptation scores corresponding one-to-one to the branch annotation results; Update the adaptation information for the corresponding branch annotation model according to the annotation adaptation scores.

11. The method according to claim 10, wherein, Scoring the multiple branch annotation results based on the comprehensive annotation result to obtain annotation adaptation scores corresponding one by one to the branch annotation results, including: When the branch annotation result is the same as the comprehensive annotation result, the branch annotation result obtains a third score; When the branch annotation result is different from the comprehensive annotation result, the branch annotation result obtains a fourth score; wherein, the fourth score is less than the third score.

12. A cell annotation device based on single-cell spatial transcriptomics, characterized in that it includes: A data acquisition unit, configured to acquire single-cell transcriptome data, spatial transcriptome data, and multiple branch annotation models; wherein, the single-cell transcriptome data and the spatial transcriptome data are derived from the same target single-cell spatial transcriptome; A first conversion unit, configured to convert the single-cell transcriptome data into a single-cell annotation file according to the first format constraint condition corresponding to each branch annotation model; A second conversion unit, configured to convert the spatial transcriptome data into a spatial annotation file according to the second format constraint condition corresponding to each branch annotation model; A branch annotation unit, configured to input the single-cell annotation file and the spatial annotation file into the corresponding branch annotation model for branch annotation to obtain branch annotation results output by each branch annotation model; An annotation integration unit, configured to perform comprehensive annotation prediction based on the multiple branch annotation results to obtain a comprehensive annotation result for the target single-cell spatial transcriptome.

13. An electronic device, including a memory and a processor, the memory stores a computer program, characterized in that when the processor executes the computer program, it implements the cell annotation method based on single-cell spatial transcriptomics according to any one of claims 1 to 11.

14. A computer-readable storage medium, the storage medium stores a computer program, characterized in that when the computer program is executed by a processor, it implements the cell annotation method based on single-cell spatial transcriptomics according to any one of claims 1 to 11.

15. A computer program product, the computer program product includes a computer program, the computer program is read and executed by a processor of a computer device, so that the computer device executes the cell annotation method based on single-cell spatial transcriptomics according to any one of claims 1 to 11.

Citation Information

Cited By

  • Single-cell multi-omics data analysis system, method and equipment and storage medium

    CN120766779A