Single cell identification method based on gene wave and related device

Through the single-cell recognition method based on gene waves, the problem of accurate definition in cell type and cell subtype annotation is solved, and rapid and accurate automated annotation is achieved, and disease-specific cell subtypes are identified, supporting the analysis of disease pathological mechanisms and the discovery of therapeutic targets.

CN119920320AActive Publication Date: 2025-05-02ZHEJIANG UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411923447.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-02
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The prior art has difficulties in the automated annotation of cell types and cell subtypes, and the problems of manual annotation leading to inconsistent results and decreased comparability, making it impossible to quickly and accurately identify disease-specific cell subtypes.

Method used

Gene wave-based single-cell recognition method is used to obtain the gene expression matrix, pre-process and sort, gene waves of each single cell to be annotated are generated, and cell types and/or cell subtypes are determined using the trained single-cell recognition model.

Benefits of technology

Fast and accurate automated annotation of cell types and cell subtypes is achieved, and the accuracy of identification is improved, and disease-specific cell subtypes can be identified, thus supporting the fine analysis of disease pathological mechanisms and the discovery of potential therapeutic targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920320A_ABST
    Figure CN119920320A_ABST
Patent Text Reader

Abstract

The invention discloses a gene wave-based single cell identification method and a related device, and relates to the technical field of single cell identification, and the method comprises the following steps: preprocessing a gene expression matrix of a to-be-annotated object to obtain a preprocessed matrix, sorting genes in the preprocessed matrix based on a human reference genome, and obtaining a gene wave of each single cell to be annotated, taking the gene waves of the single cells to be annotated as input, and determining an identification result of the single cells to be annotated by using the trained single cell identification model, the identification result being a cell type and / or a cell subtype. According to the method, automatic annotation of the cell type and the cell subtype can be rapidly and accurately completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of single cell identification technology, and in particular to a single cell identification method based on gene waves and related devices. Background Art

[0002] A major limitation in the development and evaluation of current clinical disease markers, target drugs, or repair treatments is the lack of rich information about the cell composition of pathological tissues. Understanding the cell composition of pathological tissues will fill the knowledge gap that hinders the identification and implementation of effective treatments in this field. The single-cell RNA sequencing (scRNA-seq) technology developed in recent years can provide in-depth understanding of the gene expression patterns and functional states of individual cells. By classifying and annotating individual cells, different cell types can be identified. However, the low-resolution classification level of cell types may not be sufficient to meet the needs. It is necessary to further refine to the high-resolution classification level of cell subtypes to more comprehensively understand the heterogeneity of cell populations. Based on this, different cell subtypes and their gene expression characteristics in healthy and diseased states are identified, and disease-specific cell subtypes (DSCSs) are identified, which provides strong support for fine-tuning the pathological mechanisms of diseases, discovering potential therapeutic targets, and providing potential therapeutic drugs.

[0003] The accurate definition of cell types and cell subtypes is the prerequisite for accurate identification of DSCSs. At present, the definition and classification of cell types and cell subtypes are still mainly based on manual annotation. The conventional strategy is to use machine learning methods to perform unsupervised clustering of cell expression profiles, and then manually assign each cluster identity (i.e., cell type and cell subtype) based on the different expression patterns of characteristic genes of different cell types and cell subtypes in each cluster and the functions enriched by highly expressed genes in each cluster in the literature, combined with empirical judgment. In view of the lack of clear standards and consensus when assigning identities, there is subjectivity between different studies. At the same time, different studies may use different naming and classification methods for the same cell types and cell subtypes. Manual annotation of DSCSs will lead to increased inconsistency in research results and decreased comparability, which in turn limits the research progress of DSCSs. Therefore, it is extremely urgent to develop tools for automatic and accurate annotation of DSCSs.

[0004] In order to meet the challenges brought by manual annotation, multiple studies have developed informatics tools to automatically assign cell identities and have made extensive efforts in automatic cell annotation algorithms. These tools are mainly divided into two categories: methods that rely on cell type-specific genes: such as CellAssign, Garnett, and scType; methods that rely on pre-configuration data: such as singleR, scMatch, and scmap. Although these tools provide powerful annotation algorithms at the cell type level, the accuracy of annotating cell types and cell subtypes in complex annotation hierarchies needs to be improved, and it is impossible to quickly and uniformly identify specific cell subtypes, which largely hinders the exploration of key therapeutic targets for diseases and the understanding of related pathogenesis. Therefore, it is urgent to develop an automated technology that can quickly and accurately complete the annotation of cell types and cell subtypes. Summary of the invention

[0005] The purpose of this application is to provide a single-cell identification method and related devices based on gene waves, which can quickly and accurately complete the automated annotation of cell types and cell subtypes.

[0006] To achieve the above objectives, this application provides the following solutions:

[0007] In a first aspect, the present application provides a single cell identification method based on gene waves, and the single cell identification method based on gene waves comprises:

[0008] Acquire a query data set; the query data set includes a gene expression matrix, the gene expression matrix includes the expression level of each gene in each single cell to be annotated included in the object to be annotated; the object to be annotated is an organ or tissue of an organism;

[0009] The gene expression matrix is ​​preprocessed so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set, so as to obtain a preprocessed matrix; the reference data set includes a reference gene expression matrix and reference cell metadata, the reference gene expression matrix includes the expression amount of each reference gene in each reference single cell included in the annotated object, the reference cell metadata includes the label of each reference single cell included in the annotated object, the label is a cell type and / or a cell subtype, and the annotated object is of the same type as the object to be annotated;

[0010] Based on the human reference genome, the genes in the preprocessed matrix are sorted, and for each of the single cells to be annotated, the expression levels of the sorted genes of the single cell to be annotated are combined into a gene wave of the single cell to be annotated; the abscissa of the gene wave is the gene, and the ordinate of the gene wave is the expression level of the gene;

[0011] For each of the single cells to be annotated, the gene wave of the single cell to be annotated is used as input, and the identification result of the single cell to be annotated is determined using a trained single cell identification model; the trained single cell identification model is a model trained based on the reference data set.

[0012] In a second aspect, the present application provides a single cell identification device based on gene waves, the single cell identification device based on gene waves comprising:

[0013] A data acquisition module, used to acquire a query data set; the query data set includes a gene expression matrix, the gene expression matrix includes the expression level of each gene in each single cell to be annotated included in the object to be annotated; the object to be annotated is an organ or tissue of an organism;

[0014] A preprocessing module, used for preprocessing the gene expression matrix so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set, and obtaining a preprocessed matrix; the reference data set includes a reference gene expression matrix and reference cell metadata, the reference gene expression matrix includes the expression amount of each reference gene in each reference single cell included in the annotated object, the reference cell metadata includes the label of each reference single cell included in the annotated object, the label is a cell type and / or a cell subtype, and the annotated object is of the same type as the object to be annotated;

[0015] A sorting module, for sorting the genes in the preprocessed matrix based on the human reference genome, and for each of the single cells to be annotated, composing the expression levels of the sorted genes of the single cell to be annotated into a gene wave of the single cell to be annotated; the abscissa of the gene wave is the gene, and the ordinate of the gene wave is the expression level of the gene;

[0016] The recognition module is used to determine the recognition result of each single cell to be annotated by using the gene wave of the single cell to be annotated as input and using a trained single cell recognition model; the trained single cell recognition model is a model trained based on the reference data set.

[0017] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned single-cell identification method based on gene waves.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned single-cell identification method based on gene waves.

[0019] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned single-cell identification method based on gene waves.

[0020] According to the specific embodiments provided in this application, this application has the following technical effects:

[0021] The present application provides a single-cell identification method and related devices based on gene waves, preprocessing the gene expression matrix of the object to be annotated so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set to obtain a preprocessed matrix, sorting the genes in the preprocessed matrix based on the human reference genome to obtain the gene wave of each single cell to be annotated, and for each single cell to be annotated, using the gene wave of the single cell to be annotated as input, using a trained single-cell recognition model to determine the recognition result of the single cell to be annotated, the recognition result is a cell type and / or cell subtype. By introducing gene sorting, the present application can consider gene position information in the recognition process, and at the same time, the gene wave can be input into the trained single-cell recognition model to complete the recognition, so that the automatic annotation of cell types and cell subtypes can be completed quickly and accurately, and then the disease-specific cell subtypes can be identified based on the cell subtypes obtained by recognition in a healthy state and a disease state. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 This is a diagram of the application environment of a single-cell identification method based on gene waves provided in Example 1 of the present application.

[0024] Figure 2 A schematic flow chart of a single-cell identification method based on gene waves provided in Example 1 of the present application.

[0025] Figure 3 A schematic diagram of the principle of a single-cell identification method based on gene waves provided in Example 1 of the present application.

[0026] Figure 4 A schematic diagram of the application process of a single-cell identification method based on gene waves provided in Example 1 of the present application.

[0027] Figure 5 This is a schematic diagram of the gene wave provided in Example 1 of the present application.

[0028] Figure 6 A schematic diagram of the network structure of the initial single-cell recognition model provided in Example 1 of the present application; wherein: Figure 6 (a) is the initial single-cell recognition model at the cellular level; Figure 6 (b) is the initial single-cell recognition model at the cell population level.

[0029] Figure 7 A schematic diagram comparing the accuracy of different methods provided in Example 1 of the present application.

[0030] Figure 8 A schematic diagram comparing the mean, median and standard deviation of the accuracy of different methods provided in Example 1 of the present application.

[0031] Fig. 9 A schematic diagram of the functional modules of a single-cell identification device based on gene waves provided in Example 2 of the present application.

[0032] Fig.10 A schematic diagram of the structure of a computer device provided in Example 3 of the present application. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0034] Example 1

[0035] The single cell identification method based on gene wave provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the query data set to be processed to the server. After the server receives the query data set to be processed, for the query data set to be processed, the server preprocesses the gene expression matrix in the query data set so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set, and obtains the preprocessed matrix; based on the human reference genome, the genes in the preprocessed matrix are sorted, and for each single cell to be annotated, the expression amount of the sorted genes of the single cell to be annotated is composed of the gene wave of the single cell to be annotated; for each single cell to be annotated, the gene wave of the single cell to be annotated is used as input, and the recognition result of the single cell to be annotated is determined using the trained single cell recognition model. The server can feedback the recognition result of each single cell to be annotated included in the obtained object to be annotated to the terminal.

[0036] In addition, in some embodiments, the single-cell identification method based on gene waves can also be implemented independently by a server or a terminal. For example, the terminal can directly process the query data set to be processed, or the server can obtain the query data set to be processed from the data storage system and process the query data set to be processed.

[0037] The terminals may be, but are not limited to, various desktop computers, laptops, smart phones, tablet computers, IoT devices and portable wearable devices. IoT devices may be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server may be implemented as an independent server or a server cluster consisting of multiple servers, or a cloud server.

[0038] In an exemplary embodiment, Figure 2 , Figure 3 and Figure 4 As shown, a single cell identification method based on gene wave is provided. The method is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The following steps are used to illustrate the server in the example.

[0039] Step S1, obtaining a query data set; the query data set includes a gene expression matrix, and the gene expression matrix includes the expression level of each gene in each single cell to be annotated included in the object to be annotated; the object to be annotated is an organ or tissue of an organism.

[0040] Step S2, preprocessing the gene expression matrix so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set, and obtaining a preprocessed matrix; the reference data set includes a reference gene expression matrix and reference cell metadata, the reference gene expression matrix includes the expression level of each reference gene in each reference single cell included in the annotated object, and the reference cell metadata includes the label of each reference single cell included in the annotated object, the label is a cell type and / or cell subtype, and the annotated object is of the same type as the object to be annotated.

[0041] Step S3, based on the human reference genome, sorting the genes in the preprocessed matrix, and for each of the single cells to be annotated, forming a gene wave of the single cell to be annotated with the expression levels of the sorted genes; the horizontal axis of the gene wave is the gene, and the vertical axis of the gene wave is the expression level of the gene.

[0042] Step S4, for each of the single cells to be annotated, the gene wave of the single cell to be annotated is used as input, and the identification result of the single cell to be annotated is determined using a trained single cell recognition model; the trained single cell recognition model is a model trained based on the reference data set.

[0043] By implementing the above steps S1 to S4, this embodiment provides a method for identifying single cell types and / or subtypes based on gene waves. Inspired by the principle of speech recognition in a noisy environment, based on the expression waveforms of reordered genes, unified annotation of disease-specific cell types and / or subtypes can be achieved in heterogeneous single cell datasets.

[0044] The object to be annotated in this embodiment can be any tissue or organ of an organism of interest to the researcher, such as tendon, muscle, pancreas, lung, liver, etc. After determining the object to be annotated, single-cell transcriptome sequencing is performed on the object to be annotated to obtain single-cell transcriptome sequencing data of the object to be annotated, that is, a query data set is obtained, and the query data set includes a gene expression matrix, and the gene expression matrix includes the expression level of each gene in each single cell to be annotated included in the object to be annotated, specifically represented in the form of a gene × cell matrix, each column of the matrix represents a single cell to be annotated, each row of the matrix represents a gene, and each element of the matrix represents the expression level of a gene in a specific single cell to be annotated.

[0045] Optionally, the query data set of this embodiment may also include cell metadata, which includes additional information of each single cell to be annotated, such as the name of the single cell to be annotated, the tissue to which it belongs, the cell type preliminarily annotated by the user, and the cell subtype preliminarily annotated by the user.

[0046] The above gene expression matrix and cell metadata constitute the query dataset, which is the single-cell transcriptomics data to be annotated provided by the user.

[0047] The present embodiment also designs a reference data set, which is a single-cell transcriptomics data provided by the user and considered by the user to meet his needs. The user believes that the reference data set annotation is correct, and the label of the cell type or cell subtype meets his expectations, which is used as a reference for the gene position on the query data set. Specifically, the reference data set includes a reference gene expression matrix and reference cell metadata, and the reference gene expression matrix includes the expression amount of each reference gene in each reference single cell included in the annotated object, and the reference cell metadata includes the label of each reference single cell included in the annotated object, and the label is a cell type and / or a cell subtype, that is, the label can be a cell type, and the annotation of the cell type is completed at this time, the label can be a cell subtype, and the annotation of the cell subtype is completed at this time, and the label can be a cell type and a cell subtype, and the mixed annotation of the cell type and the cell subtype is completed at this time. The annotated object is the same as the object type to be annotated.

[0048] After introducing the reference data set, this embodiment further preprocesses the gene expression matrix so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set, and obtains a preprocessed matrix to ensure that the query data set takes genes that are consistent with the reference data set. If some reference genes do not exist in the query data set, they are filled with zeros.

[0049] The gene expression matrix is ​​preprocessed so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set to obtain a preprocessed matrix, which specifically includes:

[0050] (1) Compare the genes in the gene expression matrix with the reference genes in the reference gene expression matrix to determine the redundant genes and missing genes. The redundant genes are genes that are present in the gene expression matrix but not in the reference gene expression matrix, and the missing genes are genes that are not present in the gene expression matrix but present in the reference gene expression matrix.

[0051] (2) The expression levels of redundant genes in the gene expression matrix are deleted, that is, the rows where the redundant genes are located are deleted, and the missing genes are filled in the gene expression matrix. At the same time, the expression levels of the missing genes are filled with zero, that is, the rows where the missing genes are located are added, and each element of the rows where the missing genes are located is zero, so as to preprocess the gene expression matrix so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set, and a preprocessed matrix is ​​obtained.

[0052] After obtaining the preprocessed matrix, this embodiment further sorts the genes in the preprocessed matrix based on the human reference genome, and for each single cell to be annotated, the expression levels of the sorted genes of the single cell to be annotated are combined into a gene wave of the single cell to be annotated, thereby obtaining a gene wave of each single cell to be annotated. The horizontal axis of the gene wave is the gene, and the vertical axis of the gene wave is the expression level of the gene.

[0053] The human reference genome specifically uses version number HG38. HG38 (also known as GRCh38, Genome Reference Consortium human genome build 38) is a human genome version released in 2013. HG38 contains complete sequence information of the human genome, including genes, regulatory regions, and repeated sequences. Based on the human reference genome, the genes in the preprocessed matrix are sorted, specifically including: sorting the genes to be sorted in the preprocessed matrix according to the gene order of the human reference genome HG38, specifically downloading the hg38.refGene.gtf file of HG38, and after loading, sorting each gene to be sorted in the preprocessed matrix according to the value of the chromosome where each gene is located and the value of the termination site on the chromosome where the gene is located from small to large, that is, first sorting according to the value of the chromosome where the gene to be sorted is located from small to large, and for multiple genes on a chromosome, sorting according to the value of the termination site of the gene to be sorted on the chromosome where the gene is located from small to large, thereby obtaining the re-sorted genes.

[0054] Among them, after obtaining the reordered genes, only the rows where the reordered genes are located are retained in the preprocessed matrix, and each row in the preprocessed matrix is ​​adjusted according to the reordered genes, and then each adjusted column is used as the gene wave of the single cell to be annotated corresponding to the column, that is, the gene wave at this time includes the expression level of each reordered gene, and the gene wave of each single cell to be annotated is obtained.

[0055] The genes to be sorted are determined based on the different training methods used by the trained single-cell recognition model. When the training method at the cell group level is used, the genes to be sorted are the characteristic genes determined during the training process. When the training method at the cell level is used, the genes to be sorted are all the genes in the preprocessed matrix. At this time, when the training method at the cell group level is used, the genes in the preprocessed matrix are sorted based on the human reference genome, specifically including: sorting all the characteristic genes in the preprocessed matrix based on the human reference genome to obtain sorted genes; when the training method at the cell level is used, the genes in the preprocessed matrix are sorted based on the human reference genome, specifically including: sorting all the genes in the preprocessed matrix based on the human reference genome to obtain sorted genes.

[0056] It should be noted that if the preprocessed matrix does not contain the corresponding characteristic genes, the preprocessed matrix is ​​deleted and filled with zeros so that the preprocessed matrix only contains all the characteristic genes.

[0057] This embodiment reorders the genes expressed by each cell from small to large according to the position of the termination site of each gene in the human reference genome HG38, and then plots the level of gene expression to obtain the gene wave to which it belongs. The gene wave is very similar to the amplitude envelope of the sound wave. Each gene can be regarded as a frame segment, and the expression amount of the gene can be regarded as the value on the amplitude envelope of the sound wave. The deep learning algorithm on sound wave recognition can be used to analyze the gene expression of a single cell. Most of the existing cell annotation algorithms directly analyze or learn the cell-gene expression matrix, while this embodiment is based on the principle of deep learning in sound wave recognition, and regards the position information of gene open expression as the time information of the sound wave, and each gene interval is regarded as a frame segment of the sound wave. The scRNA-seq data is converted into one-dimensional data for learning and annotation according to the order of the human reference genome, which has a new feature compared to other algorithms-gene position information. As we all know, DNA is arranged in order on each chromosome, and the histones in the nucleosomes, as the spools around which DNA is wound, can affect the transcription of genes through post-translational modifications at specific sites. Moreover, genes are clustered on nucleosomes. Genes on the same nucleosome may be co-regulated and show similar expression patterns, forming a clustered expression phenomenon. Therefore, after the expressed genes are sorted according to the human reference genome, there may be correlations between adjacent genes. Therefore, adding position features is meaningful for each cell. Adding the position features of this dimension to the algorithm reshapes the gene expression pattern into a one-dimensional form of gene waves, adds new features, and is closer to the form of gene expression. In principle, it will have a certain effect on improving the accuracy of cell annotation. Figure 5 As shown, the black line is the expression curve of each gene, the red line is the expression curve processed by the Gaussian filter, the gene wave is the black line part, and the red line is a state that is easier to understand.

[0058] After obtaining the gene wave of each single cell to be annotated, for each single cell to be annotated, the gene wave of the single cell to be annotated is used as input, and the recognition result of the single cell to be annotated is determined using the trained single cell recognition model, which is a model trained based on the reference data set. When the label is a cell type, the recognition result is the cell type, when the label is a cell subtype, the recognition result is the cell subtype, and when the label is a cell type and a cell subtype, the recognition result is the cell type and the cell subtype.

[0059] The reference data set is introduced to train the initial single-cell recognition model to obtain a trained single-cell recognition model, and then the trained single-cell recognition model is used to complete the recognition of the query data set. Before the gene wave of the single cell to be annotated is used as input and the trained single-cell recognition model is used to determine the recognition result of the single cell to be annotated, the single-cell recognition method based on gene wave of this embodiment also includes:

[0060] (1) Obtain a reference dataset.

[0061] In the single-cell data, for the reference data set, genes with a total expression value less than 2 (if less than 2, it may be a false positive) are deleted, and labels (cell types and / or cell subtypes) with less than 10 cells (if less than 10, the corresponding features cannot be extracted) are deleted, because too few cells will cause the trained model to overfit.

[0062] At this point, obtain the reference data set, including:

[0063] 1) Obtaining the original data set, the original data set includes the original gene expression matrix and the original cell metadata, the original gene expression matrix includes the expression level of each original gene in each original single cell included in the annotated object, the original cell metadata includes the label of each original single cell included in the annotated object, and can also include the name of each original single cell included in the annotated object.

[0064] 2) For each original gene and each label, determine whether the total expression of the original gene is less than the first preset value. If so, the original gene is recorded as a gene to be deleted; determine whether the number of original single cells belonging to the label is less than the second preset value. If so, the label is recorded as a label to be deleted.

[0065] Among them, if the label is a cell type, each label refers to each cell type; if the label is a cell subtype, each label refers to each cell subtype; if the label is a cell type and a cell subtype, each label refers to each cell type and each cell subtype.

[0066] The total expression of the original gene is the sum of the expression of the original gene of each original single cell, which is equivalent to the sum of the elements in the row of the original gene in the original gene expression matrix. The first preset value may be 2, and the second preset value may be 10. It should be noted that the original genes other than the genes to be deleted are the reference genes, and the original single cells other than the original single cells belonging to the label to be deleted are the reference single cells.

[0067] 3) Delete the expression amount of the gene to be deleted and the expression amount of each original gene in the original single cell belonging to the label to be deleted in the original gene expression matrix, that is, delete the row where the gene to be deleted and the column where the original single cell belonging to the label to be deleted in the original gene expression matrix to obtain a reference gene expression matrix.

[0068] 4) Based on each reference single cell included in the reference gene expression matrix, the original cell metadata is processed, that is, only the label of each reference single cell is retained in the original cell metadata to obtain the reference cell metadata.

[0069] 5) Combining the reference gene expression matrix and reference cell metadata into a reference dataset.

[0070] (2) Use the reference data set to train the initial single-cell recognition model to obtain a trained single-cell recognition model.

[0071] This embodiment provides two training methods, one is a training method at the cell group level, that is, a training method for single-cell identification based on gene waves from cell group to cell level, and the other is a training method at the cell level, that is, a training method for single-cell identification based on gene waves from cell to cell level. After the training is completed, the subsequent cell-to-cell level single-cell identification method based on gene waves can be prepared to run.

[0072] If the cell population-level training method is used, the FindAllmarkers() function in the Seurat package is used to calculate the characteristic genes of each cell population (i.e., cell type and / or cell subtype), and then these characteristic genes are input into the initial single-cell recognition model at the cell population level for training according to the order of the human reference genome.

[0073] If a cellular level training method is used, the reference genes are directly input into the initial single-cell recognition model at the cellular level for training according to the order of the human reference genome.

[0074] The initial single cell recognition model can adopt any existing neural network model. As an example, Figure 6 The specific options are as follows:

[0075] If the training method at the cellular level is adopted, based on the gene wave of single cells, a one-dimensional convolutional neural network (1D CNN) is used, which includes an input layer, five hidden layers and an output layer connected in sequence. The five hidden layers are 2 convolutional layers, 1 pooling layer, 1 Flatten layer (flat layer) and 1 fully connected layer, that is, the first hidden layer is a convolutional layer, the second hidden layer is a convolutional layer, the third hidden layer is a pooling layer, the fourth hidden layer is a Flatten layer, and the fifth hidden layer is a fully connected layer. The number of reference genes in the reference data set is used as the dimension of the model input, the number of convolution kernels of the convolution layer is 16, the convolution kernel size is 3, and the Softmax function is used for the fully connected layer.

[0076] If a training method at the cell group level is used, a back propagation (BP) neural network is used based on the gene wave of the cell group.

[0077] The training process of the initial single-cell recognition model is mainly divided into two stages. The first stage is the forward propagation of the signal, and the second stage is the back propagation of the error. The steepest descent method is used as the learning rule. The weights and thresholds of the initial single-cell recognition model are continuously adjusted through back propagation to minimize the sum of squared errors of the initial single-cell recognition model.

[0078] If the cell group level training method is used, the initial single cell recognition model is trained using the reference data set to obtain a trained single cell recognition model, specifically including:

[0079] 1) Feature extraction is performed based on the reference gene expression matrix to determine the characteristic genes corresponding to each label.

[0080] Specifically, you can use the FindAllmarkers() function in the Seurat package to complete the feature extraction process. Take the reference gene expression matrix and the corresponding reference metadata as input. After using the Seurat package to build the corresponding Object, use the FindAllmarkers() function to determine the characteristic genes corresponding to each label. Each label corresponds to a cell population, so the characteristic genes related to the cell population can be determined.

[0081] If the label is a cell type, the characteristic genes of each cell type can be determined; if the label is a cell subtype, the characteristic genes of each cell subtype can be determined; if the label is a cell type and a cell subtype, the characteristic genes of each cell type and the characteristic genes of each cell subtype can be determined.

[0082] 2) Based on the human reference genome, all characteristic genes are sorted, and for each reference single cell, the expression levels of the sorted characteristic genes of the reference single cell are combined into the gene wave of the reference single cell.

[0083] 3) Taking the gene wave of each reference single cell as input and the label of each reference single cell as output, the initial single cell recognition model is trained to obtain a trained single cell recognition model.

[0084] If a cell-level training method is used, the initial single-cell recognition model is trained using a reference data set to obtain a trained single-cell recognition model, specifically including:

[0085] 1) Based on the human reference genome, the reference genes in the reference gene expression matrix are sorted, and for each reference single cell, the expression amounts of the sorted reference genes of the reference single cell are combined into a gene wave of the reference single cell.

[0086] 2) Taking the gene wave of each reference single cell as input and the label of each reference single cell as output, the initial single cell recognition model is trained to obtain a trained single cell recognition model.

[0087] When the label is a cell type, the trained single-cell recognition model can be used to generate recognition results for the cell type. When the label is a cell subtype, the trained single-cell recognition model can be used to generate recognition results for the cell subtype. When the label is a cell type and a cell subtype, two single-cell recognition models are trained, one for generating recognition results for the cell type and the other for generating recognition results for the cell subtype.

[0088] Optionally, in this embodiment, before training, the reference data set may be standardized and normalized, and before recognition, the query data set may be standardized and normalized.

[0089] After the training is completed, this embodiment uses indicators such as precision, recall, correctness, F1-score and receiver operating curve to evaluate the trained single-cell recognition model. Among them, precision refers to the proportion of samples of the true class among the samples predicted to be positive, recall refers to the proportion of samples predicted to be positive among all positive classes, correctness is a performance indicator for evaluating classification problems, that is, for given data, the proportion of the number of correctly classified samples to the total number of samples, F1-score is the harmonic mean of precision and recall, and is a measurement indicator for classification problems. The receiver operating curve uses the false positive rate as the horizontal axis and the true positive rate as the vertical axis, where the area under the curve can be used to comprehensively evaluate the classification performance of the model.

[0090] The training methods at the cell level and the cell group level have their own characteristics. The cell-to-cell level is more precise, while the cell group-to-cell level can capture group characteristics. Users can use the trained single-cell recognition models obtained by the two training methods for recognition, respectively, and choose the result that is more convincing to them from the recognition results, or make a final choice after comparing the two.

[0091] In this example, all scRNA-seq data (i.e., reference dataset and query dataset) can be preprocessed using R (version 3.6.1).

[0092] In this embodiment, the reference data set can be designed as a data set obtained under a healthy state, and the query data set can be designed as a data set obtained under a certain disease state. After completing the cell type / subtype annotation of each single cell in the query data set, based on the proportion of each cell type / subtype in a healthy state (that is, the proportion of the number of single cells of this cell type / subtype to the total number of single cells) and the proportion of each cell type / subtype in a certain disease state, the cell type / subtype whose proportion has changed significantly (which may be greater than a preset value) can be used as a disease-specific cell type / subtype.

[0093] This embodiment can obtain the label of the cell subtype or type after annotating each cell, and realize the unified annotation of specific cell types or subtypes for heterogeneous single-cell datasets across diseases. By introducing gene waves and trained single-cell recognition models, the recognition accuracy is higher. This embodiment compares the method used in this embodiment with tools such as Azimuth and Garnett on single-cell data on tissues such as Blood, Brain, Liver, and Pancreas in the HCL (Human Cell Landscape) database. Figure 7 and Figure 8 As shown, Figure 7In , the number after each tissue represents the number of single cells included in the tissue, for example, human blood 9649 represents that the human blood tissue includes 9649 single cells, our-method-cell represents the corresponding recognition method when the present embodiment adopts the cell-level training method, and our-method-group represents the corresponding recognition method when the present embodiment adopts the cell group-level training method. From the average, median, and standard deviation of the accuracy of each data set, it can be seen that the average (91.15%) and median (97.16%) of the accuracy of the method used in the present embodiment are higher than those of other tools (such as the average (67.95%) and median (89.51%) of the accuracy of Azimuth), and the standard deviation (0.1959) is lower than that of other tools such as Azimuth, indicating that the method of the present embodiment has a high accuracy, little fluctuation, and stable performance.

[0094] The current tools can realize the annotation of cell types across data sets and the annotation of cell subtypes within data sets. For example, algorithms such as scDeepsort and SingleR can annotate unknown cells to the large subgroup level based on the information of cells in large databases (HCL, MAC, etc.); algorithms such as scType and CellAssign can complete the annotation of cell types based on the existing cell marker gene database; algorithms such as scmap and Cellid can complete the cell type mapping between small data sets; the devcellpy algorithm can complete the annotation of cell subtypes on the heart development model, but the existing tools cannot accurately perform the annotation of cell types and cell subtypes across data sets. This may be caused by factors such as batch effects between data sets and the absence of cell subtype marker genes. Therefore, there are still problems in the rapid and unified annotation of DSCSs across data sets. It is necessary to develop a more suitable annotation algorithm. The use of gene waves in this embodiment solves the problem of annotation failure caused by the absence of marker genes. The use of neural networks reduces the influence of noise on cell gene expression characteristics and improves the problems caused by batch effects. The development inspiration of this embodiment comes from the application of deep learning in sound recognition. Just as everyone has a unique voiceprint, each cell has its own unique one-dimensional gene expression pattern: a feature containing gene expression patterns and location information. As is known to all, DNA is ordered and arranged in clusters on each chromosome, and genes on the same nucleosome may be co-regulated, showing similar expression patterns, forming a phenomenon of clustered expression. Therefore, the inclusion of gene position information is of great significance for the accurate comparison of expression patterns between cells. At the same time, referring to the application of filters in eliminating noise, filtering out unnecessary frequencies, etc., the filtering principle is introduced, and the position information of gene open expression is regarded as the time information of sound waves, and each gene interval is regarded as the frame segment of sound waves. This point can better predict and supplement missing values ​​in the processing of single-cell data with a high incidence of dropout events, and reduce the impact caused by missing information, thereby more effectively improving the accuracy of cell annotation. In summary, compared with the traditional annotation method that only considers the expression amount, the algorithm based on gene expression position and expression amount information proposed in this embodiment can more effectively improve the accuracy of cell annotation in principle.

[0095] This embodiment proposes a method for single cell type and subtype identification based on cell waves based on R and Python software packages to complete the unified mapping of DSCSs obtained from any tissue that may have different annotation levels. Based on the principle of sound recognition in a noisy environment, the gene expression of each cell is first reordered according to the position of each cell in HG38, and the expression of each gene in each cell is converted into the form of gene waves. Then, referring to the filter principle, the neural network method is used to reduce the impact of background noise on DSCSs identification in single cell analysis, and finally achieves rapid, unified and accurate annotation of DSCSs.

[0096] The present application also provides an application scenario, which applies the above-mentioned single-cell identification method based on gene waves. Specifically, the single-cell identification method based on gene waves provided in this embodiment can be applied in the cell annotation scenario. The cell annotation scenario includes a single-cell sequencing link and a single-cell annotation link. The single-cell sequencing link is used to sequence the object to be annotated to obtain single-cell transcriptome sequencing data, and the single-cell annotation link is used to process the single-cell transcriptome sequencing data to obtain the identification result of each single cell to be annotated included in the object to be annotated. The single-cell identification method based on gene waves provided in this embodiment belongs to the single-cell annotation link.

[0097] Example 2

[0098] Based on the same inventive concept, the embodiment of the present application also provides a single cell identification device based on gene waves for implementing the single cell identification method based on gene waves involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more single cell identification device embodiments based on gene waves provided below can refer to the limitations of the single cell identification method based on gene waves above, and will not be repeated here.

[0099] In an exemplary embodiment, Fig. 9 As shown, a single cell identification device based on gene wave is provided, and the single cell identification device based on gene wave comprises:

[0100] The data acquisition module M1 is used to acquire a query data set; the query data set includes a gene expression matrix, and the gene expression matrix includes the expression level of each gene in each single cell to be annotated included in the object to be annotated; the object to be annotated is an organ or tissue of an organism.

[0101] The preprocessing module M2 is used to preprocess the gene expression matrix so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set to obtain a preprocessed matrix; the reference data set includes a reference gene expression matrix and reference cell metadata, the reference gene expression matrix includes the expression level of each reference gene in each reference single cell included in the annotated object, and the reference cell metadata includes the label of each reference single cell included in the annotated object, the label is a cell type and / or cell subtype, and the annotated object is of the same type as the object to be annotated.

[0102] The sorting module M3 is used to sort the genes in the preprocessed matrix based on the human reference genome, and for each of the single cells to be annotated, the expression levels of the sorted genes of the single cell to be annotated are combined into a gene wave of the single cell to be annotated; the horizontal axis of the gene wave is the gene, and the vertical axis of the gene wave is the expression level of the gene.

[0103] The recognition module M4 is used to determine the recognition result of each single cell to be annotated by taking the gene wave of the single cell to be annotated as input and using a trained single cell recognition model; the trained single cell recognition model is a model trained based on the reference data set.

[0104] Example 3

[0105] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Fig.10 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a single cell recognition method based on gene waves is implemented.

[0106] Those skilled in the art will understand that Fig.10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0107] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the single-cell identification method based on gene waves in Example 1 is implemented.

[0108] Example 4

[0109] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which, when executed by a processor, implements the single-cell identification method based on gene waves in Example 1.

[0110] Example 5

[0111] In an exemplary embodiment, a computer program product is provided, including a computer program, which implements the single-cell identification method based on gene waves in Example 1 when executed by a processor.

[0112] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0113] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0114] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A single cell identification method based on gene waves, characterized in that: The single cell identification method based on gene wave includes: Acquire a query data set; the query data set includes a gene expression matrix, the gene expression matrix includes the expression level of each gene in each single cell to be annotated included in the object to be annotated; the object to be annotated is an organ or tissue of an organism; The gene expression matrix is ​​preprocessed so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set, so as to obtain a preprocessed matrix; the reference data set includes a reference gene expression matrix and reference cell metadata, the reference gene expression matrix includes the expression amount of each reference gene in each reference single cell included in the annotated object, the reference cell metadata includes the label of each reference single cell included in the annotated object, the label is a cell type and / or a cell subtype, and the annotated object is of the same type as the object to be annotated; Based on the human reference genome, the genes in the preprocessed matrix are sorted, and for each of the single cells to be annotated, the expression levels of the sorted genes of the single cell to be annotated are combined into a gene wave of the single cell to be annotated; the abscissa of the gene wave is the gene, and the ordinate of the gene wave is the expression level of the gene; For each of the single cells to be annotated, the gene wave of the single cell to be annotated is used as input, and the identification result of the single cell to be annotated is determined using a trained single cell identification model; the trained single cell identification model is a model trained based on the reference data set.

2. The single cell identification method based on gene wave according to claim 1, characterized in that: Preprocessing the gene expression matrix so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set to obtain a preprocessed matrix, specifically comprising: Compare the genes in the gene expression matrix with the reference genes in the reference gene expression matrix to determine redundant genes and missing genes; the redundant genes are genes that are present in the gene expression matrix but not in the reference gene expression matrix, and the missing genes are genes that are not present in the gene expression matrix but present in the reference gene expression matrix; The expression levels of redundant genes in the gene expression matrix are deleted, and the missing genes are filled in the gene expression matrix, and the expression levels of the missing genes are filled with zero, so as to preprocess the gene expression matrix so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set, and obtain a preprocessed matrix.

3. The single cell identification method based on gene wave according to claim 1, characterized in that: Before using the gene wave of the single cell to be annotated as input and using the trained single cell recognition model to determine the recognition result of the single cell to be annotated, the single cell recognition method based on gene wave also includes: Acquire the reference data set; The reference data set is used to train the initial single-cell recognition model to obtain a trained single-cell recognition model.

4. The single cell identification method based on gene wave according to claim 3 is characterized in that: Acquiring the reference data set specifically includes: Acquire an original data set; the original data set includes an original gene expression matrix and original cell metadata, the original gene expression matrix includes the expression amount of each original gene in each original single cell included in the annotated object, and the original cell metadata includes the label of each original single cell included in the annotated object; For each original gene and each label, determine whether the total expression of the original gene is less than a first preset value, if so, record the original gene as a gene to be deleted; determine whether the number of original single cells belonging to the label is less than a second preset value, if so, record the label as a label to be deleted; Deleting the expression amount of the gene to be deleted and the expression amount of each original gene in the original single cell belonging to the tag to be deleted from the original gene expression matrix to obtain a reference gene expression matrix; Based on each reference single cell included in the reference gene expression matrix, the original cell metadata is processed to obtain reference cell metadata; The reference gene expression matrix and the reference cell metadata are combined into a reference data set.

5. The single cell identification method based on gene wave according to claim 3 is characterized in that: When using the reference data set to train the initial single-cell recognition model, a cell population-level training method or a cell-level training method is adopted; When a cell population-level training method is used, the initial single-cell recognition model is trained using the reference data set to obtain a trained single-cell recognition model, specifically including: Perform feature extraction based on the reference gene expression matrix to determine the characteristic genes corresponding to each label; Based on the human reference genome, all the characteristic genes are sorted, and for each of the reference single cells, the expression amounts of the sorted characteristic genes of the reference single cell are combined into a gene wave of the reference single cell; Taking the gene wave of each reference single cell as input and the label of each reference single cell as output, the initial single cell recognition model is trained to obtain a trained single cell recognition model; When a cell-level training method is used, the initial single-cell recognition model is trained using the reference data set to obtain a trained single-cell recognition model, specifically including: Based on the human reference genome, the reference genes in the reference gene expression matrix are sorted, and for each of the reference single cells, the expression amounts of the sorted reference genes of the reference single cell are combined into a gene wave of the reference single cell; The gene wave of each reference single cell is used as input, and the label of each reference single cell is used as output to train the initial single cell recognition model to obtain a trained single cell recognition model.

6. The single cell identification method based on gene wave according to claim 5, characterized in that: When a cell population-level training method is used, the genes in the preprocessed matrix are sorted based on the human reference genome, specifically including: Based on the human reference genome, all characteristic genes in the preprocessed matrix are sorted to obtain sorted genes; When a cell-level training method is used, the genes in the preprocessed matrix are sorted based on the human reference genome, specifically including: Based on the human reference genome, all genes in the preprocessed matrix are sorted to obtain sorted genes.

7. A single cell identification device based on gene waves, characterized in that: The single cell identification device based on gene wave comprises: A data acquisition module, used to acquire a query data set; the query data set includes a gene expression matrix, the gene expression matrix includes the expression level of each gene in each single cell to be annotated included in the object to be annotated; the object to be annotated is an organ or tissue of an organism; A preprocessing module, used for preprocessing the gene expression matrix so that the genes in the gene expression matrix are completely consistent with the reference genes in the reference data set, and obtaining a preprocessed matrix; the reference data set includes a reference gene expression matrix and reference cell metadata, the reference gene expression matrix includes the expression amount of each reference gene in each reference single cell included in the annotated object, the reference cell metadata includes the label of each reference single cell included in the annotated object, the label is a cell type and / or a cell subtype, and the annotated object is of the same type as the object to be annotated; A sorting module, for sorting the genes in the preprocessed matrix based on the human reference genome, and for each of the single cells to be annotated, composing the expression levels of the sorted genes of the single cell to be annotated into a gene wave of the single cell to be annotated; the abscissa of the gene wave is the gene, and the ordinate of the gene wave is the expression level of the gene; The recognition module is used to determine the recognition result of each single cell to be annotated by using the gene wave of the single cell to be annotated as input and using a trained single cell recognition model; the trained single cell recognition model is a model trained based on the reference data set.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the single-cell identification method based on gene waves as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the single-cell identification method based on gene waves described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the single-cell identification method based on gene waves described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Single cell automatic classification method and device based on characteristic genes

    CN112837754A

  • Cell data annotation method, device, equipment and medium

    CN115116549A

  • Tumor single cell transcriptome sequencing data cell type identification method and system

    CN117037907A

  • Single cell automatic annotation method and device based on selective domain discriminator

    CN117789837A

  • Single cell type detection method, apparatus, device, and storage medium

    WO2020154885A1