Cell Data Annotation Method, Apparatus, Device and Medium
By obtaining the cellular data of the transcriptome to be predicted, determining the annotated cell object and using the prediction model to correct it, the problem of low cell annotation accuracy in the prior art is solved, and more accurate cell type annotation is achieved.
Patent Information
- Application Number
- CN202210442634.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-04-25
AI Technical Summary
In the prior art, the cell annotation accuracy is low due to the single-sided characteristics of the reference data set during cell type annotation.
By obtaining the cellular data of the transcriptome to be predicted, the annotated cell object is determined, and the initial cell annotation results are determined based on the gene expression information. The initial results are corrected using the first prediction model, and the gene expression and spatial information are fused for self-supervised correction.
It improves the accuracy of cell annotation results and combines more comprehensive feature information to improve the accuracy of spatial transcriptome cell annotation.
Smart Images

Figure CN115116549B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of bioinformatics analysis technology, and specifically relates to a method, device, equipment and medium for annotating cell data. Background Art
[0002] With the continuous development of bioinformatics technology, spatial transcriptome technology has been widely applied to various medical research fields, such as different fields like tumor research, neuroscience, developmental biology, molecular pathology, etc. During the medical research process, in order to conduct research such as gene differential expression analysis, cell development trajectory analysis, and Gene Ontology (GO) enrichment analysis based on spatial transcriptome cell data, it is necessary to annotate the spatial transcriptome cell data.
[0003] Currently, in the related art, a classification training method can be used to build a model based on a reference data set, and then the cell type of the data set to be annotated can be predicted by using the built model.
[0004] However, the reference data set used in this method for determining the cell type has the problem of one-sided features, which leads to a low accuracy of cell type annotation. Summary of the Invention
[0005] In view of the above-mentioned defects or deficiencies in the prior art, it is desired to provide a method, device, equipment and medium for annotating cell data, which can obtain the cell annotation result of the transcriptome to be predicted based on cell data with relatively comprehensive features, thereby improving the accuracy of the cell annotation result of the spatial transcriptome. The technical solution is as follows:
[0006] According to one aspect of the present application, a method for annotating cell data is provided, and the method includes:
[0007] Obtain the cell data of the transcriptome to be predicted; the cell data includes the gene expression information of multiple sequencing points in the transcriptome to be predicted and the spatial information of the multiple sequencing points;
[0008] Determine the annotated cell object corresponding to the transcriptome to be predicted, and determine the initial cell annotation result of the transcriptome to be predicted according to the cell object and the gene expression information;
[0009] Input the cell data into the first prediction model, and correct the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result to obtain the cell annotation result of the transcriptome to be predicted.
[0010] According to another aspect of the present application, a device for annotating cell data is provided, and the device includes:
[0011] An acquisition module for acquiring cell data of a transcriptome to be predicted; the cell data includes gene expression information of multiple sequencing points in the transcriptome to be predicted and spatial information of the multiple sequencing points.
[0012] A processing module for determining an annotated cell object corresponding to the transcriptome to be predicted, and determining an initial cell annotation result of the transcriptome to be predicted according to the cell object and the gene expression information.
[0013] A cell annotation module for inputting the cell data into a first prediction model, and correcting the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result to obtain a cell annotation result of the transcriptome to be predicted.
[0014] According to another aspect of the present application, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the cell data annotation method as described above.
[0015] According to another aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and the computer program is used to implement the cell data annotation method as described above.
[0016] According to another aspect of the present application, there is provided a computer program product, on which instructions are included, and when the instructions are executed, the cell data annotation method as described above is implemented.
[0017] The cell data annotation method, device, equipment and medium provided in the embodiments of the present application determine annotated cell objects corresponding to the transcriptome to be predicted, and determine the initial cell annotation result of the transcriptome to be predicted according to the cell objects and gene expression information. Without any manual annotation, guidance information for the cell annotation result of the transcriptome to be predicted can be obtained, and the cell data is predicted by the first prediction model, and the initial cell annotation result is corrected according to the loss between the output of the first prediction model and the initial cell annotation result. The gene expression information of multiple sequencing points and the spatial information of multiple sequencing points are effectively fused, and self-supervised correction can be performed in combination with the guidance information. On the one hand, the spatial information of the spatial transcriptome can characterize the spatial distribution characteristics of each sequencing point in the spatial transcriptome, and the gene expression information can characterize the gene characteristics of each sequencing point. Determining the cell annotation result of the spatial transcriptome based on the spatial information and gene expression information combines more comprehensive features to determine the cell annotation result compared with the prior art, and also significantly improves the accuracy of the cell annotation result determined by the method provided in the present application compared with the prior art. On the other hand, the cell annotation result is further corrected according to the prediction loss of the first prediction model, and with the help of the self-supervised learning algorithm of the first prediction model, the accuracy of the cell annotation result is further improved.
[0018] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:
[0020] Figure 1 It is a system architecture diagram of the application system for cell data annotation provided in the embodiments of the present application;
[0021] Figure 2 It is a flowchart of the cell data annotation method provided in the embodiments of the present application;
[0022] Figure 3 It is a structural diagram of the cell data annotation provided in the embodiments of the present application;
[0023] Figure 4 It is a flowchart of the method for obtaining the initial cell annotation result provided in the embodiments of the present application;
[0024] Figure 5 It is a flowchart of the method for obtaining the cell annotation result provided in the embodiments of the present application;
[0025] Figure 6Schematic structural diagram of obtaining an output result through a first prediction model provided by an embodiment of the present application;
[0026] Figure 7 Schematic structural diagram of obtaining an output result through a first prediction model provided by another embodiment of the present application;
[0027] Figure 8 Schematic structural diagram of obtaining an output result through a first prediction model provided by yet another embodiment of the present application;
[0028] Figure 9 Schematic flowchart of a cell data annotation method provided by an embodiment of the present application;
[0029] Figure 10 Schematic flowchart of a method for training a first prediction model provided by an embodiment of the present application;
[0030] Figure 11 Schematic flowchart of a method for training a first prediction model provided by an embodiment of the present application;
[0031] Figure 12 Schematic flowchart of a method for predicting a transcriptome to be predicted provided by an embodiment of the present application;
[0032] Figure 13 Schematic structural diagram of a cell data annotation device provided by an embodiment of the present application;
[0033] Figure 14 Schematic structural diagram of a cell data annotation device provided by another embodiment of the present application;
[0034] Figure 15 Schematic structural diagram of a computer device shown in an embodiment of the present application. Detailed implementation manners
[0035] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of convenience of description, only parts related to the invention are shown in the drawings.
[0036] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments. For the sake of easy understanding, some technical terms related to the embodiments of the present application are explained below:
[0037] (1) Artificial Intelligence (AI): It uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce an intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0038] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software mainly includes several major directions such as computer vision, speech processing technology, natural language technology, and machine learning / deep learning.
[0039] (2) Machine Learning (ML): It is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0040] (3) Deep Neural Networks (DDN): It is a neural network with at least one hidden layer. Similar to shallow neural networks, deep neural networks can model complex nonlinear systems, but the additional layers provide a higher level of abstraction for the model, thus improving the model's capabilities.
[0041] (4) Spatial Transcriptomics (ST): It can be the collection of all transcripts within cells under a certain physiological condition. For example, it can be messenger RNA, ribosomal RNA, transfer RNA, and other non-coding RNAs. Or, it can also be considered that spatial transcriptomics is the collection of all mRNAs.
[0042] (5) Cell annotation result: It refers to the processing result obtained after analyzing and annotating the spatial transcriptome data, which is used to identify the identity information of the cell data in the spatial transcriptome, so as to quickly obtain the cell type information and cell attribute characteristics of the spatial transcriptome data, etc.
[0043] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart home, smart wearable social security, virtual assistant, smart speaker, smart marketing, driverless, autonomous driving, drone, robot, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0044] The solution provided in the embodiments of this application relates to technologies such as neural networks of artificial intelligence, which will be specifically described through the following embodiments.
[0045] Currently, in the related art, a classification training method can be used to build a model based on a reference data set, and then the cell type of the data set to be annotated can be predicted through the built model. However, in the process of predicting the cell type by this automatic annotation method, the use of the reference data set is relatively one-sided, resulting in a low accuracy of cell data annotation.
[0046] Based on the above defects, this application provides a cell data annotation method, device, equipment and medium. Compared with the prior art, it can obtain the cell annotation result of the transcriptome to be predicted based on cell data with relatively comprehensive features, thereby improving the accuracy of the cell annotation result of the spatial transcriptome.
[0047] Figure 1 It is the implementation environment architecture diagram of a cell data annotation method provided by the embodiments of this application. As Figure 1 shown, this implementation environment architecture includes: terminal 10 and server 20.
[0048] Among them, in the field of sequencing analysis, the process of annotating the cell data of the spatial transcriptome can be executed either on the terminal 10 or on the server 20. For example, by receiving the cell data input by the user through the terminal 10, cell annotation can be performed locally on the terminal 10 to obtain the cell annotation result; or the cell data can be sent to the server 20, so that the server 20 receives the cell data, performs cell annotation according to the cell data to obtain the cell annotation result, and then sends the cell annotation result to the server 20 to implement the cell annotation of the cell data of the transcriptome to be predicted.
[0049] The cell data annotation solution provided by the embodiments of the present application improves the accuracy of the computer device in annotating spatial transcriptome cells. The above cell annotation results can be applied to multiple sequencing analysis fields involved in computer devices such as terminals or servers. For example, gene differential expression analysis, cell development trajectory analysis, Gene Ontology (GO) enrichment analysis, etc. It can also be provided to users as preliminary cell annotation results for modification and fine-tuning to obtain a complete cell annotation result that the user deems more credible.
[0050] In addition, the terminal 10 can display an application interface through which cell data of the to-be-predicted transcriptome uploaded by the user can be obtained, or the cell data of the to-be-predicted transcriptome uploaded can be sent to the server 20.
[0051] Optionally, the terminal 10 can be a terminal device in various AI application scenarios. For example, the terminal 10 can be a smart TV, a smartphone, a tablet computer, a television, a laptop computer, a desktop computer, etc. The embodiments of the present application do not specifically limit this.
[0052] The server 20 can be a single server, a server cluster or a distributed system composed of several servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.
[0053] A communication connection is established between the terminal 10 and the server 20 through a wired or wireless network. Optionally, the above wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network.
[0054] For ease of understanding and explanation, the following Figures 2 to 15 elaborates in detail the cell data annotation method, device, equipment and medium provided by the embodiments of the present application.
[0055] Figure 2 The following shows a schematic flowchart of the cell data annotation method according to the embodiments of the present application. This method can be executed by a computer device, and the computer device can be the above Figure 1Either the server 20 or the terminal 10 in the system shown, or the computer device may also be a combination of the terminal 10 and the server 20. As Figure 2 shown, the method includes:
[0056] S101. Obtain cell data of the transcriptome to be predicted; the cell data includes gene expression information of multiple sequencing points in the transcriptome to be predicted and spatial information of the multiple sequencing points.
[0057] The above cell data of the transcriptome to be predicted can be information obtained by gene sequencing of the sample to be predicted (tissue or part). For example, sequencing points can be selected on the section of the sample to be predicted, and gene sequencing is performed on the sequencing points to obtain the cell data of the transcriptome to be predicted. Among them, the above tissues can be brain tissue, heart tissue, lung tissue, etc.
[0058] It should be noted that the above sequencing points can be some cells in the section used to obtain the transcript information of the transcriptome to be predicted.
[0059] In a possible implementation manner, the information obtained by gene sequencing of the sequencing points can include gene expression information of the sequencing points and spatial information of the sequencing points. The gene expression information is used to characterize the gene information of each sequencing point, and the spatial information of the sequencing points is used to characterize the spatial position information of each sequencing point. That is to say, the cell data of the transcriptome can include the spatial information of each sequencing point in the transcriptome and the gene expression information of each sequencing point.
[0060] In a possible implementation manner, the acquisition method of the above cell data of the spatial transcriptome can include: (1) Combining microscopic imaging technology and sequencing technology to obtain transcriptome data in the section of a certain tissue, and then calling the 10x Genomics Visium platform to analyze, process and visualize the data;
[0061] (2) Obtained based on transcriptome sequencing and transcriptome analysis methods based on laser capture microdissection (LCM). This sequencing method can, for example, include a new spatial transcriptome method (transcriptome in vivo analysis, TIVA), an analysis method of padlock probe in situ sequencing (in situ sequencing, ISS), fluorescence in situ sequencing (fluorescence in situ sequencing, FISSEQ);
[0062] (3) Obtained by directly importing through an external device.
[0063] S102. Determine the annotated cell object corresponding to the transcriptome to be predicted, and determine the initial cell annotation result of the transcriptome to be predicted according to the cell object and gene expression information.
[0064] Specifically, the annotated cell object corresponding to the transcriptome to be predicted may be a cell object that is homologous to the transcriptome to be predicted and has a known cell annotation result. For example, the annotated cell object may be a single cell or a cell group included in the gene sequencing object corresponding to the transcriptome to be predicted. Among them, a single cell refers to cell data formed by sequencing each sample on a cell-by-cell basis. A cell group refers to cell data formed by sequencing each sample on a basis of multiple cells, such as bulk data, or each sample may be subcellular-level cell data.
[0065] It should be noted that the gene sequencing object corresponding to the transcriptome to be predicted may be a gene sequencing object homologous to the transcriptome to be predicted. For example, a gene sequencing object taken from the same tissue or part as the transcriptome to be predicted. Exemplarily, the transcriptome to be predicted taken from the brain corresponds to the gene sequencing object taken from the brain, and does not correspond to the gene sequencing object taken from the heart. The cell data of the transcriptome to be predicted taken from the heart corresponds to the gene sequencing object taken from the heart. Among them, if the cell data of the transcriptome to be predicted is a brain slice, the determined gene sequencing object should also be taken from the same area of the brain; if the cell data of the transcriptome to be predicted is a heart slice, the determined corresponding gene sequencing object should also be taken from the same area of the heart.
[0066] In this embodiment, by determining the cell object corresponding to the transcriptome to be predicted, it is possible to determine the guiding information for obtaining the cell annotation result of the transcriptome to be predicted for the correct cell object, and thus make the determined initial cell annotation result more accurate.
[0067] In the process of determining the annotated cell object corresponding to the transcriptome to be predicted in the embodiment of the present application, as a possible implementation manner, it may be to identify the annotated cell data and determine whether the region or tissue of the cell data is the same as the region or tissue corresponding to the cell data of the transcriptome to be predicted. In one possible implementation manner, the cell object may be determined by observing the morphological and structural characteristics of cells in a large dataset. For example, epithelial cells are composed of dense cells and extracellular matrix, muscle tissue is composed of muscle cells and is fibrous, and nerve tissue is composed of neurons and glial cells. In another possible implementation manner, the annotated cell data taken from the same tissue or region as the transcriptome to be predicted may be queried and obtained from a preset cell database, and the annotated cell data is determined as the annotated cell object corresponding to the transcriptome to be predicted. Among them, the preset cell database may be constructed by summarizing and classifying cell data with different cell types, cell morphologies, and structural characteristics.
[0068] Further, when determining the initial cell annotation result of the transcriptome to be predicted based on the cell object and gene expression information, as an implementable method, the method of single-cell migration can be adopted. According to the single cell and the gene expression information in the transcriptome to be predicted, the initial cell annotation result of the transcriptome to be predicted can be determined. For example, a neural network model can be trained in advance based on the single cell and the corresponding cell annotation result of the single cell. The neural network model has the ability to annotate the cell type of the cell data of the transcriptome to be predicted, and then the neural network model is used to predict the cell data of the transcriptome to be predicted to obtain the initial cell annotation result.
[0069] As another implementable method, other non-single-cell migration methods can be adopted. For example, when the non-single cell is multiple cells for each sample, the corresponding algorithm can be a method combining deconvolution and migration to obtain the initial cell annotation result of the transcriptome to be predicted; for example, when the non-single cell is at the subcellular level, the corresponding algorithm is a method combining aggregation and migration to obtain the initial cell annotation result of the transcriptome to be predicted.
[0070] Another possible implementation method is also to comprehensively analyze the gene expression information and spatial information by identifying the mark marker genes specifically expressed by cell types to determine the initial cell annotation result of the transcriptome to be predicted; the method of introducing prior knowledge in the corresponding field can also be adopted, and the cell data of the transcriptome to be predicted is clustered through a clustering algorithm to obtain a clustering result, and then the initial cell annotation result of each clustering result in the transcriptome to be predicted is determined by using artificial prior biological knowledge; it is also possible to query the pre-established transcriptome database of known cell types, compare the cell characteristics of the cell data of the unknown type of transcriptome to be predicted with the cell characteristics of the transcriptome database of known cell types, and determine the cell type with the same cell characteristics as the initial cell annotation result of the transcriptome to be predicted.
[0071] It should be noted that the above implementation methods for determining the initial cell annotation result of the transcriptome to be predicted based on the cell object and gene expression information are only examples, and the embodiments of the present application do not limit this.
[0072] In the embodiments of the present application, after determining the cell object corresponding to the transcriptome to be predicted, determining the initial cell annotation result of the transcriptome to be predicted according to the cell object and gene expression information can provide guiding information for determining the final cell annotation result, so as to provide an accurate data basis for the final cell annotation result, thereby improving the accuracy of cell data annotation.
[0073] S103. Input the cell data into the first prediction model, and correct the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result, so as to obtain the cell annotation result of the transcriptome to be predicted.
[0074] It should be noted that the above first prediction model is a neural network model with the gene expression information and spatial information of the transcriptome to be predicted as the input, the cell annotation result of the transcriptome to be predicted as the output, and the ability to integrate the gene expression information and spatial information of the transcriptome to be predicted, and can predict the cell annotation result. The first prediction model can be the initial model during iterative training through a self-supervised algorithm, that is, the model parameters of the first prediction model are in the initial state, or it can be the model adjusted in the previous round of iterative training, that is, the model parameters of the first prediction model are in the intermediate state. The cell data can be input into the first prediction model to obtain an output result. The output result includes the cell type of the transcriptome to be predicted, or may include multiple cell attributes of the transcriptome to be predicted under the cell type.
[0075] Specifically, please refer to Figure 3 As shown, after determining the cell object corresponding to the transcriptome to be predicted, the initial cell annotation result 310 of the transcriptome to be predicted can be determined according to the cell object and the gene expression information. For example, the initial cell annotation result of the transcriptome to be predicted can be obtained according to the cell annotation method of the above-mentioned annotated cell object.
[0076] Optionally, when inputting cell data into the first prediction model 320, it can be to obtain fused feature information by fusing gene expression information and spatial information, and then perform classification processing on the fused feature information to obtain the output result 330 of the first prediction model, that is, the cell annotation result of the above-mentioned transcriptome to be predicted predicted by the first prediction model. It is also possible to correct the initial cell annotation result 310 according to the loss between the output result 330 of the first prediction model and the initial cell annotation result 310 to obtain the cell annotation result 340 of the transcriptome to be predicted. Specifically, the process of correcting the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result can be to use the initial cell annotation result as a pseudo-label, integrate the gene expression information and spatial information of the transcriptome to be predicted using the network framework of a Graph Convolutional Network (GCN), and perform classification through a classification model to obtain the output result of the first prediction model. Then, the initial cell annotation result is corrected in a self-supervised training manner according to the output result, so as to obtain an adjusted first prediction model. Furthermore, the cell data is predicted through the adjusted first prediction model to obtain the cell annotation result of the transcriptome to be predicted. Optionally, the above-mentioned first prediction model may include an encoding-decoding module and a classification module. The encoding-decoding module can be, for example, a Deep Auto-Encoder (DAE) and a Variational Graph Auto-Encoder (VGAE), and can also be other auto-encoding modules. For example, VGAE can be replaced by other graph network frameworks such as GraphSAGE. Among them, GraphSAGE is equivalent to a framework for aggregating adjacent node information. In this framework, different aggregation functions can be used to combine the information of adjacent nodes and the information from the previous layer of the current node. Among them, the aggregation functions can be Mean, Pool, LSTM, etc.
[0077] Among them, the above cell annotation results are used to identify the transcriptome to be predicted, so that information, characteristics, etc. of the transcriptome to be predicted can be quickly obtained through the annotation results of the transcriptome to be predicted. For example, the annotation results of the transcriptome to be predicted may include the cell types of the transcriptome to be predicted, or may include multiple cell attributes of the transcriptome to be predicted under cell types. For example, cell types can be epithelial cells, T cells, fibroblast cells, shape glial cells, endothelial cells, etc. Among them, the cell attributes corresponding to T cells can be, for example, helper T cells, inhibitory T cells, effector T cells, cytotoxic T cells, memory T cells, etc. The functions of T cells corresponding to different cell attributes are different. Helper T cells have the function of assisting humoral immunity and cell immunity; inhibitory T cells have the function of inhibiting cell immunity and humoral immunity; effector T cells have the function of releasing lymphokines; cytotoxic T cells have the function of killing target cells; memory T cells have the function of remembering specific antigen stimulation.
[0078] In the process of performing cell data annotation on the transcriptome to be predicted, as a feasible implementation, when performing cell data annotation on the transcriptome to be predicted, cell data of the transcriptome to be predicted can be obtained. The cell data includes gene expression information of multiple sequencing points in the transcriptome to be predicted and spatial information of multiple sequencing points. Then, an annotated cell object corresponding to the transcriptome to be predicted is determined. By obtaining the tissue or region from which the cell data is taken from a large dataset, and then determining the cell data with the same tissue or region as the tissue or region from which the transcriptome to be predicted is taken as the cell object. The cell object can be a single cell or a cell group corresponding to the transcriptome to be predicted. And an initial cell annotation result of the transcriptome to be predicted is determined according to the cell object and the gene expression information. The initial cell annotation result of the transcriptome to be predicted can be determined by the method of single cell migration according to the single cell and the gene expression information in the transcriptome to be predicted.
[0079] Exemplarily, for example, a neural network model is pre-trained based on single cells and the corresponding cell annotation results of single cells. Then, the cell data of the transcriptome to be predicted is input into the neural network model, and an initial cell annotation result is obtained through prediction processing. The initial cell annotation result can be used as the final cell annotation result of the transcriptome to be predicted. In this embodiment, only the single cell data with known annotation results is used to annotate the cell data of the transcriptome to be predicted, the sequencing differences between single cells and spatial data cannot be captured, and the most important spatial information is not utilized, so the accuracy of the obtained cell annotation result is relatively poor.
[0080] As another implementable approach, the manner of single-cell migration or other manners of single-cell migration can be used to determine the initial cell annotation result according to the gene expression information of the transcriptome to be predicted, and the initial cell annotation result can be used as the pseudo-label of the self-supervised model. Then, the cell data of the transcriptome to be predicted is input into the first prediction model to obtain an output result, and the initial cell annotation result is corrected according to the loss between the output result of the first prediction model and the pseudo-label, so as to obtain the cell annotation result of the transcriptome to be predicted.
[0081] In the cell data annotation method provided by the embodiments of the present application, by determining the cell object corresponding to the transcriptome to be predicted and determining the initial cell annotation result of the transcriptome to be predicted according to the cell object and the gene expression information, guidance information of the cell annotation result of the transcriptome to be predicted can be obtained without any manual annotation. And by predicting the cell data through the first prediction model and correcting the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result, the gene expression information of multiple sequencing points and the spatial information of multiple sequencing points are effectively fused, and self-supervised correction can be performed in combination with the guidance information, so as to more accurately and comprehensively utilize all the information in the cell data of the transcriptome to be predicted, and further obtain the cell annotation result of the transcriptome to be predicted more accurately, greatly improving the accuracy of the cell data annotation of the transcriptome to be predicted. It can also be applied to a sequencing analysis system to accurately predict the cell data of the transcriptome to be predicted, greatly improving the quality and efficiency of cell annotation, and providing strong support for the data analysis of spatial transcriptomics.
[0082] In another embodiment of the present application, the initial cell annotation result of the transcriptome to be predicted can be obtained by means of a neural network model applicable to the annotated cell object. Figure 4 A specific implementation manner for determining the initial cell annotation result of the transcriptome to be predicted is provided. Please refer to Figure 4 as shown, which specifically includes:
[0083] S201. Determine a second prediction model that matches the cell object; the second prediction model is trained based on the historical gene expression information of the cell object and the historical cell annotation result corresponding to the historical gene expression information.
[0084] Specifically, the second prediction model that matches the annotated cell object is a neural network model capable of predicting the annotation result of the annotated cell object. This second prediction model is trained based on the historical gene expression information of the cell object and the corresponding historical cell annotation results of the historical gene expression information. Among them, it can be the true cell type or true cell attribute of the cell object. This second prediction model can obtain an initial cell annotation result based on the gene expression information. First, the feature information in the gene expression information can be extracted, then the feature information is passed through a fully connected layer to obtain a fully connected vector, and finally the fully connected vector is processed through an activation function to obtain the initial cell annotation result.
[0085] In a possible implementation manner, to determine the second prediction model that matches the cell object, the model file of the second prediction model is obtained when performing step S201, so as to subsequently run the model file to predict the annotation result of the transcriptome to be predicted.
[0086] In another possible implementation manner, the historical gene expression information of the annotated cell object and the corresponding historical cell annotation results of the historical gene expression information can be obtained when performing step S201, and then the above-mentioned second prediction model is obtained through real-time training according to the historical gene expression information and the corresponding historical cell annotation results of the historical gene expression information.
[0087] Among them, the cell object can be a single cell or a cell group. This second prediction model has the ability to annotate the cell type of cell data. This second prediction model can be trained through a preset training algorithm, that is, the model parameters of the second prediction model are in an optimal state. The historical cell annotation results can be the true cell type or true cell attribute of the cell object.
[0088] Optionally, the above neural network model can be a deep neural networks (DNN) model, a Convolutional neural networks (CNN) model, a Long short term memory cells (LSTM) model, a Recurrent neural networks (RNN) model, etc.
[0089] In a possible implementation, the training process of the second prediction model specifically includes: randomly dividing the historical gene expression information into a training set and a validation set according to a certain ratio, and then constructing the second prediction model using the training set and the validation set according to the training learning algorithm. Among them, the training set is used to train the initial second prediction model to obtain a trained second prediction model, and the validation set is used to verify the trained second prediction model to verify the performance of the second prediction model.
[0090] Among them, during the process of the computing device training the second prediction model, the second prediction model to be verified in the validation set is used to optimize the second prediction model to be verified according to the minimization of the loss function, and a second prediction model is obtained. According to the difference between the result obtained by inputting the validation set into the second prediction model to be verified and the historical cell annotation result corresponding to the historical gene expression information, the parameters in the second prediction model to be constructed are updated to achieve the purpose of training the second prediction model. Among them, the above historical cell annotation result can be the result obtained by manually annotating the historical gene expression information.
[0091] In the embodiment of the present application, when training the second prediction model, the loss function can be used to calculate the loss value between the result obtained by inputting the validation set into the second prediction model to be verified and the historical cell annotation result, so as to update the parameters in the second prediction model to be verified. Optionally, the loss function can use the cross-entropy loss function, the normalized cross-entropy loss function, or Focalloss, etc.
[0092] S202: Input the gene expression information into the second prediction model to obtain the initial cell annotation result of the transcriptome to be predicted.
[0093] After constructing the second prediction model, the gene expression information of the transcriptome to be predicted can be input into the second prediction model for prediction processing to obtain the initial cell annotation result of the transcriptome to be predicted. Among them, the initial cell annotation result can include the cell type of the transcriptome to be predicted, or can include multiple cell attributes of the transcriptome to be predicted under the cell type.
[0094] In one possible implementation, the processing of the gene expression information by the second prediction model specifically includes: during the process of processing the gene expression information, it may first extract features from the transcriptome to be predicted to obtain feature information, and then perform operations on the feature information through a multi-classification function to output cell types. It is also possible to perform operations on the feature information through multiple binary classifications to output cell attributes. Optionally, the multi-classification function may be a softmax function, and the multiple binary classification functions may be multiple sigmoid functions. One sigmoid function can achieve a binary classification prediction. Among them, the role of the multi-classification function is to add non-linear factors because the expression ability of the linear model is insufficient, and it can transform the input continuous real values into outputs between 0 and 1.
[0095] Exemplarily, when the gene expression information is input into the second prediction model, the prediction results of the second prediction model may include any one of cell types such as "T cell", "epithelial cell", and "fibroblast". Among them, the prediction results may also include cell attributes. For example, the cell attributes corresponding to the cell type "T cell" may be multiple of "helper T cell", "suppressor T cell", "effector T cell", "cytotoxic T cell", and "memory T cell".
[0096] Among them, taking the three-classification as an example, the output of the multi-classification function is introduced. For example, the cell types that the multi-classification function can predict are respectively "T cell", "epithelial cell", and "fibroblast". The output result of the neural network model can be represented by a vector. For example, it can be a 3*1-dimensional vector. Each element in the vector corresponds to a cell type, and the value of each element in the vector represents the probability that the transcriptome to be predicted is the corresponding label type. Assuming that the output vector of the multi-classification function is [0.61, 0.31, 0.08], it means that the probability that the transcriptome to be predicted is a "T cell" is 0.61, the probability that the transcriptome to be predicted is an "epithelial cell" is 0.31, and the probability that the transcriptome to be predicted is a "fibroblast" is 0.08. The element value with the largest probability can be selected as the prediction result of the transcriptome to be predicted, that is, taking "T cell" as the initial cell annotation result of the transcriptome to be predicted.
[0097] Taking the ternary binary classification as an example, the output of the multi-class binary classification function is introduced. For example, the cell attributes that the multi-class binary classification function can predict are represented by "helper T cells", "suppressor T cells", "effector T cells", and "memory T cells", and its output result can be represented by a vector. For example, it can be a 4*1-dimensional vector. Each element in this vector corresponds to a cell attribute, and the value of each element in the vector represents the probability that the transcriptome to be predicted is the corresponding cell attribute. Suppose the output vector of the ternary binary classification function is [0.51, 0.15, 0.22, 0.62], then it means that the probability that this transcriptome to be predicted is "helper T cells" is 0.51, the probability that the transcriptome to be predicted is "suppressor T cells" is 0.15, the probability that the transcriptome to be predicted is "effector T cells" is 0.62, and the probability that the transcriptome to be predicted is "memory T cells" is 0.12. Suppose the preset threshold is 0.5, and the element values with probabilities greater than the preset threshold are used as the prediction results of the transcriptome to be predicted, that is, "helper T cells" and "effector T cells" are used as the initial cell annotation results of the multi-class binary classification function.
[0098] In the embodiments of the present application, by inputting the transcriptome to be predicted into the second prediction model for prediction processing, the accuracy of determining the cell annotation result is greatly improved, and the initial cell annotation result can be obtained more accurately, providing accurate guiding information for the determination of the final cell annotation result and realizing higher-precision cell data annotation.
[0099] In another embodiment of the present application, a specific implementation manner for correcting the initial cell annotation result based on the first prediction model is also provided. For example, Figure 5 A specific implementation manner for obtaining the cell annotation result of the transcriptome to be predicted is provided. Please refer to Figure 5 As shown, the method includes:
[0100] S301. Input the cell data into the first prediction model for encoding processing to obtain fused encoding information, and input the fused encoding information into the classification module of the first prediction model to obtain the output result of the first prediction model.
[0101] Specifically, the cell data of the transcriptome to be annotated may include the gene expression information of multiple sequencing points in the transcriptome to be predicted and the spatial information of multiple sequencing points, where the gene expression information and the spatial information may be matrices or vectors. Please refer to Figure 6 As shown, inputting the transcriptome to be predicted into the first prediction model for encoding processing may include: inputting the cell data into the encoding module 410 in the first prediction model to perform information fusion and encoding based on the gene expression information and the spatial information, so as to extract fused feature information and obtain fused encoding information. The representation form of the fused encoding information may be in the form of a matrix or a vector, where the encoding module may include a first encoding module and a second encoding module.
[0102] Input the fusion coding information into the classification module 420 of the first prediction model to obtain the output result 430 of the first prediction model. This classification module is responsible for establishing the relationship between the fusion feature information and the target cell type. Among them, this classification module may include, but is not limited to, a fully connected layer and an activation function. The fully connected layer may include one layer, or may also include multiple layers. The fully connected layer is mainly used to classify the fusion feature information. The fusion feature information can be processed through the fully connected layer to obtain a fully connected vector, and then the fully connected vector is processed through the activation function to obtain the output result of the first prediction model, and this output result can be the cell type.
[0103] Among them, the above activation function can be a Sigmoid function, or a Tanh function, or a ReLU function. By processing the fully connected vector through the activation function, its result can be mapped between 0 and 1.
[0104] As an alternative implementation, the specific processing flow of the above coding module is provided. Please refer to Figure 7 As shown, in the process of inputting cell data into the first prediction model for coding processing, the gene expression information can be input into the first coding module 510 of the first prediction model to obtain the first coding information, and then the adjacency matrix is determined according to the spatial information. The adjacency matrix is a matrix used to characterize the adjacent relationship between each sequencing point, including the adjacent sequencing points of each sequencing point in the transcriptome to be predicted. And input the adjacency matrix and the first coding information into the second coding module 520 of the first prediction model for processing to obtain the second coding information. Finally, the first coding information and the second coding information are fused to obtain the fusion coding information. Then input the fusion coding information into the classification module 530 for classification processing to obtain the output result 540.
[0105] Among them, the above first coding module is used to perform feature coding on the gene expression information, and it can be a deep autoencoder DAE. The second coding module is used to perform feature coding on the spatial information. The first coding information refers to the coding information obtained by performing feature coding on the gene expression information, and it can be a variational graph autoencoder VGAE. The second coding information refers to the coding information obtained by coding the gene expression information and the spatial information. The above first coding information and second coding information can be represented by vectors or matrices.
[0106] It should be noted that in the process of fusing the first coding information and the second coding information, when the first coding information and the second coding information are represented by vectors, information fusion can be performed by means of vector combination; when the first coding information and the second coding information are represented by matrices, information fusion can be performed by means of matrix splicing, so as to obtain the corresponding fused coding information.
[0107] S302. Determine the loss between the output result and the initial cell annotation result according to the loss function of the first prediction model, iteratively train the classification module according to the loss, and obtain the adjusted first prediction model when the loss function is minimized.
[0108] Optionally, the loss function of the first prediction model described above may be a classification loss, for example, the loss between the initial cell annotation result and the output of the first prediction model; it may also be a reconstruction loss, for example, the loss between the reconstructed gene expression information and the gene expression information input to the first prediction model, and may also be the loss between the reconstructed spatial information and the spatial information input to the first prediction model; it may also include a classification loss and a reconstruction loss, including the sum of one or two of the reconstruction losses and the classification loss.
[0109] After obtaining the output result, the loss between the output result and the initial cell annotation result can be determined according to the loss function, the classification module can be iteratively trained according to the loss, and the adjusted first prediction model can be obtained when the loss function is minimized. For example, the gradient descent method can be used for iterative training. Among them, when iteratively training the first prediction model according to the loss, the parameters in the classification module can be updated, for example, the matrix parameters such as the weight matrix and the bias matrix in the classification module can be updated. Among them, the above-mentioned weight matrix and bias matrix include, but are not limited to, the matrix parameters in the self-attention layer, the feed-forward network layer, and the fully connected layer of the classification module.
[0110] Among them, when updating the parameters in the first prediction model through the loss function, it can be determined that the first prediction model has not converged according to the loss function, and the parameters in the first prediction model can be adjusted, for example, the parameters in the classification module can be adjusted, so that the first prediction model converges, so as to obtain the adjusted first prediction model. Among them, the convergence of the first prediction model may mean that the output result of the first prediction model is less than a preset threshold compared with the initial cell annotation result, or the change rate of the difference between the output result and the initial cell annotation result approaches a certain lower value. When the calculated loss function is small, or the difference from the loss function output in the previous round of iteration approaches 0, it is considered that the classification module converges, and then it is considered that the first prediction model converges.
[0111] S303. Input the cell data into the adjusted first prediction model to obtain the cell annotation result of the transcriptome to be predicted.
[0112] It should be noted that the above-adjusted first prediction model can be a model obtained through iterative training of a training algorithm, that is, the model parameters of the adjusted first prediction model are in an optimal state. Among them, the training algorithm can be a self-supervised learning algorithm. Optionally, the adjusted first prediction model can include an adjusted encoding module and a classification module. Among them, the adjusted encoding module includes a third encoding module and a fourth encoding module. The third encoding module is used to perform feature encoding on gene expression information, and the fourth encoding module is used to perform feature encoding on spatial information.
[0113] Specifically, after obtaining the adjusted first prediction model, the cell data of the transcriptome to be predicted can be input into the adjusted first prediction model. The cell data includes the gene expression information and spatial information of the transcriptome to be predicted. The gene expression information can be input into the third encoding module in the adjusted first prediction model to obtain third encoding information. Then, an adjacency matrix is determined according to the spatial information, and the adjacency matrix and the third encoding information are input into the fourth encoding module of the adjusted first prediction model to obtain fourth encoding information, and the third encoding information and the fourth encoding information are fused to obtain combined encoding information. The combined encoding information is input into the classification module of the adjusted first prediction model for classification processing to obtain the cell annotation result of the transcriptome to be predicted.
[0114] In this embodiment, by inputting cell data into the first prediction model for encoding processing and classification processing to obtain an output result, and iteratively training to obtain the adjusted first prediction model according to the loss between the output result and the initial annotation result, it is possible to effectively extract large gene expression information and spatial information through self-supervised training, making the constructed adjusted first prediction model better, and further making the accuracy of the obtained cell annotation result higher.
[0115] In another embodiment of the present application, the first decoding module in the first prediction model can also be iteratively trained. Specifically, after obtaining the fused encoding information, the fused encoding information can also be decoded to obtain the reconstructed gene expression information, and the loss between the reconstructed gene expression information and the gene expression information is determined according to the loss function, and the parameters of the first encoding module are adjusted according to the loss to make the reconstructed gene expression information close to the gene expression information.
[0116] Specifically, please refer to Figure 8As shown, during the process of decoding the fused encoded information, it can first perform linear feature extraction on the fused encoded information to obtain first intermediate feature information, then perform feature extraction through a linear layer, and then input the first intermediate feature information into the first decoding module 610 for feature restoration to obtain the reconstructed gene expression information. Among them, the number of the above linear layers can be two or more. The more the number of linear layers, the more accurate the feature information obtained after feature extraction from the fused encoded information. The first intermediate feature information can be the linearly processed gene expression information.
[0117] It should be noted that the role of the above first decoding module is to accurately restore the reconstructed gene expression information back to the original input gene expression information.
[0118] Furthermore, after obtaining the reconstructed gene expression information, the loss between the reconstructed gene expression information and the gene expression information can be determined according to the loss function, and the first encoding module can be iteratively trained according to the loss. When the loss function is minimized, an adjusted first prediction model can be obtained. For example, the gradient descent method can be used for iterative training.
[0119] Among them, when iteratively training the first prediction model according to the loss, the parameters in the first encoding module can be updated. For example, the matrix parameters such as the weight matrix and bias matrix in the first encoding module can be updated. Among them, the above weight matrix and bias matrix include but are not limited to the matrix parameters in the input layer and hidden layer of the first encoding module.
[0120] Among them, when updating the parameters in the prediction module through the loss function, when it is determined that the prediction module has not converged according to the loss function, the parameters in the first encoding module can be adjusted. That is, by adjusting the parameters in the first encoding module, the first prediction model can converge, so as to obtain an adjusted first prediction model. The convergence of the first prediction model can refer to that the reconstructed gene expression information is less than a preset threshold compared with the gene expression information input to the first prediction model, or the change rate of the difference between the reconstructed gene expression information and the gene expression information input to the first prediction model approaches a certain low value. When the calculated loss function is small, or the difference from the loss function output in the previous round of iteration approaches 0, it is considered that the first encoding module converges, and further that the first prediction model converges.
[0121] In another embodiment of the present application, the second decoding module in the first prediction model can also be iteratively trained. Specifically, after obtaining the fused encoded information, the fused encoded information can also be decoded to obtain the reconstructed spatial information, and the loss between the reconstructed spatial information and the spatial information can be determined according to the loss function, and the parameters of the second encoding module can be adjusted according to the loss to make the reconstructed spatial information close to the spatial information.
[0122] Specifically, during the process of decoding the fused encoded information, the fused encoded information can be first subjected to linear feature extraction to obtain corresponding second intermediate feature information, and then feature extraction is performed through a linear layer. Then, the second intermediate feature information is input into the second decoding module 620 for feature restoration to obtain the reconstructed spatial information. Among them, the number of the above linear layers can be two or more. The more the number of linear layers, the more accurate the feature information obtained by feature extraction from the fused encoded information. The second intermediate feature information can be the spatially information after linear processing.
[0123] It should be noted that the role of the second decoding module is to accurately restore the reconstructed spatial information back to the original input spatial information.
[0124] Further, after obtaining the reconstructed spatial information, the loss between the reconstructed spatial information and the spatial information can be determined according to the loss function, and the second encoding module is iteratively trained according to the loss. When the loss function is minimized, an adjusted first prediction model is obtained. For example, the gradient descent method can be used for iterative training.
[0125] Among them, when iteratively training the first prediction model according to the loss, the parameters in the second encoding module can be updated. For example, the matrix parameters such as the weight matrix and the bias matrix in the second encoding module can be updated. Among them, the above weight matrix and bias matrix include but are not limited to the matrix parameters in the input layer and the hidden layer of the second encoding module.
[0126] Among them, when updating the parameters in the prediction module through the loss function, when it is determined that the prediction module has not converged according to the loss function, the parameters in the second encoding module can be adjusted. That is, the parameters in the second encoding module can be adjusted to make the first prediction model converge, so as to obtain an adjusted first prediction model. The convergence of the first prediction model can mean that the difference between the reconstructed spatial information and the spatial information input to the first prediction model is less than a preset threshold, or the change rate of the difference between the reconstructed spatial information and the spatial information input to the first prediction model approaches a certain lower value. When the calculated loss function is small, or the difference from the loss function output in the previous round of iteration approaches 0, it is considered that the second encoding module converges, and thus it is considered that the first prediction model converges.
[0127] In this embodiment, by decoding the fused encoded information to obtain the reconstructed gene expression information and the reconstructed spatial information, the parameters in the first encoder and the second encoder can be accurately adjusted in combination with the gene expression information, so that the adjusted first prediction model has higher accuracy, and thus the accuracy of determining the cell annotation result is improved.
[0128] It should be noted that the iterative training of the first encoding module and the second encoding module in the embodiments of the present application are two independent processing processes. It is possible to perform only the iterative training of the first encoding module, or only the iterative training of the second encoding module. Of course, it is also possible to perform iterative training on both the first encoding module and the second encoding module. The execution order of the two is not limited. It can be executed in series in one iterative training, or in parallel. Taking parallel execution as an example, after obtaining the fused encoding information, it is also possible to perform decoding processing on the fused encoding information to obtain the reconstructed spatial information and the reconstructed gene expression information, and determine the loss between the reconstructed spatial information and the spatial information according to the loss function, and determine the loss between the reconstructed gene expression information and the gene expression information. According to the loss, the parameters of the first encoding module and the second encoding module are adjusted simultaneously, so that the reconstructed spatial information is close to the spatial information, and the reconstructed gene expression information is close to the gene expression information.
[0129] In another embodiment of the present application, a specific implementation of the loss function of the first prediction model is also provided. In a possible implementation manner, the loss function of the first prediction model includes a first component, a second component, and a third component. The first component is used to characterize the loss between the initial cell annotation result and the output of the first prediction model; the second component is used to characterize the loss between the reconstructed gene expression information and the gene expression information in the input of the first prediction model; the reconstructed gene expression information is obtained after decoding the fused encoding information; the third component is used to characterize the loss between the reconstructed spatial information and the spatial information in the input of the first prediction model; the reconstructed spatial information is obtained after decoding the fused encoding information.
[0130] Among them, the constructed loss function described above can be the sum of the first component, the second component, and the third component. By setting the second component and the third component in the loss function, the difference between the reconstructed gene expression information and the gene expression information input to the first prediction model, and the difference between the reconstructed spatial information and the spatial information input to the first prediction model can be reduced, thereby ensuring that the fused feature information fuses the gene expression information and the spatial information of the transcriptome to be predicted. And by setting the first component, it is possible to provide guiding information for obtaining the final cell type of the fused feature information, so that the adjusted first prediction model obtained is better.
[0131] In the embodiments of the present application, when constructing a loss function to obtain an adjusted first prediction model, the differences between the initial cell annotation results and the outputs of the first prediction model, the differences between the reconstructed gene expression information and the gene expression information input into the first prediction model, and the differences between the reconstructed spatial information and the spatial information in the input of the first prediction model are comprehensively considered. Training the first prediction model based on this loss function can iteratively train the model parameters in the first prediction model more accurately and comprehensively, making the obtained adjusted first prediction model better, and further making the determined cell annotation results more accurate.
[0132] In another embodiment of the present application, during the model training process, reasonable weight coefficients can also be assigned to the first component, the second component, and the third component in the loss function, so that the prediction differences of the model highly match the actual business requirements, which can also bring about an improvement in model performance. In a possible implementation manner, when determining the loss function, the weight coefficients of the first component, the second component, and the third component can be determined, and the loss function can be determined according to the weight coefficients of the first component, the second component, and the third component, the first component, the second component, and the third component. Among them, the weight coefficient of the first component is related to the importance of the transcriptomic cell annotation results; the weight coefficient of the second component is related to the importance of the gene expression information; the weight coefficient of the second component is related to the importance of the transcriptomic spatial information.
[0133] Among them, the weight coefficient of the first component is positively correlated with the importance of the transcriptomic cell annotation results, that is, the larger the weight coefficient of the first component, the higher the importance of the transcriptomic cell annotation results. Similarly, the weight coefficient of the second component is positively correlated with the importance of the transcriptomic gene expression information, that is, the larger the weight coefficient of the second component, the higher the importance of the transcriptomic gene expression information. The weight coefficient of the third component is positively correlated with the importance of the transcriptomic spatial information, that is, the larger the weight coefficient of the third component, the higher the importance of the transcriptomic spatial information.
[0134] The above first component, second component, third component, and the loss function satisfy the following formula:
[0135] Y = a1 * y1 + a2 * y2 + a3 * y3
[0136] Wherein, Y is the loss function during the training process of the above first prediction model, y1 is the first component, a1 is the weight coefficient of the first component, y2 is the second component, a2 is the weight coefficient of the second component, y3 is the third component, and a3 is the weight coefficient of the third component.
[0137] In the embodiments of the present application, the weight coefficients of the first component, the second component, and the third component can also be reasonably determined according to the business scenarios to which the cell data annotation method is applied. In one possible implementation, the importance ratio of the transcriptome spatial information, the transcriptome gene expression information, and the transcriptome cell annotation result is determined according to the business requirements, and the weight coefficients of the first component, the second component, and the third component are determined according to the importance ratio.
[0138] Exemplarily, when the cell data annotation method provided in the embodiments of the present application is applied to certain scenarios, more attention is paid to the transcriptome cell annotation result. For example, the importance ratio of the transcriptome cell annotation result, the transcriptome gene expression information, and the transcriptome spatial information is 80%:10%:10%. Then, the weight coefficient of the first component can be 0.8, the weight coefficient of the second component can be 0.1, and the weight coefficient of the third component can be 0.1.
[0139] In another embodiment of the present application, a specific implementation manner for determining the adjacency matrix according to the spatial information of the transcriptome to be predicted is also provided. The neighboring sequencing points of each sequencing point among multiple sequencing points can be determined according to the spatial information, and then the adjacency matrix is generated according to the spatial distance between the neighboring sequencing points and the corresponding sequencing points. The adjacency matrix is an N*N matrix, where N is the number of sequencing points in the transcriptome to be predicted.
[0140] Specifically, the above-mentioned spatial information is used to represent the position information of each sequencing point in the slice, and the spatial information can be represented by a matrix. For example, it can be a two-dimensional matrix or a three-dimensional matrix. When the spatial information is a two-dimensional matrix, it can be an N*2-dimensional coordinate matrix, where the element values in each row of the matrix represent the abscissa value and the ordinate value of each sequencing point, and the element values in each column represent multiple sequencing points. When the spatial information is a three-dimensional matrix, it can be an N*3-dimensional coordinate matrix. It should be noted that the number of the above-mentioned neighboring sequencing points can be multiple.
[0141] Among them, when determining the neighboring sequencing points of each sequencing point among multiple sequencing points according to the spatial information, a corresponding empty matrix can be constructed based on the spatial information first. For example, if the spatial information is an N*2-dimensional coordinate matrix, the constructed empty matrix is an N*N-dimensional coordinate matrix, where N is the number of sequencing points. Then, the neighborhood range of each sequencing point is determined in the empty matrix, and the sequencing points within the neighborhood range are determined as the neighboring sequencing points of the sequencing point. It can be in the order from top to bottom in the empty matrix. Starting from the first sequencing point, the neighboring sequencing points corresponding to the first sequencing point are determined first, and then the neighboring sequencing points corresponding to the second sequencing point are determined, and so on, until the neighboring sequencing points corresponding to the last sequencing point are determined, so as to determine the neighboring sequencing points corresponding to each sequencing point. The neighborhood range can be determined according to the position information of the sequencing point according to a preset rule. For example, it is an area centered on the position information of the sequencing point and with a preset distance as the radius.
[0142] As an implementable method, when determining the neighboring sequencing points of each sequencing point among multiple sequencing points according to the spatial information, for each sequencing point among the multiple sequencing points, the position information of the sequencing point can be determined according to the spatial information, and the distance between the sequencing point and the other sequencing points can be calculated according to the position information of the sequencing point. Then, the sequencing points with a distance less than a preset threshold from the sequencing point are determined as the neighboring sequencing points of the sequencing point. Among them, the preset threshold can be custom-set according to actual needs.
[0143] Furthermore, for each sequencing point among the multiple sequencing points, after determining the corresponding neighboring sequencing points, the spatial distance between the neighboring sequencing points and the corresponding sequencing point can be determined. For example, the spatial distance can be obtained by calling a distance calculation function. Then, according to the spatial distance between the neighboring sequencing points and the corresponding sequencing point, the spatial distance is normalized to obtain the elements of the sequencing point and other sequencing points. These elements are used to represent the neighboring relationship between the sequencing point and other sequencing points, so as to construct an adjacency matrix.
[0144] Specifically, it can be in the order from top to bottom in the empty matrix. Starting from the first sequencing point, the spatial distance between the first sequencing point and the corresponding neighboring point is determined first, and then the spatial distance is normalized to obtain a processing result. This processing result is used as the element at the corresponding position in the empty matrix between the first sequencing point and its neighboring sequencing points. Then, the spatial distance between the second sequencing point and the corresponding neighboring point is determined and normalized to obtain the element at the corresponding position in the empty matrix between the second sequencing point and its neighboring sequencing points, and so on, until the element at the corresponding position in the empty matrix between the last sequencing point and its neighboring sequencing points is determined, so as to obtain the adjacency matrix.
[0145] As an implementable manner, for each sequencing point among multiple sequencing points, the element in the adjacency matrix corresponding to the sequencing point can be determined according to the spatial distance corresponding to the neighboring sequencing points of the sequencing point. By determining the K elements in the i-th row or the i-th column of the adjacency matrix according to the K neighboring sequencing points corresponding to the i-th sequencing point among the multiple sequencing points. Then set the values of the remaining elements in the adjacency matrix to zero.
[0146] Specifically, after determining the neighboring sequencing points corresponding to each sequencing point, for each sequencing point, the position information of the neighboring sequencing points can be determined, and the spatial distance between the two can be calculated according to the position information of the sequencing point and the neighboring sequencing points. Among them, when the spatial information is a two-dimensional matrix and the constructed empty matrix is an N*N-dimensional matrix, the position information of any sequencing point can be represented by the abscissa value and the ordinate value. When determining the spatial distance between the neighboring sequencing point and the corresponding sequencing point, it can be by calculating the Euclidean distance between the sequencing point and each neighboring sequencing point, and taking this Euclidean distance as the spatial distance. Then take the opposite of the Euclidean distance and normalize it to a value within the range of [0,1], and use this value as the element in the adjacency matrix corresponding to the sequencing point, and set the remaining elements in the adjacency matrix to zero.
[0147] Exemplarily, when the spatial information is an N*2-dimensional matrix, an empty N*N-dimensional matrix can be established, where N represents the number of sequencing points, that is, the number of cells. Each row in this matrix represents a sequencing point or a cell. For each sequencing point i among the multiple sequencing points, the position information of sequencing point i is determined according to the spatial information. This position information can include the abscissa value and the ordinate value. Then calculate the distance between sequencing point i and the remaining sequencing points according to the position information of sequencing point i, and determine the sequencing points whose distance from sequencing point i is less than a preset threshold as the neighboring sequencing points of sequencing point i. For example, K sequencing points j are determined as neighboring sequencing points, including j1, j2, …, j K ,K. Then for each neighboring sequencing point among the multiple neighboring sequencing points, calculate the Euclidean distance between sequencing point i and neighboring sequencing point j, and calculate the opposite of the Euclidean distance and normalize it to a value within the range of [0,1], so as to obtain the element value corresponding to the position (i,j) in the empty matrix. For the points not in the neighborhood, that is, the element values at the corresponding positions of the non-neighboring sequencing points in the adjacency matrix are set to zero, thus generating a distance-weighted adjacency matrix. Denote this adjacency matrix as S, which is an N-order matrix. It should be noted that all elements in this adjacency matrix are within the range of [0,1]. Each row of the adjacency matrix represents the neighborhood relationship of each sequencing point or cell. There are K non-zero elements in each row. Being zero means that the cell or sequencing point at this position is not adjacent to the cell or sequencing point represented by this row. When the element is non-zero, the larger the value, the closer the cell or sequencing point at this position is to the cell or sequencing point of this row.
[0148] To better understand the embodiments of the present application, the following further describes the complete flowchart method of the cell data annotation method proposed by the present application.
[0149] As Figure 9 shown, the method may include the following steps:
[0150] S401. Obtain the cell data of the transcriptome to be predicted; the cell data includes the gene expression information of multiple sequencing points in the transcriptome to be predicted and the spatial information of multiple sequencing points.
[0151] Specifically, please refer to Figure 10 shown, the cell data (ST dataset) of the above-mentioned transcriptome to be predicted includes the gene expression information of multiple sequencing points and the spatial information of multiple sequencing points. Among them, the gene expression information of each sequencing point can be represented as an N×G-dimensional matrix, where N is the number of sequencing points and G is the number of genes detected. The element in the i-th row and j-th column of the matrix represents the expression level of gene j at sequencing point i. The spatial information of each sequencing point among the above-mentioned multiple sequencing points refers to the coordinate information of each sequencing point in the slice, and this coordinate information can be two-dimensional data or three-dimensional data. Taking two-dimensional data as an example, the coordinate information can be an N×2-dimensional matrix, where N is the number of sequencing points. Among them, the i-th row in the matrix represents the position coordinates of sequencing point i in the slice, which can include two abscissa values and an ordinate value. Among them, the matrix corresponding to the gene expression information of the transcriptome to be predicted can be denoted as X.
[0152] S402. Determine the annotated cell object corresponding to the transcriptome to be predicted and the second prediction model matching the cell object.
[0153] S403. Input the gene expression information into the second prediction model to obtain the initial cell annotation result of the transcriptome to be predicted.
[0154] Specifically, the cell object corresponding to the transcriptome to be predicted refers to a cell object whose tissue or region is the same as that of the transcriptome to be predicted. For example, if the transcriptome to be predicted is a brain slice taken from brain tissue, the cell object corresponding to this transcriptome to be predicted is also taken from brain tissue. The above-mentioned cell object can be a single cell or a cell group.
[0155] Taking the cell object corresponding to the transcriptome to be predicted as single-cell data and the second prediction model matching the single cell as a DNN model as an example, the single-cell data (Single-cell dataset) can include two parts. The first part is the gene expression information of the single cell, which is an M*N-dimensional matrix, where N is the number of sequencing points or cells, and G is the number of genes detected. The second part is the cell annotation result corresponding to each single cell, that is, the cell type, which is an M-dimensional vector. Among them, the matrix corresponding to the gene expression information of the single cell can be denoted as A, and its corresponding cell type can be denoted as Y.
[0156] Among them, in the single-cell migration stage, a DNN model can be obtained through training using a single-cell dataset and the corresponding cell annotation results. The single-cell dataset is randomly divided into a training set and a validation set according to a certain ratio, and then the DNN model is constructed using the training set and the validation set according to the training learning algorithm.
[0157] Specifically, after the DNN model is constructed, the cell data (Single-cell dataset) of the transcriptome to be predicted is input into the DNN model for prediction (Inference) processing. It can be by extracting the feature information of the transcriptome to be predicted, and then performing operations on the feature information to obtain the initial cell annotation result of the transcriptome to be predicted, that is, the cell type. This cell type can be represented by the probability of each cell type (class probabilities for each cells of ST dataset) in the transcriptome to be predicted. The initial cell annotation result can be represented by an N*C-dimensional matrix, where N is the number of sequencing points or cells in the transcriptome to be predicted, and C is the number of cell types. This initial cell annotation result can be denoted as L.
[0158] It can be understood that each row in the above N*C-dimensional matrix corresponds to a sequencing point or cell, and each column corresponds to a cell type. If the value of a certain row and a certain column is large, it means that the cell corresponding to that row may belong to the cell type corresponding to that column. Among them, the initial cell annotation result L can be used as the pseudo-label of the first prediction model.
[0159] S404. Input the cell data into the first prediction model for encoding processing to obtain fusion encoding information, and input the fusion encoding information into the classification module of the first prediction model to obtain the output result of the first prediction model.
[0160] S405. Determine the loss between the output result and the initial cell annotation result according to the loss function of the first prediction model, and iteratively train the classification module according to the loss to obtain the adjusted first prediction model when the loss function is minimized.
[0161] It should be noted that the above three parts: the matrix A corresponding to the gene expression information of the single cell, the corresponding cell type Y, and the matrix X corresponding to the gene expression information of the transcriptome to be predicted can be directly used by the corresponding first prediction model. However, the spatial information of the transcriptome to be predicted, that is, the N*2-dimensional coordinate matrix, needs to be converted into a corresponding distance-weighted adjacency matrix, and this adjacency matrix is used as a part of the input of the first prediction model.
[0162] In this embodiment, the spatial information of the transcriptome to be predicted is an N×2-dimensional coordinate matrix. An N×N-dimensional empty matrix can be established first, where each row of the empty matrix represents a sequencing point. Then, for each sequencing point i, K sequencing points j1, j2, …, j within its preset distance range are taken K as the neighboring sequencing points of the sequencing point i. Then, for each neighboring sequencing point among the multiple neighboring sequencing points, the Euclidean distance between the sequencing point i and the neighboring sequencing point j is calculated, and the negative value of the Euclidean distance is calculated and normalized to a value within the range of [0, 1], so as to obtain the element value corresponding to the position (i, j) in the empty matrix. For points not within the neighborhood, that is, the element values corresponding to the positions of non-neighboring sequencing points in the adjacency matrix are set to zero, thereby generating a distance-weighted adjacency matrix, which is denoted as S and is an N-order matrix.
[0163] Please refer to Figure 11 as shown. The first prediction model can be composed of two pairs of encoding-decoding modules and a classification module. Among them, one encoding-decoding module can be a DAE, which is used to obtain feature encoding of gene expression information to obtain the first encoded information. The DAE includes a Deep Encoder and a Deep Decoder; the other encoding-decoding module can be a VGAE, which is used to obtain feature encoding of spatial information to obtain the second encoded information. The VGAE includes a Graph Encoder and a Graph Decoder; the above classification module is responsible for establishing the correspondence between the complete feature encoding and the target cell type, and can include a Cluster / Classifier.
[0164] The matrix X corresponding to the gene expression information of the transcriptome to be predicted and the adjacency matrix S are input into the two encoding-decoding modules and a classification module in the first prediction model. The matrix X corresponding to the gene expression information of the transcriptome to be predicted is encoded by the first encoding module (Deep Encoder) of the DAE to obtain the first encoded information E X , and the adjacency matrix S and the first encoded information E X are encoded by the second encoding module (Graph Encoder) in the VGAE to obtain the second encoded information E g . Then, the first encoded information E X and the second encoded information E g are subjected to fusion encoding processing, which can be to obtain the fusion encoded information E through splicing processing. It can be understood that the above fusion encoded information E contains both the matrix X corresponding to the gene expression information in the transcriptome to be predicted and the information encoded by the adjacency matrix S corresponding to the spatial information.
[0165] It should be noted that the ultimate goal of training the first prediction model is to enable the fusion feature encoding E to effectively encode all the information in the transcriptome to be predicted and be accurately classified by the classification module to obtain the corresponding cell type. To this end, two corresponding decoding modules are introduced to improve the robustness of the fusion feature encoding E in a self-supervised manner. Among them, self-supervised means using its own information as a label and making the result obtained after decoding as consistent as possible with the result input before encoding through an encoding-decoding form.
[0166] Furthermore, the fusion encoded information E can be processed through two linear layers for feature extraction to obtain the first intermediate feature information and the second intermediate feature information. Then, the first intermediate feature information is subjected to feature reduction processing through the first decoding module (Deep Decoder) of the DAE to obtain the reconstructed gene expression information X'. The reconstructed gene expression information is represented by an N*G-dimensional matrix, where N is the number of sequencing points or cells, and G is the number of genes measured. In parallel, the second intermediate feature information is subjected to feature reduction processing through the second decoding module (Graph Decoder) of the VGAE to obtain the reconstructed spatial information S'. The reconstructed spatial information is represented by an N*N-dimensional matrix, where N is the number of sequencing points or cells. The fusion encoded information E is classified through the classification module (Cluster / Classifier) to generate the output result L' of the first prediction model. The output result can be represented by an N*C-dimensional matrix, where N is the number of sequencing points or cells in the transcriptome to be predicted, and C is the number of cell types.
[0167] Further, determine the loss between the output result and the initial cell annotation result L according to the loss function of the first prediction model, and iteratively train the first prediction model according to the loss. Obtain the adjusted first prediction model when the loss function is minimized. The above loss function includes a first component, a second component, and a third component. The first component is the classification loss between the initial cell annotation result L and the output L' of the first prediction model. The second component is the reconstruction loss between the reconstructed gene expression information X' and the gene expression information X in the input of the first prediction model. The third component is the reconstruction loss between the reconstructed spatial information S' and the adjacency matrix S corresponding to the spatial information in the input of the first prediction model. Optionally, during the model training process, reasonable weight coefficients can also be assigned to the first component, the second component, and the third component in the loss function. It can be determined that the weight coefficient of the first component is 0.8, the weight coefficient of the second component can be 0.1, and the weight coefficient of the third component can be 0.1. Then, the loss function can be determined according to the first component and its weight coefficient, the second component and its weight coefficient, and the third component and its weight coefficient. Then, iteratively train the first prediction model according to the minimization of the loss function. It can be to adjust the parameters in the first coding module, the second coding module, and the classification module in the first prediction model by using the gradient descent method, so as to obtain the adjusted first prediction model.
[0168] In this embodiment, by setting the first component, the second component, and the third component to construct the loss function, the model parameters in the first prediction model can be accurately adjusted by optimizing the three-part loss function, so that the adjusted first prediction model is better, and further, more accurate annotation of the cell data of the transcriptome to be predicted can be achieved.
[0169] S406. Input the cell data into the adjusted first prediction model to obtain the cell annotation result of the transcriptome to be predicted.
[0170] Specifically, please refer to Figure 12As shown in the figure, the first prediction model after the above adjustment includes an adjusted encoding module and a classification module Cls. Among them, the adjusted encoding module includes a third encoding module (D-Enc) and a fourth encoding module (G-Enc). After obtaining the first prediction model after adjustment, the cell data ST dataset can be input into the first prediction model after adjustment, and the gene expression information X can be input into the third encoding module (D-Enc) in the first prediction model after adjustment to obtain the third encoding information. Then, the adjacency matrix S is determined according to the spatial information. The adjacency matrix S is an N-order matrix, and the adjacency matrix S and the third encoding information are input into the fourth encoding module (G-Enc) of the first prediction model after adjustment to obtain the fourth encoding information. The third encoding information and the fourth encoding information are fused to obtain the combined encoding information. The combined encoding information is input into the classification module (Cls) of the first prediction model after adjustment for classification processing to obtain the cell annotation result of the transcriptome to be predicted. The cell annotation result may include the cell type (Cell type annotation). The cell annotation result can be represented by an N*C-dimensional matrix, where N is the number of sequencing points or cells of the transcriptome to be predicted, and C is the number of cell types.
[0171] In addition, in order to test the performance of the solution in this application, the accuracy metrics of different solutions such as Seurat, scmap, scNym, and sciBet for the two spatial transcriptome data of MERFISH and Slide-seq are calculated respectively. Among them, Seurat, scmap, scNym, and sciBet represent the solutions obtained in the prior art. Seurat and scmap represent the annotation solutions based on the reference dataset. This solution means that the dataset to be predicted is first clustered to obtain the clustering result (i.e., divided into multiple clusters), and then the cell annotation result corresponding to each cluster is determined according to the reference dataset; scNym and sciBet represent the methods based on classification training. The accuracy metrics include the following data:
[0172] This solution Seurat scmap scNym sciBet MERFISH 92.21% 86.80% 72.52% 87.72% 90.43% Slide-seq 60.76% 57.76% 20.71% 52.69% 47.87%
[0173] The above data shows that the performance metric of this solution for annotating the cell data of the spatial transcriptome is the best, that is, by integrating the gene expression information and spatial information in the spatial transcriptome, all the information in the cell number data of the spatial transcriptome can be used more accurately and comprehensively, thereby improving the accuracy of determining the cell annotation result.
[0174] In this embodiment, the gene expression information of multiple sequencing points and the spatial information of multiple sequencing points can be effectively integrated, and self-supervised correction is performed in combination with the guidance information, so as to more accurately and comprehensively utilize all the information in the cell data of the transcriptome to be predicted, and then more accurately obtain the cell annotation result of the transcriptome to be predicted, greatly improving the accuracy of the cell data annotation of the transcriptome to be predicted.
[0175] On the other hand, Figure 13 FIG. is a schematic structural diagram of a cell data annotation device provided by an embodiment of the present application. This device can be a device in a terminal or a server, such as Figure 13 As shown, the device 700 includes:
[0176] An acquisition module 710, configured to acquire cell data of a transcriptome to be predicted; the cell data includes gene expression information of multiple sequencing points in the transcriptome to be predicted and spatial information of multiple sequencing points;
[0177] A processing module 720, configured to determine an annotated cell object corresponding to the transcriptome to be predicted, and determine an initial cell annotation result of the transcriptome to be predicted according to the cell object and the gene expression information;
[0178] A cell annotation module 730, configured to input the cell data into a first prediction model, and correct the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result, so as to obtain the cell annotation result of the transcriptome to be predicted.
[0179] In some embodiments, please refer to Figure 14 As shown, the above-mentioned processing module 720 includes:
[0180] A first determination unit 721, configured to determine a second prediction model matching the cell object; the second prediction model is trained based on the historical gene expression information of the cell object and the historical cell annotation result corresponding to the historical gene expression information;
[0181] A first processing unit 722, configured to input the gene expression information into the second prediction model to obtain an initial cell annotation result of the transcriptome to be predicted.
[0182] In some embodiments, the cell object is a single cell or a cell group included in the gene sequencing object corresponding to the transcriptome to be predicted.
[0183] In some embodiments, the above-mentioned cell annotation module 730 includes:
[0184] A second processing unit 731, configured to perform encoding processing on the cell data input into the first prediction model to obtain fusion encoding information, and input the fusion encoding information into the classification module of the first prediction model to obtain the output result of the first prediction model;
[0185] A training unit 732, configured to determine the loss between the output result and the initial cell annotation result according to the loss function of the first prediction model, iteratively train the classification module according to the loss, and obtain an adjusted first prediction model when the loss function is minimized;
[0186] A second determination unit 733, configured to input the cell data into the adjusted first prediction model to obtain the cell annotation result of the transcriptome to be predicted.
[0187] In some embodiments, the second processing unit 731 is specifically configured to:
[0188] Input the gene expression information into the first encoding module of the first prediction model to obtain first encoded information;
[0189] Determine an adjacency matrix according to the spatial information; the adjacency matrix is used to characterize the neighboring sequencing points of each sequencing point in the transcriptome to be predicted;
[0190] Input the adjacency matrix and the first encoded information into the second encoding module of the first prediction model for processing to obtain second encoded information;
[0191] Perform a fusion process on the first encoded information and the second encoded information to obtain fused encoded information.
[0192] In some embodiments, the second processing unit 731 is further configured to:
[0193] Perform a decoding process on the fused encoded information to obtain reconstructed gene expression information;
[0194] Determine the loss between the reconstructed gene expression information and the gene expression information according to the loss function, and adjust the parameters of the first encoding module according to the loss so that the reconstructed gene expression information is close to the gene expression information.
[0195] In some embodiments, the second processing unit 731 is further configured to:
[0196] Perform a linear feature extraction process on the fused encoded information to obtain intermediate feature information;
[0197] Input the intermediate feature information into the first decoder for feature reduction processing to obtain reconstructed gene expression information.
[0198] In some embodiments, the second processing unit 731 is further configured to:
[0199] Perform a decoding process on the fused encoded information to obtain reconstructed spatial information;
[0200] Determine the loss between the reconstructed spatial information and the spatial information according to the loss function, and adjust the parameters of the second encoding module according to the loss, so that the reconstructed spatial information is close to the spatial information.
[0201] In some embodiments, the loss function includes a first component, a second component, and a third component;
[0202] The first component is used to characterize the loss between the initial cell annotation result and the output of the first prediction model;
[0203] The second component is used to characterize the loss between the reconstructed gene expression information and the gene expression information in the input of the first prediction model; the reconstructed gene expression information is obtained after decoding the fused coding information;
[0204] The third component is used to characterize the loss between the reconstructed spatial information and the spatial information in the input of the first prediction model; the reconstructed spatial information is obtained after decoding the fused coding information.
[0205] In some embodiments, the above device is further configured to:
[0206] Determine the weight coefficients of the first component, the second component, and the third component;
[0207] Determine the loss function according to the weight coefficients of the first component, the second component, the third component, the first component, the second component, and the third component;
[0208] Among them, the weight coefficient of the first component is related to the importance degree of the transcriptome cell annotation result; the weight coefficient of the second component is related to the importance degree of the transcriptome gene expression information; the weight coefficient of the third component is related to the importance degree of the transcriptome spatial information.
[0209] In some embodiments, the above second processing unit 731 is further configured to:
[0210] Determine the neighboring sequencing points of each sequencing point among the multiple sequencing points according to the spatial information;
[0211] Generate an adjacency matrix according to the spatial distances corresponding to the neighboring sequencing points of all sequencing points in the transcriptome to be predicted; the adjacency matrix is an N*N matrix, and N is the number of sequencing points in the transcriptome to be predicted.
[0212] In some embodiments, the above second processing unit 731 is further configured to:
[0213] For each sequencing point among the multiple sequencing points, determine the position information of the sequencing point according to the spatial information, and calculate the distance between the sequencing point and the remaining sequencing points in the transcriptome to be predicted according to the position information of the sequencing point;
[0214] Sequencing points with a distance less than a preset threshold from a sequencing point are determined as neighboring sequencing points of the sequencing point.
[0215] In some embodiments, the above-mentioned second processing unit 731 is further configured to:
[0216] For each sequencing point among multiple sequencing points, determine the element in the adjacency matrix corresponding to the sequencing point according to the spatial distance between the sequencing point and the neighboring sequencing points;
[0217] Set the remaining elements in the adjacency matrix to zero.
[0218] It can be understood that the functions of the functional modules of the cell data annotation device in this embodiment can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the relevant descriptions of the above method embodiments, which will not be elaborated here.
[0219] In summary, the cell data annotation device provided by the embodiments of the present application determines the annotated cell objects corresponding to the transcriptome to be predicted, and determines the initial cell annotation result of the transcriptome to be predicted according to the cell objects and gene expression information. Without any manual participation in annotation, it can obtain the guiding information of the cell annotation result of the transcriptome to be predicted, and predicts the cell data through the first prediction model, and corrects the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result, effectively integrating the gene expression information of multiple sequencing points and the spatial information of multiple sequencing points, and being able to perform self-supervised correction in combination with the guiding information, so as to more accurately and comprehensively utilize all the information in the cell data of the transcriptome to be predicted, and thus obtain the cell annotation result of the transcriptome to be predicted more accurately, greatly improving the accuracy of the cell data annotation of the transcriptome to be predicted. It can also be applied to a sequencing analysis system to accurately predict the cell data of the transcriptome to be predicted, greatly improving the quality and efficiency of cell annotation, and providing strong support for the data analysis of spatial transcriptomics.
[0220] On the other hand, the device provided by the embodiments of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the cell data annotation method as described above.
[0221] Next, refer to Figure 15 , Figure 15 which is a schematic structural diagram of the computer system of the terminal device according to the embodiments of the present application.
[0222] As Figure 15As shown, the computer system 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage section 303 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the system 300 are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0223] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the storage section 308 as needed.
[0224] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 303, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the above functions defined in the system of the present application are executed.
[0225] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0226] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0227] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. For example, it can be described as a processor including an acquisition module, a processing module, and a cell annotation module. Among them, the names of these units or modules do not constitute a limitation to the units or modules themselves in some cases. For example, the acquisition module can also be described as "configured to acquire cell data of a transcriptome to be predicted; the cell data includes gene expression information of multiple sequencing points in the transcriptome to be predicted and spatial information of the multiple sequencing points".
[0228] As another aspect, this application also provides a computer-readable storage medium, which can be included in the electronic device described in the above embodiments; or it can exist alone without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above-mentioned programs are executed by one or more processors to perform the cell data annotation method described in this application:
[0229] Acquire cell data of a transcriptome to be predicted; the cell data includes gene expression information of multiple sequencing points in the transcriptome to be predicted and spatial information of the multiple sequencing points;
[0230] Determine an annotated cell object corresponding to the transcriptome to be predicted, and determine an initial cell annotation result of the transcriptome to be predicted according to the cell object and the gene expression information;
[0231] Input the cell data into a first prediction model, and correct the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result to obtain a cell annotation result of the transcriptome to be predicted.
[0232] In summary, the cell data annotation method, apparatus, device, and medium provided in the embodiments of the present application determine annotated cell objects corresponding to the transcriptome to be predicted, and determine the initial cell annotation result of the transcriptome to be predicted based on the cell objects and gene expression information. Without any manual annotation, it can obtain the guiding information of the cell annotation result of the transcriptome to be predicted, and predict the cell data through the first prediction model. The initial cell annotation result is corrected according to the loss between the output of the first prediction model and the initial cell annotation result, effectively integrating the gene expression information of multiple sequencing points and the spatial information of multiple sequencing points, and being able to perform self-supervised correction in combination with the guiding information, so as to more accurately and comprehensively utilize all the information in the cell data of the transcriptome to be predicted, and then obtain the cell annotation result of the transcriptome to be predicted more accurately, greatly improving the accuracy of the cell data annotation of the transcriptome to be predicted. It can also be applied to a sequencing analysis system to accurately predict the cell data of the transcriptome to be predicted, greatly improving the quality and efficiency of cell annotation, and providing strong support for the data analysis of spatial transcriptomics.
[0233] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.
Claims
1. A method for cell data annotation, characterized in that, Including: Obtaining cell data of a transcriptome to be predicted; the cell data includes gene expression information of multiple sequencing points in the transcriptome to be predicted and spatial information of the multiple sequencing points; Determining an annotated cell object corresponding to the transcriptome to be predicted, and determining an initial cell annotation result of the transcriptome to be predicted according to the cell object and the gene expression information; Inputting the cell data into a first prediction model, and correcting the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result to obtain a cell annotation result of the transcriptome to be predicted; Wherein, determining the initial cell annotation result of the transcriptome to be predicted according to the cell object and the gene expression information includes: Determining a second prediction model matching the cell object; the second prediction model is trained based on historical gene expression information of the transcriptome to be predicted and historical cell annotation results corresponding to the historical gene expression information; Inputting the gene expression information into the second prediction model to obtain an initial cell annotation result of the transcriptome to be predicted.
2. The method according to claim 1, wherein The cell object is a single cell or a cell group included in a gene sequencing object corresponding to the transcriptome to be predicted.
3. The method according to claim 1, wherein The step of inputting the cell data into the first prediction model, and correcting the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result to obtain a cell annotation result of the transcriptome to be predicted includes: Inputting the cell data into the first prediction model for encoding processing to obtain fused encoding information, and inputting the fused encoding information into a classification module of the first prediction model to obtain an output result of the first prediction model; Determining the loss between the output result and the initial cell annotation result according to a loss function of the first prediction model, iteratively training the classification module according to the loss, and obtaining the adjusted first prediction model when the loss function is minimized; Inputting the cell data into the adjusted first prediction model to obtain a cell annotation result of the transcriptome to be predicted.
4. The method according to claim 3, characterized in that, The step of inputting the cell data into the first prediction model for encoding processing includes: Inputting the gene expression information into a first encoding module of the first prediction model to obtain first encoding information; Determining an adjacency matrix according to the spatial information; the adjacency matrix is used to represent neighboring sequencing points of each sequencing point in the transcriptome to be predicted; Inputting the adjacency matrix and the first encoding information into a second encoding module of the first prediction model for processing to obtain second encoding information; Performing a fusion process on the first encoding information and the second encoding information to obtain the fused encoding information.
5. The method according to claim 4, wherein After obtaining the fused encoding information, the method further includes: Performing decoding processing on the fused encoding information to obtain reconstructed gene expression information; Determine the loss between the reconstructed gene expression information and the gene expression information according to the loss function, and adjust the parameters of the first encoding module according to the loss so that the reconstructed gene expression information is close to the gene expression information.
6. The method according to claim 5, wherein Perform decoding processing on the fused encoding information to obtain reconstructed decoded information, including: Perform linear feature extraction processing on the fused encoding information to obtain intermediate feature information; Input the intermediate feature information into the first decoder for feature restoration processing to obtain the reconstructed gene expression information.
7. The method according to any one of claims 4 to 6, characterized in that After obtaining the fused encoding information, the method further includes: Perform decoding processing on the fused encoding information to obtain reconstructed spatial information; Determine the loss between the reconstructed spatial information and the spatial information according to the loss function, and adjust the parameters of the second encoding module according to the loss so that the reconstructed spatial information is close to the spatial information.
8. The method according to claim 4, characterized in that, The loss function includes a first component, a second component, and a third component; The first component is used to characterize the loss between the initial cell annotation result and the output of the first prediction model; The second component is used to characterize the loss between the gene expression information in the input of the first prediction model and the reconstructed gene expression information; the reconstructed gene expression information is obtained by performing decoding processing on the fused encoding information; The third component is used to characterize the loss between the spatial information in the input of the first prediction model and the reconstructed spatial information; the reconstructed spatial information is obtained by performing decoding processing on the fused encoding information.
9. The method according to claim 8, characterized in that The method further includes: Determine the weight coefficients of the first component, the second component, and the third component; Determine the loss function according to the weight coefficients of the first component, the second component, the third component, the first component, the second component, and the third component; Wherein, the weight coefficient of the first component is related to the importance degree of the transcriptome cell annotation result; the weight coefficient of the second component is related to the importance degree of the transcriptome gene expression information; the weight coefficient of the third component is related to the importance degree of the transcriptome spatial information.
10. The method according to claim 4, wherein The determining the adjacency matrix according to the spatial information includes: Determine the neighboring sequencing points of each sequencing point among the multiple sequencing points according to the spatial information; Generate the adjacency matrix according to the spatial distances corresponding to the neighboring sequencing points of all sequencing points in the transcriptome to be predicted; the adjacency matrix is an N*N matrix, and N is the number of sequencing points in the transcriptome to be predicted.
11. The method according to claim 10, characterized in that, Determine the neighboring sequencing points of each sequencing point among the multiple sequencing points according to the spatial information, including: For each sequencing point among the multiple sequencing points, determine the position information of the sequencing point according to the spatial information, and calculate the distance between the sequencing point and the remaining sequencing points in the transcriptome to be predicted according to the position information of the sequencing point; Determine the sequencing points with a distance less than a preset threshold from the sequencing point as the neighboring sequencing points of the sequencing point.
12. The method according to claim 10 or 11, characterized in that, Generating the adjacency matrix according to the spatial distances corresponding to the neighboring sequencing points of all the sequencing points in the transcriptome to be predicted, includes: For each of the multiple sequencing points, determining the element corresponding to the sequencing point in the adjacency matrix according to the spatial distance between the sequencing point and the corresponding neighboring sequencing point; Setting the remaining elements in the adjacency matrix to zero.
13. A cell data annotation device, characterized in that, The device includes: An acquisition module, configured to acquire cell data of the transcriptome to be predicted; the cell data includes gene expression information of multiple sequencing points in the transcriptome to be predicted and spatial information of the multiple sequencing points; A processing module, configured to determine an annotated cell object corresponding to the transcriptome to be predicted, and determine an initial cell annotation result of the transcriptome to be predicted according to the cell object and the gene expression information; A cell annotation module, configured to input the cell data into a first prediction model, and correct the initial cell annotation result according to the loss between the output of the first prediction model and the initial cell annotation result, to obtain a cell annotation result of the transcriptome to be predicted; Wherein, determining the initial cell annotation result of the transcriptome to be predicted according to the cell object and the gene expression information includes: Determining a second prediction model matching the cell object; the second prediction model is trained based on historical gene expression information of the transcriptome to be predicted corresponding to the cell object and historical cell annotation results corresponding to the historical gene expression information; Inputting the gene expression information into the second prediction model to obtain the initial cell annotation result of the transcriptome to be predicted.
14. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor is configured to implement the cell data annotation method according to any one of claims 1-12 when executing the program.
15. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program is configured to implement the cell data annotation method according to any one of claims 1-12.
16. A computer program product, characterized in that, The computer program product includes instructions, and when the instructions are executed, the cell data annotation method according to any one of claims 1-12 is implemented.
Citation Information
Patent Citations
Cell subset annotation method based on single cell transcriptome sequencing
CN112700820A
Spatial transcriptome cell clustering and analyzing method
CN114091603A