Method and device for constructing scRNA-seq cell type annotation database and electronic equipment

CN115579069BActive Publication Date: 2026-09-25BGI TECH SOLUTIONS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211328761.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-27
Publication Date
2026-09-25
Estimated Expiration
2042-10-27

AI Technical Summary

Benefits of technology

[0057]根据本公开的第五方面,提供了一种计算机程序产品,包括计算机程序,所述计算机程序在被处理器执行时实现如前述第一方面所述的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115579069B_ABST
    Figure CN115579069B_ABST
Patent Text Reader

Abstract

The application discloses a kind of scRNA-Seq cell type annotation database construction method, device and electronic equipment, it is related to scRNA-Seq cell type annotation technical field, main technical scheme includes: in each published data set, the single cell data of target tissue type of target species is extracted;Each pre-constructed cell classification model is trained based on each training set;After training, the gene expression profile in each test data set is respectively input into each trained cell classification model, and the classification prediction result of each cell in each test data set is obtained;According to each classification prediction result and each cell annotation label, the pre-constructed ensemble learning model is trained. By training a cell classification model for each single cell data respectively, when the data in database is added, only a cell classification model needs to be trained for the added data, and the original data does not need to be retrained, so that the scalability of the model is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of scRNA-Seq cell type annotation technology, and in particular to a method, apparatus and electronic device for constructing an scRNA-Seq cell type annotation database. Background Technology

[0002] Single-cell RNA sequencing (scRNA-Seq) technology can analyze the transcriptome (RNA) expression profile of cells at the resolution of a single cell, greatly facilitating the study of tissue heterogeneity. Deciphering and identifying the specific cell types contained in scRNA-Seq data is a crucial step in scRNA-Seq data mining.

[0003] Currently, the following method is commonly used to determine the type of single cells based on scRNA-Seq data: a cell classification model is trained based on scRNA-Seq data and cell types in an existing cell type database, and the type of single cells with undetermined cell types is predicted based on the trained cell classification model.

[0004] While the above method can predict the type of single cells, it does not take into account the scalability of the model. That is, when the data in the cell type database is updated, the cell classification model needs to be retrained based on the updated cell type database, which will consume a lot of human and material resources. Therefore, how to achieve the scalability of the cell classification model is an urgent problem to be solved. Summary of the Invention

[0005] This disclosure provides a method, apparatus, and electronic device for constructing an scRNA-Seq cell type annotation database. Its main purpose is to achieve scalability of cell classification models.

[0006] According to a first aspect of this disclosure, a method for constructing an scRNA-Seq cell type annotation database is provided, comprising:

[0007] Based on the target tissue type of the target species, single-cell data of each tissue type are extracted from each publicly published dataset. The single-cell data includes the gene expression profile of each cell and the annotation tags of each cell. The annotation tags are used to identify the cell type.

[0008] The data for each single cell were divided into training datasets and test datasets.

[0009] Each pre-built cell classification model is trained based on each training dataset, wherein each cell classification model corresponds one-to-one with the training dataset.

[0010] The gene expression profiles of each cell in each test dataset are input into the cell classification models after training, and the classification prediction results of each cell in the test dataset are obtained.

[0011] The pre-built ensemble learning model is trained based on the cell annotation labels of each cell in each of the test datasets and the corresponding cell classification prediction results.

[0012] Optionally, before dividing the single-cell data into training datasets and test datasets, the method further includes:

[0013] The single-cell data described above are preprocessed;

[0014] Unify the annotation labels for single-cell data of the same cell type in different datasets.

[0015] Optionally, the preprocessing of each of the single-cell data includes:

[0016] The gene expression profiles in each single-cell data are screened based on preset screening conditions, and the gene expression profiles are standardized.

[0017] Select a characteristic gene from each of the gene expression profiles, the characteristic gene being the gene with the highest degree of variation among cells.

[0018] Optionally, unifying the annotation labels for single-cell data of the same cell type in different datasets includes:

[0019] Remove batch effects between individual cell data;

[0020] Calculate the average gene expression levels for each cell type;

[0021] Based on the average gene expression levels described, determine whether the cell types of different single-cell data are the same in pairs, and unify the annotation labels of the same cell types until the cell types of all different single-cell data are determined.

[0022] Optionally, before inputting the gene expression profiles of each cell in each test dataset into the trained cell classification model to obtain the classification prediction results of each cell in the test dataset, the method further includes:

[0023] The classification ability of each trained cell classification model is evaluated based on the test dataset and the preset evaluation function.

[0024] Based on the evaluation results, the cell classification models are screened according to a preset screening threshold.

[0025] Optionally, the method further includes:

[0026] Obtain the cell gene expression profile of the cell type to be identified, and preprocess the cell gene expression profile;

[0027] The processed cell gene expression profiles are input into each of the cell classification models to obtain the prediction results.

[0028] The prediction results are input into the ensemble learning model to obtain the final predicted type of the cell type to be identified.

[0029] According to a second aspect of this disclosure, an apparatus for constructing an scRNA-Seq cell type annotation database is provided, comprising:

[0030] The extraction unit is used to extract single-cell data of each tissue type from each publicly published dataset according to the target tissue type of the target species. The single-cell data includes the gene expression profile of each cell and the annotation tag of each cell. The annotation tag is used to identify the cell type.

[0031] The segmentation unit is used to divide the single-cell data into training datasets and test datasets.

[0032] The first training unit is used to train each pre-built cell classification model based on each training dataset, wherein each cell classification model corresponds one-to-one with the training dataset.

[0033] The first acquisition unit is used to input the gene expression profile of each cell in each test dataset into each trained cell classification model to obtain the classification prediction results of each cell in the test dataset.

[0034] The second training unit is used to train the pre-built ensemble learning model based on the cell annotation labels of each cell in each of the test datasets and the corresponding classification prediction results of each cell.

[0035] Optionally, the device further includes:

[0036] A preprocessing unit is used to preprocess each single-cell data before the segmentation unit divides each single-cell data into each training dataset and each test dataset.

[0037] A unified unit is used to standardize the annotation labels of single-cell data of the same cell type from different datasets.

[0038] Optionally, the preprocessing unit further includes:

[0039] The preprocessing module is used to screen the gene expression profiles in each single cell data based on preset screening conditions, and to standardize the data of each gene expression profile.

[0040] The selection module is used to select a characteristic gene from each of the gene expression profiles, wherein the characteristic gene is the gene with the highest degree of variation among cells.

[0041] Optionally, the unified unit includes:

[0042] The removal module is used to remove batch effects between individual cell data.

[0043] The calculation module is used to calculate the average gene expression levels for each cell type.

[0044] The judgment module is used to determine whether the cell types in different single-cell data are the same based on the average gene expression levels, and to unify the annotation labels of the same cell types until all cell types in different single-cell data are judged.

[0045] Optionally, the device further includes:

[0046] The evaluation unit is used to evaluate the classification ability of each cell classification model after training based on the test dataset and a preset evaluation function before the first acquisition unit inputs the gene expression profile of each cell in each test dataset into each trained cell classification model and obtains the classification prediction results of each cell in the test dataset.

[0047] The screening unit is used to screen the cell classification model according to the evaluation results and a preset screening threshold.

[0048] Optionally, the device further includes:

[0049] The second acquisition unit is used to acquire the cell gene expression profile of the cell type to be identified and to preprocess the cell gene expression profile.

[0050] The input unit is used to input the processed cell gene expression profile into each of the cell classification models to obtain prediction results.

[0051] The third acquisition unit is used to input the prediction results into the ensemble learning model to obtain the final predicted type of the cell type to be identified.

[0052] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0053] At least one processor; and

[0054] A memory communicatively connected to the at least one processor; wherein,

[0055] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0056] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

[0057] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in the first aspect above.

[0058] The method, apparatus, and electronic device for constructing the scRNA-Seq cell type annotation database disclosed herein mainly include the following technical solutions: Based on the target tissue type of the target species, extracting single-cell data corresponding to each tissue type from publicly published datasets, wherein the single-cell data includes gene expression profiles and annotation tags for each cell, the annotation tags being used to identify cell types; dividing each single-cell data into training datasets and test datasets; training each pre-constructed cell classification model based on each training dataset, wherein each cell classification model corresponds one-to-one with the training dataset; inputting the gene expression profiles of each cell in each test dataset into each trained cell classification model to obtain classification prediction results for each cell in the test dataset; and training a pre-constructed ensemble learning model based on the cell annotation tags and corresponding classification prediction results for each cell in each test dataset. Compared with related technologies, this application trains a corresponding cell classification model for each database separately. When new data is added to the database, only a new cell classification model needs to be trained for the new data, eliminating the need to retrain the existing data, thus achieving model scalability.

[0059] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0060] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0061] Figure 1 This is a flowchart illustrating a method for constructing an scRNA-Seq cell type annotation database according to an embodiment of this disclosure.

[0062] Figure 2This is a flowchart illustrating a method for processing single-cell data from different datasets, provided in an embodiment of this disclosure.

[0063] Figure 3 A schematic flowchart illustrating a cell type identification method provided in an embodiment of this disclosure;

[0064] Figure 4 A schematic diagram of a device for constructing an scRNA-Seq cell type annotation database provided in this embodiment of the disclosure;

[0065] Figure 5 A schematic diagram of a device for constructing an scRNA-Seq cell type annotation database provided in this embodiment of the disclosure;

[0066] Figure 6 A schematic block diagram of an example electronic device provided for embodiments of this disclosure. Detailed Implementation

[0067] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0068] The following describes, with reference to the accompanying drawings, a method, apparatus, and electronic device for constructing an scRNA-Seq cell type annotation database according to embodiments of the present disclosure.

[0069] Figure 1 This is a flowchart illustrating a method for constructing an scRNA-Seq cell type annotation database provided in this embodiment of the disclosure.

[0070] like Figure 1 As shown, the method includes the following steps:

[0071] Step 101: Based on the target tissue type of the target species, extract single-cell data of each tissue type from each publicly published dataset. The single-cell data includes the gene expression profile of each cell and the annotation tags of each cell. The annotation tags are used to identify the cell type.

[0072] First, the species and tissue type of the cells to be classified are determined. For example, if the cells to be classified are determined to be human blood tissue cells, then data related to human blood tissue cells are extracted from the dataset. Data sources include, but are not limited to, publicly available datasets such as Gene Expression Omnibus (NCBI), Single Cell Expression Atlas (EMBL-EBI), and Single Cell PORTAL (Broad Institute). Publicly available datasets contain data on various tissue cells, and each tissue cell contains various types of single-cell data. Furthermore, since the gene expression levels of cells differ, each single-cell dataset contains gene expression profiles of multiple different cells. In this embodiment, the collection of human-blood cell related data from two datasets is used as an example for illustration. However, it should be noted that this description is not a specific limitation on specific cell types and specific datasets, and this embodiment does not impose any limitations on them.

[0073] Step 102: Divide the single-cell data into training datasets and test datasets.

[0074] The single-cell data extracted from the two datasets are divided into training datasets and test datasets, resulting in two training datasets and two test datasets. During allocation, random allocation or specified allocation can be performed. The data in the training dataset must be much larger than that in the test dataset. The ratio of the number of data in the training dataset to the number of data in the test dataset can be set to 9:1 or 8:2. Specifically, this application embodiment does not limit the data allocation method and allocation ratio.

[0075] Step 103: Train each pre-built cell classification model based on each training dataset, wherein each cell classification model corresponds one-to-one with the training dataset.

[0076] The cell classification model can be constructed using deep neural network models, such as the TensorFlow framework, and this application embodiment does not limit this method. After the data is randomly allocated in step 102, two training datasets are obtained. The training datasets correspond one-to-one with the cell classification model for training, so that the cell classification model can learn the ability to classify cell types.

[0077] Step 104: Input the gene expression profile of each cell in each test dataset into the cell classification model after training to obtain the classification prediction results of each cell in the test dataset.

[0078] Once the cell classification model is trained, it will have a certain cell classification ability. At this point, the data from the two test datasets are input into the two cell classification models in turn. For a single cell dataset, there will be three label information data: the classification results obtained by the two classification models and the cell's original annotation label.

[0079] Step 105: Train the pre-built ensemble learning model based on the cell annotation labels of each cell in each of the test datasets and the corresponding classification prediction results of each cell.

[0080] When building an ensemble learning model, a logistic regression classifier can be constructed using the Scikit-learn framework. This classifier serves as the ensemble learning model. Based on the actual cell annotation labels and the classification results of the two classifiers, the ensemble learning model is trained to output a final, confirmed cell classification result.

[0081] The method for constructing the scRNA-Seq cell type annotation database disclosed herein mainly includes the following technical solutions: Based on the target tissue type of the target species, extract single-cell data corresponding to each tissue type from various publicly published datasets. The single-cell data includes gene expression profiles and annotation tags for each cell, with the annotation tags identifying the cell type. Divide each single-cell data into training datasets and test datasets. Train each pre-constructed cell classification model based on each training dataset, wherein each cell classification model corresponds one-to-one with the training dataset. Input the gene expression profiles of each cell in each test dataset into the trained cell classification model to obtain the classification prediction results for each cell in the test dataset. Train the pre-constructed ensemble learning model based on the cell annotation tags and corresponding classification prediction results for each cell in each test dataset. Compared with related technologies, this application trains a corresponding cell classification model for each database separately. When new data is added to the database, only a new cell classification model needs to be trained for the new data, eliminating the need to retrain the existing data and achieving model scalability.

[0082] As an extension to the above-described embodiments, since different datasets were created by different people and the cell data in different datasets also differ, the data needs to be processed before training the model using data from different data sources; please refer to... Figure 2 , Figure 2 This is a flowchart illustrating a method for processing single-cell data from different datasets, provided by an embodiment of this disclosure; including:

[0083] Step 201: Preprocess the single-cell data.

[0084] The gene expression profiles in each single-cell data are screened based on preset screening criteria, and the gene expression profiles are then standardized.

[0085] Because there are some differences in gene expression in individual cells, gene expression profile data of multiple cells will be collected for each cell type. Among them, there will be some low-quality gene expression profile data, such as data with too low gene expression levels or data with undetected gene sequences. Therefore, in practical applications, cell data need to be screened in advance. The preset screening conditions can be set to normal gene expression levels, etc. In practical applications, the settings can be made according to the experimental precision requirements and cell specificity. This application embodiment does not limit this.

[0086] The expression levels of different cells in a cell are different. If a sample has a high expression level, it will occupy an absolute dominant position in the overall expression, thus masking the role of samples with low expression levels. However, this does not mean that samples with low expression levels are unimportant. It is also possible that the sample contains a lot of low-expressed genes. The data in the gene expression spectrum needs to be standardized. For specific processing methods, please refer to any implementation method in the prior art. The embodiments of this application will not be described in detail here.

[0087] Select a characteristic gene from each of the gene expression profiles, the characteristic gene being the gene with the highest degree of variation among cells.

[0088] After standardizing the gene expression profile data, the genes with the highest degree of variation among cells are selected as the characteristic genes of the current cell. During the selection, the gene expression level can be selected based on indicators such as variance, dispersion, or F-score. This application does not limit this.

[0089] Step 202: Unify the annotation labels of single-cell data of the same cell type in different datasets.

[0090] Because the datasets were edited by different people, the annotations for the same cell type may differ in terms of capitalization or naming conventions, resulting in multiple annotations for the same cell type. Furthermore, batch effects exist between single-cell datasets; therefore, it is necessary to first remove batch effects between individual cell datasets. Since different cell type datasets are influenced by non-biological factors such as different times, locations, and experimenters, differences in cell data may exist. These differences caused by non-biological factors can be removed using methods such as Combat, BBKNN, and Harmony. This application does not limit the method for removing batch effects.

[0091] After removing batch effects between single-cell data, the differences between single-cell data are differences in biological factors. Based on this, it is possible to identify whether different single-cell data belong to the same type. First, the average gene expression level of each cell type is calculated. The difference in gene expression level of single cells will lead to slight differences in gene expression level of cells of the same type. However, the gene expression level of cells of the same type is the same overall. Therefore, by calculating the average gene expression level and removing the differences in gene expression level of individual cells, it is possible to determine whether they belong to the same cell type based on the average gene expression level. Based on the average gene expression levels, the cell types of different single-cell data are determined pairwise to see if they are the same, and the annotation labels of the same cell types are unified until all cell types of different single-cell data are determined. When determining whether two cell types are the same cell type, the following method can be used: After determining the average gene expression level of each cell type, the Euclidean distance between the average gene expression levels of each cell type is calculated, and a minimum spanning tree between cell types is constructed based on the results. In the minimum spanning tree, the nodes represent cell types, the edges represent the approximate relationship between cell types, and the weight of the edges is inversely proportional to the distance between cell types. The Pearson correlation coefficient between two adjacent nodes in the minimum spanning tree is calculated. If the calculated Pearson correlation coefficient is greater than a preset threshold, it can be determined that the two cell types are the same cell type. That is, two cell types that are adjacent in the minimum spanning tree and have a Pearson correlation coefficient greater than the preset threshold can be confirmed as the same cell type. The preset threshold is an empirical value and should not be set too low. It can be set to 0.95 or 0.97, and this application embodiment does not limit this.

[0092] In one possible implementation of this application embodiment, when pre-constructing the cell classification model, it can be built based on the TensorFlow framework. The number of hidden layers and the dimension of each layer can be adjusted according to the data in the training set. During training, the activation function of the hidden layers is the ReLU function, the activation function of the output layer is the Softmax function, the cross-entropy loss function is used as the loss function, and the Adam optimization method is used to control and optimize the cell classification model process. After training, the classification ability of each trained cell classification model needs to be evaluated based on the test dataset and the preset evaluation function. Based on the evaluation results, the cell classification models are screened according to the preset screening threshold. The evaluation index can be the classification accuracy index, the ARI index, etc. The screening index can be set to 0.9 or 0.92. Cell classification models with a screening index lower than the screening index are discarded, and only cell classification models with a screening index higher than the screening index are retained to ensure the accuracy of cell classification. It should be noted that this description is not a limitation on specific screening indicators, and this application embodiment does not limit the evaluation indicators and screening index.

[0093] Once the ensemble learning model has been trained, it can be used for cell type prediction based on cell location; please refer to [reference needed]. Figure 3 , Figure 3 This is a flowchart illustrating a cell type identification method provided in an embodiment of the present disclosure; including:

[0094] Step 301: Obtain the cell gene expression profile of the cell type to be identified, and preprocess the cell gene expression profile.

[0095] The embodiments of this application can be used to determine the type of single cells in tissue cells. After determining the gene expression profile of a single cell, this part of the data is input into the pre-trained data to obtain the predicted classification of the cell. The preprocessing method for the cell gene expression profile can be referred to step 201, and will not be described in detail here.

[0096] Step 302: Input the processed cell gene expression profile into each of the cell classification models to obtain the prediction results.

[0097] Step 303 inputs each prediction result into the ensemble learning model to obtain the final predicted type of the cell type to be identified.

[0098] Based on the classification results of each cell classification model, a final predicted type is determined and output as the type of the cell to be identified.

[0099] Corresponding to the above-described method for constructing an scRNA-Seq cell type annotation database, this invention also proposes an apparatus for constructing an scRNA-Seq cell type annotation database. Since the apparatus embodiments of this invention correspond to the method embodiments described above, details not disclosed in the apparatus embodiments can be referred to in the method embodiments described above, and will not be repeated here.

[0100] Figure 4 A schematic diagram of a device for constructing an scRNA-Seq cell type annotation database provided in this disclosure embodiment is shown below. Figure 4 As shown, it includes:

[0101] Extraction unit 41 is used to extract single-cell data of each tissue type from each publicly published dataset according to the target tissue type of the target species. The single-cell data includes the gene expression profile of each cell and the annotation tag of each cell. The annotation tag is used to identify the cell type.

[0102] Segmentation unit 42 is used to divide each single cell data into each training dataset and each test dataset;

[0103] The first training unit 43 is used to train each pre-built cell classification model based on each training dataset, wherein each cell classification model corresponds one-to-one with the training dataset.

[0104] The first acquisition unit 44 is used to input the gene expression profile of each cell in each test dataset into each trained cell classification model to obtain the classification prediction results of each cell in the test dataset.

[0105] The second training unit 45 is used to train the pre-built ensemble learning model based on the cell annotation labels of each cell in each of the test datasets and the corresponding classification prediction results of each cell.

[0106] The scRNA-Seq cell type annotation database construction device disclosed herein mainly includes the following technical solutions: Based on the target tissue type of the target species, extracting single-cell data corresponding to each tissue type from various publicly published datasets; the single-cell data includes gene expression profiles and annotation tags for each cell, with the annotation tags identifying the cell type; dividing each single-cell data into training datasets and test datasets; training each pre-constructed cell classification model based on each training dataset, wherein each cell classification model corresponds one-to-one with the training dataset; inputting the gene expression profiles of each cell in each test dataset into each trained cell classification model to obtain classification prediction results for each cell in the test dataset; and training a pre-constructed ensemble learning model based on the cell annotation tags and corresponding classification prediction results for each cell in each test dataset. Compared with related technologies, this application trains a corresponding cell classification model for each database separately. When new data is added to the database, only a new cell classification model needs to be trained for the new data, eliminating the need to retrain the existing data, thus achieving model scalability.

[0107] Furthermore, in one possible implementation of this embodiment, such as Figure 5 As shown, the device further includes:

[0108] Preprocessing unit 46 is used to preprocess each single cell data before the segmentation unit 42 divides each single cell data into each training dataset and each test dataset;

[0109] Unified Unit 47 is used to unify the annotation labels of single-cell data of the same cell type in different datasets.

[0110] Furthermore, in one possible implementation of this embodiment, such as Figure 5 As shown, the preprocessing unit 46 further includes:

[0111] The preprocessing module 461 is used to screen the gene expression profiles in each single cell data based on preset screening conditions, and to standardize the data of each gene expression profile.

[0112] Selection module 462 is used to select a characteristic gene from each of the gene expression profiles, the characteristic gene being the gene with the highest degree of variation among cells.

[0113] Furthermore, in one possible implementation of this embodiment, such as Figure 5 As shown, the unified unit 47 includes:

[0114] Module 471 is used to remove batch effects between individual cell data.

[0115] Calculation module 472 is used to calculate the average gene expression level of each cell type in each single cell data;

[0116] The judgment module 473 is used to determine whether the cell types in different single-cell data are the same based on the average gene expression levels, and to unify the annotation labels of the same cell types until all cell types in different single-cell data are judged.

[0117] Furthermore, in one possible implementation of this embodiment, such as Figure 5 As shown, the device further includes:

[0118] The evaluation unit 48 is used to evaluate the classification ability of each cell classification model after training based on the test dataset and a preset evaluation function before the first acquisition unit 44 inputs the gene expression profile of each cell in each test dataset into each trained cell classification model to obtain the classification prediction results of each cell in the test dataset.

[0119] The screening unit 49 is used to screen the cell classification model according to the evaluation results and a preset screening threshold.

[0120] Furthermore, in one possible implementation of this embodiment, such as Figure 5 As shown, the device further includes:

[0121] The second acquisition unit 410 is used to acquire the cell gene expression profile of the cell type to be identified and to preprocess the cell gene expression profile.

[0122] Input unit 411 is used to input the processed cell gene expression profile into each of the cell classification models to obtain each prediction result;

[0123] The third acquisition unit 412 is used to input the prediction results into the ensemble learning model to obtain the final predicted type of the cell type to be identified.

[0124] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and the principle is the same, so it is not limited in this embodiment.

[0125] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0126] Figure 6 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0127] like Figure 6 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 502 or a computer program loaded from storage unit 508 into RAM (Random Access Memory) 503. RAM 503 can also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. I / O (Input / Output) interface 505 is also connected to bus 504.

[0128] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0129] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the method for constructing an scRNA-Seq cell type annotation database. For example, in some embodiments, the method for constructing an scRNA-Seq cell type annotation database can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, computing unit 501 may be configured by any other suitable means (e.g., by means of firmware) to perform the aforementioned method for constructing the scRNA-Seq cell type annotation database.

[0130] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0131] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0132] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0134] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0135] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0136] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0137] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0138] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for constructing an scRNA-Seq cell type annotation database, characterized in that, include: Based on the target tissue type of the target species, single-cell data of each tissue type are extracted from each publicly published dataset. The single-cell data includes the gene expression profile of each cell and the annotation tags of each cell. The annotation tags are used to identify the cell type. The single-cell data described above are preprocessed; Standardize the annotation labels for single-cell data of the same cell type in different datasets; The data for each single cell were divided into training datasets and test datasets. Each pre-built cell classification model is trained based on each training dataset, wherein each cell classification model corresponds one-to-one with the training dataset. The gene expression profiles of each cell in each test dataset are input into the cell classification models after training, and the classification prediction results of each cell in the test dataset are obtained. The pre-constructed ensemble learning model is trained based on the cell annotation labels of each cell in each of the test datasets and the corresponding cell classification prediction results; so as to obtain a trained ensemble learning model that can output the final confirmed cell type based on the prediction results of multiple cell classification models. By training a cell classification model for each single cell data point, when new data is added to the database, only a new cell classification model needs to be trained for the new data, without having to retrain the existing data. The method further includes: Obtain the cell gene expression profile of the cell type to be identified, and preprocess the cell gene expression profile; The processed cell gene expression profiles are input into each of the cell classification models to obtain prediction results. The prediction results are input into the ensemble learning model to obtain the final predicted type of the cell type to be identified.

2. The method according to claim 1, characterized in that, The preprocessing of the single-cell data includes: The gene expression profiles in each single-cell data are screened based on preset screening conditions, and the gene expression profiles are standardized. Select a characteristic gene from each of the gene expression profiles, the characteristic gene being the gene with the highest degree of variation among cells.

3. The method according to claim 1, characterized in that, The process of unifying the annotation labels for single-cell data of the same cell type from different datasets includes: Remove batch effects between individual cell data; Calculate the average gene expression levels for each cell type; Based on the average gene expression levels described, determine whether the cell types of different single-cell data are the same in pairs, and unify the annotation labels of the same cell types until the cell types of all different single-cell data are determined.

4. The method according to claim 1, characterized in that, Before inputting the gene expression profiles of each cell in each test dataset into the trained cell classification models to obtain the classification prediction results for each cell in the test dataset, the method further includes: The classification ability of each trained cell classification model is evaluated based on the test dataset and the preset evaluation function. Based on the evaluation results, the cell classification models are screened according to a preset screening threshold.

5. An apparatus for constructing an scRNA-Seq cell type annotation database, characterized in that, include: The extraction unit is used to extract single-cell data of each tissue type from each publicly published dataset according to the target tissue type of the target species. The single-cell data includes the gene expression profile of each cell and the annotation tag of each cell. The annotation tag is used to identify the cell type. The segmentation unit is used to divide the single-cell data into training datasets and test datasets. The first training unit is used to train each pre-built cell classification model based on each training dataset, wherein each cell classification model corresponds one-to-one with the training dataset. The first acquisition unit is used to input the gene expression profile of each cell in each test dataset into each trained cell classification model to obtain the classification prediction results of each cell in the test dataset. The second training unit is used to train the pre-constructed ensemble learning model based on the cell annotation labels of each cell in each of the test datasets and the corresponding classification prediction results of each cell; so as to obtain a trained ensemble learning model that can output the final confirmed cell type based on the prediction results of multiple cell classification models; by training a cell classification model for each single cell data separately, when new data is added to the database, only a cell classification model needs to be trained on the new data, without having to retrain the original data; The device further includes: The second acquisition unit is used to acquire the cell gene expression profile of the cell type to be identified and to preprocess the cell gene expression profile. The input unit is used to input the processed cell gene expression profile into each of the cell classification models to obtain prediction results. The third acquisition unit is used to input the prediction results into the ensemble learning model to obtain the final predicted type of the cell type to be identified. The device further includes: A preprocessing unit is used to preprocess each single-cell data before the segmentation unit divides each single-cell data into each training dataset and each test dataset. A unified unit is used to standardize the annotation labels of single-cell data of the same cell type from different datasets.

6. The apparatus according to claim 5, characterized in that, The preprocessing unit further includes: The preprocessing module is used to screen the gene expression profiles in each single cell data based on preset screening conditions, and to standardize the data of each gene expression profile. The selection module is used to select a characteristic gene from each of the gene expression profiles, wherein the characteristic gene is the gene with the highest degree of variation among cells.

7. The apparatus according to claim 5, characterized in that, The unified unit includes: The removal module is used to remove batch effects between individual cell data. The calculation module is used to calculate the average gene expression levels for each cell type. The judgment module is used to determine whether the cell types in different single-cell data are the same based on the average gene expression levels, and to unify the annotation labels of the same cell types until all cell types in different single-cell data are judged.

8. The apparatus according to claim 5, characterized in that, The device further includes: The evaluation unit is used to evaluate the classification ability of each cell classification model after training based on the test dataset and a preset evaluation function before the first acquisition unit inputs the gene expression profile of each cell in each test dataset into each trained cell classification model and obtains the classification prediction results of each cell in the test dataset. The screening unit is used to screen the cell classification model according to the evaluation results and a preset screening threshold.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-4.

11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Method and device for establishing classification model

    CN114328936A

  • Cell type automatic classification method based on ensemble learning

    CN114882954A