A method, apparatus, equipment, and medium for identifying cell subtypes.

By processing and comparing single-cell transcriptome sequencing data, and combining them with a pre-set model, the problem of low accuracy and efficiency in the identification of disease-related abnormal cell subtypes in existing technologies has been solved, and rapid and accurate cell subtype identification has been achieved.

CN117079717BActive Publication Date: 2026-04-03CHANGSHA KINGMED MEDICAL DIAGNOSTICS INST
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-18
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing flow cytometry and single-cell transcriptome sequencing technologies have low accuracy and efficiency in identifying disease-related abnormal cell subtypes, making it difficult to quickly and accurately identify cell subtypes with unknown markers.

Method used

By acquiring single-cell transcriptome sequencing data and preset transcriptome sequencing data, data processing and cluster analysis are performed, proportion matrix and differential analysis are calculated, gene feature files are extracted, and comparisons are made using preset models to determine the target cell subtype.

Benefits of technology

It improves the accuracy and efficiency of identifying disease-associated immune cell subtypes, enabling rapid and accurate identification of abnormal cell subtypes and enhancing the reliability of identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117079717B_ABST
    Figure CN117079717B_ABST
Patent Text Reader

Abstract

This application proposes a method, apparatus, device, and medium for identifying cell subtypes, comprising: acquiring single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and the preset transcriptome sequencing data are data from the same type of disease; processing the single-cell transcriptome sequencing data to obtain a first cell subtype; processing the first cell subtype to obtain a gene feature file for each cell subtype within the first cell subtype; and comparing the first cell subtype with the preset transcriptome sequencing data to obtain a target cell subtype. Through secondary data processing of the single-cell transcriptome sequencing data, a more accurate abnormal cell subtype can be obtained. Further calculation of the gene feature file of the abnormal cell subtype and comparison with the preset transcriptome sequencing data yields a definitive cell subtype through verification, thus improving the accuracy and efficiency of cell subtype identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cell classification technology, and in particular to a method, apparatus, device and medium for identifying a cell subtype. Background Technology

[0002] Cells are the basic structural and functional units of organisms. Changes in cell state are closely related to the development and occurrence of diseases. Cells in a certain state are called cell subtypes. A large number of studies have clearly linked cell subtypes with specific diseases.

[0003] Identifying disease-associated cell subtypes is crucial for understanding the mechanisms of disease development and developing corresponding treatments. With continuous advancements in medicine, we have come to recognize that immune cells, derived from hematopoietic stem cells, play a vital role in eliminating harmful substances both inside and outside the body, as well as clearing away aging, degenerated, and dead cells. They possess the ability to detect external abnormalities and examine internal balance, earning them the reputation of being the body's health guardians. There are many types of immune cells, each playing a different role in the immune response, including dendritic cells, macrophages, natural killer cells, T lymphocytes, B lymphocytes, and monocytes. Based on gene expression patterns, different immune cells can be further subdivided into different subtypes.

[0004] Currently, the main methods for identifying disease-related abnormal cell subtypes are flow cytometry and single-cell transcriptome sequencing (scRNA-seq). However, since flow cytometry identifies corresponding cell subpopulations through specific fluorescent labels, the target must be determined before sample collection, and it can only identify cell subpopulations with known labels. Furthermore, single-cell transcriptome sequencing technology has high sensitivity but low accuracy, resulting in low accuracy and lower efficiency in the identification results.

[0005] Therefore, there is an urgent need for a method that can identify a cell subtype based on an abnormal cell subtype. Summary of the Invention

[0006] To address the shortcomings of the existing technologies, the present invention aims to provide a method for identifying cell subtypes. This method utilizes single-cell sequencing data and conventional transcriptome sequencing data to further identify disease-related immune cell subtypes, with faster speed and higher accuracy. According to an embodiment of the present invention, a first approach is provided as follows: acquiring single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and the preset transcriptome sequencing data represent data from the same disease; processing the single-cell transcriptome sequencing data to obtain a first cell subtype; processing the first cell subtype to obtain a gene feature file for each cell subtype within the first cell subtype; and comparing the first cell subtype with the preset transcriptome sequencing data to obtain a target cell subtype.

[0007] Furthermore, as a more preferred embodiment of the present invention, the step of processing the single-cell transcriptome sequencing data to obtain a first cell subtype includes: clustering the single-cell transcriptome sequencing data to obtain at least two major cell classes; calculating the proportion of cells in the at least two major cell classes to the total number of cells, obtaining a proportion matrix for each major cell class; performing difference analysis on the proportion matrix of each major cell class to obtain a target cell class; clustering the target cell class to obtain at least two cell subtypes; calculating the proportion of cells in the at least two cell subtypes to the total number of cells in that major cell class, obtaining a proportion matrix for each cell subtype; and performing difference analysis on the proportion matrix of each cell subtype to obtain a first cell subtype.

[0008] Furthermore, as a more preferred embodiment of the present invention, the step of comparing the first cell subtype with the preset transcriptome sequencing data to obtain the target cell subtype includes: constructing a gene expression matrix from the preset transcriptome sequencing data; inputting the gene feature file of each cell subtype and the gene expression matrix into a preset model; predicting the proportion of each cell subtype in each sample of the preset transcriptome sequencing data based on the preset model; and determining the target cell subtype by comparing the proportion of each cell subtype in each sample of the preset transcriptome sequencing data with the proportion of the first cell subtype.

[0009] Furthermore, as a more preferred embodiment of the present invention, the step of predicting the proportion of each cell subtype in each sample of the preset transcriptome sequencing data according to the preset model includes: in the preset model, performing deconvolution calculation on the gene feature file of each cell subtype and the gene expression matrix to obtain the proportion of each cell subtype in each sample of the preset transcriptome sequencing data.

[0010] Further, as a more preferred embodiment of the present invention, the step of determining the target cell subtype by comparing the proportion of each cell subtype in each sample of the first cell subtype with that in the preset transcriptome sequencing data includes: obtaining a first difference between the disease group value and the control group value in the first cell subtype; calculating the difference between the disease group value and the control group value corresponding to each cell subtype in the preset transcriptome sequencing data to obtain a plurality of second differences; comparing the first difference with each of the plurality of second differences one by one to obtain at least one second difference within the same numerical range; comparing the first difference with the at least one second difference one by one to obtain the second difference with the smallest difference as the target difference; obtaining the second cell subtype corresponding to the target difference in the preset transcriptome sequencing data; and determining the first cell subtype and the second cell subtype as the target cell subtype.

[0011] Furthermore, as a more preferred embodiment of the present invention, the step of performing differential analysis on the proportion matrix of each cell subtype to obtain the first cell subtype includes: acquiring control group data and disease group data from the single-cell transcriptome sequencing data; performing differential analysis on the proportion matrix of each cell subtype in the control group data and the disease group data to obtain the difference value of the proportion of each subtype in the control group data and the disease group data to that cell class; correcting the difference value to obtain a target value; and selecting cell subtypes with a value less than a preset threshold as the first cell subtype.

[0012] Furthermore, as a more preferred embodiment of the present invention, the step of processing the data of the first cell subtype to obtain the gene feature file of each cell subtype in the first cell subtype includes: screening out the first gene of each subtype in the first cell subtype according to a preset difference threshold; obtaining the proportion set of the first gene in the total number of cells; screening out target genes with a proportion value greater than a preset proportion value in the proportion set; calculating the average expression level of the first gene in each subtype; and extracting the average expression level corresponding to the target gene in each subtype as the gene feature file of each subtype cell.

[0013] According to an embodiment of the present invention, a second solution provided by the present invention is: a cell subtype identification device, characterized in that the device includes an acquisition unit for acquiring single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and the preset transcriptome sequencing data are data of the same type of disease; a processing unit for processing the single-cell transcriptome sequencing data to obtain a first cell subtype; the processing unit is further configured to process the first cell subtype to obtain a gene feature file of each cell subtype in the first cell subtype; and a comparison unit for comparing the first cell subtype with the preset transcriptome sequencing data to obtain a target cell subtype.

[0014] According to an embodiment of the present invention, a third solution is provided by the present invention: a computer device including a memory and a processor, the memory storing a computer program, the computer program being executed by the processor, the program including instructions for use as in the first solution.

[0015] According to an embodiment of the present invention, a fourth solution is provided by the present invention: a computer storage medium is provided, the computer storage medium storing one or more instructions, the one or more instructions being adapted to be loaded by a processor and executed as described in the first solution above and any possible implementation thereof.

[0016] This application obtains single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and preset transcriptome sequencing data are data from the same type of disease; data processing is performed on the single-cell transcriptome sequencing data to obtain a first cell subtype; data processing is performed on the first cell subtype to obtain a gene feature file for each cell subtype within the first cell subtype; the target cell subtype is obtained by comparing the first cell subtype with the preset transcriptome sequencing data. This invention, through secondary data processing of single-cell transcriptome sequencing data, can obtain a more accurate abnormal cell subtype, further calculate the gene feature file of the abnormal cell subtype, and compare it with the preset transcriptome sequencing data. The precise cell subtype is obtained through comparison and verification, improving the accuracy of cell subtype identification. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the accompanying drawings used in the embodiments of this application or the background art will be described below.

[0018] Figure 1 A schematic flowchart illustrating a method for identifying cell subtypes provided in an embodiment of this application;

[0019] Figure 2AA schematic diagram of a scaling matrix provided for an embodiment of this application;

[0020] Figure 2B A schematic diagram of a scaling matrix provided for an embodiment of this application;

[0021] Figure 3 A schematic flowchart illustrating a method for identifying cell subtypes provided in an embodiment of this application;

[0022] Figure 4 A schematic diagram of a cell subtype identification device provided in an embodiment of this application;

[0023] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0024] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] It should be noted that when a component is referred to as being "fixed to" or "set on" another component, it can be directly on or indirectly set on the other component; when a component is referred to as being "connected to" another component, it can be directly connected to or indirectly connected to the other component.

[0026] It should be understood that the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0027] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" or "several" means two or more, unless otherwise explicitly specified.

[0028] It should be noted that the structures, proportions, sizes, etc., shown in the accompanying drawings of this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed in the specification, and are not intended to limit the conditions under which this application can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size should still fall within the scope of the technical content disclosed in this application, provided that they do not affect the effects and purposes that this application can produce.

[0029] The embodiments of this application are described below with reference to the accompanying drawings.

[0030] Please see Figure 1 , Figure 1 This is a schematic flowchart illustrating a method for identifying cell subtypes provided in an embodiment of this application. The method may include:

[0031] 101. Obtain single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and preset transcriptome sequencing data are data of the same type of disease.

[0032] The aforementioned preset transcriptome sequencing data is conventional transcriptome sequencing data (bulk RNA-seq).

[0033] The single-cell transcriptome sequencing data mentioned above are divided into normal group and disease group, and the conventional transcriptome sequencing data are divided into normal group and disease group. Both groups of data are data of the same type of disease, but the specific data may be different.

[0034] 102. The single-cell transcriptome sequencing data is processed to obtain the first cell subtype.

[0035] In one optional implementation, the step of processing the single-cell transcriptome sequencing data to obtain a first cell subtype includes: clustering the single-cell transcriptome sequencing data to obtain at least two major cell classes; calculating the proportion of cells in the at least two major cell classes to the total number of cells, obtaining a proportion matrix for each major cell class; performing difference analysis on the proportion matrix of each major cell class to obtain a target cell class; clustering the target cell class to obtain at least two cell subtypes; calculating the proportion of cells in the at least two cell subtypes to the total number of cells in that major cell class, obtaining a proportion matrix for each cell subtype; and performing difference analysis on the proportion matrix of each cell subtype to obtain a first cell subtype.

[0036] In practice, after clustering the single-cell transcriptome sequencing data to obtain at least two major cell categories, the single-cell transcriptome sequencing data (i.e., the cell-gene expression matrix of each cell output by CellRanger) is first filtered. This process selects cells expressing no more than 2500 genes but at least 200 genes, while filtering out cells with a mitochondrial percentage greater than 5%. Then, the cell-gene expression matrix is ​​standardized using the "lognormalize" global scaling normalization method. This normalizes the gene expression value of each cell by the total expression value and multiplies it by a scaling factor (set to 10,000). Finally, a logarithmic transformation is performed on the result. Next, a feature subset showing high variability in the dataset (genes expressed highly in some cells and poorly in others) is calculated. `nfeatures = 2000` is set, meaning 2000 genes from each dataset are used for subsequent analysis. The `ScaleData` function is used to normalize the data, ensuring all marker genes have statistically equal significance. The `RunPCA` function is used to reduce the dimensionality of the data, selecting the top 10 principal components for subsequent analysis. Finally, the first 10 principal components were clustered using existing functions, with a resolution of 0.08. Existing techniques were then used to annotate the various cell types, and the annotation results are shown in Table 1.

[0037]

[0038]

[0039] Table 1

[0040] In calculating the proportion of cells in at least two cell categories to the total number of cells, the proportion matrix for each cell category can be as follows: Figure 2A As shown, this is for illustrative purposes only and is not the only data.

[0041] In practice, after performing a difference analysis on the proportion matrix of each cell class to obtain the target cell class, existing functions are first used to perform a difference analysis on the proportion of each cell class in the normal sample group and the disease sample group, respectively, to calculate the difference pvalue of the proportion of each cell class in the total number of cells in the two groups of samples, and the cell class with pvalue <= 0.01 is selected. Then, the pvalues ​​are sorted, and the cell class with the smallest pvalue is selected as the cell class with the most significant difference between the two groups of samples. As shown in Table 2, using the data in Table 1 above, Cluster0, Cluster3 and Cluster6 are obtained. Then, the pvalues ​​of these three clusters are sorted, and the cell class with the smallest pvalue, namely Cluster0 (Classical Monocytes), is selected as the cell class with the most significant difference between the two groups of samples.

[0042]

[0043] Table 2

[0044] In specific implementation, when clustering the target cell class to obtain at least two cell subtypes, the target cell class is first extracted from the total data using existing functions. Then, the top 10 principal components of the extracted data are re-clustered, with resolution = 0.08, to divide Classical Monocytes into three subtypes: subcluster0, subcluster1, and subcluster2.

[0045] For example, in calculating the proportion of cells in at least two cell subtypes to the total number of cells in that cell class, and obtaining the proportion matrix for each cell subtype, the proportion of cells in each cell subtype (subcluster0, subcluster1, subcluster2) in each sample to the total number of cells in that cell class in that sample is calculated, thus obtaining the proportion matrix for each cell subtype across all samples. Figure 2B As shown, Figure 2B This is a schematic diagram of a proportional matrix.

[0046] In one optional implementation, the step of performing differential analysis on the proportion matrix of each cell subtype to obtain a first cell subtype includes: acquiring control group data and disease group data from the single-cell transcriptome sequencing data; performing differential analysis on the proportion matrix of each cell subtype in the control group data and the disease group data to obtain the difference value of the proportion of each subtype in the control group data and the disease group data to that cell class; correcting the difference value to obtain a target value; and selecting cell subtypes with a value less than a preset threshold as the first cell subtype.

[0047] Among them, the above difference value is the FDR value. The smaller the FDR value, the greater the difference.

[0048] Among them, the preset threshold can be set manually or at the factory, and there is no unique limitation here.

[0049] For example, first use existing functions to perform differential analysis on the proportions of subcluster0, subcluster1, and subcluster2 in the disease group and the normal control group respectively, and calculate the differential pvalue of the proportion of each subtype in the two groups of samples in this major cell type. Then use existing functions to correct the pvalue to obtain the FDR value. Screen for FDR <= 0.05 to obtain subcluster0 and subcluster1. Sort the FDR values of subcluster0 and subcluster1, and FDR(subcluster1) < FDR(subcluster0). We infer that subcluster1 is the abnormal cell subtype most relevant to sepsis. Moreover, from the average percentages of subcluster0, subcluster1, and subcluster2 in the two groups of samples respectively, it can be seen that compared with the C normal control group, the proportion of subcluster0 in the disease group increases, the proportion of subcluster1 decreases, the proportion of subcluster2 is the least in both groups, and the difference change is not significant. The results are shown in Table 3:

[0050]

[0051] Table 3

[0052] 103. Perform data processing on the first cell subtype to obtain a gene feature file for each cell subtype in the first cell subtype.

[0053] In an optional implementation manner, the performing data processing on the first cell subtype to obtain a gene feature file for each cell subtype in the first cell subtype includes: screening out the first genes of each subtype in the first cell subtype according to a preset difference threshold; obtaining a proportion set of the first genes in the total number of cells; screening out target genes greater than a preset proportion value in the proportion set; calculating the average expression level of the first genes in each subtype; and extracting the average expression level corresponding to the target genes in each subtype as the gene feature file for each subtype of cells.

[0054] For example, using existing functions to analyze marker genes for each subtype, we set min.pct = 0.25 and logfc.threshold = 0.25. min.pct represents the proportion of the number of cells expressing this gene relative to the total number of cells in that subtype. logfc.threshold represents the fold change in gene expression levels between subtypes. We calculate the average expression level of the gene in each subtype and extract the average expression level of the top 50 marker genes for each subtype as the gene feature file for each subtype of cells.

[0055] 104. The target cell subtype is obtained by comparing the first cell subtype with the preset transcriptome sequencing data.

[0056] In one optional implementation, the step of comparing the first cell subtype with the preset transcriptome sequencing data to obtain the target cell subtype includes: constructing a gene expression matrix from the preset transcriptome sequencing data; inputting the gene feature file of each cell subtype and the gene expression matrix into a preset model; performing prediction based on the preset model to obtain the proportion of each cell subtype in each sample of the preset transcriptome sequencing data; and comparing the proportion of each cell subtype in each sample of the preset transcriptome sequencing data with the proportion of the first cell subtype to determine the target cell subtype.

[0057] The aforementioned preset model can be pre-trained on a target dataset, which mainly contains the gene expression characteristics of each different immune cell.

[0058] For example, bulk RNA sequencing data from the GSE11755 chip was first downloaded from the NCBI database. RNA expression data isolated from monocytes was then extracted based on the sample information, including monocyte RNA expression data from three normal samples and monocyte RNA expression data from eleven sepsis samples. Then, the gene signature files of the subtype cells and routine transcriptome sequencing data were input to predict the proportion of each cell subtype in each sample. The results showed that compared to the normal sample group, the proportion of subcluster0 increased in the sepsis group, the proportion of subcluster1 decreased, and the proportion of subcluster2 was the lowest in both groups, with little difference. This result is consistent with the analysis results of single-cell sequencing data. (See Table 4.)

[0059]

[0060] Table 4

[0061] In one optional implementation, the step of predicting the proportion of each cell subtype in each sample of the preset transcriptome sequencing data according to the preset model includes: performing deconvolution calculation on the gene feature file of each cell subtype and the gene expression matrix in the preset model to obtain the proportion of each cell subtype in each sample of the preset transcriptome sequencing data.

[0062] As can be seen, the technical solution of this invention, through computer methods, not only predicts the types of immune cells related to diseases, but also further predicts the changes in immune cell subtypes during diseases, and verifies the reliability of the results using convolutional analysis of conventional transcriptome sequencing data.

[0063] In one optional implementation, the step of determining the target cell subtype by comparing the proportion of each cell subtype in each sample of the first cell subtype with that in the preset transcriptome sequencing data includes: obtaining a first difference between the disease group value and the control group value in the first cell subtype; calculating the difference between the disease group value and the control group value corresponding to each cell subtype in the preset transcriptome sequencing data to obtain a plurality of second differences; comparing the first difference with each of the plurality of second differences one by one to obtain at least one second difference within the same numerical range; comparing the first difference with the at least one second difference one by one to obtain the second difference with the smallest difference as the target difference; obtaining the second cell subtype corresponding to the target difference in the preset transcriptome sequencing data; and determining the first cell subtype and the second cell subtype as the target cell subtype.

[0064] The above-mentioned same numerical range refers to both being positive or both being negative. If the first difference is -0.2, and among the multiple second differences, A is -0.2, B is 0.2, and C is 1, then the second difference of A being -0.2 is in the same numerical range as the first difference.

[0065] The second difference with the smallest difference is determined by comparing the first difference with each of the at least one second difference. The smaller the difference, the more likely the second difference is to be the target difference. This can be illustrated by the data in Tables 3 and 4: In Table 3, the first cell subtype is subcluster1, with a corresponding first difference of -0.3022; in Table 4, the second difference for subcluster0 is 0.1825, for subcluster1 it is -0.1807, and for subcluster2 it is -0.0018. Therefore, by comparing the first difference with each of the multiple second differences, the subcluster is determined. The second difference value corresponding to ster1 is -0.1807, and the second difference value corresponding to subcluster2 is -0.0018, which are within the same range. Then, the first difference value is compared with the second difference value corresponding to subcluster1, which is -0.1807, to obtain the first difference value. The first difference value is compared with the second difference value corresponding to subcluster2, which is -0.0018, to obtain the second difference value. It is determined that the first difference value is less than the second difference value. Therefore, the second difference value corresponding to subcluster1 is the value closest to the first cell subtype, which is subcluster1. That is, subcluster1 is the target cell subtype.

[0066] This application first obtains single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and preset transcriptome sequencing data are data from the same type of disease; data processing is performed on the single-cell transcriptome sequencing data to obtain a first cell subtype; data processing is performed on the first cell subtype to obtain a gene feature file for each cell subtype within the first cell subtype; the target cell subtype is obtained by comparing the first cell subtype with the preset transcriptome sequencing data. This invention, through secondary data processing of single-cell transcriptome sequencing data, can obtain a more accurate abnormal cell subtype, further calculate the gene feature file of the abnormal cell subtype, and compare it with the preset transcriptome sequencing data. The precise cell subtype is obtained through comparison and verification, improving the accuracy and efficiency of cell subtype identification.

[0067] Figure 3 This is a schematic flowchart of another method for identifying cell subtypes provided in an embodiment of this application, as shown below. Figure 3 As shown, the method includes:

[0068] 301. Obtain single-cell transcriptome sequencing data and preset transcriptome sequencing data;

[0069] 302. Cluster the single-cell transcriptome sequencing data to obtain at least two major cell categories;

[0070] 303. Calculate the proportion of cells in the at least two major cell categories to the total number of cells, and obtain the proportion matrix for each major cell category;

[0071] 304. Perform a difference analysis on the proportion matrix of each cell category to obtain the target cell category;

[0072] 305. Perform clustering on the target cell class to obtain at least two cell subtypes;

[0073] 306. Calculate the proportion of cells in at least two cell subtypes to the total number of cells in the cell class, and obtain the proportion matrix for each cell subtype;

[0074] 307. Perform differential analysis on the proportion matrix of each cell subtype to obtain the first cell subtype;

[0075] 308. Construct a gene expression matrix from the preset transcriptome sequencing data;

[0076] 309. Input the gene feature file of each cell subtype and the gene expression matrix into the preset model;

[0077] 310. Based on the preset model, predict the proportion of each cell subtype in each sample of the preset transcriptome sequencing data;

[0078] 311. The target cell subtype is determined by comparing the proportion of each cell subtype in each sample of the first cell subtype with that in the preset transcriptome sequencing data.

[0079] This application first obtains single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and preset transcriptome sequencing data are data from the same type of disease; data processing is performed on the single-cell transcriptome sequencing data to obtain a first cell subtype; data processing is performed on the first cell subtype to obtain a gene feature file for each cell subtype within the first cell subtype; the target cell subtype is obtained by comparing the first cell subtype with the preset transcriptome sequencing data. This invention, through secondary data processing of single-cell transcriptome sequencing data, can obtain a more accurate abnormal cell subtype, further calculate the gene feature file of the abnormal cell subtype, and compare it with the preset transcriptome sequencing data. The precise cell subtype is obtained through comparison and verification, improving the accuracy and efficiency of cell subtype identification.

[0080] Based on the description of the above-described methods for identifying cell subtypes, this application also discloses a device for identifying cell subtypes, such as... Figure 4 As shown, the cell subtype identification device 400 includes:

[0081] Acquisition unit 401 is used to acquire single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and preset transcriptome sequencing data are data of the same type of disease;

[0082] Processing unit 402 is used to process the single-cell transcriptome sequencing data to obtain a first cell subtype;

[0083] The processing unit 402 is also used to perform data processing on the first cell subtype to obtain the gene feature file of each cell subtype in the first cell subtype;

[0084] The comparison unit 403 is used to compare the first cell subtype with the preset transcriptome sequencing data to obtain the target cell subtype.

[0085] This application first obtains single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and preset transcriptome sequencing data are data from the same type of disease; data processing is performed on the single-cell transcriptome sequencing data to obtain a first cell subtype; data processing is performed on the first cell subtype to obtain a gene feature file for each cell subtype within the first cell subtype; the target cell subtype is obtained by comparing the first cell subtype with the preset transcriptome sequencing data. This invention, through secondary data processing of single-cell transcriptome sequencing data, can obtain a more accurate abnormal cell subtype, further calculate the gene feature file of the abnormal cell subtype, and compare it with the preset transcriptome sequencing data. The precise cell subtype is obtained through comparison and verification, improving the accuracy and efficiency of cell subtype identification.

[0086] This application also provides a computer storage medium (memory), which is a memory device in an electronic device used to store programs and data. It is understood that the computer storage medium here can include both built-in storage media in the electronic device and extended storage media supported by the electronic device. The computer storage medium provides storage space that stores the operating system of the electronic device. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one computer storage medium located remotely from the aforementioned processor.

[0087] In one embodiment, a processor may load and execute one or more instructions stored in a computer storage medium to implement the corresponding steps in the above embodiments; specifically, one or more instructions in the computer storage medium may be loaded and executed by a processor. Figure 1 And / or any step of the method in Figure 2, which will not be described in detail here.

[0088] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0089] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the division of modules is merely a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling, direct coupling, or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, apparatuses, or modules, and may be electrical, mechanical, or other forms.

[0090] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0091] Figure 5 An internal structural diagram of a computer device is shown in one embodiment. This computer device may be a terminal. Figure 5 As shown, the computer device includes a processor, memory, and network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium may store an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement the aforementioned cell subtype identification method. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to perform the aforementioned cell subtype identification method. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the device to which the present application is applied. Specific devices may include more or fewer components than those shown in the figure, or may combine certain components, or may have different component arrangements.

[0092] A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the cell subtype identification method in any of the above embodiments.

[0093] A computer device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the cell subtype identification method in any of the above embodiments.

[0094] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be read-only memory (ROM), random access memory (RAM), or magnetic media, such as floppy disks, hard disks, magnetic tapes, magnetic disks, or optical media, such as digital versatile discs (DVDs), or semiconductor media, such as solid state disks (SSDs).

[0095] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for identifying cell subtypes, characterized in that, The method includes: Acquire single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and preset transcriptome sequencing data are data of the same type of disease; Data processing was performed on the single-cell transcriptome sequencing data to obtain the first cell subtype; Data processing is performed on the first cell subtype to obtain the gene feature file of each cell subtype in the first cell subtype; The target cell subtype is obtained by comparing the first cell subtype with the preset transcriptome sequencing data; wherein, the step of obtaining the target cell subtype by comparing the first cell subtype with the preset transcriptome sequencing data includes: constructing a gene expression matrix from the preset transcriptome sequencing data; inputting the gene feature file of each cell subtype and the gene expression matrix into a preset model; predicting the proportion of each cell subtype in each sample of the preset transcriptome sequencing data based on the preset model; and determining the target cell subtype by comparing the proportion of each cell subtype in each sample of the preset transcriptome sequencing data with the proportion of the first cell subtype. The step of determining the target cell subtype by comparing the proportion of each cell subtype in each sample of the first cell subtype with that in the preset transcriptome sequencing data includes: obtaining a first difference between the disease group value and the control group value in the first cell subtype; calculating the difference between the disease group value and the control group value corresponding to each cell subtype in the preset transcriptome sequencing data to obtain multiple second differences; comparing the first difference with each of the multiple second differences to obtain at least one second difference within the same value range; comparing the first difference with the at least one second difference to obtain the second difference with the smallest difference as the target difference; obtaining the second cell subtype corresponding to the target difference in the preset transcriptome sequencing data; and determining the first cell subtype and the second cell subtype as the target cell subtype.

2. The method according to claim 1, characterized in that, The data processing of the single-cell transcriptome sequencing data to obtain the first cell subtype includes: Clustering was performed on the single-cell transcriptome sequencing data to obtain at least two major cell categories; Calculate the proportion of cells in the at least two major cell categories to the total number of cells, and obtain the proportion matrix for each major cell category; Difference analysis was performed on the proportion matrix of each cell category to obtain the target cell category; Clustering the target cell class yields at least two cell subtypes; Calculate the proportion of cells in the at least two cell subtypes to the total number of cells in the cell class, and obtain the proportion matrix for each cell subtype; Differential analysis was performed on the proportion matrix of each cell subtype to obtain the first cell subtype.

3. The method according to claim 1, characterized in that, The step of predicting the proportion of each cell subtype in each sample of the preset transcriptome sequencing data based on the preset model includes: In the preset model, the gene feature file of each cell subtype and the gene expression matrix are deconvolved to obtain the proportion of each cell subtype in each sample of the preset transcriptome sequencing data.

4. The method according to claim 2, characterized in that, The differential analysis of the proportion matrix of each cell subtype to obtain the first cell subtype includes: Obtain the control group data and disease group data from the single-cell transcriptome sequencing data; A difference analysis was performed on the proportion matrix of each cell subtype in the control group data and the disease group data to obtain... The difference in the proportion of each subtype to that cell class between the control group data and the disease group data; The target value is obtained by correcting the difference value; Cell subtypes with values ​​below a preset threshold are selected as the first cell subtype.

5. The method according to claim 1, characterized in that, The data processing of the first cell subtype to obtain the gene feature file of each cell subtype in the first cell subtype includes: The first gene of each subtype in the first cell subtype is selected based on a preset difference threshold. Obtain the set of proportions of the first gene in the total number of cells; Target genes with a ratio greater than a preset ratio value are selected from the set of ratios. Calculate the average expression level of the first gene in each subtype; The average expression level of the target gene in each subtype is extracted as the gene feature file for each subtype cell.

6. A device for identifying cell subtypes, characterized in that, The apparatus for using the cell subtype identification method according to any one of claims 1 to 5, the apparatus comprising: The acquisition unit is used to acquire single-cell transcriptome sequencing data and preset transcriptome sequencing data, wherein the single-cell transcriptome sequencing data and preset transcriptome sequencing data are data of the same type of disease. A processing unit is used to process the single-cell transcriptome sequencing data to obtain a first cell subtype. The processing unit is also used to process the data of the first cell subtype to obtain the gene feature file of each cell subtype in the first cell subtype; The comparison unit is used to compare the first cell subtype with the preset transcriptome sequencing data to obtain the target cell subtype.

7. A computer device comprising a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method for identifying cell subtypes as claimed in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The device contains a computer program that, when executed by a processor, causes the processor to perform the steps of the method for identifying cell subtypes as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Cell subtype intelligent judgment method

    CN113918786A

  • Differential analysis method and system based on single cell samples of mixed experimental group and control group

    CN114864003A

  • Analysis method, device and equipment based on single cell transcriptome sequencing data

    CN116189764A