Double-cell recognition method, system and equipment and storage medium
Through the COSG method of cyclic iterative combined with statistical strategies, the two-cell data in single-cell sequencing data were identified, solving the problems of data sparseness and noise impact in the prior art, and achieving more stable and accurate two-cell recognition.
Patent Information
- Application Number
- CN202510121135.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing two-cell recognition methods are prone to loss of information due to the sparseness of data in single-cell transcriptome analysis, resulting in blind spots in quality control, and the prior art is not stable enough under the influence of background noise, clustering and expression levels.
The COSG method with cyclic iterative combined with statistical strategies was used to identify the two-cell data in single-cell sequencing data through the steps of clustering, identification of the Yin-Yang cell population of mutex pairs, demarcation threshold calculation and double-cell data identification.
More stable and accurate double-cell recognition is achieved, the influence of background noise and clustering is avoided, and supporting proteins and genes are not required. Only ADT data can be used for identification and analysis.
Smart Images

Figure CN119917879A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of single-cell omics, and specifically relates to a dual-cell identification method, system, device and storage medium. Background Art
[0002] The development of single-cell omics technology has played a key role in promoting people's understanding of health and disease, and has provided an important basis for discovering disease mechanisms and potential therapeutic targets. Quality control (QC) of single-cell multi-omics is crucial to ensure the reliability and accuracy of data.
[0003] In a single-cell sequencing experiment, if a reaction body happens to contain two cells, a double cell will be formed. Since double cells are not real cells, their existence will seriously interfere with the analysis of single-cell sequencing data. Therefore, it is necessary to develop some computational methods to identify double cell data in single-cell sequencing data.
[0004] However, existing dual-cell identification methods are usually limited in use. Single-cell transcriptome analysis is prone to information loss due to data sparsity, resulting in blind spots in quality control. The ADT data of single-cell RNA and ADT co-sequencing is bimodal, and the missing rate is not that serious. Therefore, for multi-omics RNA+ADT co-sequencing data, it is necessary to further identify and supplement dual-cell identification at the protein level. Summary of the invention
[0005] The invention provides a dual-cell identification method, system, device and storage medium to more stably and accurately identify dual-cell data in single-cell sequencing data. The invention is implemented by the following technical solutions:
[0006] A double cell identification method comprises the following steps:
[0007] Step 1) obtaining a data set to be subjected to dual-cell identification, wherein the data set to be subjected to dual-cell identification includes protein data generated by single-cell transcriptome sequencing and single-cell membrane protein sequencing; then clustering the data set to be subjected to dual-cell identification to obtain multiple cell populations;
[0008] Step 2) determining mutually exclusive protein pairs according to the cell types of the data set to be subjected to dual cell identification, and then determining the negative expression cell population and the positive expression cell population corresponding to each protein in the mutually exclusive protein pair in the cell population obtained in step 1) using the cyclic iterative COSG method;
[0009] Step 3) calculating the demarcation threshold of the positive and negative expression of each protein in the mutually exclusive protein pair according to the data in the negative expression cell population and the positive expression cell population;
[0010] Step 4) identifying the double-cell data in the data set to be subjected to double-cell identification according to the demarcation threshold.
[0011] The mutually exclusive protein pair is customized according to the biological cell type, for example, the mutually exclusive protein pair of B cells and T cells is CD19 and CD3.
[0012] Optionally, the data set to be used for dual-cell identification is a single-cell data set generated by RNA+ADT combined sequencing technology.
[0013] Optionally, the dual-cell identification method only identifies dual cells at the protein data level of the single-cell data set generated by the RNA+ADT combined sequencing technology.
[0014] Optionally, the clustering in step 1) is performed based on single-cell transcriptome sequencing data or single-cell membrane protein sequencing data;
[0015] Preferably, the dataset to be used for dual-cell identification is a single-cell dataset generated by CITE-seq technology.
[0016] Optionally, in step 2), it can be selected whether to perform background denoising on the protein data in each cell population, and the denoised data can be used to calculate the corresponding expression-negative cell population and expression-positive cell population for each protein in the mutually exclusive protein pair.
[0017] Optionally, the background denoising is performed using a scCDC method.
[0018] Optionally, the method of calculating the corresponding negative expression cell population and positive expression cell population of each protein in the mutually exclusive protein pair in step 2) includes:
[0019] For each protein in the mutually exclusive protein pair, the percentage of the protein detected in each cell population and the average expression level are calculated; the protein percentage and the average expression level are multiplied to obtain the penalty mean of the cell population;
[0020] The cell population with the highest penalty mean is used as the initial seed of the cell population combination, and is iteratively added to the cell population combination from high to low according to the penalty mean of the cell population; in each iteration, the COSG score of the combination of the added cell populations is calculated using the COSG method; when the COSG score does not increase after iteration or the frequency of the corresponding marker gene in the current cell population is less than the frequency threshold, the iteration is stopped; the cell population in the added cell population combination is the expression-positive cell population, and the remaining cell populations are the expression-negative cell populations;
[0021] The frequency threshold is 0.1-0.2.
[0022] Optionally, the COSG score is calculated using the following publicly available COSG algorithm formula:
[0023]
[0024] Among them, g i Represents each protein, G k represents the combination of cell populations, λ k Represents G k Expression pattern of putative marker genes for cell population combinations, λ t Represents except G k The expression pattern of putative marker genes of other cell populations outside the cell population, u represents the penalty factor.
[0025] Optionally, in the step 3), calculating the demarcation threshold between positive expression and negative expression of each protein in the mutually exclusive protein pair based on the data in the negative expression cell population and the positive expression cell population includes:
[0026] At least one of ROC and Naive Bayes method is used to calculate the demarcation threshold between positive expression and negative expression of each protein in the mutually exclusive protein pair;
[0027] The ROC method uses the labels of the positive and negative groups of the protein generated in step 3) and the expression value of the protein to generate the ROC curve and its object, and uses the Youden index to find the optimal demarcation threshold of the ROC curve, that is, the demarcation threshold of the positive expression and negative expression of the protein.
[0028] The naive Bayes method uses the label of the positive and negative groups of the protein generated in the step 3) as the dependent variable of the model, and the expression value of the protein as the independent variable of the model to construct a naive Bayes classification model. The model is then used to predict the expression value of the protein to obtain the label of the newly predicted positive and negative cells, and then the minimum value of the protein expression of the new positive group and the maximum value of the protein expression of the negative group are calculated, and the mean of the minimum and maximum values is taken as the demarcation threshold of the positive and negative expression of the protein.
[0029] Optionally, the method for identifying the two-cell data in the data set to be subjected to two-cell identification according to the demarcation threshold comprises:
[0030] The cell data whose expression level of each protein in the mutually exclusive protein pair was higher than the respective threshold were identified as double cell data.
[0031] A dual cell recognition system, comprising:
[0032] Clustering module: used to obtain a data set to be used for dual cell identification, and cluster the data set to obtain multiple cell populations;
[0033] A positive and negative cell population identification module for mutually exclusive protein pairs: using a cyclic iterative COSG method to determine the negative expression cell population and the positive expression cell population corresponding to each protein in the mutually exclusive protein pair in the cell population obtained in step 1);
[0034] A demarcation threshold calculation module: used to calculate the demarcation threshold of positive expression and negative expression of each protein in the mutually exclusive protein pair according to the data in the negative expression cell population and the positive expression cell population;
[0035] A two-cell data identification module is used to identify two-cell data in the data set to be subjected to two-cell identification according to the demarcation threshold.
[0036] An electronic device comprises a memory and a processor, wherein the memory and the processor are connected;
[0037] The memory stores computer instructions, and the processor executes the above-mentioned double-cell identification method by executing the computer instructions.
[0038] A computer-readable storage medium stores computer instructions, wherein the computer instructions are used to enable a computer to execute the above-mentioned double-cell identification method.
[0039] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0040] The dual cell identification method based on the COSG method of the present invention is more stable and effective in identifying positive and negative cell population modules, while the MLTiplet CITE-seq method in the prior art uses the mean to distinguish positive and negative cell populations and is easily affected by background noise, clustering and expression levels. In addition, each mutually exclusive protein of the MLTiplet CITE-seq in the prior art needs to contain a matching gene to identify dual cell populations. This method does not require matching proteins and genes, and only requires ADT data to perform dual cell identification analysis. In addition, it also provides support for background correction strategies and a variety of statistical discrimination methods, which can better adapt to different user groups.
[0041] The iterative COSG+ statistical strategy of this method can effectively divide the bimodal characteristics of proteins in different data sets and proteins, while the MLTiplet CITE-seq method of the prior art can only perform well in some samples and proteins. Secondly, this method also maintains good division performance on the corrected data. In addition, the threshold fluctuations of different proteins in the iterative COSG+ statistical method strategy of the present invention are more stable than MLTipletCITE-seq in clustering dimensions of different resolutions, which avoids the influence of subjective selection of resolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0043] Figure 1 Figure 2 is a scatter plot of protein expression threshold division without background correction; A is a scatter plot of expression threshold division of mutually exclusive protein pairs of CD3-CD20 and CD19-CD3 of TP1 sample, and the three figures correspond to different method strategies. Red represents the division threshold of each method, among which the first two are the strategies of this method, and the last one is the MLTiplet CITE-seq method strategy. The red dot ambiguous represents the TB double cells identified at the VDJ level; B is a scatter plot of expression threshold division of mutually exclusive protein pairs of TP2 samples; C is a scatter plot of expression threshold division of mutually exclusive protein pairs of TP3 samples; D is a scatter plot of expression threshold division of mutually exclusive protein pairs of TP3-rep samples; E is a scatter plot of expression threshold division of mutually exclusive protein pairs of pbmc5K samples; F is a scatter plot of expression threshold division of mutually exclusive protein pairs of pbmc10K samples.
[0044] Figure 2 It is a scatter plot of protein expression threshold division under background correction; A is the scatter plot of expression threshold division of CD3-CD20 and CD19-CD3 mutually exclusive protein pairs of TP1 sample, and the two figures correspond to different method strategies, red represents the division threshold of each method, and the red dot ambiguous represents the TB double cells identified at the VDJ level; B is the scatter plot of expression threshold division of mutually exclusive protein pairs of TP2 samples; C is the scatter plot of expression threshold division of mutually exclusive protein pairs of TP3 samples; D is the scatter plot of expression threshold division of mutually exclusive protein pairs of TP3-rep samples; E is the scatter plot of expression threshold division of mutually exclusive protein pairs of pbmc5K samples; F is the scatter plot of expression threshold division of mutually exclusive protein pairs of pbmc10K samples.
[0045] Figure 3 Figure 2 shows the threshold changes of various methods with different clustering resolutions; A shows the threshold fluctuations of different methods for TP1 samples; B shows the threshold fluctuations of different methods for TP2 samples; C shows the threshold fluctuations of different methods for TP3 samples; D shows the threshold fluctuations of different methods for TP3-rep samples; E shows the threshold fluctuations of different methods for pbmc5K samples; F shows the threshold fluctuations of different methods for pbmc10K samples.
[0046] Figure 4 4 is a structural block diagram of a dual-cell identification device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0047] Now, various exemplary embodiments of the present invention are described in detail, and this detailed description should not be considered as a limitation of the present invention, but should be understood as a more detailed description of certain aspects, characteristics and embodiments of the present invention. It should be understood that the terms described in the present invention are only for describing specific embodiments and are not used to limit the present invention.
[0048] The terms "include", "comprising", "having", "containing", etc. used in the present invention are open-ended terms, which mean including but not limited to. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as those commonly understood by those skilled in the art in the field of the present invention. Although the present invention describes only preferred methods, any methods similar or equivalent to those described in the present invention may also be used in the implementation or testing of the present invention.
[0049] Example 1
[0050] The present invention proposes a double cell recognition method, comprising the following steps:
[0051] Step 1) Obtain a data set to be subjected to dual cell identification, and then cluster the data set to obtain multiple cell populations. The data set is a single cell data set generated by RNA+ADT combined sequencing technology, and in this embodiment, a single cell data set generated by CITE-seq technology can generally be used.
[0052] Step 2) using the iterative COSG method to determine the negative expression cell population and the positive expression cell population corresponding to each protein in the mutually exclusive protein pair in the cell population obtained in step 1); the mutually exclusive protein pair is customized according to the biological cell type, for example, the mutually exclusive protein pair of B and T cells is CD19 and CD3;
[0053] Specifically include:
[0054] For each protein in the mutually exclusive protein pair, the percentage of the protein detected in each cell population and the average expression level are calculated; the protein percentage and the average expression level are multiplied to obtain the penalty mean of the cell population;
[0055] The cell population with the highest penalty mean is used as the initial seed, and it is iterated and added in descending order according to the penalty mean of the cell population; in each iteration, the COSG score of the combination of the added cell populations is calculated using the public COSG method. This public COSG method calculates the cosine similarity between the expression pattern of a feature in a certain cell population and its assumed marker pattern, thereby obtaining the specific COSG score of the feature in a certain cell population. The specific formula is as follows:
[0056]
[0057] Among them, g i Represents each protein, G k represents the combination of cell populations, λ k Represents G k Expression pattern of putative marker genes for cell population combinations, λ t Represents except G k The expression pattern of the assumed marker gene of other cell groups outside the cell group, u represents the penalty factor (in this embodiment, u takes the value of 1). Thus, the COSG specificity score of the cell group combination of the protein in each iteration can be obtained; when the COSG score does not increase after iteration or the frequency of the corresponding marker gene in the current cell group is less than 0.2 (the data is generally sparse, and it is generally believed in the industry that markers less than 0.1 or 0.2 are basically not very useful. In order to identify more accurately, 0.2 is selected), the iteration is stopped; the cell group in the added cell group combination is the expression positive cell group, and the remaining cell groups are the expression negative cell groups.
[0058] The COSG algorithm (Cosine Similarity-based Gene Identification) mentioned above is a marker gene identification method based on cosine similarity. It is mainly used in single-cell data analysis to identify cell marker genes more accurately and quickly. The COSG algorithm uses cosine similarity to measure the relationship between gene expression vectors. Cosine similarity evaluates their similarity by calculating the cosine value of the angle between two vectors in the vector space. Unlike traditional statistical test methods, cosine similarity compares the direction of two vectors rather than their absolute values, which means that even if two genes have the same expression pattern but different expression abundance, cosine similarity will consider them equivalent. The COSG algorithm is particularly suitable for cell type annotation in single-cell data analysis. In single-cell sequencing technology, cell type annotation is a key step, and traditional statistical methods may encounter challenges in identifying cell marker genes because statistical tests tend to identify candidate genes with systematic differences between two groups rather than true cell markers. The COSG algorithm can more accurately identify cell marker genes by calculating the cosine similarity of gene expression vectors, thereby improving the accuracy of cell type annotation.
[0059] Step 3) calculating the demarcation threshold between positive expression and negative expression of each protein in the mutually exclusive protein pair based on the data in the negative expression cell population and the positive expression cell population;
[0060] Step 4) identifying the dual-cell data in the data set to be subjected to dual-cell identification according to the demarcation threshold. Generally speaking, the cell data whose expression level of each protein in the mutually exclusive protein pair is higher than the respective threshold is identified as dual-cell data.
[0061] Furthermore, in step 1), the clustering is performed by RNA clustering or ADT clustering. The present invention supports RNA and ADT single-omics clustering. When the number of ADT-omics proteins is relatively small, RNA clustering can be used instead. ADT clustering will be used by default.
[0062] Furthermore, in the step 2), it can be selected whether to first use the scCDC method to perform background denoising on the protein data in each cell population, and the denoised data can be used to calculate the corresponding expression-negative cell population and expression-positive cell population for each protein in the mutually exclusive protein pair.
[0063] Furthermore, in the step 3), according to the data in the negative expression cell population and the positive expression cell population, calculating the demarcation threshold of the positive expression and the negative expression of each protein in the mutually exclusive protein pair includes:
[0064] At least one of ROC and Naive Bayes methods is used to calculate the threshold value of positive expression and negative expression of each protein in the mutually exclusive protein pair; the ROC method uses the label of the positive and negative groups of the protein generated in the step three) and the expression value of the protein to generate the ROC curve and its object, and uses the Youden index to seek the optimal threshold value of the ROC curve, that is, the threshold value of positive expression and negative expression of the protein. The Naive Bayes method uses the label of the positive and negative groups of the protein generated in the step three) as the dependent variable of the model, and the expression value of the protein as the independent variable of the model to construct a Naive Bayes classification model. Then, the model is used to predict the expression value of the protein to obtain the label of the newly predicted positive and negative cells, and then the minimum value of the protein expression of the new positive group and the maximum value of the protein expression of the negative group are calculated, and the average of the minimum and maximum values is taken as the threshold value of the positive and negative expression of the protein.
[0065] In order to better compare the advantages and disadvantages of the methods, we tested them on four sets of self-tested data and two sets of public CITE-seq data sets. The four sets of self-tested data include TP1, TP2, TP3 and TP3-rep (the source is the same healthy donor, and the sampling interval is three days each time. Three peripheral blood mononuclear cell (PBMCs) samples are obtained, namely TP1, TP2, and TP3. Among them, the third node sample (TP3) initially failed to pass QC, and then TP3-rep was re-measured). The public data includes pbmc5K and pbmc10K, which are from the 10x Genomics official website. Since the 6 sets of CITE-seq data sets all contain matching VDJ sequencing data, in order to better reflect the performance difference, we selected the common B and T cell mutually exclusive protein pairs CD19-CD3 and CD20-CD3. It can be seen that the method of the present invention can effectively divide the bimodal characteristics of proteins in different data sets and proteins, while MLTiplet CITE-seq can only perform well in some samples and proteins ( Figure 1 ). Secondly, this method also maintains good segmentation performance on the corrected data of the scCDC method ( Figure 2 ).
[0066] In addition, the present invention further analyzes that the threshold fluctuations of different proteins in the iterative COSG+ statistical method strategy at different resolutions of clustering dimensions are more stable than those of MLtiplet CITE-seq ( Figure 3 ), which avoids the influence of subjective choice of resolution.
[0067] Compared with the existing MLTiplet CITE-seq method, this method is more stable in identifying positive and negative cell population modules. MLTiplet CITE-seq uses the mean to distinguish positive and negative cell populations, which is easily affected by background noise, clustering and expression levels. In addition, each mutually exclusive protein of MLTiplet CITE-seq needs to contain a matching gene to identify dual cell populations. This method does not require matching proteins and genes, but only requires ADT data for dual cell identification and analysis. In addition, it also provides support for background correction strategies and a variety of statistical discrimination methods, which can better adapt to different user groups.
[0068] The present invention also proposes a dual-cell recognition system, comprising:
[0069] Clustering module: used to obtain a data set to be used for dual cell identification, and cluster the data set to obtain multiple cell populations;
[0070] A positive and negative cell population identification module for mutually exclusive protein pairs: used to obtain the cell population obtained by the clustering module and calculate the corresponding negative expression cell population and positive expression cell population for each protein in the mutually exclusive protein pair using the cyclic iteration COSG method;
[0071] A demarcation threshold calculation module: used to calculate the demarcation threshold of positive expression and negative expression of each protein in the mutually exclusive protein pair according to the data in the negative expression cell population and the positive expression cell population;
[0072] A two-cell data identification module is used to identify two-cell data in the data set to be subjected to two-cell identification according to the demarcation threshold.
[0073] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0074] The dual-cell identification device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0075] An embodiment of the present invention further provides a computer device having the above-mentioned dual-cell identification device.
[0076] See also Figure 4, is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface).
[0077] In some optional embodiments, if desired, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 A processor 10 is taken as an example.
[0078] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0079] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the double-cell identification method shown in the above embodiment.
[0080] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0081] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0082] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means.
[0083] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0084] The computer device also includes a communication interface, which is used for the computer device to communicate with other devices or a communication network.
[0085] An embodiment of the present invention also provides a computer-readable storage medium, and the above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented by downloading through a network and originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware.
[0086] The storage medium may be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid state drive, etc.; further, the storage medium may also include a combination of the above-mentioned types of memories. It is understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, processor or hardware, the method shown in the above embodiment is implemented.
[0087] Embodiments of the present application may also provide a computer program product, including computer program instructions, which, when executed by a processor, cause the processor to perform the steps in the above method. The computer program product may be written in any combination of one or more programming languages to perform program codes for performing the operations of the disclosed embodiments, wherein the programming languages include object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language 10 or similar programming languages. The program code may be executed entirely on a user computing device, partially on a user computing device, as an independent software package, partially on a user computing device, partially on a remote computing device, or entirely on a remote computing device or server.
[0088] The above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them. Although the embodiments of the present invention are described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A double cell identification method, characterized in that: The steps include: Step 1) obtaining a data set to be subjected to dual-cell identification, wherein the data set to be subjected to dual-cell identification includes protein data generated by single-cell transcriptome sequencing and single-cell membrane protein sequencing; then clustering the data set to be subjected to dual-cell identification to obtain multiple cell populations; Step 2) determining mutually exclusive protein pairs according to the cell types of the data set to be subjected to dual cell identification, and then determining the negative expression cell population and the positive expression cell population corresponding to each protein in the mutually exclusive protein pair in the cell population obtained in step 1) using the cyclic iterative COSG method; Step 3) calculating the demarcation threshold of the positive and negative expression of each protein in the mutually exclusive protein pair according to the data in the negative expression cell population and the positive expression cell population; Step 4) identifying the double-cell data in the data set to be subjected to double-cell identification according to the demarcation threshold.
2. The double cell identification method according to claim 1, characterized in that: In the step 1), clustering is performed based on single-cell transcriptome sequencing data or single-cell membrane protein sequencing data; Preferably, the dataset to be used for dual-cell identification is a single-cell dataset generated by CITE-seq technology.
3. The double cell identification method according to claim 1, characterized in that: In the step 2), background denoising is first performed on the protein data in each cell population, and the denoised data is used to calculate the corresponding negative expression cell population and positive expression cell population for each protein in the mutually exclusive protein pair; Preferably, the background denoising is performed using a scCDC method.
4. The double cell identification method according to claim 1, characterized in that: The method for calculating the corresponding negative expression cell population and positive expression cell population of each protein in the mutually exclusive protein pair in step 2) comprises: For each protein in the mutually exclusive protein pair, the percentage of the protein detected in each cell population and the average expression level are calculated; the protein percentage and the average expression level are multiplied to obtain the penalty mean of the cell population; The cell population with the highest penalty mean is used as the initial seed of the cell population combination, and is iteratively added to the cell population combination from high to low according to the penalty mean of the cell population; in each iteration, the COSG score of the combination of the added cell populations is calculated using the COSG method; when the COSG score does not increase after iteration or the frequency of the corresponding marker gene in the current cell population is less than the frequency threshold, the iteration is stopped; the cell population in the added cell population combination is the expression-positive cell population, and the remaining cell populations are the expression-negative cell populations; Preferably, the frequency threshold is 0.1-0.
2.
5. The double cell identification method according to claim 4, characterized in that: The COSG score is calculated using the following publicly available COSG algorithm formula: Among them, g i Represents each protein, G k represents the combination of cell populations, λ k Represents G k Expression pattern of putative marker genes for cell population combinations, λ t Represents except G k The expression pattern of putative marker genes of other cell populations outside the cell population, u represents the penalty factor.
6. The double cell identification method according to claim 1, characterized in that: The method of calculating the demarcation threshold of positive and negative expression of each protein in the mutually exclusive protein pair according to the data in the negative expression cell population and the positive expression cell population in the step 3) comprises: using at least one method of ROC and naive Bayes to calculate the demarcation threshold of positive expression and negative expression of each protein in the mutually exclusive protein pair; The ROC method uses the labels of the positive and negative groups of the protein generated in step 3) and the expression value of the protein to generate a ROC curve and its object, and uses the Youden index to find the optimal demarcation threshold of the ROC curve, that is, the demarcation threshold of the positive expression and negative expression of the protein; The naive Bayes method uses the label of the positive and negative groups of the protein generated in the step three) as the dependent variable of the model, and the expression value of the protein as the independent variable of the model to construct a naive Bayes classification model; then the model is used to predict the expression value of the protein to obtain the label of the newly predicted positive and negative cells, and then the minimum value of the protein expression of the new positive group and the maximum value of the protein expression of the negative group are calculated, and the average of the minimum and maximum values is taken as the demarcation threshold of the positive and negative expression of the protein.
7. The double cell identification method according to claim 1, characterized in that: The method for identifying the double-cell data in the data set to be subjected to double-cell identification according to the demarcation threshold comprises: Cell data in which the expression level of each protein in the mutually exclusive protein pair was higher than the respective thresholds were identified as double cell data.
8. A dual cell recognition system, characterized in that: include: Clustering module: used to obtain a data set to be used for dual-cell identification, wherein the data set to be used for dual-cell identification includes protein data generated by single-cell transcriptome sequencing and single-cell membrane protein sequencing; Then, the data set to be subjected to dual-cell identification is clustered and grouped to obtain multiple cell populations; A module for identifying the positive and negative cell populations of mutually exclusive protein pairs: used to determine the mutually exclusive protein pairs according to the cell types of the data set to be subjected to dual cell identification, and then determine the negative expression cell population and the positive expression cell population corresponding to each protein in the mutually exclusive protein pair using the iterative COSG method in the cell population obtained in step 1); A demarcation threshold calculation module: used to calculate the demarcation threshold of positive expression and negative expression of each protein in the mutually exclusive protein pair according to the data in the negative expression cell population and the positive expression cell population; A two-cell data identification module is used to identify two-cell data in the data set to be subjected to two-cell identification according to the demarcation threshold.
9. An electronic device, characterized in that: comprising a memory and a processor, wherein the memory and the processor are connected; The memory stores computer instructions, and the processor executes the dual-cell identification method according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the double-cell identification method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Single cell identification method and device based on graph neural network self-supervised clustering
CN115798593A
PDX-based single cell transcriptome data analysis method, system, equipment and medium
CN116153401A
Cell type high-precision identification method based on artificial intelligence
CN116451077A
Dual cytokine fusion proteins comprising multi-subunit cytokines
CN118715239A