Method for identifying tissue-derived cells in body fluid based on single-cell sequencing technology

By employing single-cell sequencing technology and Rogers regression model, the universality and throughput issues of tissue-derived cell identification methods in body fluids have been resolved, enabling high-dimensional data acquisition and supporting clinical applications.

WO2025222503A1PCT designated stage Publication Date: 2025-10-30SHENZHEN HUADA GENE INST

Patent Information

Application Number
PCT/CN2024/090137
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-26
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing methods for identifying tissue-derived cells in body fluids have poor universality, low throughput, and cannot acquire high-dimensional data information.

Method used

By employing single-cell sequencing technology, combined with reference component analysis and Rogers regression model, a tissue-derived cell prediction model was constructed using high-throughput single-cell RNA sequencing data to achieve specific identification of tissue cells in body fluids.

Benefits of technology

It enables high-throughput and universally applicable tissue and cell identification, and can obtain high-dimensional cellular transcriptome information to support clinical disease diagnosis and treatment guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024090137_30102025_PF_FP_ABST
    Figure CN2024090137_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the field of tissue-derived cell identification, in particular to a method for identifying tissue-derived cells in the body fluid based on single-cell sequencing technology. In the present invention, on the basis of high-throughput single-cell RNA sequencing data of cell samples obtained from body fluids, by means of reference component analysis (RCA), expression of known tissue-related marker genes, and a tissue-derived cell prediction model constructed based on logistic regression, the specific identification of tissue-derived cells in body fluids is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

A method for identifying tissue-derived cells in body fluids based on single-cell sequencing technology. Technical Field

[0001] This invention relates to the field of tissue-derived cell identification, and in particular to a method for identifying tissue-derived cells in body fluids based on single-cell sequencing technology. Background Technology

[0002] Currently, the main methods for identifying tissue-derived cells in body fluids include: immunomagnetic bead analysis for positive screening, microfluidic positive screening, RT-PCR, immunofluorescence, fluorescence in situ hybridization, and next-generation sequencing.

[0003] Table 1

[0004] Current detection technologies mostly rely on a few known cellular markers for identification, resulting in relatively poor universality in identifying cells from different tissue origins. Furthermore, some methods cannot achieve single-cell identification or have low throughput. In addition, the information obtained after detection is relatively limited, mostly consisting of simple identification, counting, or phenotypic analysis. In recent years, the development of single-cell omics technology has provided a powerful tool for characterizing cell omics, offering features at the transcriptomic, genomic, and epigenomic levels. This facilitates further research into cell phenotypes and their clinical application in disease monitoring and treatment guidance.

[0005] Table 2

[0006] In summary, the following problems urgently need to be addressed: 1) The methods for identifying tissue-derived cells in body fluids have poor universality and rely on the expression of specific proteins on the surface of specific cells and specific physical properties; 2) The identification throughput is low; 3) The acquisition of high-dimensional data information of single cells identified as tissue-derived cells.

[0007] Summary of the Invention

[0008] In view of this, the present invention provides a method for identifying tissue-derived cells in body fluids based on single-cell sequencing technology. The present invention utilizes high-throughput single-cell RNA sequencing data from cell samples obtained from body fluids, and employs Reference Component Analysis (RCA), known tissue-related marker gene expression, and a tissue-derived cell prediction model constructed based on Logistic Regression to achieve specific identification of tissue cells in body fluids.

[0009] To achieve the above-mentioned objectives, the present invention provides the following technical solution:

[0010] In a first aspect, the present invention provides a method for identifying tissue-derived cells based on single-cell sequencing technology, comprising the following steps:

[0011] S1: Take the enriched cell suspension from the sample to be tested, construct a single-cell RNA sequencing library, sequence the cells, preprocess the sequencing data, and obtain the single-cell expression matrix.

[0012] S2: Project the single-cell expression matrix described in S1 onto the reference dataset to obtain the correlation matrix;

[0013] S3: Based on the trained Rogers regression prediction model, the prediction probability matrix is ​​obtained according to the single-cell expression matrix of the sample to be tested;

[0014] S4: Obtain the marker gene expression matrix of cell-related marker genes in the sample to be tested;

[0015] S5: Perform a comprehensive cluster analysis on the correlation matrix, prediction probability matrix, and marker gene expression matrix of the sample to be tested to obtain the identification result of the sample to be tested.

[0016] In some specific embodiments of the present invention, the dataset mentioned in S2 includes, but is not limited to:

[0017] The BioGPS database built into the RCA software package comes from two large-sample transcriptome datasets covering different tissue and cell types: HumanU133A / GNF1H GeneAtlas and the Primary CellAtlas; or

[0018] The BioGPS database contains two tumor-related datasets: Human NCI60 Cell Lines and Human Primary Tumors (U95).

[0019] In some specific embodiments of the present invention, S2 specifically includes:

[0020] Based on the clustering of cells to which they belong;

[0021] Obtain the known cell types of the reference sample;

[0022] The cell type to which each group belongs is defined based on the correlation between each group and the cell type of the reference sample.

[0023] In some specific embodiments of the present invention, the method for obtaining the group of classes described in S2 includes the following steps:

[0024] S2-1: Standardize the relative expression levels of each reference transcriptome dataset;

[0025] S2-2: Select feature genes based on gene-specific expression in the dataset;

[0026] S2-3: By calculating the log per cell from scRNA–seq 10 Standardized feature gene expression vectors and reference sets for each sample's logarithm 10 The Pearson correlation coefficient between the characteristic gene expression vectors projects the expression profile of each single cell onto each sample in the reference set; then the correlation coefficient is raised to the fourth power and centered and normalized to define the mapping vector between each cell and each reference sample.

[0027] S2-4: The cutreeDynamic function in the WGCNA package is used to perform hierarchical clustering on the mapping vector of each cell. The resulting tree diagram identifies the cell groups, obtains the group corresponding to each cell, and obtains the correlation matrix.

[0028] In some specific embodiments of the present invention, the method for obtaining the group of classes described in S2 includes the following steps:

[0029] S2-1: Relative expression levels are standardized for each reference transcriptome dataset, which is calculated by dividing the gene read count for each sample by the total number of reads for that sample, as shown in Formula 1:

[0030] S2-2: Select feature genes based on gene-specific expression in the dataset; for example, the logarithm of the expression level of a gene in any sample with respect to the median expression level across all samples. 10 If the ratio is greater than 1, it is considered a characteristic gene, and the log of the reference data is finally obtained. 10 The standardized feature matrix is ​​shown in Formula 2:

[0031] S2-3: By calculating the log per cell from scRNA–seq 10 Standardized feature gene expression vectors and reference sets for each sample's logarithm 10 The Pearson correlation coefficient between the characteristic gene expression vectors projects the expression profile of each single cell onto each sample in the reference set; then the correlation coefficient is raised to the fourth power and centered and normalized to define the mapping vector between each cell and each reference sample.

[0032] S2-4: The cutreeDynamic function in the WGCNA package is used to perform hierarchical clustering on the mapping vector of each cell. The parameters deepSplit=1 and mingroupSize=5 are set. The cell groups are identified in the final generated dendrogram, and the corresponding group of each cell is obtained to obtain the correlation matrix.

[0033] In some specific embodiments of the present invention, S3 includes the following steps:

[0034] Collect single-cell data from human fluids or solid tissues;

[0035] Obtain gene expression profiles related to the training set;

[0036] Obtain the optimal number of features;

[0037] Construct Rogers regression and validate the evaluation.

[0038] In some specific embodiments of the present invention, S3 specifically includes the following steps:

[0039] S3-1: Collect single-cell data from human body fluids or solid tissues, and obtain single-cell expression matrices and / or cell type annotation information for each cell;

[0040] S3-2: Data preprocessing, including standardization and / or normalization methods;

[0041] S3-3: Obtain the expression profile of tumor-related genes in the training set. Tumor-related genes in different tissues are obtained from the CancerSEA database.

[0042] S3-4: Randomly group the training set cells for cross-validation; perform a t-test on each gene as a basis for separability, and use the individual optimal feature combination method to obtain the optimal number of features;

[0043] S3-5: Construct Rogers regression using the entire training set.

[0044] In some specific embodiments of the present invention, S3 specifically includes the following steps:

[0045] S3-1: Collect single-cell data from human body fluids or solid tissues, and obtain single-cell expression matrices and / or cell type annotation information for each cell;

[0046] S3-2: Data preprocessing, including standardization and / or normalization methods;

[0047] The standardization method includes calculating Transcripts Per Kilobase of exon model per Million mapped reads (TPM) or counts per million (CPM);

[0048] The normalization method standardizes the maximum and minimum values ​​and scales them to [0, 10]; missing values ​​are padded with 0.

[0049] S3-3: Obtain the expression profile of tumor-related genes in the training set. Tumor-related genes in different tissues are obtained from the CancerSEA database.

[0050] S3-4: Randomly divide the training set cells into 5 parts for cross-validation; perform a t-test on each gene as a basis for separability, and use the individual optimal feature combination method to obtain the optimal number of features;

[0051] S3-5: Construct Rogers regression using the entire training set, and perform validation evaluation using the validation set; g(p) = β0 + β1x1 + ... + β k x k Where p represents a given feature x1, ..., x k The probability of identifying the cell as a tissue-derived cell is calculated. The logit function is chosen as the correlation function g(p).

[0052] In some specific embodiments of the present invention, the z-score standardization is as follows:

[0053] The variable X is the feature, E|X| is the expected value of X, and σ(X) is the standard deviation.

[0054] In some specific embodiments of the present invention, the sample to be tested includes, but is not limited to, body fluids or solid tissues.

[0055] Based on the above research, in a second aspect, the present invention also provides an identification model obtained by the method described above.

[0056] Thirdly, the present invention also provides the application of the identification model in any of the following:

[0057] (I) Identification of tissue-derived cells;

[0058] (II) Preparation of products for identifying tissue-derived cells.

[0059] Fourthly, the present invention also provides an apparatus, wherein the apparatus has a built-in module or component corresponding to the identification model.

[0060] Fifthly, the present invention also provides a computer-readable storage medium comprising a stored program, wherein the program, when executed, controls the device on which the storage medium is located to perform the method.

[0061] In a sixth aspect, the present invention also provides a processor for running a program, wherein the program executes the method during runtime.

[0062] In a seventh aspect, the present invention also provides an apparatus comprising: a memory and a processor; wherein one or more computer programs are stored in the memory, the one or more computer programs comprising instructions; and when the instructions are executed by the processor, the electronic device performs the method of the present invention.

[0063] This invention provides a method for identifying tissue-derived cells in body fluids based on single-cell sequencing technology. Based on high-throughput single-cell RNA sequencing data from cell samples obtained from body fluids, this invention achieves specific identification of tissue-derived cells in body fluids through Reference Component Analysis (RCA), known tissue-related marker gene expression, and a tissue-derived cell prediction model constructed based on Logistic Regression. The beneficial effects of this invention include, but are not limited to:

[0064] This invention provides a high-throughput, highly universal method for identifying tissue-derived cells in body fluids. It combines high-throughput single-cell omics sequencing technology with a highly sensitive tissue-derived cell identification model to identify tissue-derived cells in body fluids. This overcomes the limitations of traditional methods that rely on cell surface-specific protein expression and specific physical properties for identifying tissue-derived cells in body fluids. Furthermore, by combining different single-cell omics technologies, transcriptomic information of tissue cells can be obtained, which has clinical significance for non-invasive diagnosis, monitoring, treatment efficacy evaluation, and guiding personalized precision medicine for clinical diseases (such as tumors). Attached Figure Description

[0065] Figure 1 shows the UMAP dimensionality reduction visualization clustering results;

[0066] Figure 2 shows a scatter plot of predicted probabilities for tissue-derived cells (non-immune cells);

[0067] Figure 3 shows cell annotation and expression of genes related to tissue-derived cells;

[0068] Figure 4 shows a heatmap of clustering based on comprehensive features;

[0069] Figure 5 shows the detection and clustering results of a mixed sample of peripheral blood and tumor cell lines. Detailed Implementation

[0070] This invention discloses a method for identifying tissue-derived cells in body fluids based on single-cell sequencing technology. Those skilled in the art can refer to this document and appropriately modify the process parameters to achieve the desired result. It is particularly important to note that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The methods and applications of this invention have been described through preferred embodiments. Those skilled in the art will clearly be able to modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit, and scope of this invention to realize and apply the technology of this invention.

[0071] The raw materials and reagents used in the method for identifying tissue-derived cells in body fluids based on single-cell sequencing technology provided by this invention are all commercially available. The invention is further illustrated below with reference to embodiments:

[0072] Terminology Explanation

[0073] The single cells involved in this invention are derived from living organisms, specifically tissue-derived cells, and prepared using conventional methods. Single cells can be obtained through in vitro culture or directly isolated from clinical samples such as body fluids or solid tissues (including plasma, serum, cerebrospinal fluid, bone marrow, lymph, ascites, pleural effusion, oral fluid, skin tissue, respiratory tract, digestive tract, reproductive tract, urinary tract, tears, saliva, blood cells, stem cells, and tumors). The samples can reflect specific cellular states, such as cell proliferation, cell differentiation, apoptosis / death, disease state, external stimulus state, and developmental stage.

[0074] The method provided by this invention can be implemented in a terminal environment that may include one or more of the following components: a processor, a memory, and a display screen. The memory stores at least one instruction, which is loaded and executed by the processor to implement the method described in the following embodiments.

[0075] A processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts of the terminal, and performs various functions and processes data by running or executing instructions, programs, code sets or instruction sets stored in memory, and by calling data stored in memory.

[0076] The display screen is used to show the user interface of each application.

[0077] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus, electronic devices (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable device, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.

[0078] In a typical configuration, an electronic device includes one or more processors (CPUs), memory, and a bus. The electronic device may also include input / output interfaces, network interfaces, etc.

[0079] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM, and memory includes at least one memory chip. Memory is an example of computer-readable media.

[0080] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0081] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0082] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0083] The present invention will be further illustrated below with reference to the embodiments:

[0084] Example 1: High-throughput single-cell RNA sequencing

[0085] Single-cell RNA sequencing libraries were constructed from enriched cell suspensions obtained from human fluid samples such as blood, ascites, and cerebrospinal fluid using a high-throughput droplet microfluidic system. After quality control, the libraries were sequenced, and the sequencing data were preprocessed to obtain single-cell expression profiles.

[0086] The throughput of this method mainly refers to its ability to perform batch analysis of single cells, with no limit on the sample size. Theoretically, it can analyze hundreds of thousands to millions of cells. This invention does not impose any limitations on this.

[0087] Example 2 Reference Component Analysis

[0088] Based on the single-cell expression profile obtained in Example 1, the single-cell expression profile to be analyzed was projected onto two large-sample transcriptome datasets covering different tissues and cell types from the BioGPS database built into the RCA software package (https: / / github.com / GIS-SP-Group / RCA): HumanU133A / GNF1H Gene Atlas and the Primary Cell Atlas, to perform cell clustering and cell type identification.

[0089] In addition, this method integrates a further improved set of reference maps from two tumor-related datasets in the BioGPS database: Human NCI60 Cell Lines and Human Primary Tumors (U95).

[0090] 1. Reference set construction

[0091] First, the relative expression levels of each reference transcriptome dataset are standardized, i.e., the number of reads corresponding to each gene in each sample is divided by the total number of reads in that sample (Formula 1). Then, feature genes are selected based on gene-specific expression in the dataset. The expression level of a gene in any sample is logarithmic with respect to the median expression level across all samples. 10 If the ratio is greater than 1 (Formula 2), it is considered a characteristic gene, and the log of the reference data is finally obtained. 10 Standardized feature matrix.

[0092] Table 3 Standardized Reference Feature Matrix

[0093] The matrix contains rows representing genes, columns representing reference samples, and numbers representing the standardized expression levels of characteristic genes.

[0094] 2. Single-cell data mapping to reference set

[0095] By calculating the log for each cell from scRNA–seq 10 Standardized feature gene expression vectors and reference sets for each sample's logarithm 10 The Pearson correlation coefficient between the expression vectors of characteristic genes (expression vectors in the reference set, one vector per cell type, where the value of one dimension of the vector represents the expression level of the corresponding characteristic gene) projects the expression profile of each single cell onto each sample in the reference set. Only selected characteristic genes are used when calculating the correlation coefficient. For each single cell, the correlation coefficient is then raised to the fourth power and centered and normalized to define the mapping vector of each cell to each reference sample (the mapping vector of a cell is a vector composed of the Pearson coefficients of a cell relative to each reference sample; for example, if the coefficient with the first reference sample is 'a', and the coefficients with the second and third reference samples are 'b' and 'c', the mapping vector of that cell is 'abc').

[0096] The `cutreeDynamic` function from the WGCNA package was used to perform hierarchical clustering on the mapping vector of each cell. With parameters `deepSplit=1` and `mingroupSize=5`, the resulting dendrogram identified cell clusters, obtaining the cluster corresponding to each cell. The cell type to which each cluster belonged was defined based on the Pearson coefficient (correlation) between each cell cluster and the reference sample cell type.

[0097] Example 3: Constructing a predictive model for tissue cells in body fluids based on Rogerster regression.

[0098] 1. Collect single-cell data from human fluids or solid tissues.

[0099] Obtain publicly available data on human fluids or solid tissues, including single-cell expression matrices (each single cell used for training requires its gene expression and cell type annotation information) and cell type annotation information for each cell.

[0100] Table 4

[0101] 2. Data Preprocessing:

[0102] Standardization method: Calculate Transcripts Per Kilobase of exon model per Million mapped reads (TPM) or counts per million (CPM);

[0103] Normalization method: standardize the maximum and minimum values ​​and scale them to [0,10]; fill missing values ​​with 0.

[0104] 3. Obtain the expression profile of related genes in the training set tissues. Tumor-related genes in different tissues were obtained from the CancerSEA database.

[0105] 4. The training set cells were randomly divided into 5 parts for cross-validation. A t-test was performed on each gene as a criterion for separability. The optimal number of feature variables was obtained using the individual optimal feature combination method.

[0106] 5. Construct a Rogers regression model using the entire training set, obtain the coefficients (β) corresponding to each feature variable in the model, and then use the validation set for validation and evaluation.

[0107] g(p) = β0 + β1x1 + ... + β k x k Where p represents a given feature x1, ..., x k The probability of identifying the cell as a tissue-derived cell is calculated. The logit function is chosen as the correlation function g(p).

[0108] The β values ​​corresponding to some feature genes (β is the coefficient of each feature variable in the model function) are as follows:

[0109] Table 5

[0110] Example 4: Based on the expression of cell-related marker genes

[0111] In addition to universally applicable identification methods, this method also incorporates the expression of cell-related marker genes to further improve the accuracy of the identification method and assist in verifying the identification results.

[0112] Example 5: Identification of Tissue Cells Based on Comprehensive Characteristics

[0113] The cell type correlations analyzed from the above relevant reference components, the predicted probabilities based on the Rogers regression tissue cell prediction model, and the expression of cell-related marker genes were z-score standardized and then subjected to comprehensive hierarchical cluster analysis for the identification of tissue-derived cells in body fluids. Specifically, the correlation matrix between each cell and different samples, along with the predicted probabilities and marker gene matrices, were merged into a new feature matrix. Furthermore, the behavioral characteristics, listed as cells, were z-score standardized.

[0114] The variable X is the feature, E|X| is the expected value of X, and σ(X) is the standard deviation.

[0115] Based on the hierarchical clustering results, the characteristics of different cell groups are compared, and groups with high tissue-derived cell correlation, high model prediction probability, and high expression of related marker genes are identified as tissue-derived cells.

[0116] Example 6

[0117] 1 Experimental Sample

[0118] Cell samples obtained from portal vein whole blood of hepatocellular carcinoma patients after negative enrichment with CD45 immunomagnetic beads and sequencing quality control yielded a total of 1660 cells.

[0119] 2. Single-cell library preparation and sequencing

[0120] This embodiment uses the DNBelab C4 series single-cell library preparation kit (MGI Tech) according to a standard procedure for single-cell RNA sequencing library preparation. The specific process mainly includes droplet generation, demulsification, collection of mRNA-capturing magnetic beads, reverse transcription, cDNA amplification, and purification, ultimately preparing a single-cell RNA library to be tagged from the single-cell suspension. Subsequently, sequencing libraries are constructed according to the manufacturer's protocol. The sequencing libraries are quantified using the Qubit ssDNA analysis kit (Thermo Fisher Scientific). Sequencing of the libraries is performed using a DIPSEQ T1 sequencer.

[0121] 3 Data Preprocessing

[0122] The original FASTQ file was filtered and single-cell barcode corresponding reads were separated using PISA (v0.2; https: / / github.com / shiquan / PISA) with default parameters. Then, STAR (v2.5.3) was used to align the reads with the GRCh38 human reference genome. Reads were sorted from highest to lowest based on the number of RNA molecular tags (UMIs), retaining cells with a UMI count greater than 1 / 10 of the tenth-ranked cell. The aligned reads were then used to generate a cell × gene UMI count matrix using PISA based on cell tags and RNA molecular tags (UMIs).

[0123] Table 6 Cell × Gene UMI Count Matrix

[0124] 4. Cell cluster analysis

[0125] The obtained cell expression matrix was clustered and visualized using the Seurat software package (Figure 1).

[0126] 5-cell RCA reference component analysis

[0127] The obtained cell expression matrix is ​​mapped and clustered onto a self-constructed reference cell set.

[0128] 6. Rogers regression predicts tissue-derived cells

[0129] The cell expression matrix to be analyzed was input into the trained prediction model, and cell group 6 was predicted to be tissue cells with the highest probability (Figure 2).

[0130] 7. Cell type annotation and expression of cell-related marker genes

[0131] Cell types were annotated based on the genes expressed specifically for each cell type. Group 6 expressed the KRT8 and KRT18 epithelial cell marker genes but did not exhibit classic hepatocyte characteristics (ALB expression) (Figure 3). Group 6 does not have classic hepatocyte characteristics but has epithelial characteristics, which may indicate its origin as liver cancer cells or intrahepatic bile duct epithelial cells.

[0132] 8. Comprehensive Feature Identification

[0133] After z-score normalization of the above-mentioned features, including the correlation of reference cell types, the prediction probability of cell prediction models, and the expression of cell-related marker genes, a comprehensive cluster analysis was performed. The hierarchical clustering results revealed a cell group with unique characteristics, mainly composed of Cluster6 cells. This cell group had a high model prediction probability and a high mapping relationship with liver cells (L40_Hepatocyte and Fetalliver) and platelets (L52_Platelet, a characteristic of circulating tumor cells), as well as CNV-related features (Figure 4). Other cells were identified as immune cells and erythrocytes based on the clustering and RCA reference component analysis results (Figures 2 and 4). Based on the clustering results of these multi-dimensional information, we ultimately identified Cluster6 cells as liver-derived cells (accounting for 4.96%).

[0134] 9. Capture Rate Detection and Verification

[0135] Peripheral blood PBMC samples were mixed with 10, 50, and 100 Huh-7 hepatocellular carcinoma cells, respectively, to simulate actual test samples. The mixed samples were then analyzed using the aforementioned testing procedure. Results showed that tumor cells could be detected in all six mixed samples with high detection rates (the number of cells used was estimated using the dilution method). Even at a positive cell level of 0.2%, a high detection rate was maintained (Figure 5).

[0136] In summary, the above analysis demonstrates that this invention can identify tissue-derived cells in body fluid samples using a simplified negative screening method. It exhibits high sensitivity and throughput, allowing for the simultaneous detection of large numbers of cells (depending on single-cell sequencing throughput), while maintaining a high detection rate even for rare tissue-derived cells (approximately 0.2%). Based on single-cell transcriptomics, reference set mapping and regression models enable the detection of diverse samples, overcoming limitations imposed by specific biomarkers and demonstrating high universality. Furthermore, the transcriptional data obtained from cell identification can be used for downstream analysis as needed.

[0137] Comparative Example 1

[0138] Positive sorting using the single antibody EpCAM to identify epithelial-derived cells is the most commonly used detection method, but its detection efficiency is not high. This is because not all cells express this molecule. For example, during tumor metastasis, epithelial-mesenchymal transition occurs, changing the characteristics of epithelial cells to those of mesenchymal cells. These cells may be missed due to decreased or absent EpCAM expression. The table below compares the identification of circulating tumor cells in body fluids based on specific biomarkers.

[0139] Table 7 Comparison of Protein Markers for Tumor Detection in Body Fluid Circulation

[0140] Comparative Example 2

[0141] Using EpCAM and CSV antibodies, via The CTC detection system was used to enrich circulating tumor cells (CTCs) in body fluids. A total of 853 CTC test results from 690 cancer patients and 72 healthy individuals were collected and analyzed.

[0142] Table 8. Fluid circulation tumor detection in different solid tumors using EpCAM or CSV antibodies.

[0143] The results showed that EpCAM had the highest CTC detection rate in colorectal cancer (84.09%), followed by breast cancer (78.32%), while liver cancer (25%), pancreatic cancer (32.5%), and ovarian cancer (33.33%) showed relatively poor detection rates. Overall, the EpCAM antibody had a CTC detection rate of less than 40% in 14 common solid tumors. CSV had the highest CTC detection rate in sarcoma (90%), followed by brain cancer (85.71%), urothelial carcinoma (84.62%), ovarian cancer (83.33%), and breast cancer (81.82%). In all 14 solid tumors, the CSV antibody had a CTC detection rate exceeding 60%, but the false positive rate was high in healthy individuals. Except for colorectal cancer, CSV outperformed EpCAM in CTC detection in most solid tumors.

[0144] Therefore, it is evident that methods for detecting tissue-derived cells based on a single marker have poor universality and can only identify enriched and screened samples, resulting in low throughput.

[0145] The above provides a detailed description of the method for identifying tissue-derived cells in body fluids based on single-cell sequencing technology, as provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are merely for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this invention.

Claims

1. A method for identifying cells based on single-cell sequencing technology, characterized in that, Includes the following steps: S1: Take the enriched cell suspension from the sample to be tested, construct a single-cell RNA sequencing library, sequence the cells, preprocess the sequencing data, and obtain the single-cell expression matrix. S2: Project the single-cell expression matrix onto the reference dataset to obtain the correlation matrix; S3: Based on the trained Rogers regression prediction model, the prediction probability matrix is ​​obtained according to the single-cell expression matrix of the sample to be tested; S4: Obtain the marker gene expression matrix of cell-related marker genes in the sample to be tested; S5: Perform a comprehensive cluster analysis on the correlation matrix, prediction probability matrix, and marker gene expression matrix of the sample to be tested to obtain the identification result of the sample to be tested.

2. The method as described in claim 1, characterized in that, The method for obtaining the group described in S2 includes the following steps: S2-1: Standardize the relative expression levels of each reference transcriptome dataset; S2-2: Select feature genes based on gene-specific expression in the dataset; S2-3: By calculating the log per cell from scRNA–seq 10 Standardized feature gene expression vectors and reference sets for each sample's logarithm 10 The Pearson correlation coefficient between the characteristic gene expression vectors projects the expression profile of each single cell onto each sample in the reference set; the correlation coefficient is then raised to the fourth power and centered and normalized to define the mapping vector between each cell and each reference sample. S2-4: Perform hierarchical clustering on the mapping vector of each cell, identify the cell group in the final generated tree diagram, obtain the group corresponding to each cell, and obtain the correlation matrix.

3. The method as described in claim 1, characterized in that, The construction of the Rogers regression prediction model includes the following steps: Collect single-cell data from human fluids or solid tissues; Obtain gene expression profiles related to the training set; Obtain the optimal number of features; Construct Rogers regression and validate the evaluation.

4. The method as described in claim 3, characterized in that, The construction of the Rogers regression prediction model includes the following steps: S3-1: Collect single-cell data from human body fluids or solid tissues, and obtain single-cell expression matrices and / or cell type annotation information for each cell; S3-2: Data preprocessing, including standardization and / or normalization methods; S3-3: Obtain the expression profile of tumor-related genes in the training set. Tumor-related genes in different tissues are obtained from the CancerSEA database. S3-4: Randomly group the training set cells for cross-validation; perform a t-test on each gene as a basis for separability, and use the individual optimal feature combination method to obtain the optimal number of features; S3-5: Construct Rogers regression using the entire training set.

5. The method according to any one of claims 1 to 4, characterized in that, The sample to be tested includes, but is not limited to, body fluids or solid tissues.

6. The identification model obtained by the method as described in any one of claims 1 to 5.

7. The application of the identification model as described in claim 6 in any of the following; (I) Identification of tissue-derived cells; (II) Preparation of products for identifying tissue-derived cells.

8. An apparatus, characterized in that, The device has a built-in module or component corresponding to the identification model as described in claim 6.

9. A computer-readable storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the storage medium is located to perform the method as described in any one of claims 1 to 5.

10. A processor, the processor being used to run a program, wherein, The program executes the method as described in any one of claims 1 to 5 when it runs.

11. An electronic device, characterized in that, The device includes: a memory and a processor; The memory stores one or more computer programs, the one or more computer programs including instructions; when the instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Library construction method for single-cell sequencing

    CN110438572A

  • Method for identifying tissue source of intermediate mesenchymal stem cells of sample and application thereof

    CN115565608A

  • Method for detecting mesenchymal stem cells with different tissue sources, generations and culture conditions based on single cell sequencing

    CN116884484A

Cited By

  • Cell annotation method and device, electronic equipment and storage medium

    CN121415883A

  • Method for detecting virus load based on single cell transcriptome sequencing

    CN121999871A

  • Large language model driven single-cell double-score iterative annotation method

    CN122392655A