A method, system, apparatus, and medium for distinguishing malignant cells from multi-omic single-cell sequencing data

By using microporous single-cell multi-omics sequencing methods, combined with microporous chips and molecular marker beads, the copy number variations and regulatory networks of tumor cells were analyzed, which solved the technical difficulties in analyzing tumor cell heterogeneity, achieved rapid identification of tumor cells and enrichment of target genes, and improved the accuracy of clinical diagnosis and treatment.

CN117476101BActive Publication Date: 2025-10-10ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311568169.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-10-10
Estimated Expiration
2043-11-22

AI Technical Summary

Technical Problem

Existing technologies lack an autonomous, low-cost, and high-throughput platform for multidimensional analysis of tumor cells, and are unable to effectively reflect the heterogeneity of cells within tumors, affecting clinical treatment outcomes and drug resistance analysis.

Method used

A microwell single-cell multi-omics sequencing method was used. By loading molecularly labeled microbeads into a microwell chip and mixing them with the cell nucleus, single-cell transcriptome, chromatin accessibility, and genome sequencing were performed. Combined with cell identity tags, copy number variation levels were analyzed, malignant cells were screened and merged, and the chromosome variation pattern map and regulatory network of malignant cells were constructed.

Benefits of technology

It achieves rapid and accurate identification of tumor cells, determines the proportion of malignant cells and copy number variation patterns, enriches target gene characteristics, provides auxiliary diagnosis and treatment reference, and improves the resolution and accuracy of tumor heterogeneity analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117476101B_ABST
    Figure CN117476101B_ABST
Patent Text Reader

Abstract

The application discloses a method, system, device and medium for distinguishing malignant cells from multi-omics single-cell sequencing data, and belongs to the technical field of tumor single-cell sequencing. The method comprises the following steps: performing high-throughput single-cell multi-omics sequencing by using molecular marker microbeads; and further performing copy number variation analysis based on multi-omics single-cell sequencing data, so as to distinguish malignant cells in tumor and tumor-adjacent tissues. By using the application, the genomic sequence characteristics and gene expression patterns of malignant cells in the tumor can be accurately distinguished at the multi-omics level, and the application has great application value in the detection and auxiliary diagnosis of clinical tumor samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tumor single-cell sequencing, and specifically relates to a method, system, device and medium for distinguishing malignant cells from multi-omics single-cell sequencing data. Background Art

[0002] Cancer is the disease with the highest morbidity and mortality worldwide. The occurrence of tumors is, to a certain extent, due to the accumulation of mutations in initial malignant cells that acquire stemness. After changes in the endogenous tumor microenvironment and the induction of exogenous conditions, the proliferation and differentiation of malignant cells produce cell types with different phenotypes and morphologies, which shape the heterogeneity of the tumor. The occurrence and development of tumors in various organs and tissues are all derived from the heterogeneity within the tumor. The evolution process of various tumors also has common characteristics. The evolution process of tumor clones with different mutations leads to one or more clone types with survival advantages, which determine the molecular characteristics of the tumor and the formation of the microenvironment. This process is dynamic and complex. Internal tumor heterogeneity is a key factor in the generation of chemotherapy, targeted drug therapy and immunotherapy resistance, as well as recurrence and mortality during clinical treatment.

[0003] With the advancement of high-throughput second-generation sequencing technology in recent years, deep genomic sequencing studies of different types of tumors have revealed genomic instability, and multiple somatic mutations are closely related to the formation of tumor heterogeneity and the evolution of survival. Multidimensional analysis of tumors at cellular resolution can help further clarify the formation of internal tumor heterogeneity and the history of clonal evolution, explore the common and differential mechanisms of tumor development, and help solve important problems such as clinical tumor recurrence and drug resistance. However, multidimensional cellular analysis and comparative studies of issues such as the development and internal heterogeneity of different types of tumors are still relatively rare, and there is a lack of autonomous, low-cost, and relatively high-throughput platforms at the technical level.

[0004] In the past, molecular characterization of various tumor tissues was typically performed through genome sequencing at the population cell level, gene expression analysis (transcriptome sequencing, gene expression chips, or fluorescence quantitative analysis), and protein localization and expression at the tissue level. Limited by the resolution of the technical means, gene expression detection at the population level cannot reflect the heterogeneity of internal cells. Single-cell sequencing technology can accurately detect differential gene expression or genomic changes in cells, providing new opportunities for analyzing internal heterogeneity and the evolutionary development trajectory of tumors. In the field of tumor research, single-cell sequencing can provide assistance from multiple omics dimensions such as the genome, transcriptome, proteome, metabolome, and epigenetic group to solve a series of problems such as primary tumor heterogeneity, tumor microenvironment, and the relationship between primary and recurrent metastatic tumor foci.

[0005] The mechanisms of tumor cell occurrence and the heterogeneity of tumor cell evolution discovered based on single-cell omics can further provide clues for the diagnosis and prevention of tumors from the molecular characteristics of malignant cell mutations and internal heterogeneous cells, and have huge application and transformation potential in the direction of mechanism research and diagnosis and prevention. At present, single-cell omics research on tumors focuses on molecular typing of cells based on transcriptome gene expression. With the development of commercial single-cell technology platforms and high-throughput sequencers, single-cell transcriptome maps of various tumor animal models and human clinical tumor samples have been published. A variety of tumor cell maps systematically characterize the heterogeneity of intratumoral cells and immune microenvironment cells. Therefore, the development of a rapid tumor cell identification method based on micropore single-cell multi-omics sequencing is of great clinical significance. Summary of the Invention

[0006] In order to solve the above technical problems, the technical solutions provided by the present invention are as follows:

[0007] A first aspect of the present invention provides a method for distinguishing malignant cells based on single-cell multi-omics sequencing, comprising the following steps:

[0008] S1: Obtain tumor samples and adjacent tumor samples, prepare single-cell nuclei suspensions, mix them with molecularly labeled microbeads, and load them into a microwell chip. The base sequences of the labeled cell nuclei are captured in situ within the microwells and labeled with cell identity tags and molecular tags.

[0009] S2, constructing a sequencing library and performing at least two of the following sequencing methods: single-cell transcriptome sequencing, single-cell chromatin accessibility sequencing, single-cell genome sequencing, and single-cell methylation sequencing to obtain different single-cell sequencing data;

[0010] S3, for each type of single-cell sequencing data, perform the following analysis:

[0011] S31, obtain the average copy number variation levels in tumor samples and adjacent tumor samples, respectively, as the malignant copy number variation expectation and normal copy number variation expectation,

[0012] S32: Divide the single-cell sequencing data of the tumor sample and the adjacent tumor sample into N subsets. For each subset, make judgments based on the following criteria:

[0013] If the average copy number variation level of the subset is less than the normal copy number variation expectation, the subset is a normal subset and its cells are normal cells; if the average copy number variation level of the subset is greater than the malignant copy number variation expectation, the subset is a malignant subset and its cells are malignant cells; if the average copy number variation of the subset is between the normal copy number variation expectation and the malignant copy number variation expectation, the subset is an intermediate state.

[0014] S33, for the intermediate state subset, re-divide into N subsets, and classify according to the standard in S32;

[0015] S34, repeat step S33 until there are no more normal subsets or malignant subsets, or the maximum number of iterations Y is reached,

[0016] Wherein N = 20-100, Y = 10-50;

[0017] S4, correlation analysis is performed on the chromosome copy number variation patterns of the malignant cells identified in step S3 from different single cell sequencing data, and the malignant cells are merged using the chromosome regions with the same copy number variation patterns.

[0018] In some embodiments of the present application, in step S1, the molecular marker microbeads are mixed with the cell nuclei at a ratio of 1:1 after proportioning, and then loaded onto the microwell chip, which can give the cell nuclei a cell identity tag, facilitating the rapid determination of cell nuclei from different cell types in the subsequent analysis process. Preferably, for transcriptome sequencing and cytoplasm accessibility sequencing, the cell nuclei are given a cell identity tag while reverse transcription / genome fragmentation is performed.

[0019] Further, in step S1, further comprising: using any one of an aldehyde fixing solution (such as paraformaldehyde), an alcohol fixing solution (such as ethanol), an acid fixing solution, and a crosslinking agent to fix the cell nucleus suspension, so that the nucleic acids / proteins in the cell nucleus are crosslinked and fixed, and the nucleic acid molecules are more effectively reacted in the cell / nucleus. Preferably, in genome sequencing, the cell nuclei are not subjected to any organic solvent fixation treatment, so that the transposase can more effectively enter the cell nucleus for reaction.

[0020] In some embodiments of the present application, in step S1, the objects of the in situ nucleic acid molecule labeling reaction in the cell nuclei are mRNA and DNA. The poly-T tail carried by the nucleic acid molecules with known base sequences on the surface of the microbeads can be hybridized and combined with the mRNA in the treated cell nuclei; the random or fixed sequence carried by the nucleic acid molecules with known base sequences on the surface of the microbeads can be hybridized and combined with the DNA in the treated cell nuclei.

[0021] In some embodiments of the present application, in step S3, before analysis, further comprising the step of performing pseudo-population processing:

[0022] According to the number of cells derived from the same sample, the single cell sequencing data is added to perform pseudo-population processing, the single cell sequencing data sets of adjacent cells are added according to the Euclidean distance to construct a pseudo-population set, and data normalization processing is performed.

[0023] In some embodiments of the present invention, in step S4, the specific step of merging malignant cells by merging chromosomal regions with the same copy number variation pattern is: screening chromosomal regions with copy number variation directions of "amplification" or "deletion", drawing a chromosomal variation pattern diagram of malignant cells, and thus merging the malignant cells.

[0024] In some embodiments of the present invention, the method further comprises the step of identifying cell subtypes based on any single-cell sequencing data:

[0025] According to the cell identity tags in the sequencing data, all microbeads are grouped into pairs to form microbead pairs;

[0026] A traversal calculation is performed on each bead pair, the calculation content is the similarity of the bead capture sequence, and the bead pairs are sorted according to the similarity;

[0027] Next, based on the actual number of wells contained in the microwells, bead pairs with sequence similarity above a preset threshold are merged;

[0028] Finally, the gene matrices of cells from tumor samples and adjacent tumor samples were merged separately, and the merged single-cell omics matrix was subjected to dimensionality reduction, feature selection, differential analysis, and cell subpopulation clustering, and the cell subpopulations were annotated based on public databases.

[0029] The above process calculates the similarity expression score of the random sequence distribution similarity in the cell identity tags carried and captured by the microbeads in the data, thereby determining which microwells have multiple microbeads located in the same microwell, and merging the genetic sequence information of all microbeads in the same microwell. For multiple cell nuclei in the same microwell, the genetic sequence information merged by the microbeads is assigned and restored to a single cell nucleus through the cell identity tag, so that multi-omics data with single-cell resolution can be obtained.

[0030] For transcriptome sequencing, the primer sequence attached to the beads consists of four parts: library adapter sequence, cell tag sequence, molecular tag sequence, and nucleic acid capture sequence. The library adapter sequence is used for subsequent sequencing; the cell tag sequence is used to identify different cells; the molecular tag sequence is a sequence composed of random bases. Each DNA molecule contains a unique molecular tag sequence, which is used to identify different DNA molecules during mixed sequencing; the nucleic acid capture sequence contains a poly T tail or random primer sequence for capturing RNA molecules.

[0031] For genomic sequencing, transposase is used during library construction to fragment open regions of genomic chromatin. The primer sequence for microbead attachment consists of four parts: a library adapter sequence, a cell tag sequence, a molecular tag sequence, and a nucleic acid capture sequence. The library adapter sequence is used for subsequent sequencing; the cell tag sequence is used to identify different cells; the molecular tag sequence is a sequence composed of random bases. Each DNA molecule contains a unique molecular tag sequence, which is used to identify different DNA molecules during mixed sequencing. The nucleic acid capture sequence contains a hybridization sequence that matches the transposase adapter sequence and is used to capture DNA molecules fragmented by the transposase.

[0032] In some embodiments of the present invention, the merging of bead pairs having sequence similarity higher than a preset threshold is specifically performed as follows:

[0033] (1) One cell and one microbead in a microwell directly use the cell identity tag and captured genetic sequence information of the microbead as the genetic information matrix of the single cell;

[0034] (2) There are multiple cells and a microbead in the microwell, and the genetic sequence information captured by the microbead is assigned to the multiple cells according to the cell identity tags of the multiple cells it is paired with, as the genetic information matrix of the multiple cells;

[0035] (3) There is a cell and multiple microbeads in the microwell, and the genetic sequences captured by the microbeads are accumulated and assigned to this cell as the genetic information matrix of the single cell;

[0036] (4) There are multiple cells and multiple microbeads in the microwell. The genetic sequences captured by the microbeads are first accumulated and then assigned to the multiple cells according to the cell identity tags of the multiple cells they are paired with, as the genetic information matrix of the multiple cells.

[0037] In some embodiments of the present invention, the method further comprises predicting key transcription factors and / or their target genes in the identified malignant cells to perform molecular typing of the malignant cells.

[0038] The present invention can cross-validate and further assist in confirming malignant cells from suspected cancer (malignant tumor) samples, and further quickly clarify the cell lineage origin, proportion, copy number variation pattern, target gene and other molecular typing indicators of malignant cells from the cytogenomic dimension, thereby providing auxiliary diagnosis.

[0039] A second aspect of the present invention provides a system for distinguishing malignant cells based on single-cell multi-omics sequencing, comprising the following modules:

[0040] Data input module: for receiving different single-cell sequencing data obtained from at least two of single-cell transcriptome sequencing, single-cell chromatin accessibility sequencing, single-cell genome sequencing and single-cell methylation sequencing of tumor sample and tumor-adjacent sample;

[0041] Malignant cell distinguishing module: connected with the data input module, for respectively performing the following analysis for each kind of single-cell sequencing data:

[0042] S31, respectively obtaining the average copy number variation level in the tumor sample and the tumor-adjacent sample as the malignant copy number variation expectation and the normal copy number variation expectation,

[0043] S32, dividing the single-cell sequencing data of the tumor sample and the tumor-adjacent sample into N subsets respectively, and for each subset, judging according to the following criteria:

[0044] If the average copy number variation level of the subset is less than the normal copy number variation expectation, the subset is a normal subset and the cells thereof are normal cells; if the average copy number variation level of the subset is greater than the malignant copy number variation expectation, the subset is a malignant subset and the cells thereof are malignant cells; if the average copy number variation level of the subset is between the normal copy number variation expectation and the malignant copy number variation expectation, the subset is an intermediate state,

[0045] S33, for the intermediate state subset, re-dividing into N subsets and classifying according to the criteria in S32;

[0046] S34, repeating step S33 until there are no more normal subsets or malignant subsets, or reaching the maximum iteration

[0047] times Y, wherein N=20-100 and Y=10-50;

[0048] S4, performing correlation analysis on the chromosome copy number variation patterns of the malignant cells identified in step S3 from different single-cell sequencing data, and merging the malignant cells with the same chromosome copy number variation patterns.

[0049] Further, it further comprises:

[0050] Cell subtype identification module, connected with the data input module and the malignant cell distinguishing module respectively, for identifying cell subtypes according to the following steps:

[0051] According to the cell identity tag in the sequencing data, all the microbeads are paired to form microbead pairs;

[0052] For each microbead pair, traversal calculation is performed, the calculation content being the similarity of microbead capture sequences, and the microbead pairs are sorted according to the similarity;

[0053] Next, based on the actual number of wells contained in the microwells, bead pairs with sequence similarity above a preset threshold are merged;

[0054] Finally, the gene matrices of cells from tumor samples and adjacent tumor samples were merged, and the merged single-cell omics matrix was subjected to dimensionality reduction, feature selection, differential analysis, and cell subpopulation clustering. The cell subpopulations were annotated based on public databases.

[0055] It is also used to determine the variation patterns of malignant cells in different cell subpopulations based on the malignant cells identified by the malignant cell differentiation module.

[0056] Furthermore, it also includes a key target gene and its regulatory network enrichment module, which is connected to the malignant cell differentiation module to predict the key transcription factors and / or their target genes in the identified malignant cells and perform molecular typing of malignant cells.

[0057] A third aspect of the present invention provides a computer device, comprising:

[0058] memory for storing computer programs;

[0059] A processor is configured to implement the steps of a method for distinguishing malignant cells based on single-cell multi-omics sequencing as described in any one of the first aspects of the present invention when executing the computer program.

[0060] A fourth aspect of the present invention provides a computer-readable storage medium,

[0061] The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for distinguishing malignant cells based on single-cell multi-omics sequencing as described in any one of the first aspects of the present invention.

[0062] Beneficial effects of the present invention

[0063] Compared with the prior art, the present invention has the following beneficial effects:

[0064] The present invention provides a method, device, and medium for rapid tumor cell identification based on microwell single-cell multi-omic sequencing. This method, based on a microwell microbead system, enables high-throughput detection of multi-omic genetic information from tumor samples at the single-cell level. Based on this multi-omic information, malignant cells in tumors can be rapidly and accurately identified and their characteristic regulatory patterns of target genes enriched, providing a reference for clinical tumor classification and auxiliary diagnosis.

[0065] The present application performs correlation analysis on each chromosome copy number variation pattern obtained from different single cell sequencing data, screens chromosome regions with "amplification" or "deletion" copy number variation direction, and draws a malignant cell chromosome variation pattern map. The malignant cells obtained by iterative grouping according to the average copy number level of each cell are combined. The core malignant cell subpopulation and its distribution ratio in each cell lineage are determined. The genomic variation pattern is further determined, and the key target genes of the malignant cells and their regulatory network are enriched for molecular typing of the malignant cells. The present application can integrate multi-omics data of tumor malignant cells to construct a regulatory network, while the regulatory network construction in the prior art is basically based on single cell transcriptome gene expression data. Integrating tumor single cell transcriptome and single cell chromatin accessibility and other genomic data to construct a single cell resolution regulatory network is the first of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0066] Figure 1 The cell subtypes defined by genomic chromatin accessibility of mouse lung tumor samples, tumor adjacent samples, and normal reference control lung samples and the sample sources of the cell subtypes are shown. Adj represents the tumor adjacent sample, Normal represents the contralateral normal lung sample, and Tumor represents the tumor sample.

[0067] Figure 2 The cell subtypes defined by genomic chromatin accessibility of mouse lung tumor samples, tumor adjacent samples, and contralateral normal lung samples, and the distribution projection of malignant cells (malignant) and non-malignant normal cells (nonmalignant) defined by the degree of copy number variation are shown.

[0068] Figure 3 The degree of copy number variation grouped by the degree of malignancy at the genomic chromatin accessibility group level identified by Copy-scAT is shown. NMF_cluster represents Non-negative matrix factorization, an unsupervised clustering method, and the 3rd group obtained by this method is a malignant cell subpopulation, which matches the distribution defined by copy number variation.

[0069] Figure 4 The predicted chromosome range copy number variation pattern of malignant cells (malignant) and non-malignant cells (nonmalignant) at the transcriptome level identified by inferCNV is shown.

[0070] Figure 5The figure shows the correlation between the average copy number variation scores at the transcriptome level identified by inferCNV and the average copy number variation scores at the genomic chromatin accessibility level identified by Copy-scAT on different chromosomal bands, and on chromosomal bands consistent with copy number deletion (del effect) and copy number amplification (dup effect).

[0071] Figure 6 This figure shows the enrichment results and regulatory network of key target genes in mouse lung tumor malignant cells. Pink genes are selected key target genes. The color of the node represents its network centrality, and the size of the node represents the number of interacting genes. DETAILED DESCRIPTION

[0072] Unless otherwise indicated, implied from the context, or customary in the art, all parts and percentages in this application are based on weight, and the test and characterization methods used are current as of the filing date of this application. Where applicable, the contents of any patents, patent applications, or publications referred to in this application are incorporated herein by reference in their entirety, and their equivalent patent families are also incorporated by reference, particularly for definitions of relevant terms in the art disclosed in such documents. If the definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition of the term provided in this application shall prevail.

[0073] Numerical ranges in this application are approximate values, so unless otherwise stated, they may include numerical values ​​outside the scope. Numerical ranges include all numerical values ​​from the lower limit to the upper limit increased by 1 unit, provided that there is an interval of at least 2 units between any lower value and any higher value. For a range comprising a numerical value less than 1 or comprising a fraction greater than 1 (e.g., 1.1, 1.5, etc.), 1 unit is appropriately considered to be 0.0001, 0.001, 0.01 or 0.1. For a range comprising a single digit less than 10 (e.g., 1 to 5), 1 unit is typically considered to be 0.1. These are merely specific examples of what is intended to be expressed, and all possible combinations of the numerical values ​​between the minimum and maximum values ​​listed are considered to be clearly recorded in this application.

[0074] The terms "comprising", "including", "having" and their derivatives do not exclude the presence of any other components, steps or processes and are irrelevant to whether these other components, steps or processes are disclosed in this application. To eliminate any doubt, all compositions using the terms "comprising", "including", or "having" in this application may include any additional additives, excipients or compounds unless expressly stated otherwise. In contrast, the term "essentially consisting of" excludes any other components, steps or processes from the scope of any description of the term below, except those necessary for operational performance. The term "consisting of" does not include any components, steps or processes that are not specifically described or listed. Unless expressly stated otherwise, the term "or" refers to the listed members alone or in any combination thereof.

[0075] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention more clearly understood, the present invention is further described in detail below with reference to the embodiments.

[0076] The following examples are provided to illustrate preferred embodiments of the present invention. Those skilled in the art will appreciate that the techniques disclosed in the following examples represent techniques discovered by the inventors that can be used to practice the present invention and, therefore, can be considered preferred embodiments of the present invention. However, those skilled in the art will appreciate from this disclosure that many modifications may be made to the specific embodiments disclosed herein while still achieving the same or similar results without departing from the spirit or scope of the present invention.

[0077] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs, and the disclosure and materials they cite are hereby incorporated by reference.

[0078] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many technical equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the claims.

[0079] The experimental methods in the following examples, unless otherwise specified, are all conventional methods. The instruments and equipment used in the following examples, unless otherwise specified, are all conventional laboratory instruments and equipment; the experimental materials used in the following examples, unless otherwise specified, are all purchased from conventional biochemical reagent stores.

[0080] Example 1 Preparation and sequencing of single-cell multi-omics libraries of lung tumor samples from aged mice based on the micropore method

[0081] 1. Sample Preparation

[0082] Tumor tissue and adjacent control tissue were isolated from aged C57BL6 mice with identified lung tumors. Both tissues were rapidly frozen in liquid nitrogen, ground into a powder, resuspended in nuclear lysis buffer, and lysed on ice. Single-nuclei suspensions were obtained after centrifugation and washing.

[0083] 2. Transcriptomics Library Construction

[0084] Cell nuclei were fixed with 4% paraformaldehyde (PFA), and single-nuclei suspensions from different tumor and adjacent tissue samples were added to separate centrifuge tubes. Reverse transcription primers carrying different cell identity tag sequences, reverse transcriptase, reverse transcription buffer, dNTPs, RNase inhibitor, 50% PEG 8000, and 10% Triton X10 were added to each tube. After mixing thoroughly, the tube was placed in a PCR instrument for a constant-temperature reverse transcription reaction. After the reverse transcription reaction, the cell nuclei were washed with 3× SSC and PBS, respectively, in preparation for chip loading.

[0085] During chip loading, cell nuclei and molecularly labeled microbeads are mixed in equal proportions, and an amplification system consisting of a constant-temperature polymerase and a high-fidelity polymerase is added. Based on the sample size, reverse-transcribed nuclei from different tumor samples are quickly and evenly loaded into the microwell chip using a pipette. The microbeads and cells are inspected under a microscope to ensure a nucleus and microbead penetration rate greater than 70%. Sealing oil is then added to seal the microwell chip, creating a separate reaction chamber, and amplification is performed in a PCR thermal cycler.

[0086] After the reaction is complete, the liquid and molecular marker beads on the chip are fully collected through multiple centrifugation. The reaction liquid is then transferred to a new centrifuge tube. DNA purification magnetic beads are then added to purify the amplified cDNA liquid. This is then added to a sequencing adapter (P5 and P7) containing a sequencing tag (index) and an amplification system using a high-fidelity polymerase to amplify the sequencing library and obtain an indexed sequencing library. The sequencing library is then purified using DNAClean Beads magnetic beads, and the library concentration is determined using a Qubit 3.0 fluorescent reagent. The library is then stored at -20°C. An appropriate amount of sequencing library is selected for sequencing according to the sequencer's requirements.

[0087] 3. Epigenomics-Chromatin Accessibility Library Construction

[0088] Single-nuclei suspensions from different tumor and paratumor tissue samples were added to separate centrifuge tubes. Each tube was then supplemented with a transposase containing a different cell identity tag sequence, a 2× enzyme digestion solution, 1% digitonin, 10% Tween-20, and 1× PBS. The tubes were thoroughly mixed and incubated at 37°C to 55°C for half an hour. The digestion reaction was terminated on ice, and the nuclei were collected and centrifuged. The nuclei were then washed twice with PBS and prepared for chip loading.

[0089] When loading the chip, cell nuclei and molecularly labeled microbeads are mixed in equal proportions, and 50mM EDTA and 2x high-fidelity polymerase are added and mixed thoroughly. Depending on the sample size, reverse-transcribed nuclei from different tumor samples are quickly and evenly loaded into the microwell chip using a pipette. Microbead and cell retention in the microwells is examined under a microscope, ensuring a retention rate of greater than 70% for both nuclei and microbeads. The tube cap is then securely fastened, and the centrifuge tube is placed in a 50°C constant temperature reaction for half an hour to release genomic fragments. An amplification system containing a constant temperature polymerase and a high-fidelity polymerase is then added to the microwell chip. Sealing oil is then added to seal the microwell chip to form a separate reaction chamber, and amplification is performed in a PCR thermal cycler.

[0090] After the reaction is complete, the liquid and molecular marker beads in the chip are fully collected through multiple centrifugation, and the reaction liquid is transferred to a new centrifuge tube. DNA purification magnetic beads are then added for purification to obtain the amplified cDNA liquid. This is then added to an amplification system containing sequencing tags P5 and P7 and a high-fidelity polymerase to amplify the sequencing library and obtain an indexed sequencing library. The sequencing library is then purified using DNA Clean Beads magnetic beads to obtain the sequencing library. The library concentration is determined using Qubit 3.0 fluorescent reagent and stored at -20°C. An appropriate amount of sequencing library is selected according to the sequencer's requirements for sequencing.

[0091] Example 2 Identification of cell subtypes in lung tumor samples from aged mice based on single-cell multiomics

[0092] The raw fastq data of the high-throughput sequencing obtained in Example 1 were extracted and screened based on the cell identity tag sequences carried by reverse transcription / transposase digestion and the molecular tag sequences carried on the microbeads, and the sequencing data were aligned to the mouse reference genome to obtain the single-cell transcriptome matrix and chromatin accessibility matrix.

[0093] First, all beads are grouped into pairs according to the cell identity tags extracted by position in the sequencing data to form bead pairs.

[0094] A traversal calculation is then performed for each bead pair, calculating the similarity of the bead-captured sequences (Jaccard expression score), and the bead pairs are ranked according to similarity. In transcriptome sequencing, the capture sequence included in the calculation is the random primer sequence; in chromatin accessibility sequencing, the capture sequence included in the calculation is the captured genetic information.

[0095] Next, based on the actual number of wells in the microwells (10,000), the bead pairs with high sequence similarity (the Jaccard expression scores corresponding to the 10,000 bead pairing positions are used as the threshold for judging high sequence similarity) are merged to determine which beads are located in the same microwell and perform the following classification processing:

[0096] (1) For the case where there is one cell and one microbead in the microwell after calculation, the cell identity tag of the microbead and the captured genetic sequence information are directly used as the genetic information matrix of the single cell.

[0097] (2) For the case where there are multiple cells and one microbead in the microwell after calculation, the genetic sequence information captured by the microbead is assigned to the multiple cells according to the cell identity tags of the reverse transcription / enzyme cleavage steps of the multiple cells it is paired with, as the genetic information matrix of the multiple cells.

[0098] (3) For the case where there is one cell and multiple microbeads in the microwell after calculation, the genetic sequences captured by the microbeads are accumulated and assigned to this cell as the genetic information matrix of the single cell.

[0099] (4) In the case where there are multiple cells and multiple microbeads in the microwell after calculation, the captured genetic sequences of the microbeads are first accumulated and then assigned to the multiple cells according to the cell identity tags of the reverse transcription / enzyme cleavage steps of the multiple cells they are paired with, as the genetic information matrix of the multiple cells.

[0100] Finally, the tumor cells and adjacent tissue cells were merged, and Seurat and ArchR software were used to generate and process the downstream gene expression matrix for the transcriptome data and chromatin accessibility data, respectively. The top 2,000 characteristic differential genes were selected, and the single-cell gene expression matrix was subjected to PCA analysis and dimensionality reduction, and the characteristic single-cell subpopulations were projected on a two-dimensional plane. For each subpopulation, differential enrichment tools such as Findmarker were used to obtain the most characteristic gene set for the subpopulation. Referring to existing gene annotation databases, such as the PanglaoDB database, the cell subpopulation typing was annotated and defined, and it was defined as a specific cell type of different lineages. Different epithelial cells, stromal cells, and immune cell types in the tumor tissue were identified, and subtype classification was performed based on genomic chromatin accessibility, such as Figure 1As shown, the sample origin of each cell is also integrated into the label.

[0101] Example 3 Identification of malignant cells in mouse lung tumor samples

[0102] After using single-cell genomic chromatin accessibility and transcriptome data to identify the cell types of tumor and adjacent tumor samples, a certain merging ratio is selected based on the number of cells sequenced from the tumor and adjacent tumor samples (in this example, 100 cells within the subpopulation are merged into a pseudo-population set of cells). The 100 neighboring cell data sets are added together based on the Euclidean distance to construct a pseudo-population set. Subsequently, the gene expression count matrix of the merged cells (columns are for each cell, behavioral genes) is added together in the Seurat software, that is, the sum of the counts of each sequenced gene (each row) in 100 cells (100 columns) is calculated, and used as the gene expression matrix of the pseudo-population set after addition. The single-cell gene expression matrix after pseudo-population processing is reduced in dimension and clustered, and data normalization is performed to normalize the pseudo-population set of single-cell data to 10 6 Order of magnitude.

[0103] The inferCNV software and Copy-scAT software were combined to perform copy number variation analysis at the transcriptome and genomic levels (pseudo-population level) of the lung tumor sample dataset, quantitatively characterizing the copy number loss (del effect) and copy number gain (dup effect) patterns within different chromosome ranges.

[0104] First, assuming that both tumors and adjacent tissues contain malignant and normal cells, the copy number region in adjacent tissues was initially used as the normal control. The average copy number variation (average copy number variation level) of tumor and adjacent tissue cells was calculated at the pseudo-population level. The average copy number variation levels of adjacent tissues and tumor tissues were used as the "normal copy number variation expectation" and "malignant copy number variation expectation," respectively. For transcriptome data, gene expression levels were quantified to a range of -1 to 1.

[0105] The hierarchical clustering algorithm was used to divide the pseudo-population set constructed from the single-cell data of lung tumor adjacent tissue and lung tumor tissue into 50 subsets, and the following were defined:

[0106] If the average copy number variation level of each subset is less than the “normal copy number variation expectation”, the subset is defined as “normal”;

[0107] If the average copy number variation level of each subset is greater than the “malignant copy number variation expectation”, the subset is defined as “malignant”;

[0108] The average copy number variation level of each subset was between the two, and the subset was defined as "intermediate".

[0109] Subsets classified as "intermediate" will enter the next round of hierarchical clustering iterations, where they will be re-divided into 50 subsets and the classification will be calculated. This continues until there are no more "normal" or "malignant" subsets or the maximum number of iterations is reached. The final copy number variation signature is then projected onto the single-cell clustering results, and the malignant cells identified by multi-omics are merged to observe whether there are subpopulations of malignant cells clustering alone or a scattered distribution pattern of malignant cells.

[0110] The results are as follows Figure 2 As shown in Figure 2, it can be seen that most of the predicted malignant cells come from tumor tissue samples, and a certain proportion of malignant transitional cells are also distributed in the adjacent tissues of the tumor. The regional and global copy number variation patterns of each chromosome level are shown in Figure 2. Figure 3 、 Figure 4 As shown, it can be seen that in the malignant cell population of the identified lung tumor sample, chromosomes 8, 16, and 17 showed significant copy number amplification; the malignant cells in the tumor and adjacent tissues showed copy number loss in chromosomes 4, 5, and 11.

[0111] To integrate and correlate the copy number variation patterns of different omics, it is necessary to divide the annotated chromosome regions into different bands, quantify the average copy number variation scores at the overlapping band level of different omics (quantified from deletion to amplification to a range of -2 to 2) and project them, and compare the copy number "amplification" and "deletion" of multi-omics copy number variations in different bands. For the inferCNV transcriptome level data and the Copy-scAT genomic chromatin accessibility level data, a total of 42 overlapping chromosome bands were obtained, and the copy number amplification and copy number loss of each chromosome band were marked, as shown in the figure. Figure 5 As shown in the figure, the correlation performance of the copy number variation results of the multi-omics joint analysis reached 0.73, which was significant (p value was 0.0018).

[0112] Example 4 Construction of a gene expression regulatory network for malignant cells

[0113] SCRIP software was used to predict key transcription factors and their target genes in lung tumor malignant cells identified by multi-omics and to construct a regulatory network.

[0114] First, the ClusterProfiler tool was used to select a significant p-value threshold of 0.1 to filter low-quality target genes and enrich potential key pathways, obtaining 21 common enriched pathways. At the same time, 49 target genes with high correlation with these pathways were screened and named key node target genes.

[0115] The target gene interaction information of the STRING database was used to depict the interaction network of these target genes, and proteins outside the network were further removed. The key node target genes of normal and malignant cells were enriched according to the threshold of average expression fold (avg.LogFC)>0.25 and significant BH adjusted p value<0.05, such as Figure 6 shown.

[0116] This framework enriched key target genes known to be highly correlated with lung tumors, such as Tp63, Foxc2, and Nkx2-1, indicating that this malignant cell population exhibits characteristics of epithelial-mesenchymal transition. Furthermore, the tumor samples detected so far are squamous cell lung carcinomas, suggesting a transition from epithelial lung adenocarcinoma to stromal lung squamous cell carcinoma. This demonstrates that the method of the present invention can rapidly and accurately identify key regulatory target genes and their regulatory networks in malignant cells in tumors.

[0117] All documents mentioned in this application are incorporated herein by reference, just as if each document were incorporated herein by reference individually. It should also be understood that after reading the above teachings of the present invention, those skilled in the art may make various changes or modifications to the present invention, and that such equivalents also fall within the scope of the claims appended hereto.

Claims

1. A method for distinguishing malignant cells based on single-cell multi-omics sequencing, characterized in that: The following steps are involved: S1: Obtain tumor samples and adjacent tumor samples, prepare single-cell nuclei suspensions, mix them with molecularly labeled microbeads, and load them into a microwell chip. The base sequences of the labeled cell nuclei are captured in situ within the microwells and labeled with cell identity tags and molecular tags. S2, constructing a sequencing library and performing at least two of the following sequencing methods: single-cell transcriptome sequencing, single-cell chromatin accessibility sequencing, single-cell genome sequencing, and single-cell methylation sequencing to obtain different single-cell sequencing data; S3, for each type of single-cell sequencing data, perform the following analysis: S31, obtain the average copy number variation levels in tumor samples and adjacent tumor samples, respectively, as the malignant copy number variation expectation and normal copy number variation expectation, S32: Divide the single-cell sequencing data of the tumor sample and the adjacent tumor sample into N subsets. For each subset, make judgments based on the following criteria: If the average copy number variation level of the subset is less than the normal copy number variation expectation, the subset is a normal subset and its cells are normal cells; if the average copy number variation level of the subset is greater than the malignant copy number variation expectation, the subset is a malignant subset and its cells are malignant cells; if the average copy number variation of the subset is between the normal copy number variation expectation and the malignant copy number variation expectation, the subset is an intermediate state. S33, for the intermediate state subset, re-divide it into N subsets and classify them according to the criteria in S32; S34, repeat step S33 until there are no more normal subsets or malignant subsets, or the maximum number of iterations Y is reached, Where N=20~100, Y=10~50; S4, performing correlation analysis on the chromosomal copy number variation patterns of the malignant cells identified by the different single-cell sequencing data in step S3, merging the malignant cells with the same chromosomal regions with the same copy number variation pattern, screening the chromosomal regions with the copy number variation direction of "amplification" or "deletion" to draw a chromosomal variation pattern diagram for the malignant cells, and merging the malignant cells obtained by iterative grouping according to the average copy number level of each cell.

2. The method for distinguishing malignant cells based on single-cell multi-omics sequencing according to claim 1, characterized in that: The method further comprises the step of identifying cell subtypes based on any single-cell sequencing data: According to the cell identity tags in the sequencing data, all microbeads are grouped into pairs to form microbead pairs; A traversal calculation is performed on each bead pair, the calculation content is the similarity of the bead capture sequence, and the bead pairs are sorted according to the similarity; Next, based on the actual number of wells contained in the microwells, bead pairs with sequence similarity above a preset threshold are merged; Finally, the gene matrices of cells from tumor samples and adjacent tumor samples were merged separately, and the merged single-cell omics matrix was subjected to dimensionality reduction, feature selection, differential analysis, and cell subpopulation clustering, and the cell subpopulations were annotated based on public databases.

3. The method for distinguishing malignant cells based on single-cell multi-omics sequencing according to claim 2, characterized in that: The merging of the microbead pairs whose sequence similarity is higher than a preset threshold is specifically as follows: (1) One cell and one microbead in a microwell directly use the cell identity tag and captured genetic sequence information of the microbead as the genetic information matrix of the single cell; (2) There are multiple cells and a microbead in the microwell, and the genetic sequence information captured by the microbead is assigned to the multiple cells according to the cell identity tags of the multiple cells it is paired with, as the genetic information matrix of the multiple cells; (3) There is a cell and multiple microbeads in the microwell. The genetic sequences captured by the microbeads are accumulated and assigned to this cell as the genetic information matrix of the single cell. (4) There are multiple cells and multiple microbeads in the microwell. The genetic sequences captured by the microbeads are first accumulated and then assigned to the multiple cells according to the cell identity tags of the multiple cells they are paired with, as the genetic information matrix of the multiple cells.

4. The method for distinguishing malignant cells based on single-cell multi-omics sequencing according to claim 1, characterized in that: In step S3, before the analysis, a pseudo-grouping process is further included: The single-cell sequencing data were combined based on the number of cells from the same sample to perform pseudo-grouping processing. The single-cell sequencing data sets of neighboring cells were combined based on the Euclidean distance to construct a pseudo-group set, and the data were normalized.

5. The method for distinguishing malignant cells based on single-cell multi-omics sequencing according to claim 1, characterized in that: The method further includes predicting key transcription factors and / or their target genes in the identified malignant cells and performing molecular typing of the malignant cells.

6. A system for distinguishing malignant cells based on single-cell multi-omics sequencing, characterized in that: Includes the following modules: A data input module is configured to receive different single-cell sequencing data obtained by performing at least two of single-cell transcriptome sequencing, single-cell chromatin accessibility sequencing, single-cell genome sequencing, and single-cell methylation sequencing on tumor samples and adjacent tumor samples; the single-cell sequencing data obtaining step comprises: preparing single-cell nucleus suspensions for each of the tumor sample and adjacent tumor sample, mixing the suspensions with molecularly labeled microbeads, and loading the suspensions into a microwell chip; in situ capturing the base sequences of the labeled cell nuclei within the microwells, adding cell identity tags and molecular tags, and constructing a sequencing library for sequencing; Malignant cell differentiation module: connected to the data input module, used to perform the following analysis on each type of single-cell sequencing data: S31, obtain the average copy number variation levels in tumor samples and adjacent tumor samples, respectively, as the malignant copy number variation expectation and normal copy number variation expectation, S32: Divide the single-cell sequencing data of the tumor sample and the adjacent tumor sample into N subsets. For each subset, make judgments based on the following criteria: If the average copy number variation level of the subset is less than the normal copy number variation expectation, the subset is a normal subset and its cells are normal cells; if the average copy number variation level of the subset is greater than the malignant copy number variation expectation, the subset is a malignant subset and its cells are malignant cells; if the average copy number variation of the subset is between the normal copy number variation expectation and the malignant copy number variation expectation, the subset is an intermediate state. S33, for the intermediate state subset, re-divide it into N subsets and classify them according to the criteria in S32; S34, repeat step S33 until there are no more normal subsets or malignant subsets, or the maximum number of iterations Y is reached, where N = 20~100, Y = 10~50, It is also used to perform correlation analysis on the chromosome copy number variation patterns of identified malignant cells, merge malignant cells with chromosomal regions with the same copy number variation pattern, screen chromosomal regions with copy number variation directions of "amplification" or "deletion" to draw chromosome variation pattern maps of malignant cells, and merge malignant cells obtained by iterative grouping according to the average copy number level of each cell.

7. The system for distinguishing malignant cells based on single-cell multi-omics sequencing according to claim 6, characterized in that: Also includes: The cell subtype identification module is connected to the data input module and the malignant cell differentiation module respectively, and is used to perform cell subtype identification according to the following steps: According to the cell identity tags in the sequencing data, all microbeads are grouped into pairs to form microbead pairs; A traversal calculation is performed on each bead pair, the calculation content is the similarity of the bead capture sequence, and the bead pairs are sorted according to the similarity; Next, based on the actual number of wells contained in the microwells, bead pairs with sequence similarity above a preset threshold are merged; Finally, the gene matrices of cells from tumor samples and adjacent tumor samples were merged, and the merged single-cell omics matrix was subjected to dimensionality reduction, feature selection, differential analysis, and cell subpopulation clustering. The cell subpopulations were annotated based on public databases. It is also used to determine the variation patterns of malignant cells in different cell subpopulations based on the malignant cells identified by the malignant cell differentiation module.

8. The system for distinguishing malignant cells based on single-cell multi-omics sequencing according to claim 6, characterized in that: Also includes: The key target gene and its regulatory network enrichment module is connected to the malignant cell differentiation module to predict the key transcription factors and / or their target genes in the identified malignant cells and perform molecular typing of the malignant cells.

9. A computer device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of a method for distinguishing malignant cells based on single-cell multi-omics sequencing as described in any one of claims 1 to 6 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for distinguishing malignant cells based on single-cell multi-omics sequencing as described in any one of claims 1 to 6.