Method, electronic equipment and system for detecting rare cell mutation based on single cell sequencing
By constructing single-cell conventional and target region enriched nucleic acid sequencing libraries and combining single-cell expression profile information, the accuracy and sensitivity problems of rare cell mutation detection in high-throughput single-cell sequencing were solved, and efficient detection and analysis of rare cell mutations were achieved.
Patent Information
- Application Number
- CN202410261964.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-07
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies make it difficult to accurately detect single-base mutations in rare cells, especially in high-throughput single-cell sequencing. Conventional methods have the risk of low detection sensitivity and missing rare cells.
By constructing single-cell conventional sequencing libraries and target region enriched nucleic acid sequencing libraries, combined with single-cell expression profile information, mutation sites of rare cells are screened out, and detection is achieved using single-cell sequencing data from multiple platforms, including DNA-seq, RNA-seq, and ATAC-seq, to achieve mutation detection in rare cells.
It improves the accuracy and sensitivity of rare cell mutation detection, can automatically obtain rare cell mutation results in batches, reduces the cost of subsequent scientific and clinical applications, and is widely applicable to a variety of single-cell sequencing platforms.
Smart Images

Figure CN120613007A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of molecular biology technology, and in particular relates to a method for detecting rare cell mutations based on single-cell sequencing, as well as electronic equipment and a system. Background Art
[0002] Rare cell mutation signatures play an important role in early cancer diagnosis, personalized treatment selection, drug resistance monitoring, prognosis assessment, and scientific research. However, there is currently no effective solution for accurately detecting these mutation signatures and localizing them to specific rare cell types. Although single-cell sequencing technology can detect rare cell mutations to some extent, the sequencing depth of individual cells in high-throughput single-cell sequencing experiments is generally very low, making mutation detection a significant challenge.
[0003] Currently, single-base mutation detection methods for single-cell DNA or RNA sequencing data can be divided into two main categories. One category is adapted to single-tube / micropore single-cell sequencing technology, but this type of data does not contain barcode sequences to distinguish cells. Detection methods such as Monoar and SCcaller have low computational cell throughput and generate a large number of computationally redundant files. The other category of methods can be used for high-throughput single-cell sequencing data analysis, but require additional traditional large-scale sequencing (bulk) genome or transcriptome data to use a priori mutation sites as a reference, such as cb_sniffer and VarTrix.
[0004] Although there are methods that have noticed the above problems and provided corresponding solutions, such as SComatic, which uses a single cell population as the minimum detection unit (SComatic) to enhance the detection signal. However, in practical applications, this method needs to first analyze cell heterogeneity and use the same cell population as the minimum analysis unit for mutation detection. The stringent filtering standards brought about by this operation may lead to the loss of rare cells and is not suitable for the identification of rare cell mutations. Therefore, it is necessary to provide a rare cell single-base mutation detection method suitable for high-throughput single-cell sequencing data, which can retain rare mutations to the greatest extent. Summary of the Invention
[0005] Embodiments of the present invention provide a method, electronic device, and system for detecting rare cell mutations based on single-cell sequencing. This method is suitable for detecting single-base mutations in rare cells using high-throughput single-cell sequencing data and can retain rare mutations to the greatest extent possible.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] In a first aspect, the present invention provides a method for detecting rare cell mutations based on single-cell sequencing, comprising the following steps:
[0008] Acquire first single-cell sequencing data of a sample to be tested and second single-cell sequencing data of a target region in the sample to be tested;
[0009] Obtaining a map of the rare cells in the sample to be tested based on the first single-cell sequencing data;
[0010] Obtaining mutation site information of the sample to be tested based on the second single-cell sequencing data;
[0011] The mutation site information of the rare cells is screened out from the mutation site information of the sample to be tested according to the atlas of the rare cells.
[0012] In some embodiments of the present invention, the map is at least one of a genomic map, an expression map, and a chromatin accessibility map.
[0013] In some embodiments of the present invention, the profile is an expression profile, and obtaining the profile of rare cells according to the first single-cell sequencing data includes obtaining the rare cells according to A1 and A2 clustering;
[0014] A1. Expression levels of characteristic genes;
[0015] A2. At least one of copy number variation, signaling pathway activation level, or none of the above.
[0016] In some embodiments of the present invention, the rare cells are tumor cells.
[0017] In some embodiments of the present invention, obtaining the mutation site information of the sample to be tested based on the second single-cell sequencing data includes obtaining single-base mutations at different positions on the genome of the sample to be tested based on the comparison results of the sequencing read length on the reference genome.
[0018] In some embodiments of the present invention, the method further comprises performing genotyping on a single cell of the sample to be tested according to the single base mutation.
[0019] In some embodiments of the present invention, the first single-cell sequencing data and the second single-cell sequencing data are droplet-based single-cell sequencing data.
[0020] According to a second aspect of an embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the aforementioned method.
[0021] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and the processor implements the aforementioned method when running the computer program.
[0022] A fourth aspect of the embodiments of the present invention provides a system for detecting rare cell mutations based on single-cell sequencing, comprising:
[0023] an acquisition module, configured to acquire first single-cell sequencing data of a sample to be tested and second single-cell sequencing data of a target region in the sample to be tested;
[0024] a map construction module, configured to obtain a map of the rare cell based on the first single-cell sequencing data;
[0025] a mutation site information acquisition module, configured to acquire mutation site information of the sample to be tested based on the second single-cell sequencing data;
[0026] A mutation site information analysis module is used to filter out the mutation site information of the rare cell from the mutation site information according to the atlas of the rare cell.
[0027] The beneficial effects of the embodiments of the present invention are:
[0028] This method fills a technological gap in detecting rare cell point mutations using single-cell sequencing data, and can maximize the retention of relevant results for site mutations in rare cells. This includes the ability to automatically and batch-acquire rare cell mutation results, significantly improving analysis efficiency; it can simultaneously detect rare cell mutations and perform other analyses, including but not limited to genomic variation, gene expression, and chromatin accessibility, enhancing the dimension of cellular cognition; this method does not require prior mutation sites as a reference, which can reduce the cost of subsequent scientific and clinical applications; and this method is not limited by various single-cell sequencing technology platforms, expanding the scope of large-scale promotion of the technology.
[0029] Specifically, the method of the present invention can use various types of single-cell sequencing data from multiple platforms to detect single-base mutations in rare cells. Taking mutation detection of single-cell RNA sequencing data as an example, when constructing the sequencing library, in order to improve the accuracy of mutation detection, a dual sequencing library of the single-cell conventional transcriptome and the enriched transcripts of the target gene will be constructed at the same time; when detecting mutation sites, the first step is to align the sequencing data to the reference genome to obtain mutation location information; the second step is to perform cell barcode analysis to trace the cell source of the mutation; finally, the mutation sites are screened and filtered, and the single-cell expression profile information is combined to identify single-base mutations carried by rare cells. In summary, this method supports cell barcode analysis; performs mutation detection based on the single cell level, retaining rare mutations to the maximum extent; does not require a priori mutation sites as a reference; and is widely applicable to major single-cell DNA-seq / ATAC-seq / RNA-seq sequencing platforms including DNBlab C4 and 10×Genomics. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 The figure is a flow chart of a method for detecting rare cell mutations based on single-cell sequencing according to an embodiment of the present invention.
[0031] Figure 2 This is the identification result of rare cells in one embodiment of the present invention. A represents the cell composition of a blood sample identified using single-cell expression profiling, with CTCs representing circulating tumor cells, Candi representing suspected circulating tumor cells, Immune representing immune cells and blood cells, and LowQ representing low-quality data. B represents the expression of the hepatocyte marker gene ALB and the immune cell marker gene PTPRC in each cell component. C represents the enrichment of the hepatocyte and epithelial cell signature gene sets in each cell component. D represents the copy number variation in each cell component. E represents the enrichment of the Hallmark signaling pathway in each cell component.
[0032] Figure 3 This is the result of the effectiveness of single-base mutation detection in cells in one embodiment of the present invention. A is the increase in target gene reads after enrichment; B is the number of single-base mutations in tumor tissue covered by each cellular component in the blood; C is the number of in situ tumor mutations covered by abnormal cells in the blood as detected by RareSNP and SComatic; and D is the difference in the number of single-base mutations in tumor tissue covered by each cellular component in the blood as detected by RareSNP and SComatic.
[0033] Figure 4This is the detection result of a single-base mutation in circulating tumor cells in one embodiment of the present invention. A represents the gene and mutation type of the single-base mutation detected in the circulating tumor cells; B represents the signaling pathway affected by the mutated gene and the number of circulating tumor cells affected in the corresponding pathway. DETAILED DESCRIPTION
[0034] In the description of the present invention, the terms "first," "second," and "third" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first," "second," or "third" may explicitly or implicitly include at least one of such features. In the description of the present invention, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0035] The first aspect of the present invention provides a method for detecting rare cell mutations based on single cell sequencing, referring to Figure 1 , including the following steps S100 to S400. It is understandable that the method is not limited to being performed in the order of S100, S200, S300, and S400. For example, the method may also be performed in the order of S100, S300, S200, and S400.
[0036] S100: Acquire first single-cell sequencing data of a sample to be tested and second single-cell sequencing data of a target region in the sample to be tested.
[0037] The sample to be tested refers to a sample obtained from an individual animal, which may be, for example, a mammal. Mammals include, but are not limited to, animals of the orders Carnivora, Perissodactyla, Artiodactyla, Rodentia, Lagomorpha, and Primates, specifically any one of dogs, cats, pigs, sheep, cattle, horses, rabbits, mice, monkeys, orangutans, gorillas, chimpanzees, and humans. Sample types include samples from the skin, lungs, intestines, epithelium (including oral cavity, genitalia, etc.), and body fluids or secretions such as blood (such as peripheral blood), saliva, sputum, alveolar lavage fluid, bronchial brushings, and urine, as well as at least one sample of feces, intestinal contents, etc.
[0038] Among them, sequencing refers to the identification of base sequences of DNA or RNA. The object of conventional sequencing is a mixed sample of a large number of cells, which masks the heterogeneity between cells. Single-cell sequencing is the sequencing of genomes, transcriptomes, etc. at the level of individual cells. In order to achieve sequencing at the single-cell level, it is usually necessary to capture single cells in advance. Capture methods include limiting dilution, flow sorting, laser cutting, microscopy, etc., and the capture costs of these methods are relatively high. The new capture method with the help of microfluidic technology has the advantages of high throughput, fast speed, low cost, and high capture efficiency, and has therefore become the current mainstream capture method. This method constructs droplets of cells and cell barcodes so that each cell corresponds to a specific cell barcode (barcode), so that reads with the same cell barcode can be classified into the same cell.
[0039] The general sequencing process typically involves constructing a sequencing library and loading it onto a sequencing platform. Depending on the selected sequencing platform, different types of sequencing libraries can be constructed, such as single-end or paired-end, linear or circular. The sequencing data obtained after the sequencing library is unloaded from the platform typically consists of multiple reads. In some embodiments, the reads include the cellular nucleotide sequence and the cell barcode sequence. In some specific embodiments, the reads also include molecular markers.
[0040] In some specific embodiments, the cell barcode includes a nucleotide sequence consisting of N1 random bases, where N1 can be 5 to 30, for example, 5, 10, 15, 20, 25, or 30, and can form up to 4 5 , 4 10 , 4 15 , 4 20 , 4 25 , 4 30 The nucleotide sequence is composed of different random bases, so the length of the verified cell barcode can be determined according to the actual number of cells that may exist in the sample to be tested.
[0041] In some embodiments, the read length also includes a molecular marker (UMI). The molecular markers of the amplified products of the same read length template are the same, while the molecular markers of the amplified products of different read length templates are different. Therefore, the introduction of UMI eliminates PCR bias and achieves absolute quantification of gene expression. In some specific embodiments, the molecular marker includes a nucleotide sequence composed of N2 random bases, and N2 can be 5 to 20, such as 5, 10, 15, and 20, which can form up to 4 5 , 4 10 , 4 15 , 4 20 A nucleotide sequence composed of different random bases.
[0042] Conventional methods, such as cb_sniffer and VarTrix, require additional traditional large-scale sequencing (bulk) genome or transcriptome data as a reference when detecting single-cell nucleic acid mutations. Although SComatic solves the above problems, in practical applications, it is necessary to first analyze cell heterogeneity and perform mutation detection using a homogeneous cell population as the minimum analysis unit. However, the stringent filtering criteria involved may result in missing rare cells and are not suitable for mutation identification of rare cells. To this end, the method provided in the embodiments of the present application improves the sensitivity of detecting low-frequency mutations in rare cells by enriching nucleic acids in target regions of rare cells.
[0043] Specifically, a conventional single-cell sequencing library and a target region enriched nucleic acid sequencing library are constructed simultaneously to obtain first single-cell sequencing data for the sample under test and second single-cell sequencing data for the target region in the sample under test. This method filters mutation sites and, combined with single-cell expression profile information, identifies single-base mutations carried by rare cells, improving the accuracy of mutation detection.
[0044] Among them, the target area refers to a characteristic area associated with rare cells, and rare cells refer to cells or cell populations with a lower proportion in the sample to be tested, for example, cells or cell populations that account for less than 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% of the total number of cells in the sample to be tested. The proportion (number) of rare cells can be determined by any method known in the art, such as flow cytometry. The target area can be, for example, a target gene, and the corresponding characteristic area is a characteristic gene. A characteristic gene refers to a gene that is specifically expressed in a cell, such as a gene that is specifically highly expressed or specifically underexpressed. It is understood that the presence of a characteristic gene in one cell does not necessarily mean that it is not expressed or overexpressed in other cells. Therefore, in order to better confirm the expression profile of rare cells, the number of target genes is generally multiple, for example, 3, 5, 10, 20, 50, 100, 200, or 500 or more target genes. In other embodiments, the target region may also be the entire exome of the sample to be tested or a gene with a reported pathogenic hotspot mutation known to those skilled in the art.
[0045] S200: Obtaining a map of rare cells in the sample to be tested based on the first single-cell sequencing data.
[0046] Among them, the atlas refers to a summary of the relevant specific information reflecting the genetic material in the sample. It is understandable that different atlases can be generated according to different sequencing data information, including genome atlases, expression profiles, chromatin openness maps, etc. For example, the genome atlas refers to the map of the genomic structure of the cell, the expression profile refers to the related information such as the types and abundance of gene expression in the cell under a specific state, and the chromatin openness map refers to the map of the open regions of the chromosome, also known as the chromatin accessibility map. For example, if the first single-cell sequencing data is DNA-seq data, a genome atlas can be obtained; if the first single-cell data is RNA-seq data, an expression profile can be obtained; if the first single-cell data is ATAC-seq data, a chromatin openness map can be obtained. In an embodiment of the present application, mutations in rare cells are extracted by atlas, thereby achieving enrichment and detection of low-frequency mutations contained in rare cells. Therefore, the atlas of rare cells in the sample to be tested obtained based on the first single-cell sequencing data can be at least one of a genome atlas, an expression profile, and a chromatin openness map.
[0047] In some embodiments, the profile is an expression profile, and obtaining the expression profile of rare cells based on the first single-cell sequencing data includes clustering rare cells based on the expression levels of genes across the transcriptome. In some embodiments, obtaining the expression profile of rare cells based on the first single-cell sequencing data includes clustering rare cells based on the expression levels of signature genes, and at least one or none of copy number variation and signaling pathway activation levels to identify rare cells. Signature genes include, for example, genes specifically expressed in a cell type, hotspot mutation genes with no specific expression, or one or more of both.
[0048] In some specific embodiments, the rare cells are tumor cells. For example, when the sample to be tested is a blood sample, the rare cells may be circulating tumor cells. It is understood that rare cells may also be other cells with low levels in the corresponding sample to be tested.
[0049] S300: Obtaining mutation site information of the sample to be tested based on the second single-cell sequencing data. It is understandable that since the second single-cell sequencing data is sequencing data of the target region in the sample to be tested, the mutation site information of the sample to be tested is naturally also the mutation site information of the target region in the sample to be tested.
[0050] In some embodiments, in order to determine the mutation site information present in the sample to be tested, it is necessary to align the reads in the original sequencing data to the reference genome. In some embodiments, before the sequencing data is aligned, it is also included in the filtering and quality control, such as removing low-quality sequences and adapters contained in the sequencing data. In some embodiments, the filtering and quality control of the sequencing data can be performed by existing methods. In some embodiments, the low-quality sequences vary according to the requirements of different sequencing platforms and software. For example, it can be a sequence in which the low-quality bases in the read are above a certain threshold (such as the proportion of bases with a base quality less than 10 in the read is not less than 50%, etc.), and the reads with too low average quality values can be further removed. In some embodiments, removing adapters includes directly removing reads containing adapter sequences or removing adapter sequences in the reads. In addition, filtering and quality control also include removing reads with a high N ratio, such as reads with an N ratio greater than 5%.
[0051] In some embodiments, the reference genome is the entire human genome or a set range of the entire human genome, or may be the genome of one or more human individuals. In some embodiments, the reference genome includes, but is not limited to, GRCh36 (hg18), GRCh37 (hg19), GRCh38 (hg38), or any of the following:
[0052] In some embodiments, obtaining the mutation site information of the sample to be tested based on the second single-cell sequencing data includes obtaining the single-base mutation situation at different positions on the genome of the sample to be tested based on the comparison result of the sequencing read length on the reference genome. In some embodiments, the comparison result of the sequencing read length on the reference genome adopts a bam file, which is a binary format of a sam (sequence alignment / image format file, Sequence Alignment / Map Format) file, including relevant data for the specific comparison situation. In some embodiments, samtools mpileup is used to compare the similarities and differences of the reference bases at the corresponding positions of the bases in the first single-cell sequencing data of the sample to be tested, and the single-base mutation of the set genomic region is obtained, for example, it can be a single-base mutation of the whole genome or a single-base mutation situation within a set interval of the whole genome.
[0053] In some embodiments, genotyping of individual cells of the sample to be tested is also included based on single-base mutations. In some embodiments, genotyping includes site depth statistics, such as counting the number of supporting reads (i.e., depth) of the reference base and the mutant base at the base site, and calculating the total depth (DP) and mutant allele depth (AD) of each site. After filtering out sites with DP=0, if AD=DP, the genotype (GP) is specified as 1 / 1, otherwise the genotype is 0 / 1.
[0054] S400: Screening out the mutation site information of rare cells from the mutation site information of the sample to be tested according to the atlas of rare cells.
[0055] In some embodiments, the mutation site information of the sample to be tested is further filtered based on the cell classification information in the atlas to screen out the mutation site information of rare cells. For example, if the rare cells are tumor cells, by filtering out the single nucleotide polymorphism site information carried by normal cells, the mutation site information of the tumor somatic cells can be obtained and used for tumor cell characterization and disease mechanism research. In some embodiments, the atlas is an expression profile, and the cell classification information in the expression profile can be, for example, at least one of the expression types and abundances of the aforementioned cell characteristic genes and the corresponding cell barcodes.
[0056] In some embodiments, the first single-cell sequencing data is any one of DNA-seq data, RNA-seq data, and ATAC-seq data, and the second single-cell sequencing data is any one of DNA-seq data or RNA-seq data. For example, the first single-cell sequencing data and the second single-cell sequencing data are DNA-seq data, the first single-cell sequencing data and the second single-cell sequencing data are RNA-seq data, and the first single-cell sequencing data is ATAC-seq data and the second single-cell sequencing data is DNA-seq data. It is understood that if the first single-cell sequencing data is DNA-seq data, a genomic map is obtained; if the first single-cell sequencing data is RNA-seq data, an expression profile is obtained; and if the first single-cell sequencing data is ATAC-seq data, a chromatin accessibility map is obtained.
[0057] Taking obtaining the expression profile of a transcriptome as an example, in some embodiments, the method includes constructing and sequencing a single-cell transcriptome sequencing library, and analyzing the sequencing data;
[0058] The construction and sequencing of the single-cell transcriptome sequencing library includes the following steps:
[0059] Separate and capture cells based on high-throughput single-cell separation technology to obtain single-cell cDNA products;
[0060] The specific target gene sequence in the cDNA product is captured according to the probe of the target gene, thereby achieving the enrichment and purification of the target gene cDNA sequence;
[0061] A sequencing library was constructed based on the cDNA sequence and high-throughput sequencing was performed.
[0062] Analysis of sequencing data includes the following steps:
[0063] The single-cell sequencing data is converted into a sequence file and aligned to the reference genome to obtain the read length comparison results. The cell barcodes in the read length are mapped to specific cells, and the corresponding read length information is classified according to the cell type.
[0064] Compare the similarities and differences of the reference bases at the corresponding positions of the sequenced bases in the test sample to obtain the base mutations at different positions on the genome of the test sample;
[0065] Individual cells are genotyped based on the location of base mutations. After obtaining the genotype of each cell, the mutation sites are further filtered based on the cell classification information suggested in the expression profile, and the mutation site information of rare cells is screened out from the mutation site information of the sample to be tested.
[0066] According to a second aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the aforementioned method.
[0067] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including a processor and a memory, wherein the memory stores a computer program that can be executed on the processor, and the processor implements the aforementioned method when executing the computer program.
[0068] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs, such as the aforementioned method described in the embodiments of the present invention. The processor implements the aforementioned method by executing the non-transitory software programs and instructions stored in the memory.
[0069] The memory may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function, while the data storage area may store and execute the aforementioned programs. Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.
[0070] In some embodiments, the memory may include a memory remotely located relative to the processor, and the remote memory may be connected to the processor via a network. Examples of the aforementioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0071] The non-transitory software program and instructions required to implement the above method are stored in the memory, and when executed by one or more processors, the above method is performed.
[0072] A fourth aspect of the present invention provides a system for detecting rare cell mutations based on single-cell sequencing, comprising:
[0073] An acquisition module, configured to acquire first single-cell sequencing data of a sample to be tested and second single-cell sequencing data of a target region in the sample to be tested;
[0074] A map construction module, configured to obtain a map of rare cells based on the first single-cell sequencing data;
[0075] A mutation site information acquisition module is used to obtain the mutation site information of the sample to be tested based on the second single-cell sequencing data;
[0076] The mutation site information analysis module is used to screen out the mutation site information of rare cells from the mutation site information of the sample to be tested based on the atlas of rare cells.
[0077] In some embodiments, the profile is at least one of a genomic profile, an expression profile, and a chromatin accessibility profile.
[0078] In some embodiments, obtaining a rare cell profile based on the first single-cell sequencing data includes obtaining the rare cells based on A1 and A2 clustering; A1. expression levels of characteristic genes; A2. at least one or none of copy number variation and signal pathway activation levels. In some embodiments, the rare cells are tumor cells. In some embodiments, obtaining mutation site information of the sample to be tested based on the first single-cell sequencing data includes obtaining single-base mutations at different positions on the genome of the sample to be tested based on alignment results of sequencing reads on a reference genome. In some embodiments, the method further includes genotyping individual cells of the sample to be tested based on the single-base mutations.
[0079] In some embodiments, the first single-cell sequencing data and the second single-cell sequencing data are droplet-based single-cell sequencing data.
[0080] In some embodiments, the system further includes a sequencing module for sequencing the sample to be tested.
[0081] In some embodiments, the system further comprises a library construction module for constructing a sequencing library based on the extracted nucleic acid sequence.
[0082] In some embodiments, the system further includes a sequence extraction module for extracting corresponding nucleic acid sequences from the sample to be tested.
[0083] The device or system implementation described above is merely illustrative. The modules described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network elements. Some or all of these modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0084] It is understood that all or some of the steps disclosed above can be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). It is understood that computer storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include but are not limited to RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, disk storage or other magnetic storage device, or any other medium that can be used to store desired information and can be accessed by a computer.
[0085] Additionally, it will be appreciated that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0086] The present invention is further described in detail below through specific examples.
[0087] It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0088] The experimental methods in the following examples, where specific conditions are not specified, were generally performed under conventional conditions or the conditions recommended by the manufacturers. The materials and reagents used in these examples were commercially available unless otherwise specified.
[0089] Example 1
[0090] This example is based on the analysis of RNA sequencing data from peripheral blood single cells of liver cancer patients. The experimental process is as follows:
[0091] 1. Sample Collection
[0092] 1.1 Collect primary cancer tissue and peripheral blood samples from patients with liver cancer. During blood sample preparation, use CD45 immunomagnetic beads to remove white blood cells. During tumor tissue sample preparation, use a tissue dissociation kit to dissociate the tissue into single cells. If red blood cells are present, use red blood cell lysis buffer to remove them.
[0093] 2. cDNA Enrichment and Single-Cell Transcriptome Sequencing Library Preparation
[0094] 2.1 Prepare single-cell suspension. Combined with BGI's independently developed DNBelab C series high-throughput, portable single-cell library preparation system, the cell suspension is prepared into droplets encapsulating single cells.
[0095] 2.2 Cell lysis is completed in the droplet, and the linker on the magnetic beads in the droplet captures the mRNA and adds the cell barcode and UMI information. The cell barcode is a nucleotide sequence consisting of two 10bp long random bases. Theoretically, there are 4 10 ×4 10 Different barcode sequences, each barcode sequence can mark a cell. UMI is a nucleotide sequence composed of random bases in length of 10bp, which can theoretically be used to mark 4 10 Different cDNA molecules can be used to eliminate PCR bias and achieve absolute quantification of gene expression.
[0096] 2.3 Convert the captured mRNA into cDNA product through reverse transcription reaction.
[0097] 2.4 Reference the QuarXeq Pan-Cancer Panel 1.0 hybridization capture RNA probe kit (523 genes, Catalog No. NY1014C) from Shanghai Diying Biotechnology Co., Ltd. to identify hotspot mutations associated with pan-solid tumors. Rescreen the panel based on gene function and expression, and select 300 signature genes. Following the manufacturer's protocol, hybridize 1 μg of cDNA product with the target gene probe panel at 65°C for approximately 16 hours.
[0098] Table 1. List of characteristic genes
[0099] ABL2 ACVR1 ACVR1B AKT2 AKT3 ALK ALOX12B ANKRD11 ARAF ARID1A ARID1B ARID2 ASXL1 ASXL2 ATF1 ATM ATR ATRX AURKA AURKB AXL B2M BACH1 BARD1 BCOR BCR BLM BRAF BRCA1 BRCA2 BRD4 BTG1 BTK CARD11 CASP8 CBFB CBL CCNE1 CD274 CD74 CDC73 CDK12 CDK4 CDK6 CDK8 CDKN1B CHD1 CHD2 CHD4 CHEK1 CHEK2 CIC CRBN CREBBP CSF1R CSF3R CTCF CTLA4 CTNNA1 CTNNB1 CUL3 CUL4B CXCR4 DAXX DDR2 DICER1 DNAJB1 DNMT1 DNMT3A DNMT3B DOT1L EGFR EIF1AX ELOC EMSY EP300 EPHA2 EPHA5 EPHA7 EPHB1 ERBB2 ERBB3 ERBB4 ERCC1 ERCC3 ERCC4 ERCC5 ERG ERRFI1 ESR1 EWSR1 FANCA FANCC PMAIP1 PMS1 PMS2 POLD1 POLE PPM1D PPP2R2A PPP6C PRDM1 PREX2 PRKAR1A PRKCI PRKDC PTCH1 PTEN PTK2 PTPRD PTPRS PTPRT QKI RAC1 RAD21 RAD51 RAD51B RAD51C RAD51D RAD52 RAF1 RANBP2 RASA1 RB1 RBM10 REL RET RHEB RHOA RICTOR RIT1 RNF43 ROS1 RPA1 RPS6KB2 RPTOR RUNX1 RYBP SDHA SDHB SDHC SDHD SF3B1 SH2B3 SH2D1A SLX4 SMAD2 SMAD3 SMARCB1 SMARCD1 SNCAIP SOX9 SPEN SPOP SPTA1 SRSF2 STAG2 STAT3 STAT5B SUFU SYK TAF1 TBX3 TCF7L2 TENT5C TET2 TGFBR2 TIPARP TMPRSS2 TNFAIP3 TNFSF11 TOP1 TP63 TRRAP TSC1 TSC2 TSHR VEGFA VEGFB VHL XPO1 YAP1 ZBTB2 ZFHX3 ZNF217
[0100] 2.5 Add streptavidin magnetic beads to capture the target cDNA sequence.
[0101] 2.6 Amplify and purify the target cDNA.
[0102] 3. Single-cell transcriptome sequencing
[0103] 3.1 After fragmentation and end-repair of the purified cDNA sequence, sequencing adapters are added. The product is circularized to construct a sequencing library.
[0104] 3.2 Perform high-throughput sequencing on the library using platforms such as DNBSEQ to obtain transcriptome information of the targeted gene and single-cell whole transcriptome information.
[0105] 4. Data Analysis
[0106] 4.1 Data preprocessing
[0107] 4.2 Use PISA software to align, split, and quality control the sequencing data to obtain the expression data of cells in the transcriptome library and the alignment information of reads in the enriched cDNA library.
[0108] 4.3 Rare Cell Identification
[0109] The rare cells in this embodiment are circulating tumor cells that exist in trace amounts in the peripheral blood of liver cancer patients. According to the expression level of characteristic genes, copy number variation, and signal pathway activation level (refer to Figure 2 Comprehensive clustering such as BE) was used to identify rare tumor cells in blood samples (reference Figure 2 A).
[0110] The rare cells identified by this method exhibit the following characteristics: 1) significantly higher expression of liver cancer-related genes compared to other cell types present in peripheral blood; 2) a significant increase in copy number; and 3) significantly higher levels of activation of oxidative phosphorylation, fatty acid metabolism, bile acid metabolism, and pro-angiogenic signaling compared to normal cells. These characteristics are consistent with the characteristics of circulating tumor cells currently identified in liver cancer.
[0111] 4.4 Single-cell single-base mutation detection
[0112] This method (RareSNP) was used to enrich cancer hotspot mutation genes and detect mutations. Based on the difference in sequencing depth between reference bases and mutant bases at each site in the target gene region of each cell, the single-cell genotype was defined, and the number of mutation sites with genotypes of 1 / 1 and 0 / 1 in each cell was counted. The results are shown in Figure 2. Figure 3 As shown in Figure A, the average sequencing depth of hotspot mutation genes after enrichment increased by ≥2800%, indicating that RareSNP is beneficial for capturing low-frequency mutation sites occurring in single cells. In addition, in the patient's peripheral blood, the number of tumor cell mutation sites detected by this method covering the patient's primary tumor mutations was significantly higher than that detected in immune cells (the number of mutations increased by 300%), and the mutations covered genes closely related to tumor occurrence and development, such as BAP1, PIK3CA, and FBXW7. Figure 3 As shown in B.
[0113] The RareSNP method provided in this example was compared with SComatic, a method for detecting rare cell mutations based on single-cell sequencing of cell populations. The specific steps of SComatic were referenced to Muyas, F., Sauer, CM, Valle-Inclán, JE et al. De novo detection of somatic mutations in high-throughput single-cell profiling data sets. Nat Biotechnol (2023). The same samples and libraries were used.
[0114] The results show Figure 3 C and D, the RareSNP of Example 1 detected a total of 183 primary tumor mutation sites in circulating tumor cells of liver cancer, while the cell population-based detection method SComatic of Comparative Example 1 could not detect them. Moreover, the number of primary tumor mutation sites detected by RareSNP in various cell components in blood samples was significantly higher than that detected by Somatic, indicating that RareSNP has better performance in detecting rare cell mutation characteristics. At the same time, Figure 4 As shown in the figure, after further filtering the above mutation features using conventional cellular single nucleotide polymorphisms, RareSNP can also detect tumor somatic mutations in non-cancer adjacent control samples.
[0115] In summary, the method provided by the present invention detected mutation sites consistent with primary tumor cells in abnormal rare cells in blood samples, which were not detected by existing single-cell base mutation detection methods such as SComatic. This result demonstrates that the method provided by the present invention is significantly superior to other methods in detecting rare mutations. In addition, the results of comparing the rare mutations detected in blood samples by the method provided by the present invention with mutations in the primary tumor site of the same patient also demonstrate that the method can stably and reliably detect rare mutations originating from primary tumor lesions in blood samples.
[0116] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A method for detecting rare cell mutations based on single-cell sequencing, characterized in that: The following steps are involved: Acquire first single-cell sequencing data of a sample to be tested and second single-cell sequencing data of a target region in the sample to be tested; Obtaining a map of the rare cells in the sample to be tested based on the first single-cell sequencing data; Obtaining mutation site information of the sample to be tested based on the second single-cell sequencing data; The mutation site information of the rare cells is screened out from the mutation site information of the sample to be tested according to the atlas of the rare cells.
2. The method according to claim 1, wherein: The map is at least one of a genome map, an expression map and a chromatin accessibility map.
3. The method according to claim 2, wherein: The profile is an expression profile, and obtaining the profile of rare cells according to the first single-cell sequencing data includes obtaining the rare cells according to clustering A1 and A2; A1. Expression levels of characteristic genes; A2. At least one of copy number variation, signaling pathway activation level, or none of the above.
4. The method according to claim 1, wherein: The rare cells are tumor cells.
5. The method according to claim 1, wherein: Obtaining mutation site information of the sample to be tested based on the second single-cell sequencing data includes obtaining single-base mutations at different positions on the genome of the sample to be tested based on the comparison results of the sequencing read lengths on the reference genome.
6. The method according to claim 5, characterized in that: The method also includes performing genotyping on a single cell of the sample to be tested according to the single base mutation.
7. The method according to claim 1, wherein: The first single-cell sequencing data and the second single-cell sequencing data are droplet-based single-cell sequencing data.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.
9. An electronic device, characterized in that The method comprises a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and the processor implements the method according to any one of claims 1 to 7 when running the computer program.
10. A system for detecting rare cell mutations based on single-cell sequencing, characterized in that: include: an acquisition module, configured to acquire first single-cell sequencing data of a sample to be tested and second single-cell sequencing data of a target region in the sample to be tested; a map construction module, configured to obtain a map of the rare cell based on the first single-cell sequencing data; a mutation site information acquisition module, configured to acquire the mutation site information of the sample to be tested based on the second single-cell sequencing data; A mutation site information analysis module is used to filter out the mutation site information of the rare cell from the mutation site information according to the atlas of the rare cell.
Citation Information
Patent Citations
Method and system for identifying tumor cell groups in single cell transcriptome sequencing data
CN115083521A
Multiplexed droplet-based sequencing using natural genetic barcodes
US20220005547A1
Systems and methods to detect rare mutations and copy number variation
US20220389489A1