Sequencing read-based methods for single-cell whole genome analysis of CpG methylation and related devices
By performing sequencing read alignment and Hamming distance screening on single-cell methylation data, the problems of low efficiency and insufficient sensitivity in single-cell methylation analysis were solved, achieving efficient and accurate methylation pattern detection and cell type identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU MEDICAL UNIV
- Filing Date
- 2025-08-14
- Publication Date
- 2026-04-24
AI Technical Summary
Existing single-cell methylation analysis methods suffer from low analysis efficiency and insufficient sensitivity. In particular, when processing high-dimensional and sparse single-cell methylation data, traditional methods are prone to losing base-level resolution and biologically relevant differences, have high computational costs, and are difficult to screen for specific methylation sites.
By comparing methylation data sequencing reads of single cells to be analyzed, methylation-specific regions are screened based on random sampling strategy and Hamming distance of sequencing reads. Hamming distance matrix is calculated for initial clustering, reducing data size and feature dimension, and improving sensitivity by combining Hamming distance metric.
It improves the efficiency and sensitivity of single-cell methylation analysis, effectively reduces computational complexity, improves clustering efficiency and accuracy, and supports more granular biological insights.
Smart Images

Figure CN121260264B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of gene analysis technology, and in particular to a method and related equipment for single-cell whole-genome analysis based on CpG methylation of sequencing reads. Background Technology
[0002] DNA methylation (DNAm) refers to the process by which nucleotide bases in DNA undergo methylation through enzymatic reactions. In animal and many plant genomes, cytosine-guanine dinucleotide (CpG) is the most common methylation target. Abnormal CpG methylation patterns reflect epigenetic heterogeneity, including in diseases. Therefore, DNA methylation data are widely used for disease typing to study the epigenetic landscape associated with clinical prognosis, diagnosis, and phenotypic characteristics. However, existing methods for single-cell methylation analysis still suffer from low analytical efficiency and insufficient sensitivity. Summary of the Invention
[0003] The main objective of this application is to propose a method and related equipment for single-cell whole-genome analysis of CpG methylation based on sequencing reads, aiming to improve the efficiency and sensitivity of single-cell methylation analysis.
[0004] To achieve the above objectives, one aspect of this application proposes a method for single-cell whole-genome analysis of CpG methylation based on sequencing reads, the method comprising the following steps:
[0005] Obtain single cells to be analyzed, and perform DNA methylation data sequencing read alignment on the single cells to obtain statistical results of methylation sites;
[0006] Screening of methylation-specific regions based on random sampling strategy and Hamming distance of sequencing reads;
[0007] Based on the statistical results of the methylation sites and the methylation-specific intervals, the Hamming distance matrix of each pair of the CpG sites in the single cells to be analyzed is calculated. Based on the Hamming distance matrix, the initial clustering is performed to obtain the initial clustering results.
[0008] In some embodiments, the method further includes:
[0009] For each single cell to be analyzed, a set of cells of the same type and a set of cells of different types are constructed based on the initial clustering results.
[0010] Calculate the first average Hamming distance between each of the single cells to be analyzed and the set of cells of the same type, and calculate the second average Hamming distance between each of the single cells to be analyzed and the set of cells of different types;
[0011] Calculate the cellular heterogeneity measure for each of the single cells to be analyzed based on the first average Hamming distance and the second average Hamming distance;
[0012] Feature extraction and clustering were performed on all the cellular heterogeneity measures of the single cells to be analyzed to obtain methylated cell clustering results.
[0013] In some embodiments, the method further includes:
[0014] Based on the cell heterogeneity measures of all the single cells to be analyzed and the cell clustering results, a set of specific intervals for all the single cells to be analyzed is selected.
[0015] Heuristic search of the specific region set based on methylation site patterns was performed to screen out differentially methylated site patterns.
[0016] Based on the differential methylation pattern sites, the methylation patterns of different cell types at all sites are calculated to obtain a panoramic view of specific methylation patterns.
[0017] In some embodiments, the method includes:
[0018] Based on the panoramic view of the specific methylation pattern, effective read segments are screened.
[0019] Based on the methylation patterns of the effective reads and the methylation patterns of each cell type, the similarity between each read and each cell type is calculated.
[0020] Based on the similarity between each read segment and each cell type, and the maximum posterior probability, the proportion of all the cell components to be analyzed is estimated to obtain the estimated proportion of each cell type.
[0021] In some embodiments, the step of performing DNA methylation data sequencing read alignment on the single cell to be analyzed to obtain methylation site statistical results includes:
[0022] The paired-end sequencing data of the single cells to be analyzed were quality controlled and pruned to remove the tag sequence. The DNA reads were aligned with the sequencing library to obtain the initial alignment results.
[0023] Low-quality alignment results in the initial alignment results are filtered out, sorted, and deduplicated to obtain statistical results of methylation sites.
[0024] In some embodiments, the screening of methylation-specific regions based on the methylation Hamming distance of sequencing reads using a random sampling strategy includes:
[0025] Based on a random sampling strategy, pseudo-batch data of methylation data is constructed, and a list of intervals is built based on the whole genome.
[0026] Calculate the methylated Hamming distance set for each interval read segment in the interval list based on the pseudo-batch data;
[0027] The methylation Hamming distance statistical characteristics of each interval read are determined based on the methylation Hamming distance set of each interval read, and the methylation-specific intervals in the interval list are screened based on the methylation Hamming distance statistical characteristics of each interval read.
[0028] In some embodiments, the Hamming distance matrix between each pair of the single-cell CpG sites to be analyzed is calculated based on the statistical results of the methylation sites and the methylation-specific regions, including:
[0029] Extract CpG site information for each of the single cells to be analyzed from the methylation-specific region;
[0030] The Hamming distance matrix between each pair of cells to be analyzed is calculated based on the CpG site information of each cell to be analyzed.
[0031] To achieve the above objectives, another aspect of this application proposes a single-cell whole-genome analysis device based on sequencing reads for CpG methylation, the device comprising:
[0032] The site statistics module is used to acquire the single cell to be analyzed, perform DNA methylation data sequencing read alignment on the single cell to be analyzed, and obtain methylation site statistics results.
[0033] The methylation-specific interval calculation module is used to screen methylation-specific intervals based on random sampling strategies and Hamming distances of sequencing reads.
[0034] The clustering analysis module is used to calculate the Hamming distance matrix between each pair of CpG sites in the single cells to be analyzed based on the statistical results of the methylation sites and the methylation-specific intervals, and to perform initial clustering based on the Hamming distance matrix to obtain the initial clustering results.
[0035] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the methods described above.
[0036] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0037] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the methods described above.
[0038] The embodiments of this application include at least the following beneficial effects: This application provides a method, device, electronic device, and program product for CpG methylation single-cell whole-genome analysis based on sequencing reads. This scheme obtains methylation site statistical results by comparing sequencing reads of methylation data of the single cell to be analyzed. Based on a random sampling strategy and the methylation Hamming distance of the sequencing reads, methylation-specific intervals are screened. The Hamming distance matrix of CpG sites of the single cell to be analyzed is calculated according to the methylation site statistical results and the methylation-specific intervals. Initial clustering is performed based on the Hamming distance matrix to obtain the initial clustering results. The methylation Hamming distance between different reads is calculated according to the interval division to screen methylation-specific intervals, reducing the data size read into memory at one time and improving analysis efficiency. By selecting methylation-specific intervals, the feature dimension can be effectively reduced. Combined with the measurement of Hamming distance to improve sensitivity, the clustering efficiency and accuracy can be effectively improved. Attached Figure Description
[0039] Figure 1 This is a flowchart of a single-cell whole-genome analysis method based on sequencing reads for CpG methylation, provided in an embodiment of this application.
[0040] Figure 2 This is a flowchart illustrating the statistical results of methylation sites provided in the embodiments of this application;
[0041] Figure 3 This is a flowchart illustrating the screening of methylation-specific regions provided in an embodiment of this application;
[0042] Figure 4 This is a flowchart of the calculation of the Hamming distance matrix provided in the embodiments of this application;
[0043] Figure 5 This is a flowchart illustrating the methylated cell clustering results provided in the embodiments of this application;
[0044] Figure 6 This is a flowchart illustrating the process of obtaining a panoramic view of specific methylation patterns, provided in an embodiment of this application.
[0045] Figure 7 This is a flowchart illustrating the process of obtaining the estimated proportion of each cell type, as provided in an embodiment of this application.
[0046] Figure 8 This is a flowchart of a specific embodiment provided in this application;
[0047] Figure 9 This is a schematic diagram of the structure of a single-cell whole genome analysis device based on sequencing reads provided in an embodiment of this application;
[0048] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0050] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0051] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0053] Among related technologies, single-cell whole-genome bisulfite sequencing (scWGBS) is a breakthrough technique that can resolve DNA methylation maps at the single-cell level, providing important evidence for epigenetic regulation, cellular heterogeneity, and gene expression dynamics research. This high-resolution method enables researchers to explore methylation patterns in different cell populations, revealing the molecular basis of cell identity, disease states, and developmental processes. However, scWGBS data is characterized by high dimensionality and high sparsity. Traditional analysis workflows, in order to address sparsity and high dimensionality, typically employ binning-based operations, which often result in the loss of base-level resolution due to excessive smoothing. This masks biologically relevant differences in sparsely covered regions or at low sequencing depths, leading to inefficiency and poor sensitivity. Therefore, effectively processing high-dimensional and sparse single-cell methylation data while maintaining high sensitivity to changes in methylation patterns is crucial for extracting biological insights from such complex data.
[0054] Furthermore, the current lack of screening for specific methylation sites under different conditions based on single-cell methylation data limits further applications in clinical settings. The following technical issues will be addressed: (1) Traditional methods (such as binning-based methylation ratio quantification strategies) often lose base-level resolution due to excessive smoothing, potentially masking biologically relevant differences in sparsely covered regions or at low sequencing depths; (2) The high computational cost of processing large-scale datasets remains a major bottleneck. Methods based on pairwise comparisons of inter-cell reads (such as Hamming distance calculation) face a surge in computational complexity as the data scale increases, with time complexity quadratically related to the number of reads or cells; (3) How to identify highly variable methylation-specific regions in cells under different conditions to provide a basis for scenarios such as deconvolution. Therefore, it is urgent to develop more efficient and sensitive scWGBS data analysis methods for single-cell methylation data.
[0055] In view of this, this application provides a method and related equipment for single-cell whole-genome analysis of CpG methylation based on sequencing reads. This method compares sequencing reads of methylation data from the single cells to be analyzed to obtain statistical results of methylation sites. Based on a random sampling strategy and the Hamming distance of methylation in the sequencing reads, methylation-specific intervals are screened. Based on the statistical results of methylation sites and the methylation-specific intervals, a Hamming distance matrix of CpG sites of each pair of single cells to be analyzed is calculated. Initial clustering is performed based on the Hamming distance matrix to obtain initial clustering results. The Hamming distance of methylation between different reads is calculated according to the interval division to screen methylation-specific intervals, reducing the data size read into memory at one time and improving analysis efficiency. By selecting methylation-specific intervals, the feature dimension can be effectively reduced, and the sensitivity can be improved by combining the measurement of Hamming distance, which can effectively improve the clustering efficiency and accuracy.
[0056] The single-cell whole-genome analysis method based on sequencing reads for CpG methylation provided in this application relates to the field of information technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the single-cell whole-genome analysis method based on sequencing reads, but is not limited to the above forms.
[0057] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0058] Figure 1 This is an optional flowchart of a single-cell whole-genome analysis method based on sequencing reads for CpG methylation provided in this application embodiment. Figure 1 The method may include, but is not limited to, steps S101 to S103.
[0059] Step S101: Obtain the single cell to be analyzed, perform DNA methylation data sequencing read alignment on the single cell to be analyzed, and obtain the statistical results of methylation sites;
[0060] Step S102: Screening methylation-specific regions based on random sampling strategy and Hamming distance of methylation in sequencing reads;
[0061] Step S103: Calculate the Hamming distance matrix of CpG sites in each pair of single cells to be analyzed based on the statistical results of methylation sites and methylation-specific intervals. Perform initial clustering based on the Hamming distance matrix to obtain the initial clustering results.
[0062] Based on GPU, methylation data sequencing reads of single cells to be analyzed are aligned to obtain methylation site statistics, i.e., the position of each read of methylation data of each single cell in the reference genome. Batch data is constructed based on a random sampling strategy, and methylation-specific regions are screened based on the Hamming distance of methylation between the batch data and the sequencing reads. Based on the screened methylation-specific regions, CpG site information of each cell in the regions is extracted. Based on the CpG site information of each cell, the Hamming distance matrix between each pair of cells is calculated. A hierarchical clustering method based on Euclidean distance is used to obtain a dendrogram of the initial cluster, and appropriate parameters are selected to obtain the initial clustering results.
[0063] In some embodiments, see Figure 2 DNA methylation data from the single cells to be analyzed was sequenced and reads were compared to obtain statistical results of methylation sites, including:
[0064] Step S201: Perform quality control and trimming on the paired-end sequencing data of the single cells to be analyzed, remove the tag sequence, and perform DNA read alignment based on the sequencing library to obtain the initial alignment results;
[0065] Step S202: Filter out low-quality alignment results in the initial alignment results, sort and remove duplicates to obtain statistical results of methylation sites.
[0066] Cutadapt was used for quality control and trimming of paired-end sequencing data to remove barcode (tag sequences). Cutadapt is a software primarily used to remove unwanted sequences such as adapter sequences, primers, and poly-A tails from high-throughput sequencing data. It supports exact and fuzzy matching modes and provides various quality control functions. Subsequently, based on sequencing library information, bismark (a methylation sequencing data alignment tool) was used for bisulfite transformation sequence alignment (PBAT mode for R1 end, standard alignment for R2 end). Then, samtools (a toolset for processing high-throughput sequencing data) was used to filter and sort low-quality alignment results. Picard (a labeling and deduplication tool) was used to remove PCR repeats. Finally, paired-end data were merged to generate the final BAM file and methylation site statistics (ALLC format). To utilize GPU acceleration, arioc (a GPU-accelerated DNA short-read alignment tool) was used to achieve GPU-accelerated alignment of single-cell methylation data, improving alignment efficiency.
[0067] To improve the alignment rate and coverage of individual cell data, two post-processing methods are supported for unaligned reads: One method involves adding the "-un" parameter during bismark alignment to save the unaligned results, then using samtools to generate new R1 and R2 values, and re-performing one-end alignment on the unaligned reads; the other method uses a fragment-based alignment approach, segmenting the reads and adding a fault tolerance parameter, then re-performing bismark alignment on the unaligned data. Finally, the results from the first and second alignments are merged using the samtools merge command to obtain the overall alignment result. Finally, deduplication is performed on the aligned data based on "deduplicate_bismark" or "picard" to obtain the final alignment result. The entire workflow is modularly designed, supports breakpoint resume and quality control, and saves log output for key steps, making it suitable for large-scale single-cell methylation data alignment.
[0068] This method integrates GPU-accelerated alignment technology and fragmented alignment strategy to optimize segment localization, comprehensively improving the efficiency of data flow calculation. It can achieve more than double the alignment efficiency and increase the alignment rate by more than 10%.
[0069] In some embodiments, see Figure 3 Based on random sampling strategies and Hamming distance of methylation from sequencing reads, methylation-specific regions were screened, including:
[0070] Step S301: Based on a random sampling strategy, construct pseudo-batch data of methylation data and construct an interval list based on the whole genome;
[0071] Step S302: Calculate the methylated Hamming distance set for each interval read segment in the interval list based on the pseudo-batch data;
[0072] Step S303: Determine the methylation Hamming distance statistical features of each interval reading segment based on the methylation Hamming distance set of each interval reading segment, and filter the methylation-specific intervals in the interval list based on the methylation Hamming distance statistical features of each interval reading segment.
[0073] The batch number N and batch size S are set based on the cell count, sequencing depth, etc. of the dataset. For each batch, S cells are randomly selected from all samples, and the methylation data of the corresponding cells are merged into pseudo-batch data. This process is repeated N times. The whole genome is divided into intervals of a fixed size (e.g., 5000), resulting in a list of genome intervals. Within these intervals, based on the merged pseudo-batch data, the Hamming distance matrix between all pairwise reads in each interval is calculated, and the union of the calculated matrices is summed to obtain N Hamming distances for each interval. Based on the statistical characteristics of the Hamming distances in each interval, including mean, variance, and coefficient of variation, methylation-specific intervals are selected according to these statistical characteristics. Since the intervals and N are usually large, a locally weighted Hamming distance based on bit operations is introduced to improve the efficiency of Hamming distance calculation. To avoid the influence of missing values, a masking strategy is used to ignore missing sites, considering only CpG sites that match between cells. CpG sites are specific regions in DNA molecules where cytosine (C) and guanine (G) alternate, and their sequence pattern is CG. Simultaneously, the distribution of CpG sites matched between each pair of cells and their corresponding weighted Hamming distances was statistically analyzed and used as a criterion for outlier screening.
[0074] Hamming distance was used to capture methylation differences at the base level (CpG sites) between reads. The number of cells and the batch number were set for each batch, and pseudo-batch data was constructed based on a random sampling strategy. Statistical characteristics of methylation Hamming distances, such as the standard deviation of Hamming distances over each fixed interval, were calculated on the pseudo-batch data. Intervals with larger variances were selected as candidate intervals (containing information on different cell types and diseases). This method significantly reduces time complexity and calculates intercellular heterogeneity without sacrificing sensitivity.
[0075] By randomly downsampling to construct pseudo-batch data and calculating the average Hamming distance between different read segments within an interval according to the interval division, the size of data read into memory at one time is effectively reduced, making the framework scalable.
[0076] In some embodiments, see Figure 4 Based on the statistical results of methylation sites and methylation-specific regions, the Hamming distance matrix of pairwise CpG sites in the single cells to be analyzed is calculated, including:
[0077] Step S401: Extract CpG site information for each single cell to be analyzed in the methylation-specific region;
[0078] Step S402: Calculate the Hamming distance matrix between each pair of cells to be analyzed based on the CpG site information of each cell to be analyzed.
[0079] Based on the selection intervals, wgbstools is used to extract the CpG site information of each cell within the intervals, generating new bed files. Based on the bed of each cell, the Hamming distance matrix dist between each pair of cells is calculated.
[0080] The statistical characteristics of Hamming distance in each interval after repeated sampling are calculated. High-information methylation intervals are selected, and clustering of single-cell methylation data is performed based on these intervals. By selecting intervals, the feature dimensionality can be effectively reduced, and the Hamming distance metric has high sensitivity, which can effectively improve the clustering efficiency and accuracy.
[0081] In some embodiments, see Figure 5 The methods also include:
[0082] Step S501: For each single cell to be analyzed, construct a set of cells of the same type and a set of cells of different types based on the initial clustering results;
[0083] Step S502: Calculate the first average Hamming distance between each single cell to be analyzed and a set of cells of the same type, and calculate the second average Hamming distance between each single cell to be analyzed and a set of cells of different types.
[0084] Step S503: Calculate the cellular heterogeneity measure for each single cell to be analyzed based on the first average Hamming distance and the second average Hamming distance;
[0085] Step S504: Feature extraction and clustering are performed based on the cellular heterogeneity measure of all single cells to be analyzed to obtain methylated cell clustering results.
[0086] While Hamming distance based on the screening interval can achieve clustering and other analyses, it cannot support more fine-grained downstream analysis scenarios, such as differential methylation pattern screening. Therefore, a cell heterogeneity measure, Diff, is designed: For each cell i, two sets are constructed based on the initial clustering results: a set of cells of the same type, Si, representing all cells whose initial clustering results are consistent with i; and a set of cells of different types, Sio, representing all cells whose initial clustering results are different from i. The average Hamming distance Diff for all cells in i and Si is calculated separately. i And the average Hamming distance Diff for all cells of i and Sio io Subtracting the two yields the single-cell methylation data representation based on Diff. io -Diff i The data is denoted as an N x M matrix, where N represents the cell data and M is the number of selected intervals. The data is then processed through dimensionality reduction, imputation, and batch de-batching to obtain an N x d matrix F, where d represents the PCA feature extraction dimension. Based on F, Leiden clustering is used to obtain methylation-based cell clustering results.
[0087] Based on the initial clustering results, a set of cells of the same type and a set of cells of different types are constructed for each cell. The average Hamming distance between a cell and cells of the same type is calculated, as well as the average Hamming distance between a cell and cells of different types. The difference between these two distances yields a measure of cell heterogeneity, the Diff. This Diff is highly sensitive and can support various analytical scenarios based on single-cell methylation data, such as clustering and identification of specific sites under different conditions.
[0088] In some embodiments, see Figure 6 The methods also include:
[0089] Step S601: Based on the cell heterogeneity measurement and cell clustering results of all single cells to be analyzed, filter the set of specific intervals of all single cells to be analyzed.
[0090] Step S602: Perform a heuristic search on the set of specific regions based on the methylation site patterns to screen out differentially methylated pattern sites;
[0091] Step S603: Based on the differential methylation pattern sites, calculate the methylation patterns of different cell types at all sites to obtain a panoramic view of specific methylation patterns.
[0092] Based on the Diff matrix and corresponding clustering information, differential intervals are selected for each cell type, satisfying the following conditions: cells of the same type maintain largely the same pattern within this interval, while cells of different types show significant differences in pattern within this interval. A set of specific intervals S is then selected based on a t-test. bin Since not all sites within the set are specific sites, a heuristic search is performed based on methylation site patterns, and statistical methods are used for verification and secondary screening to obtain differentially methylated site patterns. Methylation patterns for all cell types at all sites are calculated, and finally, an atlas of specific methylation patterns for different cell types is obtained.
[0093] To extract type-specific methylation sites, a t-test was used based on cellular diffs to first identify type-specific methylation regions. Since these regions may contain non-specific sites, a heuristic search was further performed using methylation features based on the selected regions to screen for type-specific methylation sites, which were then further tested. This cascaded screening method improves the efficiency of specific site screening while maintaining CpG site-level accuracy.
[0094] Based on the Diff distribution combined with t-tests, type-specific differential methylation intervals were screened. Furthermore, based on the screened intervals, the patterns of CpG sites within each interval were statistically analyzed to identify type-specific differential methylation arrangements. Preliminary interval screening ensured the efficiency of the process, while further screening of CpG sites improved sensitivity.
[0095] In some embodiments, see Figure 7 The methods include:
[0096] Step S701: Based on the specific methylation pattern panorama, filter valid reads;
[0097] Step S702: Calculate the similarity between each read and each cell type based on the methylation pattern of the effective read and the methylation pattern of each cell type;
[0098] Step S703: Estimate the proportion of all cell components to be analyzed based on the similarity between each read segment and each cell type and the maximum a posteriori probability, and obtain the estimated proportion of each cell type.
[0099] For batch data, based on the aforementioned differential methylation pattern panorama, each read in the batch is matched to the site corresponding to the differential methylation pattern. If a read does not contain any pattern in the panorama, it is considered an invalid read. For valid reads, the Hamming distance between the read's methylation pattern and the methylation pattern of each cell type is calculated, serving as the difference between the read and that class. Since the standard Hamming distance is for binary vectors, the methylation pattern at specific sites for each class does not meet the condition. Here, a soft Hamming distance is used: the absolute value of the difference between two vectors is used instead of the traditional Hamming distance. The similarity between each read and each cell type is calculated based on the soft Hamming distance. Considering the sparsity and randomness of sequencing, a softmax-based activation function is added to avoid homogenization of read classification results, thereby minimizing the entropy value of the read classification results. Reads are further filtered based on the entropy value of the classification results to select reads with specificity, providing a basis for subsequent cell component estimation. Based on the above read classification results (including the probability that each read belongs to each class), the cell components of the batch data are estimated based on the maximum a posteriori probability. A parameterized prior distribution is set based on prior knowledge; a log-posterior function is constructed, combining the likelihood term with the prior term; finally, optimization algorithms such as gradient descent or variational inference are used to find the maximum posterior solution to obtain the estimated proportion of each cell type. To reduce the influence of the prior distribution on the results, two methods are introduced: one is based on a grid search strategy, which samples the prior distribution with a fixed step size. This method can guarantee the optimal result when the number of parameters (cell types) is small, but the search cost increases exponentially as the number of cell types increases; the other method adopts a heuristic tree search strategy, which can effectively reduce invalid searches based on existing historical results, reduce the search space, significantly reduce search costs, and usually achieve better results.
[0100] To apply this method to the deconvolution task, a simulated bulk dataset was constructed using a randomized mixed single-cell approach. Based on the arrangement patterns of continuously specific CpG methylation states, reads were first screened and then classified according to their methylation patterns, achieving sequencing read-level classification. To estimate the components of each cell type in the bulk dataset, a maximum a posteriori probability-based method was used, estimating the components of the pseudo-bulk dataset based on the read classification results. To eliminate the influence of prior distributions, a tree-based heuristic search was further introduced to optimize the final results. A similar deconvolution process was used for real cfDNA data.
[0101] See below Figure 8 The following is a detailed introduction and explanation of the scheme of the present invention, combined with a specific scenario of single-cell whole-genome analysis of CpG methylation based on sequencing reads:
[0102] The algorithm employs GPU-accelerated alignment and multi-stage alignment with read segment segmentation. Multiple pseudo-batch datasets are generated based on random sampling. The average weighted Hamming distance between each pair of reads within a given interval is calculated using these pseudo-batch datasets. Based on the statistical characteristics of the Hamming distance within each interval, high-information intervals are selected. For the selected high-information intervals, the methylation pattern of each cell is extracted, and the weighted Hamming distance matrix between each pair of cells is calculated. Initial clustering is achieved using a hierarchical clustering method with Euclidean distance based on the Hamming distance matrix. The heterogeneity representation (Diff) of each cell is recalculated based on the initial clustering results. Clustering is then performed using TSNE based on the Diff, combined with padding, PCA, and normalization. Differential intervals are selected based on the Diff, and differential sites are selected within each interval to construct a panoramic view of methylation patterns representing different methylation patterns. Based on this panoramic view, soft Hamming distance similarity is used to classify the reads in the batch data. The cell type composition of the batch data is estimated using maximum a posteriori probability, combined with tree search to achieve optimal results, thus deconvolving the batch data.
[0103] Please see Figure 9 This application also provides a single-cell whole-genome analysis device based on sequencing reads for CpG methylation, which can implement the above-mentioned method. The device includes:
[0104] The site statistics module is used to acquire the single cell to be analyzed, perform DNA methylation data sequencing read alignment on the single cell to be analyzed, and obtain methylation site statistics results.
[0105] The methylation-specific interval calculation module is used to screen methylation-specific intervals based on random sampling strategies and Hamming distances of sequencing reads.
[0106] The clustering analysis module is used to calculate the Hamming distance matrix of CpG sites in each pair of single cells to be analyzed based on the statistical results of methylation sites and methylation-specific intervals, and to perform initial clustering based on the Hamming distance matrix to obtain the initial clustering results.
[0107] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0108] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0109] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0110] Please see Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0111] The processor 101 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0112] The memory 102 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 102 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 102 and is called and executed by the processor 101 using the methods described in the embodiments of this application.
[0113] Input / output interface 103 is used to implement information input and output;
[0114] The communication interface 104 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0115] Bus 105 transmits information between various components of the device (e.g., processor 101, memory 102, input / output interface 103, and communication interface 104);
[0116] The processor 101, memory 102, input / output interface 103 and communication interface 104 are connected to each other within the device via bus 105.
[0117] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0118] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0119] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0120] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0121] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0122] The methods, apparatus, electronic devices, and program products for single-cell whole-genome analysis of CpG methylation based on sequencing reads provided in this application embodiment obtain statistical results of methylation sites by comparing sequencing reads of methylation data of the single cells to be analyzed. Methylation-specific intervals are screened based on a random sampling strategy and the methylation Hamming distance of the sequencing reads. Hamming distance matrices of pairwise CpG sites of the single cells to be analyzed are calculated based on the statistical results of methylation sites and the methylation-specific intervals. Initial clustering is performed based on the Hamming distance matrix to obtain the initial clustering results. The methylation Hamming distance between different reads is calculated according to the interval division to screen methylation-specific intervals. This reduces the amount of data read into memory at one time and improves analysis efficiency. By selecting methylation-specific intervals, the feature dimensionality can be effectively reduced. The sensitivity can be improved by combining the measurement of Hamming distance, which can effectively improve the clustering efficiency and accuracy.
[0123] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0124] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0126] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0127] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0128] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0129] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0130] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0131] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0132] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0133] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for single-cell whole-genome analysis based on CpG methylation of sequencing reads, characterized in that, The method includes the following steps: Obtain single cells to be analyzed, and perform DNA methylation data sequencing read alignment on the single cells to obtain statistical results of methylation sites; Based on random sampling strategies and Hamming distance of sequencing reads, methylation-specific regions were screened, including: Based on a random sampling strategy, a pseudo-batch of methylation data is constructed, and a list of intervals is constructed based on the whole genome. The methylation Hamming distance set of each interval read in the interval list is calculated based on the pseudo-batch data. The methylation Hamming distance statistical features of each interval read are determined based on the methylation Hamming distance set of each interval read, and the methylation-specific intervals in the interval list are screened based on the methylation Hamming distance statistical features of each interval read. Based on the statistical results of the methylation sites and the methylation-specific regions, the Hamming distance matrix of each pair of CpG sites in the single cells to be analyzed is calculated. Initial clustering is then performed based on the Hamming distance matrix to obtain the initial clustering results, including: Extract CpG site information of each single cell to be analyzed in the methylation-specific region; calculate the Hamming distance matrix between each pair of single cells to be analyzed based on the CpG site information of each single cell to be analyzed; obtain the dendrogram of the initial clustering using the hierarchical clustering method of Euclidean distance based on the Hamming distance matrix, and select parameters to obtain the initial clustering results.
2. The method according to claim 1, characterized in that, The method further includes: For each single cell to be analyzed, a set of cells of the same type and a set of cells of different types are constructed based on the initial clustering results. Calculate the first average Hamming distance between each of the single cells to be analyzed and the set of cells of the same type, and calculate the second average Hamming distance between each of the single cells to be analyzed and the set of cells of different types; Calculate the cellular heterogeneity measure for each of the single cells to be analyzed based on the first average Hamming distance and the second average Hamming distance; Feature extraction and clustering were performed on all the single cells to be analyzed based on the cellular heterogeneity measures, resulting in methylated cell clustering results.
3. The method according to claim 2, characterized in that, The method further includes: Based on the cell heterogeneity measures of all the single cells to be analyzed and the cell clustering results, a set of specific intervals for all the single cells to be analyzed is selected. Heuristically search the specific region set based on methylation site patterns to screen out differentially methylated sites; Based on the differential methylation pattern sites, the methylation patterns of different cell types at all sites are calculated to obtain a panoramic view of specific methylation patterns.
4. The method according to claim 3, characterized in that, The method includes: Based on the panoramic view of the specific methylation pattern, effective read segments are screened. Based on the methylation patterns of the effective reads and the methylation patterns of each cell type, the similarity between each read and each cell type is calculated. Based on the similarity between each read segment and each cell type, and the maximum posterior probability, the proportion of all the cell components to be analyzed is estimated to obtain the estimated proportion of each cell type.
5. The method according to any one of claims 1-4, characterized in that, The step of performing DNA methylation data sequencing read alignment on the single cell to be analyzed to obtain methylation site statistical results includes: The paired-end sequencing data of the single cells to be analyzed were quality controlled and pruned, the tag sequence was removed, and the NDA reads were aligned with the sequencing library to obtain the initial alignment results. Low-quality alignment results in the initial alignment results are filtered out, sorted, and deduplicated to obtain statistical results of methylation sites.
6. A single-cell whole-genome analysis device based on CpG methylation of sequencing reads, characterized in that, The device includes: The site statistics module is used to acquire the single cell to be analyzed, perform DNA methylation data sequencing read alignment on the single cell to be analyzed, and obtain methylation site statistics results. The specific interval calculation module is used to screen methylation specific intervals based on a random sampling strategy and the methylation Hamming distance of sequencing reads; the specific interval calculation module is specifically used for: Based on a random sampling strategy, a pseudo-batch of methylation data is constructed, and a list of intervals is constructed based on the whole genome. The methylation Hamming distance set of each interval read in the interval list is calculated based on the pseudo-batch data. The methylation Hamming distance statistical features of each interval read are determined based on the methylation Hamming distance set of each interval read, and the methylation-specific intervals in the interval list are screened based on the methylation Hamming distance statistical features of each interval read. The clustering analysis module is used to calculate the Hamming distance matrix between each pair of CpG sites in the single cells to be analyzed based on the statistical results of the methylation sites and the methylation-specific regions, and to perform initial clustering based on the Hamming distance matrix to obtain the initial clustering results; the clustering analysis module is specifically used for: Extract CpG site information of each single cell to be analyzed in the methylation-specific region; calculate the Hamming distance matrix between each pair of single cells to be analyzed based on the CpG site information of each single cell to be analyzed; obtain the dendrogram of the initial clustering using the hierarchical clustering method of Euclidean distance based on the Hamming distance matrix, and select parameters to obtain the initial clustering results.
7. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of claims 1-5.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Single cell methylation data clustering method based on multi-distance spectrum embedding fusion
CN114298201A
Method capable of improving methylation detection accuracy of single-molecule 5mC based on CCS
CN119479824A