A method and system for detecting cancer biomarkers based on BS-seq data methylation vectors

By generating methylation vector clusters through window cutting and clustering algorithms based on BS-seq data, the problem of failing to fully utilize the information of adjacent CpG sites in existing technologies is solved, thereby improving the accuracy and robustness of cancer biomarker detection at low sequencing depth.

CN119418770BActive Publication Date: 2025-11-14ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411316653.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-11-14
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

Existing BS-seq data methylation analysis only considers the methylation status of a single CpG site, failing to fully utilize the information from adjacent CpG sites, and has high requirements for sequencing depth and tumor purity, affecting detection accuracy.

Method used

By cutting BS-seq data into windows containing 4 CpGs, methylation vectors are generated. Clustering algorithms are used to form methylation vector clusters, and vector clusters originating only from cancer groups are selected as biomarkers, reducing sequencing depth requirements and minimizing the impact of tumor purity.

Benefits of technology

This technology enables the use of adjacent CpG site information at lower sequencing depths, improving detection accuracy, reducing sensitivity to tumor purity, and enhancing the accuracy of cancer biomarker detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418770B_ABST
    Figure CN119418770B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for detecting cancer biomarkers based on BS-seq data methylation vectors, comprising: Step S1, acquiring multiple test samples from a cancer group and a test group, and segmenting the BS-seq data of each test sample into multiple fragments using multiple windows; Step S2, for each fragment, representing methylated CpG sites within the fragment with 1 and unmethylated CpG sites with 0, to obtain a methylation vector; Step S3, clustering all methylation vectors in each window to form multiple methylation vector clusters; Step S4, for each methylation vector cluster, determining whether all methylation vectors in the cluster originate from the cancer group: if so, using each methylation vector as a biomarker for cancer detection; otherwise, exiting the process. The beneficial effect is that this invention can fully utilize the methylation status information of adjacent CpG sites, reduce the sequencing depth requirement for a single test sample, and reduce the impact of tumor purity on cancer tumor samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of biomarker detection, and more specifically, to a method and system for detecting cancer biomarkers based on BS-seq data methylation vectors. Background Technology

[0002] Clinically, there has always been a huge demand for non-invasive detection, risk stratification, and monitoring of treatment progress in tumors. In this regard, circulating tumor cells, extracellular vesicles, RNA, proteins, and metabolites have made significant progress in recent years as potential non-invasive biomarkers. However, cell-free DNA (cfDNA) has attracted the most widespread attention. In a typical biological plasma sample, the total cfDNA population represents a very rich source of biological and pathological information and has shown great potential as a multifunctional biomarker in oncology, non-invasive prenatal testing, and transplant monitoring.

[0003] Currently, the main methylation analysis metrics for BS-seq data rely on the methylation rate of CpG sites. CpG sites are consecutive cytosine (C) and guanine (G) sequences on DNA. DNA methylation detection typically involves detecting the methylation status of cytosine in CpGs. However, the methylation rate only considers the methylation status of a single CpG site and cannot utilize the methylation information of adjacent CpG sites. Metrics such as methylation entropy and methylation haplotype can take into account the methylation status of adjacent CpG sites. Fragments are DNA fragments used for library construction during the BS-seq experiment. These scalar metrics treat all fragments in each sample uniformly, requiring each sample to have a high sequencing depth. Furthermore, these methylation scalar metrics are significantly affected by the tumor purity of cancer samples, which can impact detection accuracy. Summary of the Invention

[0004] The technical problem to be solved by this invention is to make full use of the methylation status information of adjacent CpG sites, reduce the sequencing depth requirements for a single test sample, and reduce the impact of tumor purity on cancer tumor samples. In order to overcome the above-mentioned defects of the prior art (or related technologies), this invention provides a method and system for detecting cancer biomarkers based on BS-seq data methylation vectors.

[0005] This invention provides a method for detecting cancer biomarkers based on BS-seq data methylation vectors, comprising the following steps:

[0006] Step S1: Obtain multiple test samples from the cancer group and the test group, and cut the BS-seq data of each test sample into multiple fragments using multiple windows containing 4 CpGs.

[0007] Step S2: For each fragment, methylated CpG sites within the fragment are represented by 1, and unmethylated CpG sites are represented by 0, to obtain the corresponding methylation vector;

[0008] Step S3: Cluster all the methylation vectors in each window to form multiple methylation vector clusters;

[0009] Step S4: For each methylation vector cluster, determine whether all methylation vectors in the methylation vector cluster originate from the cancer group.

[0010] If so, each of the methylation vectors in the methylation vector cluster will be used as a biomarker for cancer detection;

[0011] If not, then exit.

[0012] This application presents a cancer biomarker detection method based on BS-seq data methylation vectors, which has the following advantages compared with existing technologies:

[0013] In this application, methylated CpG sites within each fragment are represented by 1, and unmethylated CpG sites are represented by 0, making full use of the methylation status information of adjacent CpG sites. Furthermore, this application uses multiple windows containing four CpGs to cut BS-seq data into multiple fragments, and then performs methylation vector generation on each fragment. This eliminates the need for high sequencing depth for each test sample, reducing the sequencing depth requirement for individual test samples. In addition, this application processes test samples from the cancer group and the test group uniformly, obtaining biomarkers for cancer detection through fragment processing, vector generation, clustering operations, and other steps, reducing the impact of tumor purity on cancer tumor samples.

[0014] In one possible implementation, in step S1, for each subject, 10 ml of blood is drawn from the subject, serum is separated, cfDNA is extracted using a cfDNA purification kit, and BS-seq library construction is performed using a BS-seq conversion kit to obtain the corresponding test sample.

[0015] In one possible implementation, the clustering method used in step S3 is:

[0016] All methylation vectors in each window are merged to form a methylation matrix, and the methylation vectors that are close to each other in each window are clustered to obtain multiple methylation vector clusters.

[0017] In one possible implementation, step S3 further includes:

[0018] During the clustering process, the source of each methylation vector in each window is recorded.

[0019] In one possible implementation, in step S3, the MRESC clustering algorithm is used to cluster all the methylation vectors in each window.

[0020] Another technical solution of the present invention is to provide a cancer biomarker detection system based on BS-seq data methylation vectors, which applies the above-mentioned cancer biomarker detection method and includes:

[0021] A sample acquisition module is used to acquire multiple test samples from the cancer group and the test group, and to cut the BS-seq data of each test sample into multiple fragments using multiple windows containing 4 CpGs.

[0022] A fragment processing module, connected to the sample acquisition module, is used to characterize methylated CpG sites in the fragment with 1 and unmethylated CpG sites with 0 for each fragment, thereby obtaining the corresponding methylation vector;

[0023] A vector processing module, connected to the fragment processing module, clusters all methylation vectors in each window to form multiple methylation vector clusters;

[0024] A biomarker analysis module, connected to the vector processing module, is used to, for each methylation vector cluster, when all methylation vectors in the methylation vector cluster originate from the cancer group, use each methylation vector in the methylation vector cluster as a biomarker for cancer detection; or

[0025] If any of the methylation vectors in the methylation vector cluster does not originate from the cancer group, exit the process.

[0026] In one possible implementation, the vector processing module includes a clustering unit for merging all the methylation vectors in each window to form a methylation matrix, and clustering the methylation vectors that are close to each other in each window into multiple methylation vector clusters.

[0027] In one possible implementation, the vector processing module further includes a recording unit connected to the clustering unit, for recording the source of each methylation vector in each window during the clustering process of the clustering unit. Attached Figure Description

[0028] Figure 1 This is a flowchart of the method steps of the present invention;

[0029] Figure 2 This is a schematic diagram of the system structure of the present invention;

[0030] Figure 3 This is a flowchart illustrating the actual operation of the present invention;

[0031] Explanation of reference numerals in the attached diagram: 1. Sample acquisition module; 2. Fragment processing module; 3. Vector processing module; 31. Clustering unit; 32. Recording unit; 4. Biomarker analysis module. Detailed Implementation

[0032] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of this application and are not intended to limit the scope of protection of the embodiments of this application. Those skilled in the art can make adjustments as needed to adapt to specific application scenarios.

[0033] The present application will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0034] See Figure 1 This application discloses a method for detecting cancer biomarkers based on BS-seq data methylation vectors, including steps S1: sample acquisition and segmentation, step S2: methylation vector generation, step S3: clustering processing, and step S4: biomarker screening. Specifically, step S1 involves acquiring multiple test samples from the cancer group and the test group, and segmenting the BS-seq data of each test sample into multiple fragments using multiple windows containing four CpGs. Step S2 involves representing each fragment with methylated CpG sites as 1 and unmethylated CpG sites as 0, obtaining the corresponding methylation vector. Step S3 involves clustering all methylation vectors in each window to form multiple methylation vector clusters. Step S4 involves determining whether all methylation vectors in each methylation vector cluster originate from the cancer group; if so, each methylation vector in the cluster is used as a biomarker for cancer detection; otherwise, the process is terminated.

[0035] See also Figure 1 This application proposes for the first time a vector index for measuring methylation status: methylation vector. Compared with existing methylation indices, it has the following advantages: 1. It can make full use of the methylation status information of adjacent CpG sites; 2. It has lower requirements for the sequencing depth of a single test sample; 3. It is not affected by the tumor purity of cancer tumor samples.

[0036] See also Figure 1 This application mainly includes three stages in its actual operation:

[0037] The first stage is vectorization: the BS-seq data of each test sample is divided into windows containing four CpGs. For the portion of each fragment within the window, methylated CpG sites are represented by 1, and unmethylated CpG sites are represented by 0. This method allows each fragment within the window to be represented by a vector, which we call the methylation vector. Multiple methylation vectors within a window can form a methylation matrix. Figure 3 The “vectorization” part in the text refers to the 0-1 matrix within the box.

[0038] The second stage, clustering, involves merging the methylation vectors from each window of all test samples in both the cancer group and the test group. For example, if the methylation vectors for a given window in each test sample are v1, v2, v3, ..., vn, these vectors are merged into a single matrix V. merge = [v1,v2,v3,...vn], cluster all methylation vectors in each window, and after clustering, the methylation vectors that are close to each other in each window are grouped into clusters;

[0039] The third stage is separation: During the second stage of merging and clustering, the source of all methylation vectors in each window is recorded. By traversing all windows, if all methylation vectors in a certain methylation vector cluster in a window originate from the cancer group, then this methylation vector cluster can be considered as a cancer-related cluster, and the methylation vectors in this methylation vector cluster can be used as biomarkers for cancer detection.

[0040] See also Figure 1 In the third stage, when performing merge clustering, existing clustering algorithms such as K-means, DBSCAN, or HDBSCAN can be selected. In this application, the MRESC multiple repeats and equal spacing clustering clustering algorithm, which is specifically designed for methylation vector analysis, is used.

[0041] See Figure 2This application discloses a cancer biomarker detection system based on BS-seq data methylation vectors, including a sample acquisition module 1, a fragment processing module 2, a vector processing module 3, and a biomarker analysis module 4. The sample acquisition module 1 acquires multiple test samples from a cancer group and a test group, and segments the BS-seq data of each test sample into multiple fragments using multiple windows containing four CpGs. The fragment processing module 2, for each fragment, represents methylated CpG sites with 1 and non-methylated CpG sites with 0, obtaining the corresponding methylation vector. The vector processing module 3 clusters all methylation vectors in each window to form multiple methylation vector clusters. The biomarker analysis module 4, for each methylation vector cluster, if all methylation vectors in the cluster originate from the cancer group, uses each methylation vector in the cluster as a biomarker for cancer detection; or, if any methylation vector in the cluster does not originate from the cancer group, it exits the process.

[0042] See also Figure 2 The vector processing module 3 includes a clustering unit 31, which is used to merge all methylation vectors in each window to form a methylation matrix, and to cluster the methylation vectors that are close to each other in each window into multiple methylation vector clusters.

[0043] See also Figure 2 The vector processing module 3 also includes a recording unit 32 connected to the clustering unit 31, used to record the source of each methylation vector in each window during the clustering process of the clustering unit.

[0044] In the description of this application, the references to terms such as "an embodiment," "some embodiments," "in this embodiment," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0045] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for detecting cancer biomarkers based on BS-seq data methylation vectors, characterized in that, Includes the following steps: Step S1: Obtain multiple test samples from the cancer group and the test group, and cut the BS-seq data of each test sample into multiple fragments using multiple windows containing 4 CpGs. Step S2: For each fragment, methylated CpG sites within the fragment are represented by 1, and unmethylated CpG sites are represented by 0, to obtain the corresponding methylation vector; Step S3: Cluster all the methylation vectors in each window to form multiple methylation vector clusters; Step S4: For each methylation vector cluster, determine whether all methylation vectors in the methylation vector cluster originate from the cancer group. If so, each of the methylation vectors in the methylation vector cluster will be used as a biomarker for cancer detection; If not, then exit.

2. The method for detecting cancer biomarkers according to claim 1, characterized in that, In step S1, for each subject, 10ml of blood is drawn from the subject, serum is separated, cfDNA is extracted using a cfDNA purification kit, and BS-seq library construction is performed using a BS-seq conversion kit to obtain the corresponding test sample.

3. The method for detecting cancer biomarkers according to claim 1, characterized in that, In step S3, the clustering method used is as follows: All methylation vectors in each window are merged to form a methylation matrix, and the methylation vectors that are close to each other in each window are clustered to obtain multiple methylation vector clusters.

4. The method for detecting cancer biomarkers according to claim 1, characterized in that, Step S3 further includes: During the clustering process, the source of each methylation vector in each window is recorded.

5. A cancer biomarker detection system based on BS-seq data methylation vectors, characterized in that, The method for detecting cancer biomarkers as described in any one of claims 1-4 includes: A sample acquisition module (1) is used to acquire multiple test samples from the cancer group and the test group, and to cut the BS-seq data of each test sample into multiple fragments using multiple windows containing 4 CpGs. A fragment processing module (2) is connected to the sample acquisition module (1) and is used to characterize methylated CpG sites in the fragment with 1 and unmethylated CpG sites with 0 for each fragment, so as to obtain the corresponding methylation vector. A vector processing module (3) is connected to the fragment processing module (2) to cluster all the methylation vectors in each window to form multiple methylation vector clusters; A biomarker analysis module (4), connected to the vector processing module (3), is used to, for each methylation vector cluster, when all the methylation vectors in the methylation vector cluster originate from the cancer group, use each methylation vector in the methylation vector cluster as a biomarker for cancer detection; or If any of the methylation vectors in the methylation vector cluster does not originate from the cancer group, exit the process.

6. The cancer biomarker detection system according to claim 5, characterized in that, The vector processing module (3) includes a clustering unit (31) for merging all the methylation vectors in each window to form a methylation matrix, and clustering the methylation vectors that are close to each other in each window into a cluster to obtain multiple methylation vector clusters.

7. The cancer biomarker detection system according to claim 6, characterized in that, The vector processing module (3) further includes a recording unit (32) connected to the clustering unit (31) for recording the source of each methylation vector in each window during the clustering process of the clustering unit.

Citation Information

Patent Citations

  • Method and system for screening pan cancer markers based on cfDNA methylation as well as combination and application of markers

    CN118064583A

  • Method, system, equipment and medium for screening marker for diagnosing cancer based on methylated cfDNA fragment

    CN118280447A