Separation or purification method of peripheral red blood cell micronucleus DNA and application thereof

CN120693409APending Publication Date: 2025-09-23TIMING BIOTECH TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480009552.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-15
Filing Date
2024-12-31
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The prior art has not yet effectively isolated and purified micronuclear DNA in peripheral blood red blood cells, and there is no method for cancer detection using it.

Method used

A method for isolating and purifying micronuclear DNA from peripheral blood red blood cells, including centrifugation, filtration and Benzonase nuclease treatment, combined with whole genome sequencing and analysis of specific gene regions, is provided for quality control and cancer detection.

Benefits of technology

It has achieved highly sensitive and specific cancer detection, especially accurate identification of thyroid cancer, breast cancer, colorectal cancer, gastric cancer and lung cancer, reducing the contamination of free DNA of monocytes and circulating cells and simplifying the operation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120693409A_ABST
    Figure CN120693409A_ABST
Patent Text Reader

Abstract

Relates to the fields of biology, medicine and bioinformatics, in particular to a peripheral red blood cell micronucleus DNA separation or purification method and application thereof.
Need to check novelty before this filing date? Find Prior Art

Description

Method for separating or purifying micronuclear DNA from peripheral blood erythrocytes and its application Technical Field

[0001] The present invention relates to the fields of biology, medicine and bioinformatics, and in particular to a method for separating or purifying micronuclear DNA of peripheral blood erythrocytes and an application thereof. Background Art

[0002] Cancer is one of the most significant threats to human health and life. In 2018, there were reportedly 18.1 million new cancer cases and 9.6 million cancer deaths worldwide, with nearly half of these cases and over half of these deaths occurring in Asia. Despite decades of progress in cancer diagnosis and treatment, significant unmet medical needs remain for cancer detection, particularly screening, diagnosis, classification, and staging.

[0003] Blood circulates continuously throughout the body. In a normal adult, the total blood volume accounts for approximately 8% of body weight in men and 7.5% in women. Peripheral blood samples are easy to collect, store, and transport, and are highly stable.

[0004] Micronuclei (MNs) are generally considered to be small nuclear structures formed when chromosomes or chromosome fragments are not incorporated into one of the daughter nuclei during cell division. They are often a sign of genotoxic events and chromosomal instability. They are usually caused by incorrect repair or unrepaired DNA breaks, or chromosome nondisjunction, resulting in lagging asymmetric chromosomes or chromatid fragments, forming small nuclear structures independent of the main nucleus.

[0005] To date, there are no reports on how to perform quality control on micronuclear DNA (MNeDNA) isolated or purified from peripheral blood red blood cells, and there are also few reports on using peripheral blood red blood cell micronuclear DNA for cancer detection. Summary of the Invention

[0006] The present invention provides a method for separating or purifying micronuclear DNA from peripheral blood erythrocytes and its application, and in particular, relates to a quality control method thereof and its application in screening, diagnosis, typing and / or staging of diseases.

[0007] According to a first aspect of the present invention, a method for separating or purifying micronuclear DNA from peripheral blood red blood cells is provided, comprising the following steps: a) providing a peripheral blood sample; b) separating mononuclear cells and red blood cells in the peripheral blood sample; c) collecting the red blood cells; d) lysing the red blood cells; and e) extracting micronuclear DNA.

[0008] Preferably, the step b) comprises: centrifuging the peripheral blood sample to obtain a mononuclear cell layer and a red blood cell layer; and filtering the red blood cell layer to obtain a filtrate containing red blood cells.

[0009] Preferably, the step b) further comprises: treating the filtrate with Benzonase nuclease.

[0010] Preferably, the peripheral blood sample is centrifuged by density gradient.

[0011] Preferably, the red blood cell layer is filtered using a 10 micron cell mesh.

[0012] According to a second aspect of the present invention, provided is the use of at least one of the ARHGAP25, PIK3R1, FLI1, PTPRC, MBNL1, TANK, THEMIS, FYB1, FYN, HIVEP2, PRKACB, INPP4B, PIP4K2A, SATB1, ZFC3H1, UBAC2, and KLF12 genes of peripheral blood erythrocyte micronuclear DNA in a system for controlling the quality of micronuclear DNA.

[0013] According to a third aspect of the present invention, there is provided a use of at least one region of peripheral blood erythrocyte micronuclear DNA as shown in Table 1 in a system for controlling the quality of micronuclear DNA.

[0014] According to the fourth aspect of the present invention, a method for determining a quality control benchmark for peripheral blood red blood cell micronuclear DNA is provided, comprising: a) providing a group of subjects with common characteristics; b) isolating or purifying peripheral blood red blood cell micronuclear DNA from the peripheral blood red blood cells of each subject; c) performing whole genome sequencing on the peripheral blood red blood cell micronuclear DNA to obtain fragment sequence information of the micronuclear DNA; d) comparing the fragment sequence information of the peripheral blood red blood cell micronuclear DNA of the subject with the fragment sequence information of the peripheral blood mononuclear cell genomic DNA of the subject to obtain the gene-enriched region of the peripheral blood red blood cell micronuclear DNA of the subject; e) determining a candidate interval in which the read abundance of the peripheral blood red blood cell micronuclear DNA of the subject is greater than the read abundance of the peripheral blood mononuclear cell genomic DNA of the subject; f) determining the quality control benchmark based on the first read abundance of mitochondrial DNA and / or the second read abundance of the gene-enriched region in the candidate interval.

[0015] Preferably, in step e), based on the median read count of the peripheral blood red blood cell micronuclear DNA of the subject in the 1Mb interval of the whole genome, the 1Mb interval in which the read abundance of the peripheral blood red blood cell micronuclear DNA of the subject is greater than 1.2 times that of the peripheral blood mononuclear cell genomic DNA of the subject is determined as the candidate interval.

[0016] Preferably, the group of subjects having a common characteristic comprises: non-cancer subjects.

[0017] Preferably, the length of the gene-enriched region in the candidate interval is between 179 bp and 32.04 kb.

[0018] Preferably, the gene-enriched region in the candidate interval includes the region shown in Table 1 or the region corresponding to the ARHGAP25, PIK3R1, FLI1, PTPRC, MBNL1, TANK, THEMIS, FYB1, FYN, HIVEP2, PRKACB, INPP4B, PIP4K2A, SATB1, ZFC3H1, UBAC2, and KLF12 genes.

[0019] Preferably, the quality control benchmark is determined based on the first read abundance of mitochondrial DNA and / or the second read abundance of the gene-enriched region in the candidate interval, including: a) providing more than one group of peripheral blood red blood cell micronuclear DNA containing peripheral blood mononuclear cell genomic DNA in different proportions; b) performing whole genome sequencing on the group of peripheral blood red blood cell micronuclear DNA to determine the first read abundance of mitochondrial DNA and / or the second read abundance of the gene-enriched region in the candidate interval; c) determining the quality control benchmark based on the impact of the ratio of the different proportions of peripheral blood mononuclear cell genomic DNA on the first read abundance of mitochondrial DNA and / or on the second read abundance of the gene-enriched region in the candidate interval.

[0020] According to a fifth aspect of the present invention, a method for quality control of a peripheral blood red blood cell micronuclear DNA extraction process is provided, comprising: comparing the first read abundance of mitochondrial DNA in the peripheral blood red blood cell micronuclear DNA and / or the second read abundance of at least one of the aforementioned regions or the corresponding region of at least one gene among ARHGAP25, PIK3R1, FLI1, PTPRC, MBNL1, TANK, THEMIS, FYB1, FYN, HIVEP2, PRKACB, INPP4B, PIP4K2A, SATB1, ZFC3H1, UBAC2, and KLF12 genes with the quality control benchmark determined above.

[0021] According to a sixth aspect of the present invention, there is provided a peripheral blood red blood cell micronuclear DNA of SARS1, MYBPHL, SORT1, PSMA5, SYPL2, CELSR2, PSRC1, POGK, TADA1, MAEL, GPA33, STYXL2, ILDR2, POU2F1, POU2F1-DT, C1orf105, SUCO, RGS16, RGSL1, RNASEL, NPL, DHX9, RGS8, ZNF281, KIF14, CAMSAP2, DDX59, DUSP10, EGLN1, OSBPL9, RAB3B, NRDC, RERE, ERRFI1-DT, SLC45A1 ,ERRFI1,PTBP2,CFAP43,GSTO1,SFR1,NRG3,PANK1,KIF20B,USP47,MICAL2,DKK3,DISC1FP1,SCAF11,TEX30,POGLUT2,BIVM-ERCC5,BIVM,ERCC5,METTL 21EP,TPP2,METTL21C,CCDC168,RNY3P9,RB1,CYSLTR2,FNDC3,PCNX4,DHRS7,ATP10A,RAB27A,DAPK2,CIAO2A,SNX1,RBFOX1,SMARCE1,KRT222,KRT20,K RT23,KRT39,KRT40,KRTAP3-3,KRT24,KRT25,KRT26,KRT27,KRT28,KRT10,KRT12,KRTAP2-3,KRTAP2-4,KRTAP3-1,KRTAP1-5,KRTAP1-4,KRTAP3-2,KRT AP1-1,KRTAP2-1,KRTAP2-2,KRTAP1-3,CDC27,MYL4,ITGB3,EFCAB13,RNU7-186P,EFCAB13-DT,FAP,GCG,IFIH1,GCA,SPC25,KCTD18,SGO2,SPATS2L,RN Y4P34,NBEAL1,ICA1L,WDR12,CARF,IKZF2,HTR2B,GPR55,SPATA3,C2orf72,PSMD1,GKN2,BMP10,SIAH2,KCNAB1,SSR3,TIPARP,LEKR1,RN7SKP177,SCN1 1A,WDR48,RPSA,SNORA6,MOBP,CSRNP1,GORASP1,TTC21A,CCR8,SLC25A38,XIRP1,CX3CR1,CCR5AS,CCR5,LTF,CCR2,CCRL2,EPHA6,ARL6,CRYBG3,BANK1,PPP3CA,FLJ20021,RN7SL446P,FBXW7,KCNIP4,GBA3,NRG2,PURA,IGIP,CYSTM1,CDH12,TTC33,ESM1,CERT1,P OLK,HMGCR,RNU7-175P,ANKDD1B,WASF1,CDC40,CALHM4,TRAPPC3L,RWDD1,RSPH4A,ZUP1,AHI1,AHI1-DT,ZDHH C14,TMEM242-DT,TMEM242,PRKN,BRPF3,PNPLA1,BNIP5,ETV7,PXT1,KCTD20,C7orf33,CUL1,BMPER,PHF20L1, TG,DNAAF11,TMEM71,PTCSC1,CCN4,NDRG1,SLA,ESCO2,PBK,CCDC25,CDH17,FSBP,VIRMA,GEM,RAD54B,ESRP1, VIRMA-DT,PANK1-AS1,TUBA3C,LRFN5,GOLGA6L2,UBBP4,FAM27E5,FLJ36000,ZNF675,NLRP12,MYADM-AS1,MYA DM,PRKCG,CACNG7,RNU1-7P,DNM3,PIGC,FMN2,ACSS1,CST7,APMAP,C1QL2,RN7SL468P,ZEB2,GTDC1,GALNT13, Use of at least one gene selected from the group consisting of SPAG16, SPAG16-DT, ATG16L1, SCARNA6, SAG, DGKD, GRM7, DCAF16, NCAPG, LCORL, ARAP2, ANKRD31, IRF1, IL5, RAD50, TH2LCRR, IL13, GFOD1, TSBP1-AS1, TSBP1, BTNL2, HLA-DRA, and SDCBP in a system for performing pan-cancer detection on a test subject.

[0022] According to a seventh aspect of the present invention, there is provided use of at least one region of peripheral blood erythrocyte micronuclear DNA as shown in Table 3 in a system for performing pan-cancer detection on a test subject.

[0023] According to an eighth aspect of the present invention, there is provided PRNLS, LIPJ, LIPF, ANKRD22, LIPK, LIPN, LIPM, STAMBPL1, KIF20B, PANK1, MIR107, FDX1, RDX, ZC3H12C, TRHDE, TUBA3C, MDGA2, GOLGA6L2, GABRG3, GNAO1-DT, GNAO1, UBBP4, FAM27E5, FLJ36000, NLRP1 2,MYADM,PRKCG,CACNG7,PTBP2,NBPF6,EEIG2,FMN2,REL,PEX13,CTNNA2,GACAT1,RGPD4,SPAG16,SPAG16-DT,GRM7 ,CDV3,TOPBP1,TF,SRPRB,RAB6B,KCNAB1,SSR3,ANKRD31,HMGCR,CERT1,POLK,FAM174A-DT,FAM174A,RN7SKP62,ST8 SIA4,GFOD1,ELOVL4,TTK,UBE3D,DOP1A,PGM3,RWDD2A,ME1,AKAP7,ING3,CPED1,C7orf33,CUL1,SGCZ,SDCBP,NSMA F,TOX,OSGIN2,NBN,CSMD3,DPYD,GSTO1,SFR1,EBLN1,SHANK2,RB1,CYSLTR2,FNDC3A,RCBTB2,SMARCE1,KRT222,KRT Use of at least one of the following genes: 24, KRT25, KRT26, KRT27, KRT28, KRT10, KRT12, KRT20, KRT23, TIMP3, FBXO30, EPM2A-DT, ABCC10, AP1G1, NFKBIA, JAKMIP2, PRKN, EXOC4, CFAP43, SYN3, IKZF2, NELL1, SHPRH, PUS10, PTPRK, PKN2 in a system for detecting colorectal cancer in a test subject.

[0024] According to a ninth aspect of the present invention, there is provided the use of at least one region of peripheral blood erythrocyte micronuclear DNA as shown in Table 4 in a system for detecting colorectal cancer in a test subject.

[0025] According to the tenth aspect of the present invention, provided is the use of at least one of the PRKN, DCDC1, TUBA3C, ZNF675, C1QL2, RN7SL468P, RNU6-725P, LINC01208, LINC00408, LINC00442, LINC02479, LINC03025, and MIR7977 genes of peripheral blood erythrocyte micronuclear DNA in a system for detecting gastric cancer in a test subject.

[0026] According to an eleventh aspect of the present invention, there is provided use of at least one region of peripheral blood erythrocyte micronuclear DNA as shown in Table 5 in a system for detecting gastric cancer in a test subject.

[0027] According to the twelfth aspect of the present invention, a method for constructing a classifier for cancer detection using peripheral blood red blood cell micronuclear DNA is provided, comprising: a) providing a control group sample and more than one different category, wherein each category represents a group of subjects having common characteristics; b) isolating or purifying peripheral blood red blood cell micronuclear DNA from the control group sample and the peripheral blood red blood cells of each subject in each category; c) performing whole genome sequencing on the peripheral blood red blood cell micronuclear DNA to obtain fragment sequence information of the micronuclear DNA; d) comparing the fragment sequence information of the peripheral blood red blood cell micronuclear DNA of the control group sample and subjects of different categories; e) selecting regions with statistically different read abundance as feature regions based on the differential distribution of the fragment sequence information of the micronuclear DNA in the peripheral blood red blood cells of subjects of different categories and the control group sample, and obtaining a classifier for cancer detection by training the feature regions.

[0028] Preferably, the different categories include cancer subjects with different cancers.

[0029] Preferably, the different categories include: thyroid cancer subjects, intestinal cancer subjects, gastric cancer subjects, lung cancer subjects and breast cancer subjects.

[0030] Preferably, a classifier for cancer detection is obtained by training the feature regions, comprising: performing XGBoost model training on the feature regions of different categories of subjects and control group samples, performing 5-fold cross-validation on each of the feature regions, and using the first feature subset of the obtained optimal model for a classifier for cancer detection, wherein the cancer is pan-cancer.

[0031] Preferably, a classifier for cancer detection is obtained by training the feature regions, comprising: performing XGBoost model training on the feature regions of the subjects with intestinal cancer and the control group samples, and performing 5-fold cross-validation on each of the feature regions, and using the second feature subset of the obtained optimal model for a classifier for cancer detection, wherein the cancer is intestinal cancer.

[0032] Preferably, a classifier for cancer detection is obtained by training the feature regions, comprising: performing XGBoost model training on the feature regions of the gastric cancer subjects and the control group samples, and performing 5-fold cross-validation on each of the feature regions, and using the third feature subset of the obtained optimal model for the cancer detection classifier, wherein the cancer is gastric cancer.

[0033] According to the thirteenth aspect of the present invention, a system for detecting cancer in a test subject is provided, comprising: a detection device for detecting micronuclear DNA in peripheral blood red blood cells from the test subject using the classifier constructed according to the aforementioned method.

[0034] Preferably, the method further comprises: a separation device for separating peripheral blood erythrocyte micronuclear DNA from the test subject; and a sequencing device for sequencing the peripheral blood erythrocyte micronuclear DNA from the test subject.

[0035] Preferably, the system performs cancer detection by the following method, which includes: a) isolating or purifying micronuclear DNA in the peripheral blood red blood cells of the test subject; b) performing whole genome sequencing on the micronuclear DNA to obtain fragment sequence information of the micronuclear DNA in the peripheral blood red blood cells of the test subject; c) detecting the fragment sequence information of the micronuclear DNA obtained in step b) by the classifier constructed according to the above, thereby classifying the test subject into one or more of the more than one different categories.

[0036] Preferably, the cancer detection includes cancer screening, diagnosis, typing and / or staging.

[0037] The present invention has achieved excellent technical effects in at least the following aspects.

[0038] Rich sample sources

[0039] The present invention uses peripheral blood as a sample source, and the material source is abundant and easy to obtain, easy to collect, store and transport, and has high stability.

[0040] Effectively isolate micronuclear DNA from red blood cells

[0041] The method of the present invention can effectively separate micronuclear DNA from red blood cells from human peripheral blood. This separation method can effectively eliminate contamination from monocytes and cfDNA.

[0042] Easy and fast operation

[0043] The present invention only requires collecting a small amount of peripheral blood from the subject, thereby alleviating the psychological pressure of the subject.

[0044] High sensitivity and specificity for cancer detection

[0045] The method correctly identified 79% (95% confidence interval: 69-88%) of cancer patients across various cancer types, including thyroid cancer (TC), breast cancer (BC), colorectal cancer (CRC), gastric cancer (GC), and lung cancer (LC), with an overall specificity of 95.7%. This remarkable detection rate included 78% of stage I, 83% of stage II, and 76% of stage III cancer patients. Notably, this predictive ability was not present in a random classifier, emphasizing the specificity of the present invention.

[0046] Further analysis revealed high sensitivity and specificity when using the present invention to distinguish specific types of cancer from HD. For example, in a pairwise comparison, the model demonstrated an AUC of 96% for CRC and 92% for GC patients, respectively. The CRC model demonstrated a sensitivity of 83% for stage I, 100% for stage II, and 81% for stage III, with a specificity of 95.7%. Similarly, the GC-specific model demonstrated a sensitivity of 75% for both stage I and stage II patients, and a sensitivity of 100% for stage III patients, with a specificity of 91.5%. Overall, these data demonstrate that the present invention can be used to distinguish between cancer patients and HD patients.

[0047] The above content is summarized and therefore contains simplifications, generalizations and omissions of details when necessary; therefore, those skilled in the art will recognize that this summary is illustrative only and is not intended to be limiting in any way. Other aspects, features and advantages of the methods, compositions and / or devices and / or other subjects described herein will become apparent from the teachings shown herein. An overview is provided to simplify the introduction of some selected concepts, which will be further described in the detailed description below. This overview is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used as an auxiliary means for determining the scope of the claimed subject matter. In addition, the contents of all references, patents and published patent applications cited throughout this application are incorporated herein by reference in their entirety. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0049] FIG1 shows a flow chart for extracting the peripheral blood mononuclear cell genome and erythrocyte micronuclear DNA respectively;

[0050] FIG2 shows the results of atomic force microscopy (AFM) imaging of erythrocyte micronucleus DNA;

[0051] FIG3 shows the experimental results of eliminating cfDNA contamination by the separation method provided by the present invention;

[0052] FIG4 shows the experimental results of detecting gDNA contamination in MN DNA samples using digital PCR;

[0053] FIG5 shows the results of staining red blood cells at different separation steps using DAPI;

[0054] Figure 6 shows the distribution results of MN DNA reads;

[0055] FIG7 shows the read density of representative MN DNA and corresponding gDNA in a 1 Mb region of the whole genome;

[0056] FIG8 shows the Spearman correlation of MN DNA or gDNA read density and the median MN DNA read density of 10 healthy individuals;

[0057] Figure 9 shows the results of a 100 kb analysis of a MN DNA-enriched region on chromosome 1;

[0058] Figure 10 shows the results of the ratio of random regions and MN DNA-enriched regions annotated by ChromHMM;

[0059] FIG11 shows the Spearman correlation results of MN DNA read density and median MN DNA read density in cancer patients and healthy individuals;

[0060] FIG12 shows the sample distribution used in the process of constructing the classifier of the present invention;

[0061] FIG13 exemplifies a method for determining tumor-associated MN DNA;

[0062] FIG14 shows the distribution results of 288 selected taMN DNA features throughout the genome;

[0063] Figure 15 shows the regions corresponding to most taMN DNA features;

[0064] FIG16 shows the recognition results of various cancer types including thyroid cancer (TC), breast cancer (BC), colorectal cancer (CRC), gastric cancer (GC) and lung cancer (LC) according to the method provided by the present invention. DETAILED DESCRIPTION

[0065] Although the present invention can be implemented in many different forms, what is disclosed here is its specific illustrative embodiment that proves the principle of the present invention. It should be emphasized that the present invention is not limited to the specific embodiment illustrated. In addition, any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0066] Unless otherwise defined herein, scientific and technical terms used in conjunction with the present invention will have the meanings commonly understood by those of ordinary skill in the art. In addition, unless the context requires otherwise, terms in the singular shall include the plural, and terms in the plural shall include the singular. More specifically, as used in this specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" include plural referents. Thus, for example, reference to "a protein" includes multiple proteins; reference to "a cell" includes a mixture of cells, etc. In this application, unless otherwise stated, the use of "or" means "and / or". In addition, the use of the term "comprising" and other forms (such as "including" and "containing") is not restrictive. In addition, the ranges provided in the specification and the appended claims include all values ​​between the endpoints and breakpoints. The terms "first" and "second" do not impose additional restrictions on the nouns that follow them and do not affect the technical solutions.

[0067] Generally, terms relating to, and techniques of, cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein are those well known and commonly used in the art. Unless otherwise indicated, the methods and techniques of the present invention are generally performed according to conventional methods well known in the art and as described in various general and more specific references that are cited and discussed throughout this specification. See, for example, Abbas et al., Cellular and Molecular Immunology, 6th ed., WB Saunders Company (2010); Sambrook J. & Russell D. Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2000); Ausubel et al., Short Protocols in Molecular Biology: A Compendium of Methods from Current Protocols in Molecular Biology Biology, Wiley, John & Sons, Inc. (2002); Harlow and Lane Using Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (1998); Detecting cell-of-origin and cancer specific methylation features of cell-free DNA from Nanopore sequencing and Coligan et al., Short Protocols in Protein Science, Wiley, John & Sons, Inc. (2003). Additionally, any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0068] The following examples are provided to facilitate a clearer understanding of the present invention for those skilled in the art. It should be noted that the following examples do not limit the scope of the present invention and are provided for illustrative purposes only. Unless otherwise specified, the raw materials, reagents, and devices mentioned in the following examples are commercially available or obtained by known methods.

[0069] definition

[0070] For a better understanding of the present invention, definitions and explanations of relevant terms are provided below.

[0071] In the context of the present disclosure, "DNA" refers to deoxyribonucleic acid.

[0072] In the context of the present disclosure, "micronucleus" refers to a small nuclear structure containing DNA in addition to the nucleus in a specific cell. In peripheral blood red blood cells, there is no nucleus, so there are only micronucleus structures.

[0073] In the context of the present disclosure, "cfDNA" or "circulating cell-free DNA" refers to nucleic acid substances circulating in blood.

[0074] In the context of the present disclosure, "mtDNA" means "mitochondrial DNA".

[0075] In the context of the present disclosure, "subject" means an object to be tested. In certain embodiments, the "subject" is a human subject.

[0076] In the context of this disclosure, "patient" means a subject suffering from a certain disease.

[0077] In the context of this disclosure, "cancer" is a general term for malignant tumors. A tumor is a lesion formed by the abnormal proliferation of cells in local tissues under the influence of various tumorigenic factors.

[0078] In the context of the present disclosure, "cancer subject" or "cancer patient" are used interchangeably to refer to a subject suffering from a certain cancer.

[0079] In the context of the present disclosure, a "non-cancer subject" refers to a subject that does not have a certain type of cancer. For example, a "non-cervical cancer subject" refers to a subject that does not have cervical cancer. In the specific embodiments and examples of the present disclosure, a "non-cancer subject" is also referred to as a "healthy individual" or "HD," which similarly means that the individual or subject does not have that type of cancer.

[0080] In the context of the present disclosure, "cancer detection" means detecting the condition of a subject suffering from cancer. "Detection" includes but is not limited to screening, diagnosis, typing, staging, etc. Among them, "screening" means preliminary detection of whether a subject has cancer or is at risk of developing cancer. "Diagnosis" or "medical diagnosis" means making a judgment on the subject's condition from a medical perspective. "Typing" means further dividing the same type of cancer into specific subtypes. For example, cervical cancer can be classified into cervical squamous cell carcinoma and cervical adenocarcinoma. "Staging" means predicting, judging or dividing the stage of a certain cancer. For example, cervical cancer (squamous cell carcinoma) can be divided into stages such as poor differentiation, moderately poor differentiation, moderate differentiation, and well differentiation.

[0081] In the context of the present disclosure, "nucleated cells" refer to cells that have a nucleus. For peripheral blood, "nucleated cells" is a general term for granulocytes, monocytes, and lymphocytes.

[0082] In the context of this disclosure, "genome" means the sum of all genetic information in a cell, particularly a complete set of haploid genetic material in a cell.

[0083] In the context of the present disclosure, "nucleated cell genomic DNA," "nucleated cell nuclear genome," "gDNA," or "nucleated cell nuclear genomic DNA" are used interchangeably to refer to all genetic information contained in the chromosomes of the nucleus of a nucleated cell.

[0084] In the context of the present disclosure, "gene classifier" or "classifier" are used interchangeably to mean a group of DNA fragments or a group of genes in genomic DNA or micronuclear DNA that are specific for a particular disease.

[0085] In the context of the present disclosure, "DNA fragment library" or "DNA library" are used interchangeably, referring to double-stranded DNA obtained by end-filling, adding a phosphate group to the 5' end, adding an adenine nucleotide (A) to the 3' end of the sample DNA fragments, and then connecting adapters and sample tags (barcodes) at both ends.

[0086] In the context of the present disclosure, "reads" refer to the sequences of sample DNA fragments in a DNA fragment library (minus the sequences ligated during the library preparation stage) determined by sequencing.

[0087] In the context of the present disclosure, "coverage depth" refers to the effective nucleic acid sequencing fragments used for base identification in a specific region, also known as the number of reads or the number of read sequences.

[0088] In the context of the present disclosure, "sequence alignment" refers to aligning reads to a reference genome (eg, a human reference genome) by the principle of sequence identity.

[0089] In the context of the present disclosure, a "reference genome" is a complete genome sequence of the same organism as the sample DNA, available from a public database. In one embodiment, the reference genome is a human reference genome. The public database is not particularly limited. In certain embodiments, the public database is GenBank from NCBI.

[0090] In the context of the present disclosure, the definition and interpretation of receiver operating characteristic (ROC) curves and area under the curve (AUC) can be found in the specific references (Receiver-operating characteristic (ROC) plots: a fundamental evaluation tool in clinical medicine).

[0091] In the context of this disclosure, "sensitivity" refers to the percentage of samples that produce a positive test among patients compared to the total number of patients. In medical diagnosis, sensitivity can be expressed by the following formula, reflecting the proportion of patients correctly diagnosed:

[0092] Sensitivity = number of true positives / (number of true positives + number of false negatives) × 100%.

[0093] In short, if true positive, false positive, true negative and false negative are represented by a, b, c and d respectively, the relationship between sensitivity, specificity, missed diagnosis rate, misdiagnosis rate and accuracy can be shown as follows.

[0094] Among the cases with positive screening results using this method, true positive (a) indicates the number of cases with a pathological diagnosis of disease and a positive result of this method; false positive (b) indicates the number of cases with a pathological diagnosis of no disease and a positive result of this method; false negative (c) indicates the number of cases with a pathological diagnosis of disease and a negative result of this method; true negative (d) indicates the number of cases with a pathological diagnosis of no disease and a negative result of this method.

[0095] Sensitivity sen = a / (a+c);

[0096] Specificity sep = d / (b + d);

[0097] Missed diagnosis rate = c / (a+c);

[0098] Misdiagnosis rate = b / (b+d);

[0099] Accuracy = (a+d) / (a+b+c+d)

[0100] As known to those skilled in the art, the higher the sensitivity and specificity values, the better; the lower the missed diagnosis rate and misdiagnosis rate values, the better.

[0101] In the context of this disclosure, "specificity" refers to the percentage of healthy individuals who receive a negative test result. In medical diagnosis, specificity can be expressed as follows, reflecting the rate of correct diagnosis of non-patients:

[0102] Specificity = number of true negatives / (number of true negatives + number of false positives) × 100%.

[0103] In the context of this disclosure, the "missed diagnosis rate," also known as the false negative rate, refers to the percentage of subjects who are actually diagnosed as non-patients according to the diagnostic criteria when screening or diagnosing a disease in a population. In medical diagnosis, the missed diagnosis rate can be expressed by the following formula:

[0104] Missed diagnosis rate = number of false negatives / (number of true positives + number of false negatives) × 100%.

[0105] In the context of this disclosure, the "misdiagnosis rate," also known as the false positive rate, refers to the percentage of subjects who, when screening or diagnosing a disease in a population, are classified as patients according to the diagnostic criteria when they do not actually have the disease. In medical diagnosis, the misdiagnosis rate can be expressed by the following formula:

[0106] Misdiagnosis rate = number of false positives / (number of true negatives + number of false positives) × 100%.

[0107] In the context of the present disclosure, "about" means not more than plus or minus 10% of the specified value or range.

[0108] In this disclosure, "peripheral blood" refers to blood released into the circulatory system by hematopoietic organs and circulates therein. "Peripheral blood" is distinct from immature blood cells in hematopoietic organs (e.g., bone marrow). For purposes of this disclosure, peripheral blood can be collected by methods known in the art, such as venous, fingertip, or earlobe sampling.

[0109] Peripheral blood typically consists of plasma and blood cells, which further include white blood cells, red blood cells, and platelets. By volume, red blood cells make up approximately 45% of peripheral blood, plasma accounts for approximately 54.3%, and white blood cells account for approximately 0.7%. White blood cells are nucleated cells, encompassing granulocytes, monocytes, and lymphocytes. Normal red blood cells, on the other hand, lack a nucleus and genomic DNA, and are therefore considered anucleated.

[0110] In the context of the present disclosure, "peripheral blood mononuclear cell" (PBMC) means cells with a single nucleus in peripheral blood, including monocytes and lymphocytes.

[0111] Example 1

[0112] Density gradient centrifugation of peripheral blood

[0113] The peripheral blood samples of each subject were subjected to density gradient centrifugation according to the following steps.

[0114] 1. Obtain 1 ml of fresh peripheral blood from the subject and dilute it with an equal volume of 1× PBS to obtain a diluted blood sample.

[0115] 2. Add 5 ml of Ficoll density gradient centrifugation medium (Stemcell, LymphoprepTM07801) to the density gradient centrifugation tube.

[0116] 3. Slowly add the diluted blood sample prepared in step 1 to the density gradient centrifugation tube in step 2, and centrifuge at 1200 g and 18°C ​​for 15 minutes to perform density gradient centrifugation.

[0117] After density gradient centrifugation, the liquid is separated into three layers: the upper layer is plasma, the middle layer is peripheral blood mononuclear cells (PBMC), and the bottom layer is red blood cells.

[0118] Example 2

[0119] Separation of blood cells

[0120] After density centrifugation in Example 1, peripheral blood mononuclear cells and erythrocytes were separated.

[0121] Specifically, the middle and upper layers of liquid in the density gradient centrifuge tube were aspirated with a pipette to separate and collect the peripheral blood mononuclear cell sample; the bottom red blood cells were extracted from the bottom of the density gradient centrifuge tube using a syringe and transferred to a 1.5 ml centrifuge tube, 1× PBS was added to 1 ml, and the tube was centrifuged at 300 g for 10 minutes at room temperature to collect the red blood cells at the bottom of the tube.

[0122] An additional red blood cell purification step is performed by filtration through a 10-micron cell mesh to obtain a filtrate containing red blood cells, which is then treated with Benzonase nuclease to obtain purified red blood cells.

[0123] Example 3

[0124] DNA extraction

[0125] In this example, the peripheral blood mononuclear cell genome and erythrocyte micronuclear DNA were extracted separately. The process is shown in Figure 1.

[0126] 1. Extraction of Genomic DNA from Peripheral Blood Mononuclear Cells

[0127] Genomic DNA was extracted from the peripheral blood mononuclear cell samples obtained in Example 2 using QIAamp DNA Blood Mini Kit (Qiagen, Cat No. / ID: 51106).

[0128] 2. Extraction of Erythrocyte Micronuclear DNA

[0129] The erythrocytes obtained in Example 2 were lysed. Specifically, the erythrocytes collected in Example 2 were lysed at room temperature in the dark for 20 minutes. Thereafter, the supernatant was centrifuged at 3000g for 10 minutes at room temperature, and the supernatant was taken and incubated at 56°C for 8 hours using 10mM EDTA (Solarbio Cat No. / ID:E1170) and 200ug / ul proteinasek (Ambion, Cat No. / ID:AM2548). After incubation, the erythrocyte micronuclear DNA was extracted using the QIAamp DNA Blood Mini Kit (Qiagen, CatNo. / ID:51106). It should be noted that this method does not require a limit on the type of erythrocyte lysate, i.e., it is not necessary to specifically lyse erythrocytes without lysing nucleated cells. The micronuclear DNA was then purified by a silica column.

[0130] Example 4

[0131] Experimental verification of the effect of blood cell separation method

[0132] Statistics of micronuclear DNA fragments in red blood cells

[0133] The product of Example 2 was subjected to PacI restriction endonuclease to remove circular mitochondrial DNA. Atomic force microscopy (AFM) was used to image erythrocyte micronuclear DNA (Figure 2). In contrast to genomic DNA (gDNA), most MN DNA produced in vivo is short DNA fragments, indicating that MN DNA in erythrocytes is the result of accumulated DNA damage in vivo. The lengths of 1496 DNA fragments were measured to quantify the frequency of fragments in each size range and the size distribution of the fragments. The weighted average of the 1496 MN DNA fragments was 1437.57 nm, and the relative length of the DNA fragment base pairs scaled by nanometer length was 3705.01 bp. The length of MN DNA is much longer than that of cfDNA, which has a median length of 167 bp, indicating that the separation method in Example 2 can effectively eliminate the contamination of cfDNA.

[0134] Effects of Benzonase Nuclease on cfDNA

[0135] Three experimental groups were set up: ① Control group: cfDNA extracted from 1 ml of plasma; ② Experimental group 1: cfDNA extracted from 1 ml of plasma treated with Benzonase nuclease; and ③ Experimental group 2: cfDNA directly treated with Benzonase nuclease. The effect of Benzonase nuclease on the purification of erythrocyte micronuclear DNA was evaluated using Qubit. Figure 3 shows that the isolation method described in Example 2 effectively eliminates cfDNA contamination.

[0136] Effects of Filtration and Benzonase on the Purification of cfDNA During Erythrocyte Micronuclear DNA Purification

[0137] Furthermore, the present invention uses a xenograft model involving H460 human lung cancer cells to eliminate the possibility of circulating cfDNA contamination in MN DNA. In this model, the ratio of human cfDNA to mouse cfDNA positively correlates with tumor weight. Four weeks after tumor cell inoculation, MN DNA was isolated and sequenced from mouse red blood cells using the isolation method described in Example 2. No significant human DNA signal was detected, confirming the absence of cfDNA contamination in MN DNA. This demonstrates that the isolation method described in Example 2 can effectively eliminate cfDNA contamination.

[0138] Effects of Filtration on Monocytes During the Purification of Erythrocyte Micronuclear DNA

[0139] DAPI was used to stain the red blood cells in different separation steps in Example 2 to detect the possibility of nuclear cell contamination of the red blood cell micronuclear DNA. As shown in Figures 5A (human) and 5B (mouse), each point in the figure represents a quantitative analysis. Ten independent experiments were performed, and the scale bar is 20 microns. The purity of the separated red blood cells was determined under a microscope. The results showed that if only the density gradient centrifugation step was performed, the collected red blood cell micronuclear DNA would reduce the concentration of monocytes but would still be mixed with monocytes. However, if an additional red blood cell purification step was performed using a 10-micron cell sieve on the basis of density gradient centrifugation, the finally collected red blood cell micronuclear DNA did not contain any monocytes. This shows that the separation method in Example 2 can effectively eliminate monocyte contamination. It should be separately explained that the order of the centrifugation and filtration steps will bring unexpected technical effects. Only by the separation method in Example 2 can the contamination of monocytes be effectively eliminated.

[0140] To further test for contamination, two mixing experiments were performed. To detect contamination of gDNA in human samples, 10 610 mouse peripheral blood mononuclear cells were added to human peripheral blood, and MNDNA was isolated from the mixed cells and digital PCR was performed using mouse β-actin primers (Figure 4A). The enrichment of mouse β-actin was used as an indicator of gDNA contamination of MNDNA. 6 Human peripheral blood mononuclear cells were added to mouse peripheral blood (Figure 4B). MNDNA was isolated from the mixed cells and digital PCR was performed using human β-actin primers. The enrichment of human β-actin was used as an indicator of gDNA contamination of MNDNA.

[0141] The results showed that if only the density gradient centrifugation step was performed, the collected erythrocyte micronuclear DNA would reduce the concentration of monocytes but would still be contaminated with monocytes. However, if the density gradient centrifugation was supplemented with an additional erythrocyte purification step using a 10-micron cell sieve, the collected erythrocyte micronuclear DNA was free of monocytes. This indicates that the separation method in Example 2 can effectively eliminate monocyte contamination.

[0142] Example 5

[0143] Method for determining quality control benchmarks for micronuclear DNA in peripheral blood erythrocytes

[0144] The method comprises: a) providing a group of subjects with common characteristics; b) isolating or purifying peripheral blood red cell micronuclear DNA from peripheral blood red cells of each subject; c) performing whole genome sequencing on the peripheral blood red cell micronuclear DNA to obtain fragment sequence information of the micronuclear DNA; d) comparing the fragment sequence information of the peripheral blood red cell micronuclear DNA of the subject with the fragment sequence information of the genomic DNA of the peripheral blood mononuclear cells of the subject to obtain gene-enriched regions of the peripheral blood red cell micronuclear DNA of the subject; e) determining, based on the read counts of the peripheral blood red cell micronuclear DNA of the subject in the whole genome, candidate intervals in which the read abundance of the peripheral blood red cell micronuclear DNA of the subject is greater than the read abundance of the peripheral blood mononuclear cell genomic DNA of the subject; and f) determining a quality control benchmark based on the first read abundance of the mitochondrial DNA and / or the second read abundance of the gene-enriched region in the candidate interval. The principle of this method is described in detail below.

[0145] The present invention unexpectedly found that the addition of gDNA was positively correlated with an increase in the percentage of mtDNA. To further understand the genomic profile of MN DNA, the present invention performed whole-genome sequencing (WGS) on MN DNA samples and corresponding genomic DNA (gDNA) from healthy donors (n = 8910). Whole-genome analysis revealed that the distribution of MN DNA reads was preferentially enriched in introns and exons, with less enrichment in intergenic regions, compared to gDNA (Figure 6). The results showed that subsampling from 6x genome coverage (equivalent to approximately 60 million reads) to 1-2x genome coverage (equivalent to approximately 10 million reads) still yielded highly correlated results with the original profile. This suggests that WGS with 10 million reads is sufficient for MN DNA profiling. Furthermore, it was further confirmed that replicates of MN DNA samples sequenced with low-coverage reads were highly correlated. In low-coverage WGS, the correlation of normalized read density within 1Mb windows indicated that the erythrocyte micronuclear DNA group had higher intragroup consistency than the gDNA group. Therefore, this finding allows the analysis of MN DNA characteristics in hundreds of patients with different cancer types and HD in a cost-effective manner. Low-coverage sequencing reads were used in the present invention. The proportion of reads mapped from MN DNA to mtDNA was 0.0085%, which was significantly lower than the corresponding percentage of 0.0709% in the gDNA group (Figure 6). Therefore, mtDNA can be used as a marker for DNA contamination in MN DNA. To further test contamination, the present invention conducted a mixing experiment in which white blood cell DNA was mixed with MN DNA. The results showed that the addition of gDNA was positively correlated with the increase in the percentage of mtDNA. The detection of 0.05% mtDNA corresponded to the addition of 12pg of gDNA, which is equivalent to two eukaryotic cells. Therefore, this mtDNA threshold was used as a quality control benchmark for subsequent studies. Based on this finding, the present invention conducted the following experiments.

[0146] Step 1: Isolate or purify peripheral blood erythrocyte micronuclear DNA from peripheral blood erythrocytes of each subject.

[0147] Peripheral blood samples were collected from a group of subjects with shared characteristics (healthy subjects, subjects with the same cancer, or subjects with other shared characteristics). Peripheral blood erythrocyte micronuclear DNA and peripheral blood mononuclear cell DNA were isolated or purified according to the methods of Examples 1-3. The following description uses healthy subjects as an example.

[0148] Step 2: Perform whole genome sequencing on peripheral blood red blood cell micronuclear DNA to obtain fragment sequence information of micronuclear DNA and peripheral blood mononuclear cell genomic DNA.

[0149] MN-containing red blood cells are relatively rare in human peripheral blood (approximately 0.1%), and MN DNA accounts for only a small fraction of the total genomic DNA in cells. Therefore, to prepare MN DNA next-generation sequencing (NGS) libraries for sequencing, Illumina MN DNA sequencing libraries were prepared by Tn5 transposase-based labeling using the TruePrep Flexible DNA library Prep Kit for Illumina (Vazyme) according to the manufacturer's instructions. 200 pg of MN DNA or corresponding gDNA from the same sample was directly labeled with Tn5 transposase and then amplified 15 times by PCR using Illumina sequencing adapters. The DNA library was sequenced at 150 bp on the Novo-seq platform (NovaSeq 6000, Novogene, Beijing). 20 million reads were sequenced for MN DNA and corresponding gDNA from the same sample. To obtain sufficient MN DNA product for analysis in animal models, Illumina MN DNA sequencing libraries were prepared using the TruePrep DNA Library Prep Kit V2 for Illumina (Vazyme) with Tn5 transposase labeling according to the manufacturer's instructions. One ng of MDA product from MN DNA was directly labeled with Tn5 transposase and then amplified 13 times using Illumina sequencing adapters. The DNA library was sequenced on the Novo-seq platform (NovaSeq 6000, Novogene, Beijing) using 150-bp paired-end sequencing. MN DNA and corresponding gDNA from the same sample were sequenced for 20 million reads.

[0150] Initial processing of the whole-genome NGS data of MN DNA samples encoded in FASTQ files included 5'-end quality trimming and shearing of sequence adapters using Trimmomatic (version 0.39), followed by mapping to the human reference genome (GRCh38) or the mouse reference genome (mm10) using BWA-MEM (version 0.7.17). Non-primary, supplementary alignments, PCR duplicates, and read pairs with a MAPQ score of 30 were removed using SAMtools (version 1.10). To ensure equal representation of each sequencing sample, the filtered reads were downsampled to the same number. The number of downsampled reads after excluding low mappability and blacklisted genomic regions was analyzed in non-overlapping bins across the entire genome (excluding chromosomes X, Y, and M). To reduce bias in each sample, the present invention estimated the correction for GC content and the mappability of the number of reads in each bin.

[0151] Step 3: Compare the fragment sequence information of the micronuclear DNA of the subject's peripheral blood red blood cells with the fragment sequence information of the genomic DNA of the subject's peripheral blood mononuclear cells to obtain the gene enrichment region of the micronuclear DNA of the subject's peripheral blood red blood cells.

[0152] MACS2 was used to find the main enriched genomic regions of multiple red blood cell micronuclear DNA relative to the genomic DNA sequencing reads of peripheral blood mononuclear cells, and the "enriched regions" determined by multiple samples were merged.

[0153] Step 4: Based on the read counts of the micronuclear DNA in the peripheral blood red blood cells of the subject in the whole genome, determine the candidate intervals where the read abundance of the micronuclear DNA in the peripheral blood red blood cells of the subject is greater than the read abundance of the genomic DNA in the peripheral blood mononuclear cells of the subject.

[0154] As a preferred approach, the present invention selects 1Mb intervals where the abundance of erythrocyte micronuclear DNA is 1.2 times greater than that of peripheral blood mononuclear cell DNA as "1Mb candidate intervals" based on the median read count in the subject sample within the 1Mb interval across the entire genome. Ultimately, 6,566 enriched regions covering the PIP4K2A, ARHGAP25, PTPRC, FYN, PIK3R1, FYB1, and FLI1 genes were identified as erythrocyte micronuclear DNA-enriched intervals.

[0155] Figure 7 shows a representative whole-genome MN DNA map (top) and the corresponding gDNA (bottom) read density in a 1 Mb region covered by the MN DNA enriched region. The median curve for each group is represented by a dark solid line. For each 1 Mb interval, the differences between individuals are represented by shading. The ribbon diagram for each chromosome is shown at the bottom. Compared with gDNA, the 1 MB region containing the MN DNA enrichment peak showed an overall read density that was more than 1.2 times higher than that of genomic DNA. Figure 8 shows the Spearman correlation of the MN DNA or gDNA read density and the median MN DNA read density of 10 healthy individuals. Figure 9 shows a 100 kb analysis of a MN DNA enriched region on chromosome 1. The median curve for each group is represented by a dark solid line. For each 1 Mb box, the differences between individuals are represented by shading. FIG10 shows the proportions of random regions (left) and MN DNA-enriched regions (right) clearly annotated by ChromHMM (where annotation to a certain class exceeds 50%) (***p≤0.001, ****p≤0.0001, Welch's t-test).

[0156] The enriched regions are shown in Table 1 below.

[0157] Table 1

[0158] Step 5: Determine the quality control benchmark based on the abundance of the first read sequence of the mitochondrial DNA and / or the abundance of the second read sequence of the gene-enriched region in the candidate interval.

[0159] A gradient experiment of mixing peripheral blood mononuclear cell DNA was set up, and 0.3%, 0.6%, 0.9%, 1.5%, 3%, 6%, 9%, 12%, 15%, 30%, 45%, 60%, 75%, 90%, and 100% (for example, the proportion of peripheral blood nucleated cell DNA in 200 ng sample) of peripheral blood nucleated cell DNA was mixed into MNDNA for sequencing. The acceptable proportion of mixed peripheral blood nucleated cell DNA was determined based on the abundance of sequence reads in specific regions of these samples. As the abundance of reads in the 6566 erythrocyte micronuclear DNA-enriched intervals (PIP4K2A, ARHGAP25, PTPRC, FYN, PIK3R1, FYB1, FLI1 genes) and mitochondria increased with the proportion of DNA mixed with peripheral blood nucleated cells, the abundance of reads in the 6566 erythrocyte micronuclear DNA-enriched intervals (PIP4K2A, ARHGAP25, PTPRC, FYN, PIK3R1, FYB1, FLI1 genes) gradually decreased (65 The abundance of reads in the 66 erythrocyte micronuclear DNA-enriched regions or the PIP4K2A, ARHGAP25, PTPRC, FYN, PIK3R1, FYB1, and FLI1 genes showed a consistent trend, indicating that each of the 6566 erythrocyte micronuclear DNA-enriched regions or the PIP4K2A, ARHGAP25, PTPRC, FYN, PIK3R1, FYB1, and FLI1 genes can be used to determine quality control benchmarks. The abundance of reads in mitochondria gradually increased. Based on this, the acceptable proportion of contamination of peripheral blood nucleated cell DNA was determined to be 6%, based on the abundance of reads in the erythrocyte micronuclear DNA-enriched regions (PIP4K2A, ARHGAP25, PTPRC, FYN, PIK3R1, FYB1, and FLI1 genes) and mitochondria.

[0160] In order to simulate the possible influence of a large number of nucleated cells in the extraction process on the MN-DNA sequencing results, the present invention mixed 2*10 5 -*10 7 The nucleated cells of human peripheral blood monocytes (PBMCs) were introduced into the purified red blood cells, and the mixed products were subjected to the subsequent red blood cell lysis, library construction and sequencing processes.

[0161] Preferably, based on the median read count of the subject sample in a 1 Mb interval across the entire genome, the present invention calculates candidate intervals where the read abundance of micronuclear DNA in peripheral blood erythrocytes is greater than the read count of micronuclear DNA in erythrocytes from mixed nucleated cells across the entire genome and greater than the read abundance of genomic DNA from peripheral blood mononuclear cells. Combined with the genes annotated in the regions in Table 1, the 10 gene regions with the greatest differences were selected: PTPRC, MBNL1, TANK, THEMIS, FYB1, FYN, HIVEP2, PRKACB, INPP4B, and PIP4K2A, as erythrocyte micronuclear DNA enriched intervals.

[0162] At the same time, the present invention calculates the difference in the read counts of erythrocyte micronuclear DNA mixed with nucleated cells in the whole genome relative to the read abundance of peripheral blood mononuclear cell genomic DNA within a 1Mb region, and selects the region with the smallest difference but the read abundance of peripheral blood erythrocyte micronuclear DNA is significantly greater than that of these two groups of samples as the core region SATB1, ZFC3H1, UBAC2, and KLF12 with increased quality control characteristics of erythrocyte micronuclear DNA relative to samples with nucleated cell contamination.

[0163] Step 6: Perform quality control on micronuclear DNA.

[0164] Compare the abundance of micronuclear DNA enriched regions of red blood cells (PIP4K2A, ARHGAP25, PTPRC, FYN, PIK3R1, FYB1, FLI1 genes) or mitochondrial sequences in the test sample to the quality control benchmark established in step 5. If the quality control benchmark is exceeded, the test sample does not meet the requirements.

[0165] Example 6

[0166] A method for constructing a classifier for cancer detection using peripheral blood erythrocyte micronuclear DNA.

[0167] The method includes: a) providing a control group sample and more than one different category, wherein each category represents a group of subjects with common characteristics; b) isolating or purifying peripheral blood red blood cell micronuclear DNA from the control group sample and the peripheral blood red blood cells of each subject in each category; c) performing whole genome sequencing on the peripheral blood red blood cell micronuclear DNA to obtain fragment sequence information of the micronuclear DNA; d) comparing the fragment sequence information of the peripheral blood red blood cell micronuclear DNA of the control group sample and subjects of different categories; e) based on the differential distribution of the fragment sequence information of the micronuclear DNA in the peripheral blood red blood cells of subjects of different categories and the control group sample, selecting regions with statistically different read sequence abundance as feature regions, and obtaining a classifier for cancer detection by training the feature regions.

[0168] Step 1: Isolate or purify peripheral blood erythrocyte micronuclear DNA from the control sample and the peripheral blood erythrocytes of each subject in each category.

[0169] In the present invention, 2 mL of whole blood was collected from HDs (n=236) and patients (n=399) diagnosed with stage I, II, or III cancers prior to treatment, including lung cancer (LC, n=15), colorectal cancer (CRC, n=174), gastric cancer (GC, n=109), thyroid cancer (TC, n=55), and breast cancer (BC, n=46) ( FIG12 ). FIG12 shows the distribution of samples used in the present invention. These cancer patients were staged according to the American Joint Committee on Cancer (AJCC) guidelines. HDs served as age- and sex-matched controls for the patient group. The numbers in parentheses indicate the number of patients or HDs in that group. Blood analysis confirmed normal hematopoietic function in most patients, indicating an intact hematopoietic system. Peripheral blood erythrocyte micronuclear DNA and peripheral blood mononuclear cell DNA were isolated or purified according to the methods of Examples 1-3.

[0170] Step 2: Perform whole genome sequencing on the micronuclear DNA of peripheral blood red blood cells to obtain the fragment sequence information of the micronuclear DNA.

[0171] MN-containing red blood cells are relatively rare in human peripheral blood (approximately 0.1%), and MN DNA accounts for only a small fraction of the total genomic DNA in cells. Therefore, to prepare MN DNA next-generation sequencing (NGS) libraries for sequencing, Illumina MN DNA sequencing libraries were prepared by Tn5 transposase-based labeling using the TruePrep Flexible DNA library Prep Kit for Illumina (Vazyme) according to the manufacturer's instructions. 200 pg of MN DNA from the same sample was directly labeled with Tn5 transposase and then amplified 15 times by PCR using Illumina sequencing adapters. The DNA library was sequenced at 150 bp on the Novo-seq platform (NovaSeq 6000, Novogene, Beijing). 20 million reads of MN DNA from the same sample were sequenced. To obtain sufficient MN DNA product for analysis in animal models, Illumina MN DNA sequencing libraries were prepared using the TruePrep DNA Library Prep Kit V2 for Illumina (Vazyme) with Tn5 transposase labeling according to the manufacturer's instructions. Briefly, 1 ng of MDA product from MN DNA was directly labeled with Tn5 transposase, followed by 13 PCR amplifications using Illumina sequencing adapters. The DNA library was sequenced on the Novo-seq platform (NovaSeq 6000, Novogene, Beijing) using 150 bp paired-end sequencing. MN DNA from the same sample was sequenced for 20 million reads.

[0172] Initial processing of whole-genome NGS data from MN DNA samples encoded in FASTQ files involved 5'-end quality trimming and shearing of sequence adapters using Trimmomatic (version 0.39), followed by mapping to the human reference genome (GRCh38) or mouse reference genome (mm10) using BWA-MEM (version 0.7.17). Non-primary, complementary alignments, PCR duplicates, and read pairs with a MAPQ score of 30 were removed using SAMtools (version 1.10). To ensure equal representation of each sequencing sample, filtered reads were downsampled to the same number. The number of downsampled reads, after excluding low mappability and blacklisted genomic regions, was analyzed in non-overlapping bins across the entire genome (excluding chromosomes X, Y, and M). To reduce bias within each sample, we estimated the GC content correction and mappability of the number of reads in each bin.

[0173] Step 3: Compare the fragment sequence information of peripheral blood red blood cell micronuclear DNA of the control group samples and subjects of different categories.

[0174] Genome-wide analysis revealed that, compared with HD, cancer patients showed significant signal enrichment in intergenic and centromeric regions of MN DNA, whereas no such enrichment was observed in genic regions. The correlation of MN DNA read distribution ratios in cancer patients deviated significantly from that in HD patients, suggesting that early malignancy leads to altered MN DNA genomic profiles (Figure 11). Figure 11 shows the Spearman correlation between MN DNA read density and median MN DNA read density in cancer patients and healthy individuals (****p ≤ 0.0001, Welch's t test). Figure 12 shows representative examples of MN DNA associated with five tumor types.

[0175] Step 4: Based on the differential distribution of micronuclear DNA fragment sequence information in peripheral blood red blood cells of different categories of subjects and control samples, regions with statistically different read abundance are selected as feature regions, and a classifier for cancer detection is obtained by training the feature regions.

[0176] We divided samples from HD and cancer patients into three independent groups: a training set containing 65% (n=408) of the total samples for model development, a validation set containing 15% (n=101) for fine-tuning hyperparameters, and a test set containing the remaining 20% ​​(n=126) for validating model performance. To further investigate the differences in MN DNA profiles between cancer patients and HD, we analyzed autosomal MN DNA read density. We identified 288 genomic regions in cancer patients from the training and validation sets, termed tumor-associated MN DNA (taMN DNA) signatures, that exhibited significant changes in normalized read counts compared to HD (Figure 13). Figures 13A-13B show regions with upregulated MN DNA read density across all cancer patients, while Figures 13C-13D show regions with downregulated MN DNA read density across all cancer patients. For each signature, the right side of each figure shows the difference in read density observed between HD and pan-cancer patients or individual cancer types, and the left side of each figure plots the genomic distribution of these intervals. Figure 13 exemplifies the method for identifying tumor-associated MN DNA.

[0177] Notably, 262 of these 288 features (90.97%) were not consistent with MN DNA-enriched regions of the genome (the sites identified in Example 5). As shown in Figure 14, the genomic distribution of the 288 selected taMN DNA features across the genome showed negligible overlap with MN DNA-enriched regions.

[0178] Furthermore, most taMN DNA features corresponded to regions marked by quiescent chromatin (Figure 15). Thirty percent of these overlapped with fragile sites, particularly CFSs, suggesting increased replication stress and the consequent DNA breakage that occurs during tumor progression. Training these signature regions could yield a classifier for cancer detection.

[0179] Specifically, the screening of tumor-related features in erythrocyte micronuclear DNA and the construction of a classifier include:

[0180] 1. Sample Grouping: Qualified samples are randomly grouped into 80% training and validation sets and 20% testing sets. This 80% training and validation set is then split into 80% training (64% of the total sample size) and 20% validation (16% of the total sample size) for model parameter adjustment during subsequent model training. Finally, the model is tested on the 20% testing set to obtain the model's prediction results.

[0181] 2. Screening of tumor-related differential features: The whole genome was divided into 10kb equal-width intervals to evaluate the abundance of micronuclear DNA. At the same time, 2-100 consecutive 10kb intervals were merged to obtain intervals of different lengths from 10kb to 1000kb, and the abundance of micronuclear DNA was evaluated. Based on the statistical differences in the read abundance of the control group samples and multiple cancer samples (including thyroid cancer, intestinal cancer, gastric cancer, lung cancer, and breast cancer) in these intervals in 80% of the training and validation sets, a total of 288 candidate tumor-related feature regions were obtained. The regions with statistically different read abundance are shown in Table 2 below.

[0182] Table 2

[0183] Among them: (1) According to the model training performance of the training validation set, the control and multiple cancer samples (including thyroid cancer, intestinal cancer, gastric cancer, lung cancer, and breast cancer) were used to train the XGBoost model, and each feature in the candidate features was cross-validated 5 times. Finally, a feature subset with the best model performance in the current sample was obtained, totaling 123 features, as candidates for the pan-cancer model. The feature subset covers the following genes: SARS1, MYBPHL, SORT1, PSMA5, SYPL2, CELSR2, PSRC1, POGK, TADA1, MAEL, GPA33, STYXL2, ILDR2, POU2F1, POU2F1-DT, C1orf105, SUCO, RGS16, RGSL1, RNASEL, NPL, DHX9, RGS8, ZNF281, KIF14, CAMSAP2, DDX59, DUSP10, EGLN1, OSBPL9, RAB3B, NRDC, RERE ,ERRFI1-DT,SLC45A1,ERRFI1,PTBP2,CFAP43,GSTO1,SFR1,NRG3,PANK1,KIF20B,USP47,MICAL2,DKK3,DISC1FP1,SCAF11,TEX30, POGLUT2,BIVM-ERCC5,BIVM,ERCC5,METTL21EP,TPP2,METTL21C,CCDC168,RNY3P9,RB1,CYSLTR2,FNDC3,PCNX4,DHRS7,ATP10A,RAB 27A,DAPK2,CIAO2A,SNX1,RBFOX1,SMARCE1,KRT222,KRT20,KRT23,KRT39,KRT40,KRTAP3-3,KRT24,KRT25,KRT26,KRT27,KRT28,K RT10,KRT12,KRTAP2-3,KRTAP2-4,KRTAP3-1,KRTAP1-5,KRTAP1-4,KRTAP3-2,KRTAP1-1,KRTAP2-1,KRTAP2-2,KRTAP1-3,CDC27,MY L4,ITGB3,EFCAB13,RNU7-186P,EFCAB13-DT,FAP,GCG,IFIH1,GCA,SPC25,KCTD18,SGO2,SPATS2L,RNY4P34,NBEAL1,ICA1L,WDR12, CARF,IKZF2,HTR2B,GPR55,SPATA3,C2orf72,PSMD1,GKN2,BMP10,SIAH2,KCNAB1,SSR3,TIPARP,LEKR1,RN7SKP177,SCN11A,WDR48,RPSA, SNORA6, MOBP, CSRNP1, GORASP1, TTC21A, CCR8, SLC25A38, XIRP1, CX3CR1, CCR5AS, CCR5, LTF, CCR2, CCRL2, EPHA6, AR L6,CRYBG3,BANK1,PPP3CA,FLJ20021,RN7SL446P,FBXW7,KCNIP4,GBA3,NRG2,PURA,IGIP,CYSTM1,CDH12,TTC33,ESM1,CER T1,POLK,HMGCR,RNU7-175P,ANKDD1B,WASF1,CDC40,CALHM4,TRAPPC3L,RWDD1,RSPH4A,ZUP1,AHI1,AHI1-DT,ZDHHC14,TM EM242-DT,TMEM242,PRKN,BRPF3,PNPLA1,BNIP5,ETV7,PXT1,KCTD20,C7orf33,CUL1,BMPER,PHF20L1,TG,DNAAF11,TMEM71 ,PTCSC1,CCN4,NDRG1,SLA,ESCO2,PBK,CCDC25,CDH17,FSBP,VIRMA,GEM,RAD54B,ESRP1,VIRMA-DT,PANK1-AS1,TUBA3C,L RFN5,GOLGA6L2,UBBP4,FAM27E5,FLJ36000,ZNF675,NLRP12,MYADM-AS1,MYADM,PRKCG,CACNG7,RNU1-7P,DNM3,PIGC,FMN2 ,ACSS1,CST7,APMAP,C1QL2,RN7SL468P,ZEB2,GTDC1,GALNT13,SPAG16,SPAG16-DT,ATG16L1,SCARNA6,SAG,DGKD,GRM7,DCAF16,NCAPG,LCORL,ARAP2,ANKRD31,IRF1,IL5,RAD50,TH2LCRR,IL13,GFOD1,TSBP1-AS1,TSBP1,BTNL2,HLA-DRA,SDCBP. The specific regions are shown in Table 3 below.

[0184] A recursive elimination algorithm (RFE) was used to identify 288 differentially expressed features from the control and multiple cancer samples (including thyroid, colorectal, gastric, lung, and breast cancer) used in the training and validation sets. The algorithm used repeated cross-validation (20 times) to evaluate the stability of the features was used. The average importance of each feature across the 20 replicates was recorded to assess its overall contribution. Features with the highest frequency of occurrence and a significant decrease in mean feature importance during the recursive elimination process were selected. For the detection of pan-cancer MNDNA (thyroid, colorectal, gastric, lung, and breast cancer), core regions were defined. These regions are a subset of the current Table 3. It is important to clarify that this section focuses on protecting the following regions and the gene intervals associated with these regions: PRKN, CERT1, PSMD1, IKZF2, GCG, CFAP43, PTBP2, and AHI1.

[0185] Table 3

[0186] (2) Based on the model training performance of the training validation set, the control and colorectal cancer samples were used for XGBoost model training, and each of the 288 candidate features was cross-validated 5-fold. Finally, a feature subset with the best model performance in the current sample was obtained, totaling 75 features, as candidates for the colorectal cancer model. This feature subset covers the following genes: PRNLS, LIPJ, LIPF, ANKRD22, LIPK, LIPN, LIPM, STAMBPL1, KIF20B, PANK1, MIR107, FDX1, RDX, ZC3H12C, TRHDE, TUBA3C, MDGA2, GOLGA6L2, GABRG3, GNAO1-DT, GNAO1, UBBP4, FAM27E5, FLJ36000, NLRP12,MYADM,PRKCG,CACNG7,PTBP2,NBPF6,EEIG2,FMN2,REL,PUS10,PEX13,CTNNA2,GACAT1,RGPD4, SPAG16,IKZF2,SPAG16-DT,GRM7,CDV3,TOPBP1,TF,SRPRB,RAB6B,KCNAB1,SSR3,ANKRD31,HMGCR,CERT1 ,POLK,FAM174A-DT,FAM174A,RN7SKP62,ST8SIA4,GFOD1,ELOVL4,TTK,UBE3D,DOP1A,PGM3,RWDD2A,ME 1,AKAP7,PRKN,ING3,CPED1,C7orf33,CUL1,SGCZ,SDCBP,NSMAF,TOX,OSGIN2,NBN,CSMD3,PKN2,DPYD, GSTO1, SFR1, CFAP43, EBLN1, NELL1, SHANK2, RB1, CYSLTR2, FNDC3A, RCBTB2, SMARCE1, KRT222, KRT24, KRT25, KRT26, KRT27, KRT28, KRT10, KRT12, KRT20, KRT23, SYN3, TIMP3, PTPRK, FBXO30, EPM2A-DT, SHPRH. The specific regions are shown in Table 4.

[0187] For the 288 candidate features that differed from those in Table 2, a recursive elimination algorithm (RFE) was performed using 20 repeated cross-validations. The number of times a feature was selected in each cross-validation was recorded as an important indicator of feature stability, and the average importance of each feature in the 20 repeated experiments was recorded to assess its overall contribution. Features with the highest frequency of occurrence and whose mean feature importance showed a significant downward trend during the recursive elimination process were selected. Regarding the addition of core regions to the definition of MNDNA as a colorectal cancer test, it is necessary to clarify that this section focuses on protecting the following regions and the gene intervals associated with these regions: ABCC10, AP1G1, NFKBIA, JAKMIP2, PRKN, EXOC4, CFAP43, SYN3, IKZF2, NELL1, SHPRH, PUS10, PTPRK, PKN2 (some of which belong to the current subset of Table 4, and the rest are from the 288 regions).

[0188] Table 4

[0189] (3) Based on the model training performance of the training and validation sets, the XGBoost model was trained using control and gastric cancer samples. Each of the 288 candidate features was cross-validated 5-fold. Finally, a feature subset with the best model performance in the current sample was obtained, totaling 14 features, as candidate gastric cancer models. This feature subset covers the following genes: PRKN, DCDC1, TUBA3C, ZNF675, C1QL2, RN7SL468P, RNU6-725P, LINC01208, LINC00408, LINC00442, LINC02479, LINC03025, MIR7977. The specific regions are shown in Table 5 below.

[0190] Table 5

[0191] 3. Classifier construction: Based on the features of multiple classification tasks confirmed, the classifier is constructed using the XGBoost model.

[0192] 4. Sample Prediction: Based on the training model obtained in the previous step, the present invention uses the test set samples that did not participate in the training and uses the classifier constructed in the previous step to predict the test set samples. The predicted results of the test set and the true labels of the samples are obtained, and the proportion of each predicted result in the two categories (i.e., the risk assessment index) is displayed. Predictions are made for unknown samples and the binary classification results are displayed.

[0193] 5. Results: In the test set, the selected MN DNA signature demonstrated impressive discriminatory power between HD and cancer patients, achieving an area under the receiver operating characteristic (ROC) curve (AUC) of 96% (Figure 16). The analyses for the training, validation, test, and random classifier cohorts are shown in red, blue, green, and gray lines, respectively, in Figure 16A. The dashed line indicates no discriminatory power (left). The overall accuracy of pan-cancer classification for each cancer type and stage in the test set, including the corresponding sensitivity and specificity, was 95.7% (CI, confidence interval) (right). Figure 16B shows the ROC curves for the binary outcome of HD and colorectal cancer patients paired with the training cohort (red), validation cohort (blue), and test cohort (green) in Figure 16A (left) and the sensitivity for distinguishing colorectal cancer clinical stages, with a specificity of 93.6%. The numbers in parentheses represent the number of samples (right). Figure 16C shows the ROC curves (left) for the binary outcomes of HD and gastric cancer patients for the paired training cohort (red), validation cohort (blue), and test cohort (green) in Figure 16A, and the sensitivity for distinguishing the clinical stage of gastric cancer with a specificity of 93.6%. The numbers in parentheses represent the number of samples (right).

[0194] Specifically, the model correctly identified 79% (95% confidence interval: 69-88%) of cancer patients across various cancer types, including thyroid cancer (TC), breast cancer (BC), colorectal cancer (CRC), gastric cancer (GC), and lung cancer (LC), with an overall specificity of 95.7%. This remarkable detection rate included 78% of stage I, 83% of stage II, and 76% of stage III cancer patients. Notably, this predictive ability was not present in a random classifier, highlighting the specificity of the taMN DNA signature.

[0195] Further analysis revealed high sensitivity and specificity when using the taMN DNA signature to distinguish specific types of cancer from HD. For example, in pairwise comparisons, the model showed AUCs of 96% and 92% for patients with CRC and GC, respectively. The CRC model demonstrated a sensitivity of 83% (5 of 6) for stage I, 100% (12 of 12) for stage II, and 81% (13 of 16) for stage III, with a specificity of 95.7%. Similarly, the GC-specific model demonstrated a sensitivity of 75% (12 of 16) for both stage I and stage II patients, and a sensitivity of 100% (4 of 4) for stage III patients, with a specificity of 91.5%. Overall, these data suggest that these cancer-associated MN DNA signatures can be used to distinguish patients with cancer from those with HD.

[0196] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for separating or purifying micronuclear DNA from peripheral blood red blood cells, characterized in that, Comprising the following steps: a) Providing a peripheral blood sample; b) Separating mononuclear cells and red blood cells in the peripheral blood sample; c) Collecting the red blood cells; d) Lysing the red blood cells; and e) Extracting micronuclear DNA.

2. The method according to claim 1, wherein The step b) comprises: Centrifuging the peripheral blood sample to obtain a mononuclear cell layer and a red blood cell layer; Filtering the red blood cell layer to obtain a filtrate containing red blood cells.

3. The method according to claim 2, wherein The step b) further comprises: Treating the filtrate with Benzonase nuclease.

4. The method according to claim 2, wherein Centrifuging the peripheral blood sample by density gradient.

5. The method according to claim 2, wherein Filtering the red blood cell layer using a 10-micron cell sieve.

6. Use of at least one of the genes ARHGAP25, PIK3R1, FLI1, PTPRC, MBNL1, TANK, THEMIS, FYB1, FYN, HIVEP2, PRKACB, INPP4B, PIP4K2A, SATB1, ZFC3H1, UBAC2, KLF12 in peripheral blood red blood cell micronuclear DNA in a system for controlling the quality of micronuclear DNA.

7. Use of at least one region as shown in Table 1 of peripheral blood red blood cell micronuclear DNA in a system for controlling the quality of micronuclear DNA.

8. A method for determining the quality control benchmark of micronucleus DNA in peripheral blood red blood cells, characterized in that, Comprising: a) Providing a group of subjects with common characteristics; b) Separating or purifying peripheral blood red blood cell micronuclear DNA from the peripheral blood red blood cells of each subject; c) Performing whole-genome sequencing on the peripheral blood red blood cell micronuclear DNA to obtain fragment sequence information of the micronuclear DNA; d) Comparing the fragment sequence information of the peripheral blood red blood cell micronuclear DNA of the subject and the fragment sequence information of the genomic DNA of the peripheral blood mononuclear cells of the subject to obtain the gene enrichment region of the peripheral blood red blood cell micronuclear DNA of the subject; e) Determining a candidate interval where the read abundance of the peripheral blood red blood cell micronuclear DNA of the subject is greater than the read abundance of the genomic DNA of the peripheral blood mononuclear cells of the subject; f) Determining the quality control benchmark according to the first read abundance of mitochondrial DNA and / or the second read abundance of the gene enrichment region in the candidate interval.

9. The method of the quality control benchmark according to claim 8, characterized in that, In the step e), according to the median of the read counts of the peripheral blood red blood cell micronuclear DNA of the subject in a 1-Mb interval of the whole genome, a 1-Mb interval where the read abundance of the peripheral blood red blood cell micronuclear DNA of the subject relative to the genomic DNA of the peripheral blood mononuclear cells of the subject is greater than 1.2 times is determined as the candidate interval.

10. The method according to claim 8, wherein The group of subjects with common characteristics includes: non-cancer subjects.

11. The method according to claim 8, wherein The length of the gene enrichment region in the candidate interval ranges from 179 bp to 32.04 kb.

12. The method according to claim 8, wherein The gene enrichment region in the candidate interval contains the region shown in Table 1 or the region corresponding to the genes ARHGAP25, PIK3R1, FLI1, PTPRC, MBNL1, TANK, THEMIS, FYB1, FYN, HIVEP2, PRKACB, INPP4B, PIP4K2A, SATB1, ZFC3H1, UBAC2, KLF12.

13. The method according to claim 8, wherein Determining the quality control benchmark according to the first read abundance of mitochondrial DNA and / or the second read abundance of the gene enrichment region in the candidate interval, including: a) Providing a set of peripheral blood erythrocyte micronuclear DNAs containing more than one genomic DNA of peripheral blood mononuclear cells with different proportions; b) Performing whole-genome sequencing on the set of peripheral blood erythrocyte micronuclear DNAs to determine the first read abundance of the mitochondrial DNA and / or the second read abundance of the gene enrichment region in the candidate interval; c) Determining the quality control benchmark according to the influence of the proportions of the genomic DNAs of peripheral blood mononuclear cells with different proportions on the first read abundance of mitochondrial DNA and / or the second read abundance of the gene enrichment region in the candidate interval.

14. A method for quality control in the process of extracting micronuclear DNA from peripheral blood red blood cells, characterized in that, Including: Comparing the first read abundance of mitochondrial DNA in peripheral blood erythrocyte micronuclear DNA and / or the second read abundance of at least one region shown in Table 1 or at least one gene corresponding region among the genes ARHGAP25, PIK3R1, FLI1, PTPRC, MBNL1, TANK, THEMIS, FYB1, FYN, HIVEP2, PRKACB, INPP4B, PIP4K2A, SATB1, ZFC3H1, UBAC2, KLF12 with the quality control benchmark determined in claim 8.

15. SARS1, MYBPHL, SORT1, PSMA5, SYPL2, CELSR2, PSRC1, POGK, TADA1, MAEL, GPA33, STYXL2, ILDR2, POU2F1, POU2F1-DT, C1orf105, SUCO, RGS16, RGSL1, RNASEL, NPL, DHX9, RGS8, ZNF281, KIF14, CAMSAP2, DDX59, DUSP10, EGLN1, OSBPL9, RAB3B, NRDC, RERE, ERRFI1-DT, SLC45A1, ERRFI1, PTBP2, CFAP43, GSTO1, SFR1, NRG3, PANK1, KIF20B, USP47, MICAL2, DKK3, DISC1FP1, SCAF11, TEX30, POGLUT2, BIVM-ERCC5, BIVM, ERCC5, METTL21EP, TPP2, METTL21C, CCDC168, RNY3P9, RB1, CYSLTR2, FNDC3, PCNX4, DHRS7, ATP10A, RAB27A, DAPK2, CIAO2A, SNX1, RBFOX1, SMARCE1, KRT222, KRT20, KRT23, KRT39, KRT40, KRTAP3-3, KRT24, KRT25, KRT26, KRT27, KRT28, KRT10, KRT12, KRTAP2-3, KRTAP2-4, KRTAP3-1, KRTAP1-5, KRTAP1-4, KRTAP3-2, KRTAP1-1, KRTAP2-1, KRTAP2-2, KRTAP1-3, CDC27, MYL4, ITGB3, EFCAB13, RNU7-186P, EFCAB13-DT, FAP, GCG, IFIH1, GCA, SPC25, KCTD18, SGO2, SPATS2L, RNY4P34, NBEAL1, ICA1L, WDR12, CARF, IKZF2, HTR2B, GPR55, SPATA3, C2orf72, PSMD1, GKN2, BMP10, SIAH2, KCNAB1, SSR3, TIPARP, LEKR1, RN7SKP177, SCN11A, WDR48, RPSA, SNORA6, MOBP, CSRNP1, GORASP1, TTC21A, CCR8, SLC25A38, XIRP1, CX3CR1, CCR5AS, CCR5, LTF, CCR2, CCRL2, EPHA6, ARL6, CRYBG3, BANK1, PPP3CA,The application of at least one gene among FLJ20021, RN7SL446P, FBXW7, KCNIP4, GBA3, NRG2, PURA, IGIP, CYSTM1, CDH12, TTC33, ESM1, CERT1, POLK, HMGCR, RNU7-175P, ANKDD1B, WASF1, CDC40, CALHM4, TRAPPC3L, RWDD1, RSPH4A, ZUP1, AHI1, AHI1-DT, ZDHHC14, TMEM242-DT, TMEM242, PRKN, BRPF3, PNPLA1, BNIP5, ETV7, PXT1, KCTD20, C7orf33, CUL1, BMPER, PHF20L1, TG, DNAAF11, TMEM71, PTCSC1, CCN4, NDRG1, SLA, ESCO2, PBK, CCDC25, CDH17, FSBP, VIRMA, GEM, RAD54B, ESRP1, VIRMA-DT, PANK1-AS1, TUBA3C, LRFN5, GOLGA6L2, UBBP4, FAM27E5, FLJ36000, ZNF675, NLRP12, MYADM-AS1, MYADM, PRKCG, CACNG7, RNU1-7P, DNM3, PIGC, FMN2, ACSS1, CST7, APMAP, C1QL2, RN7SL468P, ZEB2, GTDC1, GALNT13, SPAG16, SPAG16-DT, ATG16L1, SCARNA6, SAG, DGKD, GRM7, DCAF16, NCAPG, LCORL, ARAP2, ANKRD31, IRF1, IL5, RAD50, TH2LCRR, IL13, GFOD1, TSBP1-AS1, TSBP1, BTNL2, HLA-DRA, SDCBP in a system for pan-cancer detection of a test subject., 16. Use of at least one region shown in Table 3 of peripheral blood erythrocyte micronuclear DNA in a system for pan-cancer detection of a test subject.

17. Use of at least one of the genes PRNLS, LIPJ, LIPF, ANKRD22, LIPK, LIPN, LIPM, STAMBPL1, KIF20B, PANK1, MIR107, FDX1, RDX, ZC3H12C, TRHDE, TUBA3C, MDGA2, GOLGA6L2, GABRG3, GNAO1-DT, GNAO1, UBBP4, FAM27E5, FLJ36000, NLRP12, MYADM, PRKCG, CACNG7, PTBP2, NBPF6, EEIG2, FMN2, REL, PEX13, CTNNA2, GACAT1, RGPD4, SPAG16, SPAG16-DT, GRM7, CDV3, TOPBP1, TF, SRPRB, RAB6B, KCNAB1, SSR3, ANKRD31, HMGCR, CERT1, POLK, FAM174A-DT, FAM174A, RN7SKP62, ST8SIA4, GFOD1, ELOVL4, TTK, UBE3D, DOP1A, PGM3, RWDD2A, ME1, AKAP7, ING3, CPED1, C7orf33, CUL1, SGCZ, SDCBP, NSMAF, TOX, OSGIN2, NBN, CSMD3, DPYD, GSTO1, SFR1, EBLN1, SHANK2, RB1, CYSLTR2, FNDC3A, RCBTB2, SMARCE1, KRT222, KRT24, KRT25, KRT26, KRT27, KRT28, KRT10, KRT12, KRT20, KRT23, TIMP3, FBXO30, EPM2A-DT, ABCC10, AP1G1, NFKBIA, JAKMIP2, PRKN, EXOC4, CFAP43, SYN3, IKZF2, NELL1, SHPRH, PUS10, PTPRK, PKN2 in a system for detecting colorectal cancer in a test subject.

18. Use of at least one region as shown in Table 4 of peripheral blood erythrocyte micronucleus DNA in a system for detecting colorectal cancer in a test subject.

19. Use of at least one of the genes PRKN, DCDC1, TUBA3C, ZNF675, C1QL2, RN7SL468P, RNU6-725P, LINC01208, LINC00408, LINC00442, LINC02479, LINC03025, MIR7977 in peripheral blood erythrocyte micronucleus DNA in a system for detecting gastric cancer in a test subject. Use of at least one region of peripheral blood erythrocyte micronuclear DNA as shown in Table 5 in a system for detecting gastric cancer in a test subject.

21. A method for constructing a classifier for cancer detection using micronucleus DNA of peripheral blood red blood cells, characterized in that, Including: a) Providing a control group sample and more than one different category, where each category represents a group of subjects with a common characteristic; b) Separating or purifying peripheral blood erythrocyte micronuclear DNA from the control group sample and peripheral blood erythrocytes of each subject in each category; c) Performing whole-genome sequencing on the peripheral blood erythrocyte micronuclear DNA to obtain fragment sequence information of the micronuclear DNA; d) Comparing the fragment sequence information of the peripheral blood erythrocyte micronuclear DNA of the control group sample and subjects in different categories; e) According to the differential distribution of the fragment sequence information of the micronuclear DNA in the peripheral blood erythrocytes of subjects in different categories and the control group sample, selecting regions with statistically significant differences in read sequence abundance as characteristic regions, and obtaining a classifier for cancer detection by training the characteristic regions.

22. The method according to claim 21, wherein The different categories include: Cancer subjects with different cancers.

23. The method according to claim 22, wherein The different categories include: Thyroid cancer subjects, intestinal cancer subjects, gastric cancer subjects, lung cancer subjects, and breast cancer subjects.

24. The method according to claim 21, wherein Obtaining a classifier for cancer detection by training the characteristic regions, including: Performing XGBoost model training on the characteristic regions of subjects in different categories and the control group sample, and performing 5-fold cross-validation on each of the characteristic regions, and using the first characteristic subset of the obtained optimal model for the classifier for cancer detection, where the cancer is pan-cancer.

25. The method according to claim 21, characterized in that, Obtaining a classifier for cancer detection by training the characteristic regions, including: Performing XGBoost model training on the characteristic regions of the intestinal cancer subjects and the control group sample, and performing 5-fold cross-validation on each of the characteristic regions, and using the second characteristic subset of the obtained optimal model for the classifier for cancer detection, where the cancer is intestinal cancer.

26. The method according to claim 21, wherein Obtaining a classifier for cancer detection by training the characteristic regions, including: Performing XGBoost model training on the characteristic regions of the gastric cancer subjects and the control group sample, and performing 5-fold cross-validation on each of the characteristic regions, and using the third characteristic subset of the obtained optimal model for the cancer detection classifier, where the cancer is gastric cancer.

27. A system for cancer detection in a test subject, characterized in that, Including: A detection device for detecting peripheral blood erythrocyte micronuclear DNA from the test subject by the classifier constructed according to claim 21.

28. The system according to claim 27, wherein Also including: A separation device for separating peripheral blood erythrocyte micronuclear DNA from the test subject; A sequencing device for sequencing the peripheral blood erythrocyte micronuclear DNA from the test subject.

29. The system according to claim 27, wherein The system performs cancer detection by the following method, and the method includes: a) Separating or purifying the micronuclear DNA in the peripheral blood erythrocytes of the test subject; b) Performing whole-genome sequencing on the micronuclear DNA to obtain fragment sequence information of the micronuclear DNA in the peripheral blood erythrocytes of the test subject; c) Detecting the fragment sequence information of the micronuclear DNA obtained in step b) by the classifier constructed according to claim 21, thereby classifying the test subject as one or more of the more than one different categories.

30. The system according to claim 29, wherein The cancer detection includes cancer screening, diagnosis, typing, and / or staging.