Cell methylation marker-based organ injury detection method and application
Through a relative quantitative method based on cell methylation markers, the proportion of cell sources in plasma cfDNA was analyzed, the type and degree of organ damage was identified, and background noise was eliminated through internal reference correction, which solved the problem of insufficient identification of organ damage signals in the prior art, and achieved highly sensitive organ damage detection.
Patent Information
- Application Number
- CN202411999515.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The prior art is difficult to develop a high sensitivity, high robustness, and accurate identification method for organ damage traceability at the cellular level, especially in identifying the target organ/cell trace damage signals.
Through a relative quantitative method based on cell methylation markers, the relative proportions of various cell sources in plasma cfDNA were analyzed, the types of damage and their degree were identified, and background noise was eliminated through the internal reference correction system to achieve true reduction of target cell signals.
It realizes high sensitivity detection of organ damage, can accurately identify the type and degree of damage, and assists in early disease screening, early diagnosis, disease course monitoring and prognosis evaluation.
Smart Images

Figure BDA0005226799400000041 
Figure BDA0005226799400000042 
Figure BDA0005226799400000051
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedicine, and specifically relates to a detection method and application of organ damage based on cell methylation markers. Background Art
[0002] Organ damage is a common manifestation of many diseases. Timely and accurate detection of organ damage and its degree can help develop more effective treatment plans and improve the treatment and prognosis of patients. For example, kidney damage often indicates that renal function is about to change or has changed, and liver damage may be a direct manifestation of diseases such as hepatitis or cirrhosis. The diagnosis of organ damage often requires multiple methods such as imaging, pathology, and laboratory tests. Different types of organ damage and their degree may have different characteristics and manifestations, which increases the complexity and difficulty of actual clinical diagnosis. For example, in the diagnosis of kidney diseases with high prevalence and low awareness, especially for early kidney damage detection, existing clinical indicators often show low sensitivity (serum creatinine, glomerular filtration rate). Such indicators are often positive and represent serious renal function problems, with a heavy lag. In addition, renal biopsy technology needs to be carefully considered in clinical application due to its invasiveness, sampling location, patient constitution, limited repetition, and complications. Therefore, the development of a non-invasive and highly sensitive organ damage detection method is crucial to improve diagnostic efficiency, formulate effective plans, efficacy detection, and improve prognosis.
[0003] Compared with invasive tissue biopsies, liquid biopsy technology has attracted widespread attention in recent years for its non-invasive and sensitive characteristics. cfDNA (Cell-free DNA, cfDNA) is a highly fragmented DNA molecule released into the circulatory system by cell apoptosis and necrosis, with a short half-life (about 5-150 minutes). In the occurrence, development and treatment of many diseases, such as cancer, autoimmune diseases, and various chronic diseases, the cell death rate can be affected, thereby affecting the proportion of cfDNA from various organs / cells in the blood. cfDNA provides a comprehensive and non-invasive health profile of all cell tissues in the body, deciphers and quantifies the sources of the proportions of various organ / cell components in cfDNA, and has great clinical potential in assisting disease diagnosis, efficacy monitoring and prognosis.
[0004] DNA methylation is a highly conserved basic epigenetic modification that controls gene expression and chromatin organization. It plays an important role in biological processes such as gene expression regulation, cell differentiation, embryonic development, and disease occurrence. Netanel Loyfer et al. demonstrated that DNA methylation affects downstream gene expression in functional regulation mainly through the interconnectedness of multiple CpG sites. In addition, they mapped the first unique DNA methylation map of major human cell types, and the vast majority of cell-specific methylation markers tend to be low-methylated or completely unmethylated. This also provides a new analytical perspective for the use of DNA methylation markers for quantitative cell tracing.
[0005] For the quantitative tracing of cfDNA methylation, some existing methods have their own limitations. Based on the deconvolution algorithm, Netanel Loyfer et al. have successively developed quantitative tracing methods characterized by methylation rate (MethAltas) and fragments (UMX / Markov). However, since the deconvolution method requires the estimation of all cell signals in cfDNA, it is impossible to trace the origin of cell types that lack methylation markers, which is prone to quantitative technical deviation. In addition, due to its own limitations, the deconvolution algorithm cannot accurately identify the trace signals of cfDNA (ten thousandths-thousandths). Christa Caggiano developed a low-depth whole-genome cell tracing quantitative method characterized by CpG sites based on the maximum likelihood (Maximum Likelihood Estimation, MLE) algorithm. However, due to the highly fragmented characteristics of cells, the robustness and specificity of using the site methylation rate as a quantitative indicator for cell tracing are poor, and it is also unable to meet the accurate identification of trace signals at low depth. In addition, Shuo Li et al. developed a tissue-level tracing quantitative model based on fragment methylation rate characteristics based on a deep neural network model. However, this model requires complex hyperparameter adjustment and training support from a large cohort of samples, has poor model interpretability, and existing methylation markers cannot meet the cellular-level tracing resolution.
[0006] In summary, it is necessary to develop a highly sensitive, robust and cellular-level resolution method for tracing the origin of organ damage based on cfDNA methylation, so as to accurately identify trace damage signals of target organs / cells and assist in clinical diagnosis and treatment. Summary of the invention
[0007] In order to achieve highly sensitive organ damage detection with cell-level precision, the present invention provides a relative quantitative method for organ damage based on cell methylation markers. By tracing the relative proportions of various cell sources in plasma cfDNA, the damage type and degree with cell-level precision can be accurately identified, which can assist in the health assessment of various organs, early screening and diagnosis of diseases, disease course monitoring and prognosis evaluation.
[0008] According to one aspect of the present disclosure, a method for quantifying cfDNA derived from target cells in a sample to be tested is provided, the method comprising the following steps:
[0009] 1) obtaining sequencing fragments of n methylation markers according to the methylation sequencing data of the sample to be tested, counting the total number of fragments and the number of positive fragments of each methylation marker, and screening the methylation markers with the total number of fragments ≥ z for subsequent steps, where z is an integer of 10 to 200, wherein the sample to be tested includes cfDNA derived from x types of cells, and the x types of cells include n methylation markers;
[0010] 2) Taking one cell in the sample to be tested as the target cell and the remaining x-1 cells as the background cells, the maximum likelihood MLE algorithm is used to obtain the relative content of cfDNA of each methylation marker corresponding to the target cell according to the probability prior of each methylation marker of the target cell releasing positive fragments, the total number of fragments ALL-counts and the number of positive fragments U-count of each methylation marker in the sample to be tested, and the weight of the noise signal contributed by the background cells to each methylation marker;
[0011] 3) using the total number of fragments of each methylation marker of the sample to be tested as a depth coefficient, performing a first correction on the relative content of cfDNA of the target cell corresponding to each methylation marker to obtain the relative content Fraction of cfDNA of the target cell in the sample to be tested;
[0012] Optionally, 4) repeating steps 2) and 3) to obtain the relative content of cfDNA of x types of target cells;
[0013] Wherein, the methylation marker includes a genomic region in which the ratio of unmethylated fragments in the target cell is higher or lower than that in the background cell,
[0014] Where x is an integer ≥ 1, n is an integer ≥ x,
[0015] Each cell type contains one or more methylation markers.
[0016] Those skilled in the art can select more different parameters as needed, all of which can be applied to the quantitative method described in the present disclosure. In some embodiments, the positive fragment is a fragment containing at least 3, at least 4, at least 5, or at least 6 CpG sites that are completely unmethylated.
[0017] In some embodiments, the weight of the noise signal contributed by the background cells to each methylation marker is obtained based on the probability prior of each methylation marker of each background cell releasing a positive fragment and the relative content of cfDNA of each background cell in the sample to be tested.
[0018] In some embodiments, the methylation marker of a cell includes a genomic region in which the positive fragment ratio of the cell is at least 0.2 greater than that of the background cell; the positive fragment ratio is the ratio of the number of positive fragments to the total fragments in the genomic region;
[0019] In some exemplary embodiments, each cell-specific cell population comprises at least 10 methylation markers.
[0020] In some embodiments, z includes but is not limited to 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195 or 200.
[0021] In some embodiments, the relative content of cfDNA of each background cell is obtained by a deconvolution algorithm based on the methylation sequencing data of the sample to be tested.
[0022] In some embodiments, after step 4), the method further comprises the following steps:
[0023] i) performing a second correction on the relative cfDNA content of each background cell according to the sum of the relative cfDNA contents of the x types of target cells, to obtain a corrected relative cfDNA content of each background cell; and
[0024] ii) Using the corrected relative cfDNA content of each background cell obtained in step ii), steps 2) to 3) are repeated to perform a third correction on the relative cfDNA content of the target cells in the sample to be tested.
[0025] In some embodiments, before step 1), the method further comprises the step of obtaining one or more methylation markers of each of the x types of cells.
[0026] In some embodiments, the step of obtaining one or more methylation markers for each of the x cells includes: the methylation markers of the cell include a genomic region in which the positive fragment ratio of the cell is at least 0.2 greater than that of the background cell; the positive fragment ratio is the ratio of the number of positive fragments in the genomic region to the total fragments.
[0027] In some embodiments, the methylation marker comprises a combined methylation marker, which is obtained by combining two methylation markers that are adjacent to each other on the same chromosome in the same cell and have a positive fragment ratio difference of <0.2.
[0028] In some embodiments, the merging includes accumulating the total number of fragments of the two methylation markers, and / or accumulating the number of positive fragments of the two methylation markers.
[0029] In some embodiments, the method further comprises the following steps: performing a fourth correction on the relative cfDNA content of the target cells of the sample to be tested using internal reference cells having a stable relative cfDNA content in samples from a healthy population.
[0030] In some embodiments, in step 2), according to the release of the positive fragments in cfDNA by the target cells, the maximum likelihood MLE algorithm is shown in Formula 2:
[0031] P k =binom(U-counts k , ALL-counts k , Fraction k ×U-ratio_target k +bdg k ) Formula 2;
[0032] In the formula 2:
[0033] k is an integer from 1 to n,
[0034] U-counts k is the number of positive fragments of the kth methylation marker;
[0035] ALL-counts k is the total number of fragments of the kth methylation marker;
[0036] bdg k The weight of the noise signal contributed by the background cells to the kth methylation marker;
[0037] U-ratio_target kis the probability prior of the kth methylation marker releasing positive fragments of the target cell obtained based on the cell sample of the target cell,
[0038] P k is the expected value, take P k Fraction with maximum value k is the relative content of cfDNA in the target cells corresponding to the kth methylation marker.
[0039] In some embodiments, the bdg k According to formula 1, we can get:
[0040]
[0041] In the formula 1:
[0042] k is an integer from 1 to n,
[0043] per i is the relative content of cfDNA of the i-th background cell among x-1 background cells in the sample to be tested,
[0044] i is an integer from 1 to (x-1);
[0045] U-ratio_bg ki is the probability prior of releasing positive fragments of the kth methylation marker of the i-th background cell obtained based on the cell sample of the i-th background cell.
[0046] In some embodiments, the per i It is obtained through the deconvolution algorithm.
[0047] In some embodiments, the target cell comprises t methylation markers, wherein t is an integer between 1 and n.
[0048] The first correction in step 3) is performed according to formula 3:
[0049]
[0050] In the formula 3,
[0051] k is an integer from 1 to t,
[0052] ALL-counts k is the total number of fragments of the kth methylation marker;
[0053] Fraction k is the relative content of cfDNA in the target cells corresponding to the kth methylation marker.
[0054] In some embodiments, the second correction in step i) comprises:
[0055] According to formula 4, the per i Correction was performed to obtain the corrected relative cfDNA content of each background cell per-adjust i ,
[0056] per-adjust i =per i ×sum formula 4;
[0057] Wherein i is an integer from 1 to (x-1).
[0058] In some embodiments, the third correction includes: based on the per-adjust i , Formulas 1, 2 and 3 are used for correction to obtain the corrected relative content of cfDNA in target cells, Fraction-adjust.
[0059] In some embodiments, the reference cell set is selected from any one or more cells of the x types of cells.
[0060] In some embodiments, the internal reference cells include a high-quality internal reference cell set and a suboptimal internal reference cell set.
[0061] In some embodiments, the high-quality internal reference cell set is one or more cells whose Fraction in ≥10 samples from healthy people is no more than 0.5%, whose lower limit of quantification LoQ is no higher than 0.2%, whose positive fragment ratio of methylation marker is higher than 0.5 and / or whose U-ratio_bg is lower than 0.1.
[0062] In some embodiments, the suboptimal internal reference cell set is one or more cells whose Fraction does not exceed 0.5%, the lower limit of quantification LoQ is not higher than 0.3%, the positive fragment ratio of methylation markers is higher than 0.3 and / or the U-ratio_bg is lower than 0.15 in ≥10 samples from healthy people.
[0063] In some embodiments, the high-quality internal reference cell set and the suboptimal internal reference cell set do not overlap with each other.
[0064] In some embodiments, the fourth correction includes:
[0065] a) counting the Fraction ratio of each high-quality internal reference cell in the sample to be tested, wherein the ratio = the Fraction of each high-quality internal reference cell / the sum of the Fractions of all high-quality internal reference cells;
[0066] b) if in the sample to be tested, the Fraction ratio of a high-quality internal reference cell is higher than the upper limit threshold of the high-quality internal reference cell ratio, then the high-quality internal reference cell is eliminated from the high-quality internal reference cell set, and / or the high-quality internal reference cells whose Fraction is higher than the absolute upper limit threshold of the Fraction are eliminated from the high-quality internal reference cell set,
[0067] Among them, the upper limit threshold of the high-quality internal reference cell ratio = the 85th to 99th percentile of the high-quality internal reference cell Fraction value in the healthy population sample / (the 85th to 99th percentile of the high-quality internal reference cell Fraction value in the healthy population sample + the 1st to 15th percentile of other high-quality internal reference cell Fraction values in the healthy population sample),
[0068] The absolute Fraction upper limit threshold is the 85th to 99th percentile of the Fraction of the high-quality internal reference cells in the healthy population sample;
[0069] c) If after step b), there are still m kinds of high-quality internal reference cells whose Fraction is above the lower limit threshold of absolute Fraction in the high-quality internal reference cell set, the target cell Fraction is corrected according to formula 5 to obtain the internal reference corrected target cell Fraction:
[0070]
[0071] In the formula 5,
[0072] j = 1 ~ m,
[0073] m is an integer from 1 to x,
[0074] Fraction of internal reference cells in healthy samples j is the 85th to 99th percentile value of the Fraciton of the jth internal reference cell obtained based on the healthy population sample in the healthy population sample,
[0075] Fraction of internal reference cells in the sample to be tested j is the Fraciton of the jth internal reference cell obtained based on the sample to be tested;
[0076] d) if after step b), there are still high-quality internal reference cells whose Fraciton is below the lower limit threshold of absolute Fraction in the high-quality internal reference cell set, a total of m types of internal reference cells including high-quality internal reference cells and suboptimal internal reference cells having a relative cfDNA content Fraction of less than 0.1%, less than 0.2%, less than 0.3%, less than 0.4%, less than 0.5%, less than 0.6%, less than 0.7%, less than 0.8%, less than 0.9% or less than 1% in the sample to be tested are used to correct the target cell Fraction according to formula 5 to obtain the internal reference corrected target cell Fraction;
[0077] e) if after step b), no high-quality internal reference cells exist, then the fourth calibration is performed without using the internal reference cells;
[0078] The absolute lower limit threshold is the 1st to 15th quantile of the high-quality internal reference cells in the Fraction of the healthy population samples.
[0079] In some embodiments, the fourth correction is performed after step 4), step i), or step ii).
[0080] In some embodiments, when one internal reference cell is a target cell, the internal reference cell is not used to perform the fourth calibration, and the fourth calibration is performed using other internal reference cells in the internal reference cell set.
[0081] In some embodiments, the healthy population sample includes at least 10 healthy human samples.
[0082] In some embodiments, the methylation sequencing includes one or more of whole genome methylation sequencing, multiplex PCR methylation sequencing and / or targeted methylation sequencing.
[0083] In some embodiments, the depth of the methylation sequencing is ≥400×, ≥500×, ≥600×, or ≥700×.
[0084] In some embodiments, the sample to be tested or the healthy population sample includes a body fluid sample.
[0085] In some embodiments, the body fluid sample includes: saliva or saliva supernatant, serum, plasma, milk, urine, lumbar or ventricular CSF, lymph fluid, prostatic fluid, semen, sputum, stool diluted supernatant, tears, alveolar lavage fluid, sputum, pus, nasopharyngeal swab, oral swab, cerebrospinal fluid, pleural effusion, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, vaginal discharge diluted supernatant and their processed products.
[0086] According to another aspect of the disclosure, a quantitative system for measuring the relative content of cfDNA in cells is provided, wherein the quantitative system is used to implement the quantitative method, and the system comprises:
[0087] The positive fragment acquisition module counts the positive fragments in the region of each methylation marker according to the methylation sequencing data; screens the methylation markers with a total fragment number U-counts ≥ z in the sample to be tested, where z is an integer from 10 to 200;
[0088] A background weight acquisition module, which accumulates and sums the probability of cfDNA derived from each background cell in the biological sample releasing positive fragments in the region of the target cell-specific methylation marker to obtain the background cell weight;
[0089] The module for obtaining the relative content of cfDNA of methylation markers constructs a likelihood formula to calculate the relative content of cfDNA of target cells corresponding to each methylation marker;
[0090] The target cell cfDNA relative content acquisition module uses the depth coefficient of each methylation marker as a weight to calculate the weighted average of all methylation markers specific to the same target cell as the relative cfDNA content of each target cell.
[0091] In some embodiments, the quantitative system also includes correction for unknown cell component ratios, calculating the sum of the relative cfDNA content Fraction of all target cells, and correcting the relative cfDNA content of each background cell in the background weight acquisition module.
[0092] In some embodiments, the system further comprises a methylation marker combination module: merging methylation markers with positive fragment ratio differences of the same target cells ≤ 0.2 and adjacent positions on chromosomes into a combined methylation marker.
[0093] According to another aspect of the disclosure, a method for tracing the source of organ damage is provided, the quantitative method comprising obtaining the relative content of target cell cfDNA in samples from a healthy population of ≥10 samples according to the quantitative method, and obtaining the 85th to 99th percentiles of the relative content of target cell cfDNA in the samples from the healthy population as a damage threshold, if the relative content of target cell cfDNA in the sample to be tested is ≥ the damage threshold, then it is judged that the target cell is damaged, the degree of damage is positively correlated with the relative content of target cell cfDNA in the sample to be tested, and the organ is further judged to be damaged based on the organ from which the target cell originates.
[0094] According to another aspect of the disclosure, a detection system for tracing the source of organ damage is provided. The detection system is used to implement the method for tracing the source of organ damage. The system comprises:
[0095] The damage threshold construction module uses the 85th to 99th percentile of the relative content of cfDNA in target cells of samples from a healthy population of ≥10 samples as the damage threshold;
[0096] Cell damage judgment module: when the relative content Fraction of target cell cfDNA in the sample to be tested is ≥ the damage threshold, it is judged that the target cell is damaged;
[0097] The organ damage judgment module further judges whether the organ is damaged according to the organ from which the target cells originate.
[0098] According to another aspect of the disclosure, a device is provided, comprising a memory for storing a program; and a processor for implementing the quantification method or the organ damage tracing method by executing the program stored in the memory.
[0099] According to another aspect of the disclosure, a computer-readable storage medium is provided, on which a program is stored, and the program can be executed to implement the quantitative method or the organ damage tracing method.
[0100] According to another aspect of the disclosure, provided is the application of the quantitative method or the organ damage tracing method in cell / organ tracing, cell / organ damage detection, cell / organ health assessment, early disease screening and diagnosis, disease course monitoring, recurrence prediction and prognosis assessment.
[0101] Beneficial effects:
[0102] The present invention provides a relative quantitative detection method for organ damage based on cell methylation markers, which improves the detection of positive signals by optimizing model parameters in depth; corrects the proportion of unknown cell components to eliminate or partially eliminate the deviation of the relative content of background cell cfDNA caused by the presence of unknown cell types; constructs an internal reference correction system, uses a group of internal reference cells with stable relative proportions and in the normal range of a healthy baseline as a benchmark, corrects the actual target cell signal to the baseline level to evaluate the real organ damage situation, and eliminates the deviation of the target cell signal in the actual sample caused by the high abundance background fluctuation. The quantitative method can truly restore the target cell signal, and through personalized dynamic internal reference correction, it can achieve high-sensitivity detection of trace damage signals in real clinical samples. By accurately analyzing the components in plasma cfDNA, early screening, early diagnosis, disease course monitoring and prognosis evaluation of various organ diseases can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1 Performance validation of simulations of important organs / cell types is shown.
[0104] Figure 2The relative quantitative component analysis components of leukocytes are shown.
[0105] Figure 3 The actual sample damage analysis process is shown.
[0106] Figure 4 The performance validation results for systemic lupus erythematosus (SLE) renal damage are shown.
[0107] Figure 5 The cell-level discriminatory performance of methylation markers is shown. DETAILED DESCRIPTION
[0108] The purpose of the present invention is to provide a relative quantitative detection method for organ damage based on methylation markers, which can accurately analyze the various components in cfDNA of body fluid samples (such as plasma) to achieve functions such as health assessment of various organs, early screening and diagnosis of diseases, disease course monitoring and prognosis evaluation.
[0109] The present invention provides a method for accurately tracing and quantifying each component in cfDNA of a body fluid sample (such as plasma) based on cell-specific methylation markers (methylation markers for short).
[0110] In some embodiments, the cell-specific methylation marker refers to a genomic region in which the proportion of unmethylated fragments in this type of cell is significantly higher than that in other types of cells, or refers to a genomic region in which the proportion of methylated fragments in this type of cell is significantly higher than that in other types of cells.
[0111] In some embodiments, the cell-specific methylation marker refers to a genomic region in which the proportion of unmethylated fragments in the cell type is significantly higher than that in other types of cells.
[0112] In an exemplary embodiment, the cell-specific methylation marker includes a genomic region in which the positive fragment ratio of the cell is at least 0.2, at least 0.3, at least 0.4 or at least 0.5 greater than that of other cells. The positive fragment ratio is the ratio of the number of positive fragments sequenced in the genomic region to the total fragments. In some embodiments, the positive fragment refers to a completely unmethylated fragment containing at least 3 CpG sites, at least 4, at least 5 or at least 6 CpG sites in the methylation marker region.
[0113] In some embodiments, the cell-specific methylation markers can be obtained according to conventional methods in the art, for example, methylation markers for a cell can be obtained based on conventional screening of genomic methylation sequencing data for a certain cell.
[0114] In some exemplary embodiments, the cells include, but are not limited to, one or more of aortic smooth muscle cells, cerebellar neuronal cells, coronary artery smooth muscle cells, cortical neuronal cells, endothelial cells, cardiac cardiomyocytes, cardiac fibroblasts, renal epithelial cells, hepatocytes, alveolar epithelial cells, pulmonary branch epithelial cells, oligodendrocytes, ovarian cells, pancreatic acinar cells, pancreatic Alpha cells, pancreatic Beta cells, pancreatic Delta cells, pancreatic duct cells, B cells, T cells, granulocytes, monocytes and macrophages, natural killer cells, osteoblasts, mammary epithelial cells, mammary basal cells, skin fibroblasts, smooth muscle cells, gallbladder cells, skin keratinocytes, endometrial epithelial cells, gastric epithelial cells, head and neck epithelial cells, smooth muscle cells, skeletal muscle cells, small intestinal epithelial cells, colon epithelial cells, prostate epithelial cells, tonsil cells, thyroid epithelial cells and bladder epithelial cells.
[0115] In some exemplary embodiments, the cells include, but are not limited to, one or more of aortic smooth muscle cells, cerebellar neuronal cells, coronary artery smooth muscle cells, cortical neuronal cells, endothelial cells, cardiac cardiomyocytes, cardiac fibroblasts, renal epithelial cells, hepatocytes, alveolar epithelial cells, pulmonary branch epithelial cells, oligodendrocytes, ovarian cells, pancreatic acinar cells, pancreatic Alpha cells, pancreatic Beta cells, pancreatic Delta cells and pancreatic duct cells.
[0116] In some embodiments, the method comprises:
[0117] S01: Acquisition of data of the sample to be tested.
[0118] The methylation data of the sample to be tested is obtained, and based on the methylation data, the number of positive fragments U-counts and the total number of fragments ALL-counts on each cell-specific methylation marker in the sample to be tested are counted.
[0119] In some embodiments, the cell-specific methylation markers with a total fragment number ALL-counts ≥ z in the sample to be tested are screened for subsequent steps, where z is an integer of 10-200.
[0120] In some embodiments, z includes but is not limited to 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195 or 200.
[0121] In some embodiments, the methylation sequencing includes whole genome methylation sequencing, multiplex PCR methylation sequencing and / or targeted methylation sequencing, etc.
[0122] In some embodiments, high-depth methylation sequencing is performed on the sample to be tested, and the depth of the high-depth methylation sequencing is, for example, ≥400×, ≥500×, ≥600×, or ≥700×.
[0123] In some embodiments, the sample to be tested includes a body fluid sample.
[0124] In some exemplary embodiments, the body fluid sample includes: saliva or saliva supernatant, serum, plasma, milk, urine, lumbar or ventricular CSF, lymph fluid, prostatic fluid, semen, sputum, stool diluted supernatant, tears, alveolar lavage fluid, sputum, pus, nasopharyngeal swab, oral swab, cerebrospinal fluid, pleural effusion, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, vaginal discharge diluted supernatant and their processed products.
[0125] In some embodiments, the methylation sequencing data of the sample to be tested includes methylation information of CpG sites.
[0126] In some embodiments, the methylation sequencing data of the sample to be tested includes methylation information of the CpG sites of the cell-specific methylation markers.
[0127] In some embodiments, the file format of the acquired methylation sequencing data of the sample to be tested is a fragment file containing CpG site related information.
[0128] S02, combination of methylation markers.
[0129] In order to improve the detection of positive signals, if there are two methylation markers in the test data of the sample to be tested that belong to the same cell, are adjacent in sorting position on the same chromosome and have a close positive fragment ratio (for example, the difference is <0.2), they are combined and used as a methylation marker to improve data utilization and detection rate. The combination includes:
[0130] The number of positive fragments U-counts observed for the two methylation markers is summed up to obtain the number of positive fragments U-counts for the combined methylation marker. joint ;
[0131] Sum the total number of fragments observed for the two methylation markers (ALL-counts k ≥z, z is an integer from 10 to 200), as the total number of fragments of the combined methylation marker ALL-counts joint .
[0132] For example, two renal cell methylation markers A and B with positive fragment ratios of 30 / 100 and 40 / 100 respectively can be combined into a renal cell methylation marker with a positive fragment ratio of 70 / 200.
[0133] In some embodiments, the positive fragment refers to a completely unmethylated fragment containing at least 3, at least 4, at least 5 or at least 6 CpG sites in the methylation marker region of the target cell.
[0134] In some embodiments, the positive fragment ratio is the ratio of the number of positive fragments U-counts in the target cell methylation marker region to the total fragments ALL-counts.
[0135] In some embodiments, step 2) is an optional step. If there are no two methylation markers belonging to the same cell, with adjacent sequencing positions on the same chromosome and close positive fragment ratios (e.g., difference <0.2) in the sample to be tested, step 2) may not be performed.
[0136] S03, accurately quantify background weights.
[0137] Considering the complex cell mixture in plasma cfDNA, the background cells contained in it generate noise that will interfere with the detection of target cells, especially for low-abundance target cells (kidney, heart, brain and other related cells). The positive signal of low-abundance target cells is released at a low frequency and is difficult to detect on the cfDNA noise base. In order to improve the positive detection of such trace cell signals, it is necessary to accurately quantify the weight of background cells for noise reduction optimization. By refining the weight of the noise signal released by each background cell, the noise coefficient bdg of each methylation marker of the sample to be tested is calculated.
[0138] In some embodiments, one or more cells are selected in turn as target cells, and the remaining cells are used as background cells.
[0139] In some embodiments, one type of cell is selected as the target cell in turn, and the remaining x-1 types of cells are selected as background cells.
[0140] In some embodiments, for the kth methylation marker among the n methylation markers of x cells (n is an integer ≥x, and x is an integer ≥1), the cell specifically corresponding to the kth methylation marker is the target cell, and the remaining x-1 cells are background cells. The decomposition formula is as follows:
[0141]
[0142] In the formula 1, k is an integer from 1 to n, and the noise coefficient bdg kU-ratio_bg is the weight of the noise signal contributed by x-1 background cells to the kth methylation marker in the sample to be tested, ki represents the probability prior that the i-th background cell releases a positive fragment on the k-th methylation marker region, where i is an integer from 1 to (x-1), and per i Represents the relative content of cfDNA in the i-th background cell in the sample to be tested.
[0143] In some embodiments, for per i The estimation is quantified by the UMX fragment-level deconvolution algorithm reported in the reference Loyfer N, Magenheim J, Peretz A, et al. A DNA methylation atlas of normal human cell types [J]. Nature, 2023, 613.
[0144] In some embodiments, the U-ratio_bg ki The positive fragment ratio is statistically analyzed based on the methylation sequencing data of nucleic acids in different cells, and then the probability prior of each type of background cell expected to release positive fragments in each target cell-specific methylation marker region is obtained.
[0145] S04, basic model construction.
[0146] In order to take into account the requirements of both high sensitivity and high robust quantitative performance of trace signals of target cells, the present invention uses the maximum likelihood MLE algorithm as the basis for constructing a relative quantitative model.
[0147] In some embodiments, based on the principle that the release of positive fragments from target cells satisfies the binomial distribution, a likelihood formula is constructed and a probability estimation is performed. Formula 2 is constructed for the kth methylation marker as follows:
[0148] P k =binom(U-counts k , ALL-counts k , Fraction k ×U-ratio_target k +bdg k ) Formula 2,
[0149] In the formula 2, k is an integer from 1 to n, P k is the expected value of the kth methylation marker, U-counts k is the number of positive fragments of the kth methylation marker of the sample to be tested, ALL-counts k is the total number of fragments of the kth methylation marker of the sample to be tested (ALL-countsk ≥100), U-ratio_target k bdg is the probability prior of the corresponding target cell releasing positive fragments in the kth methylation marker region. k is the noise coefficient obtained in step 2), and the expected value P is taken k Fraction at maximum k is the relative content of cfDNA in target cells corresponding to the kth methylation marker.
[0150] In some embodiments, the U-ratio_target k It is obtained by counting the positive fragment ratios of methylation sequencing samples of different types of target cells.
[0151] Based on the above model construction, in order to further improve the quantitative performance, subsequent iterative optimization is carried out to correct the Fraction.
[0152] S05, use the depth coefficient to correct and statistically obtain the target cell Fraction.
[0153] To further optimize the model parameters in depth, after obtaining the target cell Fraction of each methylation marker, the weight correction of the depth coefficient is then performed. The final quantitative result of the cfDNA relative content Fraction of the target cells containing t methylation markers in the sample to be tested is shown in Formula 3:
[0154]
[0155] The Fraction in Formula 3 k It is expressed as the relative content of cfDNA when the target cell expectation is maximized for each methylation marker, and the depth coefficient weight of each methylation marker is expressed as the total number of fragments ALL-counts k Indicates that k is an integer from 1 to t, and t is an integer from 1 to n.
[0156] S06. Optionally, perform correction for unknown cell component ratios.
[0157] The methylation marker set cannot exhaust all cell types in plasma cfDNA. The presence of unknown cell types will affect the correct estimation of known cell signals. In addition, the UXM deconvolution algorithm normalization process will allocate the unknown cell proportion to the known cell type, and the relative content of background cell cfDNA per iIt is easy to produce deviations, so it is necessary to consider the impact of unknown cell components in actual cfDNA and design a correction scheme. Considering that MLE is a non-normalized quantification and the problem of unknown cell ratio being allocated has little impact, MLE can be used to reversely correct the relative content of background cell cfDNA in the sample to be tested. i Correction of unknown cell component ratio is achieved. That is, the target cell Fraction is obtained through the MLE algorithm of S04 and S05, and the target cell Fractions of x cells are added to obtain the sum, and then the relative content of cfDNA of various background cells in the sample to be tested is further corrected as shown in Formula 4:
[0158] per-adjust i =per i ×sum formula 4,
[0159] In Formula 4, i is an integer from 1 to (x-1).
[0160] The relative content of cfDNA in the corrected background cells was then used per-adjust i Re-substitute into formula 1 to estimate the bdg weight and get bdg-adjust k , and then bdg-adjust k Substituting into formula 2, we get Fraction-adjust k , and finally Fraction-adjust k Substitute the relative content of target cell cfDNA in the sample to be tested into formula 3 to obtain Fraction-adjust k .
[0161] S07, construct an internal reference correction system to correct the relative content of cfDNA in target cells.
[0162] There is a large heterogeneity in the signal release of high-abundance cells (various blood cells, liver, and endothelial-related cells) between different individuals. The fluctuation of such high-abundance cells in the relative quantitative system will affect the true identification of low-abundance target cell damage signals in the actual sample; in addition, in the actual occurrence and development of various diseases, it is often accompanied by the occurrence of inflammatory reactions, and in most cases the level of immune cells will also be increased, which will cause trace amounts of target cell damage signals to be "squeezed", and even under the premise of accurate quantification, the actual observed signal will be lower than the true value. However, due to the heterogeneity of cfDNA content between different individuals, the performance of horizontal comparison between samples is often poor. Therefore, in order to restore the true signal abundance of target cells released in the actual sample cfDNA to the greatest extent, and to eliminate the influence of high-abundance background fluctuations on the deviation of target cell signals in the actual sample as much as possible, optionally, on the basis of the relative quantitative system, an internal reference correction system is also constructed, and the internal reference correction is used to achieve the true restoration of the target cell observation signal. That is, a group of internal reference cells with a relatively stable ratio and in the normal range of the healthy baseline are used as a benchmark, and the actual target cell signal is corrected to the baseline level to evaluate the real organ damage.
[0163] In some embodiments, constructing an internal reference correction system to correct the relative content of cfDNA in target cells comprises:
[0164] S07-01, healthy baseline cohort construction.
[0165] Using samples from a healthy population as a baseline cohort, extract cfDNA molecules from body fluid samples (such as plasma) of the baseline cohort, perform methylation sequencing on the target region based on the methylation marker region, and after obtaining the fragment methylation information of the healthy baseline cfDNA molecules, use a relative quantitative model to perform traceability quantitative analysis on the healthy baseline, and statistically calculate the normal reference range of the relative content of cfDNA in each target cell in the baseline cohort, which serves as a benchmark for subsequent internal reference selection and damage judgment.
[0166] In some embodiments, the methylation sequencing is high-depth methylation sequencing, and the depth of the high-depth methylation sequencing is, for example, ≥400×, ≥500×, ≥600×, or ≥700×.
[0167] In some embodiments, the normal reference range includes 1 to 15 percentiles - 85 to 99 percentiles.
[0168] S07-02, construction of internal reference cell cohort.
[0169] Based on the principles of stable Fraction of healthy population samples, satisfactory cell quantification accuracy and high specificity of methylation markers, a group of stable high-quality internal reference cells (referred to as high-quality internal reference cells) were selected from x types of cells for correction of the relative content of target cell cfDNA, and suboptimal internal reference cells with possible risks (referred to as suboptimal internal reference cells) were used as supplements.
[0170] In some embodiments, the high-quality internal reference cells are screened from the x types of cells according to the following standards: 1. Fraction stability of healthy population samples: the maximum Fraction or Fraction-adjust of target cells in healthy population samples does not exceed 0.5%; 2. Cell quantification accuracy: the lower limit of quantification LoQ of target cell simulation experiment is not higher than 0.2%; 3. Methylation marker selection has high specificity: the positive fragment ratio in the methylation marker region of the target cell is higher than 0.5 and the U-ratio_bg of each background cell is lower than 0.1.
[0171] In some embodiments, the suboptimal reference cells are screened from the x types of cells according to the following criteria: 1. Fraction stability of healthy population samples: the maximum Fraction or Fraction-adjust of target cells in healthy population samples does not exceed 0.5%; 2. Cell quantification accuracy: the lower limit of quantification LoQ of target cell simulation experiment is not higher than 0.3%; 3. High specificity of methylation marker selection: the positive fragment ratio in the methylation marker region of the target cell is higher than 0.3 and the U-ratio_bg of each background cell is lower than 0.15.
[0172] In some embodiments, the simulation experiment of the internal reference cell refers to the condition where the sequencing depth is ≥400×, ≥500×, ≥600× or ≥700×.
[0173] In some embodiments, the simulation experiment includes performing a relative quantification experiment according to the method described in the present disclosure using sequencing data of known relative cfDNA contents of different cells.
[0174] In some embodiments, the sequencing data of known relative cfDNA contents from different cells are from a wet experiment or a dry experiment.
[0175] In some embodiments, the wet experiment includes performing methylation sequencing using standard samples with known relative cfDNA content from different cells.
[0176] In some embodiments, the dry experiment includes mixing methylation sequencing data from different cells in a desired ratio.
[0177] In some exemplary embodiments, the simulation experiment is as shown in Example 2.
[0178] In some embodiments, the quantitative lower limit LoQ refers to the lowest known relative cfDNA content of the cell that can be relatively quantified using the method described in the present disclosure, under the condition that the difference between the relative content of cfDNA quantified using the method described in the present disclosure and the known relative content of cfDNA does not exceed 10%, and the CV between repeated samples is less than 20%.
[0179] For example, it is known that the relative cfDNA content of the gradient mixed samples of renal cells is 0.02%, 0.2%, 1.0%, and 1.5%, respectively, and the relative cfDNA content obtained using the method disclosed herein differs from the known relative cfDNA content by no more than 10%, and the CV between repeated samples is less than 20%. When the known relative cfDNA content is 0.02%, the relative cfDNA content obtained using the method disclosed herein differs from the known relative cfDNA content by more than 10%, or the CV between repeated samples is greater than or equal to 20%, then the quantitative lower limit LoQ of the renal cells is 0.2%.
[0180] S07-03, baseline internal reference value / threshold determination and use of internal reference cells to correct the relative content of cfDNA in target cells in the sample to be tested.
[0181] According to the baseline cohort of healthy population samples, the internal reference cell signal of each sample is constructed, and the relative content Fraction (or Fraction-adjust) of cfDNA of each internal reference cell of each healthy person is calculated, and the ratio between high-quality internal reference cells is preferentially used.
[0182] The high-quality internal reference cell ratio of each sample in the healthy population samples is calculated, and then the upper threshold of the high-quality internal reference cell ratio in the healthy population samples is calculated. When using the internal reference cells to calibrate the relative content of cfDNA in the sample to be tested, firstly, based on the ratio between the high-quality internal reference cells, the high-quality internal reference cells in the sample to be tested that are higher than the upper threshold of the ratio of the healthy population samples are eliminated. Such cells are mostly considered to be damaged internal references and are not suitable as internal references.
[0183] In addition, an absolute threshold is constructed to further determine the degree of damage to the internal reference. The absolute upper threshold uses the 85th to 99th percentile threshold of each type of high-quality internal reference cells in the baseline Fraction (or Fraction-adjust). If the high-quality internal reference cells of the sample to be tested are higher than the absolute upper threshold, the internal reference cells are determined to be damaged in combination with the ratio index and need to be eliminated; the absolute lower threshold uses the 1st to 15th percentile level of each type of high-quality internal reference cells in the baseline Fraction (or Fraction-adjust). If after elimination, there are high-quality internal reference cells in the sample to be tested that are between the absolute upper threshold and If the target cell Fraction (or Fraction-adjust) value is between the absolute lower limit and the absolute lower limit, the high-quality internal reference cells can be used to correct the target cell Fraction (or Fraction-adjust) value; if there are high-quality internal reference cells after elimination, but the actual quantitative result is lower than the lower limit, it is considered that the high-abundance background squeezes out the internal reference abundance, or the internal reference abundance is actually low, making the internal reference cell quantitative result lower than the quantitative performance. All internal references at the same level (including high-quality internal references and suboptimal internal references) are used to correct the target cell Fraction (or Fraction-adjust) value; however, there are risks in using this type of internal reference correction. If there is no high-quality internal reference in the sample to be tested after elimination, no correction is performed.
[0184] In some embodiments, being at the same level includes a relative cfDNA content of less than 0.1%, less than 0.2%, less than 0.3%, less than 0.4%, less than 0.5%, less than 0.6%, less than 0.7%, less than 0.8%, less than 0.9% or less than 1%.
[0185] In some embodiments, the ratio of each high-quality internal reference cell in each sample is calculated as follows: ratio of each high-quality internal reference cell = Fraction (or Fraction-adjust) of the high-quality internal reference cell / sum of Fraction (or Fraction-adjust) of all high-quality internal reference cells.
[0186] In some embodiments, the upper limit threshold of the ratio of each high-quality internal reference cell = the baseline 85th to 99th percentile of the Fraction (or Fraction-adjust) value of the high-quality internal reference cell / (the baseline 85th to 99th percentile of the Fraction (or Fraction-adjust) value of the high-quality internal reference cell + the baseline 1st to 15th percentile of the Fraction (or Fraction-adjust) value of the remaining types of high-quality internal reference cells).
[0187] In some embodiments, since the internal reference cells also belong to x types of cells, when each type of internal reference cell is used as a target cell for calibration, the remaining types of internal reference cells are used as internal reference cells for calibration. For example, for 3 types of high-quality internal reference cells and 5 types of suboptimal internal reference cells, when any of the high-quality internal reference cells is used as a target cell for calibration, the remaining 2 types of high-quality internal reference cells and 5 types of suboptimal internal reference cells are used as internal reference cells for calibration; when one of the suboptimal internal reference cells is used as a target cell for calibration, the remaining 3 types of high-quality internal reference cells and the remaining 4 types of suboptimal internal reference cells are used as internal reference cells for calibration.
[0188] In some embodiments, for the correction of various types of high-abundance cells, considering that the relative level Fraciton (or Fraction-adjust) of cfDNA in body fluid samples is much higher than that of internal reference cells (at least 1-3 orders of magnitude), the correction risk is relatively high. Therefore, for such high-abundance cells, internal reference cells are not used for correction, and the test samples and healthy population samples only use the relative cfDNA content Fraction (or Fraction-adjust) obtained by S05 or S06 as the final cell Fraction.
[0189] In some exemplary embodiments, the high-abundance cells include, but are not limited to, blood cells, endothelial cells, hepatocytes, and the like.
[0190] In some specific embodiments, the use of internal reference cells in S07-03 to calibrate the relative content of cfDNA in target cells in the sample to be tested comprises the following steps:
[0191] ① Eliminate the high-quality internal reference cells in the sample to be tested whose ratio is higher than the upper limit threshold of the high-quality internal reference cells, and further eliminate the internal reference cells whose ratio is higher than the absolute upper limit threshold, as these internal references are considered to be damaged;
[0192] ② If after step ①, the remaining internal reference cells contain the high-quality internal reference and the Fraction (or Fraction-adjust) of at least one high-quality internal reference cell is above the absolute lower limit threshold, these high-quality internal references are used for correction, and the correction formula is:
[0193]
[0194] The target cell Fraciton (or Fraction-adjust) in Formula 5 is the relative content of target cell cfDNA in the sample to be tested obtained in S05 or S06, and the internal reference cell Fraction of the healthy population sample is j (or Fraction-adjust j) is the 85th to 99th percentile of the relative content of cfDNA of the jth reference cell in the healthy population sample. j (or Fraction-adjust j ) is the relative content of cfDNA of the jth reference cell in the actual sample, j is an integer from 1 to m, and m is the number of reference cell types involved in the calibration;
[0195] ③ If after step ①, the remaining internal reference cells contain high-quality internal reference cells but the relative cfDNA content Fraciton (or Fraction-adjust) is below the absolute lower limit threshold, it is considered that the high-abundance background cells in the sample are severely "squeezed", so that the actual quantification of target cells cannot meet the accuracy requirements. For this type of situation, use all internal references (including high-quality internal references and suboptimal internal references) with a relative cfDNA content of less than 0.1%, less than 0.2%, less than 0.3%, less than 0.4%, less than 0.5%, less than 0.6%, less than 0.7%, less than 0.8%, less than 0.9% or less than 1% to perform correction according to formula 5 to minimize the correction deviation;
[0196] ④If there are no high-quality internal reference cells among the remaining internal reference cells after step ① elimination, it is considered that the sample has multiple organ damage. In this case, the internal reference ratio value is distorted and cannot be further corrected based on the internal reference ratio value. Therefore, both the test sample and the healthy population sample use the Fraction (or Fraction-adjust) obtained in S05 or S06 as the relative content of cfDNA in the final target cells.
[0197] The present invention also provides an organ damage tracing detection method, which includes obtaining the relative content Fraction of target cell cfDNA in samples from a healthy population of ≥10 samples according to a quantitative method for the relative content of cfDNA in the cells, and statistically obtaining the 85th to 99th percentiles of the target cell Fraction in the samples from the healthy population as a damage threshold. If the relative content Fraction of target cell cfDNA in the sample to be tested is ≥ the damage threshold, it is judged that the target cell is damaged, and the degree of damage is positively correlated with the Fraction value of the sample to be tested. The organ is further judged to be damaged based on the organ from which the target cell originates.
[0198] definition
[0199] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly used in the field to which the present invention belongs. For the purpose of interpreting this specification, the following definitions will apply, and where appropriate, terms used in the singular will also include the plural form, and vice versa.
[0200] As used herein, the articles "a," "an," and "an" include plural referents unless the context clearly dictates otherwise.
[0201] The expression "about" as used herein is understood by one of ordinary skill in the art and varies within a certain range depending on the context in which it is used. If one of ordinary skill in the art is not aware of the use of the term based on the context in which it is used, "about" will mean up to plus or minus 10% of the specified value.
[0202] The term "cfDNA" herein refers to cell-free DNA, including DNA molecules naturally present in an extracellular form in a subject (e.g., in blood, serum, plasma, or other body fluids such as lymph, cerebrospinal fluid, urine, or sputum). Although cfDNA is originally derived from one or more cells in a large complex biological organism (e.g., mammal), cfDNA undergoes release from the cells into the fluid present in the organism and can therefore be obtained by obtaining a sample of the fluid without the need for an in vitro cell lysis step.
[0203] The term "methylation" herein refers to the presence of methyl in the position where there is no methyl in the generally recognized typical nucleotide base on the nucleotide base. For example, cytosine does not contain a methyl moiety on its pyrimidine ring, but 5-methylcytosine contains a methyl moiety at the 5th position of its pyrimidine ring. In this case, cytosine is not a methylated nucleotide, and 5-methylcytosine is a methylated nucleotide. In another example, thymine contains a methyl moiety at the 5th position of its pyrimidine ring. In this article, when thymine is present in DNA, it is not considered to be a methylated nucleotide because thymine is a typical nucleotide base of DNA. The typical nucleoside bases of DNA are thymine, adenine, cytosine and guanine. The typical bases of RNA are uracil, adenine, cytosine and guanine. Accordingly, "methylation site" is a position where methylation occurs or may occur in the target gene nucleic acid region. For example, a position containing CpG is a methylation site, wherein cytosine may be methylated or may not be methylated. A "site" may correspond to a single site, which is a single base position or a group of related base positions, for example, a CpG site. A methylation site may refer to a CpG site or a non-CpG site of a DNA molecule that is likely to be methylated.
[0204] The term "CpG site" as used herein is a region of a DNA molecule in which a cytosine nucleotide is followed by a guanine nucleotide in a linear sequence of bases along its 5' to 3' direction, which region is methylated by events occurring naturally in vivo or by chemical action performed in vitro. A non-CpG site is a region that does not have a CpG dinucleotide sequence but is also susceptible to methylation by events occurring naturally in vivo or by chemical action performed in vitro.
[0205] The term "methylation sequencing" herein includes whole genome bisulfite sequencing (WGBS), whole genome enzymatic methyl sequencing, whole epigenome sequencing, methylation arrays, degenerate representation bisulfite sequencing (RRBS-Seq), TET-assisted pyridine borane sequencing (TAPS), Tet-assisted bisulfite sequencing (TAB-Seq), APOBEC-coupled epigenetic sequencing (ACE-seq), oxidative bisulfite sequencing (oxBS-Seq), sedimentation or methylated DNA immunoprecipitation sequencing, or cytosine 5-hydroxymethylation sequencing (e.g., via Bluestar).
[0206] The term "maximum likelihood MLE algorithm" used in this article is a statistical method used to find the parameters of the probability density function of a sample set. It is known that a random sample satisfies a certain probability distribution, but the specific parameters are unknown. Parameter estimation is to observe the results through several experiments and use the results to deduce the approximate value of the parameter. Maximum likelihood estimation is based on the assumption that if a certain parameter is known to maximize the probability of the occurrence of this sample, then of course other samples with small probabilities will no longer be selected, so the parameter is selected as the true value of the estimate.
[0207] The term "abundance" used in this article refers to the proportion of a certain type of cell in a sample, including target cells, background cells or internal reference cells.
[0208] As used herein, the term "sequencing depth" refers to the number of times a locus is covered by a consensus sequence read corresponding to a unique nucleic acid target molecule ("nucleic acid fragment") aligned to the locus.
[0209] In the present application, the term "methylation marker" refers to a parameter associated with one or more biological molecules (ie, a "gene methylation marker"), such as nucleic acids produced naturally or artificially.
[0210] As used herein, the term "sequencing" refers to the process of determining the sequence (eg, the identity and order of monomeric units) of a biological molecule, eg, a nucleic acid, such as DNA or RNA. Exemplary sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy terminator sequencing, whole genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single base extension sequencing, solid phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, low denaturation temperature co-amplification PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, DNA nanoball sequencing (DNBSEQ), composite probe anchor polymerase sequencing (cPAS), and combinations thereof. In some embodiments, sequencing can be performed by a genetic analyzer, such as a genetic analyzer commercially available from Illumina, Inc., Pacific Biosciences, Inc., Applied Biosystems / Thermo Fisher Scientific, or Shenzhen MGI Intelligent Manufacturing Technology Co., Ltd., etc. For example, BGI DNBseq sequencing platforms such as BGISEQ-500, BGISEQ-50, MGISEQ-2000, MGISEQ-200, DNBSEQ-T7, DNBSEQ-G99, DNBSEQ-T20X2, or Illumina's HiSeq2000, HiSeq2500, HiSeq4000, HiSeqX10, NovaSeq6000, etc.
[0211] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the examples. The specific embodiments described herein are only used to explain the present invention and are not intended to constitute any limitation to the present invention. The actual protection scope of the present invention is set forth in the claims. In the following description, the description of known structures and technologies is omitted to avoid unnecessary confusion of the concepts of the present disclosure. Such structures and technologies are also described in many publications. The equipment, instruments, reagents and / or kits used in the following examples are not mentioned in terms of their sources, and are all commercially available on the market or obtained by conventional methods known to those skilled in the art.
[0212] Example 1: Relative Quantification Method
[0213] 1. Methylation marker screening.
[0214] The cell methylation data were obtained from various cell whole genome methylation sequencing data (Whole-Genome Bisulfite Sequencing, WGBS) disclosed in (GSE186458), involving a total of 41 cell types from 28 organs, as shown in Table 1.
[0215] Table 1. 28 organs and 41 cell types
[0216]
[0217]
[0218] The cell-specific methylation markers were screened as follows. 1. Methylation region segmentation: Using the whole genome methylation sequencing data of various cells disclosed in the GSE186458 dataset as input, the wgbstools segment tool was used to segment the methylation regions in the whole genome for the CpG sites in the whole genome, and further extracting regions containing at least 3 CpGs and with a distance between adjacent CpG sites <200bp as the target regions for subsequent methylation marker screening; 2. Initial screening of methylation markers based on methylation rate: Using any of the 41 cell types as target cells and the rest as background cells, wgbstools was used to screen the methylation markers in the whole genome. The find_markers tool was used to screen the methylation markers of 41 types of cells based on the methylation rate index in the target screening area. The screening criteria required that the average methylation rate of target cells was ≤0.33, the average methylation rate of background cell population was ≥0.66, and the difference between the methylation rate 2.5 percentile value of all samples of the background cell population and the methylation rate 75 percentile value of all samples of the target cells was greater than 0.3, and the primary screening methylation markers based on the methylation rate were obtained; 3. Fine screening of all non-methylated fragments: repeated samples of the same type of cells were combined to increase the depth, and the fragments containing at least 3 CpGs in the methylation marker region of the previous step methylation primary screening and being fully non-methylated were taken as positive fragments, and U-counts were calculated. All fragments containing at least 3 CpGs in the methylation marker region of the previous step methylation primary screening were taken as total fragments, and ALL-counts were calculated. The ratio obtained by the positive fragment ratio of each cell in each methylation primary screening methylation marker region was used as the main indicator for subsequent fine screening of fragments. Any one of the 41 types of cells was used as the target cell, and the rest were used as background cells. The difference in the positive fragment ratio between the target cell and any background cell was required to be at least greater than 0.2, and the target cell positive fragment ratio was higher than 0.3, while the positive fragment ratio of each background cell was lower than 0.2. In addition, when various types of high-abundance blood cells were used as background cells, their positive fragment ratio was required to be lower than 0.1 to further reduce the background noise caused by high-abundance cells, and finally 41 types of cell methylation marker regions were obtained after fine screening of fragments.
[0219] Target cell-specific methylation markers were screened for each target cell, and a total of 1466 methylation markers for 41 target cell types in 28 organs were screened. Figure 5 The positive fragment refers to a completely unmethylated fragment containing at least 3 CpG sites in the target cell methylation marker region. The positive fragment ratio is the ratio of the number of positive fragments U-counts in the target cell methylation marker region to the total fragments ALL-counts.
[0220] 2. Data acquisition of samples to be tested
[0221] Acquire the methylation sequencing data of the sample to be tested, and screen the methylation markers with a total fragment number U-counts ≥ 100 in the sample to be tested for subsequent steps.
[0222] 3. If there are two methylation markers in the sample to be tested that belong to the same target cell, are adjacent in sequence position on the same chromosome, and have a positive fragment ratio difference of <0.2, they are combined and used as a methylation marker;
[0223] 4. Select one type of cell as the target cell in turn, and the remaining 40 types of cells as the background cells. According to the formula 1, the noise coefficient bdg of each methylation marker is calculated for the kth methylation marker of the total 1466 methylation markers, where n=1466, i is an integer from 1 to 40, and the U-ratio_bg ki By performing statistics on the positive fragment ratios of the methylation sequencing data of various cells in the GSE186458 dataset, we can obtain the probability that each background cell will release positive fragments in each target cell methylation marker region.
[0224] 5. For each methylation marker, construct the formula 2 to calculate the Fraciton corresponding to each methylation marker in the sample to be tested, where k is an integer from 1 to 1466, and the U-ratio_target k It is obtained by statistically analyzing the positive fragment ratios of WGBS sequencing samples of various cells in the GSE186458 database, and then obtaining the probability of each target cell releasing positive fragments in each methylation marker region.
[0225] 6. Perform a deep correction on the relative content Fraction of cfDNA of the 41 target cells according to Formula 3.
[0226] 7. According to the formula 4, the relative content of cfDNA in various background cells is corrected, and the relative content of cfDNA in various background cells after correction is used. per-adjust i Resubstitute into the formula 1 to estimate the bdg weight and get bdg-adjust k , and then bdg-adjust k Substituting into the above formula 2, we get Fraction-adjust k , and finally Fraction-adjust k Substitute into formula 3 to obtain the relative content of target cell cfDNA Fraction-adjust.
[0227] 8. Internal reference correction to achieve target cell signal restoration
[0228] 8.1 Construction of a healthy baseline cohort. Twenty-eight healthy subjects were selected as the baseline cohort. cfDNA molecules were extracted from plasma samples of the baseline cohort. Probes were designed based on 1466 methylation marker regions. The target regions were subjected to high-depth bisulfite sequencing (BS) (>500×) using the bisulfite conversion method. After obtaining the fragment methylation information of the healthy baseline cfDNA molecules, the healthy baseline was traced and quantitatively analyzed using a relative quantitative model. The normal reference range (5th percentile-95th percentile) of the relative content of cfDNA in each target cell was calculated and used as the benchmark for subsequent internal reference selection and damage judgment.
[0229] 8.2 Construction of internal reference cell queue: A group of stable high-quality internal reference cells were selected from 41 cell types for target cell signal correction, and suboptimal internal reference cells that may have risks were used as supplements.
[0230] The high-quality internal reference cells were screened from the 41 cells according to the following standards: 1. Fraction stability of healthy population samples: the maximum Fraction-adjust of target cells in healthy population samples does not exceed 0.5%; 2. Cell quantitative accuracy: the lower limit of quantitative simulation experiment LoQ of target cells at sequencing depth 500× is not higher than 0.2%; 3. High specificity of methylation marker selection: the positive fragment ratio in the methylation marker region of target cells is higher than 0.5 and the U-ratio_bg of each background cell is lower than 0.1. According to the test results of 28 healthy human plasma, three types of high-quality internal reference cells were screened according to the above standards: cerebral cortical neurons, cardiomyocytes, and oligodendrocytes.
[0231] The suboptimal internal reference cells were screened from the 41 cells according to the following standards: 1. Fraction stability of healthy population samples: the maximum Fraction-adjust of target cells in healthy population samples does not exceed 0.5%; 2. Cell quantitative accuracy: the target cells are not higher than 0.3% in the simulated experimental quantitative lower limit LoQ at a sequencing depth of 500×; 3. Methylation markers are selected with high specificity: the positive fragment ratio in the target cell methylation marker region is higher than 0.3 and the U-ratio_bg of each background cell is lower than 0.15. According to the test results of 28 healthy human plasma, 5 types of suboptimal internal reference cells were screened according to the above standards for supplementary use: alveolar epithelial cells, lung bronchial epithelial cells, pancreatic duct cells, pancreatic Alpha cells, and gallbladder cells.
[0232] 8.3 Determination of baseline internal reference ratio / threshold. Based on the baseline cohort of 28 healthy subjects, the internal reference cell signal of each sample was constructed, and the relative content Fraction-adjust of each internal reference cell of each healthy subject was calculated. The high-quality internal reference cell ratio and the upper threshold of the internal reference cell ratio were further determined; the upper threshold of each high-quality internal reference cell ratio = the baseline 95th percentile of the Fraction (or Fraction-adjust) value of the high-quality internal reference cell / (the baseline 95th percentile of the Fraction (or Fraction-adjust) value of the high-quality internal reference cell + the baseline 5th percentile of the Fraction (or Fraction-adjust) value of the remaining types of high-quality internal reference cells). The absolute threshold was constructed to further determine the degree of internal reference damage. The absolute upper threshold used the 95th percentile of each of the three internal reference cells in the baseline Fraction-adjust. The absolute lower threshold used the 5th percentile of each of the three internal reference cells in the baseline Fraction-adjust. The calculation results are shown in Table 2.
[0233] Table 2. Baseline internal reference value threshold and absolute threshold
[0234]
[0235] 8.4 Use internal reference cells to calibrate the relative content of cfDNA in target cells in the sample to be tested, including:
[0236] 8.4.1 Eliminate high-quality internal reference cells whose ratios are higher than the upper limit threshold of high-quality internal reference cells, and eliminate internal reference cells whose ratios are higher than the absolute upper limit threshold;
[0237] 8.4.2 If there are high-quality internal references in the remaining cells and the signal Fraction-adjust of at least one high-quality internal reference cell is above the absolute lower threshold, correction is performed using these high-quality internal references according to Formula 5;
[0238] 8.4.3 If there are high-quality internal reference cells in the remaining components but the Fraction-adjust signals are all below the absolute lower threshold, all internal references (including high-quality internal references and suboptimal internal references) with Fraction-adjust below 1% are used for correction according to Formula 5;
[0239] 8.4.4 If there are no high-quality internal reference cells in the remaining components, no internal reference cells are used for correction, and the Fraction-adjust obtained in step 6 is used for both the test sample and the healthy population sample as the relative cfDNA content of the final target cells;
[0240] 8.4.5 When the internal reference cells are used as target cells for calibration, first remove the cells themselves as internal references, and then use the remaining high-quality internal reference cells and suboptimal internal reference cells for calibration according to steps 8.4.1 to 8.4.4.
[0241] 8.4.6 For high-abundance cells such as blood cells, endothelial cells, and hepatocytes whose relative cfDNA levels (Fraction-adjust) are much higher than those of internal reference cells (approximately 1-3 orders of magnitude), internal reference cells are not used for correction. The Fraction-adjust obtained in step 6 is used as the relative cfDNA content of the final target cells for both the test samples and samples from healthy subjects.
[0242] Example 2: Relative Quantification Method Performance Verification
[0243] Based on the publicly published WGBS sequencing data of various cells, in order to simulate the real cfDNA situation, the methylation data of various cell fragments were mixed in the expected proportion. Under the premise of knowing the mixing ratio of each cell, the performance of the relative quantitative method was verified.
[0244] Specific mixing simulation method. Except for blood cells, the remaining cell gradients are mixed in equal proportions, and the mixing gradients are set to: 0.02%, 0.04%, 0.06%, 0.08%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1.0%, 1.1%, 1.2%, 1.3%, 1.4%, 1.5%, a total of 19 gradients, and the remaining cfDNA is distributed to various types of blood cells (B cells, T cells, monocytes and macrophages, granulocytes, erythrocytes, etc.); two gradients of 500× and 750× are designed in depth, with 20 repetitions for each gradient.
[0245] Quantitative performance verification. Linear fit the true mixing ratio (mix rate) and the actual relative quantitative results, statistical slope, intercept, R2, CV (variation between repeated mixing samples of each gradient) and other evaluation indicators. The performance of the qualified gradient is defined as: the difference between the true mixing ratio (mix rate) and the actual relative quantitative result does not exceed 10%, and the CV between repeated samples is less than 20%. The quantitative performance results of important cell types are shown in Figure 1 The quantitative lower limit LoQ of the gradient that meets the performance qualification is shown in Table 3.
[0246] The lower limit of target cell quantification at 750× depth was between 0.02% and 0.3%, and most components could reach 0.1%. The slopes of important organ cell types were all above 0.97, and R 2 The quantification limit of most target cells at 500× depth was between 0.1% and 0.5%, and the slopes of important organ cell types were all above 0.94.2 The above performances prove the rationality of the method model, and the accurate detection of trace target cell signals (ten thousandths to thousandths) can be achieved at about 500×.
[0247] Table 3 Performance of relative quantitative detection method for organ damage
[0248]
[0249]
[0250] Further verification method under the premise of the presence of high-abundance unknown cell components, can still achieve accurate quantification of known cell types, this embodiment uses the simulation of high-abundance natural killer cells (Natural Killer Cell), mononuclear macrophages (Mono-Macrophages Cell) component ratio unknown, low-abundance target cells quantitative accuracy test. The specific operation is: remove the two types of simulated unknown cell types in the probe group related methylation markers, and at the same time mix two cell fragments according to the expected proportion, and the remaining simulation specific implementation method is compared with the method without step 2). The results are shown in Table 4. After the unknown component ratio is corrected at 750× depth, the lower limit of target cell quantification is between 0.04%-0.3%; after the unknown component ratio is corrected at 500× depth, the lower limit of most target cells is between 0.1%-0.6%. It shows that the quantitative performance of target cells has not decreased significantly after noise reduction optimization, and verifies that this method can still accurately detect trace signals of target cells under the premise of the presence of unknown components.
[0251] Table 4 Performance comparison before and after unknown component ratio correction
[0252]
[0253] Based on the above relative quantitative system, the relative proportion range of each important cell type in the healthy baseline was further analyzed, as shown in Table 5. The relative proportions of most solid organ cells are mostly in the ten-thousandths to thousandths. In addition, the relative proportions of liver and endothelial cells are between 1% and 5% in most samples, which is about 1-2 orders of magnitude higher than the other solid organ / cell signals. To further analyze the white blood cell components, the 23 white blood cell WGBS sequencing data downloaded from the public open source were merged for quantitative analysis, as shown in Table 5. Figure 2 The main cell types and their proportions are as follows: granulocytes (52.8%), lymphocytes (31.7%), and monocytes (8.6%), which are consistent with the results of routine blood analysis. The relative proportion of healthy baseline white blood cells fluctuates greatly, with large differences between different individuals, and may show further differences during the development of the disease.
[0254] Table 5. Injury threshold corrected by healthy baseline internal reference (95th percentile)
[0255]
[0256] Example 3: Internal reference calibration and performance verification of real samples of renal injury
[0257] In this example, real plasma cfDNA data was used for internal reference correction signal restoration and clinical performance verification. The target organ was selected as kidney and the target cell was selected as renal epithelial cell. In this example, 13 cases of systemic lupus erythematosus (SLE) data were selected for renal cell signal internal reference correction and clinical performance verification. The analysis process is as follows: Figure 3 .
[0258] Healthy baseline construction. The target probe group region of the healthy baseline cohort was subjected to high-depth sequencing (>500×) of methylation (BS). After data preprocessing and relative quantification, the internal reference calibration results of healthy baseline renal epithelial cells were obtained, such as Figure 4 As shown in the figure, the robustness of the baseline correction signal of the renal epithelium was excellent, and most samples were stable between 0.1% and 0.3%. The 95th percentile injury threshold of the renal epithelium (0.26%) was then used to analyze the clinical performance of the injury samples.
[0259] Clinical performance analysis of renal injury samples. All samples were divided into three groups: SLE active stage (7 cases), inactive stage (6 cases) and healthy baseline. Figure 4 As shown. The results of internal reference correction quantitative analysis showed that under the baseline specificity of 95%, SLE active samples could achieve 100% sensitivity detection (7 / 7), and under the specificity of 95% in the inactive period, SLE active samples could achieve 85.7% sensitivity detection (6 / 7), indicating the high sensitivity of the organ damage quantitative method in the application scenario of renal damage; for inactive samples, the damage signal decreased significantly compared with the active samples, but compared with the healthy baseline level, it still showed a certain upward trend, indicating that further management and monitoring are still needed for inactive cases. The overall validation of the relative quantitative detection method can provide an important reference for the diagnosis and treatment of renal damage.
[0260] The technical solution of the present invention is not limited to the above-mentioned specific embodiments. All technical variations made according to the technical solution of the present invention fall within the protection scope of the present invention.
Claims
1. A method for quantifying cfDNA derived from target cells in a sample to be tested, the method comprising the following steps: 1) obtaining sequencing fragments of n methylation markers according to the methylation sequencing data of the sample to be tested, counting the total number of fragments and the number of positive fragments of each methylation marker, and screening the methylation markers with the total number of fragments ≥ z for subsequent steps, where z is an integer of 10 to 200, wherein the sample to be tested includes cfDNA derived from x types of cells, and the x types of cells include n methylation markers; 2) Taking one cell in the sample to be tested as the target cell and the remaining x-1 cells as the background cells, the maximum likelihood MLE algorithm is used to obtain the relative content of cfDNA of each methylation marker corresponding to the target cell according to the probability prior of each methylation marker of the target cell releasing positive fragments, the total number of fragments ALL-counts and the number of positive fragments U-count of each methylation marker in the sample to be tested, and the weight of the noise signal contributed by the background cells to each methylation marker; 3) using the total number of fragments of each methylation marker of the sample to be tested as a depth coefficient, performing a first correction on the relative content of cfDNA of the target cell corresponding to each methylation marker to obtain the relative content Fraction of cfDNA of the target cell in the sample to be tested; Optionally, 4) repeating steps 2) and 3) to obtain the relative content of cfDNA of x types of target cells; Wherein, the methylation marker includes a genomic region in which the ratio of unmethylated fragments in the target cell is significantly higher or lower than that in the background cell, Where x is an integer ≥ 1, n is an integer ≥ x, Each cell type contains one or more methylation markers.
2. The method according to claim 1, characterized in that The positive fragments are fragments containing at least 3, at least 4, at least 5, or at least 6 completely unmethylated CpG sites.
3. The method according to claim 1 or 2, characterized in that: The weight of the noise signal contributed by the background cells to each methylation marker is obtained based on the probability prior of each methylation marker of each background cell releasing a positive fragment and the relative content of cfDNA of each background cell in the sample to be tested. Preferably, the relative content of cfDNA of each background cell is obtained by a deconvolution algorithm based on the methylation sequencing data of the sample to be tested.
4. The method according to any one of claims 1 to 3, characterized in that After step 4), the method further comprises the following steps: i) performing a second correction on the relative cfDNA content of each background cell according to the sum of the relative cfDNA contents of the x types of target cells, to obtain a corrected relative cfDNA content of each background cell; and ii) Using the corrected relative cfDNA content of each background cell obtained in step ii), steps 2) to 3) are repeated to perform a third correction on the relative cfDNA content of the target cells in the sample to be tested.
5. The quantitative method according to any one of claims 1 to 4, characterized in that Before step 1), the method further comprises the step of obtaining one or more methylation markers of each of the x types of cells, Preferably, the step of obtaining one or more methylation markers of each of the x types of cells comprises: The methylation marker of the cell includes a genomic region in which the positive fragment ratio of the cell is at least 0.2 greater than that of the background cell; the positive fragment ratio is the ratio of the number of positive fragments to the total fragments in the genomic region; Preferably, the methylation marker comprises a combined methylation marker, which is obtained by combining two methylation markers that are adjacent to each other on the same chromosome in the same cell and have a positive fragment ratio difference of <0.
2. Preferably, the merging comprises accumulating the total number of fragments of the two methylation markers, and / or accumulating the number of positive fragments of the two methylation markers.
6. The quantitative method according to any one of claims 1 to 5, characterized in that: The method further comprises the following steps: Using internal reference cells with a stable relative cfDNA content in samples from a healthy population, a fourth correction is performed on the relative cfDNA content of the target cells of the sample to be tested.
7. The quantitative method according to any one of claims 1 to 6, characterized in that According to the rule that the release of the positive fragments by the target cells in cfDNA satisfies the binomial distribution, the maximum likelihood MLE algorithm in step 2) is shown in Formula 2: P k =binom(U-counts k , ALL-counts k Fraction k ×U-ratio_target k +bdg k ) Formula 2; In the formula 2: k is an integer from 1 to n, U-counts k is the number of positive fragments of the kth methylation marker; ALL-counts k is the total number of fragments of the kth methylation marker; bdg k The weight of the noise signal contributed by the background cells to the kth methylation marker; U-ratio_target k is the probability prior of the kth methylation marker releasing positive fragments of the target cell obtained based on the cell sample of the target cell, P k is the expected value, take P k Fraction with maximum value k is the relative content of cfDNA in the target cell corresponding to the kth methylation marker; Preferably, the bdg k According to formula 1, we can get: In the formula 1: k is an integer from 1 to n, per i is the relative content of cfDNA of the i-th background cell among x-1 background cells in the sample to be tested, i is an integer from 1 to (x-1); U-ratio_bg ki is the probability prior of the kth methylation marker release positive fragment of the i-th background cell obtained based on the cell sample of the i-th background cell; Preferably, the per i It is obtained through the deconvolution algorithm.
8. The quantitative method according to any one of claims 1 to 7, characterized in that The target cell includes t methylation markers, wherein t is an integer between 1 and n. The first correction in step 3) is performed according to formula 3: In the formula 3, k is an integer from 1 to t, ALL-counts k is the total number of fragments of the kth methylation marker; Fraction k is the relative content of cfDNA in the target cells corresponding to the kth methylation marker.
9. The quantitative method according to any one of claims 4 to 8, characterized in that The second correction in step i) comprises: According to formula 4, the per i Correction was performed to obtain the corrected relative cfDNA content of each background cell per-adjust i , per-adjust i = per i × sum formula 4; Where i is an integer from 1 to (x-1); Preferably, the third correction includes: based on the per-adjust i , Formulas 1, 2 and 3 are used for correction to obtain the corrected relative content of cfDNA in target cells, Fraction-adjust.
10. The quantitative method according to claim 6, characterized in that: The internal reference cell set is selected from any one or more cells among the x types of cells; Preferably, the internal reference cells include a high-quality internal reference cell set and a suboptimal internal reference cell set; Preferably, the high-quality internal reference cell set is one or more cells whose Fraction in ≥10 samples of healthy people is no more than 0.5%, whose lower limit of quantification LoQ is no higher than 0.2%, whose positive fragment ratio of methylation marker is higher than 0.5 and / or whose U-ratio_bg is lower than 0.1; Preferably, the suboptimal internal reference cell set is one or more cells whose Fraction is no more than 0.5%, whose lower limit of quantification LoQ is no more than 0.3%, whose positive fragment ratio of methylation marker is higher than 0.3 and / or whose U-ratio_bg is lower than 0.15 in ≥10 samples of healthy people. Preferably, the high-quality internal reference cell set and the suboptimal internal reference cell set do not overlap with each other. Preferably, the fourth correction comprises: a) counting the Fraction ratio of each high-quality internal reference cell in the sample to be tested, wherein the ratio = the Fraction of each high-quality internal reference cell / the sum of the Fractions of all high-quality internal reference cells; b) if in the sample to be tested, the Fraction ratio of a high-quality internal reference cell is higher than the upper limit threshold of the high-quality internal reference cell ratio, then the high-quality internal reference cell is eliminated from the high-quality internal reference cell set, and / or the high-quality internal reference cells whose Fraction is higher than the absolute upper limit threshold of the Fraction are eliminated from the high-quality internal reference cell set, Among them, the upper limit threshold of the high-quality internal reference cell ratio = the 85th to 99th percentile of the high-quality internal reference cell Fraction value in the healthy population sample / (the 85th to 99th percentile of the high-quality internal reference cell Fraction value in the healthy population sample + the 1st to 15th percentile of other high-quality internal reference cell Fraction values in the healthy population sample), The absolute Fraction upper limit threshold is the 85th to 99th percentile of the Fraction of the high-quality internal reference cells in the healthy population sample; c) If after step b), there are still m kinds of high-quality internal reference cells whose Fraction is above the lower limit threshold of absolute Fraction in the high-quality internal reference cell set, the target cell Fraction is corrected according to formula 5 to obtain the internal reference corrected target cell Fraction: In the formula 5, j = 1 ~ m, m is an integer from 1 to x, Fraction of internal reference cells in healthy samples j is the 85th to 99th percentile value of the Fraciton of the jth internal reference cell obtained based on the healthy population sample in the healthy population sample, Fraction of internal reference cells in the sample to be tested j is the Fraciton of the jth internal reference cell obtained based on the sample to be tested; d) if after step b), there are still high-quality internal reference cells whose Fraciton is below the lower limit threshold of absolute Fraction in the high-quality internal reference cell set, a total of m types of internal reference cells including high-quality internal reference cells and suboptimal internal reference cells having a relative cfDNA content Fraction of less than 0.1%, less than 0.2%, less than 0.3%, less than 0.4%, less than 0.5%, less than 0.6%, less than 0.7%, less than 0.8%, less than 0.9% or less than 1% in the sample to be tested are used to correct the target cell Fraction according to formula 5 to obtain the internal reference corrected target cell Fraction; e) if after step b), no high-quality internal reference cells exist, then the fourth calibration is performed without using the internal reference cells; The absolute lower limit threshold is the 1st to 15th quantile of the Fraction of the high-quality internal reference cells in the healthy population sample; Preferably, the fourth calibration is performed after step 4), step i) or step ii). Preferably, when one internal reference cell is a target cell, the internal reference cell is not used for the fourth calibration, and the fourth calibration is performed using other internal reference cells in the internal reference cell set.
11. The quantitative method according to claim 1, characterized in that: The methylation sequencing includes one or more of whole genome methylation sequencing, multiplex PCR methylation sequencing and / or targeted methylation sequencing; Preferably, the depth of the methylation sequencing is ≥400×, ≥500×, ≥600× or ≥700×; Preferably, the sample to be tested or the sample of healthy people includes a body fluid sample; Preferably, the body fluid sample includes: saliva or saliva supernatant, serum, plasma, milk, urine, lumbar or ventricular CSF, lymph fluid, prostatic fluid, semen, sputum, stool diluted supernatant, tears, alveolar lavage fluid, sputum, pus, nasopharyngeal swab, oral swab, cerebrospinal fluid, pleural effusion, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, vaginal discharge diluted supernatant and their processed products.
12. A quantitative system for the relative content of cfDNA in cells, characterized in that: The quantitative system is used to implement the quantitative method according to claims 1 to 11, and the system comprises: The positive fragment acquisition module counts the positive fragments in the region of each methylation marker according to the methylation sequencing data; screens the methylation markers with a total fragment number U-counts ≥ z in the sample to be tested, where z is an integer from 10 to 200; A background weight acquisition module, which accumulates and sums the probability of cfDNA derived from each background cell in the biological sample releasing positive fragments in the region of the target cell-specific methylation marker to obtain the background cell weight; The module for obtaining the relative content of cfDNA of methylation markers constructs a likelihood formula to calculate the relative content of cfDNA of target cells corresponding to each methylation marker; The target cell cfDNA relative content acquisition module uses the depth coefficient of each methylation marker as a weight to calculate the weighted average of all methylation markers specific to the same target cell as the relative cfDNA content of each target cell.
13. The quantitative system according to claim 12, characterized in that: The quantitative system also includes unknown cell component ratio correction, counting the sum of the relative cfDNA content Fraction of all target cells, and correcting the relative cfDNA content of each background cell in the background weight acquisition module; Preferably, the system further comprises a methylation marker combination module: merging methylation markers with a positive fragment ratio difference of the same target cell of ≤0.2 and adjacent positions on the chromosome into a combined methylation marker.
14. A method for tracing the source of organ damage, the quantitative method comprising obtaining the relative content of target cell cfDNA in samples from a healthy population of ≥10 samples according to the quantitative method according to any one of claims 1 to 11, and statistically obtaining the 85th to 99th percentiles of the relative content of target cell cfDNA in the samples from the healthy population as a damage threshold; if the relative content of target cell cfDNA in the sample to be tested is ≥ the damage threshold, it is judged that the target cell is damaged, the degree of damage is positively correlated with the relative content of target cell cfDNA in the sample to be tested, and the organ is further judged to be damaged based on the organ from which the target cell originates.
15. An organ damage tracing detection system, characterized in that: The detection system is used to implement the organ damage tracing method according to claim 14, and the system comprises: The damage threshold construction module uses the 85th to 99th percentile of the relative content of cfDNA in target cells of samples from a healthy population of ≥10 samples as the damage threshold; Cell damage judgment module: when the relative content Fraction of target cell cfDNA in the sample to be tested is ≥ the damage threshold, it is judged that the target cell is damaged; The organ damage judgment module further judges whether the organ is damaged according to the organ from which the target cells originate.
16. A device, comprising a memory for storing a program; and a processor for implementing the quantitative method of claims 1 to 11 or the organ damage tracing method of claim 14 by executing the program stored in the memory.
17. A computer-readable storage medium having a program stored thereon, wherein the program can be executed to implement the quantitative method of claims 1 to 11, or the organ damage tracing method of claim 14.
18. Application of the quantitative method according to claims 1 to 11 or the organ damage tracing method according to claim 14 in cell / organ tracing, cell / organ damage detection, cell / organ health assessment, early disease screening and diagnosis, disease course monitoring, recurrence prediction, and prognosis assessment.
Citation Information
Patent Citations
Method and system for predicting tumor tissue cfDNA
CN112086129A
Evaluation system for predicting tissue specific source and related disease probability of cfDNA and application
CN113539355A
Method and device for layered screening of methylation markers
CN115497561A
Multi-cancer-species methylation detection kit and application thereof
CN117535404A
Quantitative method for plasma ctDNA level of subject
CN117604086A
Cited By
Application of mononuclear macrophages in breast cancer visceral crisis judgment
CN121087171A
Methylation marker panel for detecting lung injury and use thereof
CN122629197A