Method for detecting organ injury based on cell methylation markers and application thereof
By developing a source quantification method for methylation biomarkers based on the maximum likelihood MLE algorithm and a deep neural network model, the sensitivity and robustness issues of existing cfDNA methylation source quantification methods in organ injury identification are resolved, achieving highly sensitive detection of target cell signals and accurate assessment of organ injury.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GENEPLUS-BEIJING CLINICAL LAB CO LTD
- Filing Date
- 2024-12-31
- Publication Date
- 2026-04-14
AI Technical Summary
Existing cfDNA methylation tracing and quantification methods suffer from low sensitivity and poor robustness when identifying organ damage, making it difficult to achieve accurate identification at the cellular level, especially under low-depth conditions where they cannot accurately identify trace signals.
By combining the maximum likelihood (MLE) algorithm and deep neural network model with a source quantification method for methylation biomarkers, specific methylation biomarkers are screened out. Relative quantitative analysis is then performed using the methylation characteristics of target cells and background cells in plasma cfDNA. An internal reference calibration system is constructed to eliminate background noise interference and achieve highly sensitive detection of target cell signals.
It achieves highly sensitive and robust detection of organ damage, accurately identifies the type and extent of damage at the cellular level, and assists in early disease screening, diagnosis, disease monitoring, and prognosis assessment.
Smart Images

Figure QLYQS_2 
Figure QLYQS_3 
Figure QLYQS_5
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedicine, specifically relating to a method and application for detecting organ damage based on cellular methylation markers. Background Technology
[0002] Organ damage is a common manifestation of many diseases. Timely and accurate detection of organ damage and its extent can help develop more effective treatment plans and improve patient treatment and prognosis. For example, kidney damage often indicates that kidney function is about to change or has already changed, while liver damage may be a direct manifestation of diseases such as hepatitis or cirrhosis. The diagnosis of organ damage often relies on multiple methods, including imaging, pathology, and laboratory tests. Different types and degrees of organ damage may have different characteristics and manifestations, which increases the complexity and difficulty of actual clinical diagnosis. For example, in the diagnosis of kidney diseases with high prevalence and low awareness, especially for early kidney damage, existing clinical indicators (serum creatinine, glomerular filtration rate) often show low sensitivity. A positive result for these indicators often indicates serious kidney function problems and has a significant lag. In addition, kidney biopsy technology requires careful consideration in clinical application due to its invasiveness, sampling location, patient condition, limited repeatability, and risks of complications. Therefore, developing a non-invasive and highly sensitive method for detecting organ damage is crucial for improving diagnostic efficiency, developing effective treatment plans, monitoring treatment efficacy, and improving prognosis.
[0003] Compared to invasive tissue biopsies, liquid biopsy technology has gained widespread attention in recent years due to its non-invasive and sensitive characteristics. Cell-free DNA (cfDNA) is a highly fragmented DNA molecule released into the circulatory system during cell apoptosis and necrosis, characterized by a short half-life (approximately 5-150 minutes). In the development and treatment of many diseases, such as cancer, autoimmune diseases, and various chronic diseases, the rate of cell death can be affected, thus influencing the proportion of cfDNA from various organs / cells in the blood. cfDNA provides a comprehensive and non-invasive overview of the health of all cells and tissues in the body. Deciphering and quantifying the source of various organ / cell components in cfDNA has enormous clinical potential in assisting disease diagnosis, monitoring treatment efficacy, and prognosis.
[0004] DNA methylation is a highly conserved, fundamental epigenetic modification controlling gene expression and chromatin organization, playing a crucial role in biological processes such as gene expression regulation, cell differentiation, embryonic development, and disease occurrence. Netanel Loyfer et al. demonstrated that DNA methylation primarily influences downstream gene expression through interconnected CpG sites in functional regulation. Furthermore, they mapped the first unique DNA methylation atlases for major human cell types, revealing that most cell-specific methylation markers are often hypomethylated or completely unmethylated. This provides a new analytical perspective for quantifying cell origins using DNA methylation markers.
[0005] Existing methods for quantifying cfDNA methylation have their limitations. Netanel Loyfer et al., based on deconvolution algorithms, developed quantification methods using methylation rate (MethAltas) and fragmentation (UMX / Markov) as features. However, because deconvolution methods require estimation of all cellular signals in cfDNA, they cannot identify cell types lacking methylation markers, thus easily leading to quantitative bias. Furthermore, deconvolution algorithms, due to their inherent limitations, cannot accurately identify trace signals (percentiles to thousandths) in cfDNA. Christa Caggiano developed a low-depth whole-genome quantification method based on the Maximum Likelihood Estimation (MLE) algorithm, using CpG sites as features. However, due to the highly fragmented nature of cells, using site methylation rate as a quantification indicator has poor robustness and specificity, and it also cannot accurately identify trace signals at low depths. In addition, Shuo Li et al. developed a tissue-level source quantification model based on fragment methylation rate as a feature, using a deep neural network model. However, this model requires complex hyperparameter tuning and training support from a large cohort of samples, resulting in poor model interpretability. Furthermore, existing methylation biomarkers cannot meet the source resolution requirements at the cellular level.
[0006] In summary, there is a need to develop a highly sensitive, robust, and cellular-level resolution-based quantitative method for tracing organ damage, enabling accurate identification of minute damage signals in target organs / cells and assisting in clinical diagnosis and treatment. Summary of the Invention
[0007] To achieve highly sensitive organ damage detection with cellular-level precision, this invention provides a relative quantitative method for organ damage based on cellular methylation markers. By tracing the relative proportions of various cell sources in plasma cfDNA, the method accurately identifies the type and extent of damage at the cellular level, assisting in functions such as health assessment of various organs, early disease screening and diagnosis, disease progression monitoring, and prognostic evaluation.
[0008] According to one aspect of this disclosure, a method for quantifying cfDNA derived from target cells in a sample to be tested is provided, the method comprising the following steps:
[0009] 1) Based on the methylation sequencing data of the sample to be tested, obtain sequencing fragments of n methylation markers, count the total number of fragments and the number of positive fragments for each methylation marker, and screen methylation markers with a total number of fragments ≥ z for subsequent steps, where z is an integer from 10 to 200, wherein the sample to be tested includes cfDNA from x types of cells, and the x types of cells include n methylation markers;
[0010] 2) Using one type of cell in the test sample as the target cell and the remaining x-1 types of cells as background cells, based on the prior probability of each methylation marker releasing a positive fragment, the total number of fragments ALL-counts and the number of positive fragments U-count for each methylation marker in the test sample, and the weight of the background cells contributing noise signal to each methylation marker, the relative content of cfDNA for each methylation marker corresponding to the target cell is obtained using the maximum likelihood (MLE) algorithm.
[0011] 3) Using the total number of fragments of each methylation marker in the sample to be tested as a depth coefficient, the relative content of cfDNA in the target cell corresponding to each methylation marker is first corrected to obtain the relative content of cfDNA in the target cell in the sample to be tested, Fraction.
[0012] Optionally, 4) repeat steps 2) and 3) to obtain the relative content of cfDNA in x types of target cells;
[0013] The methylation markers include genomic regions in the target cells where the proportion of unmethylated fragments is higher or lower than that in the background cells.
[0014] Where x is an integer ≥ 1, and n is an integer ≥ x.
[0015] Each cell type includes one or more methylation markers.
[0016] Those skilled in the art can select more different parameters as needed, all of which can be applied to the quantitative methods described in this disclosure. In some embodiments, the positive fragment is a fragment containing at least 3, at least 4, at least 5, or at least 6 completely unmethylated CpG sites.
[0017] In some implementations, the weight of the background cells contributing noise signal to each methylation marker is obtained based on the prior probability of each background cell releasing a positive fragment for each methylation marker, and the statistical analysis of the relative cfDNA content of each background cell in the sample to be tested.
[0018] In some embodiments, the methylation markers of the cells include genomic regions in which the positive fragment ratio of the cells is at least 0.2 greater than that of the background cells; the positive fragment ratio is the ratio of the number of positive fragments in the genomic region to the total number of fragments.
[0019] In some exemplary embodiments, each cell specificity contains at least 10 methylation markers.
[0020] In some implementations, z includes, but is not limited to, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200.
[0021] In some implementations, the relative cfDNA content of each background cell is obtained based on the methylation sequencing data of the sample to be tested, using a deconvolution algorithm.
[0022] In some implementations, after step 4), the method further includes the following steps:
[0023] i) Based on the sum of the relative cfDNA contents of the x target cells, perform a second correction on the relative cfDNA contents of each background cell to obtain the corrected relative cfDNA contents of each background cell; and
[0024] ii) Using the corrected relative cfDNA content of each background cell obtained in step ii), repeat steps 2) to 3) to perform a third correction on the relative cfDNA content of the target cells in the test sample.
[0025] In some embodiments, prior to step 1), the method further includes the step of obtaining one or more methylation markers for each of the x types of cells.
[0026] In some embodiments, the step of obtaining one or more methylation markers for each of the x cell types includes: the methylation markers of the cells comprising genomic regions in which the positive fragment ratio of that cell type is at least 0.2 greater than that of the background cells; the positive fragment ratio being the ratio of the number of positive fragments in the genomic region to the total number of fragments.
[0027] In some embodiments, the methylation markers include co-methylation markers, which are obtained by combining two methylation markers located adjacent to each other on the same chromosome in the same cell and having a positive fragment ratio difference of <0.2.
[0028] In some implementations, the merging includes summing the total number of fragments from the two methylation markers and / or summing the number of positive fragments from the two methylation markers.
[0029] In some embodiments, the method further includes the step of: using internal reference cells with stable relative cfDNA content in healthy human samples to perform a fourth correction on the relative cfDNA content of the target cells in the test sample.
[0030] In some implementations, in step 2), based on the binomial distribution of the positive fragment released from cfDNA by the target cells, the maximum likelihood MLE algorithm is as shown in Formula 2:
[0031] P k =binom(U-counts) k ALL-counts k Friction k ×U-ratio_target k +bdg k ) Formula 2;
[0032] In formula 2:
[0033] k is an integer from 1 to n.
[0034] U-counts k The number of positive fragments for the k-th methylation marker;
[0035] ALL-counts k The total number of fragments for the k-th methylation marker;
[0036] bdg k The weights for the noise signal contributed by the background cells to the k-th methylation marker;
[0037] U-ratio_target kThis is a priori probability of the release of a positive fragment of the k-th methylation marker of the target cell obtained from a cell sample of the target cell.
[0038] P k Let P be the expected value. k Friction at its maximum value k The relative content of cfDNA in the target cell corresponding to the k-th methylation marker.
[0039] In some implementations, the bdg k Calculated according to Formula 1:
[0040]
[0041] In Formula 1:
[0042] k is an integer from 1 to n.
[0043] per i The relative content of cfDNA in the i-th background cell among x-1 background cells in the sample to be tested.
[0044] i is an integer from 1 to (x-1);
[0045] U-ratio_bg ki This is the probability prior of the release of a positive fragment of the k-th methylation marker in a cell sample of the i-th background cell.
[0046] In some implementations, the per i It was obtained through the deconvolution algorithm.
[0047] In some embodiments, the target cells include t methylation markers, where t is an integer between 1 and n.
[0048] The first correction in step 3) is performed according to formula 3:
[0049]
[0050] In formula 3,
[0051] k is an integer from 1 to t.
[0052] ALL-counts k The total number of fragments for the k-th methylation marker;
[0053] Fraction k The relative content of cfDNA in the target cell corresponding to the k-th methylation marker.
[0054] In some implementations, the second correction in step i) includes:
[0055] According to Formula 4, per i Correction was performed to obtain the corrected relative content of cfDNA per-adjustment for each background cell type. i ,
[0056] per-adjust i =per i ×sum formula 4;
[0057] Where i is an integer from 1 to (x-1).
[0058] In some implementations, the third correction includes: based on the per-adjustment i Formulas 1, 2, and 3 are used for correction to obtain the corrected relative content of cfDNA in the target cells (Fraction-adjust).
[0059] In some implementations, the internal control cell set is selected from any one or more cells among the x types of cells.
[0060] In some embodiments, the internal control cells include a high-quality internal control cell set and a suboptimal internal control cell set.
[0061] In some embodiments, the high-quality internal control cell set is one or more cell types that have a fractionation of no more than 0.5%, a lower limit of quantitation (LoQ) of no more than 0.2%, a positive fragment ratio of methylation markers of more than 0.5, and / or a U-ratio_bg of less than 0.1 in ≥10 healthy human samples.
[0062] In some embodiments, the suboptimal internal control cell set is one or more cell types that, in ≥10 healthy human samples, have a fractionation rate not exceeding 0.5%, a lower limit of quantitation (LoQ) not exceeding 0.3%, a positive fragment ratio of methylation markers exceeding 0.3, and / or a U-ratio_bg below 0.15.
[0063] In some implementations, the high-quality internal control cell set and the suboptimal internal control cell set do not overlap.
[0064] In some implementations, the fourth correction includes:
[0065] a) Calculate the Fraction ratio of each type of high-quality internal reference cell in the test sample, wherein the ratio = Fraction of each type of high-quality internal reference cell / the sum of Fraction of all high-quality internal reference cells;
[0066] b) If, in the test sample, the Fraction ratio of a high-quality internal control cell is higher than the upper limit threshold for that high-quality internal control cell ratio, then that high-quality internal control cell is removed from the set of high-quality internal control cells, and / or high-quality internal control cells with Fraction values higher than the absolute upper limit threshold for Fraction are removed from the set of high-quality internal control cells.
[0067] Wherein, the upper limit threshold of the ratio of high-quality internal reference cells = the 85th to 99th percentile of the Friction value of the high-quality internal reference cell in the healthy population sample / (the 85th to 99th percentile of the Friction value of the high-quality internal reference cell in the healthy population sample + the 1st to 15th percentile of the Friction values of other high-quality internal reference cells in the healthy population sample).
[0068] The absolute upper limit threshold for Fraction is the 85th to 99th percentile of the Fraction of the high-quality internal control cells in the healthy population sample.
[0069] c) If, after step b), the high-quality internal reference cell set still contains m types of high-quality internal reference cells with Fraciton values above the absolute Fraction lower limit threshold, then the target cell Fraction is corrected according to Formula 5 to obtain the internal reference corrected target cell Fraction:
[0070]
[0071] In formula 5,
[0072] j = 1 ~ m,
[0073] m is an integer from 1 to x.
[0074] Fraction cells from healthy human samples j This refers to the 85th to 99th percentile value of Fraciton in the j-th type of internal reference cell, obtained from the healthy population sample.
[0075] Friction of internal control cells in the test sample j Fraciton is the type of internal reference cell obtained based on the sample to be tested;
[0076] d) If, after step b), there are still high-quality internal control cells in the high-quality internal control cell set whose Fraciton is below the absolute Fraction lower limit threshold, then use a total of m types of internal control cells, including high-quality internal control cells and suboptimal internal control cells, whose relative cfDNA content Fraction in the test sample is below 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, or 1%, to correct the target cell Fraction according to Formula 5, and obtain the internal control corrected target cell Fraction.
[0077] e) If no high-quality internal control cells are available after step b), then the fourth calibration is not performed using internal control cells;
[0078] The absolute lower limit threshold is the 1st to 15th quantile of the high-quality internal reference cells in the fraction of the healthy population sample.
[0079] In some implementations, the fourth correction is performed after step 4), step i), or step ii).
[0080] In some implementations, when an internal reference cell is the target cell, the internal reference cell is not used for the fourth calibration, which is performed using other internal reference cells in the set of internal reference cells.
[0081] In some implementations, the healthy population sample includes at least 10 healthy individuals.
[0082] In some embodiments, the methylation sequencing includes one or more of whole-genome methylation sequencing, multiplex PCR methylation sequencing, and / or targeted methylation sequencing.
[0083] In some embodiments, the methylation sequencing depth is ≥400×, ≥500×, ≥600×, or ≥700×.
[0084] In some implementations, the sample to be tested or the sample from a healthy population includes a body fluid sample.
[0085] In some embodiments, the body fluid samples include: saliva or saliva supernatant, serum, plasma, milk, urine, lumbar or ventricular CSF, lymph, prostatic fluid, semen, sputum, diluted fecal supernatant, tears, bronchoalveolar lavage fluid, sputum, pus, nasopharyngeal swabs, oral swabs, cerebrospinal fluid, pleural effusion, peritoneal fluid, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, diluted vaginal discharge supernatant, and their processed forms.
[0086] According to another aspect of the disclosure, a quantitative system for the relative content of cfDNA in cells is provided, the quantitative system being used to implement the quantitative method, the system comprising:
[0087] The positive fragment acquisition module counts positive fragments within the region of each methylation marker based on methylation sequencing data; it then filters methylation markers in the sample to be tested whose total fragment count U-counts ≥ z, where z is an integer from 10 to 200.
[0088] The background weight acquisition module sums up the probability of cfDNA from each background cell in the biological sample releasing positive fragments in the region of the target cell-specific methylation marker to obtain the background cell weight.
[0089] The module for obtaining the relative content of methylation biomarkers cfDNA constructs a likelihood formula to calculate the relative content of cfDNA in target cells corresponding to each methylation biomarker.
[0090] The module for obtaining the relative content of cfDNA in target cells calculates the weighted average of all methylation markers specific to the same target cell, using the depth coefficient of each methylation marker as the weight, and uses this average as the relative content of cfDNA for each target cell.
[0091] In some embodiments, the quantification system further includes unknown cell component ratio correction, summing the relative cfDNA content (Fraction) of all target cells, and correcting the relative cfDNA content of each background cell in the background weight acquisition module.
[0092] In some embodiments, the system further includes a methylation marker concatenation module: merging methylation markers with a positive fragment ratio difference ≤0.2 and adjacent positions on the chromosome into a concatenation methylation marker.
[0093] According to another aspect disclosed, a method for tracing the source of organ damage is provided. The quantitative method includes obtaining the relative content of target cell cfDNA in ≥10 healthy human samples according to the quantitative method, and statistically obtaining the 85th to 99th percentile of the relative content of target cell cfDNA in the healthy human samples as the damage threshold. If the relative content of target cell cfDNA in the test sample is ≥ the damage threshold, it is determined that the target cell is damaged. The degree of damage is positively correlated with the relative content of target cell cfDNA in the test sample. The organ from which the target cell originated is further determined to be damaged.
[0094] According to another aspect of the disclosure, an organ injury tracing and detection system is provided, the detection system being used to implement the organ injury tracing method, the system comprising:
[0095] The damage threshold construction module uses the 85th to 99th percentile of the relative content of target cell cfDNA in the fraction of Fraction in healthy human samples with ≥10 samples as the damage threshold.
[0096] The cell damage assessment module determines that the target cell is damaged when the relative content of cfDNA in the target cell in the test sample is greater than or equal to the damage threshold.
[0097] The organ damage assessment module further determines whether the organ is damaged based on the organ from which the target cells originate.
[0098] According to another aspect of the disclosure, an apparatus is provided, the apparatus including a memory for storing a program; and a processor for implementing the quantitative method or the organ injury tracing method by executing the program stored in the memory.
[0099] According to another aspect of the disclosure, a computer-readable storage medium is provided, on which a program is stored, the program being executable to implement the quantitative method or the organ injury tracing method described herein.
[0100] According to another aspect disclosed, the quantitative method or the organ damage tracing method described herein is provided for application in cell / organ tracing, cell / organ damage detection, cell / organ health assessment, early disease screening and diagnosis, disease progression monitoring, recurrence prediction, and prognostic assessment.
[0101] Beneficial effects:
[0102] This invention provides a method for the relative quantitative detection of organ damage based on cellular methylation markers. It improves the detection of positive signals by deeply optimizing model parameters; corrects for the proportion of unknown cell components to eliminate or partially eliminate biases caused by the presence of unknown cell types in the relative content of background cell cfDNA; and constructs an internal reference calibration system using a set of relatively stable internal reference cells within the normal range of a healthy baseline as a benchmark to correct the actual target cell signal to the baseline level for assessing real organ damage, eliminating biases caused by high-abundance background fluctuations in the target cell signal in the actual sample. This quantitative method can accurately reproduce the target cell signal, and through personalized dynamic internal reference calibration, it can achieve highly sensitive detection of trace damage signals in real clinical samples. Precise analysis of the components in plasma cfDNA enables early screening, early diagnosis, disease monitoring, and prognostic assessment of various organ diseases. Attached Figure Description
[0103] Figure 1 Performance validation of simulations for important organ / cell types is shown.
[0104] Figure 2The analysis of the relative quantitative components of white blood cells is shown.
[0105] Figure 3 The procedure for damage analysis of a real sample is shown.
[0106] Figure 4 The results of the performance verification of renal impairment in systemic lupus erythematosus (SLE) are shown.
[0107] Figure 5 The cellular-level distinguishing properties of methylation markers were demonstrated. Detailed Implementation
[0108] The purpose of this invention is to provide a method for relatively quantitative detection of organ damage based on methylation markers. By accurately analyzing various components in cfDNA of body fluid samples (e.g., plasma), this method enables functions such as health assessment of various organs, early screening and diagnosis of diseases, disease monitoring, and prognosis assessment.
[0109] This invention provides a method for accurately tracing and quantifying the components of cfDNA in body fluid samples (e.g., plasma) based on cell-specific methylation markers (hereinafter referred to as methylation markers).
[0110] In some embodiments, the cell-specific methylation marker refers to a genomic region in which the proportion of unmethylated fragments in this type of cell is significantly higher than that in other types of cells, or a genomic region in which the proportion of methylated fragments in this type of cell is significantly higher than that in other types of cells.
[0111] In some embodiments, the cell-specific methylation marker refers to a genomic region in which the proportion of unmethylated fragments in that cell type is significantly higher than that in other cell types.
[0112] In an exemplary embodiment, the cell-specific methylation marker includes genomic regions where the positive fragment ratio of that cell type is at least 0.2, at least 0.3, at least 0.4, or at least 0.5 greater than that of other cells. The positive fragment ratio is the ratio of the number of positive fragments sequenced within the genomic region to the total number of fragments. In some embodiments, the positive fragment refers to a completely unmethylated fragment within the methylation marker region containing at least 3, at least 4, at least 5, or at least 6 CpG sites.
[0113] In some embodiments, the cell-specific methylation markers can be obtained using methods conventional in the art, such as methylation markers for a particular cell obtained through routine screening based on genomic methylation sequencing data of that cell.
[0114] In some exemplary embodiments, the cells include, but are not limited to, one or more of the following: aortic smooth muscle cells, cerebellar neurons, coronary smooth muscle cells, cortical neurons, endothelial cells, cardiac cardiomyocytes, cardiac fibroblasts, renal epithelial cells, hepatocytes, alveolar epithelial cells, bronchial epithelial cells, oligodendrocytes, ovarian cells, pancreatic acinar cells, pancreatic alpha cells, pancreatic beta cells, pancreatic delta cells, pancreatic duct cells, B cells, T cells, granulocytes, monocytes / macrophages, natural killer cells, osteoblasts, mammary epithelial cells, breast basal cells, skin fibroblasts, smooth muscle cells, gallbladder cells, skin keratinocytes, endometrial epithelial cells, gastric epithelial cells, head and neck epithelial cells, smooth muscle cells, skeletal muscle cells, small intestinal epithelial cells, colonic epithelial cells, prostate epithelial cells, tonsil cells, thyroid epithelial cells, and bladder epithelial cells.
[0115] In some exemplary embodiments, the cells include, but are not limited to, one or more of the following: aortic smooth muscle cells, cerebellar neurons, coronary smooth muscle cells, cortical neurons, endothelial cells, cardiac cardiomyocytes, cardiac fibroblasts, renal epithelial cells, hepatocytes, alveolar epithelial cells, pulmonary bronchial epithelial cells, oligodendrocytes, ovarian cells, pancreatic acinar cells, pancreatic alpha cells, pancreatic beta cells, pancreatic delta cells, and pancreatic duct cells.
[0116] In some embodiments, the method includes:
[0117] Data acquisition for sample S01 to be tested.
[0118] Acquire methylation data of the sample to be tested, and based on the methylation data, count the number of positive fragments (U-counts) and the total number of fragments (ALL-counts) on each cell-specific methylation marker in the sample to be tested.
[0119] In some implementations, cell-specific methylation markers with a total number of fragments (ALL-counts) ≥ z in the sample to be tested are selected for subsequent steps, where z is an integer from 10 to 200.
[0120] In some implementations, z includes, but is not limited to, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200.
[0121] In some embodiments, the methylation sequencing includes whole-genome methylation sequencing, multiplex PCR methylation sequencing, and / or targeted methylation sequencing.
[0122] In some implementations, the sample to be tested is subjected to high-depth methylation sequencing, the depth of which is, for example, ≥400×, ≥500×, ≥600× or ≥700×.
[0123] In some implementations, the sample to be tested includes a body fluid sample.
[0124] In some exemplary embodiments, the bodily fluid samples include: saliva or saliva supernatant, serum, plasma, milk, urine, lumbar or ventricular CSF, lymph, prostatic fluid, semen, sputum, diluted fecal supernatant, tears, bronchoalveolar lavage fluid, sputum, pus, nasopharyngeal swabs, oral swabs, cerebrospinal fluid, pleural effusion, peritoneal fluid, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, diluted vaginal discharge supernatant, and their processed forms.
[0125] In some implementations, the methylation sequencing data of the sample to be tested includes methylation information of CpG sites.
[0126] In some embodiments, the methylation sequencing data of the sample to be tested includes methylation information of the CpG sites of the cell-specific methylation marker.
[0127] In some implementations, the methylation sequencing data of the sample to be tested is obtained in the form of a fragment file containing information related to CpG sites.
[0128] SO2, a combination of methylation markers.
[0129] To improve the detection rate of positive signals, if two methylation markers belonging to the same cell type, with adjacent sequencing positions on the same chromosome, and with similar positive fragment ratios (e.g., difference <0.2) exist in the test data of the sample, they are combined and used as a single methylation marker to improve data utilization and detection rate. The combination includes:
[0130] The sum of the U-counts of positive fragments observed by the two methylation markers is taken as the U-counts of positive fragments of the combined methylation marker. joint ;
[0131] Sum the total number of fragments observed for the two methylation markers (ALL-counts). k ≥z, where z is an integer from 10 to 200), is the total number of fragments (ALL-counts) representing the co-methylation marker. joint .
[0132] For example, two renal cell methylation markers, A and B, with positive fragment ratios of 30 / 100 and 40 / 100 respectively, can be combined into a single renal cell methylation marker with a positive fragment ratio of 70 / 200.
[0133] In some embodiments, the positive fragment refers to a completely unmethylated fragment containing at least 3, at least 4, at least 5, or at least 6 CpG sites within the target cell methylation marker region.
[0134] In some embodiments, the positive fragment ratio is the ratio of the number of positive fragments (U-counts) within the methylation marker region of the target cell to the total number of fragments (ALL-counts).
[0135] In some implementations, step 2) is an optional step. If there are no two methylation markers in the sample that belong to the same cell type, are adjacent in the same chromosome, and have a close ratio of positive fragments (e.g., difference <0.2), step 2 can be omitted.
[0136] S03 accurately quantifies background weights.
[0137] Given the complex cellular composition of plasma cfDNA, background cells can generate noise that interferes with target cell detection, especially for low-abundance target cells (such as those from the kidney, heart, and brain). Positive signals from low-abundance target cells are released at low frequencies and are difficult to detect on a noisy cfDNA substrate. To improve the positive detection of these trace cell signals, it is necessary to accurately quantify the weight of background cells for noise reduction optimization. By refining the weight of noise signals released by each type of background cell, the noise coefficient (bdg) of each methylation marker in the sample can be calculated.
[0138] In some implementations, one or more types of cells are selected alternately as target cells, while the remaining cells are used as background cells.
[0139] In some implementations, one type of cell is selected alternately as the target cell, while the remaining x-1 types of cells are used as background cells.
[0140] In some implementations, for the k-th methylation marker among n methylation markers of x cell types (n is an integer ≥ x, x is an integer ≥ 1), the cell specifically corresponding to the k-th methylation marker is the target cell, and the remaining x-1 cell types are background cells. The decomposition formula is as follows:
[0141]
[0142] In Formula 1, k is an integer from 1 to n, and the noise figure bdg kU-ratio_bg represents the weight of the noise signal contribution of x-1 background cells to the k-th methylation marker in the sample to be tested. ki Let represent the prior probability of the i-th background cell releasing a positive fragment in the k-th methylation marker region, where i is an integer from 1 to (x-1), per i This represents the relative content of cfDNA in the i-th background cell in the sample to be tested.
[0143] In some implementations, for per i The estimation was quantized using the UMX fragment-level deconvolution algorithm reported in the reference Loyfer N, Magenheim J, Peretz A, et al. A DNA methylation atlas of normal human cell types[J]. Nature, 2023, 613.
[0144] In some implementations, the U-ratio_bg ki It is based on the statistical analysis of positive fragment ratios using nucleic acid methylation sequencing data from different cells, thereby obtaining the prior probability of each type of background cell releasing a positive fragment in each target cell-specific methylation marker region.
[0145] S04, Basic Model Construction.
[0146] To simultaneously meet the requirements of high sensitivity and robust quantitative performance for trace signals from target cells, this invention uses the maximum likelihood (MLE) algorithm as the basis for constructing the relative quantitative model.
[0147] In some implementations, a likelihood formula is constructed based on the principle that the release of positive fragments from target cells follows a binomial distribution, and probability estimation is performed. Formula 2 is constructed for the k-th methylation marker as follows:
[0148] P k =binom(U-counts) k ALL-counts k Friction k ×U-ratio_target k +bdg k ) Formula 2,
[0149] In Formula 2, k is an integer from 1 to n, P k U-counts represents the expected value of the k-th methylation marker. k ALL-counts represents the number of positive fragments of the k-th methylation marker in the sample to be tested. k The total number of fragments of the k-th methylation marker in the sample to be tested (ALL-counts)k ≥100), U-ratio_target k This represents the prior probability of a target cell releasing a positive fragment at the k-th methylation marker region. (bdg) k The noise coefficient obtained in step 2) is taken as the expected value P. k Friction at Maximization k The relative content of cfDNA in the target cell corresponding to the k-th methylation marker.
[0150] In some implementations, the U-ratio_target k This was obtained by statistically analyzing the positive fragment ratios of methylation sequencing samples from different types of target cells.
[0151] Based on the model built above, subsequent iterative optimizations were carried out to further improve quantitative performance and to correct the fractionation.
[0152] S05, target cell fractionation is obtained by correcting for depth coefficient and statistical analysis.
[0153] To further optimize the model parameters at a deeper level, after obtaining the target cell fraction for each methylation marker, a weight correction of the depth coefficient was performed. The final quantitative result of the relative content fraction of cfDNA in target cells containing t methylation markers in the test sample is shown in Equation 3:
[0154]
[0155] In Formula 3, Friction k This represents the relative cfDNA content in the target cell at the expected maximum for each methylation marker, with the depth coefficient weighted by the total number of fragments (ALL-counts). k This means that k is an integer from 1 to t, and t is an integer from 1 to n.
[0156] S06, optionally, perform unknown cell component ratio correction.
[0157] The set of methylation markers cannot exhaustively represent all cell types in plasma cfDNA; the presence of unknown cell types can affect the accurate estimation of known cell signals. Furthermore, the UXM deconvolution algorithm, under normalization processing, will allocate the proportion of unknown cells to known cell types, resulting in a relative background cell cfDNA content per... iThis method is prone to bias, therefore the influence of unknown cellular components in the actual cfDNA needs to be considered and a correction scheme designed. Given that MLE is a non-normalized quantification method and the impact of unknown cell proportion allocation is relatively small, the relative content of background cell cfDNA per 100% of the sample can be reversed using MLE. i To achieve the correction of unknown cell component ratios, the target cell fraction is obtained through the MLE algorithm in S04 and S05. After summing the target cell fractions of x types of cells to obtain a total sum, the relative content of cfDNA of various background cells in the sample to be tested is further corrected as shown in Formula 4:
[0158] per-adjust i =per i ×sum formula 4,
[0159] In Formula 4, i is an integer from 1 to (x-1).
[0160] Subsequently, the corrected relative content of background cell cfDNA was used. per-adjust i Substituting back into Equation 1 to estimate the bdg weights, we obtain bdg-adjust. k Then adjust bdg-adjust k Substituting into Formula 2, we get Friction-adjust k Finally, Friction-adjust k Substituting the relative content of target cell cfDNA in the test sample into Formula 3, we obtain the Friction-adjustment. k .
[0161] S07, construct an internal reference calibration system to correct the relative content of cfDNA in target cells.
[0162] Significant heterogeneity exists in the signal release of high-abundance cells (various blood cells, liver cells, and endothelial-associated cells) among different individuals. Fluctuations in the abundance of these high-abundance cells under relative quantification can affect the accurate identification of low-abundance target cell damage signals in actual samples. Furthermore, the actual development of various diseases is often accompanied by inflammatory responses, and in most cases, the level of immune cells is also upregulated. This can "squeeze out" trace amounts of target cell damage signals, resulting in observed signals lower than the true value even with accurate quantification. Absolute quantification, on the other hand, often performs poorly in cross-sample comparisons due to the heterogeneity of cfDNA content among different individuals. Therefore, to maximize the restoration of the true signal abundance of target cells released from cfDNA in actual samples and to minimize the impact of high-abundance background fluctuations on target cell signals in actual samples, an internal reference correction system is optionally constructed based on the aforementioned relative quantification system. This system achieves accurate restoration of the observed target cell signals. Specifically, a set of relatively stable internal reference cells within the normal range of a healthy baseline is used as a benchmark to correct the actual target cell signals to the baseline level for assessing actual organ damage.
[0163] In some embodiments, the construction of the internal reference calibration system to calibrate the relative content of cfDNA in target cells includes:
[0164] S07-01, Construction of the Healthy Baseline Cohort.
[0165] Using healthy human samples as the baseline cohort, cfDNA molecules were extracted from bodily fluid samples (e.g., plasma) in the baseline cohort. Based on the methylation marker regions, methylation sequencing was performed on the target regions to obtain fragment methylation information of healthy baseline cfDNA molecules. Then, a relative quantification model was used to perform source quantification analysis on the healthy baseline, and the normal reference range of the relative content of cfDNA in each target cell in the baseline cohort was calculated and used as the benchmark for subsequent internal control selection and damage assessment.
[0166] In some embodiments, the methylation sequencing is high-depth methylation sequencing, and the depth of the high-depth methylation sequencing is, for example, ≥400×, ≥500×, ≥600× or ≥700×.
[0167] In some implementations, the normal reference range includes the 1st to 15th percentile and the 85th to 99th percentile.
[0168] S07-02, Construction of internal control cell cohort.
[0169] Based on the principles of stable fractionation in healthy human samples, satisfactory cell quantification accuracy, and high specificity of methylation markers, a group of stable high-quality internal control cells (hereinafter referred to as high-quality internal control cells) were selected from x types of cells for the correction of relative cfDNA content in target cells, and suboptimal internal control cells (hereinafter referred to as suboptimal internal control cells) that may pose risks were used as a supplement.
[0170] In some embodiments, the high-quality internal control cells are selected from the x types of cells according to the following criteria: 1. Friction stability in healthy human samples: the target cells' friction or friction-adjustment in healthy human samples does not exceed 0.5%; 2. Cell quantification accuracy: the lower limit of quantification (LoQ) of the target cell simulation experiment is not higher than 0.2%; 3. High specificity of methylation marker selection: the ratio of positive fragments in the methylation marker region of the target cells is higher than 0.5 and the U-ratio_bg of each background cell is lower than 0.1.
[0171] In some embodiments, the suboptimal internal control cells are selected from the x types of cells according to the following criteria: 1. Friction stability in healthy human samples: the target cells' friction or friction-adjustment in healthy human samples does not exceed 0.5%; 2. Cell quantification accuracy: the lower limit of quantification (LoQ) of the target cell simulation experiment is not higher than 0.3%; 3. High specificity of methylation marker selection: the ratio of positive fragments in the methylation marker region of the target cells is higher than 0.3 and the U-ratio_bg of each background cell is lower than 0.15.
[0172] In some embodiments, the simulation experiment of the internal reference cells refers to the conditions of sequencing depth ≥400×, ≥500×, ≥600× or ≥700×.
[0173] In some embodiments, the simulation experiment includes an experiment to perform relative quantification using sequencing data of known relative cfDNA content in different cells according to the methods described in this disclosure.
[0174] In some implementations, the sequencing data of the known relative content of cfDNA from different cells comes from wet or dry experiments.
[0175] In some embodiments, the wet experiment includes methylation sequencing using standard samples with known relative cfDNA content from different cells.
[0176] In some embodiments, the dry experiment includes mixing methylation sequencing data from different cells in a predetermined ratio.
[0177] In some exemplary embodiments, the simulation experiment is as shown in Example 2.
[0178] In some embodiments, the lower limit of quantification (LoQ) refers to the lowest known relative cfDNA content of the cells relatively quantified according to the method of this disclosure, provided that the difference between the relative cfDNA content quantified using the method of this disclosure and the known relative cfDNA content does not exceed 10% and the CV between replicate samples is less than 20%.
[0179] For example, if the relative cfDNA content of a gradient mixture of renal cells is known to be 0.02%, 0.2%, 1.0%, and 1.5%, and the relative cfDNA content obtained using the method described in this disclosure differs from the known relative cfDNA content by no more than 10%, and the CV between replicate samples is less than 20%, then the lower limit of quantification (LoQ) for the renal cells is 0.2%. However, if the relative cfDNA content is known to be 0.02%, and the relative cfDNA content obtained using the method described in this disclosure differs from the known relative cfDNA content by more than 10%, or the CV between replicate samples is greater than or equal to 20%, then the lower limit of quantification (LoQ) for the renal cells is 0.2%.
[0180] S07-03, Determination of baseline internal reference ratio / threshold and correction of the relative content of cfDNA in target cells in the test sample using internal reference cells.
[0181] Based on the baseline cohort of the healthy population samples, the internal reference cell signal of each sample was constructed, and the relative content of cfDNA of each type of internal reference cell in each healthy person was calculated (or the relative content of cfDNA in each type of internal reference cell, Fraction (or Fraction-adjust), with priority given to the ratio between high-quality internal reference cells.
[0182] The ratio of high-quality internal control cells in each sample from the healthy population was calculated, and then the upper limit threshold of the ratio of high-quality internal control cells in the healthy population samples was determined. When using internal control cells to correct for the relative cfDNA content of the test sample, firstly, based on the ratio among high-quality internal control cells, high-quality internal control cells in the test sample whose ratio exceeds the upper limit threshold of the ratio in the healthy population samples are removed. These cells are considered to be damaged internal control cells and are not suitable as internal control cells.
[0183] In addition, absolute thresholds are constructed to further determine the degree of damage to the internal control. The upper absolute threshold is the 85th to 99th percentile of each high-quality internal control cell in the baseline fraction (or fraction-adjustment). If the number of high-quality internal control cells in the test sample is higher than the upper absolute threshold, the internal control cells are determined to be damaged based on the ratio index and need to be removed. The lower absolute threshold is the 1st to 15th percentile of each high-quality internal control cell in the baseline fraction (or fraction-adjustment). If, after removal, there are high-quality internal control cells in the test sample that are between the upper absolute threshold and the lower absolute threshold, the damage is determined. Within the absolute lower limit threshold, high-quality internal control cells can be used to correct the Friction (or Friction-adjustment) value of the target cells. If high-quality internal control cells exist after removal, but the actual quantitative result is lower than the lower limit, it is considered that the high abundance background is squeezing out the internal control abundance, causing it to be too low, or that the actual internal control abundance is too low, resulting in the quantitative result of the internal control cells being lower than the quantitative performance. In this case, all internal controls at the same level (including high-quality and suboptimal internal controls) are used to correct the Friction (or Friction-adjustment) value of the target cells; however, this type of internal control correction carries risks. If no high-quality internal control is found in the sample after removal, no correction is performed.
[0184] In some embodiments, being at the same level includes a relative cfDNA content of less than 0.1%, less than 0.2%, less than 0.3%, less than 0.4%, less than 0.5%, less than 0.6%, less than 0.7%, less than 0.8%, less than 0.9%, or less than 1%.
[0185] In some implementations, the ratio of each type of high-quality internal control cell in each sample is calculated as follows: Ratio of each type of high-quality internal control cell = Friction (or Friction-adjustment) of that type of high-quality internal control cell / Sum of Friction (or Friction-adjustment) of all high-quality internal control cells.
[0186] In some embodiments, the upper limit threshold of the ratio of each type of high-quality internal control cell is equal to the 85th to 99th percentile of the baseline Friction (or Friction-adjustment) value of that type of high-quality internal control cell / (the 85th to 99th percentile of the baseline Friction (or Friction-adjustment) value of that type of high-quality internal control cell + the 1st to 15th percentile of the baseline Friction (or Friction-adjustment) value of the remaining types of high-quality internal control cells).
[0187] In some implementations, since the reference cells also belong to the x cell types, when each type of reference cell is used as a target cell for calibration, the remaining types of reference cells are used as reference cells for calibration. For example, for 3 high-quality reference cells and 5 suboptimal reference cells, when any one of the high-quality reference cells is used as a target cell for calibration, the remaining 2 high-quality reference cells and 5 suboptimal reference cells are used as reference cells for calibration; when one of the suboptimal reference cells is used as a target cell for calibration, the remaining 3 high-quality reference cells and 4 suboptimal reference cells are used as reference cells for calibration.
[0188] In some implementations, for the calibration of various high-abundance cells, considering that their relative cfDNA level Fraciton (or Fraction-adjust) in body fluid samples is much higher than that of internal control cells (at least 1-3 orders of magnitude), the calibration risk is relatively high. Therefore, for these high-abundance cells, internal control cell calibration is not used. The relative cfDNA content Fraction (or Fraction-adjust) obtained in S05 or S06 is used as the final cell Fraction for both the test sample and the healthy population sample.
[0189] In some exemplary embodiments, the high-abundance cells include, but are not limited to, blood cells, endothelial cells, and hepatocytes.
[0190] In some specific embodiments, the correction of the relative content of target cell cfDNA in the test sample using internal reference cells in S07-03 includes the following steps:
[0191] ① Remove high-quality internal reference cells from the test sample that are higher than the upper limit threshold of the ratio of high-quality internal reference cells, and further remove internal reference cells that are higher than the absolute upper limit threshold, considering that these internal reference cells have been damaged;
[0192] ② If, after step ①, the remaining reference cells contain the high-quality reference cells and at least one high-quality reference cell has a Fraciton (or Fraction-adjust) value above the absolute lower limit threshold, then these high-quality reference cells are used for calibration. The calibration formula is as follows:
[0193]
[0194] In Formula 5, the target cell Fraciton (or Fraction-adjust) refers to the relative content of target cell cfDNA in the test sample obtained from S05 or S06, and the Fraction of internal reference cells in the healthy population sample. j (or Friction-adjust) j() represents the 85th to 99th percentile of the relative content of cfDNA in the j-th type of internal reference cell in the healthy population sample, and () represents the Friction of the internal reference cells in the sample to be tested. j (or Friction-adjust) j ) represents the relative content of cfDNA in the j-th type of internal reference cell in the actual sample, where j is an integer from 1 to m, and m is the number of internal reference cell types involved in the calibration.
[0195] ③ If, after step ①, the remaining internal control cells contain high-quality internal control cells but the relative content of cfDNA (Fraciton or Fraction-adjust) is below the absolute lower limit threshold, it is considered that the high abundance of background cells in the sample is severely "squeezing out" the target cells, making it impossible to meet the accuracy requirements for actual target cell quantification. In this case, all internal controls (including high-quality and suboptimal internal controls) with relative cfDNA content below 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, or 1% should be corrected according to Formula 5 to minimize correction bias.
[0196] ④ If, after step ①, there are no high-quality internal reference cells remaining in the internal reference cells, it is considered that the sample has suffered multi-organ damage. In this case, the internal reference ratio is distorted and cannot be further corrected based on the internal reference ratio. Therefore, the Fraction (or Fraction-adjust) obtained in S05 or S06 is used as the relative content of cfDNA in the final target cells for both the test sample and the healthy population sample.
[0197] This invention also provides a method for tracing and detecting organ damage. The method includes obtaining the relative content of target cell cfDNA (Fraction) in ≥10 healthy human samples using a quantitative method for the relative content of cfDNA in the cells. The 85th to 99th percentile of the target cell Fraction in the healthy human samples is statistically determined as a damage threshold. If the relative content of target cell cfDNA (Fraction) in the test sample is greater than or equal to the damage threshold, the target cell is determined to be damaged. The degree of damage is positively correlated with the Fraction value of the test sample. The organ from which the target cell originates is further determined to be damaged.
[0198] definition
[0199] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly used in the field to which this invention pertains. For the purposes of interpreting this specification, the following definitions will apply, and where appropriate, terms used in the singular will also include the plural forms, and vice versa.
[0200] Unless the context clearly indicates otherwise, the terms “a” and “an” as used herein include plural references.
[0201] The term "about" as used herein is as understood by one of ordinary skill in the art and varies within a certain range depending on the context in which it is used. If one of ordinary skill in the art is unfamiliar with the use of this term in the context in which it is used, "about" will mean a particular value plus or minus 10%.
[0202] The term "cfDNA" as used in this article refers to cell-free DNA, including DNA molecules that are naturally present in the subject in an extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum). Although cfDNA originally originates from one or more cells in a large, complex biological organism (e.g., a mammal), it undergoes a process of release from the cell into the fluid present in the organism, and can therefore be obtained by obtaining a sample of the fluid without requiring an in vitro cell lysis step.
[0203] In this article, the term "methylation" refers to the presence of a methyl group at a position in a nucleotide base that is not typically found in a typical nucleotide base. For example, cytosine does not contain a methyl moiety on its pyrimidine ring, but 5-methylcytosine does contain a methyl moiety at position 5 of its pyrimidine ring. In this case, cytosine is not a methylated nucleotide, while 5-methylcytosine is. In another example, thymine contains a methyl moiety at position 5 of its pyrimidine ring. In this article, when thymine is present in DNA, it is not considered a methylated nucleotide because thymine is a typical nucleotide base of DNA. The typical nucleoside bases of DNA are thymine, adenine, cytosine, and guanine. The typical bases of RNA are uracil, adenine, cytosine, and guanine. Accordingly, a "methylation site" is a location in the nucleic acid region of a target gene where methylation occurs or may occur. For example, a site containing CpG is a methylation site where cytosine may or may not be methylated. A "site" can correspond to a single site, which is a single base position or a group of related base positions, such as a CpG site. A methylation site can refer to a CpG site or a non-CpG site in a DNA molecule that may be methylated.
[0204] As used herein, the term "CpG site" refers to a region in a DNA molecule in which a cytosine nucleotide is followed by a guanine nucleotide in a basic linear sequence along its 5' to 3' direction, and this region is susceptible to methylation of the nucleotide by events occurring naturally in vivo or by chemical action in vitro. A non-CpG site is a region that does not possess a CpG dinucleotide sequence but is also susceptible to methylation by events occurring naturally in vivo or by chemical action in vitro.
[0205] The term “methylation sequencing” in this article includes whole-genome bisulfite sequencing (WGBS), whole-genome enzymatic methylation sequencing, whole epigenome sequencing, methylation arrays, degenerate representative bisulfite sequencing (RRBS-Seq), TET-assisted pyridine borane sequencing (TAPS), Tet-assisted bisulfite sequencing (TAB-Seq), APOBEC-coupled epigenetic sequencing (ACE-seq), oxidized bisulfite sequencing (oxBS-Seq), precipitation or methylated DNA immunoprecipitation sequencing, or cytosine 5-hydroxymethylation sequencing (e.g., via Bluestar).
[0206] The term "Maximum Likelihood Estimation (MLE)" used in this paper refers to a statistical method used to determine the parameters of the relevant probability density function of a sample set. Given a random sample that follows a certain probability distribution, but with unknown parameters, parameter estimation involves conducting several trials, observing the results, and using those results to deduce approximate values for the parameters. Maximum likelihood estimation is based on the assumption that a certain parameter maximizes the probability of this sample occurring; therefore, other samples with lower probabilities are not considered, and this parameter is chosen as the estimated true value.
[0207] The term “abundance” as used in this article refers to the proportion of a particular cell type in a sample, including target cells, background cells, or internal control cells.
[0208] As used in this article, “sequencing depth” refers to the number of times a locus is covered by a total sequence reading corresponding to a unique nucleic acid target molecule (“nucleic acid fragment”) aligned with the locus.
[0209] In this application, the term "methylation marker" refers to a parameter associated with one or more biomolecules (i.e., "gene methylation marker"), such as naturally or artificially synthesized nucleic acids.
[0210] As used in this article, the term “sequencing” refers to the process of determining the sequence (e.g., the identity and order of monomeric units) of a biomolecule, such as a nucleic acid, like DNA or RNA. Exemplary sequencing methods include, but are not limited to, targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, double-strand sequencing, cyclic sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, low denaturing temperature co-amplification PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, synthetic sequencing, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa genome analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, DNA nanosphere sequencing (DNBSEQ), complex probe anchored polymerization sequencing (cPAS), and combinations thereof. In some implementations, sequencing can be performed using a gene analyzer, such as those commercially available from Illumina, Inc., Pacific Biosciences, Inc., Applied Biosystems / Thermo Fisher Scientific, or BGI Genomics Co., Ltd. Examples include BGI's DNBseq sequencing platforms such as BGISEQ-500, BGISEQ-50, MGISEQ-2000, MGISEQ-200, DNBSEQ-T7, DNBSEQ-G99, and DNBSEQ-T20X2, or Illumina's HiSeq2000, HiSeq2500, HiSeq4000, HiSeqX10, and NovaSeq6000.
[0211] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention in any way. The actual scope of protection of this invention is set forth in the claims. In the following description, descriptions of well-known structures and techniques are omitted to avoid unnecessarily obscuring the concepts of this disclosure. Such structures and techniques have also been described in many publications. Unless otherwise specified, the equipment, instruments, reagents, and / or kits used in the following embodiments are commercially available or obtained through conventional methods known to those skilled in the art.
[0212] Example 1: Relative Quantification Method
[0213] 1. Screening of methylation markers.
[0214] The cell methylation data were obtained from various whole-genome bisulfite sequencing (WGBS) data disclosed in (GSE186458), involving 41 cell types in 28 organs, as shown in Table 1.
[0215] Table 1. 28 organs and 41 cell types
[0216]
[0217]
[0218] The screening of cell-specific methylation biomarkers is as follows: 1. Methylation region segmentation: Using whole-genome methylation sequencing data of various cell types publicly available in the GSE186458 dataset as input, methylation regions across the entire genome were segmented using the wgbstools segmentation tool for CpG sites. Regions containing at least 3 CpG sites with a distance of <200 bp between adjacent CpG sites were further extracted as target regions for subsequent methylation biomarker screening; 2. Initial screening of methylation biomarkers based on methylation rate: Using any one of the 41 cell types as target cells and the rest as background cells, wgbstools segmentation was performed. The `find_markers` tool screens 41 cell types for methylation markers based on methylation rate in the target screening region. The screening criteria require that the average methylation rate of the target cells be ≤0.33, the average methylation rate of the background cell population be ≥0.66, and the difference between the 2.5 quantile of the methylation rate of all samples in the background cell population and the 75 quantile of the methylation rate of all samples in the target cell population is greater than 0.3, thus obtaining the initial screening methylation markers based on methylation rate. 3. Fine screening of all-unmethylated fragments: Duplicate samples of the same cell type are merged to increase the depth. Fragments containing at least 3 CpGs and being all-unmethylated within the methylation marker region of the previous step are considered positive fragments, and U-counts are calculated. All fragments containing at least 3 CpGs within the methylation marker region of the previous step are considered as total fragments, and ALL-counts are calculated. The ratio of positive fragments for each cell type within each methylation marker region of the previous step is used as the main indicator for subsequent fine screening of fragments. Using any one of the 41 cell types as target cells and the rest as background cells, the difference in the positive fragment ratio between the target cells and any background cell type was required to be at least greater than 0.2, and the positive fragment ratio of the target cells was required to be higher than 0.3, while the positive fragment ratio of each background cell type was required to be lower than 0.2. In addition, when various high-abundance blood cells were used as background cells, their positive fragment ratio was required to be lower than 0.1, in order to further reduce the background noise brought by high-abundance cells. Finally, the methylation marker regions of the 41 cell types after fragment screening were obtained.
[0219] Target-cell specific methylation markers were screened for each type of target cell, resulting in a total of 1466 methylation markers targeting 41 target cell types in 28 organs. (See attached text.) Figure 5 The positive fragments refer to completely unmethylated fragments containing at least three CpG sites within the methylation marker region of the target cell. The positive fragment ratio is the ratio of the number of positive fragments (U-counts) within the methylation marker region of the target cell to the total number of fragments (ALL-counts).
[0220] 2. Data Acquisition of the Sample to be Tested
[0221] Obtain methylation sequencing data of the sample to be tested, and screen methylation markers with a total fragment count U-counts ≥ 100 in the sample to be tested for subsequent steps.
[0222] 3. If two methylation markers belonging to the same target cell type, with adjacent sequencing positions on the same chromosome and a difference of <0.2 in the positive fragment ratio exist in the sample to be tested, they shall be combined and used as a single methylation marker;
[0223] 4. One cell type is selected alternately as the target cell, and the remaining 40 cell types are used as background cells. The noise coefficient bdg of each methylation marker is calculated according to Formula 1 for the k-th methylation marker out of a total of 1466 methylation markers, where n = 1466, i is an integer from 1 to 40, and the U-ratio_bg... ki By statistically analyzing the positive fragment ratio of methylation sequencing data of various cells in the GSE186458 dataset, the probability of each background cell releasing a positive fragment in each target cell methylation marker region was obtained.
[0224] 5. For each methylation marker, construct Formula 2 to calculate the Fraciton corresponding to each methylation marker in the sample to be tested, where k is an integer from 1 to 1466, and where U-ratio_target k This was obtained by statistically analyzing the positive fragment ratio of WGBS sequencing samples of various cells in the GSE186458 database, thereby determining the probability of each target cell releasing a positive fragment in each methylation marker region.
[0225] 6. Perform depth correction on the relative content of cfDNA in the 41 target cells according to Formula 3.
[0226] 7. Correct the relative content of cfDNA in various background cells according to Formula 4, and then use the corrected relative content of cfDNA in background cells. per-adjust i Substituting back into Equation 1, we obtain bdg-adjust by estimating the bdg weights. k Then adjust bdg-adjust k Substituting into Formula 2, we obtain Friction-adjust k Finally, Friction-adjust k Substituting into Formula 3, we obtain the relative content of cfDNA in the target cells using reaction-adjust.
[0227] 8. Internal reference calibration enables target cell signal restoration.
[0228] 8.1 Construction of the Healthy Baseline Cohort. A baseline cohort of 28 healthy individuals was established. cfDNA molecules were extracted from plasma samples of the baseline cohort. Probes were designed based on 1466 methylation marker regions. High-depth bisulfite sequencing (BS) (>500×) was performed on the target regions using the bisulfite conversion method to obtain fragment methylation information of the healthy baseline cfDNA molecules. A relative quantification model was then used to perform source quantification analysis on the healthy baseline, and the normal reference range (5th percentile-95th percentile) of the relative cfDNA content in each target cell was calculated and used as a benchmark for subsequent internal control selection and damage assessment.
[0229] 8.2 Construction of the internal control cell cohort. A stable group of high-quality internal control cells was selected from 41 cell types for target cell signal calibration, and suboptimal internal control cells that may pose a risk were used as a supplement.
[0230] The high-quality internal control cells were selected from the 41 cell types according to the following criteria: 1. Friction stability in healthy human samples: the maximum friction-adjustment of target cells in healthy human samples does not exceed 0.5%; 2. Cell quantification accuracy: the lower limit of quantification (LoQ) of target cells in simulated experiments at a sequencing depth of 500× is not higher than 0.2%; 3. High specificity of methylation marker selection: the ratio of positive fragments in the methylation marker region of target cells is higher than 0.5, and the U-ratio_bg of each background cell is lower than 0.1. Based on the plasma test results of 28 healthy individuals, three types of high-quality internal control cells were selected according to the above criteria: cerebral cortical neurons, cardiomyocytes, and oligodendrocytes.
[0231] The suboptimal internal control cells were selected from the 41 cell types according to the following criteria: 1. Friction stability in healthy human samples: the maximum friction-adjustment of target cells in healthy human samples does not exceed 0.5%; 2. Cell quantification accuracy: the lower limit of quantification (LoQ) of target cells in simulated experiments at a sequencing depth of 500× is not higher than 0.3%; 3. High specificity of methylation marker selection: the ratio of positive fragments in the methylation marker region of target cells is higher than 0.3 and the U-ratio_bg of each background cell is lower than 0.15. Based on the plasma test results of 28 healthy individuals, five types of suboptimal internal control cells were selected according to the above criteria: alveolar epithelial cells, bronchial epithelial cells, pancreatic duct cells, pancreatic alpha cells, and gallbladder cells, for supplementary use.
[0232] 8.3 Determination of Baseline Internal Reference Ratio / Threshold. Based on the healthy baseline cohort of 28 healthy individuals, internal reference cell signals were constructed for each sample, and the relative content of cfDNA in each type of internal reference cell (Fraction-adjustment) was calculated for each healthy individual. Furthermore, the ratio of high-quality internal reference cells and the upper threshold of the internal reference cell ratio were determined; where the upper threshold of the ratio of each high-quality internal reference cell is calculated as follows: Upper threshold = 95th percentile of the baseline Friction (or Friction-adjustment) value of that high-quality internal reference cell type / (95th percentile of the baseline Friction (or Friction-adjustment) value of that high-quality internal reference cell type + 5th percentile of the baseline Friction (or Friction-adjustment) value of the remaining high-quality internal reference cell types). Absolute thresholds were constructed to further determine the degree of internal reference cell damage. The upper absolute threshold used the 95th percentile of each of the three internal reference cell types in the baseline Friction-adjustment. The lower absolute threshold used the 5th percentile of each of the three internal reference cell types in the baseline Friction-adjustment. The calculation results are shown in Table 2.
[0233] Table 2. Baseline Internal Reference Thresholds and Absolute Thresholds
[0234]
[0235] 8.4 Use internal control cells to correct the relative content of target cell cfDNA in the test sample, including:
[0236] 8.4.1 Remove high-quality internal control cells that exceed the upper limit threshold of the ratio of high-quality internal control cells, and remove internal control cells that exceed the absolute upper limit threshold;
[0237] 8.4.2 If the remaining cells contain high-quality internal controls and the signal fraction-adjustment of at least one high-quality internal control cell is above the absolute lower limit threshold, then these high-quality internal controls shall be used for correction according to Formula 5.
[0238] 8.4.3 If the remaining components contain high-quality internal control cells but the signal Fraciton-adjustment is below the absolute lower limit threshold, then all internal controls with Fraciton-adjustment below 1% (including high-quality and suboptimal internal controls) should be corrected according to Formula 5.
[0239] 8.4.4 If no high-quality internal control cells are available for the remaining components, internal control cells will not be used for calibration. The Friction-adjust obtained in step 6 will be used as the relative cfDNA content of the final target cells for both the test sample and the healthy population sample.
[0240] 8.4.5 When internal control cells are used as target cells for calibration, the cells themselves are first removed from the internal control. Then, the remaining high-quality and suboptimal internal control cells are used for calibration according to steps 8.4.1 to 8.4.4.
[0241] 8.4.6 For high-abundance cells such as blood cells, endothelial cells, and hepatocytes, where the relative level of cfDNA (Fraction-adjust) is much higher than that of internal control cells (approximately 1-3 orders of magnitude), internal control cells are not used for correction. The Friction-adjust obtained in step 6 is used as the relative cfDNA content of the final target cells for both the test sample and the healthy population sample.
[0242] Example 2: Performance Verification of Relative Quantitative Method
[0243] Based on publicly available WGBS sequencing data of various cells, to simulate real-world cfDNA conditions, methylation data of various cell fragments were mixed in expected proportions. Under the premise of known mixing proportions for each cell type, the performance of the relative quantification method was validated.
[0244] Specific mixing simulation method: Except for blood cells, other cells were mixed in equal proportions at the following mixing gradients: 0.02%, 0.04%, 0.06%, 0.08%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1.0%, 1.1%, 1.2%, 1.3%, 1.4%, and 1.5%, for a total of 19 gradients. The remaining cfDNA was proportionally allocated to various blood cell types (B cells, T cells, monocytes / macrophages, granulocytes, erythrocyte progenitors, etc.). Two depth gradients of 500× and 750× were designed, with 20 replicates for each gradient.
[0245] Quantitative performance validation. Linear fitting was performed between the true mix rate and the actual relative quantitative results. Evaluation metrics included statistical slope, intercept, R², and CV (variability between repeated samples for each gradient). A satisfactory gradient was defined as one where the difference between the true mix rate and the actual relative quantitative results did not exceed 10%, and the CV between repeated samples was less than 20%. Quantitative performance results for important cell types are shown below. Figure 1 The quantitative lower limit LoQ for meeting the performance qualification gradient is shown in Table 3.
[0246] At 750× depth, the lower limit of quantification for target cells was between 0.02% and 0.3%, with most samples reaching 0.1%. The slope of cell types in important organs was above 0.97. 2 At a depth of 500×, the quantification limits for most target cells were above 0.99; at a depth of 500×, the quantification limits for most target cells were between 0.1% and 0.5%, and the slopes for cell types in important organs were all above 0.94.2 All values are above 0.98. The above performance demonstrates the rationality of the method model, and the accurate detection of trace target cell signals (ten-thousandths to thousandths) can be achieved at around 500×.
[0247] Table 3 Performance of Relative Quantitative Detection Methods for Organ Injury
[0248]
[0249]
[0250] To further verify that the method can still achieve accurate quantification of known cell types even in the presence of high-abundance unknown cell components, this embodiment simulates the unknown proportions of high-abundance Natural Killer Cells and Mono-Macrophages Cells to test the quantification accuracy of low-abundance target cells. Specifically, the following steps were taken: methylation markers related to two types of simulated unknown cell types were removed from the probe group, and two cell fragments were mixed in at the expected proportions. The remaining steps simulated the specific implementation method and were compared with the method without step 2). The results are shown in Table 4. At 750× depth, the lower limit of target cell quantification after correction for the unknown component proportion was between 0.04% and 0.3%; at 500× depth, the lower limit of quantification for most target cells was between 0.1% and 0.6%. This indicates that the target cell quantification performance did not significantly decrease after noise reduction optimization, verifying that the method can still accurately detect trace signals of target cells even in the presence of unknown components.
[0251] Table 4. Performance Comparison Before and After Correction for Unknown Component Ratios
[0252]
[0253] Based on the above relative quantitative system, the relative proportion ranges of various important cell types at the healthy baseline were further analyzed, as shown in Table 5. The relative proportions of cells in most solid organs were mostly in the ten-thousandth to thousandth percentiles. Additionally, the relative proportions of liver and endothelial cells in most samples were between 1% and 5%, which were approximately 1-2 orders of magnitude higher than the signals of other solid organs / cells. Further analysis of leukocyte components involved merging and quantitatively analyzing 23 leukocyte WGBS sequencing data downloaded from publicly available open-source databases. Figure 2 As shown. The main cell types and their proportions were analyzed as follows: granulocytes (52.8%), lymphocytes (31.7%), and monocytes / macrophages (8.6%), consistent with the results of routine blood tests. The relative proportion of white blood cells at the healthy baseline fluctuates considerably, with significant differences between individuals, and may exhibit further differences during the development of the disease.
[0254] Table 5. Intrinsic reference damage thresholds at healthy baselines (95th percentile)
[0255]
[0256] Example 3: Internal reference calibration and performance verification in real samples of kidney injury
[0257] In this embodiment, real plasma cfDNA data was used for internal reference signal restoration and clinical performance verification. The kidney was selected as the target organ, and renal epithelial cells were selected as the target cells. This embodiment used data from 13 cases of systemic lupus erythematosus (SLE) for renal cell signal internal reference correction and clinical performance verification. The analysis process is as follows: Figure 3 .
[0258] Healthy baseline construction. Methylated (BS) high-depth sequencing (>500×) was performed on the target probe region of the healthy baseline cohort. After data preprocessing and relative quantification, the internal reference correction results for renal epithelial cells of the healthy baseline were obtained, such as... Figure 4 As shown, the baseline-corrected signal robustness of the renal epithelium was excellent, with most samples remaining stable between 0.1% and 0.3%. Subsequently, the clinical performance of the damaged samples was analyzed using the 95th percentile damage threshold (0.26%) of the renal epithelium.
[0259] Clinical performance analysis of kidney injury samples. All samples were divided into three groups: active SLE (7 cases), inactive SLE (6 cases), and healthy baseline. Figure 4 As shown in the figure, the results of the internal control-corrected quantitative analysis showed that, at a baseline specificity of 95%, SLE active phase samples achieved 100% sensitivity detection (7 / 7), and at a non-active phase specificity of 95%, SLE active phase samples achieved 85.7% sensitivity detection (6 / 7), demonstrating the high sensitivity of the organ injury quantitative method in the context of kidney injury application. For non-active phase samples, the injury signal was significantly reduced compared to active phase samples, but still showed a certain upward trend compared to the healthy baseline level, indicating that further management and monitoring are still needed for non-active phase cases. Overall, the validation of the relative quantitative detection method can provide important reference for the diagnosis and treatment of kidney injury.
[0260] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.
Claims
1. A method for quantifying cfDNA derived from target cells in a sample for non-disease diagnostic purposes, the method comprising the following steps: 1) Based on the methylation sequencing data of the sample to be tested, obtain sequencing fragments of n methylation markers, count the total number of fragments and the number of positive fragments for each methylation marker, and screen methylation markers with a total number of fragments ≥ z for subsequent steps, where z is an integer from 10 to 200. The sample to be tested includes cfDNA from x types of cells, and the x types of cells include n methylation markers. Positive fragments are fragments containing at least 3, at least 4, at least 5, or at least 6 completely unmethylated CpG sites. 2) Using one type of cell in the test sample as the target cell and the remaining x-1 types of cells as background cells, based on the prior probability of each methylation marker releasing a positive fragment in the target cell, the total number of fragments of each methylation marker in the test sample is determined. ALL-counts and number of positive fragments U-count The relative cfDNA content of each methylation marker in the target cell was obtained by using the maximum likelihood MLE algorithm, along with the weight of the background cells contributing noise to each methylation marker. The maximum likelihood MLE algorithm in step 2) is shown in Equation 2: formula 2 ; In formula 2: k is an integer from 1 to n. U-counts k The number of positive fragments for the k-th methylation marker; ALL-counts k The total number of fragments for the k-th methylation marker; bdg k The weights for the noise signal contributed by the background cells to the k-th methylation marker; U-ratio_target k This is a priori probability of the release of a positive fragment of the k-th methylation marker of the target cell obtained from a cell sample of the target cell; wherein... U-ratio_target k This was obtained by statistically analyzing the positive fragment ratios of WGBS sequencing samples from various cells in the GSE186458 database, thereby determining the probability of each target cell releasing a positive fragment in each methylation marker region. P k For the expected value, take P k When the value is at its maximum Fraction k The relative content of cfDNA in the target cell corresponding to the kth methylation marker; The bdg k Calculated according to Formula 1: Formula 1 ; In Formula 1: k is an integer from 1 to n. per i The relative content of cfDNA in the i-th background cell among x-1 background cells in the sample to be tested. i is an integer from 1 to (x-1); U-ratio_bg ki This is a priori probability of the release of a positive fragment of the k-th methylation marker in a cell sample of the i-th background cell type; wherein... U-ratio_bg ki By statistically analyzing the positive fragment ratio of methylation sequencing data of various cells in the GSE186458 dataset, the probability of each background cell releasing a positive fragment in each target cell methylation marker region was obtained. The per i It was obtained through the deconvolution algorithm; 3) Using the total number of fragments of each methylation marker in the test sample as a depth coefficient, perform a first correction on the relative cfDNA content of each methylation marker corresponding to the target cell, to obtain the relative cfDNA content of the target cell in the test sample. Fraction ; The methylation markers include genomic regions in the target cells where the proportion of unmethylated fragments is significantly higher or lower than that in the background cells. Where x is an integer greater than 1, and n is an integer greater than or equal to x. Each cell type includes one or more methylation markers.
2. The method according to claim 1, characterized in that, The method further includes: 4) Repeat steps 2) and 3) to obtain the relative content of cfDNA in x types of target cells.
3. The method according to claim 1, characterized in that, The weight of the noise signal contributed by the background cells to each methylation marker is obtained based on the prior probability of each background cell releasing a positive fragment for each methylation marker, and the statistical analysis of the relative content of cfDNA of each background cell in the sample to be tested.
4. The method according to claim 1, characterized in that, The relative cfDNA content of each background cell is obtained based on the methylation sequencing data of the sample to be tested, using a deconvolution algorithm.
5. The method according to any one of claims 1 to 4, characterized in that, Following step 4), the method further includes the following steps: i) Based on the sum of the relative cfDNA contents of the x target cells, perform a second correction on the relative cfDNA contents of each background cell to obtain the corrected relative cfDNA contents of each background cell; and ii) Using the corrected relative cfDNA content of each background cell obtained in step i), repeat steps 2) to 3) to perform a third correction on the relative cfDNA content of the target cells in the test sample.
6. The method according to claim 5, characterized in that, Prior to step 1), the method further includes the step of obtaining one or more methylation markers for each of the x types of cells.
7. The method according to claim 6, characterized in that, The step of obtaining one or more methylation markers for each of the x cell types includes: Cellular methylation markers include genomic regions in which the positive fragment ratio of the cell type is at least 0.2 greater than that of background cells; the positive fragment ratio is the ratio of the number of positive fragments in the genomic region to the total number of fragments.
8. The method according to claim 6, characterized in that, The methylation markers include co-methylation markers, which are obtained by combining two methylation markers located adjacent to each other on the same chromosome in the same cell and with a positive fragment ratio difference of <0.
2.
9. The method according to claim 8, characterized in that, The merging includes summing the total number of fragments from the two methylation markers and / or summing the number of positive fragments from the two methylation markers.
10. The method according to claim 5, characterized in that, The method further includes the following steps: A fourth correction was performed on the relative cfDNA content of the target cells in the test sample using an internal reference cell set with stable relative cfDNA content in healthy human samples.
11. The method according to claim 10, characterized in that, The target cells include t methylation markers, where t is an integer between 1 and n. The first correction in step 3) is performed according to formula 3: Official 3; In formula 3, k is an integer from 1 to t. ALL-counts k The total number of fragments for the k-th methylation marker; Fraction k The relative content of cfDNA in the target cell corresponding to the k-th methylation marker.
12. The method according to claim 11, characterized in that, The second correction in step i) includes: According to Formula 4, the above... per i Correction was performed to obtain the corrected relative content of cfDNA for each background cell type. per- adjust i , Official 4; Where i is an integer from 1 to (x-1).
13. The method according to claim 12, characterized in that, The third correction includes: based on the per- adjust i Formulas 1, 2, and 3 are used for correction to obtain the corrected relative content of cfDNA in the target cells. Fraction-adjust .
14. The method according to claim 13, characterized in that, The internal reference cell set is selected from any one or more of the x types of cells.
15. The method according to claim 14, characterized in that, The internal reference cell set includes a high-quality internal reference cell set and a suboptimal internal reference cell set.
16. The method according to claim 15, characterized in that, The high-quality internal control cell set is defined as a sample of ≥10 healthy individuals. Fraction All values were no more than 0.5%, the lower limit of quantification (LoQ) was no higher than 0.2%, and the positive fragment ratio of methylation markers was higher than 0.5 and / or U-ratio_bg One or more cells, all below 0.
1.
17. The method according to claim 15, characterized in that, The suboptimal internal control cell set is defined as a sample of ≥10 healthy individuals. Fraction All values were no more than 0.5%, the lower limit of quantification (LoQ) was no higher than 0.3%, and the positive fragment ratio of methylation markers was higher than 0.3 and / or U-ratio_bg One or more cells, all below 0.
15.
18. The method according to claim 15, characterized in that, The high-quality internal control cell set and the second-best internal control cell set do not overlap.
19. The method according to claim 15, characterized in that, The fourth correction includes: a) Statistically analyze each type of high-quality internal reference cell in the test sample. Fraction The ratio, whereby each high-quality internal control cell... Fraction / All high-quality internal reference cells Fraction The sum; b) If a high-quality internal reference cell is present in the sample to be tested... Fraction If the ratio is higher than the upper limit threshold for the ratio of high-quality internal control cells, then the high-quality internal control cells are removed from the set of high-quality internal control cells, and / or removed from the set of high-quality internal control cells. Fraction Higher than absolute Fraction High-quality internal control cells with upper limit threshold, Wherein, the upper limit threshold of the ratio of high-quality internal reference cells = the ratio of high-quality internal reference cells Fraction The value is in the 85th-99th percentile of the healthy population sample / (the value of the high-quality internal reference cells). Fraction Values in the 85th-99th percentile of the healthy population sample + other high-quality internal control cells Fraction The value is in the 1st to 15th percentile of the healthy population sample. The absolute Fraction The upper limit threshold is the percentage of the high-quality internal reference cells in the healthy population sample. Fraction The 85th to 99th quantiles; c) If, after step b), the concentration of high-quality internal reference cells still exists... Fraciton In absolute Fraction For m types of high-quality internal control cells that are above the lower threshold, the target cells are corrected according to Formula 5. Fraction To obtain internal control correction target cells Fraction : Official 5; In formula 5, j=1~m, m is an integer from 1 to x. Internal control cells from healthy individuals Fraction j For the j-th type of internal reference cells obtained based on the healthy population sample Fraciton The 85th to 99th percentile values in a healthy population sample. Internal control cells of the test sample Fraction j For the j-th type of internal reference cells obtained based on the test sample Fraciton ; d) If, after step b), the concentration of the high-quality internal reference cells still exists... Fraciton In absolute Fraction For high-quality internal control cells below the lower threshold, the relative content of cfDNA in the test sample is used. Fraction Among the m types of internal control cells, including high-quality and sub-high-quality internal control cells at concentrations below 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, or 1%, the target cells were calibrated according to Formula 5. Fraction To obtain internal control correction target cells Fraction ; e) If no high-quality internal control cells are available after step b), then the fourth calibration is not performed using internal control cells; The Fraction The absolute lower limit threshold is the percentage of the high-quality internal reference cells in the healthy population sample. Fraction The 1st to 15th quantiles.
20. The method according to claim 19, characterized in that, When one internal reference cell is the target cell, the internal reference cell is not used for the fourth calibration, which is performed using other internal reference cells in the set of internal reference cells.
21. The method according to claim 1, characterized in that, The methylation sequencing includes one or more of whole-genome methylation sequencing, multiplex PCR methylation sequencing, and / or targeted methylation sequencing.
22. The method according to claim 1, characterized in that, The methylation sequencing depth is ≥400×, ≥500×, ≥600×, or ≥700×.
23. The method according to claim 1, characterized in that, The samples to be tested or samples from healthy individuals include bodily fluid samples.
24. The method according to claim 23, characterized in that, The fluid samples include: saliva or saliva supernatant, serum, plasma, milk, urine, lumbar or ventricular CSF, lymph, prostatic fluid, semen, sputum, diluted fecal supernatant, tears, bronchoalveolar lavage fluid, sputum, pus, nasopharyngeal swabs, oral swabs, cerebrospinal fluid, pleural effusion, peritoneal fluid, amniotic fluid, peritoneal fluid, aqueous humor, vitreous humor, diluted vaginal discharge supernatant, and their processed forms.
25. A quantitative system for the relative content of cfDNA in cells, characterized in that, The quantitative system is used to implement the method described in claims 1-24, and the system comprises: The positive fragment acquisition module counts positive fragments within the region of each methylation marker based on methylation sequencing data; it then filters methylation markers in the sample to be tested whose total fragment count U-counts ≥ z, where z is an integer from 10 to 200. The background weight acquisition module sums up the probability of cfDNA from each background cell in the biological sample releasing positive fragments in the region of the target cell-specific methylation marker to obtain the background cell weight. The module for obtaining the relative content of methylation biomarkers cfDNA constructs a likelihood formula to calculate the relative content of cfDNA in target cells corresponding to each methylation biomarker. The module for obtaining the relative content of cfDNA in target cells calculates the weighted average of all methylation markers specific to the same target cell, using the depth coefficient of each methylation marker as the weight, and uses this average as the relative content of cfDNA for each target cell.
26. The quantitative system according to claim 25, characterized in that, The quantitative system also includes correction for the proportion of unknown cell components and statistical analysis of the relative content of cfDNA in all target cells. Fraction The sum of the values is calculated, and the relative content of cfDNA for each type of background cell in the background weight acquisition module is corrected.
27. The quantitative system according to claim 25, characterized in that, The system also includes a methylation marker combination module: merging methylation markers with a positive fragment ratio difference ≤0.2 and adjacent positions on the chromosome of the same target cell into a combined methylation marker.
28. A method for tracing the source of organ damage for non-disease diagnosis purposes, the method comprising obtaining the relative content of target cell cfDNA in ≥10 healthy human samples according to any one of claims 1 to 24, statistically obtaining the 85th to 99th percentile of the relative content of target cell cfDNA in the healthy human samples as a damage threshold, and if the relative content of target cell cfDNA in the test sample is ≥ the damage threshold, then the target cell is determined to be damaged, the degree of damage is positively correlated with the relative content of target cell cfDNA in the test sample, and the organ from which the target cell originates is further determined to be damaged.
29. An organ injury tracing and detection system, characterized in that, The detection system is used to implement the organ injury tracing method of claim 28, and the system includes: The damage threshold construction module uses the relative content of target cell cfDNA in ≥10 healthy human samples. Fraction The 85th to 99th quantiles are used as the damage threshold; The cell damage assessment module determines the relative content of target cell cfDNA in the sample being tested. Fraction If the value is greater than or equal to the damage threshold, then the target cell is considered to be damaged. The organ damage assessment module further determines whether the organ is damaged based on the organ from which the target cells originate.
30. An apparatus comprising a memory for storing a program; and a processor for implementing the method of claims 1-24, or the organ injury tracing method of claim 28, by executing the program stored in the memory.
31. A computer-readable storage medium storing a program that can be executed to implement the quantitative method of claims 1-24, or the organ injury tracing method of claim 28.
32. The application of the method described in claims 1 to 24 or the organ damage tracing method described in claim 28 in cell / organ tracing and cell / organ damage detection for non-disease diagnosis purposes.
33. The application according to claim 32, characterized in that, The applications include non-disease diagnostic purposes such as cell / organ health assessment, disease progression monitoring, recurrence prediction, and prognostic evaluation.
Citation Information
Patent Citations
Evaluation system for predicting tissue specific source and related disease probability of cfDNA and application
CN113539355A
Multi-cancer-species methylation detection kit and application thereof
CN117535404A