Methods for determining the target detection frequency of target signal molecules in a sample to be tested

CN122811341APending Publication Date: 2026-09-25SHANGHAI WEIHE MEDICAL LAB CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510321271.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2026-09-25

Smart Images

  • Figure CN122811341A_ABST
    Figure CN122811341A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method for determining a target frequency of a target signal molecule in a sample under test. The method comprises target enrichment of a target site of the sample under test based on probe hybridization capture. The method comprises performing sequencing on the enriched sample under test using high-throughput sequencing to obtain at least one detection frequency of the target signal molecule at the at least one target site. The method comprises determining a median of the at least one detection frequency as a detection frequency to be calibrated. The method comprises determining the target frequency of the target signal molecule in the sample under test based on the detection frequency to be calibrated and an enrichment coefficient associated with the target enrichment. In this way, the accuracy of the target signal molecule proportion evaluation of the sample under test is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gene detection technology, and in particular to a method for determining the proportion of target signal molecules in a sample to be tested. Background Technology

[0002] Signal molecule proportion assessment has important applications in the diagnosis and treatment of various diseases, such as tumor detection, autoimmune diseases, orthopedic diseases, cardiovascular diseases, and neurodegenerative diseases. In tumor-related research or clinical practice, tumor-derived signal molecule proportion assessment is a quantitative method for evaluating the proportion of tumor tissue in the entire diseased tissue, organ, or body structure. For example, tumor proportion assessment refers to calculating the content of tumor-originating deoxyribonucleic acid (DNA) in a sample, such as in blood or tissue, the ratio of tumor-originating DNA to normal tissue / cell DNA. For example, in blood, tumor proportion is usually expressed as a percentage of circulating tumor DNA (ctDNA) to cell-free DNA (cfDNA) in the sample. Tumor-derived signal molecule proportion assessment can be used for: 1. early cancer screening; 2. monitoring treatment response; 3. detecting minimal residual disease; 4. guiding personalized treatment decisions, etc. Summary of the Invention

[0003] In a first aspect of this disclosure, a method is provided for determining the target frequency of a target signal molecule in a sample to be tested, comprising: targeting and enriching at least one target site in the sample to be tested based on probe hybridization capture; performing sequencing on the targeted enriched sample using high-throughput sequencing to obtain at least one detection frequency of the target signal molecule at at least one target site; determining the median of the at least one detection frequency as a detection frequency to be calibrated; and determining the target frequency of the target signal molecule in the sample to be tested based on the detection frequency to be calibrated and an enrichment coefficient associated with the targeting enrichment.

[0004] In some embodiments, the target signaling molecule includes at least one of the following: methylated haplotype, mutant haplotype, or fusion gene haplotype.

[0005] In some embodiments, the target site of the methylation haplotype includes one or more CpGs.

[0006] In some embodiments, the state of CpG includes methylation and demethylation.

[0007] In some embodiments, mutated haplotypes include base mutations (SNVs), insertions, and deletions.

[0008] In some embodiments, a fusion haplotype is formed by a structural variation process in which all or part of the sequences of two genes are fused together to form a new gene. The structural variation process includes the result of chromosomal translocation, intermediate deletion, or chromosomal inversion.

[0009] In some embodiments, the detection frequency in at least one detection frequency has a range of 0 to 1.

[0010] In some embodiments, the target frequency of signal molecules in the sample to be tested, based on the detection frequency to be calibrated and the enrichment coefficient associated with target enrichment, and the proportion of signal molecules in the sample to be tested, includes: determining the correspondence between the detection frequency to be calibrated and the target frequency based on the enrichment coefficient; and determining the target frequency based on the detection frequency to be calibrated and the correspondence.

[0011] In some embodiments, the method further includes: enriching multiple reference sites of multiple reference samples with different reference frequencies using targeted enrichment to obtain multiple enriched reference samples; sequencing the multiple enriched reference samples using high-throughput sequencing to obtain multiple sets of detection frequencies of the multiple enriched reference samples at multiple reference sites; determining multiple medians of the multiple sets of detection frequencies as multiple median detection frequencies; and determining an enrichment coefficient by fitting based on the multiple reference frequencies of the multiple reference samples and the multiple median detection frequencies.

[0012] In some embodiments, determining the correspondence between a detection frequency to be calibrated and a target frequency based on enrichment coefficients includes: determining a first representation of the detection frequency to be calibrated based on the enrichment coefficients and the target frequency; and transforming the first representation to obtain the correspondence between the target frequency and the detection frequency to be calibrated based on the enrichment coefficients.

[0013] In some embodiments, enriching at least one target site in a sample to be tested using targeted enrichment via probe hybridization capture includes: preparing a library for the sample to be tested to establish a pre-library; and performing targeted enrichment based on probe hybridization capture on the pre-library to obtain a final library.

[0014] In some embodiments, performing sequencing on at least one target site using high-throughput sequencing to obtain at least one detection frequency of the target haplotype at at least one target site includes: performing high-throughput sequencing on the final library to obtain sequencing data; and analyzing the sequencing data to obtain multiple target detection frequencies.

[0015] In some embodiments, the probes used in the hybridization capture process include probes that hybridize as a whole with all haplotype characteristic sequences in the target DNA region.

[0016] In some embodiments, the probes include equal-length probes that simultaneously capture fully methylated and fully demethylated upper OT (Original Top) and lower OB (Original Bottom) chains, equal-length probes that specifically capture fully methylated or fully demethylated OT and OB chains, and short probes of unequal lengths that specifically capture methylated haplotypes.

[0017] In some embodiments, the probe has a length ranging from 10 nt to 2000 nt.

[0018] In some embodiments, the probe includes at least one of the following: a DNA probe, an RNA probe, a modified base LNA, a modified backbone PNA, or a probe-protein complex capture.

[0019] In some embodiments, the number of probes ranges from 1 to 10. 12 Within the range.

[0020] In some embodiments, targeted enrichment is performed by an assay system comprising: one or more probes binding to at least one target site, blocking primers, a hybridization reaction solution, and a hybridization temperature for hybridization; and wherein the assay system comprises: a capture reaction solution, an elution solution, and a capture temperature for capture.

[0021] In some embodiments, the sample to be tested includes deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0022] In some embodiments, the sample to be tested includes circulating free DNA, cell line DNA, or tissue DNA.

[0023] In some embodiments, circulating cell-free DNA includes circulating cell-free DNA in plasma, circulating cell-free DNA in urine, circulating cell-free DNA in pleural effusion, circulating cell-free DNA in ascites, or circulating cell-free DNA in cerebrospinal fluid.

[0024] In some embodiments, high-throughput sequencing includes second-generation sequencing or third-generation sequencing.

[0025] It should be understood that the description in the Summary of the Invention section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0026] To better understand the above and other objects, features, advantages, and functions of the present invention, reference can be made to the preferred embodiments shown in the accompanying drawings. The same reference numerals in the drawings refer to the same parts. Those skilled in the art should understand that the drawings are intended to schematically illustrate preferred embodiments of the invention and do not limit the scope of the invention in any way; the parts in the drawings are not drawn to scale.

[0027] Figure 1A A schematic diagram of an experimental procedure that can be implemented therein according to embodiments of the present disclosure is shown;

[0028] Figure 1B A schematic diagram of an example methylated haplotype according to an embodiment of the present disclosure is shown;

[0029] Figure 2 A schematic diagram of an example method for determining the proportion of target signal molecules according to an embodiment of the present disclosure is shown;

[0030] Figure 3 A schematic diagram of an example method for determining enrichment coefficients according to embodiments of the present disclosure is shown;

[0031] Figure 4 A schematic diagram illustrating an example process for determining enrichment coefficients according to embodiments of the present disclosure is shown;

[0032] Figure 5 A schematic diagram illustrating the derivation process for determining the correspondence between the detection frequency to be calibrated and the target frequency according to an embodiment of the present disclosure is shown.

[0033] Figure 6 A schematic diagram illustrating an example process for determining the proportion of target signal molecules according to an embodiment of the present disclosure is shown;

[0034] Figure 7 A schematic diagram of an example calibration curve according to an embodiment of the present disclosure is shown;

[0035] Figure 8 A schematic diagram of an example frequency curve according to an embodiment of the present disclosure is shown; and

[0036] Figure 9 A schematic diagram of a calibration curve for verification according to an embodiment of the present disclosure is shown. Detailed Implementation

[0037] Unless otherwise indicated, the practice of this invention will employ conventional techniques from molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and synthetic biology, etc., which are within the scope of the art. Such techniques are well explained in the literature: "Molecular Cloning: A Laboratory Manual," 2nd edition (Sambrook et al., 1989); "Oligonucleotide Synthesis" (edited by M.J. Gait, 1984); "Animal Cell Culture" (edited by R.R. Freshney, 1987); "Methods in Enzymology" (Academic Press, Inc.); "Current Protocols in Molecular Biology" (edited by F.M. Usubel et al., 1987, and regularly updated); "PCR: The Polymerase Chain Reaction" (edited by Mullis et al., 1994); Singleton et al., Dictionary of Microbiology and Molecular Biology, 2nd edition, J. Wiley & Sons (New York, NY 1994); and March's Advanced Organic Chemistry Reactions, Mechanisms and Structure, 4th edition, John Wiley & Sons (New York, NY 1992); "Biophysics of DNA" (I. Tinoco... "Thermodynamics and Kinetics of DNA Melting" (edited by Santa Lucia et al., Biochemistry, 1996); "Molecular Biology" (edited by David P. Clark et al., Elsevier Inc., 2019); "The Basics of Molecular Biology" (edited by Alexander Vologodskii, Springer Cham, 2022) provides a general guide for those skilled in the art to many of the terms used in this application.

[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. For the purposes of this invention, the following terms are defined.

[0039] The articles “a” and “the” are used herein to refer to one or more (i.e., at least one) grammatical objects of the article “the”. The use of alternatives (e.g., “or”) should be understood to mean any, both, or any combination of the alternatives. The term “and / or” should be understood to mean any or both of the alternatives.

[0040] As used herein, the term “about” or “approximately” means a quantity, level, value, number, frequency, percentage, scale, size, amount, weight, or length that has changed by as much as 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% compared to a reference quantity, level, value, quantity, frequency, percentage, scale, size, amount, weight, or length.

[0041] Throughout this specification, unless the context otherwise requires, the terms "comprising," "including," "containing," and "having" shall be construed as implying the inclusion of the stated steps or elements or groups of steps or elements, but not excluding any other steps or elements or groups of steps or elements. In certain embodiments, the terms "comprising," "including," "containing," and "having" are used synonymously.

[0042] "Composed of" means including, but not limited to, anything that follows the phrase "composed of". Therefore, the phrase "composed of" indicates that the listed elements are required or mandatory, and that no other elements may exist.

[0043] "Substantially composed of..." means including any element listed after the phrase "substantially composed of..." and is limited to other elements that do not interfere with or contribute to the activities or actions specified in the public content of the listed elements. Thus, the phrase "substantially composed of..." indicates that the listed elements are necessary or mandatory, but no other elements are optional and may be present or absent depending on whether they affect the activities or actions of the listed elements.

[0044] Throughout this specification, references to "an embodiment," "some embodiments," "a specific embodiment," and similar expressions indicate that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in at least one embodiment of the invention. Therefore, the appearance of the aforementioned phrases in various places throughout this specification does not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics can be combined in any suitable manner in at least one embodiment.

[0045] The term "base" as used in this article, also known as a nucleobase or nitrogenous base, refers to nitrogenous compounds that form nucleosides and constitute the basic building blocks of nucleic acids. There are five common bases found in organisms: adenine (A), guanine (G), cytosine (C), thymine (T), and uracil (U), also known as classical bases. They are the basic units that make up the genetic code. Bases A, G, C, and T are found in DNA, while A, G, C, and U are found in RNA. Additionally, bases can also be non-classical bases. These bases are mostly derivatives formed by methylation or other chemical modifications at different sites of the aforementioned purine or pyrimidine bases, including, for example, hypoxanthine and xanthine.

[0046] As used in this article, “methylation” refers to the covalent transfer of a methyl group to a specific base by S-adenosylmethionine (SAM) as a methyl donor, under the action of methyltransferases. This typically occurs on cytosine (C), forming 5-methylcytosine (5-mC). The “CpG” site is the most common methylation site, but methylation sites are not limited to CpG sites. For example, DNA methylation can occur in the cytosine of CHG and CHH, where H is adenine, cytosine, or thymine. As used in this article, the term “CpG site” refers to a region in a DNA molecule where a cytosine nucleotide is followed by a guanine nucleotide in a linear sequence of bases along its 5' to 3' orientation. “CpG” is short for 5'-C-phosphate-G-3', which consists of cytosine and guanine separated by only one phosphate group. Multiple cytosines in a CpG dinucleotide can be methylated to form 5-methylcytosine.

[0047] A "biological sample" is any sample taken from an individual (e.g., a human, such as a cancer patient or suspected cancer patient) and containing one or more relevant nucleic acid molecules. Biological samples can be bodily fluids, such as blood, plasma, serum, urine, vaginal fluid, uterine or vaginal lavage fluid, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, etc. Fecal samples may also be used.

[0048] The terms "cell-free nucleic acid," "cell-free DNA," or "cfDNA" refer to multiple nucleic acid fragments that circulate in an individual's body (e.g., in the bloodstream) and originate from one or more healthy cells and / or one or more cancer cells. Furthermore, cfDNA may originate from other sources, such as viruses, fetuses, etc.

[0049] The term “circulating tumor DNA” or “ctDNA” refers to nucleic acid fragments originating from tumor cells that may be released into an individual’s bloodstream due to biological processes such as apoptosis or necrosis of dead or living tumor cells.

[0050] The method and probe of this invention can be used to diagnose diseases, particularly cancer and neurological disorders. Cancers can include melanoma, non-small cell lung cancer, small cell lung cancer, lung cancer, liver cancer, retinoblastoma, astrocytoma, glioblastoma, gingival cancer, tongue cancer, leukemia, neuroblastoma, head cancer, neck cancer, breast cancer, pancreatic cancer, prostate cancer, kidney cancer, bone cancer, testicular cancer, ovarian cancer, mesothelioma, cervical cancer, gastrointestinal cancer, lymphoma, brain cancer, colon cancer, sarcoma, or bladder cancer. Cancers can include tumors composed of tumor cells. Neurological disorders can include Alzheimer's disease, Parkinson's disease, or multiple sclerosis, etc.

[0051] As discussed above, assessing the proportion of signaling molecules with indicative functions is crucial. For example, assessing the proportion of disease-related molecules (e.g., tumors) is used in corresponding disease diagnosis. In related techniques, tumor proportion assessment methods may include, for example, detecting mutation frequency (VAF), which calculates the proportion of readings with mutations to the total number of readings. In this method, a higher VAF indicates a higher tumor proportion. Tumor proportion assessment methods may also include copy number variation (CNA). Tumors typically exhibit different copy number characteristics compared to normal tissues; therefore, tumor proportion can be estimated by assessing changes in DNA copy number. However, low tumor proportion scores can complicate mutation detection and quantification, especially in early-stage cancers or minimal residual disease. Biological differences and the presence of normal cfDNA can introduce noise into the data.

[0052] Besides mutation frequency and copy number variation (CNA)-based methods, current techniques also include methylation sequencing to provide more robust statistical methods for accurate estimation. However, methods for assessing tumor proportion in related technologies are still immature. For example, in some related technologies, the use of two-stranded, equal-length probes that specifically capture fully methylated or fully unmethylated OT and OB chains during hybridization capture can cause a shift in methylation haplotype frequencies. The amount of shift varies significantly depending on the probe's capture efficiency at different sites. For instance, a methylation haplotype site with 5% methylation in a real sample might be observed as having a 20% frequency in the hybridization capture data. Another methylation haplotype site with 5% methylation in a real sample might have an 80% frequency observed after hybridization capture. With hundreds or thousands of haplotype sites in a single experiment, this variation leads to inaccurate estimations of methylation haplotype frequencies, and consequently, inaccurate estimations of tumor proportion.

[0053] In view of this, embodiments of this disclosure propose a scheme for data calibration using the evaluation of enrichment coefficients for a specific assay system and the calculation of the median of the target haplotype. In this scheme, after target site enrichment and sequencing of the sample to be tested using an experimental system, and obtaining sequencing data, the median detection frequency of the sequencing data for multiple sites is determined. The target frequency of the target haplotype in the sample is calculated using this median detection frequency and the enrichment coefficient associated with the experimental system used, thereby evaluating the true signal molecule frequency of the sample as the proportion of the target molecule. In this way, by considering the enrichment coefficient, data bias caused by different enrichment efficiencies can be calibrated, thereby improving the accuracy of data processing.

[0054] The following will refer to Figures 1 to 12. Figure 9 This invention describes in detail a scheme for calculating the molecular proportion of an unknown target sample by measuring the frequency of a target haplotype in the sample according to embodiments of the present disclosure. Figure 1A A schematic diagram of a measurement process 100 that can be implemented according to an embodiment of this disclosure is shown. Figure 1A As shown, process 100 may include sample preparation and extraction. For example, high-quality genomic samples can be extracted from biological samples, such as cells, tissues, etc. Here, the integrity and purity of the extracted genomic samples are maintained as much as possible to avoid interference from impurities in subsequent experiments. This yields sample 102. Sample 102 can be any sample containing signaling molecules. Signaling molecules include methylated haplotypes, mutant haplotypes, and fusion gene haplotypes. Sample 102 may include, for example, deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). At 110, process 100 includes constructing a library for sample 102. Library preparation 110 may include, for example, genomic sample fragmentation. During genomic sample fragmentation, the genome can be fragmented into fragments of appropriate length using methods such as sonication or enzyme digestion. Library preparation 110 may also include, for example, end repair and A-tailing. In this process, the fragmented genomic fragments can be end-repaired to blunt ends, and then an A base can be added to the 3' end for subsequent ligation of sequencing adapters. Furthermore, during library preparation, process 100 includes performing bisulfite treatment on the sample to obtain a methylated library. During library preparation, process 100 also includes ligating sequencing adapters onto the sample.

[0055] At position 120, process 100 includes hybridization capture. In this process, a certain amount of the constructed methylated library can be taken. Hybridization capture includes, for example, vacuum concentration of the methylated library, hybridization, streptavidin magnetic bead capture and elution of the library, PCR amplification after elution, and magnetic bead purification of the amplification product. The purified product obtained can then be used for subsequent sequencing.

[0056] At point 130, process 100 includes high-throughput sequencing. In this process, the amplified library can be quantified and quality-checked. If it passes the quality check, it is sequenced using a high-throughput sequencing platform to obtain a large number of sequencing reads. After high-throughput sequencing 130, sequencing data 140 is obtained. Sequencing data 140 may include processed data, including the frequencies of methylation haplotypes at multiple sites. For example... Figure 1A As shown, sequencing data 140 includes multiple reads 142 at site 141, including read 142-1 indicating methylation at site 151, read 142-2 indicating unmethylation at site 141, read 142-3 indicating unmethylation at site 141, and read 142-4 indicating unmethylation at site 141. Then, at 150, for the detected sequencing data, data analysis is performed according to the method for determining the target frequency of methylation haplotypes according to embodiments of this disclosure. For example, the detection frequency of the target methylation haplotype at multiple sites can be calculated. Taking site 141 as an example, the methylation haplotype includes a CpG position. Here, "M" represents the methylation state and "U" represents the unmethylation state, so the target methylation haplotype can be "M". Since one reading indicating methylation and three readings indicating demethylation were observed at site 151, the methylation haplotype frequency of the target methylation haplotype "M" at that site is 1 / 4 = 25%. Similarly, the detection frequency of the target methylation haplotype "M" at a single site in the sample 102 can be obtained. Similarly, the detection frequencies of other methylation haplotypes at multiple sites in the sample 102 can be obtained. As mentioned above, due to data bias caused by enrichment efficiency in the performed sequencing process, 25% cannot be directly considered as the true frequency at that site. According to the scheme of this embodiment, the median of methylation haplotypes at all sites is selected, and the median of the detection frequency is calibrated using enrichment coefficients associated with the experimental system (also referred to as the experimental system or assay system) and the sequencing system (also referred to as the sequencing system) to obtain a target estimated frequency that can be considered as the true frequency. Here, the enrichment coefficient is associated with each step of the sequencing process 100. For example, the salt solution concentration used in the bisulfite treatment during library preparation at position 110, and the probe and salt solution concentration used in the hybridization capture process at position 120.

[0057] It should be understood that although the methylated haplotype in this embodiment includes only one CpG position, the scheme according to the embodiments of this disclosure is also applicable to methylated haplotypes with multiple CpG positions. Figure 1B A schematic diagram of example methylated haplotype 180 is shown.

[0058] like Figure 1B As shown, methylation haplotype 180 includes methylation haplotype 181 having a CpG position. Methylation haplotype 181 includes methylation haplotype 181-1 with methylation pattern "M". Methylation haplotype 181 also includes methylation haplotype 181-2 with methylation pattern "U".

[0059] Methylation haplotype 180 includes methylation haplotype 182, which has two CpG positions. Methylation haplotype 182 includes methylation haplotype 182-1 with methylation pattern "MM", methylation haplotype 182-2 with methylation pattern "MU", methylation haplotype 182-3 with methylation pattern "UM", and methylation haplotype 182-4 with methylation pattern "UU".

[0060] Methylation haplotype 180 includes methylation haplotype 183 with three CpG positions. Methylation haplotype 183 includes methylation haplotype 183-1 of methylation pattern “MMM”, methylation haplotype 183-2 of methylation pattern “UMM”, methylation haplotype 183-3 of methylation pattern “MUM”, methylation haplotype 183-4 of methylation pattern “MMU”, methylation haplotype 183-5 of methylation pattern “UUM”, methylation haplotype 183-6 of methylation pattern “UMU”, methylation haplotype 183-7 of methylation pattern “MUU”, and methylation haplotype 183-8 of methylation pattern “UUU”.

[0061] Similarly, methylation haplotype 180 may include methylation haplotype 184 having N CpG sites. Here, N is an integer greater than 3, and the number of methylation haplotype patterns corresponding to N CpG sites is 2^N.

[0062] The following will combine Figure 2 The following describes a scheme for determining the proportion of disease-related molecules in a sample according to embodiments of the present disclosure, using numbers 9 and 9. Figure 2 A schematic diagram of an example method 200 for determining a target frequency of methylation haplotypes according to an embodiment of this disclosure is shown. For discussion purposes, it will be combined with... Figure 1A To describe method 200. For example... Figure 2 As shown, at 202, method 200 includes targeted enrichment of at least one target site in the test sample based on probe hybridization capture. At 204, method 200 includes performing sequencing on the enriched test sample using high-throughput sequencing to obtain at least one detection frequency of the target signal molecule at at least one target site. For example, in Figure 1AIn the illustrated embodiment, the sequencing system may include all the equipment and reagents involved in performing the entire sequencing process 100. The sample to be tested may be a sample 102 comprising the genome. After performing the steps of library preparation 110, hybridization capture 120, and high-throughput sequencing 130, a detection frequency or multiple detection frequencies of the target methylation haplotype are obtained at one or more target sites.

[0063] At 206, method 200 includes determining the median of at least one detection frequency as the detection frequency to be calibrated. In some embodiments, the at least one detection frequency may include a single detection frequency. In such embodiments, the median of the detection frequencies is itself. Alternatively, the resulting at least one detection frequency may also include multiple detection frequencies.

[0064] At 208, method 200 includes determining the target frequency of the target haplotype in the test sample based on the detection frequency to be calibrated and the enrichment coefficient associated with target enrichment. The obtained target frequency can be used as the proportion of disease-derived signaling molecules in the test sample. In this way, by using the enrichment coefficient to calibrate the median of multiple detection frequencies of the target sample, highly accurate target frequency values ​​can be obtained, thereby assessing the proportion of disease-related molecular signals.

[0065] In some embodiments, the enrichment coefficient is determined for an experimental system with a predetermined target enrichment. Figure 3 A schematic diagram of an example method 300 for determining enrichment coefficients according to an embodiment of the present disclosure is shown. For discussion purposes, it will be combined with... Figure 4 The method 300 is described using embodiments. Figure 3 As shown, at 302, method 300 includes using targeted enrichment to enrich multiple reference sites of a reference sample to obtain a targeted-enriched sample. At 304, method 300 includes using high-throughput sequencing to sequence the targeted-enriched sample to obtain multiple detection frequencies at multiple reference sites. For example, in Figure 4 In the illustrated embodiment, the methylation haplotype frequencies of the reference sample were known, specifically 50% for the target methylation haplotype at site 1, 50% for site 2, and 50% for site 3. After obtaining the reference sample, it was sequenced to obtain detection frequencies of 57% for the target methylation haplotype at site 1, 94% for site 2, and 75% for site 3.

[0066] return Figure 3 At position 306, method 300 includes determining the median of multiple detection frequencies at multiple reference sites as the median detection frequency. For example, in Figure 4In the illustrated embodiment, the median of the three detection frequencies—57%, 94%, and 75%—is 75%. This means that the amount of "M" molecules is 75 units, while the amount of "U" molecules is 25 units.

[0067] At 308, method 300 includes acquiring multiple reference frequencies for multiple reference sites of a reference sample. At 310, method 300 includes determining the median of the multiple reference frequencies as the median reference frequency. For example, in Figure 4 In the illustrated embodiment, the median of the three reference frequencies at 50% is also 50%, meaning that the amount of both "M" and "U" molecules is 50 units.

[0068] Finally, at 312, method 300 includes determining the enrichment coefficients based on the median reference frequency and the median detection frequency. For example, in Figure 4 In the illustrated embodiment, the enrichment coefficient can be determined by the median of the detection frequency and the median of the reference frequency, i.e., (75 / 25) / (50 / 50) = 3. Figure 3 and Figure 4 In the illustrated embodiment, after extensive data observation, the inventors discovered that although the enrichment efficiency resulted in different target methylation haplotype frequencies at sites where the frequencies should be the same, the median of the detected methylation haplotype frequencies had a stable correspondence with the median of the predetermined or actual methylation haplotype frequencies under the same assay system or experimental system or system, i.e., under the same sequencing conditions. Thus, the enrichment coefficient associated with the sequencing system could be determined using reference samples with known frequencies at each site.

[0069] It should be understood that, although in Figure 3 and Figure 4 In the illustrated embodiment, enrichment coefficients are calculated for only one reference sample. Enrichment coefficients can also be obtained by fitting the detection frequencies of reference samples with different reference frequencies. During the fitting process, steps 302, 304, 306, 308, and 310 of method 300 can be performed for each reference sample.

[0070] In some embodiments, after obtaining the enrichment coefficient, the determination frequency of the target haplotype of the test sample can be determined based on the detection frequency to be calibrated and the enrichment coefficient associated with target enrichment. This process includes, for example, determining the correspondence between the detection frequency to be calibrated and the target frequency based on the enrichment coefficient; and determining the target frequency based on the detection frequency to be calibrated and the correspondence. In such embodiments, the correspondence between the detection frequency to be calibrated and the target frequency can be determined based on the enrichment coefficient. For example, the process of determining the correspondence between the detection frequency to be calibrated and the target frequency based on the enrichment coefficient may include: determining a first representation of the detection frequency to be calibrated based on the enrichment coefficient and the target frequency; and transforming the first representation to obtain the correspondence between the target frequency and the detection frequency to be calibrated based on the enrichment coefficient.

[0071] The following will combine Figure 5 This describes the derivation process of the calibration formula, which is a form of correspondence. Figure 5 A schematic diagram of an example process 500 for determining a calibration formula according to an embodiment of the present disclosure is shown. Figure 5 As shown, unknown sample 1 has an enrichment factor of 3. In Figure 5 In the illustrated embodiment, it is assumed that the median target detection frequency for sample 1 is 1%, which includes 1 unit of the target methylated haplotype "M" and 99 units of the remaining methylated haplotype "U". Conversely, after sequencing, the median detection frequency indicates an amount of 1 x 3 = 3 units of the target methylated haplotype "M" and 99 units of the remaining methylated haplotype "U". Therefore, the median detection frequency of the target methylated haplotype is 3 / (3+99) = 2.94%.

[0072] Similarly, when using a system with an enrichment coefficient k to detect an unknown sample N, if the median frequency of the target methylation haplotype in sample N includes i units of the target methylation haplotype "M" and 100-i units of the remaining methylation haplotypes "U", then the median frequency of the target haplotype in sample N is:

[0073] f1 = i% (1)

[0074] Where f1 represents the median of the target frequency.

[0075] In contrast, after targeted enrichment and subsequent sequencing, the median detection frequency indicates that the target methylation haplotype "M" has i*k units, and the remaining methylation haplotype "U" has 100-i units. Therefore, the median detection frequency of the target methylation haplotype should be f2 = i*k / (i*k + 100-i), i.e.

[0076] f2= i*k / (i*k+100-i) (2)

[0077] Where f2 is the median of the detection frequency, also known as the detection frequency to be calibrated. The amount of the target methylated haplotype "M" can be obtained from formula (1) as i = 100 * f1. Substituting this into (2) yields:

[0078] f2=(100*f1*k) / (100*f1*k+100- 100*f1) (3)

[0079] After simplification, we get:

[0080] f2=(f1*k) / (f1*k+1- f1) (4)

[0081] The result of the conversion is:

[0082] f1*f2*k+ f2- f2*f1=f1*k (5)

[0083] Moving the terms including f1 to one side of the equals sign, we get:

[0084] f1*k - f1*f2*k+ f2*f1=f2 (6)

[0085] Extracting f1 yields:

[0086] f1*(k + f2 - f2*k)=f2 (7)

[0087] The final result is:

[0088] f1= f2 / (k + f2 - f2*k) (8)

[0089] Figure 6 A schematic diagram of an example process 600 for determining a target frequency according to an embodiment of the present disclosure is shown. Figure 6 As shown, at position 602, based on Figure 5 The median target frequency f1 and median detection frequency f2 are derived from formulas (1) and (2) with respect to the enrichment coefficient k and the target value i, respectively, to obtain formula (8) for f1 with respect to k and f2. At 604, the actual median detection frequency f2' of the target haplotype of the unknown sample is determined. At 606, the median target frequency f1' of the target haplotype of the unknown sample is determined according to the derived formula (8), as the proportion of the signal molecules of the disease source in the sample to be tested.

[0090] In an embodiment where enrichment coefficients are determined through fitting, high-throughput sequencing can be used to sequence multiple enriched reference samples to obtain multiple detection frequencies at multiple reference sites. Then, multiple medians of these multiple detection frequencies can be determined as multiple median detection frequencies. Finally, enrichment coefficients can be determined through fitting based on the multiple reference frequencies and the multiple median detection frequencies of the multiple reference samples.

[0091] Figure 7 A schematic diagram of an example calibration curve 700 according to an embodiment of the present disclosure is shown. Figure 7 As shown, calibration curve 700 includes calibration curves 710, 720, 730, and 740 with enrichment coefficients of 50, 10, 5, and 1 obtained from the above fitting methods. When the observation frequency Y is 50, the calibrated or original target frequency X can be obtained as 16.6 from calibration curve 730.

[0092] Figure 8 A schematic diagram of an example frequency curve 800 according to an embodiment of the present disclosure is shown. Figure 8 As shown, after setting the horizontal axis to the detection frequency to be calibrated and the vertical axis to the target frequency, the frequency curve 810 is obtained according to formula (4).

[0093] Figure 9 A schematic diagram of a calibration curve 900 for verification according to an embodiment of the present disclosure is shown. Figure 9 The frequency curves in the figure are based on real experimental data. In this experiment, the samples were mixed samples of different gradients of the HCT116 commercial cell line, namely 0%, 20%, 40%, 60%, 80%, and 100% methylation standards. Hybridization capture experiments were performed using three experimental systems. Sequencing system 1 was unbiased capture, which used a four-stranded long probe to hybridize with the upper OT strand and lower OB strand of the two methylation characteristic sequences (fully methylated and fully unmethylated) in the target DNA region. Sequencing systems 2 and 3 were biased capture with different enrichment efficiencies, which used two-stranded long and short probes to enrich the target DNA. Based on the real experimental data, different fitting coefficients k were obtained. Figure 9 As shown, the enrichment coefficient of frequency curve 910 obtained using sequencing system 1 is 0.94. The enrichment coefficient of frequency curve 920 obtained using sequencing system 2 is 2.83. The enrichment coefficient of frequency curve 930 obtained using sequencing system 3 is 38.94.

[0094] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to the technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for determining the target frequency of a target signal molecule in a sample to be tested, comprising: Based on probe hybridization capture, at least one target site of the sample to be tested is targeted for enrichment; The enriched sample to be tested is sequenced using high-throughput sequencing to obtain at least one detection frequency of the target signal molecule at at least one target site; The median of the at least one detection frequency is determined as the detection frequency to be calibrated; as well as Based on the detection frequency to be calibrated and the enrichment coefficient associated with the target enrichment, the target frequency of the target signal molecule in the sample to be tested is determined.

2. The method according to claim 1, wherein the target signaling molecule comprises at least one of the following: methylated haplotype, mutant haplotype, or fusion gene haplotype.

3. The method of claim 2, wherein the target site of the methylated haplotype comprises one or more CpGs.

4. The method of claim 3, wherein the state of the CpG includes methylation and demethylation.

5. The method according to claim 2, wherein the mutated haplotype includes base mutation, insertion, and deletion.

6. The method according to claim 2, wherein the fusion gene haplotype is formed by a structural variation process in which all or part of the sequences of two genes are fused together to form a new gene, the structural variation process including the result of chromosomal translocation, intermediate deletion or chromosomal inversion.

7. The method according to claim 1, wherein the detection frequency in the at least one detection frequency has a range of 0 to 1.

8. The method of claim 1, wherein the target frequency of the signal molecule in the test sample, based on the detection frequency to be calibrated and the enrichment coefficient associated with the target enrichment, comprises the following as the proportion of the signal molecule in the test sample: Based on the enrichment coefficient, the correspondence between the detection frequency to be calibrated and the target frequency is determined; as well as The target frequency is determined based on the detection frequency to be calibrated and the corresponding relationship.

9. The method according to claim 8, further comprising: Using the targeted enrichment, multiple reference sites of multiple reference samples with different reference frequencies are enriched to obtain multiple enriched reference samples. The high-throughput sequencing was used to sequence the multiple enriched reference samples to obtain multiple sets of detection frequencies of the multiple enriched reference samples at multiple reference sites. Multiple medians of the multiple sets of detection frequencies are determined as multiple median detection frequencies; as well as The enrichment coefficient is determined by fitting based on multiple reference frequencies of the multiple reference samples and the multiple median detection frequencies.

10. The method of claim 8, wherein determining the correspondence between the detection frequency to be calibrated and the target frequency based on the enrichment coefficient comprises: The detection frequency to be calibrated is determined based on the enrichment coefficient and a first representation of the target frequency; as well as The first representation is transformed to obtain the correspondence between the target frequency and the enrichment coefficient and the detection frequency to be calibrated.

11. The method of claim 1, wherein enriching at least one target site of the test sample using targeted enrichment via probe hybridization capture comprises: Library preparation was performed on the sample to be tested to establish a pre-library; as well as The pretext is subjected to targeted enrichment based on probe hybridization capture to obtain the final text.

12. The method of claim 8, wherein sequencing the enriched sample to be tested using high-throughput sequencing to obtain at least one detection frequency of the target signal molecule at the at least one target site comprises: The high-throughput sequencing is performed on the final library to obtain sequencing data; as well as The sequencing data is analyzed to obtain the detection frequencies of the multiple targets.

13. The method according to claim 11, wherein the probes used in the hybridization capture process comprise probes that hybridize as a whole with all haplotype characteristic sequences in the target DNA region.

14. The method of claim 13, wherein the probe comprises an equal-length probe that simultaneously captures fully methylated and fully demethylated upper OT chains and lower OB chains, an equal-length probe that specifically captures fully methylated or fully demethylated OT chains and OB chains, and a short probe of unequal length that specifically captures methylated haplotypes.

15. The method of claim 13, wherein the probe has a length from 10 nt to 2000 nt.

16. The method of claim 13, wherein the probe comprises at least one of the following: a DNA probe, an RNA probe, a modified base LNA, a modified backbone PNA, or a probe-protein complex capture.

17. The method of claim 13, wherein the number of probes is from 1 to 10. 12 Within the range.

18. The method of claim 13, wherein the targeted enrichment is performed by an assay system comprising: one or more probes binding to the at least one target site, blocking primers, a hybridization reaction solution, and a hybridization temperature for hybridization, and wherein the assay system comprises: a capture reaction solution, an elution solution, and a capture temperature for capture.

19. The method according to claim 1, wherein the sample to be tested comprises: Deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

20. The method of claim 19, wherein the sample to be tested comprises circulating free DNA, cell line DNA, or tissue DNA.

21. The method according to claim 20, wherein the circulating cell-free DNA includes circulating cell-free DNA in plasma, circulating cell-free DNA in urine, circulating cell-free DNA in pleural effusion, circulating cell-free DNA in ascites, or circulating cell-free DNA in cerebrospinal fluid.

22. The method of claim 1, wherein the high-throughput sequencing comprises second-generation sequencing or third-generation sequencing.