Methods of constructing libraries for haplotype analysis and uses thereof
Through the combination of strand read sequencing technology and probe-targeted capture methods, a library for haplotype analysis is constructed, which solves multiple challenges in haplotype analysis in the prior art, and achieves efficient and accurate haplotype construction and variant detection, reducing costs.
Patent Information
- Application Number
- CN202510223416.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art faces problems such as missing family members, limited detection scope, difficulty in gene pathogenicity interpretation, chain imbalance caused by misdiagnosis, and prone to misdiagnosis of new variants in haplotype analysis. The third-generation sequencing technology has high error rate and high cost, which limits its application.
String read sequencing technology combined with probe targeted capture is used to construct a library for haplotype analysis. Independent haplotype construction and analysis are achieved through sample pre-processing, interruption reaction, molecular tag labeling, fragmented labeling enzyme release, sequencing linker ligation and target region probe hybridization capture steps.
This method does not require probands and can independently construct haplotypes, which improves the accuracy of haplotype construction and the accuracy of variant detection, reduces costs, and is suitable for large-scale applications and clinical diagnosis.
Smart Images

Figure CN120060438A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of genotype detection and analysis, and specifically relates to a method for constructing a library for haplotype analysis and its uses. Background Art
[0002] Preimplantation genetic testing (PGT) is an embryo detection technology that combines assisted reproduction and genetic diagnosis. It refers to the detection of chromosomes or specific genes in preimplantation embryos, enabling couples facing a high genetic risk of pregnancy to reduce the risk of natural miscarriage or avoid genetic birth defects after selecting genetically normal embryos for uterine implantation. PGT technology is generally divided into three categories: PGT-A (PGT for aneuploidies) for screening aneuploid embryos; PGT-M (PGT for monogenic) for detecting diseases caused by a single gene; and PGT-SR (PGT for structural rearrangements) for detecting chromosomal abnormalities caused by genomic structural rearrangements.
[0003] Currently reported detection methods for PGT-M and PGT-SR are mainly divided into direct detection and indirect detection. The direct detection method refers to directly detecting variations using a series of techniques, mainly including PCR, WGA technology, and fluorescence in situ hybridization (FISH), comparative genomic hybridization (CGH), or array comparative genomic hybridization (CGH-array), single nucleotide polymorphism array (SNP array), high-throughput sequencing technology (HTS), etc. based on WGA; while the most common indirect detection method is haplotype linkage analysis. Currently, direct detection methods including single-cell whole-genome sequencing technology cannot effectively solve the problem of allele dropout (ADO), and haplotype linkage analysis, as an indirect detection and analysis method that can avoid this problem, is widely used in clinical practice.
[0004] With the development of high-throughput sequencing technology, NGS-based haplotype analysis technology has been widely used in clinical practice as a solution for preimplantation genetic testing. However, haplotype analysis still faces challenges in special cases, such as when family members are missing or difficult to obtain, and haplotypes cannot be inferred. In addition, population inference methods are affected by linkage disequilibrium and may miss low-mutation-rate individuals, leading to missed diagnoses. For de novo variants, haplotype analysis is also prone to missed diagnoses. Third-generation long-read sequencing can make up for the deficiencies of haplotype analysis technology based on second-generation sequencing in constructing haplotypes. Third-generation sequencing technology can generate read lengths of hundreds of thousands or even millions of base pairs, which enables them to span large genomic regions, including repetitive sequences and structural variant regions that are difficult to resolve by short-read sequencing technology. Due to the long read lengths, third-generation sequencing technology can directly read multiple closely linked SNP sites, thereby determining the phase relationship between these sites and can independently complete haplotype construction without relying on proband information. However, although correction can be performed through algorithms and multiple sequencing, the error rate of current third-generation sequencing technology is usually high; at the same time, the analysis of long-read data requires more complex bioinformatics tools and computing resources, increasing the difficulty of data processing; and the cost of long-read sequencing technology is high, which limits its popularization in large-scale applications.
[0005] Linked-read sequencing is a new type of high-throughput sequencing technology that provides a solution between short-read and long-read by linking long DNA fragments with barcodes. This technology not only solves the defect of short read lengths in second-generation sequencing but also has lower costs and higher sequencing accuracy compared to third-generation sequencing. However, these technologies are currently only applicable to the detection of whole genomes or larger genomes and are not suitable for targeted detection of smaller genomes. Conventional targeted sequencing methods, including multiplex PCR and probe capture methods, cannot provide long-range haplotype information due to their own technical principles and sequencing platform limitations. The multiplex PCR method has preferences, with small fragments always being preferentially amplified; and the optimal PCR conditions for each pair of primers are different, and longer fragments require longer extension times; additional mutations are easily generated during the amplification process; multiplex PCR is more suitable for point mutations and small insertions and deletions.
[0006] Preimplantation genetic haplotype analysis (PGH) has also been used in the prior art for haplotype construction and analysis. However, this technology also has many limitations. Specifically, 1) Impact on family structure: PGH technology first requires high-throughput whole-genome SNP (single nucleotide polymorphism) genotyping of core family members (usually parents and at least one affected family member) to determine their haplotypes. For families with complex or incomplete family structures, such as those lacking a proband or key family members, it is difficult to perform accurate linkage analysis. 2) Limitations in detection scope: PGH technology has limitations in dealing with long-fragment variations and highly similar sequences, which may lead to inaccurate detection. In particular, for complex variations such as large fragment insertions, deletions, or duplications, short-read technologies are difficult to detect. 3) Limitations in gene pathogenicity interpretation: For some genes, there are cases where multiple variant sites jointly affect, and it is necessary to judge the cis / trans of the variant sites to further confirm pathogenicity. Conventional detection methods require genotype results of family members for variant source analysis, and when new cases occur, the above problems cannot be solved. 4) Limitations in linkage disequilibrium: Affected by linkage disequilibrium, low mutation rate individuals may be missed, resulting in missed diagnoses. 5) Requirements for data quality and quantity: PGH technology requires high-quality genetic marker data and a sufficient number of family samples. The data collection and analysis process may be very complex and costly. Moreover, if the data is biased or incomplete, it may affect the accuracy of the analysis. 6) Computational resources and time: Whole-genome SNP analysis and linkage analysis involve a large amount of data and require complex calculations and professional bioinformatics support. As the number of family samples increases, the required computational resources and time costs increase significantly, especially for the analysis of large-scale family data. 7) Cost issues: The cost of whole-genome high-throughput sequencing and SNP genotyping is relatively high, which limits the application of PGH technology in some regions.
[0007] Therefore, there is an urgent need in the art to develop a new haplotype construction method. Summary of the Invention
[0008] Based on this, it is necessary to provide at least one method for constructing a library for haplotype analysis and its uses.
[0009] In the first aspect of the present application, a method for constructing a library for haplotype analysis is provided, which includes the following steps:
[0010] Sample pretreatment: Lyse the sample to obtain the released sample DNA;
[0011] Fragmentation reaction: Mix the sample DNA with a fragmenting enzyme to obtain the sample DNA combined with the fragmenting enzyme;
[0012] Molecular tag labeling: Mix the sample DNA bound to the fragmentation enzyme with magnetic beads with different molecular tags. The sample DNA and the molecular tags hybridize and ligate through a bridging oligo to obtain the labeled sample DNA.
[0013] Fragmentation enzyme release: Mix the labeled sample DNA with a denaturing buffer to inactivate the fragmentation enzyme and obtain fragmented DNA molecules. Fragments from the same DNA carry the same molecular tag.
[0014] Sequencing adapter ligation: Add a sequencing adapter to the 3'-end of the fragmented DNA molecules to obtain a pre-capture library.
[0015] Target region probe hybridization capture: Add labeled probes to the pre-capture library, hybridize and incubate to obtain a labeled pre-capture library; mix with magnetic beads that specifically bind to the label on the probe to capture the target fragments, and optionally perform amplification enrichment to obtain a post-capture library.
[0016] In some embodiments, in the step of sequencing adapter ligation, a sequencing adapter is added to the 3'-end of the DNA molecules by one or more of DNA ligase and transposase.
[0017] In some embodiments, it further includes a step of amplifying and enriching the pre-capture library to meet the input amount for target region probe hybridization capture.
[0018] In some embodiments, amplification enrichment is achieved by PCR. In some embodiments, the reaction conditions for the PCR are: (i) preheat at 98°C for 3 min; (ii) denature at 95°C for 30 s, anneal at 55°C - 65°C for 30 s, extend at 72°C for 2 min, cycle 12 - 15 times; and, (iii) 72°C for 10 min.
[0019] In some embodiments, the target region probe hybridization capture step meets one or more of the conditions shown in 1) - 5) below:
[0020] 1) The label is biotin, and the magnetic beads that specifically bind to the label on the probe are streptavidin magnetic beads;
[0021] 2) Amplification enrichment is achieved by PCR; and,
[0022] 3) Hybridization heat elution is performed 2 - 4 times.
[0023] In some embodiments, the reaction conditions for the PCR in condition 2) above include: (i) preheat at 95°C for 1 min; (ii) denature at 98°C for 20 s, anneal at 60°C for 30 s, extend at 72°C for 30 s, cycle 8 - 13 times; and, (iii) 72°C for 5 min.
[0024] In some embodiments, the sample is a liquid sample.
[0025] In some embodiments, the liquid sample is selected from the group consisting of blood, semen, tissue, and cultured cells.
[0026] A second aspect of the present application provides a method for constructing a haplotype of a pathogenic variant carrier, which includes the following steps:
[0027] Construct a library using the method described in the first aspect, and the sample used for constructing the library is from a pathogenic variant carrier;
[0028] Sequence the constructed library to obtain long-read sequencing data.
[0029] In some embodiments, the method further includes:
[0030] Align the long-read sequencing data of the pathogenic variant carrier's parent to the reference genome and perform variant analysis in the target region to obtain the genotype information of its pathogenic variant;
[0031] Screen for heterozygous SNPs in the pathogenic variant carrier's parent from the genotype information;
[0032] Assemble the long-read sequencing data of the pathogenic variant carrier's parent to obtain its genomic contigs and label them with numbers. Compare the genomic contigs with the reference genome to obtain alignment information and label the coverage interval of the genomic contigs on the reference genome;
[0033] Construct a haplotype typing of SNPs in the target region;
[0034] Match the assembled contig information where the pathogenic variant is located with the result of the haplotype typing to obtain the haplotype typing result of the pathogenic locus; label the haplotypes of the pathogenic variant carrier's parent according to the typing result of the pathogenic locus, and construct the pathogenic / normal haplotypes of the pathogenic variant carrier's parent.
[0035] In some embodiments, the design principle of the probe includes: according to the position information of the pathogenic variant, select high-frequency mutant SNP sites within 1 Mb upstream and downstream of the pathogenic variant, and set the interval range between adjacent SNP sites to be 1 kb to 5 kb.
[0036] In some embodiments, the reference genome includes the human reference genome GRCh37 or GRCh38, etc.
[0037] In some embodiments, the method does not rely on the proband or reference sample.
[0038] In some embodiments, the reference sample includes one or more of the affected embryo and the direct relatives of both spouses.
[0039] The advantages of the haplotype construction method of the present application at least include:
[0040] 1) Without a proband, no pedigree-linked haplotypes are required: The strand sequencing technology has the ability to construct independent haplotypes and can generate long DNA sequence reads, which helps to span complex regions in the genome, such as repetitive sequences and structural variations, thereby more accurately constructing haplotypes and assisting patients with de novo mutations, no pedigree information, and complex genomic structures, etc., who cannot identify the source of variations through traditional pedigree linkage analysis, to independently construct haplotypes and judge pathogenicity;
[0041] 2) Precise haplotype construction: By combining the high specificity of the probe for targeted capture with the long read length advantage of strand sequencing, local haplotypes can be constructed more accurately, especially in complex gene regions or highly similar genomic regions;
[0042] 3) Improved accuracy of pathogenicity rating: By analyzing the haplotypes between different variations, the cis-trans situation of complex structural variations can be distinguished at the same time, and the pathogenicity rating of the variations can be determined without the assistance of pedigree information, greatly reducing the detection cost and analysis cost;
[0043] 4) Improve the accuracy of variant detection and solve complex genetic structures: The strand sequencing technology helps to analyze complex structural variations in the genome, such as long fragment insertions and deletions, etc. Combining probe targeted capture can further precisely determine the position and type of these variations, improving the accuracy and reliability of variant detection in these regions;
[0044] 5) High throughput and high efficiency: The combination of strand sequencing technology and probe targeted capture can achieve high-throughput sequencing, detect a large number of samples at the same time, quickly obtain a large amount of genetic information, and greatly improve the efficiency of research;
[0045] 6) Easy to automate and scale up for application: The technical combination process is simple, the operation is convenient, it is easy to achieve automation, and it is suitable for large-scale application, especially in clinical diagnosis and large-scale population genetic research;
[0046] 7) Cost-effectiveness: Compared with traditional whole-genome sequencing, this technical combination can specifically capture the research region. On the basis of detecting common variations of the target gene, typing SNPs within 1 Mb upstream and downstream of the target gene are designed, and the interval between the captured SNPs is 1-5 kb, greatly reducing the sequencing cost, while meeting the requirements of variant detection and long fragment haplotype assembly, and maintaining high data quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] To more clearly illustrate the technical solutions in the embodiments and examples of the present application and to more fully understand the present application and its beneficial effects, the following will briefly introduce the drawings required for the description of the embodiments or examples. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings. It should also be noted that the drawings are all drawn in a simplified form and are only used to conveniently and clearly assist in explaining the present application.
[0048] Figure 1 Showing the thalassemia α-variant and β-variant typing diagrams of sample B1F0 (CDs41-42(-CTTT) / N) obtained by using the method of the present invention in an embodiment of the present application;
[0049] Figure 2 Showing the thalassemia α-variant and β-variant typing diagrams of sample B1M0 (-SEA / αα, CDs41-42(-CTTT) / N) obtained by using the method of the present invention in an embodiment of the present application;
[0050] Figure 3 Showing the thalassemia α-variant and β-variant typing diagrams of sample B2F0 (IVS-II-654(C-T)) obtained by using the method of the present invention in an embodiment of the present application;
[0051] Figure 4 Showing the thalassemia α-variant and β-variant typing diagrams of sample B2M0 (IVS-II-654(C-T)) obtained by using the method of the present invention in an embodiment of the present application;
[0052] Figure 5 Showing the thalassemia α-variant and β-variant typing diagrams of sample B4F0 (-α3.7 / αα) obtained by using the method of the present invention in an embodiment of the present application;
[0053] Figure 6 Showing the thalassemia α-variant and β-variant typing diagrams of sample B4M0 (-SEA / αα) obtained by using the method of the present invention in an embodiment of the present application;
[0054] Figure 7 Showing the thalassemia α-variant and β-variant typing diagrams of sample B5F0 (-SEA / αα) obtained by using the method of the present invention in an embodiment of the present application;
[0055] Figure 8 Showing the thalassemia α-variant and β-variant typing diagrams of sample B5M0 (-SEA / HKαα) obtained by using the method of the present invention in an embodiment of the present application;
[0056] Figure 9Show the thalassemia α-variation and β-variation typing diagrams of sample B3F0(-SEA / ααWS) obtained by using the method of the present invention and the WGS method in an embodiment of the present application;
[0057] Figure 10 Show the thalassemia α-variation and β-variation typing diagrams of sample B3M0(-SEA / αα) obtained by using the method of the present invention and the WGS method in an embodiment of the present application;
[0058] Figure 11 Show the family typing diagrams of the upstream and downstream of the thalassemia variation sites of sample B3 family obtained by using the method of the present invention and the WGS method in an embodiment of the present application;
[0059] Figures 1 - 10 In each of the pictures above and below, there are different haplotypes. The colored connecting lines are the coverage areas of long fragment DNA molecules, and the red vertical lines are the positions of the target variations. In addition, Figures 9 - 11 "The method" in means the method in an embodiment of the present application. Detailed Embodiments
[0060] To facilitate the understanding of the present application, the present application will be described more comprehensively below with reference to the relevant drawings. The preferred embodiments of the present application are given in the drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the understanding of the disclosure of the present application more thorough and comprehensive.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used in the specification of the present application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0062] In the present application, unless otherwise specified, "one or more" means any one of the listed items or any combination of the listed items. Similarly, in other cases where "one or more" and the like are used to represent "one or more", the same understanding shall be made unless otherwise specified.
[0063] "Their combinations", "any of their combinations", "any combination modes thereof", etc. used in the present application include all suitable combination modes of any two or more of the listed items.
[0064] In the present application, "suitable combination modes", "suitable modes", "any suitable modes", etc., the "suitable" shall be subject to being able to implement the technical solution of the present application, solve the technical problems of the present application, and achieve the expected technical effects of the present application.
[0065] In this application, terms such as "further", "furthermore", "particularly", "for example", "such as", "example", "exemplification", etc. are used for descriptive purposes, indicating an association in terms of the covered content between the different technical solutions before and after. However, it should not be construed as a limitation on the previous technical solution, nor as a limitation on the protection scope of this application. In this application, unless otherwise specified, A (such as B) means that B is a non-restrictive example of A, and it can be understood that A is not limited to B.
[0066] In this application, "optionally", "optional", "option" mean that it can be either present or absent, that is, it refers to either of the two parallel options of "present" or "absent". If the term "optional" appears multiple times in a technical solution, unless otherwise specified and there are no contradictions or mutual constraints, each "optional" is independent. Unless otherwise specified, descriptions such as "optionally include" and "optionally contain" in this application, taking "optionally include" as an example, mean "may include or may not include".
[0067] The terms "contain", "include" and "comprise" used in this application are synonyms, and they are inclusive or open-ended, not excluding additional, unmentioned members or features. Members or features include, for example, materials or components, structures, elements, instruments, etc.; non-restrictive examples of members or features also include actions, conditions under which actions occur, timing, states, etc.
[0068] In this application, in a technical feature or technical solution described in an open language, it includes a closed technical feature or technical solution composed of the listed content, and also includes an open technical feature or technical solution containing the listed content.
[0069] In this application, exemplary descriptions such as "in some embodiments" or "in one embodiment" can cover, but are not limited to, the following meanings: These solutions can be combined with other solutions in a suitable manner to form new technical solutions.
[0070] In this application, in "the first aspect", "the second aspect", "the third aspect", "the fourth aspect", etc., the terms "first", "second", "third", "fourth", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or quantity, nor can it be understood as implicitly indicating the importance or quantity of the indicated technical features. Moreover, "first", "second", "third", "fourth", etc. only serve the purpose of non-exhaustive enumerative description, and it should be understood that they do not constitute a closed limitation on the quantity.
[0071] In this application, when it comes to numerical intervals (i.e., numerical ranges), unless otherwise specified, the distribution of the selectable numerical values within the numerical interval is considered continuous, and includes the two numerical endpoints of the numerical interval (i.e., the minimum value and the maximum value), as well as each numerical value between these two numerical endpoints. Unless otherwise specified, when the numerical interval only refers to the integers within the numerical interval, it includes the two endpoint integers of the numerical range, as well as each integer between the two endpoints, which is equivalent to directly listing each integer. When multiple numerical ranges are provided to describe features or characteristics, these numerical ranges can be combined. In other words, unless otherwise specified, the numerical ranges disclosed herein should be understood to include any and all sub-ranges subsumed therein. The "numerical value" in the numerical interval can be any quantitative value, such as a number, a percentage, a ratio, etc. The "numerical interval" is allowed to broadly include numerical interval types such as percentage intervals, ratio intervals, and ratio value intervals.
[0072] In this application, when there are multiple steps involved in a method process, unless there are clear different descriptions in this article, the execution of these steps has no strict order limit, and they can be executed in an order other than the described one. Moreover, any one step can include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments, and their execution order does not necessarily need to be sequential, but can be executed alternately or simultaneously with other steps or a part of the sub-steps or stages of other steps.
[0073] In this application, the inventors combined the chain reading sequencing technology with probe-based target capture and applied it to haplotype construction, showing significant advantages and application potential. Its main advantages include:
[0074] 1. Independent haplotype construction ability: It can complete haplotype construction independently without relying on proband or family linkage information, and is applicable to de novo mutations and cases without family information.
[0075] 2. Precise haplotype construction: Utilize the advantages of long read lengths and the high specificity of probes to improve the construction accuracy of haplotypes in complex gene regions.
[0076] 3. Improve the accuracy of variant detection: Use target capture to increase the sequencing depth, effectively resolve complex structural variants in the genome, and improve the accuracy and reliability of variant detection.
[0077] 4. Improve the accuracy of pathogenicity rating: The present invention can distinguish the cis-trans situations of complex structural variants, provide a basis for pathogenicity assessment in clinical practice, and thus improve the accuracy of pathogenicity rating.
[0078] 5. Cost-effectiveness: Targetedly capture the research area, reduce costs, and at the same time maintain high data quality.
[0079] 6. High throughput and high efficiency: Enabling high-throughput sequencing to rapidly obtain a large amount of genetic information and improve research efficiency.
[0080] 7. Easy to automate and scale: With a simple process and convenient operation, it is suitable for automated and large-scale applications, especially for clinical diagnosis and large-scale population genetic research.
[0081] In summary, the strand-reading sequencing technology combined with probe-targeted capture provides an efficient, accurate, and cost-effective technical solution for haplotype construction and detecting disease-causing variants.
[0082] This application provides a method for constructing a library for haplotype analysis, which includes the following steps:
[0083] Sample pretreatment: Lysing the sample to obtain the released sample DNA;
[0084] Fragmentation reaction: Mixing the sample DNA with a fragmentation enzyme to obtain the sample DNA bound with the fragmentation enzyme;
[0085] Molecular tag labeling: Mixing the sample DNA bound with the fragmentation enzyme with magnetic beads with different molecular tags, and hybridizing and ligating the sample DNA with the molecular tags through a bridging oligo to obtain the labeled sample DNA;
[0086] Fragmentation and release of the labeling enzyme: Mixing the labeled sample DNA with a denaturing buffer to inactivate the fragmentation enzyme and obtain fragmented DNA molecules, and the fragments from the same DNA carry the same molecular tag;
[0087] Sequencing adapter ligation: Adding a sequencing adapter to the 3' end of the fragmented DNA molecules to obtain a pre-capture library;
[0088] Probe hybridization capture of the target region: Adding labeled probes to the pre-capture library, hybridizing and incubating to obtain a labeled pre-capture library; mixing with magnetic beads specifically binding to the labels on the probes to capture the target fragments, and optionally performing amplification and enrichment to obtain a post-capture library.
[0089] In this application, the sample pretreatment may include rinsing, lysis reaction, etc. Exemplarily, rinsing can be performed by centrifugation. In some embodiments, the sample (such as peripheral blood) is centrifuged at a speed of, for example, 11000 rpm for several minutes, and the supernatant is discarded; after adding PBS for rinsing, it is centrifuged again.
[0090] In the cleavage reaction, the cleavage reaction solution used includes T4 DNA ligation buffer, thermosensitive proteinase K and nuclease-free water. For example, the volume of a standard reaction is 10 μL, which includes 1 μL T4 DNA ligation buffer, 0.4 μL thermosensitive proteinase K (120 units / mL) and 8.6 μL nuclease-free water.
[0091] In some embodiments, the conditions for the cleavage reaction include: 75°C heated lid, 37°C for 15 min, 65°C for 10 min, and hold at 25°C.
[0092] In the present application, the fragmentation enzyme used in the interruption reaction is contained in the interruption reaction solution. In some embodiments, the interruption reaction solution comprises components such as cleavage product, fragmentation buffer, fragmentation marker enzyme working solution (for example, 16-fold dilution) and nuclease-free water. Exemplarily, the volume of a reaction standard is 50 μL, which includes 10 μL cleavage product, 10 μL fragmentation buffer, 19.2 μL fragmentation marker enzyme working solution (16-fold dilution) and 10.8 μL nuclease-free water.
[0093] In some embodiments, the conditions for interrupting the reaction include: 60°C heated lid, 55°C for 10 min, and holding at 4°C.
[0094] In some embodiments, the molecular tagging step includes a hybridization reaction and a ligation reaction. The conditions of the hybridization reaction may include a 65°C hot cover, 60°C for 10 min, 45°C for 50 min, and optionally maintained at 4°C. The conditions of the ligation reaction may include 25°C for 60 min.
[0095] In some embodiments, the molecular tagging further comprises a digestion reaction.
[0096] In some embodiments, the reaction conditions for releasing the fragmentation marker enzyme include: maintaining at 20° C. to 25° C. for 10 min.
[0097] In some embodiments, the fragmentation enzyme is a transposase with a barcode tag to facilitate sample differentiation when mixing samples on a machine.
[0098] In some embodiments, for low-throughput sequencers, when mixing samples is not required, the fragmentation enzyme used may not have a barcode label.
[0099] In some embodiments, in the step of ligating the sequencing adapter, the sequencing adapter is added to the 3' end of the DNA molecule by one or more of DNA ligase and transposase.
[0100] In some embodiments, the method further comprises the step of amplifying and enriching the pre-capture library to meet the input amount of hybridization capture of the target region probe.
[0101] In some embodiments, amplification enrichment is achieved by PCR.
[0102] In some embodiments, the reaction conditions for the PCR are as follows: (i) preheating at 98°C for 3 min; (ii) denaturing at 95°C for 30 s, annealing at 55°C - 65°C for 30 s, extending at 72°C for 2 min, and cycling 12 - 15 times; and (iii) extending at 72°C for 10 min.
[0103] In some embodiments, the annealing temperature is 55°C, 56°C, 57°C, 58°C, 59°C, 60°C, 61°C, 62°C, 63°C, 64°C, or 65°C, or a range or value between any two of these values. In some embodiments, the annealing temperature is 58°C. Without wishing to be bound by any theory, it has been found that the amplification sensitivity, specificity, and amplification efficiency are optimal at an annealing temperature of 58°C.
[0104] In some embodiments, the target region probe hybridization capture step satisfies one or more of the conditions shown in 1) - 2) below:
[0105] 1) The label is biotin, and the magnetic bead that specifically binds to the label on the probe is a streptavidin magnetic bead;
[0106] 2) Amplification enrichment is achieved by PCR; optionally, the reaction conditions for the PCR include: (i) preheating at 95°C for 1 min; (ii) denaturing at 98°C for 20 s, annealing at 60°C for 30 s, extending at 72°C for 30 s, and cycling 8 - 13 times; and (iii) extending at 72°C for 5 min.
[0107] In some embodiments, hybridization thermal elution is also included. Exemplarily, the number of times of hybridization thermal elution is 2 - 4 times. Without wishing to be bound by any theory, it has been found that the elution efficiency and specificity are better when the hybridization thermal elution is performed 4 times.
[0108] Unless otherwise specified, the term "streptavidin magnetic bead" in the present application refers to a magnetic bead modified with streptavidin.
[0109] In some embodiments, the temperature of hybridization thermal elution is 50°C - 60°C. Exemplarily, for example, 50°C, 51°C, 52°C, 53°C, 54°C, 55°C, 56°C, 57°C, 58°C, 59°C, 60°C, or a range or value between any two of these values.
[0110] In some embodiments, the number of cycles is 8, 9, 10, 11, 12, or 13. Without wishing to be bound by any theory, it has been found that the effect is better when the number of cycles is 12, with the largest haplotype Block upstream and downstream of the pathogenic mutation, the lowest standard deviation SD value, and the best uniformity.
[0111] In some embodiments, the sample is selected from the group consisting of blood, semen, tissue, and cultured cells.
[0112] The present application also provides a method for constructing a haplotype of a pathogenic variant carrier, which includes the following steps:
[0113] Construct a library using the method described above, and the sample used for constructing the library is from a pathogenic variant carrier;
[0114] Sequence the constructed library to obtain long-read sequencing data.
[0115] In some embodiments, the method further includes:
[0116] Align the long-read sequencing data of the pathogenic variant carrier's parents to the reference genome and perform variant analysis in the target region to obtain the genotype information of the pathogenic variant;
[0117] Screen for heterozygous SNPs in the pathogenic variant carrier's parents from the genotype information;
[0118] Assemble the long-read sequencing data of the pathogenic variant carrier's parents to obtain their genomic contigs and label them. Compare the genomic contigs with the reference genome to obtain alignment information and label the coverage interval of the genomic contigs on the reference genome;
[0119] Construct a haplotype typing of SNPs in the target region;
[0120] Match the information of the assembled contigs where the pathogenic variant is located with the results of the haplotype typing to obtain the haplotype typing results of the pathogenic variant locus; Mark the haplotypes of the pathogenic variant carrier's parents according to the typing results of the pathogenic locus, and construct the pathogenic / normal haplotypes of the pathogenic variant carrier's parents.
[0121] In some embodiments, the design principles of the probes include: According to the position information of the pathogenic variant, select high-frequency mutant SNP sites within 1 Mb upstream and downstream of the pathogenic variant, covering the entire length of the pathogenic gene, and set the interval range between adjacent SNP sites to be 1 kb to 5 kb.
[0122] Unless otherwise specified, "1 Mb" in the present application refers to 1000 kb, that is, 1000000 bp; 1 kb refers to 1000 bp.
[0123] In some embodiments, the reference genome includes the human reference genome GRCh37 or GRCh38, etc.
[0124] In some embodiments, the method does not rely on a proband or a reference sample; the reference sample includes one or more of affected embryos and immediate relatives of both spouses.
[0125] Some examples are provided below.
[0126] The embodiments of the present application will be described in detail below in conjunction with the examples. It should be understood that these examples are only used to illustrate the present application and not to limit the scope of the present application. For the experimental methods without specified conditions in the following examples, the guidance given in the present application is preferably referred to, and it can also be carried out according to the experimental manuals or conventional conditions in the art, or according to the conditions recommended by the manufacturer, or by referring to the experimental methods known in the art.
[0127] In the following examples, the thalassemia gene is taken as an example, RNA probes are designed within 1M upstream and downstream of the thalassemia gene, covering the full lengths of the HBA1, HBA2, and HBB genes, and high-frequency mutant SNP sites within 1Mb upstream and downstream of the three genes are selected. At the same time, the interval range between adjacent SNP sites is set between 1 - 5kb.
[0128] Unless otherwise specified, the term "high-frequency mutation" in the present application refers to a mutation frequency significantly higher than the background mutation rate (BMR). In some embodiments, the high-frequency mutant SNP site is a SNP site with a mutation frequency of 0.4 - 0.6.
[0129] The reagents used in the following examples:
[0130] The sample pretreatment kit is the "Cell Lysis Solution Kit" independently developed by Suzhou Beikang Medical Devices Co., Ltd.;
[0131] The library construction kit before capture is the "Whole Genome Haplotype Typing Detection Kit" independently developed by the company;
[0132] The target region hybridization capture kit is from AgenaBio Eco Universal Blocking Oligo (for MGl Dl);
[0133] TargetSeg Hyb&Wash Kit v2.0 (for MGl Dl);
[0134] Cap Beads&Nuclease-Free Water;
[0135] The probes are customized by AgenaBio.
[0136] HBA1 / HBA2 and HBB haplotypes of thalassemia pathogenic genes obtained by using the method of the present invention in Example 1
[0137] I. Library construction before capture
[0138] 1. Sample pretreatment
[0139] 1.1 Rinsing
[0140] 1) Centrifuge 30 μL of peripheral blood at 11000 rpm for 2 min, and discard the supernatant;
[0141] 2) Add 200 μL of PBS for rinsing, centrifuge at 11000 rpm for 2 min, and discard the supernatant;
[0142] 3) Repeat step 2 once;
[0143] 1.2 Lysis reaction
[0144] After discarding the supernatant, add the lysis reaction system prepared in advance according to Table 1 to the tube:
[0145] Table 1 Preparation of lysis reaction solution
[0146] Component Standard amount for one reaction (μL) T4 DNA Ligation Buffer 1 Thermolabile Proteinase K 0.4 Nuclease - free Water 8.6 Total Volume 10
[0147] Note: The concentration of thermosensitive proteinase K is 120 units / ml (units are units).
[0148] Flick and vortex to mix well and disperse the cell pellet.
[0149] 1.3 Place the above reaction tube on a gene amplifier and perform a lysis reaction according to the reaction conditions in Table 2.
[0150] Table 2 Lysis reaction conditions
[0151] Temperature Time Hot lid at 75°C Open 37℃ 15 min 65℃ 10 min 25℃ Take out the reactants for the next step
[0152] 2. Fragmentation reaction
[0153] 2.1 Dilute the fragmentation labeling enzyme 16-fold in two steps
[0154] 1) First, take 6 μL of TE buffer into a new 0.2 mL PCR tube, add 2 μL of fragmentation labeling enzyme, and intermittently vortex at medium speed 4 times, 2 s each time, and record it as the first dilution of fragmentation labeling enzyme (4-fold dilution).
[0155] 2) Then take another new 0.2 mL PCR tube, add 18 μL of TE buffer, and add 6 μL of the well-mixed first dilution of fragmentation labeling enzyme, and intermittently vortex at medium speed 4 times, 2 s each time, and label it as the working solution of fragmentation labeling enzyme (16-fold dilution).
[0156] Among them, the preparation operation of the fragmentation reaction needs to be carried out on ice throughout the process; the dilution solution of the fragmentation labeling enzyme needs to be prepared and used immediately.
[0157] 2.2 Pipette 10 μL of the lysis product into the fragmentation reaction system prepared in advance according to Table 3:
[0158] Table 3 Preparation of the fragmentation reaction solution
[0159]
[0160]
[0161] 2.3 Place the above reaction tube on a gene amplifier and carry out the fragmentation reaction according to the reaction conditions in Table 4.
[0162] Table 4 Fragmentation reaction conditions
[0163] Temperature Time Hot lid at 60°C Open 55℃ 10 min 4℃ Take out for the next step
[0164] After the reaction is completed, centrifuge briefly and place it on an ice box for standby. Take 35 μL of TE buffer into a new 0.2 mL PCR tube, slowly pipette 15 μL of the fragmented product into the 35 μL of TE buffer with a narrow-mouth pipette tip, invert and mix 10 times, and centrifuge at low speed and place it on an ice box for standby.
[0165] 3. Hybridization reaction
[0166] 3.1 Labeled magnetic bead washing
[0167] 1) Vigorously shake and mix the labeled magnetic beads evenly. Take 30 μL of the labeled magnetic beads for each sample into a 0.2 mL PCR tube. If multiple samples (n) are carried out simultaneously, the labeled magnetic beads required for the samples (n × 1.1 × 30 μL) can be taken into the same 0.2 mL PCR tube or 1.5 mL EP tube. It is recommended to use a 0.2 mL PCR tube when n ≤ 4.
[0168] 2) Centrifuge the PCR tube or EP tube and place it on a magnetic stand for 2 min until the liquid is clear, and carefully pipette and discard the supernatant.
[0169] 3) Keep the magnetic beads adsorbed on the magnetic stand, and add 50 μL of washing buffer I to the labeled magnetic beads for each sample for washing. Rotate the tube at 180 degrees to let the magnetic beads swim in the washing buffer I to wash the magnetic beads, and rotate it back and forth 2 times to fully wash the magnetic beads. The volume of washing buffer I can be appropriately increased according to the amount of magnetic beads, and the washing buffer I should at least completely cover all the magnetic beads.
[0170] 4) After the supernatant becomes clear (about 1 min), carefully aspirate and discard the supernatant, and remove the tube from the magnetic stand. Resuspend the labeled beads of each sample by adding 50 μL of resuspension buffer. If the tube contains labeled beads of multiple (n) samples, add the corresponding volume of resuspension buffer (n × 1.1 × 50 μL) to resuspend the beads.
[0171] 3.2 Hybridization reaction
[0172] 1) Vortex the labeled beads to prevent precipitation. Aspirate 50 μL of the labeled beads and add them to the sample in step 2.3. Mix gently by inverting or flicking at least 10 times, and do not vortex vigorously. After mixing, place the reaction tube in a palm centrifuge, gently press the centrifuge button, and centrifuge at low speed multiple times to ensure that there is no liquid remaining on the tube cap and that the beads are fully suspended.
[0173] 2) After centrifuging the above hybridization product (100 μL in total) at low speed instantaneously, quickly place it on a gene amplifier and react according to the reaction conditions in Table 5.
[0174] Table 5 Hybridization reaction conditions
[0175] Temperature Time Hot lid (65°C) Open 60℃ 10 min 45℃ 50 min 4℃ Take out for the next step
[0176] 4. Ligation reaction 1
[0177] 1) Prepare the reaction solution for ligation reaction 1 on an ice box according to the formulation in Table 6. Vortex to mix and place it on the ice box for standby.
[0178] Table 6 Preparation of the reaction solution for ligation reaction 1
[0179] Component Standard amount for one reaction (μL) Ligation Buffer I 26 Ligase I 4 Total Volume 30
[0180] 2) After completing step 3 of hybridization, take out the hybridization product from the gene amplifier, centrifuge it instantaneously and place it at room temperature. Wait for the product to cool to room temperature, then add 30 μL of the ligation reaction 1 reaction solution to the hybridized product that has cooled to room temperature (final volume 130 μL). Mix by flicking or gently inverting, and centrifuge at low speed instantaneously for 1 s to ensure that there is no liquid remaining on the tube cap and that the product is fully suspended.
[0181] 3) Place the above reaction tube on a gene amplifier and perform ligation reaction 1 according to the reaction conditions in Table 7.
[0182] Table 7 Ligation reaction 1 conditions
[0183] Temperature Time Lid Temperature Close 25℃ 60 min 4℃ Take out for the next step
[0184] 4) After the reaction is completed, flick the product to mix, centrifuge it at low speed instantaneously, and then place it on the magnetic stand for 1 min - 2 min until the liquid becomes clear. Carefully aspirate and discard the supernatant.
[0185] 5) Keep the PCR tube on the magnetic stand, add 180 μL of Wash Buffer II into the tube, rotate the PCR tube at 180 degrees to let the labeled beads swim in the Wash Buffer II for bead washing, rotate it back and forth twice, wash the beads thoroughly. After the supernatant becomes clear, carefully aspirate and discard all the supernatant, ensuring no residue of Wash Buffer II remains in the tube.
[0186] 5. Digestion Reaction
[0187] 1) Prepare the reaction solution for digestion reaction on an ice box according to the recipe in Table 8, mix it by oscillation and keep it on the ice box for standby.
[0188] Table 8 Preparation of Reaction Solution for Digestion Reaction
[0189] Component Standard amount for one reaction (μL) Digestion Buffer 95 Digestive Enzyme 5 Total Volume 100
[0190] 2) Add the reaction solution obtained after digestion reaction into the PCR tube in step 5) of ligation reaction on the ice box, flick or gently invert it about 10 times to resuspend and mix evenly, and centrifuge it instantaneously at low speed for 1 s.
[0191] 3) Place the above PCR tube on the gene amplifier and carry out the digestion reaction under the conditions shown in Table 9 below.
[0192] Table 9 Reaction Conditions for Digestion Reaction
[0193] Temperature Time Lid temperature (42°C) Open 37℃ 10 min
[0194] 6. Release of Fragmentation Labeling Enzyme
[0195] 1) After the digestion reaction is completed, take out the reaction tube from the gene amplifier, centrifuge it instantaneously and place it at room temperature. Immediately add 11 μL of denaturation buffer into the digestion reaction product.
[0196] 2) Ensure that the lid of the PCR tube is tightly closed, mix the mixture added with denaturation buffer by medium-speed oscillation for 3 - 5 s, centrifuge it instantaneously and place it at room temperature (20 °C - 25 °C), and carry out the release of fragmentation labeling enzyme according to the reaction conditions shown in Table 10.
[0197] Table 10 Reaction Conditions for Release of Fragmentation Labeling Enzyme
[0198] Temperature Time 20℃-25℃ 10 min
[0199] 3) After the reaction is completed, centrifuge the product instantaneously, place it on the magnetic stand for 2 min until the liquid becomes clear, and carefully aspirate and discard the supernatant.
[0200] 4) Keep the reaction tube on the magnetic stand, add 150 μL of Wash Buffer II to the tube for washing. When washing, note that at least all the labeled magnetic beads should be covered. Vigorously shake and mix well for about 10 s, centrifuge instantaneously for 3 s, then place it on the magnetic stand for 2 min until the liquid becomes clear, and carefully aspirate and discard the supernatant.
[0201] 5) Repeat step 4 twice. After washing is completed, ensure that there is no residue of Wash Buffer II.
[0202] 7. Pretreatment of Ligation Reaction 2
[0203] 1) Prepare the reaction solution for the pretreatment of Ligation Reaction 2 on an ice box according to the recipe in Table 11. Shake and mix well and place it on the ice box for standby.
[0204] Table 11 Preparation of the Reaction Solution for the Pretreatment of Ligation Reaction 2
[0205] Component Standard amount for one reaction (μL) Ligation Buffer II 20 Ligase II 4 Total Volume 24
[0206] 2) Add the reaction solution for the pretreatment of Ligation Reaction 2 to the PCR tube in step 5) of the Fragmentation Labeling Enzyme Release. Shake and fully resuspend the labeled magnetic beads. After centrifuging instantaneously for 1 s, quickly place it in the gene amplifier and react according to Table 12.
[0207] Table 12 Reaction Conditions for the Pretreatment of Ligation Reaction 2
[0208] Temperature Time Lid temperature (42°C) Open 37℃ 30 min 4℃ Take out for the next step
[0209] 3) After the reaction is completed, immediately take out the product from the gene amplifier, centrifuge instantaneously, and place it at room temperature for the next step.
[0210] 8. Ligation Reaction 2
[0211] 1) Prepare the reaction solution for Ligation Reaction 2 on an ice box according to the recipe in Table 13. Note that the reaction solution for Ligation Reaction 2 needs to be shaken and mixed well and placed on the ice box for standby.
[0212] Table 13 Preparation of the Reaction Solution for Ligation Reaction 2
[0213] Component Standard amount for one reaction (μL) Ligation Buffer III 48 Adapter 18 Ligase I 10 Total Volume 76
[0214] 2) Add all 76 μL of the reaction solution for Ligation Reaction 2 to the sample tube that has cooled to room temperature (final volume 100 μL). Shake and fully resuspend the labeled magnetic beads. Centrifuge instantaneously for 1 s to ensure that there is no liquid remaining on the tube cap and ensure that the magnetic beads are fully suspended.
[0215] 3) Place the reaction tube in the gene amplifier and carry out Ligation Reaction 2 according to the reaction conditions shown in Table 14.
[0216] Table 14 Reaction Conditions for Ligation Reaction 2
[0217] Temperature Time Lid Temperature Close 25℃ 120 min 4℃ Take out for the next step
[0218] 4) After the reaction is completed, take out the reaction tube, centrifuge the product at low speed instantaneously, add 80 μL of washing buffer II into the tube, place it on the magnetic stand for 2 min until the liquid becomes clear, and carefully aspirate and discard the supernatant.
[0219] 5) Keep the PCR tube on the magnetic stand, add 180 μL of washing buffer II into the tube for washing, rotate the PCR tube at 180 degrees to let the labeled magnetic beads swim in the washing buffer II to wash the magnetic beads, and repeat rotating the PCR tube 2 times. After the supernatant becomes clear, carefully aspirate and discard the supernatant to ensure that there is no residue of washing buffer II.
[0220] 9. PCR Reaction
[0221] 9.1 Prepare the PCR reaction solution on an ice box according to the formula shown in Table 15, mix it by oscillation and keep it on the ice box for standby.
[0222] Table 15 Preparation of PCR Reaction Solution
[0223] Component Standard amount for one reaction (μL) PCR Enzyme Mix 75 PCR Primer Mix 7.5 Nuclease - free Water 67.5 Total Volume 150
[0224] 9.2 Add the PCR reaction solution into the reaction tube from which the supernatant was removed in step 5) of the ligation reaction, resuspend the labeled magnetic beads, mix them by pipetting, and after instantaneous centrifugation, aspirate half of the mixed solution (75 μL) into a new 0.2 mL PCR tube.
[0225] 9.3 Place the above PCR tube on a gene amplifier and perform PCR amplification reaction according to the reaction conditions shown in Table 16.
[0226] Table 16 PCR Amplification Reaction Conditions
[0227]
[0228]
[0229] 9.4 After the PCR reaction is completed, centrifuge the product instantaneously, place it on the magnetic stand for 2 min until the liquid becomes clear, and take out the supernatant of the 2 tubes of PCR products of the same sample and mix them in a new 1.5 mL EP tube.
[0230] 9.5 After ensuring that the supernatant is completely recovered, discard the original PCR tube with the dried magnetic beads retained.
[0231] 10. Purification of PCR Products
[0232] 11. Library Quality Inspection
[0233] 1) The concentration of the 31 μL library should be ≥ 25 ng / μL, and the input amount of the pre-library for the subsequent hybridization capture reaction is 750 ng (single hybridization). For multiple hybridizations, the input amount of each pre-library is 500 ng.
[0234] 2) The fragment distribution is between 200 bp and 2000 bp.
[0235] II. Hybridization Capture
[0236] 1. Preparation before Hybridization Capture Experiment
[0237] 1.1 Take out Hyb Human Block and RNase Block from the -20°C refrigerator, place them on an ice box to melt, briefly vortex and mix evenly, then centrifuge instantaneously, and store them temporarily on the ice box;
[0238] 1.2 Take out the supporting Blocking Oligo from the -20°C refrigerator, place it on an ice box to melt, briefly vortex and mix evenly, then centrifuge instantaneously, and store it temporarily on the ice box;
[0239] 1.3 Take out the probes to be used from the -20°C refrigerator, place them on an ice box to melt, briefly vortex and mix evenly, then centrifuge instantaneously, and store them temporarily on the ice box;
[0240] 1.4 Take out the library to be hybridized and captured from the -20°C refrigerator, place it on an ice box to melt, briefly vortex and mix evenly, then centrifuge instantaneously, and store it temporarily on the ice box;
[0241] 1.5 Take out TargetSeq Hyb Buffer v2, melt it at room temperature, briefly vortex and mix evenly, then centrifuge instantaneously. It is necessary to place TargetSeq Hyb Buffer v2 in a 37°C water bath and heat it until the reagent is completely dissolved before use.
[0242] 2. Hybridization of Library and Probe
[0243] 2.1 Take 750 ng of the library and add it to a PCR tube, and make a mark; when multiple libraries are hybridized together, add 500 ng of each library;
[0244] 2.2 Add 1.8 times the volume of purified magnetic beads to the library, gently pipette and mix evenly, and incubate at room temperature for 5 min;
[0245] 2.3 Place the PCR tube on a magnetic rack for 3 min until the solution becomes clear;
[0246] 2.4 Keep the PCR tube on the magnetic rack, discard the supernatant, add 200 μL of 80% ethanol solution to the PCR tube, and let it stand for 30 s;
[0247] 2.5 Keep the PCR tube on the magnetic stand, discard the supernatant, and add 200 μL of 80% ethanol solution to the PCR tube again. After standing for 30 s, completely discard the supernatant (you can perform a short centrifugation to centrifuge the liquid adhering to the wall to the bottom of the tube, and then use a 10 μL pipette to discard the residual ethanol solution at the bottom).
[0248] 2.6 Ensure that the PCR tube is on the magnetic stand and let it stand at room temperature for 3 - 5 min to air-dry the magnetic beads and completely volatilize the residual ethanol.
[0249] 2.7 Add the hybridization reaction solution to the PCR tube according to Table 17:
[0250] Table 17
[0251]
[0252] 2.8 Pipette and mix well, and let it stand at room temperature for 3 min.
[0253] 2.9 Perform a short centrifugation, place the PCR tube on the magnetic stand for 3 min until the solution becomes clear.
[0254] 2.10 Use a pipette to aspirate 28 μL of the supernatant into a new PCR tube, gently pipette and mix well, and perform a short centrifugation.
[0255] 2.11 Set the parameters of the PCR instrument as shown in Table 18, place the hybridization reaction solution on the PCR instrument, and run the program.
[0256] Table 18
[0257] Temperature Time Hot lid temperature 85℃ 80℃ 5 min 50℃ Hold
[0258] 2.12 The hybridization reaction time is 16 h.
[0259] 3. Preparation before the capture experiment
[0260] 3.1 Take out the Cap Beads from the 4 °C refrigerator in advance, mix well and place them at room temperature for 30 min to equilibrate.
[0261] 3.2 Prepare 80% ethanol with absolute ethanol and Nuclease-Free Water in advance, store it at room temperature for temporary use, and use it in the purification step.
[0262] 3.3 Take out the Wash Buffer 1. If there is precipitation, place the Wash Buffer 1 in a 37 °C water bath and heat it until the precipitation completely dissolves before use.
[0263] 3.4 Take out the TargetSeq Wash Buffer 2 v2, and preheat it on a 50 °C water bath.
[0264] 3.5 Add 50 μL of Cap Beads into a new PCR tube, place it on a magnetic stand for 1 min. Wait until the solution becomes clear, then discard the supernatant.
[0265] 3.6 Remove the PCR tube from the magnetic stand, add 180 μL of Binding Buffer, pipette or vortex to mix well to resuspend the magnetic beads.
[0266] 3.7 After a brief centrifugation, place the PCR tube on the magnetic stand for 1 min. Wait until the solution becomes clear, then discard the supernatant.
[0267] 3.8 Repeat steps 3.6 - 3.7 twice, washing the magnetic beads with Binding Buffer three times in total.
[0268] 3.9 Remove the PCR tube from the magnetic stand, add 180 μL of Binding Buffer, pipette or vortex to mix well, and immediately proceed to the next step.
[0269] 4. Target Region DNA Capture
[0270] 4.1 Keep the hybridization product obtained in step "2. Library and Probe Hybridization" on the PCR instrument. Add the 180 μL of Cap Beads prepared in step "3. Preparation before Capture Experiment" to the hybridization product, and pipette to mix well.
[0271] 4.2 Close the tube lid, remove the PCR tube from the PCR instrument, place it on a vertical rotator mixer with a rotation speed not exceeding 10 rpm, and bind at room temperature for 30 min (if there is no vertical rotator mixer in the laboratory, it can be bound at room temperature for 30 min, and invert the tube up and down several times every 5 min during this period for mixing).
[0272] 4.3 Remove the PCR tube, briefly centrifuge, place it on the magnetic stand for 2 min. After the solution becomes clear, discard the supernatant.
[0273] 4.4 Remove the PCR tube from the magnetic stand, add 150 μL of Wash Buffer 1 to the PCR tube, gently pipette to mix well to resuspend the magnetic beads, replace the tube lid with a new one, then place it on a vertical rotator mixer and wash at room temperature for 15 min with a rotation speed not exceeding 10 rpm.
[0274] 4.5 Remove the PCR tube, briefly centrifuge, place it on the magnetic stand for 2 min. After the solution becomes clear, discard the supernatant.
[0275] 4.6 Remove the PCR tube from the magnetic stand, add 150 μL of TargetSeq WashBuffer2 v2 pre - heated to 50 °C, gently pipette to mix well, briefly centrifuge, then place it on a thermostatic shaker mixer or metal bath and incubate at 50 °C for 10 min.
[0276] 4.7 Remove the PCR tube, centrifuge briefly, place it on the magnetic stand for 2 min. After the solution becomes clear, discard the supernatant;
[0277] 4.8 Repeat steps 4.6 - 4.7 once; for the Panel with the probe coverage area less than 200 kb, steps 4.6 - 4.7 can be repeated twice, which can further improve the specificity and stability of capture;
[0278] 4.9 Remove the PCR tube from the magnetic stand, add 150 μL of TargetSeq Buffer2 v2 preheated at 50 °C, gently pipette and mix well, centrifuge briefly, place it on a thermostatic shaker or metal bath, and incubate at 50 °C for 10 min;
[0279] 4.10 Remove the PCR tube, centrifuge briefly, gently pipette and mix well, transfer all the liquid (including magnetic beads) to a new PCR tube, place it on the magnetic stand for 2 min. After the solution becomes clear, discard the supernatant;
[0280] 4.11 Keep the PCR tube on the magnetic stand, add 200 μL of 80% ethanol to the PCR tube, let it stand for 30 s and then completely discard the ethanol solution (the residual ethanol can be discarded with a 10 μL pipette), and air-dry the magnetic beads at room temperature to completely volatilize the residual ethanol;
[0281] 4.12 Add 22.5 μL of Nuclease-Free Water to the PCR tube, remove the PCR tube from the magnetic stand, briefly vortex to resuspend and mix the magnetic beads well, and proceed to the next amplification reaction.
[0282] 5. Post-capture PCR amplification
[0283] 5.1 Take out the Post PCR Master Mix and PCR primer mixture from the -20 °C refrigerator in advance, place them on the ice box to melt, and store them temporarily on the ice box after melting;
[0284] 5.2 Please double-check whether the PCR primer mixture is used correctly before PCR, briefly vortex the Post PCR MasterMix and PCR primer mixture, and centrifuge briefly;
[0285] 5.3 Prepare the PCR reaction solution according to Table 19 below:
[0286] Table 19
[0287] Reagent Volume Magnetic bead suspension obtained in Step 4 22.5 μL PCR Primer Mix 2.5 μL Post PCR Master Mix 25 μL Total Volume 50 μL
[0288] Note: Step 4 refers to "4. Target region DNA capture".
[0289] 5.4 After preparation, use a pipette to aspirate and mix well, and then quickly transfer it to a PCR instrument. Do not use the method of vortexing followed by centrifugation to mix;
[0290] 5.5 Set the PCR instrument program as shown in Table 20 below. Place the PCR reaction solution on the PCR instrument and run the program;
[0291] Table 20
[0292]
[0293]
[0294] 5.6 After the program is completed, proceed to the next step of magnetic bead purification.
[0295] 6. Post-amplification purification
[0296] 6.1 Take out the purification magnetic beads, mix well and place them at room temperature for 30 min to equilibrate;
[0297] 6.2 Add 1.1 times the volume of magnetic beads (55 μL) to the PCR product in STEP 5, aspirate or vortex to mix well, and let it stand at room temperature for 5 min;
[0298] 6.3 Centrifuge briefly and place the PCR tube on a magnetic rack for 3 min until the solution becomes clear;
[0299] 6.4 Keep the PCR tube on the magnetic rack, discard the supernatant, add 200 μL of 80% ethanol solution to the PCR tube, and let it stand for 30 s;
[0300] 6.5 Keep the PCR tube on the magnetic rack, discard the supernatant, add 200 μL of 80% ethanol solution to the PCR tube again, and after standing for 30 s, completely discard the supernatant (you can centrifuge briefly to centrifuge the liquid adhering to the wall to the bottom of the tube, and then use a 10 μL pipette to discard the residual ethanol solution at the bottom);
[0301] 6.6 Ensure that the PCR tube is on the magnetic rack and let it stand at room temperature for 3 - 5 min to dry the magnetic beads and completely volatilize the residual ethanol;
[0302] 6.7 Add 25 μL of Nuclease-Free Water, remove the PCR tube from the magnetic rack, aspirate or vortex to mix well, and let it stand at room temperature for 2 min;
[0303] 6.8 Centrifuge briefly and place the PCR tube on the magnetic rack for 2 min until the solution becomes clear;
[0304] 6.9 Use a pipette to aspirate 23 μL of the supernatant and transfer it to a new PCR tube; store the captured library in a -20°C refrigerator. The captured library can be stored in a -20°C refrigerator for one month;
[0305] 6.10 Take 1 μL of the library and measure the library concentration on a Qubit 4.0 Fluorometer using the Qubit dsDNA HS Assay Kit reagent, and record the library concentration.
[0306] 6.11 Take 1 μL of the library and perform fragment quality inspection using a Fragment Analyzer. The fragment size should be basically the same as that of the pre-library.
[0307] III. Sequencing the Captured Library on the Machine
[0308] Perform machine sequencing with reference to the sequencing kit instructions.
[0309] IV. Bioinformatics Analysis
[0310] The results of haplotype construction detection are shown in Table 21.
[0311] Table 21
[0312]
[0313]
[0314] For the genotyping map, see Figures 1 - 8 , and the results show that the methods of this application have successfully obtained the pathogenic genes HBA1 / HBA2 and HBB haplotypes of thalassemia, and the SNP sites of the genotyping results are sufficient and the genotyping accuracy is relatively high.
[0315] Example 2
[0316] The B3 family was detected using this method and the multiplex PCR method respectively, and the results were compared.
[0317] Multiplex Polymerase Chain Reaction (mPCR) technology, also known as multiplex PCR, is a molecular biology technique that simultaneously uses multiple pairs of primers in a single reaction system to specifically amplify multiple targets. In pre-implantation testing, multiplex PCR technology can be used to simultaneously detect multiple gene loci, including mutation sites of single-gene genetic diseases, chromosomal abnormalities, etc. Multiplex PCR can detect multiple targets in one reaction, significantly improving the experimental efficiency. Compared with performing multiple PCR detections separately, multiplex PCR reduces reagent and labor costs, shortens the total time from sample to result, and speeds up the detection process. However, there are technical complexities in multiplex PCR detection. Designing appropriate primers and reaction conditions is relatively complex, and the compatibility between multiple pairs of primers needs to be considered. At the same time, there is a risk of non-specific amplification. The simultaneous reaction of multiple pairs of primers may increase the risk of non-specific amplification and affect the accuracy of the results.
[0318] The comparison between the results of conventional WGS genotyping and the genotyping results of this method is shown in Table 22.
[0319] Table 21
[0320]
[0321]
[0322] The results of the haplotype comparison of the parents in family B3 are shown in Figure 9 and Figure 10 . The results show that the WGS haplotype blocks are shorter, and there are unassembled regions between the downstream blocks, which cannot be connected into larger haplotype blocks, resulting in a reduction in the number of effective genotyping sites and a risk of genotyping failure.
[0323] The results of the haplotype comparison of the embryos in family B3 are shown in Figure 11 . Figure 11 The genomic region shown in the figure is chr16: 0 - 700000, which is the 500kb range region upstream and downstream of the target gene HBA. The regions within the two red lines are the regions where the HBA2 gene and the HBA1 gene are located. The left part of the map is the sample ID, and the scatter plot on the right is the genotyping sites of the two haplotypes of each sample. In the figure, the yellow scatter points represent the variant-carrying haplotype of the female, the green is the normal haplotype of the female, the blue scatter points are the variant-carrying haplotype of the male, and the red is the normal haplotype of the female. There are two haplotypes in the haplotype of the embryo, and the color of the scatter points is used to distinguish which haplotype of the parent the embryo's haplotype is inherited from. The genotyping results of the embryos by this method are: U24011701P01 is of the --SEA / --SE A type, U24011701P02 is of the --SEA / --SE A type, and the genotyping results of the embryos by the WGS method are that U24011701P01 is of the --SEA / --SE A type and U24011701P02 is of the --SEA / --SE A type.
[0324] Based on the above figures, it can be seen that the haplotype results of the method of this application are used for haplotype genotyping with embryo SNPs, and the genotyping results are consistent with the conventional WGS results, and no proband is required; the number of final effective genotyping sites of the embryos by the method of this application is close to that of the conventional WGS method, and no proband is required; in the results of WGS, there are a certain number of undetected SNPs. In the method of this application, based on high-depth sequencing, there are no undetected cases in the SNP results of the target region; the sequencing depth of the multiplex PCR method is uneven, and the sequencing depth in some regions is relatively low, resulting in smaller genotyping blocks and prone to assembly breakage areas, affecting haplotype genotyping.
[0325] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0326] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims, and the specification and drawings can be used to explain the content of the claims.
Claims
1. A method for constructing a library for haplotype analysis, characterized in that: It includes the following steps: Sample pretreatment: lyse the sample to obtain the released sample DNA; Interruption reaction: mixing the sample DNA with a fragmentation enzyme to obtain a sample DNA bound to the fragmentation enzyme; Molecular tag labeling: the sample DNA bound to the fragmentation enzyme is mixed with magnetic beads with different molecular tags, and the sample DNA and the molecular tags are hybridized and connected through a bridging oligo to obtain labeled sample DNA; Fragmentation labeling enzyme release: the labeled sample DNA is mixed with a denaturation buffer to denature and inactivate the fragmentation enzyme, and fragmented DNA molecules are obtained. Fragments from the same DNA carry the same molecular label; Sequencing adapter ligation: adding a sequencing adapter to the 3' end of the fragmented DNA molecule to obtain a pre-capture library; Target region probe hybridization capture: adding labeled probes to the pre-capture library, hybridization incubation, to obtain a labeled pre-capture library; The pre-capture library is mixed with magnetic beads that specifically bind to the label on the probe to capture the target fragment, and optionally amplified and enriched to obtain a post-capture library.
2. The method according to claim 1, characterized in that In the step of sequencing adapter ligation, a sequencing adapter is added to the 3' end of the DNA molecule by one or more of DNA ligase and transposase.
3. The method according to claim 2, characterized in that The method also includes a step of amplifying and enriching the pre-capture library to meet the input amount of hybridization capture of the target region probe; Optionally, amplification and enrichment are achieved by PCR; Further optionally, the reaction conditions of the PCR are: (i) preheating at 98°C for 3 min; (ii) denaturation at 95°C for 30 s, annealing at 55°C to 65°C for 30 s, extension at 72°C for 2 min, and 12 to 15 cycles; and (iii) maintaining at 72°C for 10 min.
4. The method according to any one of claims 1 to 3, characterized in that The target region probe hybridization capture step satisfies one or more of the following conditions 1) to 3): 1) The label of the probe is biotin, and the magnetic beads that specifically bind to the label on the probe are streptavidin magnetic beads; 2) Amplification and enrichment by PCR; Optionally, the reaction conditions of the PCR include: (i) preheating at 95°C for 1 min; (ii) denaturation at 98°C for 20 s, annealing at 60°C for 30 s, and extension at 72°C for 30 s, for 8 to 13 cycles; and (iii) maintaining at 72°C for 5 min; 3) The number of hybridization hot elutions is 2 to 4 times; optionally, the number of hybridization hot elutions is 4 times; optionally, the temperature of hybridization hot elution is 50°C to 60°C.
5. The method according to any one of claims 1 to 4, characterized in that: The sample is selected from the group consisting of blood, semen, tissue and cultured cells.
6. A method for constructing a haplotype of a pathogenic variant carrier, characterized in that: It includes the following steps: Constructing a library using the method according to any one of claims 1 to 5, wherein the samples used to construct the library are from carriers of pathogenic mutations; The constructed library was sequenced to obtain long-fragment sequencing data.
7. The method according to claim 6, characterized in that It also includes: The long-fragment sequencing data of the parents of the pathogenic variant carriers were compared with the reference genome and the target region variation analysis to obtain the genotype information of the pathogenic variant; selecting heterozygous SNPs in the pathogenic variant carrier parent from the genotype information; Assembling the long-fragment sequencing data of the parents of the pathogenic variant carriers to obtain their genomic contigs (overlapping groups) and marking them with numbers, comparing the genomic contigs with the reference genome to obtain alignment information, and marking the coverage interval of the genomic contigs in the reference genome; Construct haplotypes of SNPs in the target region; The contig information is assembled based on the long fragment where the pathogenic variation is located, and matched with the haplotype typing result to obtain the haplotype typing result of the pathogenic variation site; the haplotype of the parent of the pathogenic variation carrier is marked according to the typing result of the pathogenic site, and the pathogenic / normal haplotype of the parent of the pathogenic variation carrier is constructed.
8. The method according to claim 6 or 7, characterized in that The design principles of the probe include: according to the location information of the pathogenic variation, selecting high-frequency mutation SNP sites within 1Mb upstream and downstream of the pathogenic variation, covering the full length of the pathogenic gene, and setting the interval range between adjacent SNP sites to 1kb to 5kb.
9. The method according to any one of claims 6 to 8, characterized in that The reference genome includes human reference genome GRCh37 or GRCh38.
10. The method according to any one of claims 7 to 9, characterized in that The method does not rely on a proband or a reference sample; the reference sample includes one or more of the affected embryo and the immediate family members of the couple.