Method for mixing high-throughput sequencing libraries in proportion

By labeling characteristic sequences on DNA molecules in high-throughput sequencing libraries and recovering specific molecule quantities using conjugates, the problem of uneven DNA ratios in libraries was solved, enabling efficient and accurate proportional mixing of libraries and improving the utilization rate of sequencing data and experimental efficiency.

CN120330309BActive Publication Date: 2026-01-02北京明识至善生物技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510790189.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2026-01-02
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

In existing high-throughput sequencing technologies, uneven library DNA ratios lead to uneven data output, affecting the accuracy of genome analysis and tumor gene mutation screening. Existing quantitative methods such as Qubit quantitative PCR and real-time quantitative PCR have errors and are cumbersome to operate.

Method used

By labeling characteristic sequences onto DNA molecules in a high-throughput sequencing library, a complex is formed between the label and a conjugate, and a set number of DNA molecules are obtained through selective recovery, thus achieving proportional mixing.

Benefits of technology

It simplifies the library quantification and sample mixing process, improves the utilization rate of sequencing data and experimental efficiency, ensures that each sample is read proportionally, and reduces experimental costs and operational difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120330309B_ABST
    Figure CN120330309B_ABST
Patent Text Reader

Abstract

The application provides a method for proportionally mixing high-throughput sequencing libraries, which comprises the following steps: labeling DNA molecules in a high-throughput sequencing library by using a marker; combining the marker-labeled DNA molecules with a binder to obtain a binder-marker-library molecule complex; recovering the complex to obtain a set number of library DNA molecules; respectively treating a plurality of high-throughput sequencing libraries by using the method, wherein, according to the expected data amount proportion in the mixed library, a corresponding molecule number proportion of the binder is added into each high-throughput sequencing library to obtain a corresponding molecule number proportion of high-throughput sequencing library DNA molecules; and mixing the library DNA molecules derived from different high-throughput sequencing libraries. The method solves the problems in the prior art, such as the dependence on corresponding professional detection equipment and complicated operation process when using Qubit fluorescence quantification and real-time quantitative PCR (qPCR) to mix libraries.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of sequencing, in particular, to a method for proportionally mixing high-throughput sequencing libraries. BACKGROUND

[0002] High-throughput sequencing technology (HTS), also known as next-generation sequencing (NGS), has completely changed the way of research in genomics and molecular biology since its inception in the early 21st century. NGS technology can quickly and in parallel sequence millions to billions of DNA molecules, greatly improving the speed of obtaining genomic data and research efficiency, but at the same time, it also puts forward new requirements for sample processing and sequencing library homogenization.

[0003] In the NGS workflow, from DNA extraction, fragmentation, end repair, adapter ligation, PCR amplification to sequencing, each step is crucial. Among them, the proportional mixing of multi-sample sequencing libraries is a key step to ensure efficient use of sequencing data. If the proportion of each sample library DNA in the sequencing library does not meet the expected value, it will directly affect the data yield of each sample during sequencing, resulting in some sample data overload and other sample data insufficient, thereby affecting the accuracy of subsequent genomic analysis, genetic disease detection and tumor gene mutation screening.

[0004] In the prior art, the methods for library homogenization mainly include Qubit fluorescence quantification method and real-time quantitative PCR (qPCR). When using the Qubit fluorescence quantification method for detection, the method cannot distinguish between library DNA and non-target DNA (such as primer dimers), which will cause quantitative errors. At the same time, since this method is sensitive to DNA fragments with special structures (such as hairpin structures), the final detection results will be significantly deviated from the true concentration. When using real-time quantitative PCR (qPCR) for detection, the experimental period will be lengthened, and the timeliness will be poor. SUMMARY

[0005] The main purpose of the present application is to provide a method for proportionally mixing high-throughput sequencing libraries to solve the problems of relying on corresponding professional detection equipment and cumbersome operation process when using the Qubit fluorescence quantification method and real-time quantitative PCR (qPCR) in the prior art.

[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for proportionally mixing high-throughput sequencing libraries is provided, comprising:

[0007] a) labeling the DNA molecules in the high-throughput sequencing library with a label to obtain DNA molecules containing the label;

[0008] b) binding the binding agent to the labeled DNA molecules to obtain a binding agent-labeled DNA molecule complex;

[0009] c) recovering the complex to obtain a set number of library DNA molecules;

[0010] d) performing a) - c) on each of a plurality of high-throughput sequencing libraries, wherein the binding agent is added to each high-throughput sequencing library in a proportion corresponding to the expected data volume of the mixed library, to obtain a proportion of the set number of high-throughput sequencing library DNA molecules; and mixing the library DNA molecules from different high-throughput sequencing libraries to obtain a proportionally mixed mixed library;

[0011] wherein the proportion of the set number is less than the number of library DNA molecules.

[0012] Further, the label is located in the region of the library universal sequence after labeling.

[0013] Optionally, the library universal sequence is a library universal primer or a linker, and the label can be modified to the library universal sequence.

[0014] Optionally, the label is a characteristic sequence that can be recognized by an enzyme.

[0015] Further, the label and the binding agent are a set of corresponding substances that can chemically react to form a covalent bond or can physically bind.

[0016] Further, the step of labeling the DNA molecules in the high-throughput sequencing library with the label comprises:

[0017] labeling the DNA molecules in the high-throughput sequencing library with the label by chemical means or enzymatic means to obtain labeled DNA molecules;

[0018] or by adding the label during chemical synthesis of the DNA primer, introducing the label through the labeled primer during PCR amplification of the library DNA molecules; or by adding the label through the terminal transferase at the end of the library DNA molecules; or by adding the label during synthesis of the library linker DNA.

[0019] Further, the method of recovering the complex is specific to retain the complex, including any one or combination of forward selection adsorption of the complex or reverse digestion to remove library DNA molecules that do not form the complex.

[0020] Further, the label is any one of an active group, a special base, or a characteristic sequence that can be specifically recognized by a DNA binding protein.

[0021] Further, the active group is any one of an azido group, an alkyne group; and / or, the special base is any one of a uracil base, a phosphorylated modified base, or a characteristic sequence that can be specifically recognized by a HUH DNA binding protein.

[0022] Further, the binding agent is any one of an alkyne-PEG-desulfitobiotin compound, an azido-PEG-desulfitobiotin compound, a biotinylated DNA binding protein, a DNA binding protein, a biotinylated B family DNA polymerase, or a B family DNA polymerase.

[0023] Further, the biotinylated DNA binding protein is any one of a biotinylated UdgX, TYLCV, RepB, Tral, or a biotinylated HUH protein, or a combination of at least two or more thereof; and / or, the DNA binding protein is any one of a UdgX, TYLCV, RepB, Tral, or a HUH protein, or a combination of at least two or more thereof.

[0024] Further, when the binding agent is any one of a biotinylated compound, a biotinylated DNA binding protein, or a biotinylated B family DNA polymerase, the method of recovering the complex is a positive selection adsorption of the complex or a negative digestion to remove the library DNA molecules that do not form the complex; when the binding agent is any one of a DNA binding protein or a B family DNA polymerase, the method of recovering the complex is a negative digestion to remove the library DNA molecules that do not form the complex.

[0025] Further, the method of positive selection adsorption of the complex is any one of using a streptavidin magnetic bead or a streptavidin purification column; and the method of negative digestion to remove is any one of using a nuclease or an exonuclease.

[0026] Further, the method of mixing the high-throughput sequencing library in proportion is applied to a high-throughput sequencing instrument platform.

[0027] According to another aspect of the present application, a kit is provided, the kit comprising a test reagent in the method of mixing the high-throughput sequencing library in proportion described above, the test reagent comprising one or more of the following;

[0028] The binding agent, the label, and the recovery reagent have a set number of molecules.

[0029] Further, the kit comprises one or more of the following: a labeled primer mixture, a labeled buffer and a labeled enzyme, the compound having a set number of molecules, the biotinylated DNA binding protein, the DNA binding protein, the biotinylated B family DNA polymerase, or the B family DNA polymerase, the streptavidin magnetic bead, the streptavidin purification column, the nuclease, or the exonuclease.

[0030] Further, the kit comprises a labeled primer mixture, a labeling buffer and a labeling enzyme, a DNA binding protein and an exonuclease.

[0031] By using the technical solution of the present application, the DNA molecules in the high-throughput sequencing library are labeled by using a label, and then a binding agent with a set number of molecules is used to bind to the labeled DNA molecules. This method can effectively recover DNA molecules with a specific number of molecules. By using this method to process multiple high-throughput sequencing libraries, a mixed library with the same ratio of the number of molecules to the expected data volume can be obtained.

[0032] After processing by this method, the number of DNA molecules in the library between different samples is consistent with the expected data volume ratio, eliminating the excessive data output of high-concentration samples and the insufficient data output of low-concentration samples. This helps to improve the effective use of sequencing data and ensures that each sample is read in the expected proportion during the sequencing process.

[0033] Through simple labeling, binding and recovery steps, the cumbersome library quantification and sample mixing process is avoided, the experimental operation process is simplified, the experimental cost and operation difficulty are reduced, and the experimental repeatability and work efficiency are improved.

[0034] The use of binding agents with a set number of molecules enables the proportional mixing method to be implemented on manual and automated operation platforms, providing an efficient and accurate library sample mixing solution for high-throughput laboratories. BRIEF DESCRIPTION OF DRAWINGS

[0035] The drawings accompanying the specification of this application serve to provide further understanding of the present application, and the illustrative embodiments of the present application and their descriptions serve to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0036] Figure 1 A method flowchart of the embodiments of the present application is shown. DETAILED DESCRIPTION

[0037] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0038] High-Throughput Sequencing (HTS), also known as Next-Generation Sequencing (NGS), has revolutionized the way of research in genomics and molecular biology since its advent in the early 21st century. NGS technology enables sequencing millions to billions of DNA molecules in parallel, greatly improving the speed and efficiency of obtaining genomic data, but at the same time, it also poses new requirements for sample processing and sequencing library equalization.

[0039] In the NGS workflow, from DNA extraction, fragmentation, end repair, adapter ligation, PCR amplification to sequencing, each step is crucial. Among them, the proportional mixing of multi-sample sequencing libraries is a key step to ensure efficient use of sequencing data. If the proportion of each sample library DNA in the sequencing library does not meet the expected value, it will directly affect the data yield of each sample during sequencing, resulting in some sample data overload and other sample data insufficient, thereby affecting the accuracy of subsequent genomic analysis, genetic disease detection and tumor gene mutation screening.

[0040] In the prior art, the methods for high-throughput sequencing library equalization mainly include Qubit fluorescence quantification method and real-time quantitative PCR (qPCR). When using the Qubit fluorescence quantification method for detection, the method cannot distinguish between library DNA and non-target DNA (such as primer dimers), which will cause quantitative errors. At the same time, since this method is sensitive to DNA fragments with special structures (such as hairpin structures), the final detection results will deviate significantly from the true concentration. When using real-time quantitative PCR (qPCR) for detection, the experimental period will be lengthened, and the timeliness will be poor.

[0041] The main purpose of the present application is to provide a method for proportional mixing of high-throughput sequencing libraries, to solve the problems of using Qubit fluorescence quantification method and real-time quantitative PCR (qPCR) for library quantification and proportional mixing in the prior art, which rely on corresponding professional detection equipment and cumbersome operation processes.

[0042] As shown in Figure 1 The method for proportional mixing of high-throughput sequencing libraries provided by the present application comprises:

[0043] a) labeling the DNA molecules in the high-throughput sequencing library with a label to obtain DNA molecules containing the label;

[0044] b) binding the binding agent to the DNA molecules containing the label, and recovering to obtain a binding agent-label-library molecule complex;

[0045] c) recovering the complex to obtain a set number of library DNA molecules.

[0046] d) separately performing a) - c) on a plurality of high-throughput sequencing libraries, wherein the binding agent is added to each high-throughput sequencing library in a proportion of the number of molecules corresponding to the proportion of the expected amount of data in the mixed library, to obtain a proportion of the number of molecules of high-throughput sequencing library DNA molecules; and mixing the library DNA molecules from different high-throughput sequencing libraries to obtain a proportionally mixed mixed library.

[0047] wherein the number of molecules corresponding is less than the number of molecules of the library DNA.

[0048] The method for proportionally mixing high-throughput sequencing libraries provided in the present application solves the problem of data bias caused by uneven library concentration in the field of high-throughput sequencing by meticulous and comprehensive labeling, binding and recovery steps, and significantly improves the utilization rate of sequencing data.

[0049] By precisely controlling the number of molecules (number of molecules corresponding) of the binding agent used, the expected proportion of molecules of library DNA can be obtained from libraries of different concentrations, achieving the effect of proportionally mixing.

[0050] The use of proportionally mixed mixed libraries avoids the waste of sequencing resources on high-concentration library fragments, ensuring the rational allocation of resources in each sequencing run.

[0051] Directly using the labeling and binding steps to achieve proportionally mixing saves the cumbersome library quantification and sample mixing process, simplifies the overall process of library production, and improves the efficiency of the laboratory.

[0052] It is suitable for both manual and automated operation platforms, meaning that the laboratory can choose the operation mode flexibly according to its own conditions, which is convenient for small-scale laboratories with manual operation, and also suitable for large facilities that pursue high-throughput.

[0053] The method for proportionally mixing high-throughput sequencing libraries provided in the present application achieves simple data volume proportion mixing by using a proportion of the number of molecules of the binding agent to specifically bind to the labeled DNA molecules, and then effectively recovering it, and its technical effects mainly reflect in the following aspects:

[0054] By setting the number of molecules corresponding to the binding agent to be less than the number of molecules of the labeled library DNA, the number of labeled DNA molecules that can be bound by each unit of binding agent is fixed, so that the amount of binding agent can be controlled in the expected proportion to obtain the expected amount of complex generation, achieving the goal of proportionally mixing the library.

[0055] The covalent or non-covalent binding methods used in the binding stage and the recovery stage, such as biotin-streptavidin magnetic bead method, click chemistry method and biotinylated DNA binding protein method, can efficiently recover the labeled DNA fragments. These methods combine the techniques of affinity purification or selective adsorption to ensure the integrity and purity of the target DNA fragments during the recovery process and reduce the interference of non-specific binders.

[0056] This method avoids over-sampling of high-concentration library fragments during sequencing and ensures that low-concentration fragments are fully sequenced, significantly improving the utilization rate of sequencing data.

[0057] The method of the present application only needs conventional PCR laboratory equipment and simple pipetting operation, which is more convenient in operation than traditional methods such as real-time quantitative PCR and Qubit detection, reduces human errors in the experimental process, and reduces the dependence on professional equipment, so that the experiment is more easy to carry out and repeat under various conditions.

[0058] By adjusting the amount of binder, the recovery amount of labeled DNA molecules can be flexibly controlled, so as to adjust the final concentration and proportion of high-throughput sequencing library according to different experimental purposes and the needs of sequencing platform, and improve the controllability of sequencing data output. The throughput of the sequencing platform can be maximized, the sequencing efficiency is improved, and the complexity in subsequent data analysis is reduced.

[0059] The method provided in the present application is implemented for high-throughput sequencing library, and the construction of high-throughput sequencing library is needed.

[0060] Example 1 Construction of high-throughput sequencing library for testing;

[0061] The specific construction steps are: 100 ng of genomic DNA is used as the initial material to ensure sufficient DNA fragments in subsequent operations, the library construction kit (QIAGEN, 180477) is used for fragmentation, and the reaction program is 4°C, 1 min; 32°C, 10 min; 65°C, 30 min (intended to control the activity of the enzyme through a series of temperature changes, 4°C is the pre-cooling stage, 32°C is the optimum temperature of the enzyme, and 65°C is the inactivation temperature of the enzyme, which ensures the uniformity of the DNA fragments and the proper termination of the enzyme activity); cool to 4°C for storage, obtain the fragmentation product; use magnetic beads (Novozyme, N411-02) to purify the fragmentation product to remove excess enzymes, reaction byproducts and unbound DNA fragments to improve the purity of the DNA fragments, add specific DNA adapters (also known as adapters or linkers) to the sample tube containing the fragmented DNA to facilitate subsequent PCR amplification and sequencing, wherein the specific DNA sequence of the adapter usually contains a primer binding site recognized by the sequencing platform and a characteristic sequence required for sequencing, add the ligation reagent of the library construction kit (QIAGEN, 180477), which includes DNA ligase and buffer, to connect the adapter to both ends of the fragmented DNA to form a labeled library that can be used for sequencing, in the ligation reaction, the working temperature needs to be set at 20°C, and the reaction time is set to 20 min to ensure sufficient binding of the adapter to the DNA fragments; cool to 4°C for storage, obtain the ligation product; use magnetic beads (Novozyme, N411-02) to purify the ligation product to remove unconnected adapters, excess ligation reagents and any byproducts to ensure high quality of the ligation product, the ligation product is a DNA fragment with an adapter, and after purification, a test library, i.e., a high-throughput sequencing library, is obtained, which can be stored at 4°C to maintain the stability and activity of the DNA.

[0062] Example 2 Test of azido active group labeling and conjugate alkyne-PEG-desulfo-biotin;

[0063] Synthesis of azido compound (N3) modified primer (Shengong synthesis);

[0064] P5-N3: 5’Azide (N3)-AATGATACGGCGACCACCGA-3’;

[0065] P7: 5’-CAAGCAGAAGACGGCATACGA-3’;

[0066] The test library of Example 1 and KAPA hot start HIFI high fidelity enzyme ReadyMix (Roche: KK2602) were used to amplify the library, and the reaction program was 98°C, 3min; (98°, 10s; 55°C, 5s; 72°C, 30s) amplification for 4 cycles; 72°C, 5min; cooling to 4°C for storage. The azido-labeled amplification product was purified using magnetic beads (Novozyme, N411-02), and the azido-labeled library was obtained. Real-time quantitative QPCR instrument was used to quantify the test library, so as to calculate the initial input amount;

[0067] Different input amount labeled libraries (400fmol, 600fmol, 800fmol, 1000fmol) and quantitative (300fmol) diphenyl cyclooctyne-PEG-desulfitobiotin were reacted, and the reaction program was 45°C, 1h; cooling to 4°C for storage; then reacted with 10ul M270 magnetic beads (Thermo Fisher, 65306) for 5min; washed twice with washing buffer, eluted with 40mM NaOH solution, and neutralized with 40mM acetic acid solution to obtain the mixed library with an expected ratio of 1:1.

[0068] Real-time quantitative QPCR instrument was used to quantitatively analyze the test library, and the ratio effect of the mixed library was evaluated. The results are shown in Table 1. In the case of initial input amount ≥600fmol, the expected ratio of the mixed library can be achieved, and the results show lower library output, indicating lower reaction efficiency, which may be related to the purity of the compound. In addition, different alkynyl structures may also affect the reaction efficiency.

[0069] Table 1

[0070]

[0071] Example 3 Special U base labeling and binder biotinylated DNA binder library test;

[0072] Synthesis of "U" base labeled primer (Shengong synthesis);

[0073] 5`P-P5-U9: 5`P-CTCAAGTG / ideoxyU / AATGATACGGCGACCACCGA-3';

[0074] 5`P-P7: 5`P-CAAGCAGAAGACGGCATACGA-3';

[0075] The library was amplified using 2xTaq PCR StarMix (GenStar, ZA012-101S), and the reaction program was 95°C, 3 min after mixing in the sample tube; 4 cycles of (95°C, 10 s; 55°C, 30 s; 72°C, 30 s); 72°C, 5 min; and cooling to 4°C for storage. The "U" base labeled amplification product was purified using magnetic beads (Novozyme, N411-02) to obtain the "U" base labeled library. The test library was quantified using a real-time quantitative QPCR instrument to facilitate calculation of the initial input amount.

[0076] The labeled library (200 fmol, 400 fmol, 600 fmol, 1000 fmol) with different input amounts was combined with 200 fmol of quantified biotinylated protein Biotin-UDGX (obtained by expression and purification in the laboratory) at 37°C for 30 min, and then reacted with 10 ul M270 magnetic beads (Thermo Fisher, 65306) for 5 min. The washing buffer was used for washing twice, eluted with 40 mM NaOH solution, and then neutralized with 40 mM acetic acid solution to obtain the mixed library with an expected ratio of 1:1.

[0077] The mixed library was quantitatively analyzed using a real-time quantitative QPCR instrument to evaluate the proportional effect of the mixed library. The results are shown in Table 2. In the case of initial input amount ≥ 400 fmol, the expected ratio of the mixed library can be achieved.

[0078] Table 2

[0079]

[0080] Example 4: Special "U" base labeling and DNA binding protein library test

[0081] The "U" labeled library prepared in Example 3 was used. Different input amounts of "U" base labeled library (200 fmol, 400 fmol, 600 fmol, 1000 fmol) were combined with 200 fmol of quantified DNA binding protein UdgX (obtained by expression and purification in the laboratory) at 37°C for 30 min, and then 1 ul of Lambda exonuclease (NEB, M0262S) was added for digestion reaction at 37°C for 30 min, 75°C for 10 min, and cooling to 4°C for storage to obtain the library with an expected ratio of 1:1.

[0082] The mixed library was quantitatively analyzed using a real-time quantitative QPCR instrument to evaluate the proportional effect of the mixed library. The results are shown in Table 3. In the case of initial input amount greater than 200 fmol, the expected ratio of the mixed library can be achieved.

[0083] Table 3

[0084]

[0085] Example 5 Application test of high-throughput sequencing library proportion mixing kit

[0086] The test library was labeled by using the labeling reagent in the kit. Specifically, 10 ul of the test library, 2 ul of the labeling primer mixture (the same as in Example 4), 25 ul of the labeling enzyme mixture, and 13 ul of water were taken, and the total system was 20 ul. After mixing and instant separation, PCR was performed according to the following procedure: 95°C, 3 min; (95°C, 10 s; 55°C, 30 s; 72°C, 30 s;) 4 cycles of amplification; 72°C, 5 min; and cooling to 4°C for storage. The labeled amplification product was purified using magnetic beads (Novozyme, N411-02) to obtain the labeled test library. Real-time quantitative QPCR was used to quantify the test library to facilitate calculation of the initial input amount.

[0087] Test 1 of mixing samples in the expected proportion 1:1

[0088] The labeled library with different input amounts (200 fmol, 400 fmol, 600 fmol, and 1000 fmol) was combined with the quantified 200 fmol binding reagent at 37°C for 30 min, and then 1 ul of recovery reagent was added for digestion reaction at 37°C for 30 min, 75°C for 10 min, and cooling to 4°C for storage to obtain the test library. Real-time quantitative QPCR was used to quantitatively analyze the library to preliminarily evaluate the proportional effect of the library.

[0089] The libraries were mixed in equal volumes to obtain library 1 for sequencing using the Novaseq sequencer, and the actual data output was analyzed to further evaluate the effect of the kit application.

[0090] Test 2 of mixing samples in the expected proportion 1:1

[0091] The labeled library with different input amounts (200 fmol, 400 fmol, 600 fmol, and 1000 fmol) was combined with the quantified 400 fmol binding reagent at 37°C for 30 min, and then 1 ul of recovery reagent was added for digestion reaction at 37°C for 30 min, 75°C for 10 min, and cooling to 4°C for storage to obtain the test library. Real-time quantitative QPCR was used to quantitatively analyze the library to preliminarily evaluate the proportional effect of the library.

[0092] The libraries were mixed in equal volumes to obtain library 2 for sequencing using the Novaseq sequencer, and the actual data output was analyzed to further evaluate the effect of the uniformization kit application.

[0093] Test 3: mix test 1 and test 2 library in equal volume to get library 3, use Novaseq sequencer to sequence, analyze actual data output, further evaluate the effect of kit application;

[0094] The test results are shown in Table 4, and the following conclusions can be drawn:

[0095] The library proportion mixing kit can realize the expected proportion of the library when the initial input amount of the labeled library is ≥400 fmol.

[0096] The library proportion mixing kit can realize the expected proportion of the library when the initial input amount of the labeled library is ≥400 fmol.

[0097] Table 4

[0098]

[0099] The high-throughput sequencing library proportion mixing method in the present application is suitable for high-throughput gene sequencing, genetic disease monitoring, or tumor detection or pathogenic microorganism detection.

[0100] The present application also provides a kit comprising the test reagents in the high-throughput sequencing library proportion mixing method, and the test reagents comprise any one or more of the following: a label, a binding agent with a set number of molecules, and a recovery reagent.

[0101] Further, the kit further comprises one or more reagents of a labeled primer mixture, a labeled buffer and a labeled enzyme, a compound with a set number of molecules, a biotinylated DNA binding protein, a DNA binding protein, a biotinylated B family DNA polymerase or a B family DNA polymerase, a streptavidin magnetic bead, a streptavidin purification column, a nuclease or an exonuclease.

[0102] Further, the kit further comprises a labeled primer mixture, a labeled buffer and a labeled enzyme, a DNA binding protein and an exonuclease.

[0103] The kit in the present application comprises the test reagents in the high-throughput sequencing library proportion mixing method, and the test reagents comprise a plurality of the above, the use of the kit simplifies multiple steps in the high-throughput sequencing library proportion mixing method, standardizes the operation process, reduces the experimental result deviation caused by operation changes, and at the same time, since all test reagents in the kit are subjected to strict quality control and performance verification, the experimenter can obtain consistent labeling efficiency, binding strength and recovery rate each time, by integrating all test reagents in one kit, the need for purchasing, storing and preparing multiple reagents during the experiment is reduced, thereby reducing the experimental cost, at the same time, the standardized operation process also greatly shortens the experimental preparation and operation time, and improves the experimental efficiency.

[0104] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments in accordance with the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, devices, components and / or combinations thereof, but do not preclude the presence or addition of one or more other features, steps, operations, devices, components and / or combinations thereof.

[0105] The relative arrangement of parts and steps, numerical expressions, and numerical values set forth in the examples are not intended to limit the scope of the application unless specifically so stated. It is also to be understood that the drawings are not necessarily drawn to scale and that the dimensions of the various parts shown in the drawings are not intended to represent the actual proportions of the parts. Techniques, methods, and apparatus known to those of ordinary skill in the relevant art can not be discussed in detail, but can be assumed by those of ordinary skill in the art to be part of the specification. In all examples shown and discussed herein, any specific values should be interpreted as merely illustrative and not limiting. Other examples of the example embodiments can have different values. It is noted that like references and designations can indicate like items in the drawings, and once an item is defined in one drawing, it need not be discussed further in subsequent drawings.

[0106] In the description of the present application, it is to be understood that the orientation or positional relationships indicated by orientation words such as "front, back, upper, lower, left, right", "horizontal, vertical, perpendicular, horizontal" and "top, bottom" are generally based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present application and simplifying the description. Without the opposite indication, these orientation words do not indicate and imply that the indicated device or element must have a particular orientation or be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the scope of protection of the present application. The orientation words "inner, outer" refer to the inner and outer relative to the contour of each component.

[0107] For purposes of the description hereinafter, the terms "upper", "lower", "right", "left", "rear", "front", "vertical", "horizontal", and derivatives thereof shall relate to the application as oriented in the drawing. The terms "on", "above", "under", "below" and derivatives thereof shall relate to the application as oriented in the drawing, with the test of "above" and "under" being determined based on the position of the device in the drawing. Where the word "comprise" or variations such as "comprises" or "comprising" are used in the following description, it specifically expressly stated that other elements can also be present. Like reference numerals refer to like elements throughout. The term "coupled" as used herein is intended to mean either a direct connection between two elements or an indirect connection through one or more intervening elements. The term "associated with" as used herein is intended to mean either a direct connection between two elements or an indirect connection through one or more intervening elements.

[0108] In addition, it should be pointed out that the use of "first", "second" and the like words to qualify parts, is only for the convenience of distinguishing the corresponding parts, and the above words have no special meaning unless otherwise stated, and therefore cannot be understood as limiting the scope of protection of the present application.

[0109] The preferred embodiments of the application are described above in detail. The application can be modified and changed by those skilled in the art without departing from the spirit and principle of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall be included in the scope of protection of the application.

Claims

1. A method of high-throughput sequencing library proportional mixing, characterized in that, The method comprises the following steps: a) labeling DNA molecules in a high-throughput sequencing library with a label to obtain labeled DNA molecules; b) adding a binding agent in a set number of molecules, and binding the binding agent to the labeled DNA molecules through specific reaction between the binding agent and the label to obtain a binding agent-label-library molecule complex in an expected number of molecules; c) recovering the complex to obtain a set number of library DNA molecules; d) performing the steps a)-c) on a plurality of high-throughput sequencing libraries respectively, wherein the binding agent is added to each high-throughput sequencing library in a corresponding proportion of the number of molecules according to the proportion of the expected data amount in a mixed library, so as to obtain a corresponding proportion of the number of molecules of the high-throughput sequencing library DNA molecules; and mixing the library DNA molecules derived from different high-throughput sequencing libraries to obtain a proportionally mixed mixed library. The corresponding number of molecules is less than the number of library DNA molecules. The label and the binding agent are a group of corresponding substances that can chemically react to form a covalent bond or can physically bind. The method for recovering the complex is specific to retain the complex, including any one or a combination of forward selection adsorption of the complex or reverse digestion to remove library DNA molecules that do not form the complex.

2. The method of claim 1, wherein, The label is located in the region of the library universal sequence after labeling. The library universal sequence is a library universal primer or a linker, and the label can be modified to the library universal sequence. The label is a characteristic sequence that can be recognized by an enzyme.

3. The method of claim 1, wherein the high-throughput sequencing library scaling is performed by, The step of labeling the DNA molecules in the high-throughput sequencing library with the label comprises: labeling the DNA molecules in the high-throughput sequencing library with the label by chemical means or enzymatic means to obtain the labeled DNA molecules; or adding a label during chemical synthesis of a DNA primer, introducing the label through a labeled primer during PCR amplification of the library DNA molecules; or adding a label through a terminal transferase at the end of the library DNA molecules; or adding a label during synthesis of a library linker DNA.

4. The method of claim 1, wherein, The label is any one of an active group, a special base, or a characteristic sequence that can be specifically recognized by a DNA binding protein.

5. The method of claim 4, wherein the high-throughput sequencing library scaling is performed by, The active group is any one of an azido group or an alkyne group; and / or the special base is any one of a uracil base, a phosphorylated modified base, or a characteristic sequence that can be specifically recognized by a HUH DNA binding protein.

6. The method of claim 1, wherein, The binding agent is any one of an alkyne-PEG-desulfitobiotin compound, an azido-PEG-desulfitobiotin compound, a biotinylated DNA binding protein, a DNA binding protein, a biotinylated B family DNA polymerase, or a B family DNA polymerase.

7. The method of scaling of high-throughput sequencing libraries of claim 6, wherein, The biotinylated DNA binding protein is any one or a combination of at least two or more of biotinylated UdgX, TYLCV, RepB, Tral or biotinylated HUH protein; and / or, the DNA binding protein is any one or a combination of at least two or more of UdgX, TYLCV, RepB, Tral or HUH protein.

8. The method of claim 1, wherein the high-throughput sequencing library scaling is performed by, When the binding agent is any one of biotinylated compound, biotinylated DNA binding protein or biotinylated B family DNA polymerase, the method of recovering the complex is a forward selection adsorption complex or a reverse digestion to remove the library DNA molecules that do not form the complex; when the binding agent is any one of DNA binding protein or B family DNA polymerase, the method of recovering the complex is a reverse digestion to remove the library DNA molecules that do not form the complex.

9. The method of scaling of high-throughput sequencing libraries of claim 8, wherein, The method of forward selection adsorption complex is any one of using streptavidin magnetic beads or streptavidin purification column; the method of reverse digestion removal is any one of using nuclease or exonuclease.

10. The method of scaling of high-throughput sequencing libraries according to any one of claims 1 to 9, wherein, The method of mixing the high-throughput sequencing library in proportion is applied to a high-throughput sequencer platform.

Citation Information

Patent Citations

  • Method for rapid homogenization or equal proportion of DNA samples

    WO2017219511A1

  • Method for preparing strand-specific library for rapid detection of various types of rnas, and high-throughput sequencing technique

    WO2025000136A1