A method for constructing and detecting a gDNA library containing a unique combination of dual - ended library tags

Through the unique quality control method of double-ended library label combination, the sample misallocation problem caused by cross-contamination in multiple library sequencing is solved, and efficient and low-cost library label detection is achieved, ensuring the accuracy of sequencing results.

CN113957123BActive Publication Date: 2025-07-25GUANGZHOU BURNING ROCK DX CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111090137.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-11-09
Publication Date
2025-07-25
Estimated Expiration
2038-11-09

AI Technical Summary

Technical Problem

In the prior art During the multi-library sequencing process, cross-contamination of library labels leads to wrong allocation of samples, resulting in diagnostic errors. The conventional quality inspection methods are costly and are not sensitive enough, making it difficult to effectively detect cross-contamination at low concentrations.

Method used

The quality control method of unique double-ended library tag combination was adopted. By constructing a gDNA library with unique double-ended library tag combination, the machine was sequenced and two quality control analysis was carried out to ensure that the largest single-sided tag pollution accounted for ≤2.5%, the largest single-sided tag combination pollution accounted for ≤0.01%, the number of label sample sequences in each group was ≥5,000, the mixed proportion of all label combinations accounted for variance coefficient ≤0.5, the comprehensive sequence pass rate ≥97%, the proportion of label sample sequences in each group was ≥0.2/logarithm of library tag combinations, and the proportion of labels with one-sided greater than 1% was ≤10%.

Benefits of technology

It improves the detection efficiency of library labels, reduces the risk of cross-contamination, ensures the accuracy of sample sequence measurement, and reduces the cost of quality inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113957123B_ABST
    Figure CN113957123B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing and detecting a gDNA library containing a unique combination of dual-end library tags, belonging to the technical field of biological detection. The construction method of the gDNA library includes the following steps: (1) Dilute the gDNA standard and fragment the gDNA; (2) Repair the ends of the fragmented gDNA; (3) Connect the two ends of the fragmented gDNA after end repair to prefabricated adapters respectively, and purify the ligation product; (4) Amplify the purified ligation product to construct a library, and purify the amplified library to obtain a gDNA library with a unique combination of library tags. The gDNA library constructed by the present invention contains unique dual-end library tags, which can effectively avoid the misassignment between samples caused by cross-contamination between tag primers and is more suitable for the accurate sequencing requirements of the library.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of a Chinese patent application with an application date of November 9, 2018, an application number of 201811337895.2, and a title of "A Quality Control Method and Application for Detecting Unique Dual-End Library Tag Combinations". Technical Field

[0002] The present invention belongs to the technical field of biological detection, and particularly relates to a method for constructing and detecting a gDNA library containing unique dual-end library tag combinations. Background Art

[0003] With the rapid development of high-throughput technologies, the throughput of sequencers is getting higher and higher. The earlier method of using physical partition methods, such as lane-based flow cells, to distinguish different sequencing libraries is no longer applicable. Multiplex sequencing has been widely used in various fields of next-generation sequencing. The key to multiplex sequencing is the library index. The library index is a special sequence marker for each sample in NGS (Next Generation Sequencing) library preparation, which is a specific sequence used to distinguish DNA from different sources, usually with a length of 4 to 12 bases. During high-throughput sequencing, libraries labeled with different known tag sequences are mixed and then subjected to a sequencing reaction. The inserted fragments and tags of the library are sequentially read out and converted into bases. In the subsequent analysis process, the software classifies the sequencing results using the expected tag sequences and splits the sequencing results into different samples.

[0004] During multiplex sequencing, if there is an incorrect assignment of library sequences, sequences that originally do not belong to a certain library will be misclassified. The occurrence of such incorrect assignments will bring incorrect analysis results for some applications. For example, when libraries from tissue samples of cancer patients and libraries from tissue samples of benign tumor patients are sequenced together, if some sequences of cancer tissue samples are misassigned to benign tumor tissue samples, the test report of the benign tumor patient will show as malignant tumor, resulting in a misdiagnosis.

[0005] There are many reasons that can lead to incorrect assignment of library sequences. Common ones include the following: 1) cross-contamination during library preparation, 2) cross-contamination during the production of tag primers, 3) cross-reactions that occur when multiplex libraries undergo cluster reactions in the flow cell, and 4) optical deviations caused by reasons such as excessive cluster density.

[0006] Library tags primers applicable to next-generation sequencing (NGS) libraries often have a length of 50-70 bases. Generally, purification is required to ensure the purity of the full-length primers. However, purification itself, due to the need for gel extraction or column chromatography, often leads to more cross-contamination. In the case of HPLC (high-performance liquid chromatography), the adsorption of library tags primers by the purification column and repeated use will inevitably bring cross-contamination. Although such contamination can be reduced by performing blank elution or elution with irrelevant samples between the purification of two different library tags primers, cross-contamination still cannot be completely avoided. According to experience, 0.5%-5% of the previous library tags primer will remain in the subsequent library tags primer after two consecutive purifications.

[0007] Due to the high sensitivity brought about by the high throughput of NGS, the quality control of library tags primers requires a very sensitive method to detect possible contamination as low as one in a thousand or even one in ten thousand. In addition, since the sequences between library tags primers are very similar, conventional methods such as qPCR are not suitable for detecting contamination in terms of either sensitivity or specificity. Generally, the conventional method is still to use the NGS platform for quality control, but the conventional method can only detect one target library tags primer per lane at most, which makes the quality control cost prohibitively high.

[0008] Therefore, it is necessary to design a new and unique quality control method for the dual-end library tag combination to improve the detection efficiency. Summary of the Invention

[0009] The object of the present invention is to overcome the deficiencies of the existing technology and provide a quality control method and application for detecting unique dual-end Index combinations, which can improve the detection efficiency of library tags and better meet the requirements of accurate sequencing of libraries.

[0010] To achieve the above object, the technical solution adopted by the present invention is as follows: A quality control method for detecting unique dual-end library tag combinations, which includes the following steps:

[0011] S1) Using library tag standards and gDNA standards as raw materials, construct a gDNA library with unique dual-end library tag combinations, sequence the constructed library on a sequencer, and read the library tag sequences;

[0012] S2) Conduct the first quality control analysis on the library tag sequences. The quality control analysis indicators include the following: the maximum proportion of single-sided tag contamination ≤ 2.5%, the maximum proportion of tag combination contamination ≤ 0.01%, the number of sequences of each group of tag samples ≥ 5000, the coefficient of variance of the mixed proportion of all tag combinations ≤ 0.5, the comprehensive sequence passing rate ≥ 97%, the proportion of each group of tag sample sequences ≥ 0.2 / the logarithm of library tag combinations, and the proportion of tags with contamination greater than 1% on one side should ≤ 10%;

[0013] S3) If the quality control analysis in step S2) shows that the indicators do not meet the requirements, re-synthesize the library tags that do not meet the quality control requirements; according to the method in step S1), using the re-synthesized library tags, the library tags that meet the requirements in the first quality control analysis, and gDNA as raw materials, construct a gDNA library with a unique double-ended library tag combination, re-sequence the constructed library on the machine, and read the library tag sequences;

[0014] S4) Conduct a second quality control analysis on the library tag sequences until all library tags meet the quality control analysis indicators;

[0015] Among the parameters of the quality control analysis, the unique double-ended library tag combination is composed of an upstream library tag and a downstream library tag. The upstream library tags are collectively called IG5, and IG5 includes A and B; the downstream library tags are collectively called IG7, and IG7 includes a and b; the matching and correct unique double-ended library tag combinations are A-a and B-b; the non-matching unique double-ended library tag combinations are A-b and B-a; after each sequencing reaction, the sequence numbers of each of the above combinations can be obtained through analysis;

[0016] The proportion of single-sided tag contamination is the cross-contamination ratio that occurs between tags within a group, and the contamination can only occur within the group, that is, contamination occurs within the IG5 group or / and within the IG7 group;

[0017] When no cross-contamination occurs to a of IG7 during the production process, for A of IG5, the contamination proportion containing B = the number of sequences containing B-a / the total number of sequences containing a,

[0018] When no cross-contamination occurs to A of IG5 during the production process, for a of IG7, the contamination proportion containing b = the number of sequences containing A-b / the total number of sequences containing A;

[0019] When B contaminates A and b contaminates a, the contamination proportion of the B-b tag combination = (the number of sequences containing B-a / the total number of sequences containing a) × (the number of sequences containing A-b / the total number of sequences containing A);

[0020] When no cross-contamination occurs to b of IG7 during the production process, for B of IG5, the contamination proportion containing A = the number of sequences containing A-b / the total number of sequences containing b,

[0021] When no cross-contamination occurs to B of IG5 during the production process, for b of IG7, the contamination proportion containing a = the number of sequences containing B-a / the total number of sequences containing B;

[0022] When A contaminates B and a contaminates b, the contamination proportion of the A-a tag combination = (the number of sequences containing A-b / the total number of sequences containing b) × (the number of sequences containing B-a / the total number of sequences containing B);

[0023] The number of sequences in each group of tag samples is the number of correctly paired sequences in each group after system filtration, that is, the number of sequences containing A-a or the number of sequences containing B-b;

[0024] The variance coefficient of the mixed proportion of all tag combinations is the variance coefficient of the proportion of the number of correctly paired sequences in each group after system filtration in the total number of correctly paired sequences after system filtration;

[0025] The comprehensive sequence passing rate is the proportion of the total number of correctly paired and valid sequences after system filtration after the sequencing reaction in the total number of all sequences after system filtration;

[0026] The proportion of each group of tag samples in the sequences is the proportion of the number of correctly paired sequences in each group after system filtration in the total sequences after system filtration;

[0027] The proportion of tags with more than 1% contamination on one side is: within the upstream library tags, the proportion of library tags with a contamination ratio greater than 1% in the total number of library tags; and, within the downstream library tags, the proportion of library tags with a contamination ratio greater than 1% in the total number of library tags.

[0028] As an improvement of the above technical solution, step S1) sequentially includes the following steps: preparation of gDNA standard, fragmentation of gDNA, end repair, adapter ligation, purification of adapter ligation products, library amplification, purification of amplified library, quality inspection of purified library, detection of fragment size of purified library, and library loading for sequencing.

[0029] As an improvement of the above technical solution, the unique double-ended library tag combination consists of IG5 group and IG7 group. The Hamming distance between library tags within IG5 and IG7 is ≥3, and the sequence Hamming distance between library tags of IG5 and IG7 groups is ≥2.

[0030] As a further improvement of the above technical solution, the library tags are purified by high performance liquid chromatography and the molecular weight is confirmed by mass spectrometry analysis, with the requirement of purity ≥85%.

[0031] As an improvement of the above technical solution, the unique double-ended library tag combination consists of 96 pairs of library tags, that is, there are 96 upstream library tags in the IG5 group and 96 downstream library tags in the IG7 group, corresponding one by one; the proportion of each group of tag samples is then adjusted accordingly to ≥0.2%.

[0032] As an improvement of the above technical solution, the unique dual - end library tag combination consists of 48 pairs of library tags, that is, there are 48 upstream library tags in the IG5 group and 48 downstream library tags in the IG7 group, corresponding one by one; the proportion of each group of tag sample sequences is correspondingly adjusted to ≥0.4%.

[0033] As an improvement of the above technical solution, when the unique dual - end library tag combination consists of 192 pairs of library tags, that is, there are 192 upstream library tags in the IG5 group and 192 downstream library tags in the IG7 group, corresponding one by one, the proportion of each group of tag sample sequences is correspondingly adjusted to ≥0.1%.

[0034] As an improvement of the above technical solution, when the unique dual - end library tag combination consists of 288 pairs of library tags, that is, there are 288 upstream library tags in the IG5 group and 288 downstream library tags in the IG7 group, corresponding one by one, the proportion of each group of tag sample sequences is correspondingly adjusted to ≥0.07%.

[0035] As an improvement of the above technical solution, when the unique dual - end library tag combination consists of 384 pairs of library tags, that is, there are 384 upstream library tags in the IG5 group and 384 downstream library tags in the IG7 group, corresponding one by one, the proportion of each group of tag sample sequences is correspondingly adjusted to ≥0.05%.

[0036] In addition, the present invention also provides the application of the quality control method in sample sequence determination.

[0037] The beneficial effects of the present invention are as follows: The present invention provides a quality control method and application for detecting a unique dual - end library tag combination. This quality control method can efficiently detect cross - contamination of library tags, and has a relatively low cost, and is more suitable for high - throughput determination of sample sequences. Description of the Drawings

[0038] Figure 1 Showing the quality control results of the first simulation in Example 1;

[0039] Figure 2 Showing the quality control results of the second simulation in Example 1;

[0040] Figure 3 Showing the results of the first quality control analysis of the IG5 end in Example 2;

[0041] Figure 4 It is a hot - spot map of the pollution proportion of the first quality control analysis of the IG7 end in Example 2, Figure 4There are 96 pairs of tagged primers. The abscissa from left to right is IG5A01 - IG5A12, IG5B01 - IG5B12, IG5C01 - IG5C12 until IG5H01 - IG5H12 in sequence. The ordinate from top to bottom is IG7A01 - IG7A12, IG7B01 - IG7B12, IG7C01 - IG7C12 until IG7H01 - IG7H12 in sequence. The points circled by the ellipse in the figure represent the tagged primers that do not meet the requirements; the following is similar;

[0042] Figure 5 It is the hotspot map of the pollution proportion for the first quality control analysis of the IG5 end in Example 2;

[0043] Figure 6 It is the distribution map of the pollution proportion for the first quality control analysis of the IG7 and IG5 ends in Example 2;

[0044] Figure 7 It shows the stability comparison results of the two quality control analyses of the IG7 end in Example 2;

[0045] Figure 8 It shows the stability comparison results of the two quality control analyses of the IG5 end in Example 2;

[0046] Figure 9 It shows the results of the first quality control analysis of the IG5 end in Example 3;

[0047] Figure 10 It is the hotspot map of the pollution proportion for the first quality control analysis of the IG7 end in Example 3;

[0048] Figure 11 It is the hotspot map of the pollution proportion for the first quality control analysis of the IG5 end in Example 3;

[0049] Figure 12 It is the distribution map of the pollution proportion for the first quality control analysis of the IG7 and IG5 ends in Example 3. Detailed implementation manners

[0050] To better illustrate the purpose, technical solution and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments and drawings.

[0051] In addition, it should be noted that in the specification of the present invention, Index, library tag and tagged primer mean the same thing; in the calculation of the proportion of each group of tag sample sequences, the result of 0.2 / the number of library tag combinations is rounded to retain one non - zero digit.

[0052] Principle of unique dual - end library tags in preventing sample contamination caused by cross - contamination

[0053] In the field of NGS, in order to distinguish different samples in the same sequencing reaction, specific "tags" (Index) are added to different samples during the library construction process, so that the data of different samples can be separated during subsequent data analysis. With the continuous improvement of the throughput of sequencers, more samples are pooled into the same flow cell (Lane) for sequencing, posing higher requirements for the quantity and discrimination of Indexes. In addition, Illumina HiSeqX / 4000 and NovaSeq adopt a clustering method different from other Illumina sequencers, and the literature reports that they have a higher risk of Index cross-contamination. Traditional single-end Index primers only split data based on one end, and it is easy to misclassify data when contamination occurs. Using unique dual-end Index primers can maximize the risk of sample contamination caused by Index cross-contamination and ensure the stability and reliability of the product. Since unique dual-end Index primers rely on unique dual-end paired Indexes to split data, it adds a "double insurance" to the sequencing sequences, and most of the contaminated sequences will be discarded. Table 1 compares the tolerances of single-end, combined dual-end, and unique dual-end Index strategies to Index cross-contamination.

[0054] Table 1

[0055]

[0056]

[0057] Principle of high - throughput contamination quality control of unique dual - end Index primers by NGS method

[0058] Due to the use of unique dual-end Indexes, each sample is labeled twice by Indexes, which greatly increases the tolerance to cross-contamination between primers with single-end labeling. For example, if the unilateral contamination ratio of 2 pairs of Indexes is 1% each, the actual risk of sample misclassification and contamination is 1% × 1% = 0.01%. This tolerance also greatly reduces the pressure on the synthesis and purification of Index primers, further controlling the manufacturing cost.

[0059] Taking advantage of the unique dual-end Index, the present invention provides a simple and feasible quality control method for detecting cross-contamination of labeling primers using NGS. Its basic principle is based on observing the proportion of unexpected dual-end Index combinations in the entire sequencing results to estimate the maximum possible cross-contamination probability and the Indexes involved, thereby avoiding misallocation between samples caused by cross-contamination between Index primers.

[0060] For example, four libraries are respectively labeled as A+a, B+b, C+c, and D+d. Therefore, when performing sequence analysis, only the above 4 combinations are considered legal combinations. Taking the combination A+b as an example, since theoretically only A pairs with a, if the combination A+b is observed, there are two possibilities: 1) The tag primer b enters primer a. Here, define S as the number of sequences containing this kind of Index, and the estimated contamination ratio is S (A+b) / S A ; 2) Primer A enters primer B, and the estimated contamination ratio is S (A+b) / S b . It should be noted that the premise of this calculation method is that there are no non-homogeneous Indexes such as a / b / c / d within the same category of Indexes such as A / B / C / D. In addition, the estimation model only considers a simple one-to-one contamination mode, rather than complex situations such as multiple contaminations. In addition, this calculation method only estimates the maximum contamination possibility and has no ability to judge the directionality of contamination. In fact, after any single-directional contamination event occurs, such as the event of "A entering B", it will be detected as two possibilities: "A entering B" or "b entering a". According to this calculation model, we can estimate the maximum combined contamination risk of the unique combination of dual-index library A+a by other primers in multiple sequencing as:

[0061]

[0062] However, since the combinations we expect are only the four combinations of A+a, B+b, C+c, and D+d, the actual effective maximum contamination risk can be calculated as:

[0063]

[0064] In a practical application example, we perform PCR operations on 48 pairs or 96 pairs of Index primers respectively to label Indexes to the libraries, and then mix them together for conventional MiSeq sequencing. After sequencing, directly call the analysis script to analyze 96×6 = 9216 sequence combinations, find the proportion of abnormal combinations and calculate their respective contamination proportions.

[0065] Index on - machine sequencing

[0066] 1. Preparation of gDNA standard

[0067] 1) 500 ng of gDNA standard is required for quality inspection of 48plex Index, and 1000 ng of gDNA standard is required for quality inspection of 96plex Index;

[0068] 2) Take 50 μl of 1×IDTE Buffer and add it to a new 1.5 ml Eppendorf LoBind tube. Then add the corresponding volume of gDNA standard to the tube: for 48plex Index plate detection, the added volume of gDNA standard is 2 μl; for 96plex Index plate detection, the added volume of gDNA standard is 4 μl. Then vortex for 10 - 15 s and briefly centrifuge to return the solution to the bottom of the tube.

[0069] 3) Transfer the standard dilution to a Covaris MicroTΜBE tube, supplement 1×IDTE Buffer to 50 μl, and then perform subsequent DNA fragmentation operations.

[0070] 2. gDNA Fragmentation

[0071] Use a Covaris M220 instrument to fragment the DNA into fragments of 170 - 200 bp. After fragmentation is completed, take out the Covaris MicroTube tube and centrifuge to return the liquid to the bottom of the tube.

[0072] 3. End Repair and Adding A at the 3' End

[0073] 1) Reagent preparation: Open the KAPA Hyper Prep 96 reaction Kit and take out the following 2 tubes and melt them on ice.

[0074] 2) Prepare the end repair and A - adding reaction system mixture on ice in a new 1.5 ml Eppendorf LoBind tube, flick it gently 3 - 5 times with your finger, invert it 2 - 3 times to mix well, and centrifuge in a centrifuge for 1 - 3 s; The reaction system configuration is shown in Table 2.

[0075] 3) Pipette 60 μl of the mixture and dispense it into 4 (for 48plex Index plate) or 8 (for 96plex Index plate) 0.2 ml flat - cap PCR tubes, and briefly centrifuge in a centrifuge for 1 - 3 s.

[0076] 4) Put it into a PCR instrument and perform the following operations: 85°C hot lid, 20°C for 30 min, 65°C for 30 min, store at 4°C, and proceed to the next step within 2 h.

[0077] Table 2

[0078]

[0079] 4. Adapter Ligation, ligate the two ends of the double - stranded DNA fragment with added A to the pre - fabricated adapter (containing T sticky ends)

[0080] 1) In a new 1.5 ml Eppendorf LoBind tube, prepare the adapter ligation reaction system mixing solution on ice, flick it gently with your finger 3 - 5 times, invert it up and down 2 - 3 times, and centrifuge it in a centrifuge for 1 - 3 seconds; the configuration of the reaction system is shown in Table 3;

[0081] 2) Pipette 50 μl of the mixing solution into the above 0.2 ml tube (4 tubes for the 48plex Index plate, 8 tubes for the 96plex Index plate), pipette up and down 5 times to mix evenly, and centrifuge for 1 - 3 s;

[0082] 3) Run the following program on a PCR instrument: 20 °C for 15 min, 70 °C for 10 min, and store at 4 °C (85 °C hot lid).

[0083] Table 3

[0084]

[0085] 5. Purification of the ligation product to remove adapter dimers and other components such as unligated adapters

[0086] 1) Invert it up and down 2 - 3 times, vortex and mix the SPB magnetic beads for 5 - 10 s to return to room temperature to make them homogeneous; take a 1.5 ml centrifuge tube, and successively add the homogenized magnetic beads and the adapter - added product according to the ratio of the ligation reaction system to the magnetic bead volume of 1:0.8; the specific strategy is as follows: for magnetic beads of 352 μl and adapter product of 440 μl, 4 tubes are combined into 1 tube for purification, a total of 1 tube; for magnetic beads of 2×352 μl and adapter product of 2×440 μl, 4 tubes are combined into 1 tube for purification, a total of 2 tubes (96plex Index); after adding, vortex and mix, rotate and incubate for 5 min, and centrifuge briefly;

[0087] 2) Place the centrifuge tube on a magnetic rack and wait for the solution to clarify; keep the centrifuge tube on the magnetic rack without moving, open the tube cap, and carefully aspirate the clear supernatant, avoiding touching the magnetic beads;

[0088] 3) Keep the tube on the magnetic rack, add 500 μL of freshly prepared 75% ethanol to each tube, wait for 1 min to allow the magnetic beads to precipitate fully, during which slowly rotate the centrifuge tube 1 circle horizontally, and aspirate the ethanol; repeat this step once;

[0089] 4) Centrifuge for 1 - 3 s, place the centrifuge tube back on the magnetic rack and let it stand for 30 s, use a pipette to remove the residual ethanol completely, keep the tube cap open; dry the magnetic beads at room temperature for 3 min, add 500 μl of EB solution to each tube, blow and mix well, incubate at room temperature for 2 min; place the centrifuge tube on the magnetic rack for 2 min until the solution clarifies, use a pipette to transfer 490 μl of the supernatant to a new Eppendorf LoBind 1.5 ml centrifuge tube (for the 96plex Index plate, the two tubes are combined into 1 tube after elution), and keep it on ice for standby.

[0090] 6. Library amplification: Amplify the library that has already been ligated with adapters.

[0091] 1) Prepare the corresponding volume of reaction system mixture (prepared on ice) in a 5 ml Eppendorf LoBind tube (or 15 ml centrifuge tube), flick it gently with your finger 3 - 5 times, invert it up and down 2 - 3 times, and let it stand vertically for 0.5 - 1 min; The configuration of the reaction system is shown in Table 4.

[0092] 2) Evenly distribute the prepared reaction system mixture into 8 - tube strips, with each aliquot volume being 138 μl (for 96Index pair Plate (refer part2#) detection, two equal distributions are required: 142 μl + 132 μl).

[0093] 3) Dispense the reaction system mixture into a new 48 - well plate (48plex Index) or 96 - well PCR plate (96plexIndex), with the dispensing volume being 22.5 μl / well.

[0094] 4) Take out 2.5 μl of Index from the IDP plate, add it to the above - dispensed reaction system mixture in the 48 - well plate or 96 - well PCR plate, pipette and mix it thoroughly 2 - 3 times, and seal the membrane; Centrifuge it at 1000 rpm for 1 min using a plate shaker (reaction volume 25 μl); Place it on a PCR instrument and run it, and the running program is shown in Table 5.

[0095] Table 4

[0096]

[0097] Table 5

[0098]

[0099]

[0100] 7. Purification of the amplified library: Remove primer dimers and the reaction system.

[0101] 1) Invert the SPB beads 2 - 3 times, mix them at the maximum speed of the VORTEX for 5 - 10 s to make them homogeneous.

[0102] 2) Pipette the corresponding amount of SPB beads into the sample loading trough, adding 20 μl of SPB beads to each sample (sample:bead = 1:0.8): For 48 samples, add about 1440 μl of beads to the sample loading trough; for 96 samples, add about 2880 μL of beads to the sample loading trough.

[0103] 3) Remove the 48-well plate from the PCR instrument, centrifuge at 1000 rpm for 3 s, and carefully tear off the sticker; Pipette 20 μl of SPB magnetic beads from the loading slot into the 48-well plate / 96-well PCR plate, and pipette up and down 10 times.

[0104] 4) Cover the 48-well plate / 96-well PCR plate with a sticker, centrifuge briefly at 1000 rpm for 3 s, and place at room temperature for 5 min; Place the 48-well plate / 96-well PCR plate on a 96-well magnetic stand until the solution is clear; Discard the sticker, pipette 45 μl of the supernatant, and discard it.

[0105] 5) Keep the 48-well plate / 96-well PCR plate on the magnetic stand, add 200 μl of freshly prepared 75% ethanol to the sample wells; Let the 48-well plate / 96-well PCR plate stand on the magnetic stand for 1 min to fully wash the magnetic beads, then discard the ethanol; Repeat this step once.

[0106] 6) Let the 48-well plate / 96-well PCR plate stand on the magnetic stand for 30 s and remove all residual ethanol; Remove the 48-well plate / 96-well PCR plate from the magnetic stand and place it on a PCR plate rack at room temperature for 2 min to dry the magnetic beads; Add 14 μl of EB to the 48-well plate / 96-well PCR plate, cover with an eight-tube cap, vortex for about 5 s, and centrifuge briefly at 1000 rpm for 3 s.

[0107] 7) Incubate the 48-well plate at room temperature for 2 min, discard the sticker, place the 48-well plate on the magnetic stand for 2 min until the solution is clear; Transfer 8 μL of the supernatant to a new 48-well plate / 96-well PCR plate, without aspirating the magnetic beads.

[0108] 8) Transfer each column of the library to the same new 0.2 ml eight-tube, then transfer the library in the 0.2 ml eight-tube to the same new 1.5 ml Eppendorf LoBind tube to combine into a pooling library, vortex and centrifuge; Take out 20 μl of the purified and mixed library and transfer it to a new 1.5 ml Eppendorf LoBind tube, then add 180 μl of EB, pipette up and down 5 - 6 times to pre-dilute the library 10-fold for subsequent detection.

[0109] 8. Quality control of the purified library

[0110] Use dsDNA HS (High Sensitivity) Assay Kit (Thermo Fisher) to measure the concentration of the diluted library and convert it back to the pre-library concentration; If the library concentration is between 9 - 60 ng / μl and the Labchip result is normal, the library construction part is qualified and can proceed with subsequent Miseq loading; If the requirements cannot be met, the library preparation needs to be repeated.

[0111] 9. Purified Library Fragment Size Detection (Library QC)

[0112] The diluted library was detected using the The LabChip DNA High Sensitivity Reagent kit (Perkin Elmer); for a qualified library, the main peak of the library fragments was at 350 - 500 bp, and there were no obvious small fragments in the 10 - 150 bp range.

[0113] 10. Library Loading Strategy (Miseq Run)

[0114] 1) Dilute the purified library to 4 nM according to the detected concentration by QC, and dilute 1N NaOH to 0.2N using nuclease - free water;

[0115] 2) Library denaturation: Take 5 μl of the library diluted to 4 nM and add it to a new 1.5 ml Eppendorf LoBind tube, then add 5 μl of 0.2N NaOH, pipette and mix well 15 - 20 times, and incubate at room temperature for 5 min;

[0116] 3) Dilute the library to 13 pM;

[0117] 4) For subsequent operations, refer to the Illumina Miseq operation guide and sequence the library using the corresponding settings of Read1 = 12 cycles, Index1 = 8 cycles, and Index2 = 8 cycles.

[0118] 11. Sequencing Data Analysis (QC Analysis)

[0119] Use the Illumina bcl2fastq software with corresponding parameters to output all sequences of index1 and index2 (Fastq format), and use the corresponding script to perform statistical analysis on the sequences to obtain various indicators.

[0120] 11. Criteria for Judging Library Sequencing Results

[0121] Miseq output indicators: Sequencing data quality 01: Q30 > 90%, Sequencing data quality 02: PF > 97%, Sequencing data quality 03: Phasing and Prephasing are both less than 0.30.

[0122] Example 1 Simulation of quality control method

[0123] 1) First simulation of unidirectional contamination: One cross-contamination was detected for the first time, and two speculated contamination directions were given. The maximum contamination ratio (i.e., the maximum proportion of unilateral tag contamination) was 4%. The simulation data generated 96 pairs of standard paired sequences. Among them, there were 48,000 contaminated normal pairs of IG7F01 + IG5F01, and 2,000 of IG7F01 + IG5E01; the remaining normal pairs were 50,000 each. The simulated library data was analyzed, and the actual test results are shown in Table 6. According to the parameters in Table 6, quality control analysis was carried out, and the quality control analysis results are as Figure 1 shown; among them, the product of the maximum paired contamination ratio (i.e., the maximum proportion of tag combination contamination) = 4% × 0 = 0, the number of correct paired sequences was 48,000, and the number of correctly paired and valid sequences was 48,000. The sequence passing rate was 100%. The number of tags with a contamination greater than 1% on one side was 1 type, and the contamination index ratio greater than 1% (i.e., the proportion of tags with a contamination greater than 1% on one side) = 1 / 96 = 1.04%.

[0124] Table 6

[0125]

[0126] 2) Second simulation of bidirectional contamination that can cause sample misclassification: Two cross-contaminations were detected for the second time, and these two cross-contaminations could cause sample misclassification. The maximum contamination ratio was 2%, and the maximum paired contamination product was 0.04%. The simulated data generated standard paired sequences of 50,000 each. Among them, there were 48,000 contaminated normal pairs of IG7F01 + IG5F01, 1,000 mispaired pairs of IG7F01 + IG5E01, and 1,000 of IG7E01 + IG5F01. The simulated library samples were analyzed, and the actual test results are shown in Table 7. According to the parameters in Table 7, quality control analysis was carried out, and the quality control analysis results are as Figure 2 shown.

[0127] Table 7

[0128]

[0129]

[0130] It can be seen that the test results of this simulation test are consistent with the expectations.

[0131] Example 2

[0132] In this example, 96 pairs of library tags were used for quality inspection. The results of the first quality control analysis are shown in Tables 8 and 9. Tables 8 and 9 only list the contamination situations.

[0133] Table 8 Sequencing results for the Index at the IG7 end

[0134] Query Expected combined object Expected combination Unexpected combination Unexpected combination object Total number of sequences Number of sequences in unexpected combination Pollution source Contaminated Pollution proportion IG7A01 IG5A01 IG7A01 - IG5A01 IG7A01 - IG5B01 IG5B01 96 45 IG5B01 IG5A01 46.88% IG7A01 IG5A01 IG7A01 - IG5A01 IG7A01 - IG5A02 IG5A02 96 51 IG5A02 IG5A01 53.13% IG7A08 IG5A08 IG7A08 - IG5A08 IG7A08 - IG5H07 IG5H07 53249 88 IG5H07 IG5A08 0.17% IG7B02 IG5B02 IG7B02 - IG5B02 IG7B02 - IG5A03 IG5A03 40825 43 IG5A03 IG5B02 0.11% IG7B10 IG5B10 IG7B10 - IG5B10 IG7B10 - IG5D08 IG5D08 46021 70 IG5D08 IG5B10 0.15% IG7B11 IG5B11 IG7B11 - IG5B11 IG7B11 - IG5C11 IG5C11 47969 68 IG5C11 IG5B11 0.14% IG7C01 IG5C01 IG7C01 - IG5C01 IG7C01 - IG5G12 IG5G12 39518 64 IG5G12 IG5C01 0.16% IG7C06 IG5C06 IG7C06 - IG5C06 IG7C06 - IG5C07 IG5C07 60810 637 IG5C07 IG5C06 1.05% IG7C08 IG5C08 IG7C08 - IG5C08 IG7C08 - IG5B08 IG5B08 67961 119 IG5B08 IG5C08 0.18% IG7D03 IG5D03 IG7D03 - IG5D03 IG7D03 - IG5E03 IG5E03 44222 48 IG5E03 IG5D03 0.11% IG7D03 IG5D03 IG7D03 - IG5D03 IG7D03 - IG5C03 IG5C03 44222 56 IG5C03 IG5D03 0.13% IG7D04 IG5D04 IG7D04 - IG5D04 IG7D04 - IG5D03 IG5D03 40521 41 IG5D03 IG5D04 0.10% IG7D07 IG5D07 IG7D07 - IG5D07 IG7D07 - IG5E08 IG5E08 39029 281 IG5E08 IG5D07 0.72% IG7D08 IG5D08 IG7D08 - IG5D08 IG7D08 - IG5C08 IG5C08 53581 85 IG5C08 IG5D08 0.16% IG7D09 IG5D09 IG7D09 - IG5D09 IG7D09 - IG5E09 IG5E09 54786 70 IG5E09 IG5D09 0.13% IG7E03 IG5E03 IG7E03 - IG5E03 IG7E03 - IG5F03 IG5F03 60714 78 IG5F03 IG5E03 0.13% IG7E07 IG5E07 IG7E07 - IG5E07 IG7E07 - IG5D07 IG5D07 57285 88 IG5D07 IG5E07 0.15% IG7F04 IG5F04 IG7F04 - IG5F04 IG7F04 - IE5D04* IE5D04* 49814 54 IE5D04* IG5F04 0.11% IG7F07 IG5F07 IG7F07 - IG5F07 IG7F07 - IG5E07 IG5E07 55273 63 IG5E07 IG5F07 0.11% IG7G08 IG5G08 IG7G08 - IG5G08 IG7G08 - IG5F08 IG5F08 43769 167 IG5F08 IG5G08 0.38% IG7G10 IG5G10 IG7G10 - IG5G10 IG7G10 - IG5F06 IG5F06 57227 60 IG5F06 IG5G10 0.10% IG7H02 IG5H02 IG7H02 - IG5H02 IG7H02 - IG5H03 IG5H03 38360 58 IG5H03 IG5H02 0.15% IG7H07 IG5H07 IG7H07 - IG5H07 IG7H07 - IG5G07 IG5G07 36388 42 IG5G07 IG5H07 0.12%

[0135] Table 9 Sequencing results for the Index at the IG5 end

[0136] Query Expected combined object Expected combination Unexpected combination Unexpected combined object Total number of sequences Number of sequences of unexpected combinations Pollution source Contaminated Pollution ratio IG5A01 IG7A01 IG5A01 - IG7A01 IG5A01 - IG7B01 IG7B01 26 26 IG7B01 IG7A01 100.00% IG5A02 IG7A02 IG5A02 - IG7A02 IG5A02 - IG7A01 IG7A01 49928 51 IG7A01 IG7A02 0.10% IG5A03 IG7A03 IG5A03 - IG7A03 IG5A03 - IG7B02 IG7B02 33067 43 IG7B02 IG7A03 0.13% IG5A08 IG7A08 IG5A08 - IG7A08 IG5A08 - IG7B08 IG7B08 53201 60 IG7B08 IG7A08 0.11% IG5B08 IG7B08 IG5B08 - IG7B08 IG5B08 - IG7C08 IG7C08 61974 119 IG7C08 IG7B08 0.19% IG5B11 IG7B11 IG5B11 - IG7B11 IG5B11 - IG7A11 IG7A11 47967 51 IG7A11 IG7B11 0.11% IG5C03 IG7C03 IG5C03 - IG7C03 IG5C03 - IG7D03 IG7D03 49273 56 IG7D03 IG7C03 0.11% IG5C07 IG7C07 IG5C07 - IG7C07 IG5C07 - IG7C06 IG7C06 45027 637 IG7C06 IG7C07 1.41% IG5C08 IG7C08 IG5C08 - IG7C08 IG5C08 - IG7D08 IG7D08 67868 85 IG7D08 IG7C08 0.13% IG5C11 IG7C11 IG5C11 - IG7C11 IG5C11 - IG7B11 IG7B11 57807 68 IG7B11 IG7C11 0.12% IG5D07 IG7D07 IG5D07 - IG7D07 IG5D07 - IG7E07 IG7E07 38866 88 IG7E07 IG7D07 0.23% IG5D07 IG7D07 IG5D07 - IG7D07 IG5D07 - IG7C08 IG7C08 38866 49 IG7C08 IG7D07 0.13% IG5D08 IG7D08 IG5D08 - IG7D08 IG5D08 - IG7B10 IG7B10 53619 70 IG7B10 IG7D08 0.13% IG5D08 IG7D08 IG5D08 - IG7D08 IG5D08 - IG7E08 IG7E08 53619 65 IG7E08 IG7D08 0.12% IG5E07 IG7E07 IG5E07 - IG7E07 IG5E07 - IG7F07 IG7F07 57203 63 IG7F07 IG7E07 0.11% IG5E08 IG7E08 IG5E08 - IG7E08 IG5E08 - IG7D07 IG7D07 72767 281 IG7D07 IG7E08 0.39% IG5E09 IG7E09 IG5E09 - IG7E09 IG5E09 - IG7D09 IG7D09 58757 70 IG7D09 IG7E09 0.12% IG5F03 IG7F03 IG5F03 - IG7F03 IG5F03 - IG7E03 IG7E03 54811 78 IG7E03 IG7F03 0.14% IG5F06 IG7F06 IG5F06 - IG7F06 IG5F06 - IG7G10 IG7G10 50348 60 IG7G10 IG7F06 0.12% IG5F08 IG7F08 IG5F08 - IG7F08 IG5F08 - IG7G08 IG7G08 67091 167 IG7G08 IG7F08 0.25% IG5G12 IG7G12 IG5G12 - IG7G12 IG5G12 - IG7C01 IG7C01 40234 64 IG7C01 IG7G12 0.16% IG5H03 IG7H03 IG5H03 - IG7H03 IG5H03 - IG7H02 IG7H02 48832 58 IG7H02 IG7H03 0.12% IG5H06 IG7H06 IG5H06 - IG7H06 IG5H06 - IG7A11 IG7A11 42784 62 IG7A11 IG7H06 0.14% IG5H07 IG7H07 IG5H07 - IG7H07 IG5H07 - IG7A08 IG7A08 36410 88 IG7A08 IG7H07 0.24% IG5H11 IG7H11 IG5H11 - IG7H11 IG5H11 - IG7E11 IG7E11 32519 50 IG7E11 IG7H11 0.15%

[0137] Statistical analysis of the results of the first quality inspection of the data in Table 8 and Table 9 can obtain relevant information on IG7A01 - IG5A01, as shown in Table 10 and Figure 3 shown; in addition, statistical analysis of 96 pairs of tagged primers can obtain the hotspot map of the contamination ratio of IG7 and IG5 ( Figure 4 and Figure 5 ), as well as the distribution map of the contamination ratio of IG7 and IG ( Figure 6 ); the conclusions drawn from the summary are as follows: 1) The sequences measured for the combination of IG7A01 - IG5A01 are extremely few. There are only 96 sequences containing IG7A01 and only 26 sequences containing IG5A01, far lower than the requirement of at least 5000 sequences and a proportion > 0.2% for quality inspection; 2) Due to the extremely few sequences in the above combination and the only measured combination being an illegal combination, the contamination ratio is very high; 3) Generally speaking, the well corresponding to IG7A01 - IG5A01 is problematic and needs to be replaced in terms of both the number of valid sequences and the possibility of being contaminated.

[0138] Table 10

[0139]

[0140] Since the well corresponding to IG7A01 - IG5A01 is problematic, the two tagged primers IG7A01 and IG5A01 are separately synthesized again. The re - synthesized primers are dissolved to the specified concentration and placed in the corresponding wells of a new deep - well plate according to the ratio. Excluding the well corresponding to the original IG7A01 - IG5A01, all the remaining liquid in the original master plate with failed quality inspection is transferred to the corresponding positions in a new deep - well plate, and a new molecular plate is used for contamination quality control detection; the results of the second quality control analysis are shown in Table 11 and Table 12, and only the contamination situations are listed in Table 11 and Table 12.

[0141] Table 11 Sequencing results for the Index at the IG7 end

[0142]

[0143]

[0144] Table 12 Sequencing results for the Index at the IG5 end

[0145] Query Expected combined object Expected combination Unexpected combination Unexpected combined object Total number of sequences Number of sequences of unexpected combinations Pollution source Contaminated Pollution ratio IG5A02 IG7A02 IG5A02 - IG7A02 IG5A02 - IG7A01 IG7A01 329005 349 IG7A01 IG7A02 0.11% IG5A03 IG7A03 IG5A03 - IG7A03 IG5A03 - IG7B02 IG7B02 204279 358 IG7B02 IG7A03 0.18% IG5A08 IG7A08 IG5A08 - IG7A08 IG5A08 - IG7B08 IG7B08 244880 246 IG7B08 IG7A08 0.10% IG5B01 IG7B01 IG5B01 - IG7B01 IG5B01 - IG7A01 IG7A01 405468 579 IG7A01 IG7B01 0.14% IG5B08 IG7B08 IG5B08 - IG7B08 IG5B08 - IG7C08 IG7C08 291485 460 IG7C08 IG7B08 0.16% IG5B11 IG7B11 IG5B11 - IG7B11 IG5B11 - IG7A11 IG7A11 336412 580 IG7A11 IG7B11 0.17% IG5C01 IG7C01 IG5C01 - IG7C01 IG5C01 - IG7D01 IG7D01 285446 313 IG7D01 IG7C01 0.11% IG5C05 IG7C05 IG5C05 - IG7C05 IG5C05 - IG7B05 IG7B05 342900 393 IG7B05 IG7C05 0.11% IG5C07 IG7C07 IG5C07 - IG7C07 IG5C07 - IG7C06 IG7C06 252462 4253 IG7C06 IG7C07 1.68% IG5C07 IG7C07 IG5C07 - IG7C07 IG5C07 - IG7D07 IG7D07 252462 255 IG7D07 IG7C07 0.10% IG5C08 IG7C08 IG5C08 - IG7C08 IG5C08 - IG7D08 IG7D08 306576 448 IG7D08 IG7C08 0.15% IG5C11 IG7C11 IG5C11 - IG7C11 IG5C11 - IG7B11 IG7B11 343767 664 IG7B11 IG7C11 0.19% IG5D03 IG7D03 IG5D03 - IG7D03 IG5D03 - IG7D04 IG7D04 230782 336 IG7D04 IG7D03 0.15% IG5D07 IG7D07 IG5D07 - IG7D07 IG5D07 - IG7E07 IG7E07 242178 468 IG7E07 IG7D07 0.19% IG5D08 IG7D08 IG5D08 - IG7D08 IG5D08 - IG7E08 IG7E08 316707 328 IG7E08 IG7D08 0.10% IG5D08 IG7D08 IG5D08 - IG7D08 IG5D08 - IG7B10 IG7B10 316707 435 IG7B10 IG7D08 0.14% IG5D11 IG7D11 IG5D11 - IG7D11 IG5D11 - IG7C11 IG7C11 392842 744 IG7C11 IG7D11 0.19% IG5E08 IG7E08 IG5E08 - IG7E08 IG5E08 - IG7D07 IG7D07 313472 1585 IG7D07 IG7E08 0.51% IG5E09 IG7E09 IG5E09 - IG7E09 IG5E09 - IG7D09 IG7D09 436613 672 IG7D09 IG7E09 0.15% IG5F01 IG7F01 IG5F01 - IG7F01 IG5F01 - IG7G01 IG7G01 314675 362 IG7G01 IG7F01 0.12% IG5F06 IG7F06 IG5F06 - IG7F06 IG5F06 - IG7G10 IG7G10 252235 533 IG7G10 IG7F06 0.21% IG5F08 IG7F08 IG5F08 - IG7F08 IG5F08 - IG7G08 IG7G08 359668 733 IG7G08 IG7F08 0.20% IG5F10 IG7F10 IG5F10 - IG7F10 IG5F10 - IG7G10 IG7G10 289803 510 IG7G10 IG7F10 0.18% IG5G07 IG7G07 IG5G07 - IG7G07 IG5G07 - IG7H07 IG7H07 345835 355 IG7H07 IG7G07 0.10% IG5H03 IG7H03 IG5H03 - IG7H03 IG5H03 - IG7H02 IG7H02 286014 294 IG7H02 IG7H03 0.10% IG5H07 IG7H07 IG5H07 - IG7H07 IG5H07 - IG7A08 IG7A08 240240 300 IG7A08 IG7H07 0.12% IG5H08 IG7H08 IG5H08 - IG7H08 IG5H08 - IG7G08 IG7G08 121432 134 IG7G08 IG7H08 0.11% IG5H09 IG7H09 IG5H09 - IG7H09 IG5H09 - IG7G09 IG7G09 250324 317 IG7G09 IG7H09 0.13%

[0146] After the primer replacement operation, there is no cross - contamination between IG7A01 and IG5A01, and all the indicators of 96 pairs of tagged primers meet the quality inspection standards.

[0147] In addition, in this embodiment, the first quality control analysis and the second quality control analysis are also compared, and the results are shown in Table 13, Figure 7 (Comparative analysis of IG7 tagged primers) and Figure 8 (Comparative analysis of IG5 tagged primers) as shown. It can be seen that the reproducibility of the two quality control analyses is good, indicating that the quality control method of the present invention has good stability.

[0148] Table 13

[0149]

[0150]

[0151] Example 3

[0152] In this embodiment, 96 pairs of library tags are used for quality inspection. By statistically analyzing the results of the first quality inspection, relevant information of one pair of tagged primers can be obtained, as Figure 9 shown; in addition, by statistically analyzing 96 pairs of tagged primers, a hot - spot map of the contamination ratio of IG7 and IG5 can be obtained ( Figure 10 and Figure 11 ), and a distribution map of the contamination ratio of IG7 and IG ( Figure 12 ); the conclusion drawn by summarization is that after the first quality control analysis of this pair of tagged primers, all 96 pairs of tagged primers meet the indicators.

[0153] Finally, it should be noted that the above embodiments are used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for detecting library tag contamination in a gDNA library containing a unique combination of dual - ended library tags, characterized in that, Construct a gDNA library with a unique combination of dual - end library tags, subject the constructed gDNA library to on - machine sequencing, and read the library tag sequences. When the sequencing results do not meet the following conditions, the library is considered contaminated: The maximum proportion of single - side tag contamination ≤ 2.5%, the maximum proportion of tag - combination contamination ≤ 0.01%, the number of sequences in each group of tag samples ≥ 5000, the coefficient of variance of the mixed proportion of all tag combinations ≤ 0.5, the comprehensive sequence passing rate ≥ 97%, the proportion of sequences in each group of tag samples ≥ 0.2 / the logarithm of library tag combinations, and the proportion of tags with more than 1% contamination on one side should ≤ 10%; The proportion of single - side tag contamination is the cross - contamination proportion that occurs between tags within a group, and the contamination only occurs within the group, that is, contamination occurs within the IG5 group or / and within the IG7 group; When there is no cross - contamination of a in IG7 during the production process, for A in IG5, the proportion of contamination containing B = the number of sequences containing B - a / the total number of sequences containing a; When there is no cross - contamination of A in IG5 during the production process, for a in IG7, the proportion of contamination containing b = the number of sequences containing A - b / the total number of sequences containing A; When B contaminates A and b contaminates a, the proportion of B - b tag - combination contamination = (the number of sequences containing B - a / the total number of sequences containing a) × (the number of sequences containing A - b / the total number of sequences containing A); When there is no cross - contamination of b in IG7 during the production process, for B in IG5, the proportion of contamination containing A = the number of sequences containing A - b / the total number of sequences containing b; When there is no cross - contamination of B in IG5 during the production process, for b in IG7, the proportion of contamination containing a = the number of sequences containing B - a / the total number of sequences containing B; When A contaminates B and a contaminates b, the proportion of A - a tag - combination contamination = (the number of sequences containing A - b / the total number of sequences containing b) × (the number of sequences containing B - a / the total number of sequences containing B); The number of sequences in each group of tag samples is the number of correctly paired sequences in each group after system filtering, that is, the number of sequences containing A - a or the number of sequences containing B - b; The coefficient of variance of the mixed proportion of all tag combinations is the coefficient of variance of the proportion of the number of correctly paired sequences in each group after system filtering in the total number of correctly paired sequences after system filtering; The comprehensive sequence passing rate is the proportion of the total number of correctly paired and valid sequences after system filtering in the total number of all sequences after system filtering after the sequencing reaction; The proportion of sequences in each group of tag samples is the proportion of the number of correctly paired sequences in each group after system filtering in the total sequences after system filtering; The proportion of tags with more than 1% contamination on one side is: within the upstream library tags, the proportion of library tags with a contamination proportion greater than 1% in the total number of library tags; and within the downstream library tags, the proportion of library tags with a contamination proportion greater than 1% in the total number of library tags; The unique dual - end library tag combinations are all composed of an upstream library tag and a downstream library tag. The upstream library tags are collectively called IG5, and IG5 contains A and B; the downstream library tags are collectively called IG7, and IG7 contains a and b. The matching and correct unique dual - end library tag combinations are A - a and B - b; the non - matching unique dual - end library tag combinations are A - b and B - a. After each sequencing reaction, the sequence counts of each of the above combinations are obtained through analysis.

2. The method according to claim 1, wherein The method for constructing the gDNA library successively includes the following steps: Step (1) Preparation of gDNA standard. Step (2) Fragmentation of gDNA. Step (3) End repair. Step (4) Adapter ligation. Step (5) Purification of adapter - ligated products. Step (6) Library amplification. Step (7) Purification of amplified library. Step (8) Quality inspection of purified library. Step (9) Detection of the fragment size of purified library. Among them, the Hamming distance between library tags within each of IG5 and IG7 is ≥3, and the sequence Hamming distance between library tags of IG5 and IG7 groups is ≥2. When the unique dual - end library tag combination consists of 96 pairs of library tags, that is, there are 96 upstream library tags in the IG5 group and 96 downstream library tags in the IG7 group, corresponding one by one; the proportion of each group of tag sample sequences is adjusted accordingly to ≥0.2%. When the unique dual - end library tag combination consists of 48 pairs of library tags, that is, there are 48 upstream library tags in the IG5 group and 48 downstream library tags in the IG7 group, corresponding one by one; the proportion of each group of tag sample sequences is adjusted accordingly to ≥0.4%. When the unique dual - end library tag combination consists of 192 pairs of library tags, that is, there are 192 upstream library tags in the IG5 group and 192 downstream library tags in the IG7 group, corresponding one by one, and the proportion of each group of tag sample sequences is adjusted accordingly to ≥0.1%. When the unique dual - end library tag combination consists of 288 pairs of library tags, that is, there are 288 upstream library tags in the IG5 group and 288 downstream library tags in the IG7 group, corresponding one by one, and the proportion of each group of tag sample sequences is adjusted accordingly to ≥0.07%. When the unique dual - end library tag combination consists of 384 pairs of library tags, that is, there are 384 upstream library tags in the IG5 group and 384 downstream library tags in the IG7 group, corresponding one by one, and the proportion of each group of tag sample sequences is adjusted accordingly to ≥0.05%.

3. The method according to claim 2, wherein The step (1) preparation of gDNA standard includes: diluting the gDNA standard with 1×IDTE Buffer.

4. The method according to claim 2, wherein The step (2) fragmentation of gDNA includes: breaking the DNA into fragments of 170 - 200bp.

5. The method according to claim 2, wherein The step (3) end repair includes: adding A to the 3’ end.

6. The method according to claim 2, wherein The step (4) adapter ligation includes: ligating both ends of the DNA double - strand fragment with added A to the pre - fabricated adapter.

7. The method according to claim 6, wherein The pre - fabricated adapter contains T sticky ends.

8. The method according to claim 2, characterized in that The step (5) purification of ligation products includes: removing adapter dimers and other components without ligated adapters.

9. The method according to claim 2, wherein The quality inspection of the purified library in step (8) includes: measuring the library concentration of the diluted library and converting it back to the pre-library concentration. If the library concentration is between 9 and 60 ng / µl and the Labchip result is normal, the library construction is qualified. If the requirements cannot be met, the library preparation needs to be carried out again.

10. The method according to claim 2, wherein The detection of the fragment size of the purified library in step (9) includes: detecting the diluted library. For a qualified library, the main peak of the library fragment is between 350 and 500 bp, and there are no obvious small fragments in the range of 10 to 150 bp.

11. A method for detecting cross - contamination of unique dual - ended library tags, characterized in that, It includes the following steps: constructing a gDNA library with a unique combination of dual-end library tags, subjecting the constructed gDNA library to on-machine sequencing, and reading the library tag sequences. When the sequencing results do not meet the following conditions, the library is considered contaminated: the maximum single-sided tag contamination ratio ≤ 2.5%, the maximum tag combination contamination ratio ≤ 0.01%, the number of sequences in each group of tag samples ≥ 5000, the variance coefficient of the mixed proportion of all tag combinations ≤ 0.5, the comprehensive sequence passing rate ≥ 97%, the proportion of sequences in each group of tag samples ≥ 0.2 / the logarithm of the library tag combinations, and the proportion of tags with a single-sided contamination greater than 1% should ≤ 10%; The single-sided tag contamination ratio is the cross-contamination ratio that occurs between tags within a group, and the contamination only occurs within the group, that is, contamination occurs within the IG5 group or / and within the IG7 group; When there is no cross-contamination during the production of a in IG7, for A in IG5, the contamination ratio containing B = the number of sequences containing B - a / the total number of sequences containing a, When there is no cross-contamination during the production of A in IG5, for a in IG7, the contamination ratio containing b = the number of sequences containing A - b / the total number of sequences containing A; When B contaminates A and b contaminates a, the contamination ratio of the B - b tag combination = (the number of sequences containing B - a / the total number of sequences containing a) × (the number of sequences containing A - b / the total number of sequences containing A); When there is no cross-contamination during the production of b in IG7, for B in IG5, the contamination ratio containing A = the number of sequences containing A - b / the total number of sequences containing b, When there is no cross-contamination during the production of B in IG5, for b in IG7, the contamination ratio containing a = the number of sequences containing B - a / the total number of sequences containing B; When A contaminates B and a contaminates b, the contamination ratio of the A - a tag combination = (the number of sequences containing A - b / the total number of sequences containing b) × (the number of sequences containing B - a / the total number of sequences containing B); The number of sequences in each group of tag samples is the number of correctly paired sequences in each group after system filtration, that is, the number of sequences containing A - a or the number of sequences containing B - b; The variance coefficient of the mixed proportion of all tag combinations is the variance coefficient of the proportion of the number of correctly paired sequences in each group after system filtration in the total number of correctly paired sequences after system filtration; The overall sequence pass rate is the ratio of the total number of correctly paired and valid sequences after passing through the system filtration in the sequencing reaction to the total number of all sequences after passing through the system filtration; The proportion of each group of tag sample sequences is the ratio of the number of correctly paired sequences in each group after passing through the system filtration to the total sequences after passing through the system filtration; The proportion of tags with more than 1% contamination on one side is: within the upstream library tags, the proportion of library tags with a contamination ratio greater than 1% to the total number of library tags; and, within the downstream library tags, the proportion of library tags with a contamination ratio greater than 1% to the total number of library tags; The unique paired-end library tag combinations are all composed of upstream library tags and downstream library tags. The upstream library tags are collectively called IG5, and IG5 includes A and B; the downstream library tags are collectively called IG7, and IG7 includes a and b. The matching and correct unique paired-end library tag combinations are A-a and B-b; the non-matching unique paired-end library tag combinations are A-b and B-a. After each sequencing reaction, the number of sequences of each of the above combinations is obtained through analysis.

12. The method according to claim 11, wherein The method for constructing the gDNA library successively includes the following steps: Step (1) Preparation of gDNA standard; Step (2) Fragmentation of gDNA; Step (3) End repair; Step (4) Adapter ligation; Step (5) Purification of adapter ligation products; Step (6) Library amplification; Step (7) Purification of amplified library; Step (8) Quality inspection of purified library; Step (9) Detection of fragment size of purified library; Among them, the Hamming distance of library tags within each of IG5 and IG7 is ≥3, and the sequence Hamming distance of library tags between IG5 and IG7 is ≥2; When the unique paired-end library tag combination consists of 96 pairs of library tags, that is, there are 96 upstream library tags in the IG5 group and 96 downstream library tags in the IG7 group, corresponding one by one; the proportion of each group of tag sample sequences is correspondingly adjusted to ≥0.2%; When the unique paired-end library tag combination consists of 48 pairs of library tags, that is, there are 48 upstream library tags in the IG5 group and 48 downstream library tags in the IG7 group, corresponding one by one; the proportion of each group of tag sample sequences is correspondingly adjusted to ≥0.4%; When the unique paired-end library tag combination consists of 192 pairs of library tags, that is, there are 192 upstream library tags in the IG5 group and 192 downstream library tags in the IG7 group, corresponding one by one, the proportion of each group of tag sample sequences is correspondingly adjusted to ≥0.1%; When the unique paired-end library tag combination consists of 288 pairs of library tags, that is, there are 288 upstream library tags in the IG5 group and 288 downstream library tags in the IG7 group, corresponding one by one, the proportion of each group of tag sample sequences is correspondingly adjusted to ≥0.07%; When the unique paired-end library tag combination consists of 384 pairs of library tags, that is, there are 384 upstream library tags in the IG5 group and 384 downstream library tags in the IG7 group, corresponding one by one, the proportion of each group of tag sample sequences is correspondingly adjusted to ≥0.05%.

13. The method according to claim 12, wherein The preparation of the gDNA standard in step (1) includes: diluting the gDNA standard with 1×IDTE Buffer.

14. The method according to claim 12, wherein The gDNA fragmentation in step (2) includes: fragmenting the DNA into fragments of 170 - 200 bp.

15. The method according to claim 12, characterized in that, The end repair in step (3) includes: adding A to the 3' end.

16. The method according to claim 12, wherein The adapter ligation in step (4) includes: ligating both ends of the double-stranded DNA fragment with added A to the prefabricated adapter.

17. The method according to claim 16, wherein The prefabricated adapter contains T sticky ends.

18. The method according to claim 12, wherein The purification of the ligation product in step (5) includes: removing adapter dimers and other components without ligated adapters.

19. The method according to claim 12, characterized in that, The quality inspection of the purified library in step (8) includes: measuring the library concentration of the diluted library and converting it back to the pre-library concentration. If the library concentration is between 9 - 60 ng / µl and the Labchip result is normal, the library construction is qualified. If the requirements cannot be met, the library preparation needs to be carried out again.

20. The method according to claim 12, characterized in that, The detection of the fragment size of the purified library in step (9) includes: detecting the diluted library. For a qualified library, the main peak of the fragment is at 350 - 500 bp, and there are no obvious small fragments in the 10 - 150 bp range.

Citation Information

Patent Citations

  • Preparation method of genome mixing sequencing library

    CN105671644A