Method for sample homogenization in high-throughput sequencing
Through the grading and mixing methods, the data imbalance caused by cell activity and quantity differences in high-throughput drug screening is solved, efficient sample uniformity and accuracy of sequencing results are achieved, and it is suitable for high-throughput automation processes, reducing cost and experimental complexity.
Patent Information
- Application Number
- CN202410038620.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-11
AI Technical Summary
During the high-throughput drug screening process, perturbations in cell activity and cell number lead to uneven data signal acquisition, affecting the accuracy and reliability of sequencing results. Especially in cell samples with low starting volume and large number spans, some samples cannot reach the detection limit, and a large amount of sequencing data is required.
Through the method of grading mixing and separating mixing, the samples are divided into multiple stages according to the cell count value, and the cDNA of equal cell volume is mixed in each stage. The final mixing volume is adjusted to achieve uniformization by calculating and adjusting it. Combined with conventional library construction and sequencing steps, it is ensured that the sequencing data volume of each sample reaches the minimum detection limit.
The homogenization of high-throughput sequencing samples is achieved, which reduces costs, improves the reproducibility and sequencing efficiency of results, is suitable for high-throughput automation processes, and reduces the need for multiple rounds of screening and repeated experiments.
Smart Images

Figure BDA0004659805270000111 
Figure BDA0004659805270000121 
Figure HDA0004659805310000011
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of high-throughput sequencing, and relates to a method for sample normalization in high-throughput sequencing, specifically a method and application for normalized sequencing of samples after high-throughput drug screening. Background Art
[0002] As one of the most important molecular biology analysis methods, DNA sequencing not only provides important data for basic biological research such as the revelation of genetic information and gene expression regulation, but also plays an important role in applied research such as gene diagnosis and gene therapy. High-Throughput Sequencing, also known as Next Generation Sequencing (NGS), is a technology that enables large-scale parallel sequencing on high-density biochips. Compared with traditional Sanger Sequencing, it has the characteristics of high data output and low cost per unit of data volume. The development of high-throughput sequencing technology has greatly promoted the development of the fields of genomics and life sciences.
[0003] High-throughput drug screening is an important experimental method for target and drug discovery. The process of high-throughput drug screening mainly includes: selecting a specific cell model, plating, adding a certain concentration of compounds, culturing for a certain period of time, and after the drug causes phenotypic and genotypic perturbations to the cells, collecting the cells in the screening plate for detection. Existing applications can perform high-throughput sequencing on high-throughput screening samples. However, due to the large range of perturbation effects of cell viability and cell number under the action of different drugs and concentrations, there will be a problem of uneven acquisition of data signals during batch processing, resulting in data that cannot be analyzed or even missing results. Therefore, there is an urgent need in the art for a universal high-throughput sequencing sample normalization method specifically for high-throughput, low starting amount, and large cell number span samples. Summary of the Invention
[0004] The present invention is completed based on the following discoveries of the inventors: When the inventors conducted high-throughput sequencing research on high-throughput screening samples, they found that the sequencing signal of the samples was directly affected by the amount of available cells. After the samples were mixed and subjected to the same library construction process, samples with a small amount of cells obtained a low amount of sequencing data; conversely, samples with a large amount of cells had a high amount of sequencing data. Thus, at a certain amount of sequencing data, some samples could not reach the detection limit, and to ensure that all samples reached the lowest detection limit, a large amount of sequencing data needed to be increased.
[0005] In the process of achieving homogenization of high-throughput sequencing samples, first, hierarchical mixing is performed. The minimum sampling volume is defined, and different samples are divided into different levels according to the cell count values of the samples. Each level is calculated and mixed separately to reduce sampling bias. Then, sequential mixing is carried out. The samples after hierarchical mixing are recalculated and mixed according to the cell count values of each level. If necessary, multiple mixings can be performed to achieve the optimal mixing of each sample. For the samples using the method of the present invention, even if the initial cell amounts differ by dozens of times, balanced sequencing results can be obtained. Based on this discovery, the present inventors have developed a method for sample homogenization in high-throughput sequencing.
[0006] In one aspect, the present invention provides a method for sample homogenization in high-throughput sequencing, comprising the following steps:
[0007] 1) Cell counting: Harvest cells, perform cell counting, and obtain the cell count value of each sample, wherein the cell count value of each sample is preferably 100 - 20000 cells / well, more preferably 500 - 10000 cells / well;
[0008] 2) cDNA synthesis: Lyse the cells, and add unique molecular tag sequences to each sample to convert mRNA into cDNA with molecular tags;
[0009] 3) Hierarchical mixing: Sort the samples according to the cell count values of each sample from low to high, divide different samples into multiple levels, take a certain volume of cDNA from each sample for mixing, and prepare multiple hierarchical tubes. Among them, in the same level, the cDNA of equal cell amounts from each sample is mixed, and the difference in cell count values of each sample in the same level is not more than 3 times, preferably not more than 2 times;
[0010] 4) Sequential mixing: According to the cDNA volume and cDNA amount after mixing in each hierarchical tube in step 3), take equal cell amounts of cDNA from multiple hierarchical tubes and mix them into one tube to obtain a homogenized sample. If, after calculation, the volume taken from some hierarchical tubes during mixing is lower than the minimum limit volume, it can be re-hierarchically mixed and then sequentially mixed.
[0011] Since the number of cells in a single sample is small, the total RNA amount is low, and it is difficult to quantify cDNA. However, after sample mixing, through the same experimental treatment, the cell amount and cDNA amount are approximately linearly correlated. Therefore, in the present invention, the cDNA amount is represented by the cell count value.
[0012] In some embodiments, the minimum cell count value of the samples in step 1) is higher than 100 cells / well, and the maximum cell count value is lower than 20000 cells / well.
[0013] In some embodiments, the difference in cell count values of each sample in step 1) is not more than 40 times.
[0014] In some embodiments, the difference in the cell count values of each sample in step 1) is not greater than 20-fold.
[0015] In some embodiments, in step 3), for each sample within the same level, equal amounts of cDNA are taken, and the sample with the highest cell count value in that level takes the lowest volume, or the sample with the lowest cell count value takes the total volume as the benchmark.
[0016] In some embodiments, the minimum sampling volume for each sample in step 3) is 10 μl.
[0017] In some embodiments, in step 4), for equal amounts of cDNA taken from multiple fractionation tubes, the fractionation tube with the highest cell count value after mixing takes the lowest volume as the benchmark.
[0018] In some embodiments, the minimum sampling volume for each fractionation tube in step 4) is 10 μl.
[0019] In some embodiments, the method further includes purifying, library constructing, and sequencing the normalized sample after sequential mixing. The library constructing step includes DNA fragmentation, end filling, adding an A tail at the end, adapter ligation, and library amplification. The library is a DNA library, and the sequencing is second-generation sequencing or third-generation sequencing.
[0020] In a specific embodiment, the cell count values of each sample in step 1) are x, 2x, 4x, 8x, 16x, 32x, 40x cells / well, respectively, where the range of x is 100 - 500 cells / well.
[0021] In a specific embodiment, according to the experimental requirements, the fractionation method can be as follows: samples with cell count values from x to 2x are classified into the first level, 2x to 4x into the second level, 4x to 8x into the third level, 8x to 16x into the fourth level, 16x to 32x into the fifth level, and 32x to 40x into the sixth level; or it can be fractionated in the following way: samples with cell count values from x to 2x are classified into the first level, 4x to 8x into the second level, 16x to 32x into the third level, and 40x into the fourth level. When fractionating and mixing the samples, for equal amounts of cDNA taken from each sample within the same level, the sample with the highest cell count value within the same level takes the lowest volume as the benchmark. For example, in the first level, the sample with a cell count value of 2x takes 10 μl, and the sample with a cell count value of x takes 20 μl.
[0022] In a specific embodiment, when the samples are mixed in batches, the cDNA with equal cell amounts is based on the lowest volume taken from the fractionation tube with the highest cell count value in the mixed samples. For example, when the fractionated samples obtained according to the above fractionation method are mixed in batches, the sampling volume ratios of each fractionation tube are: first level: second level: third level: fourth level: fifth level: sixth level = 267:133:67:33:17:10 or first level: second level: third level: fourth level = 300:75:19:10.
[0023] In a specific embodiment, if the cell count value of the sample is too low, such as less than 100 cells / well, it can be selectively excluded according to the experimental requirements.
[0024] In another aspect, the present invention provides a method for constructing a sequencing library, which includes the above method for sample normalization and the construction of the sequencing library, wherein the sequencing library is a second-generation sequencing library or a third-generation sequencing library.
[0025] In another aspect, the present invention provides the application of the above method in high-throughput drug screening.
[0026] The excellent technical effects of the method of the present invention are mainly in the following aspects:
[0027] 1. No additional reagents and consumables are required for nucleic acid purification and quantification, reducing costs and shortening the process;
[0028] 2. By using the method of the present invention, a normalized sample for high-throughput sequencing is obtained, avoiding multiple rounds of screening and repeated experiments, and improving the result reproducibility and sequencing cost efficiency;
[0029] 3. Through fractionation and batch mixing, high-throughput sample normalization is achieved, which is suitable for high-throughput automated processes and greatly improves the experimental efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 : Process of high-throughput sequencing using a normalized sample.
[0031] Figure 2 : Sequencing results of cell samples after fractionation and batch mixing.
[0032] Figure 3 : Sequencing results of drug screening cell samples after fractionation and batch mixing.
[0033] Figure 4 : Sequencing results of drug screening cell samples mixed in equal volumes. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To further understand the present invention, the preferred embodiments of the present invention will be described below in conjunction with examples. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention. Unless otherwise defined herein, scientific and technical terms used in connection with the present invention will have the meanings commonly understood by those of ordinary skill in the art.
[0035] Definition
[0036] For a better understanding of the present invention, the definitions and explanations of relevant terms are provided as follows.
[0037] As used herein, the terms "comprising", "including", "having", "containing" or any other variation thereof are intended to cover non-exclusive inclusion. For example, a composition, step, method, article or apparatus containing the listed elements is not necessarily limited to those elements, but may include other elements not expressly listed or elements inherent to such composition, step, method, article or apparatus.
[0038] As used herein, the connecting word "consisting of" excludes any unstated element, step or component. If used in a claim, this phrase will render the claim closed, excluding materials other than those described, except for conventional impurities associated therewith. When the phrase "consisting of" appears in a clause of the claim body rather than immediately following the subject, it only limits the elements described in that clause; other elements are not excluded from the claim as a whole.
[0039] As used herein, when an equivalent, concentration, or other value or parameter is expressed as a range, a preferred range, or a range defined by a series of upper preferred values and lower preferred values, this should be understood to specifically disclose all ranges formed by any pairing of any range upper limit or preferred value with any range lower limit or preferred value, whether or not the ranges are separately disclosed. For example, when the range "1 to 5" is disclosed, the described range should be interpreted as including the ranges "1 to 4", "1 to 3", "1 to 2", "1 to 2 and 4 to 5", "1 to 3 and 5", etc. When a numerical range is described herein, unless otherwise indicated, the range is intended to include its end values and all integers and fractions within the range.
[0040] As used herein, "and / or" is used to indicate that one or both of the stated situations may occur. For example, A and / or B includes (A and B) and (A or B).
[0041] As used herein, "a plurality of" means two or more.
[0042] The above terms or definitions are provided only to assist in understanding the present invention. These definitions should not be construed as having a scope less than that understood by those skilled in the art.
[0043] Experimental procedure
[0044] 1. Cell counting
[0045] 1) Use a cell counter to count, such as automated cell counter Celigo bright field counting, to obtain the number of cells in each sample well.
[0046] Any cell type is feasible and is not specifically limited here. It can be selected according to the experimental needs. If it is a suspension cell, it needs to be centrifuged at 1000 rpm for 5 min before automatic counting to allow the cells to settle and then cell counting is performed.
[0047] 2. Cell lysis and tagging
[0048] Add the lysis reaction solution, mix well, and incubate at room temperature for 10 min. After cell lysis, capture mRNA and at the same time add a unique sequence tag to each sample well. The lysis reaction solution is 10 μl of β-mercaptoethanol / ml, proteinase K, 100 μl of 10× Lysisbuffer (Takara, 635013), 5 μl of RNase Inhibitor (Takara, 635013) and 700 μl of water. It is necessary to control the lysis conditions here to avoid nucleic acid degradation and affect the experimental results.
[0049] The sequence tag is: 5’-AAGCAGTGGTATCAACGCAGAGTACAACAAGGTAC NNNNNN NNNN TTTTTTTTTTTTTTTTTTTTTTTTV-3’
[0050] The underlined part is the molecular tag sequence, N is any nucleotide among A, T, C, G; V is other bases except T such as A, G, C.
[0051] 3. cDNA synthesis
[0052] Prepare the cDNA synthesis reaction reagents, add the lysis reaction products, mix well, and perform the reaction on a PCR instrument to synthesize cDNA and pre-amplify. The present invention uses a reverse transcription buffer (Thermo Fisher Maxima H Minus Reverse Transcriptase EP0753), reverse transcriptase 8 U / μL Maxima, 8 mM MgCl (AM9530G), 0.8 μM template switching oligonucleotide TSO, 0.08 μM dNTP, 0.3 U / μL RNase inhibitor (Thermo Fisher), 300 Unit / μL ExoI (NEBM0239S), the amplification enzyme is 2×Kapa Hifi PCR ReadyMix (Kapa Biosystems, KK2602), and 10 μM full-length cDNA PCR primers. The reaction conditions are 42 °C for 1 h, 95 °C for 3 min, 4 cycles (98 °C for 20 s, 65 °C for 45 s, 72 °C for 3 min), several cycles (98 °C for 20 s, 67 °C for 20 s, 72 °C for 3 min), 72 °C for 5 min, and briefly store at 4 °C. The present invention is optimized to 10 - 12 reaction cycles.
[0053] Sequences for reaction
[0054] primer TSO: 5’-AAGCAGTGGTATCAACGCAGAGTGAATrGrGrG-3’
[0055] cDNA PCR primer: 5’-AAGCAGTGGTATCAACGCAGA-3’
[0056] 4. Gradient mixing and sequential mixing
[0057] 1) Gradient mixing: According to the initial cell amount of each sample (for samples with a shorter culture time) or the cell counting result (for samples with a longer culture time), grade them in multiples of 2 from low to high, and mix the samples at the same level in a single gradient tube. For example, in the first gradient tube, mix the samples with the least cell amount to the samples with a cell amount 2 times the lowest value. The sampling amount of the sample with the highest cell count value in each gradient tube is not less than the minimum sampling amount (10 μl), or the sampling amount of the sample with the lowest cell count value is not higher than the total volume of the sample. Other samples within the same gradient are mixed with equal cell amounts in proportion. The number of gradients needs to ensure that the sampling volume of all samples is not less than the minimum sampling volume.
[0058] 2) Sequential mixing: For the samples well-mixed by gradient, according to the DNA amount and volume obtained by mixing in the gradient tube, mix them again from high to low in the same method as gradient mixing until the sampling volume of all samples is not less than the minimum sampling volume.
[0059] 3) Purification after mixing
[0060] After the samples are mixed, the residual enzymes, buffers, and short fragments (such as primer adapters) in the system are removed by purification. The purification method is a conventional method in the art. In the present invention, magnetic bead purification is used: First, the magnetic beads are fully mixed, and 0.6 times the volume of the magnetic beads of the mixed sample volume is added to the mixed sample. After the sample is adsorbed onto the magnetic beads, the magnetic beads are washed twice with 80% ethanol, and finally the sample is eluted with 20 μl of nuclease-free water.
[0061] 5. Library construction
[0062] 1) Fragmentation, end filling, and A-tailing of DNA samples: Take 1.4 μL of ULtra II reaction buffer (NEB), 0.4 μL of ULtra II Enzyme Mix (NEB); 1.7 μL of TE buffer are mixed and 10 - 50 ng of cDNA is added. Incubate at 37 °C for 5 min, incubate at 65 °C for 30 min, and store temporarily at 4 °C. The fragmentation enzyme is a mixture of various types of endonucleases.
[0063] 2) Ligation of adapters: Using a ligase, the adapter sequences are ligated to both ends of the DNA fragments. Any reagent capable of DNA ligation in the art is feasible and can be selected according to actual needs. In the present invention, T4 DNA ligase, 6 μL of NEBNext ULtra II Ligation Master Mix (NEB), 0.2 μL of NEBNext Ligation Enhancer, (NEB), 0.2 μL of Adapter (1.5 μM) are added to the product of 1). The ligation is completed by reacting at 25 °C for 15 min and stored temporarily at 4 °C.
[0064] Adapter sequences:
[0065] Adapter Read1:
[0066] 5’-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3’
[0067] Adapter Read2:
[0068] 5’-CTGACCTCAAGTCTGCACACGAGAAGGCTAGA-3’
[0069] The product after ligation of adapters is purified. The purification method is a conventional method in the art. In the present invention, magnetic bead purification is used: First, the magnetic beads are fully mixed, and 0.5× times the volume of the magnetic beads of the mixed sample volume is added to the mixed sample. After the sample is adsorbed onto the magnetic beads, the magnetic beads are washed twice with 80% ethanol, and finally the sample is eluted with 20 μl of nuclease-free water.
[0070] 3) Library amplification: Using DNA polymerase, primers and buffer, while amplifying the ligation product, oligo sequences for sequencing are extended at both ends. The product can be purified and sequenced on a sequencer. Any reagent capable of DNA library amplification in the art is feasible and can be selected according to actual needs. In this invention, 25 μL of 2×Kapa Hifi PCR ReadyMix (Kapa Biosystems), 1 μL of Index Primer (i7, 5 μM), 1 μL of Index Primer (TruSeq i5, 5 μM), and 12.5 μL of nuclease-free water are added to the sample. Incubate at 98°C for 3 min, followed by 10 cycles (98°C for 15 s, 65°C for 30 s, 72°C for 4 min), then extend at 72°C for 10 min, and store temporarily at 4°C.
[0071] i7 Index Primer:
[0072] 5’-CAAGCAGAAGACGGCATACGAGAT[i7 index]GTCTCGTGGGC TCGG-3’
[0073] i5 Index Primer:
[0074] 5’-AATGATACGGCGACCACCGAGATCTACAC[i5 index]ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3’
[0075] 4) Purification of amplification products: The purification method is a conventional method in the art. In this invention, magnetic bead purification is used: First, thoroughly mix the magnetic beads, add 0.5 times the volume of the magnetic beads of the mixed sample volume to the mixed sample. After the sample is adsorbed onto the magnetic beads, wash the magnetic beads twice with 80% ethanol, and finally elute the sample with 20 μL of nuclease-free water. Fragment sorting is carried out when necessary: In this invention, it is optimized that 0.5 times the volume of the magnetic beads is added to the mixed sample. After the sample is adsorbed onto the magnetic beads, remove the supernatant, add 0.15 times the volume of the magnetic beads. After the sample is adsorbed onto the magnetic beads, wash the magnetic beads twice with 80% ethanol, and finally elute the sample with 20 μL of nuclease-free water.
[0076] 6. Sequencing and data analysis
[0077] The library is denatured and diluted, and then immediately sequenced on the machine. The amount of sequencing data is determined according to the number of mixed samples, and is generally calculated as the sum of the data amounts required for each sample. The data obtained from the machine is quality controlled according to the standard operating procedure, removing adapter sequences and low-quality reads. After obtaining clean reads, the data is split according to the tag sequences of each sample, and the read counts are statistically analyzed. In the present invention, at least 1M reads for each sample are required to meet the downstream data processing to obtain the analysis results. Samples that meet the minimum read count requirement are subjected to subsequent data alignment, quantification, and normalization, and then the clustering and differential expression analysis processes are carried out.
[0078] 7. Results
[0079] After sequencing, the data is split according to the sample tags to obtain the effective read counts for each sample. The minimum acceptable data amount for the sample is set, and samples with a data amount lower than this will undergo downstream data processing and analysis. In the present invention, the minimum acceptable data amount is set to 1M reads.
[0080] Example 1
[0081] In a 96-well cell culture plate, different cell seeding amounts were set as 1250, 2500, 5000, 10000, 20000 cells / well respectively for each group, with 3 replicates set for each group. Two cell lines were selected without drug treatment. A total of 30 cell samples were obtained. After culturing for 6 hours, the cells were collected.
[0082] 1. Cell counting
[0083] Since the cell culture time was short, according to the cell seeding amount of each sample, it was recorded as the cell count value.
[0084] 2. Cell lysis and addition of tags
[0085] The cells were lysed according to the optimized method in the experimental procedure, and molecular tags were added to each sample.
[0086] 3. cDNA synthesis
[0087] The cDNA synthesis reagents were prepared according to the optimized method in the experimental procedure, and the cell lysates were taken and mixed, and the reaction was carried out on a PCR instrument.
[0088] 4. Fractional mixing and sequential mixing
[0089] According to the cell seeding amount of each sample, the specific methods of fractional mixing and sequential mixing were calculated as follows:
[0090] 1) Hierarchical mixing: According to the cell seeding density, it is divided into three levels according to the seeding densities of 1250, 2500, 5000, 10000, and 20000 cells / well. Each level is mixed into 1 tube, resulting in a total of 3 mixed "pool" tubes. Specifically, for the 1250 group, 20 μl is taken from each sample; for the 2500 group, 10 μl is taken from each sample and added to pool-1 tube, for a total of 180 μl; for the 5000 group, 20 μl is taken from each sample; for the 10000 group, 10 μl is taken from each sample and added to pool-2 tube, for a total of 180 μl; for the 20000 group, 10 μl is taken from each sample and added to pool-3 tube, for a total of 60 μl.
[0091] 2) Fractional mixing: Take 120 μl from pool-1 tube, 30 μl from pool-2 tube, and 10 μl from pool-3 tube, and mix them into 1 tube, for a total of 160 μl, and then perform downstream experiments.
[0092] According to the optimized method in the experimental procedure, after thoroughly mixing the magnetic beads, take 96 μl of magnetic beads and add them to the sample for mixing. Incubate for 5 min. After the nucleic acid sample binds to the magnetic beads, wash the magnetic beads twice with 200 μl of 80% ethanol, open the lid and let it dry for 5 min to remove the residual ethanol, and then add 20 μl of nuclease-free water to elute the DNA sample.
[0093] 5. Library construction
[0094] Perform DNA fragmentation, end repair, adapter ligation, purification, library amplification, purification, and fragment sorting according to the optimized method in the experimental procedure.
[0095] 6. Sequencing and data analysis
[0096] After the library preparation is completed, sequence on the machine according to the method in the experimental procedure. The sequencing mode is PE150, and the designed effective sequencing volume for each sample is 2.5M reads. Perform data quality control and splitting according to the method in the experimental procedure, and count the number of reads for each sample. The minimum acceptable sequencing volume is 1M reads. Samples that meet the minimum reads number requirement enter the subsequent analysis process.
[0097] 7. Results
[0098] After hierarchical and fractional mixing ( Figure 2 ), all samples in this example are above the minimum detection line of 1M reads. Due to the large span of cell seeding densities, according to the actually detected number of reads and the volume of the mixed samples input, it is inversely calculated that if hierarchical and fractional mixing is not performed and direct equal-volume mixing is carried out, the reads distribution is as shown in Figure 2 the right figure below. The range of sequencing reads of the samples is widened, and 15 samples are below the minimum number of reads required for detection and cannot be analyzed subsequently. The data and results of these samples need to be obtained by redesigning the experiment.
[0099] Example 2
[0100] Set the cell seeding density to 2000 cells / well, 20 μl per well, add 4 compound drugs for treatment, set 3 concentration gradients for each drug, and set 3 replicates for each treatment. Additionally, add a control sample with DMSO, also set 3 replicates. There are a total of 39 cell samples. Collect the cells after 24 h of culture.
[0101] 1. Cell counting
[0102] Perform cell counting according to the optimized method in the experimental procedure, export the counting result file, as shown in Table 1. According to the counting values, calculate the specific methods for grading and sub - mixing.
[0103] 2. Cell lysis and tagging
[0104] Lyse the cells according to the optimized method in the experimental procedure and add molecular tags to each sample.
[0105] 3. cDNA synthesis
[0106] Prepare the cDNA synthesis reagents according to the optimized method in the experimental procedure, take the cell lysates for mixing, and perform the reaction on a PCR instrument.
[0107] 4. Grading mixing and sub - mixing
[0108] According to the cell counting values of each sample, calculate the specific methods for grading and sub - mixing as shown in Table 1.
[0109] Table 1: Cell counting results and sample mixing methods
[0110]
[0111]
[0112] Finally, mix each sample into 1 tube according to the grading and sub - mixing methods and the mixing volume shown in Table 1, with a total volume of 67 μl for downstream experiments.
[0113] According to the optimized method in the experimental procedure, after thoroughly mixing the magnetic beads, take 40.2 μl and add it to the sample for mixing. After incubating for 5 min until the nucleic acid sample binds to the magnetic beads, wash the magnetic beads twice with 200 μl of 80% ethanol, open the lid for 5 min to dry the residual ethanol, and add 20 μl of nuclease - free water to elute the DNA sample.
[0114] 5. Library construction
[0115] Fragment the DNA, repair the ends, ligate adapters, purify, amplify the library, purify, and sort the fragments according to the optimized method in the experimental procedure.
[0116] 6. Sequencing and Data Analysis
[0117] After the library preparation is completed, sequence on the machine according to the method in the experimental procedure. The sequencing mode is PE150, and the designed effective sequencing amount for each sample is 2.5M reads. Perform data quality control and splitting according to the method in the experimental procedure, and count the number of reads for each sample. The minimum acceptable sequencing amount is 1M reads. Samples that meet the minimum read number requirement enter the subsequent analysis process.
[0118] 7. Results
[0119] After hierarchical and fractional mixing ( Figure 3 ), all samples in this example are above the minimum detection line of 1M reads. Due to the drug effect, cells with the same inoculation amount showed a large span after drug treatment, and the proportion of dead cells varied. According to the actual number of reads and the input mixed sample volume, the reads distribution after direct equal-volume mixing without hierarchical and fractional mixing was calculated retrospectively as shown in Figure 3 the right figure. The sequencing reads range of the samples widened, and 7 samples were below the minimum number of reads required for detection and could not be analyzed further. The data and results of these samples need to be obtained by redesigning the experiment.
[0120] Comparative example
[0121] Set the cell inoculation amount to 4000 cells / well, 20 μl per well, add drug treatment, add 4 compound drugs, each with 3 concentration gradients, and set 3 replicates for each treatment. Additionally, add a control sample with DMSO, also set 3 replicates. There are a total of 39 cell wells. After 24 hours of culture, collect the cells.
[0122] 1. Cell Lysis and Tagging
[0123] Without cell counting, directly lyse the cells according to the optimized method in the experimental procedure, and add molecular tags to each sample.
[0124] 2. cDNA Synthesis
[0125] Prepare the cDNA synthesis reagents according to the optimized method in the experimental procedure, take the cell lysate and mix, and perform the reaction on a PCR instrument.
[0126] 3. Without hierarchical mixing and fractional mixing, mix equal volumes according to the cell inoculation amount. Take 10 μl from each sample and mix them into 1 tube, a total of 390 μl, for downstream experiments.
[0127] Purify the cDNA products according to the optimized method in the experimental procedure.
[0128] 4. Library construction
[0129] Fragment the DNA, perform end repair, ligate adapters, purify, amplify the library, purify again, and perform fragment sorting according to the optimized method in the experimental procedure.
[0130] 5. Sequencing and data analysis
[0131] After the library preparation is completed, sequence the samples on the machine according to the method in the experimental procedure. The sequencing mode is PE150, and the designed effective sequencing volume for each sample is 2.5M reads. Perform data quality control and demultiplexing according to the method in the experimental procedure, and count the number of reads for each sample. The minimum acceptable sequencing volume is 1M reads. Samples that meet the minimum read count requirement enter the subsequent analysis process.
[0132] 6. Results
[0133] The sequencing data shows that in this example, a total of 11 samples are below the minimum detection line of 1M and are excluded from the subsequent analysis process ( Figure 4 right figure). To obtain valid data and results for more samples, after doubling the sequencing data volume in this example, it can be seen that the data volume increases basically in proportion. There is 1 sample with less than 1M reads, and at the same time, the data volume of 5 samples exceeds 10M, reaching nearly 14M reads at most ( Figure 4 left figure). Considering the sequencing cost efficiency, the data volume is not increased anymore, and only this sample can be excluded from the subsequent analysis process. It can be seen from this example that for samples without normalization, the sequencing data volume has a large deviation and a high rejection rate, and it will also increase the sequencing cost.
[0134] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for sample homogenization in high-throughput sequencing, comprising the following steps: 1) Cell counting: harvest cells, perform cell counting, and obtain the cell count value of each sample; 2) cDNA synthesis: Lyse the cells and add a unique molecular tag sequence to each sample to obtain cDNA with molecular tags; 3) Grading and mixing: Different samples are divided into multiple levels according to the cell count value of each sample. A certain volume of cDNA is taken from each sample and mixed to prepare multiple graded tubes, in which equal amounts of cDNA are taken from each sample in the same level and mixed; 4) Mixing in batches: According to the cDNA volume and cDNA amount after mixing in each graded tube in step 3), cDNA of equal cell amount is taken from multiple graded tubes and mixed into one tube to obtain a homogenized sample.
2. The method according to claim 1, wherein, In the step 1), the lowest cell count value of the sample is higher than 100 cells / well, and the highest cell count value is lower than 20,000 cells / well.
3. The method according to claim 1, wherein, The difference in cell count values of each sample in step 1) is no more than 40 times.
4. The method according to claim 1, wherein, The difference in cell count values of each sample in step 1) is no more than 20 times.
5. The method according to claim 1, wherein In the step 3), the different samples are divided into multiple levels according to the cell count value of each sample, which is sorted from low to high, wherein the difference in the cell count value of samples in the same level is not greater than 3 times.
6. The method according to claim 1, wherein, In the step 3), the cell count values of each sample are sorted from low to high, and different samples are divided into multiple levels, wherein the cell count values of samples in the same level differ by no more than 2 times.
7. The method according to claim 1, wherein In the step 3), the cDNA of equal cell amounts is based on the lowest volume of the sample with the highest cell count value in that level, or the total volume of the sample with the lowest cell count value.
8. The method according to claim 1, wherein The minimum sampling volume of each sample in step 3) is 10 μl.
9. The method according to claim 1, wherein In step 4), the cDNA of equal cell amounts is based on the lowest volume in the graded tube with the highest sample cell count value after mixing.
10. The method according to claim 1, wherein, The minimum sampling volume of each classification tube in step 4) is 10 μl.
11. The method according to claim 1, wherein, The method also includes purifying, constructing a library and sequencing the homogenized sample after mixing in batches.
12. The method according to claim 11, wherein, The library construction steps include DNA fragmentation, end-filling, end-overhang A, adapter ligation and library amplification.
13. The method according to claim 11, wherein, The library is a DNA library.
14. The method according to claim 11, wherein, The sequencing is second generation sequencing or third generation sequencing.
15. A method for constructing a sequencing library, characterized in that, The method comprises the sample homogenization method according to any one of claims 1 to 14, and the construction of a sequencing library.
16. The method according to claim 15, wherein, The sequencing library is a second-generation sequencing library or a third-generation sequencing library.
17. Use of the method according to any one of claims 1 to 16 in high-throughput drug screening.