Normalization before sequencing

By standardizing nucleic acid quantities through equal volume sampling and adjusting based on initial sequencing reads, the method addresses the challenges of cost, time, and equipment requirements in existing normalization methods, ensuring consistent and efficient sequencing outcomes.

FR3165893A1Pending Publication Date: 2026-03-06ZIWIG
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
FR2024009234
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Existing methods for normalizing nucleic acid quantities before sequencing are costly, time-consuming, and require specialized equipment, leading to potential errors and inconsistencies in sequencing results due to variations in nucleic acid concentrations across samples.

Method used

A method for normalizing nucleic acid quantities by taking equal volumes from each sample, performing an initial sequencing to determine the number of reads, and adjusting subsequent volumes based on these reads to achieve a target number of reads, eliminating the need for specialized reagents and equipment.

Benefits of technology

This approach ensures consistent and efficient sequencing results by standardizing nucleic acid amounts across samples, reducing costs and time, and minimizing errors, while maintaining operational efficiency and resource optimization.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

[Method for normalizing the quantity of nucleic acid between at least two samples, characterized in that it comprises the following steps: - having at least two samples, each having a volume V containing nucleic acid; - taking a known volume VCO from each of the at least two samples; - from said volumes VCO taken, performing initial sequencing to obtain, for each volume VCO, a number of reads; - based on the number of reads for each volume VCO, calculating a sample volume VEP for each sample such that each VEP corresponds to the same theoretical number of reads NLT. Figure to be published with the abstract: Fig. 1]
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Normalization before sequencing

[0001] The present invention relates to the technical field of sequencing, and more specifically to the technical field of next-generation sequencing, sometimes called "next-generation sequencing" or "NGS".

[0002] For the purposes of this description, an interval described by the expression "between [...] and [...]" includes the bounds. For example, the interval "between 1 and 2" includes the two values ​​"1" and "2" as well as all values ​​that are both strictly greater than 1 and strictly less than 2.

[0003] In the context of this description, the term "approximately" preceding a numeric value means that the value can be changed by plus or minus 10%. In the specific case of a numeric value being an interval limit, the term "approximately" means that the lower limit can be reduced by 10% or the upper limit can be increased by 10%. It is also possible to omit the term "approximately" preceding a numeric value.

[0004] Next-generation sequencing (NGS) is a revolutionary method for reading nucleic acid sequences, particularly DNA, much faster and at a lower cost than traditional sequencing methods. This technology has radically transformed genomics, molecular biology, and various areas of biomedical research. Its impact on research and medicine is immense, paving the way for innovative discoveries and applications in many fields of science and health.

[0005] NGS sequencing relies on interdependent steps, from sample preparation to bioinformatic analysis and interpretation of results.

[0006] The NGS process therefore comprises several key steps, each playing a crucial role in the generation of accurate and detailed sequencing data.

[0007] The first step of NGS is library preparation. During this step, the target DNA or RNA is extracted from a sample (blood, saliva, etc.) and fragmented into small pieces. During this process, UDI (Unique Dual Indexing) sequences are attached to the fragments in order to link each fragment to the source sample.

[0008] After library preparation, the sample is subjected to purification and amplification by PCR or in situ amplification to increase the amount of target genetic material. This step is crucial to ensure that sufficient genetic material is available for sequencing.

[0009] Next, the prepared nucleic acid sample, in particular DNA, is loaded onto a next-generation sequencer. The actual sequencing then begins, using various technologies such as sequencing by synthesis (SBS), the ligation sequencing, or other methods based on nanopores or semiconductors.

[0010] Each technology has its own unique mechanisms for reading DNA sequences. For example, in SBS, fluorescent nucleotides are added sequentially and their incorporation is detected by light signals.

[0011] After sequencing, the raw data undergo a bioinformatics processing procedure. This includes aligning the obtained sequences with known genomic references and assembling these sequences to reconstruct the original sequence. This step is crucial for correctly interpreting the NGS data. The data are also analyzed to detect genetic variations such as point mutations, insertions, deletions, and chromosomal rearrangements.

[0012] Finally, the sequencing results are interpreted in a clinical or research context. In biomedical research, NGS is used to understand the genetic basis of diseases, discover new biomarkers, and develop personalized therapies. Clinically, it aids in the diagnosis and treatment of various conditions, including genetic diseases, cancers, and infections.

[0013] UDI sequences are essential in next-generation sequencing (NGS) to improve data accuracy and reliability. This technique involves assigning two unique indices to each sample, thereby reducing the risk of misassignment and cross-contamination.

[0014] During library preparation, the DNA or RNA fragments of each sample are labeled with these indices, enabling specific and precise identification during sequencing. This doubly indexed method is particularly important during multiplexing, where many samples are mixed and sequenced simultaneously, as it minimizes the risk of incorrect sequence assignment, thus improving the quality and reliability of genomic sequencing data.

[0015] In the context of NGS, a "read" is a sequence of nucleotide bases that is read.

[0016] The number of reads is the total number of individual reads obtained during the Sequencing. It influences the accuracy and resolution of the results. A high number of readings generally improves the quality of the results.

[0017] After purification and amplification, the concentrations of nucleic acid, in particular DNA, are not the same for the different samples treated.

[0018] These samples are nevertheless intended to be mixed and sequenced in a single step.

[0019] It follows that one of the crucial steps in the NGS process conditioning the final results lies in the introduction into the sequencer of the same quantity of genetic material for each sample taken.

[0020] To do this, a number of techniques are commonly used, such as fluorescence assay or real-time PCR assay.

[0021] These dosages are complex to implement, they are costly, they require the ordering of specific equipment, and they also lengthen the overall handling time.

[0022] In the prior art, a number of steps are disclosed.

[0023] In application WO18136526 filed on behalf of COUNSYL INC., a method for preparing libraries of improved capture probes is disclosed. This application addresses several aspects of sequencing, including sequencing depth and performance-based selection of capture probes.

[0024] In application WO23034090 filed on behalf of NATERA, INC., a method for non-invasive prenatal testing is disclosed. This method includes cell-free DNA extraction from a pregnant woman's blood sample, amplification, and sequencing to determine the ploidy state of chromosomes of interest.

[0025] In application WO19170773 on behalf of CANCER RESEARCH TECHNOLOGY LIMITED, a computerized method for the detection of circulating tumor DNA from a sample taken from a patient is disclosed.

[0026] In application WO22197864 filed on behalf of NATERA, INC., a method for detecting graft rejection is disclosed. The amount of donor DNA present in the recipient is one of the parameters taken into account to determine whether there is a graft rejection problem in the recipient.

[0027] Applications WO22182878 and WO21243045, in the name of the same applicant, contain similar disclosures.

[0028] In application US20230399679 filed on behalf of GTSEEK, LLC., a method using ligand-modified primers to normalize DNA concentrations is disclosed.

[0029] In the publication Bruinsma, S., Burgess, J., Schlingman, D., Czyz, A., Morrell, N., Ballenger, C.,... Gormley, NA (2018). Bead-linked transposomes enable a normalization-free workflow for NGS library preparation. BMC Genomics, 19, 722. https: / / doi.org / 10.1186 / sl2864-018-5096-9, a method for preparing libraries for next-generation sequencing is disclosed. This method uses bead-linked transposomes to capture a fixed amount of DNA, allowing the direct preparation of libraries from blood and saliva samples. It is It is presented as being effective, even with small amounts of DNA. It eliminates the need for DNA quantification before sequencing.

[0030] Surprisingly, it was highlighted by the applicant that a new method for normalizing the quantity of nucleic acid before sequencing could be used.

[0031] Thus, the invention relates to a method for standardizing the quantity of nucleic acid between at least two samples, characterized in that it comprises the following steps:

[0032] - to have at least 2 samples, each having a volume V comprising nucleic acid;

[0033] - take a known volume VCO from each of at least 2 samples;

[0034] - from said VCO volumes collected, proceed with a first sequencing allowing each of the VCO volumes to obtain a number of readings;

[0035] - depending on the number of readings of each VCO volume, calculate a volume sample to be taken VEP for each sample such that each VEP corresponds to an identical number of theoretical NLT readings.

[0036] In one embodiment, the nucleic acid is chosen from the group consisting of ribonucleic acid (RNA) and in particular messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), small interference RNA (siRNA), microRNA (miRNA) and / or long non-coding RNA (lncRNA), deoxyribonucleic acid (DNA) and in particular chromosomal DNA (chDNA), mitochondrial DNA (mDNA), plasmid DNA (pDNA), circular DNA (cDNA), linear DNA (1DNA) and / or recombinant DNA (rDNA), and mixtures thereof.

[0037] In one embodiment, the nucleic acid is chosen from the group consisting of ribonucleic acid (RNA) and in particular messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), small interference RNA (siRNA), microRNA (miRNA) and / or long non-coding RNA (lncRNA).

[0038] In one embodiment, the nucleic acid is chosen from the group of bacterial RNAs.

[0039] In one embodiment, the nucleic acid is chosen from the group consisting of deoxyribonucleic acid (DNA) and in particular chromosomal DNA (chDNA), mitochondrial DNA (mDNA), plasmid DNA (pDNA), circular DNA (cDNA), linear DNA (DNA1) and / or recombinant DNA (rDNA).

[0040] In a preferred embodiment, the nucleic acid is circular DNA (cDNA).

[0041] In a preferred embodiment, the nucleic acid is DNA.

[0042] In a preferred embodiment, the nucleic acid is complementary DNA to RNA.

[0043] In a preferred embodiment, the nucleic acid is DNA complementary to RNA selected from the group consisting of mRNA, tRNA, rRNA, snRNA, siRNA, miRNA and lncRNA,

[0044] In a preferred embodiment, the nucleic acid is DNA complementary to bacterial RNA,

[0045] In a preferred embodiment, the nucleic acid is DNA complementary to miRNA, preferably salivary.

[0046] In a preferred embodiment, the samples comprise purified nucleic acid. In a preferred embodiment, the samples comprise amplified nucleic acid. In a particularly preferred embodiment, the samples comprise purified and amplified nucleic acid.

[0047] In one embodiment, each sample comprises nucleic acid labeled with an identification means. This identification means allows each nucleic acid fragment to be linked to its original sample. In a preferred embodiment, this identification means is a UDI sequence label.

[0048] In one embodiment, the number of samples is between 2 and approximately 1000, preferably between 2 and approximately 950, preferably between 2 and approximately 900, preferably between 2 and approximately 850, preferably between 2 and approximately 800, preferably between 2 and approximately 750, preferably between 2 and approximately 700, preferably between 2 and approximately 650, preferably between 2 and approximately 600, preferably between 2 and approximately 550, preferably between 2 and approximately 500, preferably between 2 and approximately 450, preferably between 2 and approximately 400, preferably between 2 and approximately 350, preferably between approximately 5 and approximately 350, preferably between approximately 20 and approximately 350, preferably between approximately 50 and approximately 300, preferably between approximately 50 and approximately 200, preferably between approximately 50 and approximately 100,preferably between approximately 75 and approximately 125, preferably between approximately 90 and approximately 110, preferably between approximately 95 and approximately 105.

[0049] In one embodiment, the VCO volumes are mixed between the step of collecting said VCO volumes and the step of performing said first sequencing.

[0050] In one embodiment, this mixing step is carried out by means of vortexing followed by centrifugation.

[0051] In one embodiment, said known volume VCO is between approximately 0.5 pl and approximately 2.5 pl, preferably between approximately 0.6 pl and approximately 2.4 pl, of preferably between about 0.7 ft and about 2.3 ft, preferably between about 0.8 ft and about 2.2 ft, preferably between about 0.9 ft and about 2.1 ft, preferably between about 1.0 ft and about 2.0 ft, preferably between about 1.1 ft and about 1.9 ft, preferably between about 1.2 ft and about 1.8 ft, preferably between about 1.3 ft and about 1.7 ft, preferably between about 1.4 ft and about 1.6 ft, preferably about 1.5 ft.

[0052] In one embodiment, said VCO volume is the same for each sample. This embodiment is generally preferred because it simplifies calculations.

[0053] In one embodiment, said VCO volume differs in at least some of the samples. This embodiment can be useful when the sample volumes are very different, and obviously, additional calculation steps must be performed in this case.

[0054] Sequencing depth in the context of NGS is an important concept for understanding the quantity and quality of genetic sequencing data obtained.

[0055] It can be expressed in several ways.

[0056] First, it can be expressed as "X", representing the average number of times a particular nucleotide base in the genome is sequenced. Therefore, for example, a sequencing depth of 30X means that, on average, each base in the genome has been sequenced 30 times.

[0057] This way of presenting sequencing depth is particularly useful when the sequencing involves a domain of known size, for example, the entire human genome. Indeed, the entire human genome is known to be approximately 3.2 billion nucleotide pairs in size. Therefore, it is easy to express the sequencing depth as "X".

[0058] Sequencing depth can also be expressed as the number of reads. This is particularly useful when the size of the genetic material being studied is unknown, difficult to quantify, and / or variable. This is the case, for example, when sequencing DNA from an miRNA sample.

[0059] In one embodiment, said first sequencing has a sequencing depth greater than 0.5X.

[0060] In one embodiment, said first sequencing has a sequencing depth greater than 1000 reads, preferably greater than 10000 reads.

[0061] In one embodiment, said first sequencing is carried out over a period of time between approximately 2 hours and approximately 18 hours, preferably between approximately 5 hours and approximately 15 hours, preferably between approximately 7 hours and approximately 13 hours, preferably between approximately 8 hours and approximately 12 hours, preferably approximately 10 hours.

[0062] In one embodiment, said VEP volume is between approximately 0.5 pl and approximately 7.5 pl, preferably between approximately 0.6 pl and approximately 7.2 pl, preferably between approximately 0.7 pl and approximately 6.9 pl, preferably between approximately 0.8 pl and approximately 6.6 pl, preferably between approximately 0.9 pl and approximately 6.3 pl, preferably between approximately 0.9 pl and approximately 6.0 pl, preferably between approximately 0.9 pl and approximately 5.7 pl, preferably between approximately 0.9 pl and approximately 5.4 pl, preferably between approximately 0.9 pl and approximately 5.1 pl, preferably between approximately 0.9 pl and approximately 4.8 pl. Obviously, this volume varies depending on the samples.

[0063] In one embodiment, said theoretical NLT read number is between about 1,000,000 reads and about 1,500,000 reads, preferably between about 1,000,000 reads and 1,300,000 reads, preferably between about 1,100,000 reads and 1,200,000 reads.

[0064] In one embodiment, for each sample, the sum of the volumes VCO and VEP is less than 0.95 times the volume V, preferably less than 0.90 times the volume V, preferably less than 0.85 times the volume V, preferably less than 0.80 times the volume V.

[0065] The invention also relates to a sequencing method, characterized in that it comprises:

[0066] - the implementation of a method for normalizing the quantity of nucleic acid between at least two samples according to the invention;

[0067] - the taking, for each at least 2 samples, of said VEP volume;

[0068] - the performance of a second sequencing.

[0069] In one embodiment, said second sequencing has a sequencing depth of between about 15X and about 45X, preferably between about 20X and about 40X, preferably between about 25X and about 35X, preferably about 30X.

[0070] In one embodiment, said second sequencing has a sequencing depth of between about 10 million reads and about 40 million reads, preferably between about 13 million reads and about 30 million reads.

[0071] In one embodiment, said second sequencing is carried out over a period of time between approximately 3 hours and approximately 23 hours, preferably between approximately 8 hours and approximately 18 hours, preferably between approximately 10 hours and approximately 16 hours, preferably between approximately 12 hours and approximately 14 hours, preferably approximately 13 hours.

[0072] In one embodiment, the samples are mixed between the step of collecting said VEP volumes and the step of performing said first sequencing.

[0073] In one embodiment, this mixing step is carried out by means of vortexing followed by centrifugation.

[0074] The invention also relates to a method for processing several series of samples, characterized in that it comprises the following steps:

[0075] - the implementation of a method for normalizing the quantity of nucleic acid between at least two samples from a first series of samples according to the invention;

[0076] - the taking, for each at least 2 samples, of said VEP volume;

[0077] - the performance of a second sequencing;

[0078] characterized in that the performance of said second sequencing takes place simultaneously, and preferably on the same sequencing plate, with the performance of a first sequencing of a method of normalizing the quantity of nucleic acid between at least two samples according to the invention, said other samples belonging to another series of samples.

[0079] In this embodiment, it is important to adjust the quantities of nucleic acids according to the target number of reads. For example, if the target number of reads for the first sequencing is 10,000 and the target number of reads for the second sequencing is 20 million, then the quantity of nucleic acid in the samples intended for said first sequencing should be equal to 0.0005 times (10,000 / 20,000,000) the quantity of nucleic acid in the samples intended for said second sequencing.

[0080] In one embodiment, the ratio of the quantity of nucleic acid in the first sample to the quantity of nucleic acid in the second sample is between 0.0000232 and 0.00232, preferably between 0.00005 and 0.0015, preferably between 0.0001 and 0.001, preferably between 0.0001 and 0.0007, preferably between 0.0001 and 0.0006, preferably between 0.0001 and 0.0005, preferably between 0.0001 and 0.0003, preferably 0.000232.

[0081] This embodiment with simultaneous steps is particularly interesting; it is distinguished by the possibility of simultaneously managing two distinct processes: the quantification of nucleic acid from several samples of one series and the sequencing of several other samples from another series.

[0082] This approach allows for remarkable resource optimization. In particular, it does not require an increase in the time allocated to the sequencing machine, thus preserving the operational efficiency of the process.

[0083] Moreover, the use of consumables, and in particular sequencing plates, remains unchanged compared to conventional sequencing, which is a considerable advantage from an economic point of view.

[0084] Next, better management of laboratory workflows is observed, allowing for more complete and diversified analysis of samples without requiring additional equipment.

[0085] Furthermore, the risk of error is limited, because the sample handling steps are also limited.

[0086] In one embodiment, the first and / or second sequencing is an NGS sequencing.

[0087] In one embodiment, the first and / or second sequencing is an NGS sequencing performed on a sequencing plate.

[0088] In one embodiment, the first and second sequencing is an NGS sequencing.

[0089] In one embodiment, the first and second sequencing is an NGS sequencing performed on a sequencing plate.

[0090] In one embodiment, said sequencing plate is chosen from the group consisting of plates using flow cells, semiconductor chips, nanopore flow cells, microtiter plates, bead-based plates.

[0091] Preferably, said sequencing plate is a plate using flow cells.

[0092] In one embodiment, the second sequencing results in a number of reads of between about 6 million and about 28 million, preferably between about 7 million and about 26 million, preferably between about 8 million and about 25 million, preferably between about 9 million and about 24 million, preferably between about 10 million and about 23 million, preferably between about 11 million and about 22 million, preferably between about 12 million and about 21 million, preferably between about 13 million and about 20 million.

[0093] By using partial or complete sequencing to adjust the amounts of nucleic acid taken, the method according to the invention has a large number of advantages.

[0094] Since the method used is the same as for subsequent sequencing, no bias is introduced between the techniques, as can be the case, for example, when using dye for quantification, and the result is very consistent between the different samples. Furthermore, this method does not require any additional equipment for quantification.

[0095] Description of the figure

[0096] Figure 1 is a graphical representation corresponding to Table 1. It is immediately deduced from this that when the first-sequencing assay according to the invention is used, the number of subsequent readings during the complete sequencing is much more homogeneous, which improves the quality of the results. In [Fig. 1], the graphical representation corresponding to the fluorimetric assay is on the right, and the graphical representation corresponding to the assay by first sequencing according to the invention is on the left.

[0097] Example 1: Comparison between a method according to the invention and a method according to the prior art

[0098] 1.1: Quantification by fluorescence spectrometry before sequencing and then sequencing

[0099] We have 96 samples, corresponding to 96 DNA libraries.

[0100] The following steps are carried out.

[0101] Step 1: Dilute the Qubit® dsDNA HS Reagent (ThermoScientific product, reference Q32854, used for the precise quantification of double-stranded DNA) at a ratio of 1:200 (=1X) in the Qubit® dsDNA HS Buffer (specific buffer solution for use with the Qubit® dsDNA HS reagent). Prepare 200 qL per sample to be assayed, including 3 additional volumes (2 standards and 1 control DNA).

[0102] Step 2: Distribute the mixture (Qubit® dsDNA Reagent + Qubit® dsDNA Buffer) into the plate containing the 96 samples: - 190 places for the standard ones, in Al and B1 (specific locations on the plate) - 199 pl for the control (in Cl) and the samples to be tested.

[0103] Step 3: Add the DNA: - 10 μl of standard DNA, in Al (0 ng / μl) and B1 (10 or 100 ng / μl depending on the kit used, which indicates the concentration of DNA in these standard samples) - 1 µl of control DNA (in Cl) or DNA from the samples (the samples to test).

[0104] Step 4: Analyze the 96 wells with a Fluorimeter (Spark M10, TECAN, an instrument used to measure fluorescence) and retrieve the concentrations of each sample / library.

[0105] Step 5: Perform equimolar calculations for each preparation based on the molarities.

[0106] Step 6: Make the equimolar mixture (96 to 1 tube).

[0107] As immediately deducible from the above, it should be noted that the method by fluorescence spectrometry assay before sequencing requires special reagents which are expensive and necessary only for said assay, in particular those sold under the Qubit® brand.

[0108] In addition, special equipment is also required: spectrofluorimeter, spectrofluorimeter plate (n=96), single-channel pipettes, multi-channel pipettes, disposable tips.

[0109] Moreover, in practice, the manipulation takes longer (an additional day).

[0110] In addition, this requires individual assays of each of the samples (n=96), as well as many calculations for equimolar pooling (mixing of all libraries with volumes calculated previously).

[0111] Finally, it is important to recognize that, although the fluorimetric assay is relatively reliable, it can be affected by various factors. Errors such as incorrectly sized libraries, failed preparation, the presence of primer dimers, or technical problems with Qubit® reagents can distort the results.

[0112] Next, the sequencing steps are as follows.

[0113] Step 1: Assay dsDNA High Sensitivity Kit DeNovix (a DeNovix brand kit for the quantification of double-stranded DNA) + TapeStation, Agilent (an Agilent brand instrument used for the analysis of the quality and size of nucleic acid samples) of the equimolar mixture.

[0114] Step 2: Sequencing NovaSeqôOOO (a next-generation sequencer produced by Illumina, used for high-throughput sequencing of DNA and RNA, here specified with a sample load at a concentration of 1.6nM) or other ILLUMINA sequencer (a leading brand in genetic sequencing technologies) depending on the number of samples (sequencer selection may depend on the volume and nature of the samples to be analyzed).

[0115] 1.2: Assay by first sequencing according to the invention followed by second sequencing

[0116] Step 1: Equivolume mix: l,5qL / library.

[0117] Step 2: dsDNA High Sensitivity Kit assay. DeNovix (ref: KIT-DSDNA-HIGH- 2) of the mixture (1 tube) (DeNovix brand double-stranded DNA precision measurement kit): - Prepare the mixture by sample: • 198pL AccuClear Buffer (a buffer for reactions) biochemicals, part of the AccuClear kit); • 2pL AccuClear Dye 100X (a dye for DNA detection, 100 times concentrated, also part of the AccuClear kit) For 2 standards + the "sample" tube(s) + 1 control - Distribute the buffer + dye mixture: • Put 190 pL of the mixture into the "standard" tubes, add 1 pL of the 0 ng / pL and 25 ng / pL standards (DNA concentration standard for calibration); • Put 199 pL of the mixture into the "sample(s)" tube(s), add IpL of pool. - Vortex for 2 or 3 seconds and centrifuge the tubes (mix and separate the components by centrifugal force). - Incubate for 5 minutes at room temperature in the dark (let it rest). - Measure the pool using the Denovix (measure the concentration of DNA in the sample with the Denovix device).

[0118] Step 3: Loading NextSeq2000 at 750pM: using the dosing data from the Pool.

[0119] Step 4: Perform equimolar calculations for each preparation based on the number of reads obtained per sample (calculate the quantity of each sample to use for sequencing, based on the number of sequences read).

[0120] Step 5: Make the equimolar mixture (96 to 1 tube).

[0121] The NGS first sequencing assay method according to the invention presents several significant advantages compared to the fluorescence spectrometry assay method.

[0122] This technique does not require the use of specific reagents, which simplifies the process and reduces costs.

[0123] The process involves equivolume sampling of libraries, followed by sequencing and mixing of libraries according to volumes based on the number of reads obtained.

[0124] The crucial aspect of this technique lies in the fact that the mixing method is identical to that used to obtain the final results, thus eliminating potential bias. The only significant variation arises from the pipetting accuracy. An additional advantage of this method is the ability to reliably identify non-viable samples during this initial sequencing, thanks to the use of the same technology.

[0125] Next, the sequencing steps are as follows.

[0126] Step 1: Assay dsDNA High Sensitivity Kit DeNovix (a DeNovix brand kit for the quantification of double-stranded DNA) + TapeStation, Agilent (an Agilent brand instrument used for the analysis of the quality and size of nucleic acid samples) of the equimolar mixture.

[0127] Step 2: Sequencing NovaSeq6000 (a next-generation sequencer produced by Illumina, used for high-throughput sequencing of DNA and RNA, here specified with a sample load at a concentration of 1.6nM) or other ILLUMINA sequencer (a leading brand in genetic sequencing technologies) depending on the number of samples (sequencer selection may depend on the volume and nature of the samples to be analyzed).

[0128] 1.3 Results and general conclusion

[0129] The results are given in the table below, which shows the number of reads for each sample during the second sequencing, following a normalization of the quantity of nucleic acid by first sequencing (according to the invention) or by fluorimetric assay.

[0130] [Tables 1] Sample number Number of reads (after first sequencing) Number of reads (after fluorometric assay) 1 13772773 9770797 2 13863142 10445568 3 14316972 10806639 4 14347603 11275529 5 14589286 11364177 6 14603015 11536654 7 14625673 11588000 8 14704039 11603961 9 14737198 11636106 10 14748153 11929070 11 14809330 12203414 12 14844884 12256330 13 14991006 12635915 14 15018358 13370649 15 15041130 13440753 16 15110883 13474357 17 15116518 13582730 18 15133132 13616744 19 15150322 13767772 20 15187309 13813604 21 15209317 13901171 22 15233587 13905558 23 15255630 13920891 24 15296413 14057539 25 15297248 14170850 26 15309179 14343860 27 15403451 14455614 28 15431354 14607150 29 15433302 14668520 30 15464819 14808004 31 15522601 14934195 32 15545008 14984924 33 15559944 15006023 34 15563610 15061737 35 15619067 15102513 36 15621936 15115873 37 15633836 15203363 38 15657656 15303527 39 15658019 15394915 40 15664827 15461201 41 15743705 15503286 42 15765732 15555775 43 15849137 15614170 44 15901133 15645482 45 15907105 15928181 46 15931352 16010643 47 15954567 16038193 48 15977256 16148859 49 16016330 16192359 50 16035673 16224620 51 16043320 16422787 52 16087154 16433204 53 16090842 16506503 54 16096344 16609753 55 16116518 16651554 56 16135694 16665419 57 16148469 16794291 58 16165091 16932257 59 16173601 16947043 60 16202311 16959275 61 16244331 17010629 62 16255591 17051696 63 16268932 17142186 64 16329770 17202746 65 16337068 17306725 66 16342549 17413147 67 16383932 17420003 68 16444198 17509493 69 16457459 17545337 70 16471521 17693610 71 16548854 17745356 72 16670426 17798230 73 16718021 17840934 74 16765387 17954595 75 16768158 18122301 76 16847326 18566551 77 16956538 18661550 78 16991861 18772935 79 16995405 18949778 80 17057082 19175806 81 17094095 19390025 82 17188459 19575642 83 17242067 19983296 84 17406648 20100141 85 17455527 20486700 86 17499241 20738191 87 17501881 20936298 88 17524801 21618013 89 17715464 22059587 90 18091437 22204819 91 18152815 23029553 92 18166799 23611906 93 18360530 25534745 94 18713888 29423582 Table 1: Comparison of the number of readings obtained

[0131] A graphical representation of this table is given in [Fig.1].

[0132] It is immediately deducible that when the first sequencing assay according to the invention is used, the subsequent number of readings obtained by the second sequencing is much more homogeneous, which improves the quality of the results.

[0133] Therefore, it is proven that the method of normalizing the quantities of nucleic acid by first sequencing according to the invention is superior to the method used in the prior art.

[0134] Example 2: Example of normalization of nucleic acid quantity by first sequencing

[0135] This example illustrates the calculations performed after the first sequencing.

[0136] The results are given in the table below.

[0137] [Tables2] Pu it Identification of Péchant illon Sequence UD I Pre-sequencing reads % first r sequencing age VCO( pi) VEP (pl) Al 55220810203934 CGTTAGGA TT-GTGTAA GGAT 1589792 1.4176 1.102140 19 B1 55220810203981 TTCCATTA CG-AGTCCT TGCG 1128464 1.0063072 19 1.5 1.5527067 38 Cl 5522081020102 TAGTAACGGTA CCA-GCGG 1409475 1.2568986 41 1.5 1.2431392 23 DI 55220810200877 GTAGCCAG GA-CCTCTA AGTA 1890293 1.6856678 57 1.5 0.926932 31 1.5 1,4106610 74 Fl 55220810200535 TATCCTCC AG-CGCAC CAATG 1363232 1,2156614 68 1.5 1,2853084 85 G1 55220810200435 TAAGTCGT TC-TATATC TCGC 1814423 1.6180108 22 1.5 0.9656919 34 H1 55220810201052 TCCGGATT GA-ATAGC AGTCA 2300823 2.0517577 84 1.5 0.7615421 34 A2 55220810202584 ACGTCTTG TT-AGCGG ACGTT 1726801 1.5398739 46 1.5 1.0146934 45 B2 55220810202579 ATGAAGTG CG-GACAT GTGCG 1351686 1.2053653 29 1.5 1.2962874 93 C2 55220510105887 CGATCACT GC-TCAGG TGTGC 76463 0.0681858 43 1.5 22.915314 02 D2 55220810202651 CCTATCGG AA-CTCTG GACAA 1275139 1.1371045 79 1,5 1,3741040 44 E2 55220810203459 CAGAGAGC TT-CAGGA AGGCT 1640228 1,4626725 16 1,5 1,0682500 58 F2 55220810204510 GCAACTTG CG-TGGCG TAAGG 1301784 1.1608652 45 1.5 1.3459787 92 G2 55220810204540 TATGGAGG AC-GTACG TATTC 1049084 0.9355201 43 1.5 1,6701938 61 H2 55220810200998 TGAGATCA GA-ACGGT GCCAA 1146975 1.0228143 95 1.5 1,5276476 44 A3 55220510106832 TCAGCCTA TT-GCCACC TAAT 1631019 1,4544603 95 1.5 1,0742815 73 B3 55220510106880 GTTGTGAG CG-TTCTTG ATCG 1001931 0.8934714 79 1.5 1.748796 73 C3 55220810203893 TCAGTAAC AC-AGAAG ACAGC 882690 0.7871383 75 1.5 1.9850385 26 D3 55220810200882 AAGGCTCA F3 55220510105890 CCGAGCTT AG-AGCTT ACCGG 762286 0.6797681 67 1.5 2.2985777 73 G3 55220510105987 ATCACGCT TC-CTATTC GAAC 711513 0.6344913 69 1.5 2.4626024 49 H3 55220810204562 TAGCTATG CA-GAGCC TGACA 1526526 1.3612788 13 1.5 1.1478177 62 A4 55220510105397 TGTTCCTC AT-AACAC TGTTG 1339543 1,1945368 14 1,5 1,308038 38 B4 55220510105986 CATACCTT CT-TTCCTC TCTT 1005735 0,8968636 94 1,5 1,7421822 41 C4 55220510105975 GCCTTCAA TG-GCTAC AACCG 1238660 1,1045744 48 1,5 1,4145719 21 D4 55220810201017 CTTGACCA GC-CACTTC AGGC 1722237 1,5358040 01 1,5 1,0173824 25 E4 55220810204501 CTACACAC AA-AGTGT CGTAA 695136 0,6198871 88 1,5 2,5206199 31 F4 55220810204028 TAGGCTGA AT-GATCTA GGCG 1358351 1,2113088 39 1,5 1,2899270 19 G4 55220810202395 TCGGAGTC CT-CATCCA GATT 1525461 1,3603291 1,5 1,148619 11 H4 55220810204207 AACATCGC GG-TGTAC CGTCG 1514266 1,3503459 64 1,5 1,1571108 75 A5 55220810204123 GTTGTCTT AC-GCAAT ACTAC 1126099 1,0041982 32 1,5 1,5559676 87 B5 55220810203975 GTGGCAAC TA-ATCCGC TGGA 1133857 1,0111164 25 1,5 1,545321 55 C5 55220510106934 GAGCAGGC AT-GAGGT GGTTG 1318047 1,1753677 67 1,5 1,329371 15 D5 55220810204621 AACGGCAC CT-CGCCTA AGCT 1184272 1,0560739 76 1,5 1,4795365 05 E5 55220810200446 AGTAACCT TG-AAGGA ACCGG 599862 0,5349266 45 1,5 2,9209612 48 F5 55220810203818 TCTCATAA 917 8GTC80GTTCA 23 1,5 1,7798322 49 G5 55220810200723 TGCTTGCC AA-TCGCCT CT AA 614706 0,5481637 75 1,5 2,8504254 98 GTTC317 552220G CTTG 1568046 1,3983042 53 1,5 1,1174249 07 A6 55220810204443 CCAAGTAG AT-CCGTTC TCCT 852336 0,76001702 31 21,01,575 1,01,5702 55220510105852 AAGGTTGG CG-GTCTA ACAGG 874389 0,7797359 63 1,5 2,0038834 62 C6 55220510105382 TGCT9CT14 AGA30 8AC-4AACGT6 92 1,5 3,8392362 08 D6 55220510105396 ACTGTAAC GA-TGTGG ACTTA 1577445 1,4066858 06 1,5 1,1107668 7TT 1,1107668 77 GACC48105 GT-ATGGTT CTTG 433198 0.3863041 04 1,5 4,0447408 72 F6 55220810204341 TTCACCAG AT-GTCATC AACT 1329562 1,183 6563162 55220810202626 ACTTCCAA GG-CAAGA GTAGG 1334023 1,1896143 57 1,5 1,313450 86 H6 55220510106439 CCGAATAT TC-TGGATT GTTC 168883 0,1506013 33 1,5 10,37507 42 A7 55220510106438 CTCTTATC CA-ACCAA CAGAA 235961 0,2104181 06 1,5 7,4256917 72 B7 55220810201051 TCACACGCG GT-CCTGAC GATG 930366 0,8296534 25 1,5 1,8833165 19 C7 55220510106460 CCTCTGTC GT-CTTCAT GCAT 89282 0.079617 18 1.5 19.625161 36 D7 55220810204145 TCTGTTCT CG-GCGAT TCACG 362375 0.3231477 29 1.5 4.8352498 28 E7 55220510107032 GATACTTC AC-AAGGC TGCTC 1447586 1,2908841 07 1,5 1,210410 75 F7 55220510106930 AGTGCTGA TA-TGGTA ATCGA 1524713 1,3596620 71 1,5 1,1491826 04 G7 55220510106927 ATCCTTCG GT-AATTG GACTG 1571219 1,4011337 74 1,5 1,1151683 22 H7 55220510107048 GACAACGA TT-CCAAGC CTCT 1884088 1,6801345 51 1,5 0,9299850 41 A8 55220510107043 GAACCGGT AG-TGAGA GCCTG 1446743 1,2901323 62 1.5 1,2111160 42 B8 55220510106957 AGCAATGA GC-GAGAG CGAAC 1071939 0,9559010 78 1.5 1.6345833 64 C8 55220510106860 CAAGACTC CA-ATTAGT CCGA 1473471 1,3139670 43 1.5 1,1891470 25 D8 55220510106846 ACCGTGTA GG-GAAGA TCTCG 1463726 1,3052769 44 1,5 1,197063 97 E8 55220810202557 AGGCACAG GT-GTCCG GTTAT 1076616 0,9600717 91 1,5 1,627482 46 F8 55220810200764 CGACAGAT CG-TAACT ACACG 644533 0.5747619 87 1.5 2.7185165 95 G8 55220510105990 ACGCGACA AC-ACCTAT GTTC 1471635 1.3123297 91 1.5 1.1906305 95 H8 55220810201069 ACTTGCGT TA-CGATGT TAGA 1451738 1.2945866 51 1.5 1.2069489 51 A9 55220810202599 CACCACTC AT-TGCCAC CGTT 585042 0.5217109 17 1.5 2.9949536 21 B9 55220810202562 CTTCGTAA CT-ACACC GTCCT 1116240 0.9954064 73 1.5 1.5697105 07 C9 55220810204081 CAGTATTC GG-CAGGT CACAG 1052717 0.9387598 69 1,5 1,6644299 05 D9 55220510105991 CAGTCTGG AC-TTGTTA CAGC 1020038 0,9096183 87 1,5 1,7177533 15 E9 55220810202362_5522081 0204554 TACCGTTC TA-GACGT CCGTA 1017146 0.9070394 47 1.5 1.7226373 17 F9 55220810202532 GTGTCCAC AG-GCTCCT TAGG 819091 0.7304239 98 1.5 2.1391684 88 G9 55220810204026 TTACGACT GT-CTGGCC8040404767 53 1.5 2.1777927 63 H9 55220810202332 GACGCGAA TG-TACAG ATGAG 1073401 0.9572048 16 1.5 1.6323570 19 A10 552 AGGAC010202013 CCTTC 902693 0.8049760 41 1.5 1.941051 56 B1O 55220510106929 AGCTCAGG AA-GCAGT AGAGA 1740560 1.5521435 28 1.5 1.068 CIO 7236 55220510107053 GATAGGCG GT-TTGTTC GGTT 1455951 1.2983435 91 1.5 1.2034564 74 D1O 55220510106858 AGTAGGAA GT-GTGGC14 GAT 134728 44 1.5 1.6840666 33 E1O 55220510106913 CATGTTGT AG-CGTTG CATGG 1178546 1.0509678 18 1.5 1.4867248 76 F1O 552820204538 CACATT TC-TACACC ATTC 662870 0.5911139 98 1.5 2.6433141 59 G1O 55220810202182 GCAGCTCG TA-ACGGC ATATA 605700 0.5401326 78.28 8177 H 55220810201057 GTTCAGAC GG-ATCTAT CGAG 1241489 1.1070972 08 1.5 1.4113485 15 Ail 55220810200934 TCCTGGAA GT-CTCTTG TGTT 914907 0.8158678 69 1.5 1.915138 54 Bll 55220810204139 GCATTGTT AG-ACCGA TTGCG 740721 0.6605375 89 1,5 2,3654974 76 CIL 55220510106956 GACCTACA GC-TCTACG CAAC 1522046 1,3572837 75 1,5 1,1511962 56 DU 55220810204417 CACCGACG TA-GATCA CTCTA 861090 0.7678765 86 1.5 2.0348321 97 Eli 55220810204377 CTCTCACC TT-GCTGCG TCTT 1201803 1.0717072 37 1.5 1.4579541 37 Fil 55220810200786 CTCGTTCA TT-TCGAGA AGTT 1459242 1.3012783 39 1.5 1.2007423 42 Gll 55220810204405 TGGTGGCA AG-CTCAG TTGCG 944553 0.8423046 75 1.5 1.8550294 76 Hll 55220810200631 GATTGCTT GA-ATTGC GGAGC 1704921 1,5203624 67 1,5 1,0277154 52 A12 55220810200761 CCGTTAAG GT-CGGAA GTTAC 1572895 1,4026283 46 1.5 1,1139800 54 B12 55220810203382 TGCTGAGA GG-TAGTC GTGAG 1389735 1,2392955 06 1.5 1,2607969 55 C12 55220810204151 TTGTCACT TG-GCCGTT GGTT 1165175 1.0390442 36 1.5 1.5037858 32 D12 55220810203873 GCTGTTAT GT-TACAG GCAGG 541383 0.4827780 25 1.5 3.2364770 53 E12 55220510107017 GCAGCAGT TG-CTGCA GCGTA 1293075 1.1530989 98 1.5 1.3550441 05 F12 55220510107183 GCAGATCA AT-TCGCA ACATT 1440799 1.2848318 03 1.5 1.2161124 88 G12 55220810204033 TGGTTCAC GG-CAGAA CGTCG 1384598 1.2347145 89 1.5 1.265474 64 H12 55220810204003 TCGACCGC AT-GTGTCC TATT 2331379 2.0790060 82 1.5 0.7515610 53 Table 2: Normalization by first sequencing according to the invention

[0138] It should be noted that normalization allows for very strong adjustment of the volumes to be taken VEP, which vary, in this series, from 0.75 pl (sample 55220810204003, well H12) to 19.62 pl (sample 55220510106460, well C7).

Claims

Demands

1. A method for normalizing the quantity of nucleic acid between at least 2 samples, characterized in that it comprises the following steps: - having at least 2 samples, each having a volume V comprising nucleic acid; - taking a known volume VCO from each of at least 2 samples; - from said volumes VCO taken, proceeding with a first sequencing allowing the obtaining, for each of the volumes VCO, of a number of reads; - depending on the number of reads of each volume VCO, calculating a sample volume to be taken VEP for each sample such that each VEP corresponds to an identical theoretical number of reads NLT.

2. The method according to claim 1, characterized in that the nucleic acid is selected from the group consisting of ribonucleic acid (RNA) and in particular messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), small interference RNA (siRNA), microRNA (miRNA) and / or long non-coding RNA (lncRNA), deoxyribonucleic acid (DNA) and in particular chromosomal DNA (chDNA), mitochondrial DNA (mDNA), plasmid DNA (pDNA), circular DNA (cDNA), linear DNA (1DNA) and / or recombinant DNA (rDNA), and mixtures thereof.

3. A method according to any one of the preceding claims, characterized in that the number of samples is between 2 and 1000, preferably between 2 and 950, preferably between 2 and 900, preferably between 2 and 850, preferably between 2 and 800, preferably between 2 and 750, preferably between 2 and 700, preferably between 2 and 650, preferably between 2 and 600, preferably between 2 and 550, preferably between 2 and 500, preferably between 2 and 450, preferably between 2 and 400, preferably between 2 and 350, preferably between 5 and 350, preferably between 20 and 350, preferably between between 50 and 300, preferably between 50 and 200, preferably between 50 and 100, preferably between 75 and 125, preferably between 90 and 110, preferably between 95 and 105.

4. Method according to any one of the preceding claims, characterized in that said known volume VCO is between 0.5 pl and 2.5 pl, preferably between 0.6 pl and 2.4 pl, preferably between 0.7 pl and 2.3 pl, preferably between 0.8 pl and 2.2 pl, preferably between 0.9 pl and 2.1 pl, preferably between 1.0 pl and 2.0 pl, preferably between 1.1 pl and 1.9 pl, preferably between 1.2 pl and 1.8 pl, preferably between 1.3 pl and 1.7 pl, preferably between 1.4 pl and 1.6 pl, preferably 1.5 pl.

5. Method according to any one of the preceding claims, characterized in that said VCO volume is the same for each sample.

6. Method according to any one of the preceding claims, characterized in that said VEP volume is between 0.5 pl and 7.5 pl, preferably between 0.6 pl and 7.2 pl, preferably between 0.7 pl and 6.9 pl, preferably between 0.8 pl and 6.6 pl, preferably between 0.9 pl and 6.3 pl, preferably between 0.9 pl and 6.0 pl, preferably between 0.9 pl and 5.7 pl, preferably between 0.9 pl and 5.4 pl, preferably between 0.9 pl and 5.1 pl, preferably between 0.9 pl and 4.8 pl.

7. Sequencing method, characterized in that it comprises: - carrying out a method for normalizing the quantity of nucleic acid between at least two samples according to any one of the preceding claims; - taking, for each at least 2 samples, said VEP volume; - carrying out a second sequencing.

8. A method for processing several series of samples, including a first series of samples and at least one other series of samples, characterized in that it comprises the following steps: - the implementation of a method for standardizing the quantity of nucleic acid between at least two samples from said first series of samples according to any one of claims 1 to 6; - the taking, for each at least 2 samples from the said first series of samples, of the said VEP volume; - the performance of a second sequencing; characterized in that the performance of said second sequencing takes place simultaneously, and preferably on the same sequencing plate, with the performance of a first sequencing of a method for normalizing the quantity of nucleic acid between at least two samples belonging to at least one other series of samples according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • System and method for ligand-limited normalizing polymerase chain reaction (LLN-PCR)

    US20230399679A1

  • Balanced capture probes and methods of use thereof

    WO2018136526A1

  • Improvements in variant detection

    WO2019170773A1

  • Methods for detection of donor-derived cell-free DNA

    WO2021243045A1

  • Methods for detection of donor-derived cell-free DNA in transplant recipients of multiple organs

    WO2022182878A1