Molecular sequence tags used for sample identification and quantification of nucleic acids within samples.
By designing the nucleic acid construct E1-(Z1)mY-(Z2)n-E2, the problems of sample addition errors and confusion in high-throughput sequencing technology were solved, enabling sample quality detection and quantification of contamination, and improving the accuracy and reliability of detection results.
Patent Information
- Application Number
- CN202410611434.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2044-05-16
AI Technical Summary
In existing high-throughput sequencing technologies, the sample transfer and barcode primer addition processes are prone to sample addition errors or confusion, affecting the accuracy of detection results. Furthermore, there is a lack of precise methods for quantifying nucleic acid information concentration or pathogen load.
Design a nucleic acid construct with a 5' to 3' structure E1-(Z1)mY-(Z2)n-E2, where E1 and E2 are restriction enzyme sequences, Z1 and Z2 are recognition sequences, and Y is a quantification sequence, for use in the preparation of quality control detection reagents or kits, and to identify and quantify sample contamination and concentration through sequencing results analysis.
It enables quality detection and contamination identification of samples, ensuring the accuracy of the experimental process, and accurately quantifies the concentration of nucleic acid substances and the load of pathogenic microorganisms in samples, thereby improving the reliability of the test results.
Smart Images

Figure BDA0004843756230000081 
Figure BDA0004843756230000082 
Figure BDA0004843756230000091
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology, and more specifically to molecular sequence tags for sample identification and quantification of nucleic acid substances within samples. Background Technology
[0002] High-throughput sequencing, also known as "next-generation" sequencing, is a technology capable of simultaneously sequencing hundreds of thousands to millions of DNA molecules. This technology currently has wide applications in various fields, including the discovery of disease-related mutations and variations, the elucidation of disease genetic mechanisms, disease prediction and prevention, evolutionary biology research, the diagnosis of infectious diseases, and drug administration.
[0003] High-throughput gene expression analysis can accurately determine the expression levels of individual genes, enabling relative quantitative analysis. This is significant for studying changes in gene expression in organisms under different physiological states or comparing differences in gene expression among different samples. High-throughput microbiome research can identify hundreds to tens of thousands of known and unknown microbial populations in different samples, analyzing the composition and structure of the microbial community. This is crucial for studying the interaction between microorganisms and the host, and the role of microorganisms in disease occurrence and development. However, currently, there is no method to accurately quantify the specific concentration of the nucleic acid being detected or the load of pathogenic microorganisms.
[0004] The high-throughput sequencing experimental workflow includes: 1. Sample collection: Samples can come from various sources, such as oral swabs, saliva, and blood. Strict adherence to project requirements for sample preservation and transportation is essential. 2. Sample processing and library preparation: After sample collection, a series of processing steps are required, including DNA or RNA extraction, fragmentation, and library construction. 3. Sequencing: This is the core step of the high-throughput sequencing experiment. During this process, the sequencer performs large-scale parallel sequencing of the target library to obtain massive amounts of sequence data, which is then analyzed. From extraction to library preparation, the entire high-throughput sequencing experimental process involves numerous steps and complex operations, including many sample transfer and barcode primer addition processes. These steps can easily lead to sample addition errors or confusion, thus affecting the accuracy of the test results.
[0005] The present invention relates to a molecular sequence tagging technology for sample identification and quantification of nucleic acid substances within samples. Summary of the Invention
[0006] The purpose of this invention is to provide a molecular sequence tag for sample identification and quantification of nucleic acid substances within a sample.
[0007] In a first aspect, the present invention provides a nucleic acid construct having a structure of formula (I) from 5' to 3':
[0008] E1-(Z1)mY-(Z2)n-E2(I)
[0009] In the formula,
[0010] E1 and E2 are each an independent restriction enzyme sequence;
[0011] Z1 and Z2 are the recognition sequences;
[0012] m and n are each independent positive integers from 2 to 10;
[0013] Y represents a quantitative sequence.
[0014] In another preferred embodiment, Y has a nucleotide sequence selected from any of SEQ ID NO: 1 to 6.
[0015] In another preferred embodiment, the identification sequence is used to identify the sequence type in the sample to be tested.
[0016] In another preferred embodiment, the length of the nucleic acid construct is 50–500 nt, more preferably 50–400 nt, and even more preferably 50–200 nt.
[0017] In another preferred embodiment, the length of Y is 10–40 nt, more preferably 15–35 nt, and even more preferably 20–30 nt.
[0018] In another preferred embodiment, the lengths of Z1 and Z2 are 5–30 nt, more preferably 10–25 nt, and even more preferably 10–20 nt.
[0019] In another preferred embodiment, the identification sequence does not have the following three characteristics:
[0020] (r1) contains three or more consecutive identical bases;
[0021] (r2) If any one base in the sequence is changed to any of the other three bases, it will be exactly the same as another recognition sequence;
[0022] (r3) Palindromic sequence.
[0023] In another preferred embodiment, the enzyme digestion sequence is selected from the group consisting of BamHI, PmeI, MreI, or combinations thereof.
[0024] In another preferred embodiment, the length of E1 or E2 is independently 5 to 15 nt, more preferably 5 to 10 nt, and even more preferably 6 to 8 nt.
[0025] In another preferred embodiment, m and n are each independently a positive integer from 2 to 8, preferably from 2 to 5, and even more preferably from 2 to 3.
[0026] In another preferred embodiment, the nucleic acid construct has a nucleotide sequence selected from any of SEQ ID NO:7 to 102.
[0027] In a second aspect, the invention provides the use of nucleic acid constructs as described in the first aspect for the preparation of quality control detection reagents or kits.
[0028] In a third aspect, the present invention provides a kit comprising (i) a container and (ii) a nucleic acid construct as described in the first aspect of the present invention.
[0029] In a fourth aspect, the present invention provides a quality control method, comprising the steps of:
[0030] (s1) The nucleic acid construct described in the first aspect of the present invention is added to the sample to be tested, and then processed and sequenced;
[0031] (s2) Analyze the sequencing results.
[0032] In another preferred embodiment, the analysis includes comparing the sequencing results with the identified sequence.
[0033] In another preferred embodiment, the method is non-diagnostic and non-therapeutic.
[0034] In another preferred embodiment, the method is in vitro.
[0035] In another preferred embodiment, the criterion for determining the method is:
[0036] (p1) If the sequencing results contain only one identification sequence and are consistent with the type added to the sample to be tested, it indicates that the experimental process is normal;
[0037] (p2) Sequencing results show multiple identification sequences but include the types of samples added to the test sample: This indicates that there is contamination in the experimental process;
[0038] (p3) If the sequencing results show that the identification sequence is inconsistent with the type of sample added to the test sample, it means that the wrong sample was taken or the experiment failed.
[0039] In another preferred embodiment, the method further includes quantitative analysis.
[0040] In another preferred embodiment, the quantitative analysis is obtained by transforming a quantitative sequence.
[0041] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description
[0042] Figure 1 The structure diagram of the molecular sequence tag is shown.
[0043] Figure 2 The sequencing validation results of the molecular sequence tag plasmid are shown.
[0044] Figure 3 A flowchart of the sequence tag implementation scheme is shown.
[0045] Figure 4 The document quality control chart is displayed. Detailed Implementation
[0046] Through extensive and in-depth research, and after numerous experiments and screenings, the inventors have unexpectedly discovered a nucleic acid construct for the first time. This nucleic acid construct has a structure of formula (I) from 5' to 3', in which E1 and E2 are each independently an enzyme digestion sequence; Z1 and Z2 are recognition sequences; m and n are each independently positive integers from 2 to 10; and Y is a quantitative sequence.
[0047] E1-(Z1)mY-(Z2)n-E2(I).
[0048] Experiments show that the nucleic acid constructs of this invention can perform quality testing on test samples and identify whether the samples are contaminated, which is beneficial for quality control of the experimental environment. Furthermore, the nucleic acid constructs of this invention can also quantitatively detect contaminants in test samples. Based on these findings, this invention was completed.
[0049] the term
[0050] To facilitate understanding of the invention, certain technical and scientific terms are specifically defined below. Unless otherwise expressly defined herein, all other technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. Before describing the invention, it should be understood that the invention is not limited to the specific methods and experimental conditions described, as such methods and conditions can vary. It should also be understood that the terminology used herein is intended only to describe particular embodiments and is not intended to be restrictive; the scope of the invention will be limited only by the appended claims.
[0051] As used herein, the term “comprising” or its variations such as “including” or “comprising” are understood to include the said element or component without excluding other elements or other components.
[0052] Nucleic acid constructs
[0053] As used herein, the terms "nucleic acid construct of the present invention" and "tag sequence of the present invention" are used interchangeably and both refer to structures of formula (I) from 5' to 3':
[0054] E1-(Z1)mY-(Z2)n-E2(I)
[0055] In the formula,
[0056] E1 and E2 are each an independent restriction enzyme sequence;
[0057] Z1 and Z2 are the recognition sequences;
[0058] m and n are each independent positive integers from 2 to 10;
[0059] Y represents a quantitative sequence.
[0060] In another preferred embodiment, Y has a nucleotide sequence selected from any of SEQ ID NO: 1 to 6.
[0061] In another preferred embodiment, Y has a nucleotide sequence selected from any of SEQ ID NO: 1 to 6.
[0062] In another preferred embodiment, the identification sequence is used to identify the sequence type in the sample to be tested.
[0063] In another preferred embodiment, the length of the nucleic acid construct is 50–500 nt, more preferably 50–400 nt, and even more preferably 50–200 nt.
[0064] In another preferred embodiment, the length of Y is 10–40 nt, more preferably 15–35 nt, and even more preferably 20–30 nt.
[0065] In another preferred embodiment, the lengths of Z1 and Z2 are 5–30 nt, more preferably 10–25 nt, and even more preferably 10–20 nt.
[0066] In another preferred embodiment, the identification sequence does not have one or more of the following characteristics:
[0067] (r1) contains three or more consecutive identical bases;
[0068] (r2) If any one base in the sequence is changed to any of the other three bases, it will be exactly the same as another recognition sequence;
[0069] (r3) Palindromic sequence.
[0070] In another preferred embodiment, the enzyme digestion sequence is selected from the group consisting of BamHI, PmeI, MreI, or combinations thereof.
[0071] In another preferred embodiment, the length of E1 or E2 is independently 5 to 15 nt, more preferably 5 to 10 nt, and even more preferably 6 to 8 nt.
[0072] In another preferred embodiment, m and n are each independently a positive integer from 2 to 8, preferably from 2 to 5, and even more preferably from 2 to 3.
[0073] In another preferred embodiment, the nucleic acid construct has a nucleotide sequence selected from any of SEQ ID NO:7 to 102.
[0074] Furthermore, the nucleic acid constructs of the present invention can be linear or circular. The nucleic acid constructs of the present invention can be single-stranded or double-stranded. The nucleic acid constructs of the present invention can be DNA, RNA, or a DNA / RNA hybrid.
[0075] Identification Sequence
[0076] In this invention, the identification sequence is used to identify the sequence type in the sample to be tested. The identification sequence is composed of four bases: A, T, C, and G. The design of the identification sequence mainly includes two steps. The first step is to generate random sequences. Taking an identification sequence length of 18 nt as an example, 6,871,947,6736 combinations are randomly generated. The second step is to screen the sequences, removing those that do not meet the following criteria: 1. The sequence contains three or more consecutive identical bases; 2. If any base in the sequence is changed to any of the other three bases, it will be completely identical to another sequence; 3. Palindromic sequences.
[0077] quantitative sequence
[0078] In this invention, quantitative sequences are sequences obtained by cutting and filtering the entire human genome sequence. The quantitative sequence design mainly includes three steps: First, human genome cutting, which involves cutting the entire human genome into sequences of equal length; Second, screening out sequences with 10, 100, 1000, or 10,000 identical sequences from the cut sequences; Third, comparing the screened sequences with the NR database, selecting human-specific sequences, and discarding non-specific sequences.
[0079] The quantitative sequence is a 17-30 nt segment of the human genome, which is widely distributed in the human genome. By adding a tag sequence with a known copy number to each sample, and then converting it with the number of reads obtained from the corresponding segment of the human genome, the copy number of the human genome in the sample can be calculated. Then, by analyzing quantitative sequences with different copy numbers on the human genome, quantitative analysis of other species in the sample can be performed to analyze the number of pathogenic microorganisms or tumor variants in the sample.
[0080] The main advantages of this invention include:
[0081] (1) The nucleic acid constructs of the present invention can perform quality detection on the test samples and identify whether the test samples are contaminated, which is beneficial to the quality control of the experimental environment.
[0082] (2) The method of the present invention can also quantitatively detect pollutants in the sample to be tested.
[0083] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are by weight.
[0084] Example 1: Design of Tag Sequence
[0085] 1.1 Tag Sequence
[0086] A schematic diagram of the structure of a molecular sequence tag is shown below. Figure 1 As shown, the total length is 60nt-104nt, consisting of 2-4 "identification sequences" and 1 "quantification sequence"; the "identification sequence" is a known named sequence of 10-18nt.
[0087] By comparing the raw sequencing reads with the fixed-named identification sequences, the types of identification sequences in the sample are counted. This is then compared with the types of tags added before sample processing. If the analysis results contain only one identification sequence that matches the type added in the experiment, the experiment is considered normal. If the analysis results contain multiple identification sequences but include the type added in the experiment, the experiment is considered contaminated and needs to be repeated. If the analysis results contain identification sequences that do not match the type added in the experiment, the experiment was considered to have failed or the wrong sample was taken.
[0088] "Quantitative sequencing" refers to a 17-30 nt segment of the human genome sequence, which is widely present in the human genome. By adding a tag sequence with a known copy number to each sample, and then converting it with the number of reads obtained from the corresponding segment of the human genome, the copy number of the human genome in the sample can be calculated. Then, by analyzing quantitative sequences with different copy numbers on the human genome, quantitative analysis can be performed on other species in the sample to analyze the number of pathogenic microorganisms or tumor variants in the sample.
[0089] 1.2 Sequence Design Methods and Selection Principles
[0090] The identification sequence is composed of four bases: A, T, C, and G. The design of the identification sequence mainly includes two steps. First, random sequences are generated. Taking an identification sequence length of 18 nt as an example, 6,871,947,6736 combinations are randomly generated. Second, sequences are filtered out, eliminating sequences that do not meet the following criteria: 1. The sequence contains three or more consecutive identical bases; 2. Any base in the sequence, when changed to one of the other three bases, would be identical to another sequence; 3. Palindromic sequences.
[0091] 1.3 Principles of Quantitative Sequence Design
[0092] Quantitative sequencing (QS) is a process of cutting and filtering the entire human genome sequence. QS design mainly involves three steps: First, cutting the human genome into equal-length sequences; second, selecting sequences with 10, 100, 1000, or 10,000 identical sequences from the cut sequences; and third, comparing the selected sequences with an NR database to select human-specific sequences and discard non-specific sequences.
[0093] 1.4 Combination of Identification and Quantification Sequences
[0094] according to Figure 1 The molecular sequence tag structure diagram combines the identification sequence and the quantitative sequence into a sequence with a length of 98nt-104nt, and assigns a unique number to each tag.
[0095] Example 2: Production and Preparation Process of Serial Tags
[0096] The production and preparation of sequence tags are as follows:
[0097] 2.1 Sequence Synthesis: Sequences with different numbers were synthesized using chemical synthesis methods according to the identified sequence structure and the quantitative sequence structure.
[0098] 2.2 Sequence amplification and production: Different types of sequences are ligated with the Escherichia coli PUC57 plasmid sequence. The ligated recombinant plasmids need to be transformed into competent cells for expression and positive clones are screened. Bacteria containing the target plasmid are cultured to amplify the plasmid in large quantities.
[0099] 2.3 Purification and quality control: Lyse bacterial cells to release plasmid DNA, remove impurities and host DNA, and increase the concentration of plasmid DNA.
[0100] 2.4 Quality Control: The plasmid concentration was quantified using Qubit sequencing, and the synthesized plasmid was sequenced using first-generation sequencing to determine if the plasmid matched the synthesized sequence. For example... Figure 2 As shown, the sequence of the first-generation sequencing result is completely consistent with the sequence tag, indicating that the identification sequence plasmid was successfully synthesized.
[0101] 2.5 Tag colonization: The synthesized tag plasmid powder was diluted 10,000 times with 1×TE, and qualified tag plasmids were quantified on a digital PCR platform using the Taqman probe method.
[0102] 2.6 Tag Dilution Indication: Dilute the tag plasmid to 7×10 using 1×TE. 4 Prepare copies / ml as the working solution, affix a clear and conspicuous label, and store in aliquots at -20℃.
[0103] Example 3: Implementation method of sequence tag
[0104] The flowchart of the sequence tag implementation scheme is as follows: Figure 3 As shown.
[0105] 3.1 Sample Processing
[0106] Break the cotton swab and place it in a 2mL centrifuge tube. Add 1mL of buffer GA. Incubate the sample on a metal bath at 56°C for 10 minutes. Discard the swab and discard some of the supernatant until the liquid level in the centrifuge tube is 500μL. Add 10μL of 7×10⁻⁶ buffer. 5 The molecular tag (copies / mL) was recorded and its sequence number was added to the experimental logbook. Nucleic acid extraction was performed using an oropharyngeal swab extraction kit. Nucleic acid concentration was quantified using a Qubit fluorometer, and the concentration was recorded.
[0107] 3.2 Library Preparation
[0108] Library construction was performed using an enzyme digestion library construction kit. The quality of the constructed libraries was evaluated using concentration and length distribution assays.
[0109] 3.3 Sequencing
[0110] like Figure 4 As shown, after the libraries passed quality control, they were sequenced using an Illumina Noveseq 6000. The sequencing mode was SE50, and 25M reads were sequenced for each library.
[0111] 3.4 Data Analysis
[0112] First, low-quality data is removed from the original data; then, the filtered data is subjected to label comparison analysis, species annotation analysis, species read and read ratio statistics, and species copy number calculation.
[0113] The copy number and correction formula are shown in Equation (II).
[0114] Copy species = j × number of species reads + (k × number of molecular tag reads - 7 × 10) 5 (II)
[0115] The formula for calculating the K value is:
[0116]
[0117] in,
[0118] Copy species:
[0119] Species Read Count: The number of reads measured for a particular species in the sample;
[0120] Molecular tag read count: The number of reads detected for molecular tags in the sample;
[0121] k: Calculated from quantitative tag sequence sequencing reads with copy numbers of 2, 10, 100, 1000, 10000, and 101056.
[0122] Example 4: Validation of Molecular Tag and Quantitative Tag Functions
[0123] 4.1 Sample Information and Methods
[0124] Using normal human pharyngeal swab samples, pathogenic microorganisms with known copy numbers were added according to Table 1, and the molecular tag QQbar_A0096 (SEQ ID NO:102) was added to serve as validation samples.
[0125] Table 1
[0126] Sample Name Sample type Pathogen count (CFU) Genome size Mb T1 oropharyngeal swab / / Listeria monocytogenes Standard products 140000000 2.944 Staphylococcus aureus Standard products 580000000 2.7 Candida albicans Standard products 2500000 14.3 Aspergillus niger Standard products 6600000 34.0
[0127] Sample processing, library preparation, sequencing, and data analysis were performed according to the method shown in Example 3.
[0128] 4.2 Results
[0129] The results are shown in Table 2.
[0130] Table 2
[0131]
[0132] The error-proof label detection and label addition were consistent, and the experimental operation was normal; the measured copy number of the reference sample deviated from the theoretical copy number by less than one order of magnitude, which can truly reflect the content of pathogens in the sample.
[0133] Example 5: Molecular tagging for contamination identification
[0134] 5.1 Sample Information and Methods
[0135] Using normal population pharyngeal swab samples, error-proofing labels were added according to the proportions shown in Table 3a, and sample processing, library preparation, sequencing, and data analysis were performed as shown in Example 3.
[0136] Table 3a
[0137]
[0138] 5.2 Results
[0139] The results are shown in Table 3b. Black represents positive (strong detection); light gray represents positive (weak detection); and diagonal lines represent negative (no detection).
[0140] Table 3b
[0141]
[0142] Within the detection range of the kit, contamination can be detected in cases of 1 / 10 or 1 / 100. However, contamination cannot be detected if the contamination exceeds the detection limit of the kit.
[0143] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.
[0144] Appendix Table 1: Correspondence between different copy numbers and quantitative sequences
[0145]
[0146]
[0147] Appendix 2: Exemplary Sequence Tags
[0148]
[0149]
[0150]
[0151]
Claims
1. A nucleic acid construct, characterized in that, The nucleic acid construct has a length of 50–500 nt, and the structure of the nucleic acid construct from 5' to 3' is shown in formula (I): E1-(Z1)mY-(Z2)n-E2(I) In the formula, E1 and E2 are each an independent restriction enzyme sequence; Z1 and Z2 are identification sequences, which are used to identify the types of sequences in the sample to be tested. m and n are each independent positive integers from 2 to 10; Y represents a quantitative sequence, which is a human-specific sequence obtained by cutting and filtering the whole human genome sequence. The identification sequence does not have the following three characteristics: (r1) contains three or more consecutive identical bases; (r2) If any one base in the sequence is changed to any of the other three bases, it will be exactly the same as another recognition sequence; (r3) palindromic sequence; The nucleotide sequences of the nucleic acid constructs are shown in either SEQ ID NO:7-13 or 102.
2. The use of the nucleic acid construct as described in claim 1, characterized in that, Test reagents or kits used for quality control.
3. A reagent kit, characterized in that, The kit includes (i) a container and (ii) the nucleic acid construct as described in claim 1.
4. A quality control method, characterized in that, Including the following steps: (s1) The nucleic acid construct as described in claim 1 is added to the sample to be tested, and then processed and sequenced; (s2) Analyze the sequencing results.
5. The method as described in claim 4, characterized in that, The judgment criteria for the method are as follows: (p1) If the sequencing results contain only one identification sequence and are consistent with the type added to the sample to be tested, it indicates that the experimental process is normal; (p2) Sequencing results show multiple identification sequences but include the types of samples added to the test sample: This indicates that there is contamination in the experimental process; (p3) If the sequencing results show that the identification sequence is inconsistent with the type of sample added to the test sample, it means that the wrong sample was taken or the experiment failed.
6. The method as described in claim 4, characterized in that, The analysis includes comparing sequencing results with identified sequences.
Citation Information
Patent Citations
Detection method of cross-contamination between samples in next-generation sequencing
JP2019131539A