Probe group and kit for capturing full-length genome of respiratory syncytial virus
By designing full-length genome probe sets and second-generation sequencing technologies covering all types of RSV, the problems of low sequencing accuracy and high cost in the prior art are solved, and efficient and low-cost full-length genome capture and mutation monitoring of respiratory syncytial virus are achieved.
Patent Information
- Application Number
- CN202510451043.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-01
AI Technical Summary
When capturing the full-length genome of respiratory syncytial virus, the prior art has problems such as low sequencing accuracy, high cost and low throughput, especially when the virus isolation rate is low, it is difficult to obtain accurate whole genome sequences.
A probe set was designed, including probes with nucleotide sequences such as those shown in SEQ ID NO: 1 to SEQ ID NO: 900, combined with streptavidin-labeled magnetic beads and library hybridization reaction reagents, and efficient capture and second-generation sequencing of the full-length genome of the respiratory syncytial virus were achieved through hybridization and PCR amplification.
It improves the accuracy and specificity of the detection, reduces costs, and can quickly obtain the full-length genome sequence of the respiratory syncytial virus, supporting viral mutation monitoring and vaccine development.
Smart Images

Figure CN120230885A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of virus detection, and particularly relates to a probe set and a kit for capturing the full-length genome of respiratory syncytial virus. Background Art
[0002] Respiratory syncytial virus (RSV) is one of the primary pathogens causing acute lower respiratory tract infections in infants and young children under 5 years old. The virus is mainly transmitted through contact of the nasopharyngeal or ocular mucosa with virus-containing secretions or contaminants. Direct contact is the most common route of transmission, but droplets and aerosols can also cause transmission. HRSV is highly contagious and often causes outbreaks in specific places and populations, such as the outbreak caused by the B subtype in mother and baby care institutions. Vaccine research and development is a hot field in RSV research, and vaccine research and development requires the basic support of virus genome data. Therefore, strengthening the monitoring of RSV virus genome variation is particularly important for the monitoring of RSV, the identification of epidemics, and the research and evaluation of vaccines, antibodies, drugs, etc. Therefore, accurately obtaining RSV full-genome data has important public health significance for the identification and prevention and control of human respiratory syncytial virus infection.
[0003] Currently, for the methods of obtaining the full-length RSV genome sequence, the main methods are as follows: 1) RT-PCR method: After designing specific primers for specific amplification of a certain fragment length, Sanger sequencing is used, and the full-length sequence is obtained after splicing. The accuracy is high, and a high virus load is required, but the cost is high and the throughput is low; 2) Next-generation sequencing: By using universal primers to sequence all nucleic acids of the sample, the cost is reduced, a relatively high accuracy is maintained, and the sequencing time is reduced. However, the read length is short, affected by the sequencing depth, and the subsequent output data volume is huge, requiring high requirements for data analysis; 3) Third-generation sequencing: The biggest feature is single-molecule sequencing. The sequencing process does not require PCR amplification, and it has the characteristics of long sequencing read length, low throughput, low accuracy, and high cost. It can be seen that the higher the virus nucleic acid concentration and the more single the category, the better the sequencing result. The isolation rate of RSV is very low, bedside inoculation is required, and often no virus strain can be obtained. Nucleic acids need to be directly extracted from clinical specimens such as nasopharyngeal swabs and then directly sequenced. Then the nucleic acid content and purity are lower than those of microorganisms such as viruses and bacteria that can be cultured, affecting the sequencing result. Due to the existence of variant sites in each type of RSV, the sequencing results of the traditional first-generation sequencing method for the RSV genome are often inaccurate. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a probe set for capturing the full-length genome of respiratory syncytial virus, which is designed for all variant sequences of RSV and can cover the full-length genomes of all types of RSV, and the detection results are accurate and reliable.
[0005] The present invention provides a probe set for capturing the full-length genome of respiratory syncytial virus, and the probe set includes probes with nucleotide sequences as shown in SEQ ID NO: 1 to SEQ ID NO: 900.
[0006] Preferably, one end of the probe is modified with biotin.
[0007] The present invention provides a kit for sequencing the full-length genome of respiratory syncytial virus, which includes the probe set, streptavidin-labeled magnetic beads, and library hybridization reaction reagents.
[0008] Preferably, the library hybridization reaction reagents include at least one of the following reagents: TargetSeq HybBuffer v2, Hyb Human Block, Blocking Oligo, and RNase enzyme inhibitor.
[0009] Preferably, the kit further includes at least one of the following: RNA fragmentation reagent, reverse transcription reaction reagent, cDNA second-strand synthesis reaction reagent, adapter ligation reagent, purification reagent, PCR pre-reaction reagent, and post-capture PCR amplification reaction reagent.
[0010] Preferably, the PCR pre-reaction reagent includes PCR Master Mix containing UDG enzyme and / or UDI primer;
[0011] The post-capture PCR amplification reaction reagent includes PCR primers and / or post-capture PCR reaction premix.
[0012] The present invention provides the application of the probe set or the kit in sequencing the full-length genome of respiratory syncytial virus.
[0013] Preferably, the sequencing includes next-generation sequencing.
[0014] The present invention provides a method for next-generation sequencing of the full-length genome of respiratory syncytial virus based on the probe set, which includes the following steps:
[0015] Hybridize the respiratory syncytial virus genomic library of the sample to be tested with the probe set to obtain a hybridization product;
[0016] Isolate the DNA fragment hybridized with the probe from the hybridization product;
[0017] Perform post-capture PCR amplification using the DNA fragment hybridized with the probe as a template, perform next-generation sequencing on the obtained PCR product, and perform data analysis on the sequencing results to obtain the full-length genome of respiratory syncytial virus.
[0018] The present invention provides a probe set for capturing the full-length genome of respiratory syncytial virus. The probe set includes probes with nucleotide sequences shown in SEQ ID NO: 1 to SEQ ID NO: 900. The probe set of the present invention is designed based on all RSV full-genome sequences in GenBank data, and probes are designed one by one for all its variant sequences, and then merged to remove redundancy. Compared with the conventional sequencing method of designing primers in fragments, the probe set has the characteristics of high density, high coverage, high throughput, and higher specificity, and can detect new mutations. The respiratory syncytial virus full-genome capture technology (next-generation sequencing) based on the probe set has the advantages of high throughput and short time consumption, and can improve the detection accuracy, specificity, and timeliness of sequence acquisition of the sample to be tested. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Comparison results of the sequencing results of the present invention and the first-generation sequencing method for the HRSV B subtype sample BHeN18-47;
[0020] Figure 2 Comparison results of the sequencing results of the present invention and the first-generation sequencing method for the HRSV A subtype sample ASH19-11. DETAILED DESCRIPTION OF THE INVENTION
[0021] The present invention provides a probe set for capturing the full-length genome of respiratory syncytial virus, including probes with nucleotide sequences shown in SEQ ID NO: 1 to SEQ ID NO: 900.
[0022] In the present invention, one end of the probe is preferably modified with biotin. The modification of biotin facilitates the formation of a complex after the subsequent capture probe hybridizes with the target region DNA, and then the amplification system of biotin-streptavidin is used to separate the target sequence.
[0023] In the present invention, the probe set includes a total of 900 probes, which can comprehensively cover the respiratory syncytial virus genome sequence, and has the characteristics of uniformity, specificity, and high capture efficiency for the captured region. The capture probe is preferably synthesized by a high-throughput electrochemical synthesis method. In the embodiment of the present invention, the capture probe is commissioned to be synthesized by Beijing E-Gene Technology Co., Ltd.
[0024] The present invention provides a kit for sequencing the full-length genome of respiratory syncytial virus, including the capture probe, streptavidin-labeled magnetic beads, and library hybridization reaction reagents.
[0025] In the present invention, the streptavidin-labeled magnetic beads bind to the biotin on the capture probe through streptavidin, and then the capture probe and the target region DNA sequence that specifically binds to the capture probe are separated from the system by using magnetic force. The library hybridization reaction reagent preferably includes at least one of the following reagents: TargetSeq Hyb Bufferv2, Hyb Human Block, Blocking Oligo and RNase enzyme inhibitor. The library hybridization reaction reagent is used for the hybridization reaction of the capture probe and for the capture probe to bind to a specific DNA sequence in the genomic library.
[0026] In the present invention, the kit preferably further includes an RNA fragmentation reagent, a reverse transcription reaction reagent, a cDNA second-strand synthesis reaction reagent, an adapter ligation reagent, a purification reagent, a PCR pre-reaction reagent, and a post-capture PCR amplification reaction reagent. The PCR pre-reaction reagent preferably includes at least one of the following: a PCR Master Mix containing UDG enzyme and a UDI primer. The RNA fragmentation reagent, the reverse transcription reaction reagent, the cDNA second-strand synthesis reaction reagent, the adapter ligation reagent, the purification reagent, and the PCR pre-reaction reagent are used for constructing a respiratory syncytial virus genomic library. The present invention has no special limitation on the reagents for constructing the respiratory syncytial virus genomic library, and well-known reagents in the art can be used. The post-capture PCR amplification reaction reagent preferably includes at least one of the following: a PCR primer and a post-capture PCR reaction premix. The post-capture PCR amplification reaction reagent is used to amplify the captured DNA sequence of the target region to obtain a large amount of DNA sequences of the target region to meet the sequencing requirements.
[0027] The present invention provides the use of the probe set or the kit in the full-length genome sequencing of respiratory syncytial virus.
[0028] In the present invention, the strain of respiratory syncytial virus preferably includes RSV A subtype and / or RSV B subtype. The sequencing preferably includes next-generation sequencing. The method for full-length genome sequencing of respiratory syncytial virus preferably involves hybridizing the respiratory syncytial virus genomic library with the capture probe, separating the target region DNA sequence bound to the capture probe from the hybridization solution. Since the capture probe can cover 100% of the respiratory syncytial virus genome, a sequence covering the full genome of respiratory syncytial virus is obtained. After PCR amplification and next-generation sequencing, the full-length genome of rubella virus is obtained. The full-length genome of respiratory syncytial virus obtained by sequencing can be used for monitoring the genomic variation of respiratory syncytial virus and / or tracing the origin of the respiratory syncytial virus genome. Using the kit for next-generation sequencing to obtain the full-length genome of respiratory syncytial virus, after comparing it with the reference genome, the variation information of the respiratory syncytial virus genome is obtained. The variation information includes single-base mutations and insertion / deletion mutations. By comparing and analyzing the variation information with the genomic sequences in the database, the purpose of monitoring the genomic variation of respiratory syncytial virus and / or tracing the origin is achieved. The method for tracing the origin of the respiratory syncytial virus genome is based on obtaining the full-length genome sequence information of respiratory syncytial virus, a full-genome database of a large number of sample strains can be established, and based on this database, the purpose of tracing the virus type in the outbreak area of respiratory syncytial virus can be achieved. The method for monitoring the genomic variation of respiratory syncytial virus preferably enables the effective detection of mutations for any new mutations at any locus on the respiratory syncytial virus genome, and accurately monitors the variation of respiratory syncytial virus.
[0029] The present invention provides a method for next-generation sequencing of the full-length genome of respiratory syncytial virus based on the capture probe, comprising the following steps:
[0030] Hybridize the respiratory syncytial virus genomic library of the sample to be tested with the capture probe to obtain a hybridization reaction solution;
[0031] Separate the DNA fragment hybridized with the probe from the hybridization reaction solution;
[0032] Perform post-capture PCR amplification using the DNA fragment hybridized with the probe as a template, subject the obtained PCR product to next-generation sequencing, and perform data analysis on the sequencing results to obtain the full-length genome of respiratory syncytial virus.
[0033] The present invention hybridizes the respiratory syncytial virus genomic library of the sample to be tested with the capture probe to obtain a hybridization solution.
[0034] The present invention has no special limitation on the method for constructing the respiratory syncytial virus genomic library of the sample to be tested, and the genomic library construction methods well-known in the art can be used, such as extraction and fragmentation of total RNA, reverse transcription and cDNA second-strand synthesis, addition of adapters and pre-PCR reaction, and fragment purification to obtain the genomic library.
[0035] In the present invention, the reaction system for the hybridization is preferably 30 μL and preferably includes the following reagents: TargetSeq Hyb Buffer v2 13 μL, Hyb Human Block 5 μL, Blocking Oligo 2 μL, RNaseBlock 5 μL, Nuclease-Free Water 3 μL, and Target Probes 2 μL. The reaction program for the hybridization is preferably 80 °C for 5 min and 50 °C for 12 - 18 h.
[0036] In the present invention, the method for separating the DNA fragment hybridized with the probe from the hybridization solution is preferably to mix the streptavidin-modified magnetic beads with the hybridization solution, and utilize the binding property of streptavidin and biotin to bind the magnetic beads to the hybridization complex to form a ternary complex, and separate the ternary complex from the hybridization solution under the action of magnetic field force.
[0037] After obtaining the DNA fragment hybridized with the probe, the present invention uses the DNA fragment hybridized with the probe as a template for post-capture PCR amplification, performs second-generation sequencing on the obtained PCR product, and analyzes the sequencing result to obtain the full-length genome of the respiratory syncytial virus.
[0038] In the present invention, the reaction system for the post-capture PCR amplification is preferably 50 μL and preferably includes the following reagents: 24 μL of the DNA fragment suspension hybridized with the probe, 1 μL of the Post PCR primer, and 25 μL of the Post PCR Master Mix. The reaction program for the post-capture PCR amplification is preferably 95 °C for 1 min; 98 °C for 20 s, 60 °C for 30 s, 72 °C for 30 s, with 7 - 12 cycles; 72 °C for 5 min.
[0039] In the present invention, the method for analyzing the sequencing result preferably includes preliminarily filtering the raw data of the sequencing to obtain Clean reads;
[0040] Removing the host gene sequence from the Clean reads, aligning the obtained sequence to the viral genome, and assembling to obtain the contig sequence;
[0041] After removing the host gene sequences from the contig sequences again, align them to the viral genome sequences to obtain data on viral typing and sorting;
[0042] Align the data on viral typing and sorting to the reference genome, and after correcting the variant sites, obtain the full-length genome of respiratory syncytial virus.
[0043] In the present invention, since there are adapter information and low-quality sequence reads in the original sequences obtained by sequencing, in order to improve the quality of the sequencing data, the original sequencing data is preliminarily filtered. The method of the preliminary filtering is preferably to cut off the sequences with an average base quality value less than 20 bp in a sliding window of 8 bp; remove the adapter sequences at the end of the sequences; if the first or last base of the sequence is less than 20 bp, then directly cut off the base; usually, if the remaining sequence length is less than 40 bp (paired-end) after the removal, then discard the pair of sequences.
[0044] In the present invention, the method for removing host gene sequences is preferably to use the Bowtie2 software to align the sequencing data to the reference genome of the host and collect the sequences that cannot be aligned.
[0045] In the present invention, it is preferably to use the bwa software to align the collected sequences that cannot be aligned to the genome of the virus. The genome of the virus refers to the genome of respiratory syncytial virus. The method for assembling the contig sequences is preferably to use the MEGAHIT software for assembly.
[0046] In the present invention, after removing the host gene sequences from the contig sequences again, align them to the viral genome to obtain data on viral typing and sorting. The viral genome refers to the genome of respiratory syncytial virus.
[0047] In the present invention, align the data on viral typing and sorting to the reference genome, and after correcting the variant sites, obtain the full-length genome of respiratory syncytial virus.
[0048] The reference genome is preferably the respiratory syncytial virus typing sequence ranked first in the virus typing ranking. The alignment is preferably performed using the bwa software. The variant site correction is preferably performed by using the samtools software for mutation analysis, and after obtaining the mutation sites, the iVar software is used for variant site correction. After the variant site correction, sequence integration is preferably further included. The method of sequence integration is preferably as follows: using the measles virus typing sequence ranked first in the virus typing ranking data as the reference genome, extracting reads that can be aligned to both the host and the virus genome from the Clean reads, and based on these sequences, finding the integration site coordinates on the host genome, and integrating the corrected sequences into the full-length measles virus genome sequence. After the variant site correction, sequence integration is preferably further included. The method of sequence integration is preferably as follows: using the measles virus typing sequence ranked first in the virus typing ranking data as the reference genome, extracting reads that can be aligned to both the host and the virus genome from the Clean reads, and based on these sequences, finding the integration site coordinates on the host genome, and integrating the corrected sequences into the full-length measles virus genome sequence. In the embodiments of the present invention, both the second-generation sequencing and the sequencing data analysis are commissioned to Beijing Genomics Institute (BGI) Tech Co., Ltd. to complete.
[0049] In the present invention, by adopting the above detection method and combining the probe hybridization capture and the second-generation sequencing technology, the full-length genome sequence of the respiratory syncytial virus can be obtained at low cost and quickly.
[0050] The following is a detailed description of a probe set and a kit for capturing the full-length genome of the respiratory syncytial virus provided by the present invention in conjunction with the embodiments, but they should not be construed as limiting the protection scope of the present invention.
[0051] Example 1
[0052] Design of a set of probes for capturing the nucleotide sequence of the RSV full-length genome
[0053] According to 5,647 full-length reference genome sequences of the respiratory syncytial virus (Human respiratory syncytial virus, taxid: 11250) downloaded from the NCBI database, oligonucleotide probes complementary to the sequences in the database are designed using the base complementary pairing principle.
[0054] Due to the complex genomic sequence structure and the existence of irregular regions, such as high-GC, high-AT regions, Alu sequences, etc., and at the same time due to the limitation of the second-generation sequencing read length and the need to obtain the full-length sequence, the probe design becomes particularly critical, and various factors such as the coverage rate, uniformity, and capture efficiency of the probes need to be comprehensively weighed. The following are the principles of probe design in the liquid-phase capture technology:
[0055] 1. Probe length: The probe is designed as a long probe with a length of 100 nt.
[0056] 2. Probe layer number: That is, the average number of probes covering the target region. The overlapping probe sequence design method is adopted to perform multi-layer probe coverage on the target region. At the same time, the thermodynamic stability is evaluated. For the q region with large variations, a probe density design of more than 10 layers is carried out so that even if there are relatively complex variations or new mutations in the pathogenic microorganism genome, the probe fault tolerance can still stably capture its genome.
[0057] 3. Specificity: That is, the uniqueness of the probe in the genomic range. The higher the specificity, the higher the capture efficiency of the probe.
[0058] 4. Binding ability: In the region with a GC content of about 50%, the probe has the strongest capture ability. In the high-GC region, the probe has a strong binding ability, but the DNA fragment itself has a stronger binding ability (longer length), and the probe competition resistance is relatively large; in the high-AT region, the DNA fragment itself has a weak binding ability, and the probe's binding ability to it is also weak. Therefore, in the high-GC and high-AT regions, it is necessary to appropriately increase the number of probes to make up for the disadvantages in terms of competition resistance and binding ability.
[0059] 5. Secondary structure: The formation of hairpin structures by the probe itself and dimers between probes should be avoided as much as possible, otherwise it will affect the binding of the probe to the target DNA.
[0060] 6. For cases such as microorganisms with multiple highly variable sequences, it is necessary to pay attention to:
[0061] (a) Design probes for all the viral whole-genome sequences that can be downloaded from the database to avoid omission;
[0062] Merge and remove redundancy for all probes to prevent the situation of too high sequencing depth in highly conserved regions. After comprehensive evaluation, a total of 900 probes are designed and synthesized, with a coverage rate of up to 100% for all sequences in the database, and the capture depth of each region of each genome is ensured to be uniform and effective. The probe sequences are shown in Table 1.
[0063] Table 1 Probe sequences
[0064]
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100] Example 2
[0101] A method for detecting the whole genome of RSV based on the combination of capture probe sets and next-generation sequencing
[0102] 1. RNA extraction from the sample to be tested
[0103] For RSV-positive clinical samples or virus strains collected in 2018-2019, RNA was extracted using the QIAamp Viral RNA Mini Kit (product number 52906, Qiagen, Germany). The RNA concentration was quantified using the Qubit RNA HS assay kit and the Qubit Fluorometer (Life Technologies, USA).
[0104] 2. Library preparation
[0105] 2.1. RNA fragmentation
[0106] 2.1.1 Take out the RNA sample and the Fast Frag Buffer from the -80 °C and -20 °C refrigerators respectively, melt them on an ice box, briefly vortex, centrifuge instantaneously, and place them on the ice box for standby.
[0107] 2.1.2 Prepare the reaction system according to Table 2 below, mix well and centrifuge instantaneously:
[0108] Table 2 Fragmentation reaction system
[0109] Reagent Volume RNA sample 13 μL (total 10 ng - 1 μg) Fast Frag Buffer 4 μL Total volume 17 μL
[0110] 2.1.3 Set the PCR instrument running program according to Table 3, and place the reaction solution on the PCR instrument:
[0111] Table 3 Reaction program
[0112]
[0113] 2.2. Reverse transcription
[0114] 2.2.1 Take out the Fast First Strand Buffer in the kit from the -20°C refrigerator, place it on an ice box to melt, briefly vortex it after melting on the ice box, centrifuge it instantaneously, and place it on the ice box for standby.
[0115] 2.2.2 Take out the Fast First Strand Enzyme in the kit from the -20°C refrigerator, invert and mix well, centrifuge it instantaneously, and place it on the ice box for standby.
[0116] 2.2.3 Prepare the reaction system according to Table 4 below, pipette and mix well (avoid vigorous shaking and mixing), centrifuge it instantaneously:
[0117] Table 4 Reverse Transcription Reaction System
[0118] Reagent Volume Sample from the end of reaction in Step 1 17 μL Fast First Strand Buffer 6 μL Fast First Strand Enzyme 2 μL Total volume 25 μL
[0119] 2.2.4 Set the PCR instrument operation program according to Table 5, and place the reaction solution on the PCR instrument:
[0120] Table 5 Reverse Transcription Reaction System
[0121]
[0122] 2.3 cDNA Second Strand Synthesis, 3'-End "A" Addition
[0123] 2.3.1 Take out the Fast Second Strand Buffer with dUTP in the kit from the -20°C refrigerator in advance, place it on an ice box to melt, briefly vortex it after melting on the ice box, centrifuge it instantaneously, and place it on the ice box for standby.
[0124] 2.3.2 Take out the Fast Second Strand Enzyme in the kit from the -20°C refrigerator, invert and mix well and centrifuge it instantaneously, and place it on the ice box for standby.
[0125] 2.3.3 Prepare the reaction system according to Table 6 below, pipette and mix well (avoid vigorous shaking and mixing), centrifuge it instantaneously:
[0126] Table 6 cDNA Second Strand Synthesis System
[0127] Reagent Volume Sample from the end of reaction in Step 2 25 μL Fast Second Strand Buffer with dUTP 30 μL Fast Second Strand Enzyme 5 μL Total volume 60 μL
[0128] 2.3.4 Set the PCR instrument operation program according to Table 7, and place the reaction solution on the PCR instrument to run the program:
[0129] Table 7 cDNA Second Strand Synthesis Reaction Program
[0130]
[0131] 2.4. Adapter Ligation
[0132] 2.4.1 Take out the Adapter from the -20°C refrigerator in advance, place it on an ice box to melt. After melting, briefly vortex and centrifuge instantaneously, then place it on the ice box for standby.
[0133] 2.4.2 According to the amount of RNA input for library construction, dilute the Adapter (15 μM) to an appropriate concentration in advance according to Table 8:
[0134] Table 8 Amount of Input RNA and Adapter Ratio
[0135] RNA input amount Adapter concentration Dilution factor 100 ng - 500 ng 7.5 μM 2-fold 50 ng 3.75 μM 4-fold 25 ng 1.5 μM 10-fold 10 ng 0.75 μM 20-fold
[0136] 2.4.3 Take out the Fast Ligation Buffer and Fast Ligase Mix in the kit from the -20°C refrigerator in advance, place them on an ice box to melt for standby.
[0137] 2.4.4 Prepare the reaction system on the ice box according to Table 9 below:
[0138] Table 9 Reaction System for Adding Adapter
[0139] Reagent Volume Sample from the end of reaction in Step 3 60 μL Fast Ligation Buffer 30 μL Fast Ligase Mix 5 μL Adapter (after dilution) 5 μL Total volume 100 μL
[0140] First, mix the diluted Adapter and the sample after the reaction in Step 3, then add the premixed reaction solution to reduce adapter self-ligation. Pipette and mix well (avoid vigorous shaking) and centrifuge instantaneously.
[0141] 2.4.5 Set the parameters of the PCR instrument according to Table 10, and place the PCR tube on the PCR instrument to run the program:
[0142] Table 10 Reaction Program for Adding Adapter
[0143]
[0144] 2.5. Purification after Ligation
[0145] 2.5.1 Use freshly prepared 80% ethanol (prepared with absolute ethanol and Nuclease-Free Water) for magnetic bead purification. (The magnetic beads used for purification are IGT TM Pure Beads or Agencourt AMPure XP).
[0146] 2.5.2 Take out the purification magnetic beads from the 4°C refrigerator, mix well and equilibrate at room temperature for 30 min, then vortex and mix for standby.
[0147] 2.5.3 Add purified magnetic beads with a volume 0.45 times that of the 4100 μL reaction solution in step 4, pipette and mix well, then let it stand for 5 min.
[0148] 2.5.4 Centrifuge briefly, place the PCR tube on the magnetic stand for 3 min until the solution becomes clear.
[0149] 2.5.5 Keep the PCR tube on the magnetic stand, discard the supernatant, add 200 μL of 80% ethanol solution to the PCR tube, let it stand for 30 s, and then discard the supernatant.
[0150] 2.5.6 Repeat the previous step.
[0151] 2.5.7 Cover the tube lid, centrifuge briefly, place it on the magnetic stand, and carefully use a 10 μL pipette to discard the residual ethanol at the bottom.
[0152] 2.5.8 Keep the PCR tube on the magnetic stand, let it stand at room temperature for 3 - 5 min to dry the magnetic beads.
[0153] 2.5.9 Remove the PCR tube from the magnetic stand, add 22 μL of Nuclease-Free Water, pipette and mix well, then let it stand for 2 min.
[0154] 2.5.10 Centrifuge briefly, place the PCR tube on the magnetic stand for 2 min until the solution becomes clear.
[0155] 2.5.11 Use a pipette to aspirate 20 μL of the supernatant and transfer it to a new PCR tube, make good marks, and prepare for step 6.
[0156] 2.6. Pre-PCR Reaction
[0157] 2.6.1 Take out the PCR Master Mix with UDG and UDIPrimer in the kit from the -20 °C refrigerator in advance, place it on an ice box to melt, mix well and then centrifuge briefly, and place it on the ice box for standby.
[0158] 2.6.2 Prepare the PCR reaction solution according to Table 11 below, and record the used Index number well:
[0159] Table 11 Pre-PCR Reaction System
[0160]
[0161] The sequence of UDIPrimer is as follows:
[0162] 5’-AATGATACGGCGACCACCGAGATCTACACNACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO: 901). In this sequence, N represents "i5Index". The i5Index adapter sequence refers to the Index sequence adapter located at the P5 end (near the P5 primer end of the Flow-cell) during library construction on the Illumina sequencing platform.
[0163] 5’-CAAGCAGAAGACGGCATACGAGATNGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 902). In this sequence, N represents "i7Index". The i7Index adapter sequence is used to distinguish different samples on the Illumina sequencing platform. The i7Index adapter sequence is usually located at the end of the sequencing library and binds to the sequencing primer to ensure correct differentiation of each sample's data during pooled sequencing. Pipette and mix well, then centrifuge briefly.
[0164] 2.6.5 Set the PCR instrument according to the records in Table 12 and Table 13 and run it:
[0165] Table 12 Pre-PCR reaction program
[0166]
[0167] Table 13 Relationship between the input sample volume and the number of PCR cycles
[0168]
[0169] 2.7. PCR amplification and purification
[0170] 2.7.1 Use freshly prepared 80% ethanol (prepared from anhydrous ethanol and Nuclease-Free Water) for magnetic bead purification. (The magnetic beads used for purification are IGT TM Pure Beads or Agencourt AMPure XP).
[0171] 2.7.2 Take out the purification magnetic beads from the 4°C refrigerator, mix well and equilibrate at room temperature for 30 min, then vortex and mix well for later use.
[0172] 2.7.3 Add 0.9 times the volume of the purification magnetic beads to the 50 μL reaction solution after step 6, pipette and mix well, and let it stand for 5 min.
[0173] 2.7.4 Centrifuge briefly, place the PCR tube on the magnetic stand for 3 min until the solution becomes clear.
[0174] Steps 2.7.5 - 2.7.8 are the same as 2.5.5 - 2.5.8.
[0175] 2.7.9 Remove the PCR tube from the magnetic stand, add 30 μL of Nuclease - Free Water, pipette up and down to mix well, and let it stand for 2 min.
[0176] 2.7.10 Centrifuge briefly, place the PCR tube on the magnetic stand for 2 min until the solution is clear. Pipette 28 μL of the supernatant and transfer it to a new PCR tube, and make a mark.
[0177] 2.7.12 Take 1 μL of the library and use the Qubit ds DNA HS Assay Kit reagent to measure the library concentration on a Qubit 4.0 Fluorometer, and record the library concentration.
[0178] 2.7.13 Take 1 μL of the sample and use a fragment analyzer to measure the fragment length. After the experiment, arrange for sequencing on the machine.
[0179] 3. Targeted Capture
[0180] 3.1. Preparation before Hybridization Capture Experiment
[0181] This experiment requires approximately two rounds of liquid - phase probe hybridization capture experiments.
[0182] 3.1.1 Take out Hyb Human Block, RNase Block and the supporting Blocking Oligo from the - 20 °C refrigerator, place them on an ice box to melt, briefly vortex and centrifuge briefly, and place them on the ice box for temporary storage;
[0183] 3.1.3 Take out the probes and the library for hybridization capture from the - 80 °C and - 20 °C refrigerators, place them on an ice box to melt, briefly vortex and centrifuge briefly, and place them on the ice box for temporary storage;
[0184] 3.1.5 Take out TargetSeq Hyb Buffer v2, melt it at room temperature, briefly vortex and centrifuge briefly. If there is precipitation, heat TargetSeq Hyb Buffer v2 in a 37 °C water bath until the reagent is completely dissolved before use.
[0185] 3.2. Hybridization of Library and Probe
[0186] 3.2.1 When hybridizing a single library, take 750 ng of the library and add it to a PCR tube, and make a mark; when hybridizing a mixture of multiple libraries, add 500 ng of each library.
[0187] 3.2.2 Place the PCR tube into a vacuum concentrator centrifuge, open the lid of the PCR tube, and concentrate it to a dry state. Prepare the hybridization reaction solution according to Table 14 below:
[0188] Table 14 Hybridization reaction system
[0189]
[0190] 3.2.4 Add the hybridization reaction solution to the dried library, vortex for 30 s to ensure that the DNA dried at the bottom of the tube dissolves, and centrifuge briefly.
[0191] 3.2.5 Set the reaction program according to Table 15, and run the hybridization reaction solution on a PCR instrument:
[0192] Table 15 Hybridization reaction program
[0193]
[0194] 3.2.6 It is recommended that the hybridization time be 12 - 18 h, and perform Step 3 30 min before the end of the program.
[0195] 3.3. Preparation before the capture experiment
[0196] 3.3.1 Take out the Cap Beads from the refrigerator at 4 °C, mix well and equilibrate at room temperature for 30 min.
[0197] 3.3.2 Take out the Wash Buffer 1. If there is precipitation, place Wash Buffer 1 in a water bath at 37 °C and heat it until the precipitation is completely dissolved before use.
[0198] 3.3.3 Take out the TargetSeq Wash Buffer2 v2 and preheat it on a water bath at 50 °C.
[0199] 3.3.4 Pipette 50 μL of Cap Beads into a new PCR tube, place it on a magnetic stand for 1 min, wait until the solution becomes clear, and discard the supernatant.
[0200] 3.3.5 Remove the PCR tube from the magnetic stand, add 180 μL of Binding Buffer, pipette or vortex to mix well to resuspend the magnetic beads. After instantaneous centrifugation, place the PCR tube on the magnetic stand for 1 min, wait until the solution becomes clear, and discard the supernatant;
[0201] 3.3.7 Repeat Steps 3.5 - 3.6 twice, and wash the magnetic beads with Binding Buffer three times in total;
[0202] 3.3.8 Remove the PCR tube from the magnetic stand, add 180 μL of Binding Buffer, pipette or vortex to mix well, and immediately proceed to step 4.
[0203] 3.4. Target Region DNA Capture
[0204] 3.4.1 Keep the hybridization product from step 2 on the PCR instrument, and add the 180 μL CapBeads prepared in step 3 to the hybridization product, and pipette to mix well.
[0205] 3.4.2 Close the tube cap, remove the PCR tube from the PCR instrument, place it on a vertical rotator mixer, with a rotation speed not exceeding 10 rpm, and bind at room temperature for 30 min (if there is no vertical rotator mixer in the laboratory, it can be bound at room temperature for 30 min, and invert the tube up and down several times every 5 min during this period to mix well).
[0206] 3.4.3 Remove the PCR tube, centrifuge briefly, place it on the magnetic stand for 2 min. After the solution becomes clear, discard the supernatant.
[0207] 3.4.4 Remove the PCR tube from the magnetic stand, add 150 μL of Wash Buffer 1 to the PCR tube, gently pipette to mix well to resuspend the magnetic beads, replace the tube cap with a new one, and then place it on a vertical rotator mixer to wash at room temperature for 15 min, with a rotation speed not exceeding 10 rpm.
[0208] 3.4.5 Remove the PCR tube, centrifuge briefly, place it on the magnetic stand for 2 min. After the solution becomes clear, discard the supernatant.
[0209] 3.4.6 Remove the PCR tube from the magnetic stand, add 150 μL of TargetSeq WashBuffer 2v2 preheated to 50 °C, gently pipette to mix well, centrifuge briefly, and place it on a metal bath to incubate at 50 °C for 10 min.
[0210] 3.4.7 Remove the PCR tube, centrifuge briefly, place it on the magnetic stand for 2 min. After the solution becomes clear, discard the supernatant.
[0211] 3.4.8 Repeat steps 4.6 - 4.7 twice, and wash the magnetic beads with TargetSeq Wash Buffer 2v2 at 50 °C three times in total.
[0212] 3.4.9 Keep the PCR tube on the magnetic stand, add 200 μL of 80% ethanol to the PCR tube, let it stand for 30 s, and then completely discard the ethanol solution (the residual ethanol can be discarded with a 10 μL pipette), and air-dry the magnetic beads at room temperature to completely volatilize the residual ethanol.
[0213] 3.4.10 Add 24 μL of Nuclease-Free Water to the PCR tube. Remove the PCR tube from the magnetic stand, briefly vortex to resuspend the magnetic beads, and proceed with the amplification reaction in step 5.
[0214] 3.4.11 Keep the PCR tube on the magnetic stand, add 200 μL of 80% ethanol to the PCR tube, let it stand for 30 s, and then completely discard the ethanol solution (residual ethanol can be discarded using a 10 μL pipette). Air-dry the magnetic beads at room temperature to completely evaporate the residual ethanol (during the drying process, observe the surface of the magnetic beads to avoid over-drying).
[0215] 3.4.12 Add 24 μL of Nuclease-Free Water to the PCR tube. Remove the PCR tube from the magnetic stand, briefly vortex to resuspend and mix the magnetic beads, and proceed with the amplification reaction in step 5.
[0216] 3.5. Post-capture PCR amplification
[0217] 3.5.1 Take out the Post PCR Master Mix and Post PCR Primer from the -20°C refrigerator in advance, place them on the ice box to melt, and store them temporarily on the ice box after melting.
[0218] 3.5.2 Before the experiment, please double-check whether the Post PCR Primer is used correctly. Briefly vortex the Post PCR Master Mix and Post PCR Primer, and perform a brief centrifugation.
[0219] 3.5.3 Prepare the PCR reaction mixture according to Table 16 below. Note that this step is a PCR reaction with magnetic beads:
[0220] Table 16 Post-capture PCR reaction system
[0221] Reagent Volume Magnetic bead suspension obtained in Step 4 24 μL Post PCR Primer 1 μL Post PCR Master Mix 25 μL Total volume 50 μL
[0222] Among them, the specific sequences of the Post PCR Primer are as follows:
[0223] 5'-AATGATACGGCGACCACCGA (SEQ ID NO: 903)
[0224] 5'-CAAGCAGAAGACGGCATACGA (SEQ ID NO: 904)
[0225] 3.5.4 After preparation, use a pipette to aspirate and mix well. After mixing, quickly transfer it to the PCR instrument. Do not use the method of vortexing and then centrifuging to mix.
[0226] 3.5.5 Set the PCR instrument program as follows. Place the PCR reaction solution on the PCR instrument and run the program according to Table 17:
[0227] Table 17 PCR reaction program
[0228]
[0229] 3.5.6 After the program is completed, perform the magnetic bead purification in Step 6.
[0230] The number of PCR cycles can refer to the parameters on the tube wall label of the probe Target Prober or the probe instruction file. The number of Post-PCR cycles is related to the total input amount of the library during hybridization. When the input amount of the library during hybridization is large, the number of Post-PCR cycles can be appropriately reduced. Since the MGI platform requires a relatively large amount of library for on-machine sequencing, it is recommended to add two more cycles.
[0231] 3.6. Post-amplification purification
[0232] 3.6.1 Take out the purification magnetic beads, mix well and equilibrate at room temperature for 30 min.
[0233] 3.6.2 To the PCR product in Step 5, add 1.1 times the volume of magnetic beads (55 μL), pipette or vortex to mix well, and let stand at room temperature for 5 min.
[0234] Steps 3.6.3 - 3.6.7 are the same as 2.7.4 - 2.7.8.
[0235] 3.6.8 Remove the PCR tube from the magnetic rack, add 25 μL of Nuclease-Free Water, pipette to mix well, and let stand for 2 min. Centrifuge briefly, place the PCR tube on the magnetic rack for 2 min until the solution is clear.
[0236] 3.6.9 Use a pipette to aspirate 23 μL of the supernatant and transfer it to a new PCR tube; store the captured library in a -20 °C refrigerator. The captured library can be stored in a -20 °C refrigerator for one month.
[0237] 3.6.10 Take 1 μL of the library and use the Qubit dsDNA HS Assay Kit reagent to measure the library concentration on a Qubit 4.0 Fluorometer, and record the library concentration.
[0238] 3.6.11 Take 1 μL of the library and use a fragment analyzer for fragment quality inspection. The fragment size should be basically the same as the size of the pre-library.
[0239] Perform on-machine sequencing.
[0240] Take 1 μL from the probe hybridization capture library and quantify it using the Qubit ds DNA HS Assay Kit, and record the library concentration; take 1 μL of the sample and determine the fragment length using the Agilent 2100 Bioanalyzer system (Agilent DNA 1000 Kit). The library length is between 250 bp and 400 bp. The probe hybridization capture library is sequenced using the NovaSeq 6000 (Illumina) sequencing platform to obtain the raw sequencing data.
[0241] 4. Data analysis
[0242] The raw data obtained by sequencing is preliminarily filtered for Raw reads to obtain Clean reads, and subsequent analyses are all based on Clean reads. The content of data filtering is to cut off the sequences with an average base quality value less than 20 in a sliding window of 8 bp; remove the adapter sequences at the end of the sequences; if the first or last base of the sequence is less than 20, then directly cut off the base; usually, if the remaining sequence length is less than 40 (paired-end) after removal, then discard the pair of sequences.
[0243] Remove host sequences: Since there will still be residual host genomic sequences in the sequencing data, first use Bowtie2 to align the sequencing data to the host reference genome to obtain the sequences that could not be aligned (unmapR1 & unmapR2). This part of the unmap sequences is the viral sequence information that we may actually need.
[0244] Viral genome alignment: Use the bwa software to align the unmap sequences to the viral genome, use the MEGAHIT software to assemble the reads aligned to the genome into contig sequences, then use blastn to align the contig sequences with the host genome, remove the host sequences again to obtain the non-human contig file, and then align the non-human contig sequences to the viral genome sequence to obtain the possible viral typing sorting results-virus.txt.
[0245] Obtain the full-length RSV genome sequence: Take the type sequence of results-virus-top1, which ranks first among the obtained possible virus genotyping, as the reference genome. Use the bwa software to align the clean reads to this reference genome to obtain the aligned bam file. Use the samtools software for mutation analysis to obtain the mutation sites. Use the iVar software to correct the mutation sites in the genome to obtain the final RSV sequence. Take the type sequence of results-virus-top1 as the reference genome, use the lumpy software to analyze the clean reads, extract the reads that can be aligned to both the host and the virus genome, and based on these sequences, find the coordinates of the integration sites on the host genome. Perform a depth statistics of the integrated reads on the final RSV sequence to obtain the full-length RSV genome sequence. The results are shown in Table 18.
[0246] Table 18 Sequencing results of the full-length RSV genome of the test samples
[0247]
[0248]
[0249] Comparative Example 1
[0250] Comparison of the results of RSV whole-genome sequencing based on the first-generation sequencing method
[0251] Design primers for RSV subtype A and subtype B respectively according to the Genbank database. After RT-PCR amplification, use the first-generation sequencing method for sequence sequencing, as follows:
[0252] 1. Sample name:
[0253] Sample 1 is the clinical sample ASH19-11 of RSV subtype A;
[0254] Sample 2 is the clinical sample BHeN18-47 of RSV subtype B.
[0255] 2. RNA extraction: Use the QIAamp Viral RNA Minikit (product number 52906, Qiagen, Germany) to extract and obtain RNA.
[0256] 3. The primer pair sequences are shown in Table 19.
[0257] Table 19 Primer pair sequences
[0258]
[0259] 4. PCR reaction
[0260] The reaction system was configured using the Super ScriptTM III One-Step RT-PCR System with Platinum Taq High Fidelity, Invitrogen, Cat. No. 12574-035, as shown in Table 20 specifically:
[0261] Table 20 Reaction System
[0262] Reagent Volume 2×Reaction Mix 12.5 μL SS III RT-Enzyme 0.5 μL Forward primer (10 μM) 0.5 μL Reverse primer (10 μM) 0.5 μL <![CDATA[ddH2O]]> 8 μL <![CDATA RNA > 3 μL Total volume 25 μL
[0263] Set the PCR instrument and run it according to the reaction procedures shown in Tables 21 - 23:
[0264] Table 21 Reaction Procedure
[0265]
[0266] Table 22 Reaction Procedure
[0267]
[0268] Table 23 Reaction Procedure
[0269]
[0270] 5. Send the PCR product to a biological company for sequencing
[0271] For the raw data obtained after sequencing, use the sequencher5.0 software for sequence editing, splicing, and alignment. The sequencing length of sample 1 is 15152bp, and that of sample 2 is 15205bp. Compare the sequences obtained by first-generation sequencing with those based on liquid-phase probe capture combined with second-generation sequencing: In the full sequences obtained by first-generation sequencing and high-throughput sequencing by the probe capture method, the full genome length of first-generation sequencing is shorter than that of the present invention, and the clarity of the peak maps at some sites is poor; moreover, compared with the present invention, there are 7 nucleotide differences (A11163G, A11164G, G11165T, A11166G, G11167T, G11167A, G11170T) in the full genome sequence of sample 1, with an accuracy rate of 99.9995%, and there is 1 nucleotide difference (A10347G) in sample 2, with an accuracy rate of 99.9999%.
[0272] Figure 1 and Figure 2Comparison of the sequencing results of the samples of HRSVB subtype and HRSVA subtype using the present invention and the first-generation sequencing method. The top complete green line is the second-generation sequencing result, and the multiple short lines in the middle represent the results of fragment sequencing in the first-generation sequencing. According to the reference strain for sequence alignment, the adjacent two regions with coverage of each other are connected together to finally form a complete first-generation sequencing whole-genome sequence. In the first-generation sequencing, 32 and 23 fragments were sequenced respectively, and there are a large number of unidirectional sequencing results (red or green in the figure) in the sequencing results, which means that the accuracy of the bases at the sites contained in this fragment is worse than that of bidirectional coverage (both red and green at this site); while in the present invention, the accuracy of sequencing at each base site has reached more than 99.99%, and the accuracy is higher.
[0273] By comparing the test results of Example 1 and Comparative Example 1, it shows that the present invention can achieve the purpose of accurately and efficiently detecting the RSV genome based on probe hybridization capture using the second-generation sequencing technology.
[0274] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A probe set for capturing the full-length genome of respiratory syncytial virus, characterized in that: The probe group includes probes with nucleotide sequences as shown in SEQ ID NO: 1 to SEQ ID NO:
900.
2. The probe set according to claim 1, characterized in that: One end of the probe is modified with biotin.
3. A kit for full-length genome sequencing of respiratory syncytial virus, characterized in that: The method comprises the probe set according to claim 1 or 2, magnetic beads labeled with streptavidin and a library hybridization reaction reagent.
4. The kit for sequencing the respiratory syncytial virus genome according to claim 3, characterized in that: The library hybridization reaction reagents include at least one of the following reagents: TargetSeq Hyb Buffer v2, Hyb Human Block, Blocking Oligo and RNase Inhibitors.
5. The kit for sequencing the respiratory syncytial virus genome according to claim 3 or 4, characterized in that: The kit also includes at least one of the following: RNA fragmentation reagent, reverse transcription reaction reagent, cDNA double-strand synthesis reaction reagent, adapter connection reagent, purification reagent, PCR pre-reaction reagent and post-capture PCR amplification reaction reagent.
6. The kit for full-length genome sequencing of respiratory syncytial virus according to claim 5, characterized in that: The PCR pre-reaction reagents include PCR MasterMix containing UDG enzyme and / or UDI primers.
7. The kit for full-length genome sequencing of respiratory syncytial virus according to claim 5, characterized in that: The post-capture PCR amplification reaction reagents include PCR primers and / or post-capture PCR reaction premix.
8. Use of the probe set according to claim 1 or 2 or the kit according to any one of claims 3 to 6 in sequencing the full-length genome of respiratory syncytial virus.
9. The use according to claim 8, characterized in that: The sequencing includes next-generation sequencing.
10. A method for performing second-generation sequencing of the full-length genome of respiratory syncytial virus based on the probe set according to claim 1 or 2, characterized in that: The following steps are involved: Hybridizing the respiratory syncytial virus genome library of the sample to be tested with the probe set to obtain a hybridization product; separating the DNA fragment hybridized with the probe from the hybridization product; The DNA fragment hybridized with the probe is used as a template for post-capture PCR amplification, the obtained PCR product is subjected to second-generation sequencing, and the sequencing result is subjected to data analysis to obtain the full-length genome of the respiratory syncytial virus.
Citation Information
Patent Citations
Capture kit and method of target gene
CN105524981A
Kit and method for detecting 25 RNA viruses in respiratory tract by NGS targeted probe capture method
CN111575405A
Human respiratory virus targeted enrichment capture probe set and application thereof
CN112342270A
Capture probe for rubella virus full-length genome detection and kit and application thereof
CN119040516A
Capture probe for detecting whole genome of human parainfluenza virus as well as kit and application thereof
CN119351631A
Cited By
Probe set for specifically enriching RSV whole genome in respiratory tract sample, kit and application
CN121344267A