A method for processing samples of lower respiratory tract aspirates for ventilator-associated pneumonia
By using saponins to remove host genes and batch extraction of DNA, combined with MinION nanopore sequencing and multi-sample PCR amplification technology, rapid pathogen identification and drug resistance detection of lower respiratory aspirate samples for ventilator-related pneumonia was achieved, solving the problems of long and low sample processing time in the prior art.
Patent Information
- Application Number
- CN202210716329.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-06-17
AI Technical Summary
The existing ventilator-related pneumonia samples have a long processing time and low processing efficiency, resulting in too long acquisition of pathogenic drug sensitivity results.
A sample processing method including saponin removal of host genes, batch extraction of DNA, construction of multi-sample PCR amplified DNA sequencing library, DNA sequencing, screening and analysis using MinION nanopore sequencer, pathogen gene splicing and drug resistance detection were used.
The time for identification of pathogens is shortened from conventional 24 to 48 hours to within 6 hours, which improves sample processing efficiency, reduces host genomic information, increases the amount of pathogen sequence acquisition, and shortens drug resistance detection time.
Smart Images

Figure CN115044689B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pathogen identification, and in particular to a method for processing samples of lower respiratory tract aspirates of ventilator-associated pneumonia. Background Art
[0002] Ventilator-associated pneumonia refers to pneumonia that occurs 48 hours after a patient receives mechanical ventilation or within 48 hours after extubation. It is one of the most common complications and causes of death in patients receiving mechanical ventilation. Studies have shown that giving appropriate antimicrobial drugs to patients suspected of having ventilator-associated pneumonia as early as possible can significantly shorten the patient's illness time and reduce the probability of disease progression and patient death.
[0003] The existing sample processing of lower respiratory tract aspirates for ventilator-associated pneumonia usually takes 24-48 hours. If you want to obtain the drug sensitivity results of the pathogens in the aspirates, it will take 48-72 hours, which means the processing time is long and the processing efficiency is low. Summary of the invention
[0004] In view of the above analysis, the present invention aims to provide a sample processing method for lower respiratory tract aspirates of ventilator-associated pneumonia, so as to solve the technical problems that the existing sample processing time of lower respiratory tract aspirates of ventilator-associated pneumonia patients is as long as 24-48 hours and the processing efficiency is low.
[0005] The purpose of the present invention is mainly achieved through the following technical solutions:
[0006] The present invention provides a method for processing samples of lower respiratory tract aspirates of ventilator-associated pneumonia, comprising the following steps:
[0007] S1, sample processing;
[0008] Including S11. Using saponin to remove sample host genes; S12. Batch extraction of DNA;
[0009] S2, constructing a multi-sample PCR-amplified DNA sequencing library;
[0010] S3, sequencing the DNA in the DNA sequencing library;
[0011] DNA sequencing was performed using the MinION nanopore sequencer, with data collected in real time;
[0012] S4. Screen and analyze the collected data;
[0013] S5. Perform pathogen gene splicing on the processed samples;
[0014] S6. Conduct drug resistance test on successfully spliced bacteria.
[0015] Furthermore, in S11, the process of removing the sample host gene using saponin includes the following sub-steps:
[0016] S111. Resuspend the clinical lower respiratory tract sample pellet with 250 μl sterile PBS buffer and mix thoroughly by pipetting;
[0017] S112, add 200 μl of 5% sterile saponin solution and dissolve it in sterile enzyme-free deionized water, vortex for 15 seconds to mix, and let stand at room temperature for 10 minutes to allow the saponin solution and sample to be fully mixed. Saponin can cause host cells to swell and rupture;
[0018] S113, add 350 μl of sterile enzyme-free deionized water and let stand at room temperature for 30 seconds;
[0019] S114, add 12 μl of 5 M NaCl solution and mix by inverting to stop the swelling effect of saponin;
[0020] S115, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant;
[0021] S116. Resuspend the precipitate with 100 μl sterile PBS buffer, add 100 μl HL-SAN buffer, vortex for 15 seconds to mix, add 10 μl HL-SAN enzyme, invert to mix, and digest free nucleic acids;
[0022] S117, 37℃ water bath, high speed shaking for 15min;
[0023] S118, centrifuge at 8000 rpm at 4°C for 5 min and discard the supernatant;
[0024] S119, add 800 μl sterile PBS buffer, mix by inversion, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant;
[0025] S1110, add 1000 μl sterile PBS buffer, invert to mix, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant.
[0026] Further, in step S116, the preparation process of HL-SAN buffer is as follows: 0.2 g MgCl 2 -6H 2 O was dissolved in 5 M NaCl, the volume was adjusted to 10 ml and filtered using a 0.22 μm filter.
[0027] Further, in S12, DNA extraction was performed using a BSCC45S1E kit and an automatic DNA extractor.
[0028] Further, in S12, the DNA extraction process includes the following sub-steps:
[0029] Dissolve all the lysozyme in the S121 and BSCC45S1E kits in TET buffer and mix thoroughly by vortexing;
[0030] S122, resuspend the pellet in 180 μl of the mixture, shake and mix, and incubate at 37°C for 30 min;
[0031] S123, add 20 μl of proteinase K and bacterial resuspension solution to columns 1 and 7 of the kit;
[0032] S124. Place the reagent kit in the machine and set the program: S125. After the extraction is completed, use Nanodrop to identify the DNA concentration and purity, and use Qubit to accurately quantify the concentration.
[0033] Furthermore, in S2, the rapid barcoding kit SQK-RBK004 was used to construct a multi-sample rapid DNA sequencing library. All operations were performed in a clean bench, and all instruments were exposed to ultraviolet light for at least 30 min. All reagents were melted on ice, centrifuged before use, and mixed by pipetting.
[0034] Furthermore, in S3, during the DNA sequencing process, the chip priming kit EXP-FLP002 was used for chip priming, all operations were performed in a clean bench, and all instruments except the chip were exposed to ultraviolet light for 30 minutes; FLO-MIN106DR9.4 chip and MinION sequencer were used for sequencing.
[0035] Furthermore, in S3, the DNA sequencing process specifically includes the following sub-steps:
[0036] S31. Insert the FLO-MIN106D R9.4 chip into the MinION sequencer, connect it to a laptop, and use the MinKNOW software Flowcell Test module to test the chip. If the number of nanopores on the new chip is greater than 800, it is a qualified chip.
[0037] S32, disconnect the sequencer from the computer, and place the sequencer in the clean bench for sample loading;
[0038] S33, add 30 μl of reagent FLT to one tube of reagent FLB, and mix by pipetting;
[0039] S34, opening the triggering hole and sucking out the air in the hole;
[0040] S35, add 800 μl FLB+FLT mixed reagent, close the initiation hole, and let stand at room temperature for 5 min;
[0041] S36, open the initiation well and the loading well, add 200 μl of FLB+FLT mixed reagent to the initiation well, and drop 75 μl of library into the loading well;
[0042] S37, the sequencer was connected to a laptop computer and sequencing was performed using MinKNOW software.
[0043] Furthermore, in S37, the program is set according to the kit, and data analysis is performed in real time. After the clinical pathogen sequence appears, sequencing is continued for 1-2 hours. If no other pathogen sequence appears, sequencing is stopped and the connection between the computer and the sequencer is disconnected.
[0044] Furthermore, the screening and analysis of the collected data in step S4 includes: the process of screening and analyzing the collected data includes: first performing base calling, filtering the data after base calling, performing data quality control after data filtering, performing host gene comparison and removal, and then performing classification analysis.
[0045] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0046] (1) The present invention can perform rapid pathogen identification on lower respiratory tract samples of patients with ventilator-associated pneumonia, shortening the time required for pathogen identification from the conventional 24 to 48 hours to less than 6 hours.
[0047] (2) The present invention can effectively remove host genome information and increase the amount of pathogen sequences obtained by nanopore metagenomic sequencing.
[0048] (3) The existing bacterial resistance test takes 48-72 hours to produce results. The processing method of the present invention only takes 6 hours to obtain preliminary resistance gene information, providing possible options for resistance and potential resistance.
[0049] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the embodiments of the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The drawings are only for the purpose of illustrating particular embodiments and are not to be considered limiting of the present invention. Like reference symbols denote like components throughout the drawings.
[0051] Figure 1 A flow chart showing the comparison between the method for removing host genes by saponin of the present invention and a blank control; Figure 2 The flowchart of sequencing and data analysis of the present invention; Figure 3 It is the overall analysis process and time-consuming schematic diagram of the present invention; Figure 4The figure shows the filtering effect of raw sequencing data. A: total number of sequences; B: total number of bases sequenced; C: maximum sequence length; D: minimum sequence length; E: average sequence length. Raw data: original sequencing data; Filtered data: filtered data.
[0052] Figure 5 This is the effect diagram of minimap2 alignment after removing host gene sequence information; Filtered data refers to the data after filtering; Without Homo Genome refers to the data after minimap2 alignment after removing host gene sequences.
[0053] Figure 6 It is a schematic diagram of the sequencing data processing flow of the present invention; DETAILED DESCRIPTION
[0054] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of the present invention and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0055] The present invention provides a method for processing samples of lower respiratory tract aspirates of ventilator-associated pneumonia, such as Figure 1-3 and Figure 6 As shown, the following steps are included:
[0056] S1, sample processing;
[0057] The method comprises the following sub-steps: S11. removing host genes of the sample using saponin; S12. extracting DNA in batches;
[0058] S2, using the extracted DNA to construct a library to obtain a DNA sequencing library;
[0059] DNA library was constructed using the SQK-RBK004 kit;
[0060] S3, sequencing the DNA in the DNA sequencing library;
[0061] Sequencing was performed using the MinION nanopore sequencer, with data collected in real time;
[0062] S4. Screen and analyze the collected data;
[0063] S5, splicing pathogen genes on processed samples;
[0064] S6. Conduct drug resistance test on successfully spliced bacteria.
[0065] In step S1, before using saponin to remove the host gene of the sample, the sample is collected and stored first. The specific process is: collect lower respiratory tract aspirates from patients with ventilator-associated pneumonia clinically, collect at least 1 ml of solids from each patient, and immediately perform subsequent processing after collection. If it cannot be processed immediately, the sample is stored at 4°C for up to 4 hours. It should be noted that, first, after sampling, it is transported on ice and 4 times the volume of sterile PBS buffer is added; secondly, the aspirate sample is placed in a 37°C water bath and shaken at high speed for at least 30 minutes until the sputum crust is completely dissolved; thirdly, the dissolved sample is divided into 1 ml tubes, placed in 1.5 ml sterile EP tubes, and centrifuged at 8000 rpm4°C for 30 minutes; finally, the supernatant is collected in a 5 ml cryotube, and the sample is frozen in liquid nitrogen for 2-3 minutes and stored at -80°C.
[0066] It should be noted that the microbial culture process includes:
[0067] S1', rewarm the sample stored at -80℃ to room temperature;
[0068] S2', pick 5 μl of sample, streak it on the enrichment medium according to the three-zone method, and culture it overnight at 37°C for 12-16 hours;
[0069] S3' directly uses single colonies grown on plates for mass spectrometry analysis to identify bacterial species and microbial resistance phenotypes;
[0070] S4' Pick 1 μl of a single colony sample, dissolve it in a growth broth medium, and culture it in a 37°C water bath with medium-speed shaking for 6-8 hours;
[0071] S5' bacterial broth was pipetted and mixed, 1 ml was dispensed into 1.5 ml sterile EP tubes, and centrifuged at 13000 rpm at 4°C for 1 minute;
[0072] The S6' supernatant was collected into a waste liquid tank dedicated for infection, and pure bacteria were collected, snap-frozen in liquid nitrogen for 2-3 minutes, and stored at -80°C for subsequent genomics research.
[0073] In step S11, after the aspirate sample is collected, saponin is used to remove the sample host gene, which specifically includes the following sub-steps:
[0074] S111. Resuspend the clinical lower respiratory tract sample pellet with 250 μl sterile PBS buffer and mix thoroughly by pipetting;
[0075] S112, add 200 μl of 5% sterile saponin solution and dissolve it in sterile enzyme-free deionized water, vortex for 15 seconds to mix, and let stand at room temperature for 10 minutes to allow the saponin solution and sample to be fully mixed. Saponin can cause host cells to swell and rupture;
[0076] S113, add 350 μl of sterile enzyme-free deionized water and let stand at room temperature for 30 seconds;
[0077] S114, add 12 μl of 5 M NaCl solution and mix by inverting to stop the swelling effect of saponin;
[0078] S115, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant;
[0079] S116. Resuspend the precipitate with 100 μl sterile PBS buffer, add 100 μl HL-SAN buffer, vortex for 15 seconds to mix, add 10 μl HL-SAN enzyme, invert to mix, and digest free nucleic acids;
[0080] S117, 37℃ water bath, high speed shaking for 15min;
[0081] S118, centrifuge at 8000 rpm at 4°C for 5 min and discard the supernatant;
[0082] S119, add 800 μl sterile PBS buffer, mix by inversion, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant;
[0083] S1110, add 1000 μl sterile PBS buffer, invert to mix, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant.
[0084] In the above step S116, the preparation process of HL-SAN buffer is as follows: 0.2 g MgCl 2 -6H 2 O was dissolved in 5 M NaCl, the volume was adjusted to 10 ml and filtered using a 0.22 μm filter.
[0085] It should be noted that the principle of the saponin host removal process of the present invention is that saponin, as a detergent, can swell and rupture the cell membrane of cells without cell walls through osmotic action, without affecting bacteria with cell walls. After the cell membrane of cells without cell walls ruptures, free DNA and RNA can be degraded by DNA hydrolases and RNA hydrolases, thereby reducing the abundance of host genes.
[0086] In step S11, when saponin was used to remove the sample host gene, a blank control test was also performed, such as Figure 1 As shown, the blank control test includes the following process:
[0087] S111', resuspend the clinical lower respiratory tract sample (ETA sample) pellet with 250 μl sterile PBS buffer and mix thoroughly by pipetting;
[0088] S112', add 200 μl sterile enzyme-free deionized water to dissolve, and let stand at room temperature for 10 min;
[0089] S113', add 350 μl sterile enzyme-free deionized water and let stand at room temperature for 30 seconds;
[0090] S114', add 12 μl sterile, mold-free water;
[0091] S115', centrifuge at 8000 rpm and 4 °C for 5 min, and discard the supernatant;
[0092] S116', the precipitate was resuspended in 100 μl sterile PBS buffer, and 110 μl sterile PBS buffer was added;
[0093] S117', 37℃ water bath, high speed shaking for 15min;
[0094] S118', centrifuge at 8000 rpm and 4 °C for 5 min, and discard the supernatant;
[0095] S119', add 800 μl sterile PBS buffer, mix by inversion, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant;
[0096] S1110', add 1000 μl sterile PBS buffer, mix by inversion, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant.
[0097] Control group: Comparison of the effect of saponin on host gene removal. Saponin on host genome removal used standard procedures, and all reagents in the control group were replaced with sterile enzyme-free water. ETA: lower respiratory tract aspirate; HL-SAN: thermosensitive salt-active nuclease.
[0098] In step S12, DNA extraction was performed using a BSCC45S1E kit (produced by Bioer Corporation) and an automatic DNA extractor (produced by Bioer Corporation).
[0099] The above step S12 includes the following sub-steps:
[0100] All lysozymes in the S121 and BSCC45S1E kits were dissolved in TET buffer and mixed by vortexing; TET buffer is the reagent included in the BSCC45S1E kit.
[0101] S122, resuspend the pellet in 180 μl of the mixture, shake and mix, and incubate at 37°C for 30 min;
[0102] S123, adding 20 μl of proteinase K and bacterial resuspension solution to the 1st column and the 7th column of the BSCC45S1E kit; the other columns are used to hold other reagents for automatic DNA extraction, for example, the automatic DNA extractor will automatically screen the required reagents;
[0103] S124. Place the reagent kit into an automatic DNA extraction instrument and set the program as shown in Table 3:
[0104] Table 1. Schedule for setting up the fully automated DNA extraction device
[0105]
[0106] S125. After the extraction is completed, use Nanodrop to identify the DNA concentration and purity to obtain a DNA solution that meets the quality requirements of MinION nanopore sequencing, and use Qubit to accurately quantify the concentration.
[0107] In step S2, the rapid barcoding kit SQK-RBK004 was used for multi-sample rapid PCR amplification and DNA sequencing library construction. All operations were performed in a clean bench, and the instruments were exposed to ultraviolet light for 30 minutes. All reagents were melted on ice and centrifuged briefly before use and mixed by pipetting.
[0108] DNA concentration determination using Qubit: dsDNA HS Assay Kit for Kit (12640ES60, Yisheng Biotechnology Co., Ltd.) was used to determine the DNA concentration. The determination steps were as follows:
[0109] 1. Prepare 0.5 ml sterile enzyme-free EP tubes and mark the standard 1, 2 and sample number on the tube caps;
[0110] 2. Prepare the buffer solution at a ratio of Qubit Reagent:Qubit HS DNA Buffer = 1:200, prepare a total volume of 200 μl for each sample, and vortex to mix;
[0111] 3. Add 190μl buffer and 10μl standard to the standard tube, add 199μl buffer and 1μl sample to the sample tube, and vortex to mix;
[0112] 4. Let stand at room temperature for 2 minutes;
[0113] 5. Turn on Qubit, select "DNA", "1HS DNA High Sensitivity", "Read Standards";
[0114] 6. Add standard 1 and click read to read;
[0115] 7. Add standard 2 and click read to read;
[0116] 8. Select "Run Samples", choose to add 1μl sample, and the output unit is ng / μl;
[0117] 9. Put in the sample tube and click "Read tube";
[0118] 10. Record the sample concentration.
[0119] In the above step S2, multiple sample PCR amplification DNA sequencing library construction is performed; specifically, the following processes are included:
[0120] S21, add 5ng DNA to sterile enzyme-free 0.2ml PCR tubes, the DNA volume should not exceed 3μl;
[0121] S22, add sterile enzyme-free deionized water to make up to 3 μl;
[0122] S23, add 1 μl of reagent FRM respectively, flick to mix, and centrifuge briefly;
[0123] S24, incubate at 30°C for 1 min;
[0124] S25, incubate at 80°C for 1 min, cool on ice;
[0125] S26, add 23.5 μl sterile enzyme-free deionized water, 1 μl PCR barcoding reagent RLB01-12A, 10 μl 5x Q5 buffer, 1 μl dNTP, 0.5 μl Q5 Polymerase, 10 μl enhancer;
[0126] S27, amplify according to the procedure described in Table 3.4;
[0127] Table 2 Multi-sample PCR amplification library construction amplification setup schedule
[0128]
[0129] S28 and AMPure XP magnetic beads were vortexed and mixed. All samples were collected in a sterile enzyme-free 1.5 ml EP tube, and the mixed magnetic beads were added according to the volume of 30 μl for each sample;
[0130] S29, shake on a horizontal shaker at room temperature at 200 rpm for 5 minutes;
[0131] Place the S210 and EP tubes on the magnetic rack, wait for the magnetic beads to be completely adsorbed and the liquid to become clear, and then gently remove the supernatant;
[0132] S211, add 200 μl 75% ethanol to wash the magnetic beads and discard the washing solution;
[0133] S212, repeat step 11;
[0134] S213, remove the EP tube from the magnetic rack, centrifuge it instantly, and remove the remaining ethanol liquid;
[0135] S214, open the lid and let stand for 2-5 minutes to dry the magnetic beads;
[0136] S215, add 10 μl TE buffer, thoroughly pipette to resuspend the magnetic beads, and let stand at room temperature for 2 minutes;
[0137] Place the S216 and EP tubes on a magnetic rack and collect the purified DNA into a sterile, enzyme-free 0.2 ml PCR tube;
[0138] S217, add 1 μl of reagent RAP, flick to mix, centrifuge briefly, and let stand at room temperature for 5 minutes;
[0139] S218. Add 34 μl of reagent SQB, 25.5 μl of reagent LB, and 4.5 μl of sterile enzyme-free deionized water, flick to mix, and centrifuge briefly.
[0140] It should be noted that in the above step S3, the DNA sequencing process specifically includes the following sub-steps:
[0141] S31. Insert the FLO-MIN106D R9.4 chip into the MinION sequencer, connect it to a laptop, and use the MinKNOW software Flowcell Test module to test the chip. If the number of nanopores on the new chip is greater than 800, it is a qualified chip.
[0142] S32, disconnect the sequencer from the computer, and place the sequencer in the clean bench for sample loading;
[0143] S33, add 30 μl of reagent FLT to one tube of reagent FLB, mix by pipetting, and obtain a mixed reagent of FLB+FLT; both FLB and FTL are reagents attached to the BSCC45S1E kit;
[0144] S34, opening the triggering hole and sucking out the air in the hole;
[0145] S35, add 800 μl of the above-mentioned FLB+FLT mixed reagent, close the initiation hole, and let stand at room temperature for 5 min;
[0146] S36, open the initiation well and the loading well, add 200 μl of FLB+FLT mixed reagent to the initiation well, and drop 75 μl of library into the loading well;
[0147] S37, the sequencer was connected to a laptop computer and sequencing was performed using MinKNOW software.
[0148] In S37, the program is set according to the kit, and data analysis is performed in real time. After the clinical pathogen sequence appears, sequencing is continued for 1-2 hours. If no other pathogen sequence appears, sequencing is stopped and the connection between the computer and the sequencer is disconnected.
[0149] In the above step 3, a chip cleaning kit (EXP-WSH004, Oxford Nanopore Technology Co., Ltd., UK) was used to clean the chip. All operations were performed in a clean bench. All instruments except the chip were exposed to ultraviolet light for 30 minutes. The specific cleaning process included: first, 2 μl of reagent WMX and 398 μl of reagent DIL were mixed into a cleaning mixture, and the mixture was mixed by blowing; second, the initiation hole was opened, a 1000 μl pipette was adjusted to a range of 800 μl, and after insertion into the initiation hole, 20-30 μl was reversed until the storage solution entered the bottom of the pipette tip and the air in the hole was sucked out; third, 400 μl of the cleaning mixture was added to the initiation hole, the initiation hole was closed, and the mixture was allowed to stand at room temperature for 30 minutes; fourth, the initiation hole was opened, 500 μl of storage reagent S was added to the initiation hole, and the initiation hole was closed; fifth, the waste liquid was sucked out from the waste liquid hole; sixth, the chip was returned to the packaging bag and stored at 4°C.
[0150] In the above-mentioned step S4, the process of screening and analyzing the collected data includes: first performing base calling, then performing data filtering after base calling, then performing data quality control after data filtering, performing host gene comparison and removal, and then performing classification analysis.
[0151] Furthermore, guppy was used for base calling, NanoFilt for data filtering, NanoPlot for data quality control, minimap2 for host gene alignment removal, and kraken2 for classification analysis. All analyses can be performed on the local server and can be completed in about 4-6 minutes, reducing the potential risk of data leakage caused by uploading data to the London server in the UK.
[0152] In the above S4 step, guppy is used for base calling, and the implementation process of base calling and data quality control is: write the script guppy.sh, and perform base calling on the raw data file of MinION sequencing. Base calling is performed using guppy (v.3.2.10), raw data quality statistics are performed using the summarizeFastq.pl program developed by Dickson Laboratory of the University of Michigan, and raw data quality visualization analysis (data quality control) is performed using NanoPlot (v.1.32.1). It should be noted that the specific implementation process of base calling is shown in Example 1.
[0153] The data filtering process in the above-mentioned S4 step includes: writing a script datafilt.sh, removing low-quality sequences with a quality value (Q value) less than 7 in the original data, and cutting off the adapter sequence (head adapter sequence 150bp, tail adapter sequence 50bp). Data filtering is performed using NanoFilt (v.2.7.1), and the data quality statistics after filtering are performed using the summarizeFastq.pl program developed by Dickson Laboratory of the University of Michigan, USA, and data quality visualization analysis is performed using NanoPlot (v.1.32.1). It should be noted that the implementation process of data filtering and data quality control is shown in Example 2.
[0154] Data filtering results: The raw data obtained by sequencing was filtered to obtain data that can be used for sequencing. After filtering, the number of sequences and the total number of bases in all samples were reduced. The maximum length of most sequences was reduced by 200bp, and the minimum length of the sequence increased from about 100bp to 500bp. The average sequence length increased to varying degrees. Figure 4 .
[0155] In the above step S4, the host gene comparison and elimination process includes:
[0156] Write a script mapping.sh to perform host gene alignment and elimination on the sequencing data after data filtering and data quality control. Data alignment is performed using minimap2 (v.2.17), and data screening is performed using samtools (v.1.11). The alignment genome is the GRCh38 human genome sequence downloaded from NCBI. It should be noted that the specific process of host gene comparison and removal is shown in Example 3.
[0157] The result of host gene data elimination is: the host gene is eliminated again by using the minimap2 alignment method on the sequencing data to obtain pathogen sequence information with higher abundance. Figure 5 As shown, the host gene information of all samples has been reduced to varying degrees. All host gene information in 20 samples has been eliminated, leaving only microbial gene sequence information, which provides more accurate data for subsequent classification comparison and species identification and reduces the time required for classification comparison.
[0158] In the above step S4, kraken2 is used for classification comparison and analysis, and the process is as follows: write the script kraken2_classification.sh and use the kraken2 (v.2.1.1) program to classify and compare the pathogen sequence information.
[0159] In the above step S5, pathogen gene splicing is performed on the processed sample, and the process is as follows:
[0160] Pathogen genome assembly from Mikong sequencing data:
[0161] The script bacAssemble_ONT.sh was written to screen the target pathogen sequence and assemble the genome of the sample Nanopore sequencing data. The assembly was performed using the flye (2.8.3) program, and the correction was performed using medaka (v.1.4.2).
[0162] In the above step S6, the successfully spliced bacteria are tested for drug resistance, including: using the web version of Resistance Gene Identifier (https: / / card.mcmaster.ca / analyze / rgi) to identify the pathogen resistance gene of the successfully spliced bacteria.
[0163] 1. Select DNA sequence at Select Data Type.
[0164] 2. Upload the spliced file at Upload FASTA sequence file(s);
[0165] 3. Select Perfect and Strict hits only in Select Criteria;
[0166] 4. Select Exclude nudge at Nudge ≥ 95% identity Loose hits to Strict;
[0167] 5. Select High quality / coverage for Sequence Quality.
[0168] 6. Click Submit to obtain the analysis results. The results are judged as Perfect and Strict. The drug-resistant genes with a matching region consistency percentage (% Identity of matching region) greater than 90% and a reference gene length percentage (% Length of reference sequence) greater than 80% are considered qualified drug-resistant genes.
[0169] In order to verify the sequencing results and observe their accuracy, the following processes are included:
[0170] (1) Validate the sequencing results using qRT-PCR.
[0171] Use Hieff Universal Blue qPCR SYBR Master Mix (11184ES08) was used for qRT-PCR. The whole operation was protected from light. Sterile, enzyme-free, and non-autoclaved pipette tips and EP tubes were used for the operation. The operation method was carried out according to the instructions:
[0172] 1. Entrust Ruibo Xingke Company to synthesize primers. The primer sequences are shown in Table 3;
[0173] 2. Prepare the mixed solution according to the volume of 3 replicate wells for each sample and each target gene. The content of the mixed solution in each well is as follows: 10μl SYBR + 0.4μl forward primer + 0.4μl reverse primer + 7.2μl sterile enzyme-free deionized water;
[0174] 3. Add samples according to the amount of 18μl mixed solution + 2μl DNA to each well. Set up 1 negative control group, 1 blank control group and 1 positive control group for each addition;
[0175] Table 3 Primer sequences
[0176]
[0177]
[0178] (2) Amplification time is performed to verify the sequencing results.
[0179] Set the PCR amplification program according to the method described in Table 4;
[0180] Table 4 PCR amplification program time setting table
[0181]
[0182] A sample with a CT value greater than 35 or a negative control group result is considered negative, a sample with a CT value between 30 and 35 or between 30 and the negative control group is considered suspicious, and a sample with a CT value less than 30 is considered positive.
[0183] It should be noted that the diagnostic criteria for pathogens in the samples of the present invention are as follows:
[0184] 1. Microbial culture: three-zone streaking, culturing at 37°C for 24-48 hours and visible colonies are defined as positive. Semi-quantitative results of bacterial culture are distinguished by the number of visible colonies in the three zones;
[0185] Table 5 Definition of bacterial semi-quantitative results by three-zone method
[0186]
[0187] 2.qRT-PCR:
[0188] 1) Positive: CT value <30, and less than the negative control group, the result is statistically significant;
[0189] 2) Positive and suspicious: the CT value is between 30 and 35, which is lower than that of the negative control group, and the result is statistically significant;
[0190] 3) Negative: CT value is higher than 35, or higher than the negative control group or there is no statistical difference with the negative control group.
[0191] 3. Metagenomics:
[0192] 1) Positive: metagenomic results show that the number of pathogen sequences is greater than 1 and greater than 1% of all pathogen sequences;
[0193] 2) Positive suspicion: There is only one pathogen sequence, but it is greater than 10% of all pathogen sequences
[0194] 4. Criteria for judging the presence of pathogens in samples:
[0195] 1) Positive microbial culture;
[0196] 2) qRT-PCR positive;
[0197] 3) qRT-PCR suspected positive combined with metagenomics positive or suspected positive.
[0198] Example 1
[0199] During the base calling process, the specific script code is as follows:
[0200] #! / bin / bash
[0201] module load conda / miniconda2
[0202] read-p"Enter Sample ID:"sampleID#Enter sample ID
[0203] #Use guppy for base calling
[0204] time guppy_basecaller_
[0205] --input_path#input_path\#Set the data input path, that is, the original fast5 data save path
[0206] --recursive\#Recursive processing
[0207] --save_path#output_path\#Set the output data saving path, that is, the fastq data saving path
[0208] --flowcell FLO-MIN106\#Sequencing chip number
[0209] --kit#kitID_for_library_built\#Kit ID for library construction
[0210] --device'cuda:0'#sequencing device;
[0211] cat*.fastq>>${sampleID}.total.fastq#Integrate all sample data into one fastq file;
[0212] summarizeFastq.pl${sampleID}.total.fastq>
[0213] ${sampleID}summarize_rawdata.txt
[0214] #Use the summarizeFastq.pl script developed by Dickson Laboratory at the University of Michigan to perform raw data statistics
[0215] The statistical data is saved in the ${sampleID}summarize_rawdata.txt file.
[0216] Example 2
[0217] This embodiment provides a specific implementation process of data filtering and data quality control. The specific script code is:
[0218] #! / bin / bash
[0219] module load conda / miniconda2
[0220] read-p"Enter Sample ID:"sampleID#Enter sample ID
[0221] source activate nanopack#Activate the data analysis environment;
[0222] time NanoFilt\#Use NanoFilt to filter data and calculate time
[0223] ${sampleID}.total.fastq\#Enter the data to be filtered
[0224] -l 500\#Filter out sequences less than 500bp
[0225] -q 7\#Minimum average quality value
[0226] Filter out low-quality sequences with quality values less than 7
[0227] --headcrop 150\#Cut off 150bp of the linker sequence from the head
[0228] --tailcrop 50\#Cut off 50bp of the end sequence from the tail
[0229] >${sampleID}.clean.fastq\#The result is output to ${sampleID}.clean.fastq file;
[0230] summarizeFastq.pl${sampleID}.clean.fastq>
[0231] ${sampleID}summarize_cleandata.txt
[0232] time NanoPlot\#Use NanoPlot to visualize data quality
[0233] --threads 6\#Analyze the number of threads used
[0234] --outdir NanoPlotSummary\#Output result directory
[0235] --N50\#Show the N50 mark in the sequence read length histogram
[0236] --title${sampleID}.clean\#Set the title of the output graph to ${sampleID}.clean
[0237] --fastq${sampleID}.clean.fastq\#The file to be analyzed is
[0238] ${sampleID}.clean.fastq
[0239] --plots hex dot kde pauvre\#File drawing type: kde (contour map), dot (point map), hex (hexagonal map with different color depths), pauvre (bar chart with coordinate axes on both sides), --color green#The color of the points and bar charts is green;, conda deactivate#Exit the data analysis environment.
[0240] Example 3
[0241] This embodiment provides a specific implementation process of host gene comparison and elimination. The specific script code is:
[0242] #! / bin / bash
[0243] module load conda / miniconda2
[0244] read-p"Enter Sample ID:"sampleID#Enter sample ID
[0245] source activate minimap2
[0246] #Download and build human genome index
[0247] There is no need to build again after a successful build.
[0248] wget
[0249] https: / / ftp.ncbi.nlm.nih.gov / refseq / H_sapiens / annotation / GRCh38_latest / refseq_identifiers / GRCh38_latest_genomic.fna.gz~ / minimap2 / #Download
[0250] GRCh38 human genome sequence to ~ / minimap2 / path;
[0251] minimap2-d~ / minimap2 / human.min
[0252] ~ / minimap2 / GRCh38_latest_genomic.fna.gz#
[0253] GRCh38_latest_genomic.fna.gz is the reference genome
[0254] Use minimap2 to build the human genome index human.min;
[0255] #Use minimap2 to perform sequencing data comparison and analysis.
[0256] minimap2-ax map-ont\#Set the sequencing instrument to ont
[0257] The output file format is sam
[0258] ~ / minimap2 / human.min\#Set reference index
[0259] ${sampleID}.clean.fastq\#Set input file
[0260] >${sampleID}.clean.sam\#Set the output file;
[0261] samtools view\#Use Samtools view to convert file formats
[0262] -Sb\# Input sam format file
[0263] Output bam format file
[0264] -o${sampleID}.clean.bam\#Set output file
[0265] ${sampleID}.clean.sam\#Set the input file;
[0266] #Use samtools to process alignment sequences and extract pathogen sequences.
[0267] samtools sort\#Use Samtools sort to sort the above files
[0268] -O bam\#Output file format
[0269] -o${sampleID}.clean.sorted.bam\#Set output file
[0270] ${sampleID}.clean.bam\#Set the input file;
[0271] samtools view-f 4${sampleID}.clean.sorted.bam>
[0272] ${sampleID}.nohuman.bam#Use Samtools view to extract the unaligned non-human genome sequences;
[0273] samtools fastq${sampleID}.nohuman.bam>
[0274] ${sampleID}.nohuman.fastq#Use Samtools fastq to
[0275] Convert the ${sampleID}.nohuman.bam file into the ${sampleID}.nohuman.fastq file to obtain the sequence information of pathogens without human genome;
[0276] conda deactivate#Exit the classification analysis environment.
[0277] Example 4
[0278] This embodiment provides a specific implementation process for the classification analysis of pathogen sequence information after host gene comparison and removal. The specific script code is:
[0279] #! / bin / bash
[0280] module load conda / miniconda2
[0281] read-p"Enter Sample ID:"sampleID#Enter sample ID
[0282] source activate classification
[0283] #Build the kraken2 standard reference genome database
[0284] Build once, no need to build again
[0285] But the database needs to be updated from time to time.
[0286] mkdir StandardK2IDX#Create index path;
[0287] kraken2-build --standard --threads 24 --db StandardK2IDX#Download all build-related taxonomy and other files from NCBI and build the index
[0288] Choose to use 24 threads for processing;
[0289] #Use kraken2 to analyze pathogen classification information
[0290] recordStats kraken2\#Use kraken2 to perform classification analysis and record the server status and time consumed during the analysis process
[0291] --db StandardK2IDX\#Use the standard reference genome database for classification comparison
[0292] --threads 10\#Use 10 threads for processing
[0293] --output${sampleID}.cleanOut.txt\#The output file of sequence analysis result is
[0294] ${sampleID}.cleanOut.txt
[0295] --report${sampleID}.cleanReport.txt\#The output file of classification analysis result is
[0296] ${sampleID}.cleanReport.txt
[0297] ${sampleID}.clean.fasta#The file to be analyzed is ${sampleID}.clean.fasta;
[0298] grep-w S${sampleID}Report.txt|sort-nk 3>
[0299] ${sampleID}_species.txt#Get the species information in the results and sort them according to the number of reads
[0300] The results are output to the ${sampleID}_species.txt file;
[0301] conda deactivate#Exit the classification analysis environment.
[0302] Example 5
[0303] This embodiment provides a gene splicing process, and the specific script code of gene splicing is:
[0304] #! / bin / bash
[0305] module load conda / miniconda2
[0306] read-p"Enter Sample ID:"sampleID#Enter sample ID
[0307] read-p"Enter Abbreviation of Pathogene:"Abbre#Enter the abbreviation of the pathogen to be detected
[0308] read-p"Enter Abbreviation of Pathogene:"TaxonyID#Enter the species taxonomy of the pathogen to be detected
[0309] #For pathogen fastq sequence screening
[0310] Pick out the fastq sequences of pathogens from the kraken2 analysis results
[0311] #Find pathogen sequences from all comparison analysis results and integrate them into a new file
[0312] mkdir${Abbre}
[0313] cd kraken2 /
[0314] awk'{if($3=='${TaxonyID}')printf$2"\n"}'
[0315] ${sampleID}_clean_nohumanOut.txt>${sampleID}${Abbre}TitleOut.txt
[0316] Line=$(cat${sampleID}${Abbre}TitleOut.txt|wc-l)
[0317] cd..
[0318] for((i=1;i<=$ 131 ; i++))#for loop statement
[0319] do
[0320] Title=$(sed-n${i}p kraken2 / ${sampleID}${Abbre}TitleOut.txt)#Get the Nth line in the Title file of the selected target pathogen
[0321] grep"${Title}"${sampleID}_clean_nohuman.fastq-A 3>>
[0322] ${Abbre} / ${sampleID}${Abbre}.fastq#Use Title to capture the sequence and grab the three lines after the title (4 lines in total) to the specified file
[0323] done
[0324] #Convert fastq file to fasta file
[0325] For the next step
[0326] cd${Abbre}
[0327] source activate nanopack
[0328] seqtk seq-a${sampleID}${Abbre}.fastq>${sampleID}${Abbre}.fasta#Convert fastq file to fasta file
[0329] conda deactivate
[0330] # Genome assembly using flye
[0331] source activate minimap2
[0332] mkdir antibiotic
[0333] mkdir antibiotic / flye
[0334] flye--nano-raw${sampleID}${Abbre}.fasta-o antibiotic / flye
[0335] # Use medaka to correct the assembly results
[0336] cd antibiotic
[0337] mkdir medaka
[0338] medaka_consensus-i.. / ${sampleID}${Abbre}.fasta-d flye / assembly.fasta-o
[0339] medaka / -t 4>medaka / medaka.log
[0340] #-i Input the original fasta file before splicing
[0341] #-d Input the spliced fasta file
[0342] #-o assembly result error correction result output directory
[0343] #-t Number of threads used
[0344] conda deactivate.
[0345] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for processing samples of lower respiratory tract aspirates for ventilator-associated pneumonia, It is characterized in that The following steps are involved: S1, sample processing; Including S11. Using saponin to remove sample host genes; S12. Batch extraction of DNA; S2, constructing a multi-sample PCR-amplified DNA sequencing library; S3, sequencing the DNA in the DNA sequencing library; DNA sequencing was performed using the MinION nanopore sequencer, with data collected in real time; S4. Screen and analyze the collected data; The screening and analysis of the collected data in step S4 includes: using guppy for base calling, using NanoFilt for data filtering, using NanoPlot for data quality control, using minimap2 for host gene alignment removal, and using kraken2 for classification analysis; The data filtering process in the above step S4 includes: removing low-quality sequences with a quality value less than 7 in the original data and cutting off the adapter sequence; after filtering, the number of sequences and the total number of bases of all samples are reduced, the maximum length of most sequences is reduced by 200 bp, the minimum length of the sequence is increased from about 100 bp to 500 bp, and the average sequence length is increased to varying degrees; The host gene alignment removal process in the above S4 step includes: using the minimap2 alignment method to remove host genes from the sequencing data again, all host gene information is removed, and only microbial gene sequence information remains, providing more accurate data for subsequent classification alignment and species identification; S5. Perform pathogen gene splicing on the processed samples.
2. The method for processing samples of lower respiratory tract aspirates of ventilator-associated pneumonia according to claim 1, It is characterized in that In S11, the process of removing the sample host gene by using saponin includes the following sub-steps: S111. Resuspend the clinical lower respiratory tract sample pellet with 250 μl sterile PBS buffer and mix thoroughly by pipetting; S112, add 200 μl of 5% sterile saponin solution and dissolve it in sterile enzyme-free deionized water, vortex for 15 seconds to mix, and let stand at room temperature for 10 minutes to allow the saponin solution and sample to be fully mixed. Saponin can cause host cells to swell and rupture; S113, add 350 μl of sterile enzyme-free deionized water and let stand at room temperature for 30 seconds; S114, add 12 μl of 5 M NaCl solution and mix by inverting to stop the swelling effect of saponin; S115, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant; S116. Resuspend the precipitate with 100 μl sterile PBS buffer, add 100 μl HL-SAN buffer, vortex for 15 seconds to mix, add 10 μl HL-SAN enzyme, invert to mix, and digest free nucleic acids; S117, 37℃ water bath, high speed shaking for 15min; S118, centrifuge at 8000 rpm at 4°C for 5 min and discard the supernatant; S119, add 800 μl sterile PBS buffer, mix by inversion, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant; S1110, add 1000 μl sterile PBS buffer, mix by inversion, centrifuge at 8000 rpm at 4°C for 5 min, and discard the supernatant.
3. The method for processing samples of lower respiratory tract aspirates of ventilator-associated pneumonia according to claim 2, It is characterized in that In step S116, the HL-SAN buffer solution is prepared by adding 0.2 g MgCl 2 -6H 2 O was dissolved in 5 M NaCl, the volume was adjusted to 10 ml and filtered using a 0.22 μm filter.
4. The method for processing samples of lower respiratory tract aspirates of ventilator-associated pneumonia according to claim 3, It is characterized in that In S12, DNA extraction is performed using a BSCC45S1E kit and an automatic DNA extractor.
5. The method for processing samples of lower respiratory tract aspirates of ventilator-associated pneumonia according to claim 4, It is characterized in that In said S12, the DNA extraction process includes the following sub-steps: Dissolve all the lysozyme in the S121 and BSCC45S1E kits in TET buffer and mix thoroughly by shaking; S122, resuspend the pellet in 180 μl of the mixture, shake and mix, and incubate at 37°C for 30 min; S123, add 20 μl of proteinase K and bacterial resuspension solution to columns 1 and 7 of the kit; S124. Place the reagent kit in the machine and set the program: S125. After the extraction is completed, use Nanodrop to identify the DNA concentration and purity, and use Qubit to accurately quantify the concentration.
6. The method for processing samples of lower respiratory tract aspirates of ventilator-associated pneumonia according to claim 1, It is characterized in that In S2, the rapid barcode kit SQK-RBK004 was used to construct a multi-sample rapid DNA sequencing library. All operations were performed in a clean bench, and all instruments were exposed to ultraviolet light for at least 30 minutes. All reagents were melted on ice, centrifuged before use, and mixed by pipetting.
7. The method for processing samples of lower respiratory tract aspirates of ventilator-associated pneumonia according to claim 6, It is characterized in that In S3, during DNA sequencing, the chip priming kit EXP-FLP002 was used for chip priming, all operations were performed in a clean bench, and all instruments except the chip were UV-irradiated for 30 min; FLO-MIN106D R9.4 chip and MinION sequencer were used for sequencing.
8. The method for processing samples of lower respiratory tract aspirates of ventilator-associated pneumonia according to claim 7, It is characterized in that In S3, the DNA sequencing process specifically includes the following sub-steps: S31. Insert the FLO-MIN106D R9.4 chip into the MinION sequencer, connect it to a laptop, and use the MinKNOW software Flowcell Test module to test the chip. If the number of nanopores on the new chip is greater than 800, it is a qualified chip. S32, disconnect the sequencer from the computer, and place the sequencer in the clean bench for sample loading; S33, add 30 μl of reagent FLT to one tube of reagent FLB, and mix by pipetting; S34, opening the triggering hole and sucking out the air in the hole; S35, add 800 μl FLB+FLT mixed reagent, close the initiation hole, and let stand at room temperature for 5 min; S36, open the initiation well and the loading well, add 200 μl of FLB+FLT mixed reagent to the initiation well, and drop 75 μl of library into the loading well; S37, the sequencer was connected to a laptop computer and sequencing was performed using MinKNOW software.
Citation Information
Patent Citations
Metagenome sample library construction method and identification method based on nanopore sequencing platform and kit
CN109487345A