Multi-target pathogenic microorganism analysis method based on third-generation targeted sequencing data
By employing multi-target design and data quality control correction methods, the accuracy and complexity issues of pathogen detection in third-generation sequencing technology have been resolved, enabling accurate identification and convenient analysis of pathogens.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN EASYDIAGNOSIS BIOMEDICINE
- Filing Date
- 2022-04-08
- Publication Date
- 2026-04-10
AI Technical Summary
Existing third-generation sequencing technologies suffer from problems such as high sequencing error rates, insufficient accuracy in single-target detection, and high complexity in multi-target detection, resulting in low accuracy in pathogen identification.
By employing a multi-target design, a multi-target-associated error-correcting reference sequence library is constructed by targeting and clustering the pathogenic microorganism reference genome. Combined with the quality control and correction of the original sequencing data, the sequence is compared one by one with the pathogenic microorganism detection library to achieve accurate detection of pathogenic microorganisms.
It improves the accuracy and sensitivity of pathogen detection, reduces the complexity of sequencing and alignment results, simplifies the analysis process, reduces false positive results, and ensures the accurate detection of key pathogens.
Smart Images

Figure BDA0003587256030000081 
Figure BDA0003587256030000091 
Figure BDA0003587256030000092
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of pathogenic microorganism detection, and particularly relates to a multi-target pathogenic microorganism analysis method based on third-generation targeted sequencing data. BACKGROUND
[0002] Microorganisms pathogenic to humans and animals are called pathogenic microorganisms, also known as pathogens, including viruses, bacteria, rickettsia, mycoplasma, chlamydia, spirochete, fungi, and actinomycetes. These pathogenic microorganisms can cause diseases such as infection, allergy, tumor, and dementia, and are one of the main factors endangering food safety. Some pathogenic microorganisms are extremely contagious and often cause worldwide epidemics, so the detection of pathogens must be rapid and accurate. With the continuous development of medical microbiology research technology, pathogenic diagnosis is no longer limited to the pathogen level, and detection methods at the molecular and gene levels are constantly emerging and being applied in clinical and laboratory settings. With the development of science and technology, gene detection technologies that do not rely on traditional microbial culture and can quickly and objectively detect suspected pathogenic microorganisms (including bacteria, fungi, and viruses) in clinical samples have gradually replaced other detection technologies and become the mainstream detection technology for pathogens in clinical laboratories and basic laboratories.
[0003] For the detection of pathogenic microorganisms, second-generation sequencing is currently the mainstream general platform. Second-generation sequencing data is highly accurate and has high sequencing throughput, but has the problems of short read length and long sequencing time. Generally, the sequencing read length is only 150 bp, and the sequencing time is more than 12 h. Compared with second-generation sequencing, third-generation sequencing has a very long read length, which can reach kb level and above, but has a high sequencing error rate. The current mainstream third-generation sequencing platforms are PacBio platform and Nanopore platform. Among them, the Nanopore Sequencing technology sequences the base composition by predicting the current change of single-molecule DNA (RNA) passing through a biological nanopore, and has the advantages of small and portable instrument, simple operation, fast sequencing speed, and longer sequencing read length (average read length up to 20 kb, maximum up to Mb level), which shows great application potential in pathogenic microorganism clinical detection.
[0004] At present, the pathogenic microorganism detection based on the second generation sequencing uses short fragment free DNA or fragmented DNA for library construction, and uses the second generation sequencing technology to generate short sequences, so that the database comparison often compares multiple positions or multiple species. The third generation sequencing represented by Oxford Nanopore Technology sequencing technology has the advantages of rapid library preparation and ultra-long read length, which can effectively improve the accuracy of pathogenic microorganism detection in clinical samples. Limited by the requirement of the third generation sequencing for the total amount of sample DNA and the integrity of DNA, most of the pathogenic microorganisms on the market are identified by sequencing the target 16S, ITS, 23S, 16S-23S or specific fragments. Since the molecular tags used for molecular identification are universal target fragments, these fragments have high consistency within the species and high conservation between species, and the length of the target region fragments between different species is required to be consistent in order to design universal target primers.
[0005] The most common target fragments at present are 16S, 23S and ITS fragments between 16S and 23S fragments. These fragments have relatively high conservation and very high homology, and even have very similar 16S, 23S or ITS sequences in some different genus species. In addition, limited by the extremely low content of pathogenic microorganisms in clinical samples, in order to improve the detection limit of the detection product, relatively small fragments are selected as molecular tags for identification during target fragment amplification. Therefore, the use of these single short molecular tags alone cannot accurately distinguish to the species level in many cases, and since these fragments can be compared to multiple similar species at the same time, the accuracy of pathogenic microorganism determination is greatly hindered.
[0006] Since there are certain limitations in the detection of individual target groups, multiple targets can theoretically more truly restore the sample composition. Therefore, the present application adopts multiple target design to realize accurate determination of pathogenic microorganisms. For the design of multiple targets, it is necessary to have the characteristics of universality between species and specificity of individual species, therefore, when selecting targets, large-scale sequence alignment is required to find specific sites that can be amplified in different species. At the same time, due to the similarity of sequences, it is easy to detect different microorganisms on multiple targets of the same species. Therefore, the design of multiple targets will inevitably increase the complexity of detection, making it difficult to restore the true situation of the sample. SUMMARY
[0007] In view of the above problems, the present application develops a pathogenic microorganism analysis method for multiple target fragments, which can truly restore the multiple target detection results of the sample to be tested, and realize accurate detection of pathogenic microorganisms.
[0008] To achieve the above object, the technical scheme of the present application is as follows:
[0009] A pathogenic microorganism analysis method based on long read sequencing data, comprising the following steps:
[0010] S1, intercepting and downloading a reference genome of a pathogenic microorganism to obtain a targeted region reference sequence, wherein each pathogenic microorganism has at least two targeted region reference sequences, the reference sequences are corrected by clustering according to similarity, and the reference sequences are associated at the species level to obtain a multi-target associated error correction reference sequence library;
[0011] S2, quality control of raw sequencing data;
[0012] S3, aligning the sequencing data obtained in step S2 to a pathogenic microorganism identification database to obtain a preliminary microorganism detection list after quality control;
[0013] S4, extracting all reference sequences corresponding to the species on the microorganism detection list in S3 from the error correction reference sequence library, constructing a pathogenic microorganism detection library with multi-target association, and then re-aligning the sequencing data obtained in step S2 to the pathogenic microorganism detection library one by one for determination to obtain a final microorganism detection list.
[0014] Further, in the above technical scheme, in step S1, the reference sequences are adjusted for direction consistency after clustering according to the matching order of the primers, so as to obtain a multi-target error correction reference database with consistent direction.
[0015] The clustering correction in step S1 is to merge completely identical reference sequences, and to count the similarity between the reference sequences, and to filter and remove the sequences with intra-species similarity greater than inter-species similarity. Specifically, the threshold of similarity is in the range of 95% to 99.9%.
[0016] The downloading method in step S1 can be to download reference genome sequences from databases such as NCBI, and to intercept the matched reference sequences according to the matching of the targeted region primers and the length of the amplified products after matching. It can be understood that the primers are primers used in the PCR amplification process before the sample is tested.
[0017] Further, in the above technical scheme, step S2 specifically comprises the following steps:
[0018] S21, filtering out sequencing sequences with length less than m or mass value less than n, preferably, m is in the range of 100-600 bp, and n is in the range of 8-11; after quality control of the original sequencing data, sequencing data with long length and high quality can be obtained; the data amount, average length, N50 length and sequencing sequence quality value of the high-quality sequencing data after filtering are taken as evaluation indexes of the sequencing data quality value;
[0019] S22, removing sequencing sequences aligned to the host genome; when the host is a human, the sequencing data obtained in step S21 is aligned to the human reference genome, the data aligned to the human reference genome is considered as non-analytical human data, and the human data is filtered out; the proportion of the human data in the sequencing data after quality control is taken as a reference index of the library construction sequencing quality reliability;
[0020] S23, aligning with the error correction reference sequence library, and correcting the sequencing sequences aligned to the same reference sequence;
[0021] S24, comparing the corrected sequencing sequences with each other, and merging similar sequences.
[0022] In step S23, the correction method can be: first, filtering out sequencing sequences with error rate higher than a or coverage rate lower than b, wherein a is in the range of 5%-10%, and b is in the range of 50%-95%; then, taking the reference sequence as a reference, the base frequency of each site in the sequencing sequence is counted, and when the sites with sequencing depth lower than c account for more than d of the sequence, the sequencing sequence is removed; in the remaining sequencing sequence, for the sites with sequencing depth lower than c, the corresponding base in the target region reference sequence is replaced, and for the sites with sequencing depth greater than c, the base with the highest base frequency is adopted, wherein c is in the range of 3-7, and d is in the range of 10%-30%. It can be understood that the error rate is the error rate when the sequencing sequence is compared with the reference sequence, and the coverage rate is relative to the target region.
[0023] In step S24, the merging method can be: for two sequences reaching a similarity threshold, the sequence with higher frequency is retained, and the frequency of the sequence is updated to the sum of the frequencies of the two sequences; the similarity threshold is any value in the range of 95%-99.9%.
[0024] In the present application, the similarity can be calculated by the Levenshtein.ratio module of python.
[0025] Further, step S2 further comprises: S25, aligning the merged sequencing sequence to the reagent background microorganism database, and removing the aligned sequencing data.
[0026] The removal parameters are the number of sequences and the number of mismatches in the statistical comparison to the background microorganism database, and the sequencing sequences with less than 3 mismatches are filtered out, so as to remove the microorganisms derived from the sequencing reagents or consumables from the identification results. In addition, the construction of the reagent background microorganism database is a prior art, which will not be described here.
[0027] Further, in the above technical solution, the parameters for quality control in step S3 are: retaining the aligned sequences with coverage greater than e and alignment error rate less than f, wherein the value range of e is 85% to 99%, and the value range of f is 0.5% to 3%.
[0028] In addition, while obtaining the preliminary microorganism detection list in step S3, the targeted region information, species group, number of aligned sequences, sequence length, and number of mismatches in the alignment are counted and the merged high-quality sequences are output as basic data quality control indicators for data analysis and determination.
[0029] Further, in the above technical solution, the parameters for determination in step S4 are: retaining the species with a sequence number ratio of multiple target simultaneous detection or single target detection to the total sequence number being h and above, wherein the value range of h is 2% to 8%.
[0030] Due to the difference in primer matching during targeted amplification, there may be a case where only part of the targeted fragments are detected in some samples, so in actual application, the pathogenic microorganisms can be classified and focused on: for the pathogenic microorganisms that are focused on, only one targeted fragment is detected to be considered as accurate detection; for background or colonized microorganisms and other non-focus groups, multiple targeted fragments are simultaneously detected to be considered as detection.
[0031] Further, due to the low quality of the original data of Nanopore sequencing data, there is a certain proportion of error splitting when the data is split, and due to the unavoidable aerosol pollution and environmental background microorganism pollution in the experimental process, the microorganism detection list obtained in step S4 can be filtered. The filtering parameters can be:
[0032] Filtering out groups in the identification results with Reads below x, preferably, the value range of x is 5 to 20; this part of data has low reliability and may be derived from environmental pollution, aerosol, reagent background or false matching due to low sequence quality during alignment;
[0033] Filtering out groups in the identification results with sequencing Reads number below the total sample sequencing Reads number y, preferably, the value range of y is 0.2% to 2%; this part of data has low reliability due to low Reads number.
[0034] The detected microorganisms can also be further determined based on the sample source to distinguish the final detected pathogenic microorganisms from the colonized microorganisms.
[0035] The beneficial effects of the present application are:
[0036] 1) The present application can realize accurate detection of species composition when multiple targets are sequenced, improving the accuracy and sensitivity of data analysis.
[0037] 2) By quality control and correction of raw sequencing data, the data analysis speed is effectively improved (2Gb third-generation sequencing data analysis can be completed in a few minutes and the final pathogenic microorganism detection report is generated), and the complexity of the sequencing alignment result is reduced, making the analysis result more simple and accurate, easy to perform subsequent screening, filtering and interpretation, and also effectively reducing the false positives of data analysis.
[0038] 3) The present application first constructs a preliminary species detection list based on the optimal alignment results of a single fragment, and then constructs a detection library based on the detection list to re-align all relevant target region databases, solving the problem of aligning multiple targets in a unified species to different microorganisms, greatly improving the accuracy of pathogenic microorganism detection.
[0039] 4) The present application can realize the key monitoring of key pathogenic microorganisms, ensure the detection of suspicious samples, and at the same time strictly control non-key microorganisms, improve the accuracy of detection of such microorganisms, and avoid interfering with clinical diagnosis. DETAILED DESCRIPTION
[0040] The technical solutions of the present application will be described below in conjunction with the embodiments in the present application. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.
[0041] Example 1
[0042] This example selects a lung lavage fluid sample to detect bacteria and fungi present in the sample simultaneously. In this example, the 16S double-target region of bacteria and the ITS full-length region of fungi are sequenced by third-generation sequencing. Then, based on the sequencing data, pathogenic microorganism analysis is performed, and the process is as follows:
[0043] (1) Establishment of multi-target error correction reference sequence library
[0044] According to the target region selected during sequencing, the target region reference sequences of pathogenic microorganisms are obtained from databases such as NCBI, and the targets are associated at the species level. The above reference sequences are clustered and corrected according to 99% similarity, and the direction of the reference sequences is adjusted to be consistent, obtaining a high-quality multi-target error correction reference sequence library.
[0045] (2) Processing of raw sequencing data
[0046] After sequencing, the sequencing data is quality controlled using NanoStat to determine the sequencing data yield accuracy and length information. The sequencing data folder is "20220113_1331_X1_FAR52513". The folder is generated by the sequencer system and contains information such as sequencing time, chip on-machine position, chip number, etc. The folder contains sequencing-related information and raw fastq files.
[0047] The sequencing data is filtered using NanoFit with the following parameters: -q 10, -l 300 (i.e. filtering out data with length less than 300 bp and quality value less than 10). The filtered fastq files are obtained as input files for subsequent analysis. Because the target length varies greatly, the fragment length is screened based on the actual length that can match the reference database. A total of 21.1M data is generated by actual sequencing, and the data amount after data filtering is 21.02M.
[0048] The filtered fastq files are aligned with the human genome to remove the sequenced human genome sequence. The data amount after removing human data is 21M.
[0049] (3) Error correction of sequencing data
[0050] The third-generation sequencing data is aligned to the error correction database using the software minimap2 to obtain the bam file aligned to the reference sequence. The alignment results are corrected, and the coverage and base frequency of each reference sequence at each site are counted. If the coverage is less than 90% or the matching error rate is greater than 8%, the sequencing sequence is discarded. The base site is corrected according to the base frequency of the retained sequence (in this embodiment, c is 5 and d is 20%). The corrected fasta file is generated, and the annotation information of each sequence in the fasta file includes the number of original sequencing sequences obtained from the corrected sequence.
[0051] (4) Merging of sequencing sequences
[0052] The sequencing sequences in the corrected fasta file are aligned two by two, and similar sequences with a similarity greater than 99% are merged to obtain the merged fasta file.
[0053] (5) Generating a preliminary microorganism detection list
[0054] The merged fasta file is aligned to the identification database using the minimap2 software, and the alignment details are counted. The aligned sequences with a coverage greater than 90% and an alignment error rate less than 1% are retained. The identification results are merged according to the species classification level, and a preliminary microorganism detection list is generated according to the annotation database. The preliminary microorganism detection information is shown in Table 1:
[0055] Table 1
[0056]
[0057]
[0058] (6) Generation of final microorganism detection list
[0059] According to the species types in Table 1, the corresponding reference sequences were extracted to construct a pathogenic microorganism detection library. The combined fasta file was re-aligned to the pathogenic microorganism detection library, and the number of aligned sequences, coverage, and alignment similarity were counted. Multiple targets were detected simultaneously, or the single target detection sequence was greater than 5% of the species.
[0060] Then, the groups with less than 10 Reads of aligned sequences were filtered out. The analysis results with less than 0.5% of the total Reads of all samples in the same batch (i.e., all samples sequenced at the same time) aligned to the species were filtered out. The analysis results with less than 1% of the total Reads of the sample were filtered out.
[0061] The final microorganism detection information is shown in Table 2:
[0062] Table 2
[0063]
[0064]
[0065] (7) Removal of reagent background microorganisms
[0066] The combined high-quality sequences were aligned to the reagent background microorganism database, and the number of sequences aligned to the background microorganism database and the number of mismatches were counted. Data with less than 3 mismatches were filtered out. The corresponding microorganism species of the filtered data were determined and deleted to obtain Table 3.
[0067] Table 3
[0068] target sequence number Latin name of species Chinese name 2 593 Chlamydia psittaci Chlamydia psittaci 1 276 Peptostreptococcusstomatis Peptostreptococcus 2 1363 Peptostreptococcusstomatis Peptostreptococcus 2 726 Gemellahaemolysans hemolytic twin cocci 1 1120 Neisseria oralis Oral Neisseria 2 987 Neisseria oralis Oral Neisseria 2 1871 Streptococcus oralis Oral streptococci
[0069] (8) Determination of pathogenic microorganisms and colonizing microorganisms
[0070] According to the sample source, all species identification results that meet the requirements were determined in the pathogenic microorganism and colonizing microorganism database to determine the final detected pathogenic microorganisms and colonizing microorganisms, as shown in Table 4:
[0071] Table 4
[0072] Latin name of species Chinese name determination Chlamydia psittaci Chlamydia psittaci Pathogenic microorganisms Peptostreptococcusstomatis Peptostreptococcus colonizing microorganisms Neisseria oralis Oral Neisseria colonizing microorganisms
[0073] In summary, the accuracy of pathogenic microorganism detection can be greatly improved.
[0074] Comparative Example 1
[0075] In this example, the same sample in Example 1 was detected using a single target method. The detection process was different from that of Example 1 in that:
[0076] In step (1), each pathogenic microorganism has only one target region reference sequence (where the target domain of bacteria is the 16S full-length region, and the target region of fungi is the ITS full-length region);
[0077] In step (6), there is no need to construct a pathogenic microorganism detection library and perform re-alignment. Only 3 types of low-confidence data need to be filtered out.
[0078] Compared with Example 1, this example cannot detect Chlamydophila psittaci.
[0079] Example 2
[0080] In this example, a blood sample was used to detect pathogenic microorganisms present in the sample using single 16S target detection and simultaneous 16S, 16S-23S, and 23S three-target detection. Each detection was repeated twice. Then, based on the sequencing data, a comparative analysis of single-target and multi-target pathogenic microorganism detection was performed.
[0081] (1) Single-target detection of blood samples
[0082] After processing the raw sequencing data, correcting the sequencing data, merging the sequencing sequences, removing the reagent background microorganisms, and detecting the pathogenic microorganisms, the single-target detected pathogenic microorganisms were as follows (process and parameters same as Comparative Example 1):
[0083] Table 5
[0084] content target sequence number Latin name of species Chinese name Test 1 16S 999 Shigellasonnei Shigella sonnei Test 2 16S 2896 Escherichiacoli Escherichia coli
[0085] In the case of different microorganisms detected in two repeated detections, the actual sequencing data was compared and found that both sequencing data were normal. The main problem was that Shigella and Escherichia coli had very similar sequences on the 16S target, and the three-generation sequencing data had a certain sequencing error rate, resulting in different detection results for the same sample.
[0086] (2) Three-target detection of blood samples
[0087] The three-target sequencing data is subjected to raw sequencing data processing, sequencing data error correction, merging of sequencing sequences, and preliminary microorganism detection (process and parameter settings are the same as in Example 1). The preliminary pathogenic microorganisms detected in the three targets are shown in Table 6:
[0088] Table 6
[0089]
[0090] After obtaining the preliminary list of pathogenic microorganisms, the pathogenic microorganism reference library with target association is extracted for re-alignment. The final microorganism detection information after re-alignment is shown in Table 7:
[0091] Table 7
[0092]
[0093] As shown above, through multi-target joint analysis, false positive detection results caused by similarity of single targets and sequencing errors can be solved, and the accuracy of pathogenic microorganism detection can be greatly improved.
[0094] The above is only a preferred embodiment of the present application, and of course cannot limit the scope of protection of the present application. It should be pointed out that any modification, equivalent replacement and improvement made by those skilled in the art within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for multi-target pathogenic microorganism analysis based on third generation targeted sequencing data, characterized in that, The multi-target pathogenic microorganism analysis method is not used for diagnosis of diseases, and the method comprises the following steps: S1, intercepting and downloading reference genomes of pathogenic microorganisms to obtain target region reference sequences, wherein each pathogenic microorganism has at least two target region reference sequences, the reference sequences are corrected in clusters according to similarity, and the reference sequences are associated at the species level to obtain a multi-target associated error correction reference sequence library; S2, quality control of original sequencing data; S3, aligning the sequencing data obtained in step S2 to a pathogenic microorganism identification database to obtain a preliminary microorganism detection list after quality control; S4, extracting all reference sequences corresponding to the species in the microorganism detection list in step S3 from the error correction reference sequence library to construct a pathogenic microorganism detection library with multi-target association, and then aligning the sequencing data obtained in step S2 to the pathogenic microorganism detection library for determination to obtain a final microorganism detection list.
2. The method of claim 1, wherein the method is based on third generation targeted sequencing data. The reference sequences in step S1 are adjusted for direction consistency according to the matching order of primers after clustering.
3. The method of claim 1, wherein the method is based on third generation targeted sequencing data. Step S2 comprises the following steps: S21, filtering out sequencing sequences with a length less than m or a mass value less than n, wherein the value range of m is 100-600 bp, and the value range of n is 8-11; S22, removing sequencing sequences aligned to host genomes; S23, aligning to the error correction reference sequence library, and correcting sequencing sequences aligned to the same reference sequence; S24, comparing the corrected sequencing sequences with each other, and merging similar sequences.
4. The method of claim 3, wherein the method is based on third generation targeted sequencing data. The correction method in step S23 is specifically: first, filtering out sequencing sequences with an error rate higher than a or a coverage rate lower than b, and then correcting the base sites in the sequencing sequences based on the reference sequence; the value range of a is 5%-10%, and the value range of b is 50%-95%.
5. The method of claim 4, wherein the method is based on third generation targeted sequencing data. The error correction method is specifically: counting the base frequency of each site in the sequencing sequence, calculating the proportion of sites with a sequencing depth lower than c in the sequencing sequence, and retaining sequencing sequences with a proportion higher than d; in the retained sequencing sequences, the corresponding base in the reference sequence is used to replace the sites with a sequencing depth lower than c, and the base with the highest frequency is used for the sites with a sequencing depth greater than c; the value range of c is 3-7, and the value range of d is 10%-30%.
6. The method of claim 3, wherein the method is based on a three-generation targeted sequencing data of the pathogenic microorganism. The similarity threshold of similar sequences in step 24 is any value within 95%-99.9%.
7. The method of claim 3, wherein the method is based on a third generation targeted sequencing data. Step S2 further comprises: S25, aligning the merged sequencing sequences to a reagent background microorganism database, and removing the aligned sequencing data.
8. The method of claim 1, wherein the method is based on a three-generation targeted sequencing data of a plurality of pathogenic microorganisms. The determination parameter in step S4 is: retaining species with multiple target simultaneous detection and species with only single target detection but with a sequence number proportion of the single target to the total sequence number being h and above, and the value range of h is 2%-8%.
9. The method for analyzing multi-target pathogenic microorganisms based on third-generation targeted sequencing data according to claim 1, characterized in that, Further comprising the following steps: S5, filtering out groups with a sequence number less than x or groups with a sequence number less than y of the total sequence number in the microorganism detection list obtained in step S4, wherein the value range of x is 5-20, and the value range of y is 0.2%-2%.
10. The method of claim 1, wherein the method is based on a three-generation targeted sequencing data of a plurality of pathogenic microorganisms. The parameters of the quality control in step S3 are: retaining the aligned sequences with the coverage greater than e and the error rate less than f, the value range of e is 85% to 99%, and the value range of f is 0.5% to 3%.
Citation Information
Patent Citations
Microorganism detection method and device based on targeted amplification and sequencing
CN110875082A
Fusion primer for third-generation sequencing library construction, and library construction method, sequencing method and library construction kit therefor
WO2020164015A1