Analysis method for sequencing data of targeted detection microorganisms and application
By removing host information, aligning the reference sequence of pathogenic microorganisms and correcting the chimeric sequence of the sequencing data of the body fluid samples, the problem of insufficient detection sensitivity and specificity in the prior art is solved, and more efficient data analysis and lower false positive rates are achieved.
Patent Information
- Application Number
- CN202311674180.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
Prior art In bodily fluid samples or samples from open respiratory tract sources, it is difficult to effectively distinguish signals from noise, resulting in insufficient detection sensitivity and specificity, and high background interference and false positive rates.
By obtaining the sequencing data of the sample to be examined, the host information is first removed by aligning with the reference sequence of the host, and then the mismatch sequence is filtered with the reference sequence of the pathogenic microorganism, and finally bioinformatics analysis is performed to generate the detection results. The method includes chimeric sequence correction, dynamic mismatch threshold filtering, and probe-specific filtering to improve data quality and reduce false positives.
It improves the utilization and quality of sequencing data, reduces background interference and false positive rates, enhances the sensitivity and specificity of detection, and shortens the running cycle of sequencing data analysis.
Smart Images

Figure BDA0004594269250000181 
Figure BDA0004594269250000191 
Figure BDA0004594269250000192
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and particularly to an analysis method, device and application for sequencing data targeting the detection of microorganisms. Background Art
[0002] High-throughput sequencing technology (also known as Next Generation Sequencing, NGS) has unique advantages in identifying rare or newly discovered microorganisms. With the increasing maturity of nucleic acid sequencing technology, it is increasingly applied to the detection of microorganisms.
[0003] High-throughput sequencing technology can simultaneously detect and analyze a large number of DNA or RNA sequences in a sample, as well as genomic variations and mutations in the sequences. Compared with traditional sequencing technology, it has higher throughput, faster speed and lower cost, and can quickly, comprehensively and accurately detect various different types of microorganisms.
[0004] In practical applications, common high-throughput sequencing technology strategies include unbiased metagenomic sequencing (Metagenomic Next Generation Sequencing, mNGS) and specific targeted sequencing (Targeted Next Generation Sequencing, tNGS). Metagenomic sequencing technology can directly sequence all DNA / RNA in a sample to detect a wider range of pathogenic microorganisms, including known and unknown microorganisms. However, in the detection of human-derived microbial samples, since nucleic acids from the human host occupy most of the sample (for example, the content of host nucleic acids in blood samples is generally above 99%), using metagenomic sequencing will generate a large amount of host sequence information, while the proportion of target microbial sequences is very small and they are discretely distributed on the genome, making it difficult to obtain comprehensive mutation information, and the positive detection rate for body fluid samples such as blood and cerebrospinal fluid is very low.
[0005] Targeted sequencing technology enriches the sequences of target pathogens before sequencing to increase their proportion in the sample, and then through NGS sequencing, the detection rate of microbial genomic sequences and variations and mutations in the sequences is significantly improved, especially the mutation information relied on for identifying drug resistance genes and virulence factors. It can more precisely detect and identify known pathogenic microorganisms, provide faster and more accurate detection results, with higher sensitivity and lower cost, and is of great significance in the study of mutations of specific genes, screening of disease-related genes and personalized medicine.
[0006] The methods commonly used to enrich target nucleic acids in targeted sequencing technologies mainly include multiplex PCR and probe hybridization capture enrichment, etc. Through the design of appropriate PCR primers, multiplex PCR can selectively amplify target nucleic acid sequences. The probe method specifically hybridizes and captures target nucleic acid sequences by using oligonucleotide probes complementary to the target nucleic acid sequences to obtain enriched target sequences. The number of target sequences that multiplex PCR can amplify is limited. While the number of probes that can be designed by the probe method is more flexible, ranging from dozens to thousands, and the diversity of the fragments that can be enriched is also higher. More target probes can be designed (in terms of breadth) on the genome to obtain richer fragments.
[0007] In the targeted sequencing technology of microorganisms, it is necessary to analyze the sequencing data to generate the final detection results. This includes first aligning with the host genome to remove host information, and then comparing the sequences with the host information removed with the reference genome and classifying them according to microbial species.
[0008] In practical applications, for body fluid samples or samples from open respiratory tracts, such as blood, plasma, serum, cerebrospinal fluid, pleural effusion, sputum, bronchoalveolar lavage fluid, nasopharyngeal swabs, oral swabs, etc., the genomic sequences of pathogenic microorganisms are short, with low abundance and a high proportion of host nucleic acids. Existing methods cannot fully identify true positive results (such as indicating pathogenic microorganisms in the subject) from the signal noise, thus weakening the ability to distinguish true positives from false positives caused by noise sources, which poses a challenge to generating reliable results from sequencing data analysis.
[0009] As an example of these challenges, the utilization rate of data needs to be improved. When sequencing sequences are aligned with the reference genome, there will be mismatches, and the reasons for the generation of mismatches are diverse, including errors of DNA polymerase, PCR amplification, library preparation, environmental factors, and sample processing, etc. If fixed sequence length or mismatch number filtering is used, some sequences will be lost.
[0010] Another example is that the detection performance and quality of the alignment are affected by various factors, including chimeric sequences generated by fragmentation with nucleic acid fragmentase. Additionally, during the preparation process of sequencing samples, nucleic acids will be fragmented first, and enzymatic digestion and fragmentation are widely used due to its speed and low cost. However, the enzymatic digestion and fragmentation method will introduce a certain proportion of chimeric sequences (about 5%) formed by the ligation of sequences upstream and downstream of the enzyme cleavage site while fragmenting nucleic acids. When the sequencing data is subjected to bioinformatics analysis and needs to be aligned to the reference genome, the sequencing data of these chimeric sequences will reduce the quality of the alignment, generate more regions with concentrated alignment mismatches and be filtered out, resulting in the loss of available data and further reducing the detection performance of sequencing.
[0011] Another example is the background interference caused by signal noise such as homologous sequences in the host sequence, non-pathogenic or conditionally pathogenic microbial genomes that are the same or similar to the target pathogenic microorganism, and human endogenous microbial sequences.
[0012] When removing host sequences, the human reference sequence used for alignment lacks population diversity information. In actual applications, population diversity will reduce the alignment quality, resulting in some sequences with large variations not being removed completely, and may be aligned to and classified as homologous parasites or fungi, leading to false positives.
[0013] Another example is that when aligning sequencing sequences with a reference genome database, if aligning with the reference genome databases of as many species as possible, relatively specific alignment results can be obtained, but it takes too long and cannot meet the requirements of rapid clinical testing; if the reference genome scope is narrowed down to the target microbial genome database, the alignment time can be greatly shortened, but low-specificity alignment results are likely to be produced, that is, the alignment with the target microbial genome sequence is not unique. For example, it may be aligned to the non-pathogenic or conditionally pathogenic homologous microbial genomes in the database.
[0014] There is also an example that the mismatch information generated by probe hybridization reduces specificity and introduces false positives caused by homologous non-pathogenic species. Since the hybridization process of the probe method often allows more mismatches, the captured sequences will be different from the probe sequences. Moreover, the microbial evolution rate is fast, there are many assembled versions of the reference genome, and there are also unvalidated low-quality sequences or misclassified sequences in the reference genome database, which affects the accuracy of detecting microorganisms by the probe method.
[0015] Various combinations of these challenges affect the performance such as the sensitivity and specificity of detection.
[0016] Therefore, there is an urgent need for a bioinformatics analysis method that can effectively improve data utilization and data quality, reduce background interference or false positives, and thus improve the detection sensitivity and specificity. Summary of the Invention
[0017] To solve at least one of the above problems, the present disclosure provides an analysis method, device and application for pathogenic microorganism sequencing data to improve the sensitivity and specificity of sequencing.
[0018] According to the first aspect of the present disclosure, there is provided an analysis method for pathogenic microorganism sequencing data, including the following steps:
[0019] S1. Obtain the sequencing data of the sample to be tested;
[0020] S2. Align the sequencing data with the reference sequence of the host, and filter out the sequences that are identical to the reference sequence of the host to obtain the first sequencing data;
[0021] S3. Align the first sequencing data with the reference sequence of the pathogenic microorganism, and filter out the sequences with mismatches to obtain the second sequencing data;
[0022] S4. Perform bioinformatics analysis on the second sequencing data to obtain the analysis result of the pathogenic microorganism.
[0023] In some embodiments, in step S1, the sequencing data of the test sample is obtained through the following steps: nucleic acid extraction, library construction, probe capture, and sequencing are performed on the test sample to obtain the sequencing data.
[0024] In some embodiments, before step S2, the analysis method further includes a step of correcting chimeric sequences.
[0025] In some embodiments, the correction step includes the following steps:
[0026] 1) Part of the sequences at one or both ends of the partially read sequences that show soft-clip alignment;
[0027] 2) Align with the reverse complementary sequences of the sequences within 100 - 400 bp upstream and downstream of the alignment position in the reference sequence;
[0028] 3) Filter out the sequences with a matching degree ≥ 80% in step 2) to obtain the sequencing data with chimeric interference removed.
[0029] In some embodiments, in step 2), align with the reverse complementary sequences of the sequences within 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, 260 bp, 270 bp, 280 bp, 290 bp, 300 bp, 310 bp, 330 bp, 330 bp, 340 bp, 350 bp, 360 bp, 370 bp, 380 bp, 390 bp or 400 bp upstream and downstream of the alignment position in the reference sequence.
[0030] In some embodiments, in step 3), filter out the sequences with a matching degree ≥ 80%, ≥ 81%, ≥ 82%, ≥ 83%, ≥ 84%, ≥≥ 85%, ≥ 86%, ≥ 87%, ≥ 88%, ≥ 89%, ≥≥ 90%, ≥ 91%, ≥ 92%, ≥ 93%, ≥ 94%, ≥ 95%, ≥ 96%, ≥ 97%, ≥ 98% or ≥ 99% in step 2) to obtain high-quality sequences with chimeric interference removed.
[0031] In some embodiments, the correction step of the chimeric sequence is to partially read the sequence at one end of the sequence and align it with the opposite strand sequence. Sequences with a matching degree of 90% or greater are judged as chimeric sequences and filtered out.
[0032] In some embodiments, the chimeric sequence is introduced during the fragmentation of the extracted nucleic acid.
[0033] In some embodiments, the fragmented sample nucleic acid is obtained by fragmenting the sample nucleic acid using a nucleic acid fragmenting enzyme or ultrasound.
[0034] In some embodiments, the fragmented sample nucleic acid is obtained by fragmenting the sample nucleic acid using a nucleic acid fragmenting enzyme.
[0035] In some embodiments, the fragmented sample nucleic acid is fragmented using deoxyribonuclease I (DNase I), endonuclease V, or Tn5 transposase digestion.
[0036] In some embodiments, the length of the fragmented nucleic acid is 50 - 300 bp.
[0037] In some embodiments, the fragmented sample nucleic acid has a length of 50bp, 51bp, 52bp, 53bp, 54bp, 55bp, 56bp, 57bp, 58bp, 59bp, 60bp, 61bp, 62bp, 63bp, 64bp, 65bp, 66bp, 67bp, 68bp, 69bp, 70bp, 71bp, 72bp, 73bp, 74bp, 75bp, 76bp, 77bp, 78bp, 79bp, 80bp, 81bp, 82bp, 83bp, 84bp, 85bp, 86bp, 87bp, 88bp, 89bp, 90bp, 91bp, 92bp, 93bp, 94bp, 95bp, 96bp, 97bp, 98bp, 99bp, 100bp, 101bp, 102bp, 103bp, 104bp, 105bp, 106bp, 107bp, 108bp, 109bp, 110bp, 111bp, 112bp, 113bp, 114bp, 115bp, 116bp, 117bp, 118bp, 119bp, 120bp, 121bp, 122bp, 123bp, 124bp, 125bp, 126bp, 127bp, 128bp, 129bp, 130bp, 131bp, 132bp, 133bp, 134bp, 135bp, 136bp, 137bp, 138bp, 139bp, 140bp, 141bp, 142bp, 143bp, 144bp, 145bp, 146bp, 147bp, 148bp, 149bp, 150bp, 151bp, 152bp, 153bp, 154bp, 155bp, 156bp, 157bp, 158bp, 159bp, 160bp, 161bp, 162bp, 163bp, 164bp, 165bp, 166bp, 167bp, 168bp, 169bp, 170bp, 171bp, 172bp, 173bp, 174bp, 175bp, 176bp, 177bp, 178bp, 179bp, 180bp, 181bp, 182bp, 183bp, 184bp, 185bp, 186bp, 187bp, 188bp, 189bp, 190bp, 191bp, 192bp, 193bp, 194bp, 195bp, 196bp, 197bp, 198bp, 199bp, 200bp, 201bp, 202bp, 203bp, 204bp, 205bp, 206bp, 207bp, 208bp, 209bp, 210bp, 211bp, 212bp, 213bp, 214bp, 215bp, 216bp, 217bp, 218bp, 219bp, 220bp,221 bp, 222 bp, 223 bp, 224 bp, 225 bp, 226 bp, 227 bp, 228 bp, 229 bp, 230 bp, 231 bp, 232 bp, 233 bp, 234 bp, 235 bp, 236 bp, 237 bp, 238 bp, 239 bp, 240 bp, 241 bp, 242 bp, 243 bp, 244 bp, 245 bp, 246 bp, 247 bp, 248 bp, 249 bp, 250 bp, 251 bp, 252 bp, 253 bp, 254 bp, 255 bp, 256 bp, 257 bp, 258 bp, 259 bp, 260 bp, 261 bp, 262 bp, 263 bp, 264 bp, 265 bp, 266 bp, 267 bp, 268 bp, 269 bp, 270 bp, 271 bp, 272 bp, 273 bp, 274 bp, 275 bp, 276 bp, 277 bp, 278 bp, 279 bp, 280 bp, 281 bp, 282 bp, 283 bp, 284 bp, 285 bp, 286 bp, 287 bp, 288 bp, 289 bp, 290 bp, 291 bp, 292 bp, 293 bp, 294 bp, 295 bp, 296 bp, 297 bp, 298 bp, 299 bp or 300 bp.
[0038] In some embodiments, the length of the fragmented sample nucleic acid is 30 - 120 bp.
[0039] In some embodiments, the method further comprises the step of filtering the sequencing data through the specificity of the probe:
[0040] By aligning the probe with the reference sequences of two or more species, the specificity of the probe is judged. When the mismatch value of the alignment is below the threshold, the probe aligned to two or more species is marked as a low-specificity probe; and
[0041] Filter out the sequencing data of the nucleic acid sequences captured by the low-specificity probe,
[0042] wherein this step is applied to the first sequencing data and / or the second sequencing data.
[0043] In some embodiments, by aligning the probe with the reference sequences of two or more species, when the mismatch value is below 40%, below 39%, below 38%, below 37%, below 36%, below 35%, below 34%, below 33%, below 32%, below 31%, below 30%, below 29%, below 28%, below 27%, below 26%, below 25%, below 24%, below 23%, below 22%, below 21%, below 20%, below 19%, below 18%, below 17%, below 16%, below 15%, below 14%, below 13%, below 12%, below 11%, below 10%, below 9%, below 8%, below 7%, below 6%, below 5%, below 4%, below 3%, below 2%, or below 1%, the probe aligned to two or more species is labeled as a low-specificity probe.
[0044] In some embodiments, the host includes a human, a monkey, an ape, a cow, a sheep, a camel, a deer, a pig, a horse, a rabbit, a cat, a dog, a chicken, a duck, a pigeon, a goose, a bird, a fish, a shrimp, or a crab.
[0045] In some embodiments, the host is a human.
[0046] In some embodiments, the reference sequence of the host is from the human reference genome or a population database.
[0047] In some embodiments, the human reference genome includes GRCH37, b37, hs37d5, hg19, GRCH38, or CHM13.
[0048] In some embodiments, the population database includes the dbSNP database, the gnomAD database, the ExAC database, the 1000Genomes, or the pan-genome reference map constructed by the Chinese Population Pan-Genome Consortium (CPC).
[0049] In some embodiments, the reference sequence of the host is from the Assembly database, the nt library, the GigaDB database, the EnsemblBacteria database, the IMG database, etc.
[0050] In some embodiments, in step S3, the reference sequence of the pathogenic microorganism is the reference sequence of the pathogenic microorganism after filtering.
[0051] In some embodiments, the filtering of the reference sequence of the pathogenic microorganism includes the following steps: filtering out the sequences derived from the host.
[0052] In some embodiments, filtering out the sequences derived from the host includes filtering out the sequences of microorganisms or viruses integrated in the host genomic sequence.
[0053] In some embodiments, in step S3, filtering out sequences with mismatches includes: filtering sequences using a dynamic mismatch threshold or using a fixed mismatch threshold.
[0054] In some embodiments, filtering sequences using a dynamic mismatch threshold includes, when the sequence length is 30 - 100 bp, using different mismatch thresholds for filtering according to different sequence lengths.
[0055] In some embodiments, when the sequence length is 30 - 40 bp, set the mismatch threshold to 0 bp, and filter out sequences with a mismatch value > 0 bp.
[0056] In some embodiments, when the sequence length is 41 - 50 bp, set the mismatch threshold to 2 bp, and filter out sequences with a mismatch value > 2 bp.
[0057] In some embodiments, when the sequence length is 51 - 79 bp, set the mismatch threshold to 3 bp, and filter out sequences with a mismatch value > 3 bp.
[0058] In some embodiments, when the sequence length is 80 - 100 bp, set the mismatch threshold to 4 bp, and filter out sequences with a mismatch value > 4 bp.
[0059] In some embodiments, filter out sequences with a length < 30 bp.
[0060] In some embodiments, performing bioinformatics analysis on the second sequencing data includes: obtaining the species corresponding to the sequences in the second sequencing data according to the alignment results, and obtaining a predicted species set;
[0061] Selecting species supported by non - repetitive sequences as predicted species; and
[0062] Merging all predicted species to obtain the complete set of predicted species of the sample, and obtaining the analysis result of pathogenic microorganisms.
[0063] In some embodiments, the analysis result of pathogenic microorganisms includes the Latin name of the species, the number of sequences at the species level, the relative abundance at the species level, the Latin name of the genus, the number of sequences at the genus level, the relative abundance at the genus level, the sequence proportion of each species within the genus, the Latin name of the family, the number of sequences at the family level, the relative abundance at the family level, the sequence proportion of each genus within the family, and taxonomic lineage information.
[0064] In some embodiments, select species supported by 2 or more non - repetitive sequences as predicted species.
[0065] In some embodiments, a species supported by two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more non-repetitive sequences is selected as a predicted species.
[0066] In some embodiments, a species supported by five or more non-repetitive sequences is selected as a predicted species.
[0067] According to a second aspect of the present disclosure, there is provided an apparatus for pathogenic microorganism sequencing data, the apparatus comprising:
[0068] A sequencing data acquisition module: configured to acquire sequencing data of a sample to be tested;
[0069] A first alignment module: aligning the sequencing data with a reference sequence of a host, filtering out sequences identical to the reference sequence of the host to obtain first sequencing data;
[0070] A second alignment module: aligning the first sequencing data with a reference sequence of the pathogenic microorganism, filtering out sequences with mismatches to obtain second sequencing data;
[0071] A bioinformatics analysis module: performing bioinformatics analysis on the second sequencing data to obtain an analysis result of the pathogenic microorganism.
[0072] In some embodiments, the apparatus further includes a chimera sequence correction module, which includes correcting chimera sequences introduced during the fragmentation of the sample nucleic acid.
[0073] In some embodiments, the apparatus further includes a reference sequence filtering module, which includes filtering out sequences of host origin.
[0074] In some embodiments, the reference sequence filtering module includes filtering out sequences of microorganisms or viruses integrated in the host genomic sequence.
[0075] In some embodiments, the filtering out of sequences with mismatches is performed using a dynamic mismatch threshold.
[0076] In some embodiments, the dynamic mismatch threshold sets different mismatch thresholds according to sequences of different lengths.
[0077] In some embodiments, the apparatus further includes a module for filtering out low-specificity sequences according to probes, which includes:
[0078] Judging the specificity of the probe: by aligning the probe with reference sequences of two or more species, judging the specificity of the probe, wherein when the mismatch value of the alignment is below the threshold, the probe aligned to two or more species is marked as a low-specificity probe; and
[0079] Filter out the sequencing data captured by the low-specificity probes, where the sequencing data is the first sequencing data and / or the second sequencing data.
[0080] In some embodiments, by aligning the probes with the reference sequences of two or more species, when the mismatch value is below 40%, below 39%, below 38%, below 37%, below 36%, below 35%, below 34%, below 33%, below 32%, below 31%, below 30%, below 29%, below 28%, below 27%, below 26%, below 25%, below 24%, below 23%, below 22%, below 21%, below 20%, below 19%, below 18%, below 17%, below 16%, below 15%, below 14%, below 13%, below 12%, below 11%, below 10%, below 9%, below 8%, below 7%, below 6%, below 5%, below 4%, below 3%, below 2%, or below 1%, the probes aligned to two or more species are marked as low-specificity probes.
[0081] According to the third aspect of the present disclosure, there is provided an analysis system for targeted detection of sequencing data of microorganisms, including a computer processor and a memory, the memory storing computer program instructions, and when the computer program instructions are executed by the computer processor, the processor is caused to execute the steps of the analysis method described in the first aspect.
[0082] According to the fourth aspect of the present disclosure, there is provided the use of the method described in the first aspect, the device described in the second aspect, or the analysis system described in the third aspect in detecting pathogenic microorganisms in a sample.
[0083] In some embodiments, the sample is selected from tissue or body fluid samples.
[0084] In some embodiments, the sample is selected from tissue, cells, blood, plasma, serum, bronchoalveolar lavage fluid, sputum, pus, nasopharyngeal swab, oral swab, cerebrospinal fluid, pleural effusion, ascites, or their processed products.
[0085] In some embodiments, the microorganism is a pathogenic microorganism.
[0086] In some embodiments, the microorganism includes eukaryotic microorganisms and prokaryotic microorganisms.
[0087] In some embodiments, the eukaryotic microorganism includes any one of yeast, fungi, or algae.
[0088] In some embodiments, the prokaryotic microorganism includes one or more of the following: Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli, Moraxella catarrhalis, Enterobacter cloacae, Serratia marcescens, Burkholderia cepacia, Klebsiella oxytoca, Enterobacter aerogenes, Proteus mirabilis, Streptococcus pyogenes, Haemophilus parainfluenzae, Citrobacter freundii, Cryptococcus neoformans, Candida auris, Aspergillus fumigatus, Candida albicans, influenza virus, Human Respiratory syncytial virus, Rhinovirus, Human parainfluenza virus, Human adenovirus, coronavirus, Human metapneumovirus, Human bocavirus, Mycobacterium tuberculosis complex, Mycobacterium avium, Mycobacterium intracellulare, Legionella pneumophila, ListeriaListeria monocytogenes, Mycoplasma pneumoniae, Chlamydia pneumoniae, Chlamydia psittaci, Streptococcus constellatus, Coxiella burnetii, Blastomyces dermatitidis, Coccidioides immitis, Cytomegalovirus, adenovirus, Lymphocytic choriomeningitis virus, Mycobacterium intracellulare, Mycobacterium avium, Mycobacteroides abscessus and its three subspecies (Mycobacteroides abscessus subsp. Abscessus, Mycobacteroides abscessus subsp. Bolletii, Mycobacteroides abscessus subsp. massiliense), Mycobacterium canettii, Mycobacterium interjectum, Mycobacterium kansasii, Mycobacterium marinum, Mycobacteroides chelonae, Mycolicibacterium fortuitum, Mycobacterium gordonae, Mycobacterium malmoense, Mycobacterium scrofulaceum, Mycobacterium simiae, Mycobacterium xenopi, Mycobacterium ulcerans, Serratia marcescens, Salmonella, Streptococcus pneumoniae, Enterococcus faeciumEnterococcus faecium and Enterococcus faecalis.
[0089] In some embodiments, the influenza virus includes influenza A virus (H1N1), influenza A virus (H3N2), influenza A virus (H5N1), influenza A virus (H7N9), influenza A virus (H9N2), and influenza B virus.
[0090] In some embodiments, the coronavirus includes human coronavirus HKU1, human coronavirus NL63, human coronavirus OC43, novel coronavirus, and SARS-CoV-2.
[0091] According to the fifth aspect of the present disclosure, there is provided a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the analysis method described in the first aspect are implemented.
[0092] The present disclosure provides an analysis method and application for targeted detection of sequencing data of microorganisms. Using the method provided by the present disclosure, it is possible to reduce the interference of the background, such as human host sequences, human endogenous microorganisms, or homologous microorganisms that are non-pathogenic or conditionally pathogenic in the sample; reduce the impact of chimeric sequences generated by enzymatic digestion on the alignment quality, improve the effective utilization rate and quality of the data, especially short sequence fragments, reduce false positives introduced by probe hybridization mismatches, improve the sensitivity and specificity of detection, and at the same time shorten the operation cycle of sequencing data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 Exemplarily shows a flowchart of an analysis method for pathogenic microorganism sequencing data.
[0094] Figure 2 Exemplarily shows a schematic flow diagram for obtaining sequencing data of a sample to be tested.
[0095] Figure 3 It is a schematic diagram for comparing the de-hosting effects of 20 blood samples.
[0096] Figure 4 It is the proportion of chimeric sequence-containing sequences that can be corrected in blood and bronchoalveolar lavage fluid. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0097] In view of the data characteristics of targeted sequencing, the present disclosure provides an analysis method for pathogenic microorganism sequencing data, which ensures the specificity of the sequencing result while improving the detection sensitivity. Figure 1 Exemplarily shows a flowchart of an analysis method for pathogenic microorganism sequencing data. As Figure 1 shown, the method includes, but is not limited to, the following steps.
[0098] Step S1: Obtaining the sequencing sequence of the sample to be tested;
[0099] Step S1-1: Chimeric sequence correction.
[0100] Correcting chimeric sequences introduced by enzyme interruption in the sequencing sequence. In a specific embodiment, the correction step of the chimeric sequence is to read a partial sequence with one or both ends of the sequence presented as a soft-clip alignment, and align it with the reverse complementary sequence of the sequence within 300bp upstream and downstream of the reference sequence. Sequences that meet a match of 90% or more are judged as chimeric sequences and are filtered out from the sequencing sequence, thereby obtaining a high-quality sequence with chimeric interference removed.
[0101] Step S2: Removal of host sequences.
[0102] Remove the host sequence in the sequencing sequence. The sequencing sequence is compared with the reference sequence of the host (e.g., human), and the sequence with the consistent comparison is removed as the host sequence. The reference sequence of the host can use the sequence library of some hosts known in the art. The representativeness of the host's reference sequence affects the removal efficiency of the host sequence.
[0103] In a specific embodiment, the host is human, and the human reference sequences commonly used are hg19 and hg38 versions, the Yanhuang No. 1 genome reference YH-1, or the full-length genome CHM13-T2T (hereinafter referred to as CHM13). The common point of these genomes is that they are obtained based on sequencing of individuals, a small number of people or single cell lines, lacking population diversity information. In practical applications, population diversity will reduce the quality of comparison, resulting in some sequences with large variations not being removed cleanly, and may be compared to parasites or fungi classified as homologous, resulting in the generation of false positives.
[0104] In a specific embodiment of the present disclosure, the pan-genome reference map constructed by the Chinese Pangenome Consortium (CPC) (https: / / pog.fudan.edu.cn / cpc / files / CPC.Phase1.CHM13v2-full / CPC.Phase1.CHM13v2-full.gfa.gz) is used as the human reference sequence. The pan-genome reference map deeply sequenced 58 samples representing 36 ethnic groups in China using the third-generation genome sequencing technology. Combining with the haplotype genome assembly method, 116 high-quality haplotype genomes were obtained, and the reference pan-genome of the Chinese population was constructed in the form of a graph genome. The CPC assembly largely matches or exceeds the continuity and base-level accuracy of the current reference human genome sequence (GRCh38). It can provide a comprehensive pan-genome reference for the Chinese population.
[0105] Step S2-1: Preprocessing of the reference genome before alignment.
[0106] Preprocessing of the reference genome database. There are human endogenous virus sequences in the reference genome database, that is, sequences formed by the integration of exogenous microorganisms into the human genome. By filtering out the virus nucleic acid sequences identified and classified as human in the database, the human endogenous virus sequences existing in the database can be excluded before alignment, which can exclude the interference of the human endogenous microbial genome on the detection of pathogenic microorganisms and improve the specific verification effect on pathogenic microorganisms.
[0107] As an example, there are human endogenous virus sequences in the nt library of NCBI. These sequences are classified and marked as human (taxonomy ID: 9606), which will affect the specific judgment of related viruses. By filtering the virus sequences with the classification marked as human (taxonomy ID: 9606) and keywords such as "integration site" and "endogenous retrovirus" in the nucleic acid name, these sequences can be filtered out, improving the specific verification effect on viruses such as HIV and HBV.
[0108] Step S3: Align the filtered or processed sequencing data with the reference genome.
[0109] Step S3-1: Filter the sequencing data using a dynamic threshold.
[0110] When the sequenced sequences are aligned with the reference genome, there will be mismatches. The reasons for the mismatches are diverse, including errors of DNA polymerase, PCR amplification, library preparation, environmental factors, sample handling, and so on. If fixed sequence lengths or numbers of mismatches are used for filtering, 10-30% of the sequences will be lost. Therefore, a dynamic mismatch filtering threshold is adopted to perform dynamic mismatch filtering on sequences ≥30bp, and only extremely short sequences (e.g., <30bp) are completely filtered. Thus, sequences ≥30bp can be reasonably utilized for subsequent analysis.
[0111] The sequences in the sequencing library are selected from the sequences generated by sequencing, the sequences after removing the host sequences in step S2, or the sequences corrected in step S1-1.
[0112] Table 1 exemplarily shows the filtering thresholds for sequence lengths and corresponding numbers of mismatches.
[0113] Table 1. Rules for dynamic mismatches
[0114] Actual length of the preprocessed sequence (bp) Number of allowed mismatches (bp) 30-40 0 41-50 2 51-79 3 80-100 4
[0115] Step S4: Analysis
[0116] According to the alignment results in step S3, the corresponding species are obtained to obtain a predicted species set;
[0117] Select the species supported by non-repetitive sequences as the predicted species; and
[0118] Merge the predicted species identified by all the sequenced sequences to obtain the complete set of predicted species of the sample, and further obtain the identification result of the microorganism.
[0119] Step S4-1: Prior probe filtering.
[0120] Perform prior filtering. Simulate the sequences of the used probes, and then align them to the microorganism genome database by setting different mismatch parameters. Count the types and numbers of species that each probe can align to under mismatch conditions. Probes with a mismatch value lower than the threshold but still able to align to two or more different genera within the same genus are marked as prior low-specificity probes. In the alignment analysis of the actual sequenced sequences, prior filtering of the detected sequences is combined with the probe prior, which can filter out false positives introduced by homologous non-pathogenic species and improve the specificity of the detection results.
[0121] Moreover, aligning the sequenced sequences of each sample to the nt library is time-consuming, resulting in an overly long time cycle for pathogen detection. The approach of probe prior is to implement this time-consuming step in advance and create a prior database to guide each actual sequencing classification, solving the problem of time consumption in directly aligning to the genome database.
[0122] As an example, the probes used in targeted sequencing are subjected to sequence simulation (e.g., simulating the sequencing type of SE100), and then by setting different mismatch parameters (e.g., 20%, 10%, and 5%), aligning to the nt library, and counting the species types and numbers that each probe can align to under mismatch conditions. Probes that can still align to multiple species within the same genus or species of different genera under conditions of low mismatch (such as 5%) are marked as low-specificity probe priors. In the alignment analysis of actual sequencing sequences, combining the probe priors to perform prior filtering on the detected sequences can filter out false positives introduced by homologous non-pathogenic species. In this example, using prior filtering can shorten the alignment time from 5 hours before filtering to about 30 minutes.
[0123] As an example of step S1: obtaining the sequencing sequence of a test sample, the process is as Figure 2 shown, including the following steps:
[0124] In step 110, a nucleic acid sample (DNA or RNA) is extracted from a subject. In the present disclosure, unless otherwise indicated, DNA and RNA can be used interchangeably. That is, the following embodiments for using error source information in variant identification and quality control can be applied to nucleic acid sequences of DNA and RNA types. However, for the purposes of clarity and explanation, the examples described herein may focus on DNA. The sample can be any subset of the human genome, including the entire genome. The sample may be extracted from a subject known or suspected of having cancer. The sample can include blood, plasma, serum, urine, feces, saliva, other types of body fluids, or any combination thereof. In some embodiments, the method for extracting a blood sample (e.g., syringe or finger prick) may be less invasive than the procedure for obtaining a tissue biopsy, which may require surgery. The extracted sample can include cfDNA and / or ctDNA. For healthy individuals, the human body may naturally clear cfDNA and other cell debris. CtDNA in the extracted sample may be present at a detectable level for diagnosis. In some embodiments, the extracted nucleic acid sample is fragmented, optionally, as is well known in the art, fragmentation is performed by a nucleic acid fragmentase (DNA fragmentase or RNA fragmentase) or sonication.
[0125] In a specific embodiment, the nucleic acid sample extracted in step 110 is fragmented. In a specific embodiment, fragmentation is performed by a DNA fragmentase or an RNA fragmentase.
[0126] In a specific embodiment, the length of the fragmented nucleic acid is 150 - 300bp.
[0127] In a specific embodiment, the fragmented nucleic acid has a length of 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, 260 bp, 270 bp, 280 bp, 290 bp, or 300 bp.
[0128] In step 120, a sequencing library is prepared. During library preparation, sequencing adapters are added to nucleic acid molecules (e.g., DNA molecules), for example, by adapter ligation (using T4 or T7 DNA ligase) or other known methods in the art. After adding the adapters, the adapter-nucleic acid construct is amplified, for example, using polymerase chain reaction (PCR). Optionally, as is well known in the art, the sequencing adapters may further include universal primers, sample-specific barcodes (for multiplexing), and / or one or more sequencing oligonucleotides for subsequent cluster generation and / or sequencing (e.g., for sequencing by synthesis (SBS)).
[0129] In a specific embodiment, targeted DNA sequences are enriched from the prepared sequencing library. During enrichment, hybrid probes (also referred to herein as "probes") are used to target and capture nucleic acid fragments that provide information about the presence or absence of microorganisms, the variant or non-variant state of microorganisms, or the classification of microorganisms. For a given workflow, the probes can be designed to anneal (or hybridize) to the targeted (complementary) strand of DNA or RNA. The targeted strand can be the "positive" strand (e.g., the strand transcribed into mRNA and then translated into protein) or the complementary "negative" strand. The length of the probes can range from 10 to 1000 base pairs, such as 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, 260 bp, 270 bp, 280 bp, 290 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1000 bp.
[0130] In one embodiment, probes are designed based on genetic markers to analyze specific mutations or target regions of a genome (e.g., of a human or another organism) suspected of corresponding to certain microorganisms. Moreover, the probes can cover overlapping portions of the target region. As will be readily understood by those skilled in the art, any known means in the art can be used for target enrichment. For example, in one embodiment, the probes can be biotinylated and streptavidin-coated magnetic beads for enriching the target nucleic acids captured by the probes. See, e.g., Duncavage et al., JMol Diagn. 13(3):325-333 (2011); and Newman et al., NatMed. 20(5):548-554 (2014). By using target genetic markers instead of sequencing the entire genome ("whole-genome sequencing") or all expressed genes of the genome ("whole-exome sequencing" or "whole-transcriptome sequencing"), method 100 can be used to increase the sequencing depth of the target region, where depth refers to the number of times a given target sequence in the sample is sequenced. Increasing the sequencing depth allows for the detection of rare sequence variants in the sample and / or increases the throughput of the sequencing process. After the hybridization step, the hybridized nucleic acid fragments are captured and can also be amplified using PCR.
[0131] In step 130, sequence reads are generated from the enriched nucleic acid molecules (e.g., DNA molecules). Sequencing sequences or sequence reads can be obtained from the enriched nucleic acid molecules by means known in the art. For example, method 100 can include next-generation sequencing (NGS) technologies, including synthesis technologies pyrosequencing (454 LIFE SCIENCES), ion semiconductor technology (Ion Torrent sequencing), single molecule real-time sequencing ligation sequencing (SOLiD sequencing), nanopore sequencing (OXFORD NANOPORE TECHNOLOGIES), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using synthesis sequencing with reversible dye terminators.
[0132] The method described in the present disclosure improves the host removal rate by using a pan-genome reference map, improves the data quality and increases the data utilization rate by using chimeric sequence correction and dynamic threshold filtering for aligning and classifying sequences, and reduces the non-pathogenic false positive rate by database preprocessing and the construction of probe priors.
[0133] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to constitute any limitation to the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessarily confusing the concepts of the present disclosure. Such structures and technologies have also been described in many publications.
[0134] Definition
[0135] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly used in the field to which the present invention belongs. For the purpose of interpreting this specification, the following definitions will be applied, and where appropriate, terms used in the singular form will also include the plural form, and vice versa.
[0136] Unless the context clearly indicates otherwise, the expressions "a" and "an" used herein include plural referents.
[0137] The expression "about" used herein is as understood by those of ordinary skill in the art and varies within a certain range according to the context in which it is used. If those of ordinary skill in the art do not understand the use of the term according to the context in which it is used, "about" will mean a specific value plus or minus 10%.
[0138] In the present disclosure, the term "reference genome" refers to any particular known, sequenced or characterized genome of any organism or virus that can be used to reference an identified sequence from a subject, whether partial or complete. Exemplary reference genomes for human subjects as well as many other organisms are provided in online genome browsers hosted by NCBI or UCSC. "Genome" refers to the complete genetic information of an organism or virus expressed as a nucleic acid sequence. As used herein, a reference sequence or reference genome is typically an assembled or partially assembled genomic sequence from one or more species.
[0139] In the present disclosure, the term "subject" refers to any living or non-living organism, including but not limited to humans (e.g., male, female, fetus, pregnant female, child, etc.), non-human animals, plants, bacteria, fungi, or protists. Any human or non-human animal can be a subject of study, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, cattle, horses, goats (e.g., sheep, goats), pigs (e.g., swine), camels (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), bears (e.g., bears), poultry, dogs, cats, rats, mice, fish, dolphins, whales, and sharks. In some embodiments, the subject is a male or female at any stage (e.g., male, female, or child). The subject from whom a sample is obtained or who is treated by any method or composition described herein can be of any age and can be an adult, infant, or child. In some embodiments, for example, the subject is a patient who is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99 years old or a patient within that range (e.g., between about 2 and 20 years old, between about 20 and 40 years old, or between about 40 and 90 years old). Particular classes of subjects, such as patients who can benefit from the methods of the present disclosure, are subjects, for example, subjects over 40 years old, subjects over 45 years old, subjects over 50 years old, subjects over 50 years old, 50 years old, subjects over 55 years old, or subjects over 60 years old.
[0140] In the present disclosure, the term "probe" refers to a primer labeled with a capture label or a detection label to detect a primer product. The probe sequence is used to hybridize with the sequence generated by the primer sequence and generally hybridizes with a sequence that does not include the primer sequence. Similar to the primer sequence, the probe sequence is also labeled with a capture label or a detection label. It should be noted that when the primer is labeled with a capture label, the probe is labeled with a detection label, and vice versa. The suitable length of the probe depends on the intended use of the probe and is generally in the range of 80 to 200 nucleotides (nt). Preferably, the suitable length of the primer includes 100 to 150 nucleotides. In the present invention, the probe can be one of TaqMan probes, MGB probes, and hybridization probes.
[0141] In the present disclosure, the term "probe set" generally refers to a collection of more than one probe that localizes and / or quantifies a target nucleic acid by recognizing and binding to a target sequence (by means of hybridization pairing). Each probe in the probe combination is usually an oligonucleotide, such as a single-stranded DNA molecule or RNA.
[0142] In the present disclosure, the term "true positive" refers to an indication that a pathogenic microorganism truly exists in a sample. A true positive is not caused by endogenous microorganisms, non-pathogenic microorganisms, or opportunistic pathogenic microorganisms in healthy individuals, nor is it a process error introduced during the preparation of the nucleic acid sample for measurement.
[0143] In the present disclosure, the term "false positive" refers to a test result that is incorrectly determined to be a true positive. Generally, false positives are more likely to occur when processing sequence reads associated with a larger average noise ratio or greater uncertainty in the noise ratio.
[0144] In the present disclosure, the term "cfNA" or "cell-free nucleic acid" refers to nucleic acid molecules that can be found outside cells, such as in body fluids such as blood, sweat, urine, or saliva. Cell-free nucleic acids can be used interchangeably with circulating nucleic acids.
[0145] In the present disclosure, the terms "cell-free nucleic acid", "cell-free DNA", or "cfDNA" refer to deoxyribonucleic acid fragments that circulate in body fluids such as blood, sweat, urine, or saliva.
[0146] In the present disclosure, the term "NT library (Nucleotide Sequence Database)" refers to a nucleic acid sequence database. The NT database is the nucleic acid sequence database of NCBI. The NT library belongs to a non-redundant nucleic acid sequence database, and the data is sourced from GenBank, EMBL, and DDBJ. The NT library stores karyotype DNA sequence information of known species and all human DNA information. By comparing with the NT library, it is possible to know which species our data comes from (the species to which it is aligned is that species). In some embodiments, the address of the NT library is: https: / / ftp.ncbi.nlm.nih.gov / blast / db / .
[0147] In the present disclosure, "alignment" refers to the process of comparing a read or tag with a reference sequence and thereby determining whether the reference sequence contains the read sequence. If the reference sequence contains the read, the read can be mapped to the reference sequence, or in certain embodiments, to a specific position in the reference sequence.
[0148] In the present disclosure, the term "hybridization" or "specific hybridization" refers to a molecule binding, duplexing, or hybridizing only with a specific polynucleotide sequence under stringent conditions, which is carried out when the sequence is present in a complex mixture (e.g., total cellular) DNA or RNA.
[0149] In the present disclosure, the term "sequencing" refers to methods for determining the sequence (e.g., the identity and order of monomeric units) of a biomolecule, such as a nucleic acid, such as DNA or RNA. Exemplary sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy chain termination sequencing, whole genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperature PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short read sequencing, single molecule sequencing, synthesis sequencing, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa genome analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, DNA nanoball sequencing (DNBSEQ), combinatorial probe anchor synthesis sequencing (cPAS), and combinations thereof. In some embodiments, sequencing can be performed using a genetic analyzer, such as a genetic analyzer commercially available from Illumina, Inc., Pacific Biosciences, Inc., Applied Biosystems / Thermo Fisher Scientific, or BGI Genomics Co., Ltd. For example, BGI DNBseq sequencing platforms such as BGISEQ-500, BGISEQ-50, MGISEQ-2000, MGISEQ-200, DNBSEQ-T7, DNBSEQ-G99, DNBSEQ-T20X2, or Illumina's HiSeq2000, HiSeq2500, HiSeq4000, HiSeqX10, NovaSeq6000, etc.
[0150] As used herein, the term "biological sample" means a sample of biological tissue or body fluid that contains nucleic acid or polypeptide. Such samples are typically from humans, but include tissues isolated from non-human primates or rodents (e.g., mice and rats). Biological samples can also include tissue secretions such as biopsy and autopsy samples, frozen sections obtained for histological purposes, cerebrospinal fluid, blood, plasma, serum, sputum, feces, tears, mucus, hair, skin, etc. Biological samples also include explants and / or primary and / or transformed cell cultures derived from patient tissue. The term "biological sample" also refers to a single cell or cell population or a quantity of tissue or body fluid from an animal. Most commonly, a biological sample has been removed from an animal, but the term "biological sample" can also refer to cells or tissue analyzed in vivo, i.e., not removed from the animal. Typically, a "biological sample" will contain cells from an animal, but the term can also refer to acellular biological material that can be used to measure the expression level of a polynucleotide or polypeptide, such as the acellular components of blood, serum, saliva, cerebrospinal fluid, or urine. Many types of biological samples can be used in the present invention, including but not limited to tissue biopsies or blood samples.
[0151] In the present disclosure, the term "chimeric sequence" refers to the fact that the enzymatic fragmentation method will introduce a certain proportion of chimeric sequences formed by the ligation of the upstream and downstream sequences of the restriction site while fragmenting the nucleic acid.
[0152] In the present disclosure, the term "soft clipping" means that when a certain segment of the genome is deleted or the transcriptome is spliced, during the sequencing process, when reads or chimeric sequences spanning the deletion site and the splicing site are aligned to the reference genome, one read is cut into two segments and matched to different regions, and such reads are called soft-clipped reads.
[0153] In the present disclosure, the term "targeted sequencing" refers to a technique that uses biotin-labeled DNA or RNA probes to capture target fragments in a DNA sample and perform sequencing. The probes can be labeled with biotin. Each nucleotide in the probes of the present invention can be chemically synthesized using, for example, a general DNA synthesizer (e.g., Model 394 manufactured by Applied Biosystems). Any other method well known in the art can also be used to synthesize oligonucleotides, such as probes.
[0154] In the present disclosure, the term "computer-readable medium" (e.g., data preservation, data storage, etc.) or "computer-readable storage medium" refers to any medium that participates in providing instructions to a processor for execution. Such a medium can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Examples of non-volatile media can include but are not limited to optical discs, solid-state drives, magnetic disks, such as storage devices. Examples of volatile media can include but are not limited to dynamic memories, such as memories.
[0155] Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, or any other magnetic media, CD-ROMs, any other optical media, punched cards, paper tapes, any other physical media with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, any other storage chip or cartridge, or any other tangible medium from which a computer can read.
[0156] In addition to computer-readable media, data can be provided as signals on transmission media included in a communication device or system to provide one or more instruction sequences to a processor of a computer system for execution. For example, a communication device can include a transceiver having signals indicating instructions and data. The instructions and data are configured to cause one or more processors to implement the functions outlined in the present disclosure. Representative examples of data communication transmission connections can include, for example, telephone modem connections, wide area networks (WANs), local area networks (LANs), infrared data connections, NFC connections, etc.
[0157] Embodiments and accompanying drawings are provided below to assist in understanding the present invention. However, it should be understood that these embodiments and drawings are only for illustrative purposes of the present invention and do not constitute any limitation. The actual scope of protection of the present invention is set forth in the claims. It should be understood that any modifications and changes can be made without departing from the spirit of the present invention.
[0158] Embodiment 1
[0159] Step 1: Obtain the sequencing sequences captured and enriched using a probe for the sample to be tested;
[0160] Step 1-1: Chimeric sequence correction.
[0161] Use the fade software to perform chimeric sequence correction on the sequencing sequences, filter out the reverse complementary sequences of the sequences within 300 bp upstream and downstream of the alignment position in the reference sequence for alignment, and sequences with a matching degree of more than 90%.
[0162] Step 2: Remove host sequences.
[0163] The pan-genome reference map constructed with reference to CPC uses the Kraken 2 software to filter the data obtained in Step 1-1 and filter out the host sequences.
[0164] Step 2-1: Preprocessing of the reference genome before alignment.
[0165] Under the linux system, use the "grep" command to filter the viral sequences in the nt library of NCBI with the taxonomic label of human (taxonomic ID: 9606) and the keywords "integration site", "endogenous retrovirus", and "HBV" in the nucleic acid name.
[0166] Step 3: Use the bwa-mem2 software to align the sequencing sequences after removing the host sequences obtained in Step 2 with the preprocessed reference genome obtained in Step 2-1.
[0167] Step 3-1: During the alignment process, use the samtools software to filter the sequencing sequences according to the dynamic mismatch rules shown in Table 1.
[0168] Step 4: Analysis
[0169] Use the bwa-mem2 software for the following analysis: Select species with more than 5 non-repetitive sequence supports as predicted species; and merge the predicted species identified by all sequencing sequences to obtain the complete set of predicted species of the sample.
[0170] Step 4-1: Prior probe filtering
[0171] Use the wgsim software to simulate single-end 100 (SE100) sequencing sequences of the probes used in targeted sequencing in Step 1, and then, by setting different mismatch parameters (20%, 10%, and 5%), use the blastn software to align to the nt library, and count the species types and numbers that each probe can align to under the mismatch conditions. Probes that can still align to multiple species within the same genus or species of different genera under the condition of low mismatch (5%) are marked as prior low-specificity probes. In the alignment analysis of the actual sequencing sequences, combine the probe prior to perform prior filtering on the detection results of Step 4 to remove false positive results.
[0172] Example 2
[0173] Collect 20 clinical blood samples, use the probes of tNGS-Max of Geneplus to capture and enrich the nucleic acids in the samples to obtain sequencing sequences (tNGS), and perform a de-host comparison test on the sequencing sequences: Use the classification software Kraken2 to construct the host reference libraries of CHM13 and CPC, and the other steps are the same as in Example 1. The results are as Figure 3As shown, CPC as a reference library can improve the identification and removal of host sequences by 1-3% compared with CHM13 as a reference, and reduce the interference of some host sequences on microbial classification.
[0174] Example 3
[0175] Blood cfDNA samples were collected from 153 patients with unexplained fever and 30 healthy people, as well as bronchoalveolar lavage fluid samples from 141 patients with respiratory tract infection. The nucleic acid in the samples was captured and enriched using Genentech's tNGS-Max probe to obtain sequencing sequences. The chimera sequences in the sequencing sequences were corrected, and the other steps were the same as in Example 1. The correctable ratio is as follows: Figure 4 As shown, the median values reached 10% and 17% respectively. Since cfDNA fragments are shorter, the binding efficiency of the fragmentation enzyme is lower than that of the alveolar lavage fluid with a full-length genome, so the proportion of fragments that can be interrupted is relatively small, and the resulting chimera effect is also lower.
[0176] Example 4
[0177] Bronchoalveolar lavage fluid samples from 10 patients with respiratory tract infection were collected, and the pathogenic microorganisms shown in Table 2 were detected by clinical culture method. The probes in tNGS-Max of Gene Plus were used to capture and enrich the nucleic acids in the 10 samples to obtain sequencing sequences. The fixed mismatch threshold (length ≤ 100, mismatch ≤ 5) and the dynamic mismatch threshold shown in Table 1 were used to perform probe method targeted sequencing to capture the sequencing sequence data processing, and the number of detected sequences was counted respectively. The other steps were the same as in Example 1. The results are shown in Table 2.
[0178] Table 2. Number of sequences detected in 10 clinical samples using fixed sequence length and mismatch threshold and dynamic threshold
[0179]
[0180] From Table 2, we can see that the use of dynamic mismatch thresholds can increase pathogen sequence detection by 10-30% compared to fixed thresholds.
[0181] Example 5
[0182] In the nt database (June 2019 version), for sequences classified as human (taxonomy ID: 9606), filtering was performed using the keywords "integration site", "endogenous retrovirus", and "HBV". Other steps were the same as in Example 1. A total of 1370 endogenous sequences, 7474 integration site sequences, 32 HBV sequences, and 6444 other viral sequences were filtered out. 100 non-redundant sequences were selected from the tNGS reference sequence alignment of the HIV-1 positive sample to the classified sequences and aligned with the nt database before filtering. The alignment steps were as follows: The reference genome of HIV from the NCBI Refseq database was simulated into single-end SE50 sequencing sequences with a step size of 1 bp and a window of 5 bp. 100 sequences were randomly selected as the sequences to be analyzed, and then they were aligned to the nt database of unfiltered human endogenous sequences and filtered human endogenous sequences using the blastn software respectively, and the number of sequences that could be specifically classified as HIV was counted. In the alignment results, there was a human sequence Homosapiens209-6HIV-1integration site genomic sequence (9606:10); after alignment with the filtered nt database, there was no longer interference from sequences classified as human in the alignment results. See Table 3.
[0183] Table 3 HIV-1 Specific Verification Results
[0184]
[0185] Example 6
[0186] In this example, false positives were removed through the probe prior step.
[0187] First, probe prior was performed. The probes used for targeted sequencing (tNGS-Max of Geneplus) were sequence-simulated (SE100), and then aligned to the nt database by setting different mismatch parameters (20%, 10%, and 5%). The species types and numbers that each probe could align to under the mismatch conditions were counted. For probes that could still align to two or more different genera within the same genus under relatively low mismatch conditions (such as 5%), such as the probes for the pathogenic bacterium Bacillus anthracis and the common Bacillus cereus in the laboratory and water environment, they were marked as low-specificity probe prior.
[0188] Collect plasma samples from 2 healthy individuals, use the probes in Geneplus's tNGS-Max to capture and enrich nucleic acids in the probe pair to obtain sequencing sequences, and do not perform the prior filtration in step S4-1 during the analysis of the sequencing sequences. Other steps are the same as in Example 1, and Bacillus anthracis is detected as positive; on the other hand, during the alignment analysis of the sequencing sequences, combine the probe prior to perform the prior filtration in step S4-1 on the detected sequences, and compare with the results without prior filtration.
[0189] The experimental results show that before prior filtration, the number of detected sequences of Sample-1 and Sample-2 are 66 and 5 respectively, and the number of probes that play the capture function are 5 and 3 respectively. After prior filtration, the number of species-specific probes for the pathogenic bacterium Bacillus anthracis (prior, <5% mismatch) are both 0, indicating that the positive results of Bacillus anthracis in the two samples are false positive results.
[0190] The homologous part between Bacillus cereus and the pathogenic bacterium Bacillus anthracis is likely to cause false positives of the latter. The experimental results of this example illustrate that the method of probe prior can effectively remove false positive results and improve the specificity of detection. See Table 4.
[0191] Table 4. Detection results of 2 samples
[0192]
[0193] The technical solution of the present invention is not limited to the limitations of the above specific embodiments. Any technical deformation made according to the technical solution of the present invention falls within the protection scope of the present invention.
Claims
1. A method for analyzing sequencing data of pathogenic microorganisms, comprising the following steps: S1. Obtain the sequencing data of the sample to be tested; S2. Align the sequencing data with the reference sequence of the host, and filter out the sequences that are identical to the reference sequence of the host to obtain the first sequencing data; S3. Align the first sequencing data with the reference sequence of the pathogenic microorganism, and filter out the sequences with mismatches to obtain the second sequencing data; S4. Perform bioinformatics analysis on the second sequencing data to obtain the analysis result of the pathogenic microorganism.
2. The analysis method according to claim 1, wherein, in step S1, the sequencing data of the sample to be tested is obtained through the following steps: extracting nucleic acid from the sample to be tested, constructing a library, probe capture and sequencing to obtain the sequencing data.
3. The analysis method according to claim 1, wherein, before step S2, the analysis method further includes a step of correcting chimeric sequences; wherein, the chimeric sequence is introduced during the fragmentation of the extracted nucleic acid of the sample to be tested; preferably, the length of the fragmented nucleic acid is 50 - 300bp.
4. The analysis method according to claim 2, wherein, the method further includes a step of filtering sequencing data through the specificity of the probe: by aligning the probe with the reference sequences of two or more species to determine the specificity of the probe, wherein when the mismatch value in the alignment is below the threshold, the probe that aligns to two or more species is marked as a low-specificity probe; and filter out the sequencing data of the nucleic acid sequences captured by the low-specificity probe, wherein, this step is applied to the first sequencing data and / or the second sequencing data.
5. The analysis method according to claim 1, wherein, the host includes human, monkey, ape, cattle, sheep, camel, deer, pig, horse, rabbit, cat, dog, chicken, duck, pigeon, goose, bird, fish, shrimp or crab; preferably, the host is human; preferably, the reference sequence of the host comes from the human reference genome or a population database, and more preferably includes those from GRCH37, b37, hs37d5, hg19, GRCH38 or the Chinese population pan-genome reference map CPC.
6. The analysis method according to claim 1, wherein, in step S3, the reference sequence of the pathogenic microorganism is the reference sequence of the pathogenic microorganism after filtering; preferably, the filtering of the reference sequence of the pathogenic microorganism includes the following steps: filtering out sequences of host origin; preferably, the step of filtering out sequences of host origin includes filtering out sequences of microorganisms or viruses integrated in the genomic sequence of the host.
7. The analysis method according to claim 1, wherein, in step S3, the filtering out of the sequences with mismatches includes: filtering using a dynamic mismatch threshold, preferably, when the length of the sequenced sequence is 30 - 100bp, filtering is performed using a dynamic mismatch threshold; preferably, when the sequence length is 30 - 40bp, the mismatch threshold is set to 0bp, and sequences with a mismatch value > 0bp are filtered out. Preferably, when the sequence length is 41 - 50 bp, set the mismatch threshold to 2 bp, and filter out sequences with a mismatch value > 2 bp; Preferably, when the sequence length is 51 - 79 bp, set the mismatch threshold to 3 bp, and filter out sequences with a mismatch value > 3 bp; Preferably, when the sequence length is 80 - 100 bp, set the mismatch threshold to 4 bp, and filter out sequences with a mismatch value > 4 bp.
8. The analysis method according to claim 1, wherein, in step S4, the bioinformatics analysis of the second sequencing data includes: obtaining the species corresponding to the sequences in the second sequencing data according to the alignment results to obtain a predicted species set; selecting species supported by non - repetitive sequences as predicted species; and merging all predicted species to obtain the complete set of predicted species of the sample, and obtaining the analysis result of pathogenic microorganisms, Preferably, the analysis result of the pathogenic microorganisms includes the Latin name of the species, the number of sequences at the species level, the relative abundance at the species level, the Latin name of the genus, the number of sequences at the genus level, the relative abundance at the genus level, the sequence proportion of each species within the genus, the Latin name of the family, the number of sequences at the family level, the relative abundance at the family level, the sequence proportion of each genus within the family, and taxonomic lineage information.
9. A device for pathogenic microorganism sequencing data, the device comprises: an acquisition module: used to acquire the sequencing data of a sample to be tested; a first alignment module: align the sequencing data with the reference sequence of the host, and filter out sequences identical to the reference sequence of the host to obtain the first sequencing data; a second alignment module: align the first sequencing data with the reference sequence of the pathogenic microorganism, and filter out sequences with mismatches to obtain the second sequencing data; a bioinformatics analysis module: perform bioinformatics analysis on the second sequencing data to obtain the analysis result of pathogenic microorganisms.
10. The device according to claim 9, wherein, the device further comprises a chimera sequence correction module, which corrects the chimera sequences introduced during the fragmentation of the sample nucleic acid; Preferably, the device further comprises a reference sequence filtering module, which filters out sequences from the host; Preferably, the reference sequence filtering module filters out sequences of microorganisms or viruses integrated in the host genomic sequence; Preferably, the filtering out of sequences with mismatches is performed using a dynamic mismatch threshold; Preferably, the device further comprises a module for filtering out low - specificity sequences according to probes, which includes: judging the specificity of the probe: by aligning the probe with the reference sequences of two or more species, judge the specificity of the probe, wherein when the mismatch value of the alignment is below the threshold, the probe aligned to two or more species is marked as a low - specificity probe; and filtering out the sequencing data captured by the low - specificity probe, where the sequencing data is the first sequencing data and / or the second sequencing data.
11. A sequencing data analysis system for targeted detection of microorganisms, comprising a computer processor and a memory, the memory storing computer program instructions which, when executed by the computer processor, cause the processor to perform the steps of the analysis method according to any one of claims 1 to 8.
12. Use of the analysis method according to any one of claims 1 to 8, the device according to any one of claims 9 to 10, or the analysis system according to claim 10 in detecting pathogenic microorganisms in a sample.
13. According to the use according to claim 12, characterized in that the sample comprises a tissue or body fluid sample, preferably selected from tissue, cells, blood, plasma, serum, bronchoalveolar lavage fluid, sputum, pus, nasopharyngeal swab, oral swab, cerebrospinal fluid, pleural effusion, ascites, or a processed product thereof; preferably, the microorganism is a pathogenic microorganism; preferably, the microorganism comprises eukaryotic microorganisms and prokaryotic microorganisms; preferably, the eukaryotic microorganism comprises any one of yeast, fungi or algae; Preferably, the prokaryotic microorganisms include one or more of the following: Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli, Moraxella catarrhalis, Enterobacter cloacae, Serratia marcescens, Burkholderia cepacia, Klebsiella oxytoca, Enterobacter aerogenes, Proteus mirabilis, Streptococcus pyogenes, Haemophilus parainfluenzae, Citrobacter freundii, Cryptococcus neoformans, Candida auris, Aspergillus fumigatus, Candida albicans, influenza virus, Human Respiratory syncytial virus, Rhinovirus, Human parainfluenza virus, Human adenovirus, coronavirus, Human metapneumovirus, Human bocavirus, Mycobacterium tuberculosis complex, Mycobacterium avium, Mycobacterium intracellulare, Legionella pneumophila, ListeriaListeria monocytogenes, Mycoplasma pneumoniae, Chlamydia pneumoniae, Chlamydia psittaci, Streptococcus constellatus, Coxiella burnetii, Blastomyces dermatitidis, Coccidioides immitis, Cytomegalovirus, adenovirus, Lymphocytic choriomeningitis virus, Mycobacterium intracellulare, Mycobacterium avium, Mycobacteroides abscessus and its three subspecies (Mycobacteroides abscessus subsp. Abscessus, Mycobacteroides abscessus subsp. Bolletii, Mycobacteroides abscessus subsp. massiliense), Mycobacterium canettii, Mycobacterium interjectum, Mycobacterium kansasii, Mycobacterium marinum, Mycobacteroides chelonae, Mycolicibacterium fortuitum, Mycobacterium gordonae, Mycobacterium malmoense, Mycobacterium scrofulaceum, Mycobacterium simiae, Mycobacterium xenopi, Mycobacterium ulcerans, SerratiaSerratia marcescens, Salmonella, Streptococcus pneumoniae, Enterococcus faecium, and Enterococcus faecalis; preferably, the influenza virus comprises influenza A virus (H1N1), influenza A virus (H3N2), influenza A virus (H5N1), influenza A virus (H7N9), influenza A virus (H9N2), influenza B virus; preferably, the coronavirus comprises human coronavirus HKU1, human coronavirus NL63, human coronavirus OC43, novel coronavirus, SARS-CoV-2.
14. A storage medium having a computer program stored thereon, the computer program, when executed by a processor, implementing the steps of the analysis method according to any one of claims 1 to 8.
Citation Information
Cited By
Analysis method, device and equipment for pathogen targeted high-throughput sequencing data and medium
CN120690291A