A method for identifying host exogenous nucleic acids based on high-throughput sequencing technology

By combining multiple comparison and filtering strategies with a variety of software tools, the problem of accuracy in identifying exogenous nucleic acids in a high host background was solved, and accurate identification of parasite, microbial and plant nucleic acid sequences was achieved, reducing the false positive rate.

CN119541642BActive Publication Date: 2025-09-26SOUTH CHINA UNIV OF TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202410967216.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2025-09-26
Estimated Expiration
2044-07-18

AI Technical Summary

Technical Problem

Existing high-throughput sequencing technologies have difficulty in accurately distinguishing and identifying low-abundance exogenous nucleic acid sequences when identifying host exogenous nucleic acids, especially in high host backgrounds, resulting in false positive results and inaccurate identification, especially in the identification of parasitic infections and plant nucleic acid sequences.

Method used

A multiple alignment filtering strategy, including local alignment, k-mer alignment, metagenomic analysis, BLAST alignment, and NT library alignment, was used in combination with various software tools such as BOWTIE2, Kraken2, Krakenuniq, BLAST, and PRINSEQ. Through multiple filtering and alignment, host sequences and low-complexity sequences were removed to improve identification accuracy.

Benefits of technology

It effectively removes host sequence background noise, improves the accuracy and credibility of identification of exogenous nucleic acid sequences, reduces false positives, and can accurately identify weak exogenous nucleic acid signals in a high host background.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541642B_ABST
    Figure CN119541642B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying host exogenous nucleic acids based on high-throughput sequencing technology. The method comprises: obtaining sequencing data of a sample to be tested; using a host reference genome database, filtering the sequencing data twice to remove host sequences; classifying the sequenced data after host filtration to reduce interference from homologous species sequences; using the identified species reference genome to align sequences to obtain accurate species sequences; removing low-complexity sequences from the aligned sequences to reduce the influence of sequence bias; and using BLAST to perform a secondary alignment of the species reference genome or align it to an NT library. The present invention can accurately and specifically identify exogenous nucleic acid sequences in samples from different hosts, enabling efficient analysis of samples with high host background and low exogenous species biomass.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and specifically to a method for identifying host exogenous nucleic acids based on high-throughput sequencing technology. Background Art

[0002] With the development and reduction in cost of high-throughput sequencing technology, clinical disease screening, identification of sequences from different species, and exploration of the sources of exogenous nucleic acid sequences in hosts have become possible through sequencing. Sequencing different host samples, such as lung cell lavage fluid, tissue samples, and urine, to identify exogenous nucleic acid sequences, can effectively assist in the diagnosis of infections, including bacterial, viral, fungal, and parasitic infections (such as echinococcosis and liver fluke disease), and can detect a wider range of pathogens than traditional PCR methods. However, because these samples come from hosts, the host sequence background is high. For example, human sequences account for over 98% of the sequences in plasma samples. This can significantly reduce the accuracy of bloodstream infections due to the influence of host sequences. In addition to microbial sequences, plants and animals, both eukaryotic organisms, are more susceptible to the influence of homologous sequences, low-complexity sequences, and sequences from closely related species. Current analytical methods, even when removing native host sequences and homologous sequences, still retain residual sequences, which can easily produce inaccurate results in the presence of high host background and low abundance of exogenous sequences. Furthermore, current analytical methods struggle to distinguish the exact animal and plant sequences of each species when processing mixed sequences from multiple species (human and parasite, human and microbial, or human and plant samples), particularly when samples contain both animal and plant sequences. Therefore, bioinformatics analysis methods for identifying the source of exogenous nucleic acids in hosts based on high-throughput sequencing technologies need further refinement.

[0003] Currently, when using high-throughput sequencing technology to detect exogenous nucleic acid sequences in a host, the focus is on studying exogenous nucleic acid sequences of bacteria and viruses that infect humans, which are present in various human samples (blood, sputum, and body fluids, etc.). For example, there are methods and devices for detecting microorganisms in blood (CN105525033A), methods and devices for performing microbial analysis on host samples (CN111009286A), methods for determining the presence of predetermined species in environmental samples through metagenomic sequencing (CN114574565A), pathogenic microorganism analysis and identification systems and applications (CN111462821A), a metagenomic data analysis method and system (CN108334750B), and metagenomics-based pathogen analysis methods, analysis devices, equipment, and storage media (CN111951895A). These methods primarily perform simple quality control on the data from high-throughput sequencing technology. After filtering the host sequences in a single step, the data is directly compared or blasted against a constructed database of pathogenic bacteria or viruses to collect nucleic acid sequence information for the relevant bacterial or viral species. Although this method can detect sequences of more infectious species as much as possible, it may cause false positives or fail to detect in the early stages of infection when the content of exogenous nucleic acid sequences is low and the background noise of host nucleic acid sequences is large. In addition, host humans and microorganisms are far apart in biological classification and are relatively easy to distinguish, but when there is a parasitic infection in the host human, it is more difficult to distinguish the nucleic acid sequences of the two. For example, during the inventor's research and development process, it was found that the blood samples of patients suspected of being infected with echinococcosis had a low content of nucleic acid sequences released in the early stages of parasitic infection, so it was difficult for current conventional analysis methods to capture specific signals. In addition, few people are currently focusing on studying whether plant nucleic acid sequences may exist in different hosts. Studies have shown that plant RNA sequences exist in mammals. Therefore, in the face of nucleic acids with stronger homology between animal and animal sequences, and animal and plant sequences, accurate nucleic acid sequence identification methods are more needed to obtain accurate and reliable results. Summary of the Invention

[0004] The present invention aims to address one of the technical problems in the related art to a certain extent. To this end, the present invention aims to provide a bioinformatics analysis method based on high-throughput sequencing technology that can rapidly and accurately identify the species of microbial, animal, and plant sequences in samples with a high host background, thereby improving the reliability of detecting non-host species sequences. By combining multiple sequence alignment strategies and algorithms, the present invention effectively improves the reliability and accuracy of species classification, and can assist in the extensive identification of sequences from different host samples, such as pathogenic microorganisms and other exogenous species.

[0005] The purpose of the present invention is achieved by at least one of the following technical solutions.

[0006] A method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology comprises the following steps:

[0007] S1. Obtain whole genome sequencing data, targeted sequencing data, or metagenomic sequencing data of the host sample and perform routine quality control on the sequencing data;

[0008] S2. Remove host sequences from sequencing data: Obtain the latest host reference genome, filter out nucleic acid sequences in the sequencing data that are identical to the host reference genome nucleic acid sequences, and complete the first filtering;

[0009] S3, filtering out the nucleic acid sequences identical to the host reference genome nucleic acid sequences in the sequencing data after the first filtration, completing the second filtration;

[0010] S4. Specify the species to be analyzed, including parasites, microorganisms, and other species that do not originate from the host, obtain the reference genome sequence of the specified analysis species, and construct a species classification database; use metagenomic analysis software to align the sequencing data after the second filtration to the species classification database to obtain the accurate nucleic acid sequence of the specified analysis species;

[0011] S5. Re-align the accurately classified nucleic acid sequence of the designated analysis species to the reference genome sequence of the designated analysis species, retain the paired-end sequences that are successfully aligned to the reference genome of the designated analysis species, filter out the results that cannot be aligned and the results of single-end alignment, and obtain the accurate species nucleic acid sequence;

[0012] S6. removing low-complexity nucleic acid sequences from the accurate species nucleic acid sequences;

[0013] S7. Using the BLAST alignment algorithm, align the accurate species nucleic acid sequence after removing the low-complexity nucleic acid sequence to the reference genome sequence of the designated analysis species, retaining the paired-end sequence alignment results with a confidence level greater than 90 and a coverage greater than 90% in the BLAST alignment to obtain a nucleic acid sequence with high confidence;

[0014] S8. Align the highly reliable nucleic acid sequences to the NCBI NT database, retain the paired-end sequence alignment results with a reliability greater than 90 and a coverage greater than 90% in the BLAST comparison, further filter the nucleic acid sequences of distantly related species, and finally obtain the nucleic acid sequences that are truly from the specified analysis species;

[0015] S9. If the actual nucleic acid sequence from the designated analysis species obtained in step S8 is the same as the accurate sequence of the designated analysis species, it indicates that the host sample contains the nucleic acid sequence of the designated analysis species; otherwise, it indicates that the host sample does not contain the nucleic acid sequence of the designated analysis species.

[0016] Furthermore, the host sample includes a body fluid sample or a tissue sample.

[0017] This method can accurately identify the presence of exogenous nucleic acid sequences in the host, such as some parasites, bacteria, viruses, and potential non-host nucleic acid sequences in the human body. This method can effectively distinguish mixed nucleic acid sequences of combinations of animals and plants, animals and microorganisms, plants and microorganisms, etc., and subdivide them into specified species. Here, the host of this method generally refers to the species with the largest sequence proportion in the mixed sequence, and the rest can be called exogenous nucleic acid sequences. The specific demarcation can be based on the needs of actual sample analysis.

[0018] Furthermore, the nucleic acid sequence includes a DNA sequence and an RNA sequence.

[0019] Furthermore, in step S1, the conventional quality control refers to removing sequencing residual adapter sequences and low-quality sequencing reads; the low-quality sequencing reads include sequencing reads with an N ratio greater than 5% and sequencing reads with a length less than 30.

[0020] Furthermore, during the second filtering, a different alignment algorithm than that used during the first filtering needs to be used for nucleic acid sequence filtering.

[0021] Furthermore, in step S4, if there is a species that may have homology with the sequence of the designated analysis species during the analysis process and affect the analysis results, the reference genome sequence of the species affecting the analysis is obtained, and the genomes of the designated analysis species and the species affecting the analysis are mixed to construct a species classification database.

[0022] Furthermore, in step S6, the low-complexity nucleic acid sequence refers to a nucleic acid sequence in which the sequencing reads have an interval of more than two consecutive identical bases with a length exceeding 5%.

[0023] Furthermore, in step S6, different software is needed to be used than in step S5 when aligning to the species reference genome to remove low-complexity nucleic acid sequences from the accurate species nucleic acid sequence; thereby, the accuracy of the sequence can be verified again, effectively reducing the systematic errors and biases caused by different software.

[0024] Furthermore, in step S8, the distant species refers to species that are annotated to different species at the family level when using the NT library for species annotation. These species may have similar sequences in evolutionary terms, but their taxa are far apart. Such sequences should not be considered as the true sequence of the specified species, such as the results of cross-kingdoms such as eukaryotes and bacteria, or eukaryotes and viruses. After filtering by distant species, the specific sequence of the species can be obtained, greatly improving the credibility of the sequence information.

[0025] Furthermore, if the reference genome sequence of the specified analysis species is not included in the NCBI NT library, no valid results will be produced. In this case, the NT library comparison can be ignored, that is, step S8 is not executed, and the highly reliable sequence obtained in step S7 is directly used as the accurate sequence of the specified analysis species, and then step S9 is executed. At this time, the interference of sequences of distant species is not removed.

[0026] Compared with the prior art, the advantages of the present invention are:

[0027] Through multiple comparison and filtering, the present invention can almost completely remove host sequences from sequencing data containing a high proportion (99%) of host sequences, capture weak exogenous nucleic acid sequence signals, and accurately identify nucleic acid sequences from exogenous sources, effectively improving the accuracy of analysis results and reducing false positives. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a schematic diagram of an information analysis process for a method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology in Example 1 of the present invention;

[0029] Figure 2 This is a schematic diagram of an information analysis process for a method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology in Example 2 of the present invention;

[0030] Figure 3 This is a schematic diagram of the genome coverage of the identified exogenous sequences according to the results of Example 3 of the present invention. DETAILED DESCRIPTION

[0031] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings, which are intended to explain the present invention but should not be construed as limiting the present invention.

[0032] Example 1 - Identification of Echinococcus Sequences in Patients Infected with Hydatid Disease

[0033] Human echinococcosis is a parasitic disease caused by tapeworms of the genus Echinococcus, also known as hydatid disease. Human infection with Echinococcus often results in the formation of hydatid cysts in the liver, and the incubation period can be as long as several years. Currently, the diagnosis of echinococcosis relies primarily on imaging, but it is difficult to detect in the early latent stages, significantly increasing the cost of treatment. Early detection of echinococcosis can be achieved by sequencing free nucleic acids from Echinococcus in the blood. However, due to the excessive background noise from host sequences, sequences from Echinococcus are difficult to distinguish, which can easily lead to false positives. More accurate analytical methods are needed to distinguish between parasite sequences and host sequences.

[0034] To evaluate the effectiveness of the inventive method in distinguishing parasite sequences from host sequences, in this example, 50 simulated data samples were generated using Wgsim (https: / / github.com / lh3 / wgsim) with default parameters using human hosts. In addition to 10 million paired-end host sequences, 1000 paired-end sequencing sequences of each of the two main infecting species, Echinococcus granulosus (GCA_021556725.1) and Echinococcus multilocularis (GCA_000469725.3), were also included.

[0035] A method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology, such as Figure 1 As shown, the following steps are included:

[0036] S1. Obtain whole genome sequencing data, targeted sequencing data, or metagenomic sequencing data of the host sample and perform routine quality control on the sequencing data;

[0037] The host sample includes a body fluid sample or a tissue sample.

[0038] In this example, we first created an alignment database for BOWTIE2 and Kraken2 using the hg38 reference genome downloaded from NCBI. Subsequently, we downloaded the reference genomes of Echinococcus granulosus and Echinococcus multilocularis from NCBI and created an alignment database using Krakenuniq, BOWTIE2, and BLAST. These databases were used for subsequent bioinformatics analysis.

[0039] The conventional quality control refers to the removal of sequencing residual adapter sequences and low-quality sequencing reads; the low-quality sequencing reads include sequencing reads with an N ratio greater than 5% and sequencing reads with a length less than 30.

[0040] S2. Removing host sequences from sequencing data: Obtain the latest host reference genome. In this embodiment, a local alignment algorithm is used to filter out nucleic acid sequences in the sequencing data that are identical to the host reference genome nucleic acid sequence, completing the first filtering.

[0041] The nucleic acid sequence includes DNA sequence and RNA sequence.

[0042] The preferred software is BOWTIE2, and the parameter is set to "-very-senstive-local" to use a local alignment method to remove sequences from the host as completely as possible.

[0043] S3. In this embodiment, the k-mer alignment algorithm software is used to filter out the nucleic acid sequences identical to the host reference genome nucleic acid sequences in the sequencing data after the first filtration, thereby completing the second filtration;

[0044] The preferred software is Kraken2. Use the Kraken2-build function to build an alignment database based on the host reference genome. Set the confidence parameter to 0, which will classify as many sequences as possible to the host and remove as many host sequences as possible.

[0045] In this example, the initial analysis data was aligned to the hg38 reference genome using BOWTIE2, and sequences on the aligned genome were filtered out, retaining unaligned paired-end sequences. Preferably, BOWTIE2 uses the "-very-sensitive-local" parameter. The retained data was then aligned to the hg38 reference genome using Kraken2, retaining unaligned paired-end sequences. Preferably, Kraken2 uses the "-paired-confidence=0" parameter.

[0046] S4. Specify the species to be analyzed, including parasites, microorganisms, and other species that do not originate from the host, obtain the reference genome sequence of the specified analysis species, and construct a species classification database; use metagenomic analysis software to align the sequencing data after the second filtration to the species classification database to obtain the accurate nucleic acid sequence of the specified analysis species;

[0047] If there are species that may have homology with the sequences of the designated analysis species and thus affect the analysis results during the analysis process, the reference genome sequence of the species affecting the analysis is obtained, and the genomes of the designated analysis species and the species affecting the analysis are mixed to construct a species classification database.

[0048] Preferred software is MetaPhlAn4 or Krakenuniq. In this paper, Krakenuniq is used as an example to establish a species classification database based on species sequences and annotate and classify the sequences. This step can utilize unique k-mers to better classify multiple homologous sequences and reduce interference between model species.

[0049] In this example, the data remaining after filtering host sequences was aligned to a constructed species database using Krakenuniq using the "-paired" parameter. Sequences annotated to Echinococcus granulosus and Echinococcus multilocularis were extracted for subsequent analysis. This step effectively separated the sequences of the two homologous Echinococcus species.

[0050] S5. In this embodiment, the end-to-end alignment algorithm is used to re-align the accurately classified nucleic acid sequence of the specified analysis species to the reference genome sequence of the specified analysis species, retaining the paired-end sequences that are successfully aligned to the reference genome of the specified analysis species, filtering out the results of unaligned and single-end alignments, and obtaining the accurate species nucleic acid sequence;

[0051] In this example, the obtained sequences of Echinococcus granulosus and Echinococcus multilocularis were aligned to their respective genomes using BOWTIE2. The paired-end sequence alignment results were then extracted using SAMTOOLS. Preferably, BOWTIE2 uses an "end-to-end" alignment mode.

[0052] S6. In this embodiment, the DUST algorithm is used to remove low-complexity nucleic acid sequences from the accurate species nucleic acid sequences;

[0053] Low-complexity nucleic acid sequences refer to sequences in which sequencing reads contain more than two consecutive identical bases with a base length exceeding 5%, such as the protein sequence "PPCDPPPPPKDKKKKDDGPP" and the nucleic acid sequence "AATAAAAAAAATAAAAAAT." Because these sequences have strong biases and a single base, they correspond to multiple species during sequence similarity searches, making effective species identification impossible. Therefore, it is necessary to remove these biased sequences to improve identification accuracy.

[0054] The preferred software is PRINSEQ. The parameters of the software are set as follows: the low complexity sequence masking algorithm is DUST, and the calculated base range is 7.

[0055] In this example, the sequences aligned to Echinococcus were used to remove low-complexity sequences using PRINSEQ. Preferably, PRINSEQ uses the parameters "-lc_method dust-lc_threshold 7".

[0056] S7. Using the BLAST alignment algorithm, align the accurate species nucleic acid sequence after removing the low-complexity nucleic acid sequence to the reference genome sequence of the designated analysis species, retaining the paired-end sequence alignment results with a confidence level greater than 90 and a coverage greater than 90% in the BLAST alignment to obtain a nucleic acid sequence with high confidence;

[0057] S8. Align the highly reliable nucleic acid sequences to the NCBI NT database, retain the paired-end sequence alignment results with a reliability greater than 90 and a coverage greater than 90% in the BLAST comparison, further filter the nucleic acid sequences of distantly related species, and finally obtain the nucleic acid sequences that are truly from the specified analysis species;

[0058] Use BLAST to compare the species reference genome or NT library, giving priority to the latest database on NCBI;

[0059] If the reference genome sequence of the specified analysis species is not included in the NCBI NT library, no valid results will be produced. In this case, the NT library comparison can be ignored, that is, step S8 is not executed, and the highly reliable sequence obtained in step S7 is directly used as the accurate sequence of the specified analysis species, and then step S9 is executed. At this time, the interference of sequences of distantly related species is not removed.

[0060] S9. If the actual nucleic acid sequence from the designated analysis species obtained in step S8 is the same as the accurate sequence of the designated analysis species, it indicates that the host sample contains the nucleic acid sequence of the designated analysis species; otherwise, it indicates that the host sample does not contain the nucleic acid sequence of the designated analysis species.

[0061] In this embodiment, BLAST uses the parameters "-evalue 1e-5-perc_identity 90-qcov_hsp_perc90".

[0062] The information analysis results of Example 1 using the present invention are shown in Table 1. The results demonstrate that the present invention completely removes host-derived sequences, while ultimately achieving recovery rates of 95.76% and 95.89% for sequences derived from Echinococcus granulosus and Echinococcus multilocularis, respectively. Even in plasma samples from patients with early-stage Echinococcus infection, even when containing low amounts of Echinococcus sequences, the present invention effectively reduces host interference and preserves trace Echinococcus sequence signals.

[0063] Table 1 Information analysis results of Example 1

[0064] Step Echinococcus_granulosus Echinococcus_multilocularis Homo_sapiens Raw 991 998 10000000 Bowtie2 991±0 998±0 4325852±1887.69 Kraken2 984.35±0.80 995.47±0.89 5013.80±1212.93 Krakenuniq 976.84±0.75 986.27±0.45 0 Bowite2 971.39±4.40 974.39±4.01 0 PRINSEQ 970.65±3.96 973.65±4.46 0 BLAST 949.04±1.76 957.27±0.45 0 BLAST recovery rate (%) 95.76% 95.89% 0

[0065] In Example 1, under the mixed sequences of all nine model species, the present invention can still well distinguish the sequences of the nine species, and maintain consistency and specificity, with a good sequence recovery rate, which confirms the effectiveness of the analysis method.

[0066] Example 2 - Identification of contaminating sequences from human host sequences mixed with different model species:

[0067] During the experiment and sequencing process, there may be sequence contamination from some model species of animals, plants or microorganisms. If you need to obtain a pure sequence after removing the contamination, this method can be used to effectively separate the specified sequence of the corresponding species and perform sequence recovery.

[0068] like Figure 2 As shown, a method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology comprises the following steps:

[0069] S1. Obtain whole genome sequencing data, targeted sequencing data, or metagenomic sequencing data of the host sample and perform routine quality control on the sequencing data;

[0070] In this example, in order to test the accuracy and consistency of the analysis method, 50 simulated data samples were generated using Wgsim (https: / / github.com / lh3 / wgsim) using default parameters with humans as the host. In addition to the 10 million double-end sequencing host sequences, it also contains mixed sequences of 2 plant species sequences (potatoes and tomatoes), 2 animal species sequences (domestic pigs and chickens), and 2 microbial species sequences (new coronavirus and Escherichia coli). In the simulated data samples, the sequence of each species accounts for only 0.01% of the host sequence, all of which are 1000 double-end sequencing reads. The specific reference genome information is shown in Table 2 below.

[0071] Table 2 Simulation data construction information

[0072] Species Total of reads (PE) Assembly_accession* Sample number Solanum tuberosum (potato) 1000 GCA_014189475.1 50 Solanum lycopersicum (tomato) 1000 GCF_000188115.5 50 Susscrofa (domestic pig) 1000 GCF_000003025.6 50 Gallusgallu (chicken) 1000 GCF_016699485.2 50 SARS-CoV-2 (Novel Coronavirus) 1000 GCF_009858895.2 50 Escherichia coli 1000 GCF_000005845.2 50 Homosapiens (people) 10000000 GCF_000001405.40 50

[0073] *Assembly accession: Assembly accession number, the sequence accession number assigned to the genome assembly format by NCBI database staff.

[0074] S2. Removing host sequences from sequencing data: Obtain the latest host reference genome, and use local alignment algorithm software to filter out nucleic acid sequences in the sequencing data that are identical to the host reference genome nucleic acid sequences, completing the first filtering;

[0075] In this embodiment, the generated simulated data (simulated data does not require quality control) is aligned to the human reference genome using alignment software, the aligned sequences are removed, and the sequences that cannot be aligned are retained. Preferably, the "-very-sensitive-local" parameter of BOWTIE2 is used for local alignment. In addition, alignment software such as BWA can also be used for sequence alignment, and then SAMTOOLS is used to extract sequences that may not be aligned. The human reference genome here uses the hg38 reference genome downloaded from NCBI, and the human T2T genome can also be selected. This step of processing can reduce a large amount of host sequences and reduce the interference caused to subsequent analysis.

[0076] S3, using k-mer alignment algorithm software to filter out the nucleic acid sequences in the sequencing data after the first filtration that are identical to the host reference genome nucleic acid sequence, completing the second filtration;

[0077] In the present embodiment, the metagenomic analysis software Kraken2 is selected, and a host reference database is established based on the reference genome of the human. Kraken2 uses the k-mer alignment algorithm, which can further filter the host sequence. So the sequence preliminarily filtered in step 2 is aligned to the host database by Kraken2, and the parameter "-paired confidence 0" is used to filter the sequence classified into the host, retain the sequence that cannot be classified, and output the unclassified sequence as a fastq format file. By filtering for the second time, sequences similar to and homologous to the host can be further filtered out, and the remaining sequences hardly exist in the sequence belonging to the host. The human reference genome of the host database set up in this step is hg38, and the T2T genome of the human can also be selected to carry out the construction of the database.

[0078] S4. Specify the species to be analyzed, including parasites, microorganisms, and other species that do not originate from the host, obtain the reference genome sequence of the specified analysis species, and construct a species classification database; use metagenomic analysis software to align the sequencing data after the second filtration to the species classification database to obtain the accurate nucleic acid sequence of the specified analysis species;

[0079] In the present embodiment, metagenome analysis software Krakenuniq is selected, and according to the above-mentioned 6 specified species reference genomes, a species classification database is established based on the above-mentioned species using Krakenuniq-build function. In the process of setting up the classification database, it is necessary to download the taxdump.tar.gz classification tree information file from NCBI to assist in establishing a complete classification result. Subsequently, the sequence after step 3 filtering is subjected to detailed species classification, and the sequence of the specified species is extracted by the extract_kraken_reads.py script of Krakentools, and output is a fastq format result. By setting up a homologous database of multiple species, the sequence classification of species is carried out, and the interference of homologous species sequences can be effectively reduced, and the accuracy and consistency of analysis are improved. The reference genomes for setting up the database in this step are all sourced from NCBI, and specific information is referenced to Table 2.

[0080] S5. Use an end-to-end alignment algorithm to re-align the accurately classified nucleic acid sequence of the specified analysis species to the reference genome sequence of the specified analysis species, retain the paired-end sequences that are successfully aligned to the reference genome of the specified analysis species, filter out the results that cannot be aligned and the results of single-end alignment, and obtain the accurate species nucleic acid sequence;

[0081] In the present embodiment, the sequence after species classification is compared to the reference genome of specified species using alignment software BOWTIE2 using default parameters. Preferably, SAMTOOLS is used, the double-ended sequence successfully compared to species genome is selected, and the result (flag value includes 4 and 8 results) of single-ended comparison cannot be compared. Preferably, the double-ended comparison result compared to species genome is retained, i.e., the specified species sequence. By comparing back to the genome of species, the sequence unrelated to the species can be filtered out, while the pollution caused from host and other species can also be further reduced. Alignment software can also use the software of different precisions such as BWA or SNAP-aligner, improves the reliability of comparison result. Selecting the higher reference species genome of assembly quality and integrity simultaneously can also improve the accuracy of result.

[0082] S6. Using the DUST algorithm to remove low-complexity nucleic acid sequences from accurate species nucleic acid sequences;

[0083] In this example, low-complexity sequences are present in multiple species, exhibit bias, and are unable to effectively distinguish between sources. Therefore, low-complexity sequences need to be filtered to reduce their impact on subsequent analysis. Preferably, the software PRINSEQ is used, with the parameters "-lc_meathod dust -lc_threshold 7" set to filter low-complexity sequences and retain sequences with higher specificity for subsequent analysis.

[0084] S7. Using the BLAST alignment algorithm, align the accurate species nucleic acid sequence after removing the low-complexity nucleic acid sequence to the reference genome sequence of the designated analysis species, retaining the paired-end sequence alignment results with a confidence level greater than 90 and a coverage greater than 90% in the BLAST alignment to obtain a nucleic acid sequence with high confidence;

[0085] In this embodiment, in order to further improve the accuracy of the sequence screened out in step 4, a BLAST database of the reference genome of the specified species was constructed. Preferably, the sequence of step 4 was subjected to a secondary alignment using BLAST alignment software, and the double-end sequence alignment results containing the BLAST alignments with a credibility greater than 90 and a coverage greater than 90% were selected, and the output results were in fasta format. The results that could not be aligned and the single-end alignments were filtered out to obtain a clean sequence of the specified species. The accuracy of the sequence was further verified by different alignment algorithms. Selecting multiple alignment algorithms can further reduce the bias brought by the software itself, improve the accuracy and uniqueness of the species sequence, and improve the precision for subsequent analysis.

[0086] S8. Align the highly reliable nucleic acid sequences to the NCBI NT database, retain the paired-end sequence alignment results with a reliability greater than 90 and a coverage greater than 90% in the BLAST comparison, further filter the nucleic acid sequences of distantly related species, and finally obtain the nucleic acid sequences that are truly from the specified analysis species;

[0087] S9. If the actual nucleic acid sequence from the designated analysis species obtained in step S8 is the same as the accurate sequence of the designated analysis species, it indicates that the host sample contains the nucleic acid sequence of the designated analysis species; otherwise, it indicates that the host sample does not contain the nucleic acid sequence of the designated analysis species.

[0088] In this embodiment, in order to further eliminate the interference of similar sequences of homologous species. The clean sequence of the specified species is preferably aligned to the NT (May 6, 2023) library downloaded from NCBI using BLAST software. The results with a credibility of less than 90 and a coverage of less than 90% in the alignment results are removed. The double-end results that can be aligned to the classification of the specified species are retained. Finally, results with high credibility and accuracy of the specified species are obtained, which are sufficient to support subsequent further analysis and downstream applications of the species sequence. Through this step, the interference of distant species sequences is removed, and the highest accuracy results are obtained compared to previous methods. Table 3 shows the analysis results of each step of the simulated 2 plant sequences. In the results, through two-step host sequence removal, the host sequence can be removed by 99.999713%, which is almost completely removed. In the third step of the analysis step, the host sequence is no longer included in the sequence of the specified species. After the second alignment, the sequence recovery rates of BLAST to the specified species were 89.2% and 91.29%, respectively, and the sequence consistency was high. Furthermore, even after rigorous BLAST alignment to NT screening, sequence recovery rates reached 71.1% and 91.39%, respectively. This high recovery rate was maintained despite eliminating interference from most distantly related species and homologous sequences, confirming the method's ability to accurately analyze mixed sequences from animals and plants.

[0089] Table 3 Sequence analysis results of two plants in Example 2

[0090]

[0091]

[0092] Table 4 shows the analysis results of each step of the simulated two animal sequences. In step S3, the host sequence is no longer included in the sequence of the specified species. After the final secondary alignment, the sequence recovery rates of BLAST to the specified species were 91.4% and 79.18%, respectively, and the sequence consistency was high. At the same time, even after the strict screening conditions of BLAST alignment to NT, the sequence recovery rates reached 57% and 24.12%, respectively, excluding the interference of most distant species and homologous sequences. Although the recovery rate is not very high, for higher animals, the NT library itself contains a large number of indistinguishable homologous and contaminated sequences. Therefore, through filtering and quality control by this method, relatively accurate results can still be obtained after the secondary alignment. It is confirmed that this method can find species-specific sequences in animals and animal mixed sequences, thereby improving the accuracy of the identification results.

[0093] Table 4 Sequence analysis results of two animals in Example 2

[0094] Step Gallus_gallus(PE) Sus_scrofa(PE) Homo sapiens(PE) Raw 1000 999 1000000 Bowtie2 997.4±0.73 974.4±4.19 4320973.26±2174.37 Kraken2 968.62±4.63 852.2±11.59 287.04±12.82 Krakenuniq 968.54±4.61 851.7±11.59 0 Bowtie2 956.54±4.93 837.74±12.72 0 PRINSEQ 950.48±3.91 831.56±12.43 0 BLAST 914.6±8.46 791.56±14.37 0 BLAST recovery rate (%) 91.40% 79.18% 0 BLAST_NT 570.24±6.99 241.76±10.76 0 BLAST_NT recovery rate (%) 57.00% 24.12% 0

[0095] Table 5 shows the analysis results of each step of the simulated two microbial sequences. Microorganisms and hosts belong to different kingdoms in biological taxonomy and have a large species gap, so a high recovery rate can be obtained. Despite this, microbial pathogen detection still requires more accurate identification methods to reduce false positive results in clinical diagnosis. After secondary alignment, the sequence recovery rates of BLAST to the specified species were 97% and 97.1%, respectively, and the sequence consistency was high. At the same time, even after the strict screening conditions of BLAST alignment to NT, the sequence recovery rates reached 97.3% and 98.7%, respectively. Even after eliminating the interference of most distant species and homologous sequences, a high recovery rate was still maintained. This confirms that this method has relatively accurate analysis results in mixed animal and microbial sequences.

[0096] Table 5 Results of sequence analysis of two microorganisms in Example 2

[0097] Step Escherichia_coli(PE) SRAS-COV-2(PE) Homo sapiens(PE) Raw 1000 1000 1000000 Bowtie2 1000±0 1000±0 4320973.26±2174.37 Kraken2 999.96±0.2 999.88±0.33 287.04±12.82 Krakenuniq 999.44±0.86 998.38±0.53 0 Bowtie2 987.3±3.54 987.98±1.53 0 PRINSEQ 987.3±3.54 987.98±1.53 0 BLAST 970.7±4.62 971.04±5.98 0 BLAST recovery rate (%) 97.00% 97.10% 0 BLAST_NT 973.52±4.68 987.84±1.77 0 BLAST_NT recovery rate (%) 97.30% 98.70% 0

[0098] Example 3 - Identification of Endobacterium Cardiomycosis Infection in Patients with Endocarditis:

[0099] Pathogen identification is fundamental to the diagnosis of endocarditis (IE) and can guide clinical treatment and antibiotic selection. Currently, 31% of IE patients have negative microbial culture results from blood or heart valve samples. Metagenomic analysis using high-throughput sequencing technology can effectively identify the infecting pathogens and improve the effectiveness of clinical microbiological diagnosis. Sequencing data from a patient with endocarditis, including two negative samples (NEC1 and NEC2) and two positive samples (Sample1 and Sample2), were downloaded from the European Nucleotide Archive (ENA), project number PRJEB25228, to identify the presence of exogenous nucleic acid sequences from Endocardial Bacilli in the patient.

[0100] The analysis of the present invention is shown in Table 6, which shows the number of paired-end reads remaining after each analysis step S1-S9. In negative samples, all human host sequences were ultimately removed after analysis, resulting in a clean control result free of Endobacterium cardiacis. In positive samples, analysis using the present invention yielded 265,834 and 2,884 paired-end sequences of Endobacterium cardiacis for Sample 1 and Sample 2, respectively, representing 19.03% and 0.19% of the original sequences, respectively. The identified sequences for Sample 1 and Sample 2 were aligned to the Endobacterium cardiacis reference genome (NCBI RefSeqassembly: GCF_900637305.1), achieving alignment rates of 99.99% and 100%, respectively. We then calculated the genome coverage and average sequencing depth of the identified sequences for the two samples, finding 85.49% coverage and 49.84% average sequencing depth for Sample 1, and 30.09% and 0.54% average sequencing depth for Sample 2. We mapped the distribution of Endocardial Bacillus sequences on the genome based on the reads identified by sample1, as shown in the figure: Figure 3 The above results confirm that the present invention can effectively identify exogenous nucleic acids in the human body and assist in clinical diagnosis.

[0101] Table 6 Information analysis results of Example 3

[0102]

[0103]

[0104] The above embodiments are only preferred implementation modes of the present invention and are only used to explain the present invention rather than to limit the present invention. Any changes, substitutions, modifications, etc. made by those skilled in the art without departing from the spirit of the present invention should fall within the scope of protection of the present invention.

Claims

1. A method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology, characterized in that: The steps include: S1. Obtain whole genome sequencing data, targeted sequencing data, or metagenomic sequencing data of the host sample and perform routine quality control on the sequencing data; the routine quality control refers to removing residual adapter sequences and low-quality sequencing reads; the low-quality sequencing reads include sequencing reads with an N ratio greater than 5% and sequencing reads with a length of less than 30; S2. Remove host sequences from sequencing data: Obtain the latest host reference genome, filter out nucleic acid sequences in the sequencing data that are identical to the host reference genome nucleic acid sequences, and complete the first filtering; S3, filtering out the nucleic acid sequences identical to the host reference genome nucleic acid sequences in the sequencing data after the first filtration, completing the second filtration; S4. Specify the species to be analyzed, including parasites and microorganisms that do not originate from the host, obtain the reference genome sequence of the specified analysis species, and construct a species classification database; use metagenomic analysis software to align the sequencing data after the second filtration to the species classification database to obtain the accurate nucleic acid sequence of the specified analysis species; S5. Re-align the accurately classified nucleic acid sequence of the designated analysis species to the reference genome sequence of the designated analysis species, retain the paired-end sequences that are successfully aligned to the reference genome of the designated analysis species, filter out the results that cannot be aligned and the results of single-end alignment, and obtain the accurate species nucleic acid sequence; S6. removing low-complexity nucleic acid sequences from the accurate species nucleic acid sequences; S7. Using the BLAST alignment algorithm, align the accurate species nucleic acid sequence after removing the low-complexity nucleic acid sequence to the reference genome sequence of the designated analysis species, retaining the paired-end sequence alignment results with a confidence level greater than 90 and a coverage greater than 90% in the BLAST alignment to obtain a nucleic acid sequence with high confidence; S8. Align the highly reliable nucleic acid sequences to the NCBI NT database, retain the paired-end sequence alignment results with a reliability greater than 90 and a coverage greater than 90% in the BLAST comparison, further filter the nucleic acid sequences of distantly related species, and finally obtain the nucleic acid sequences that are truly from the specified analysis species; S9. If the nucleic acid sequence obtained in step S8 that is actually from the designated analysis species is identical to the accurate sequence of the designated analysis species, it indicates that the host sample contains the nucleic acid sequence of the designated analysis species; otherwise, it indicates that the host sample does not contain the nucleic acid sequence of the designated analysis species; If the reference genome sequence of the specified analysis species is not included in the NCBI NT library, no valid results will be generated. In this case, the NT library comparison will not be considered, that is, step S8 will not be executed. The highly reliable sequence obtained in step S7 will be directly used as the accurate sequence of the specified analysis species, and then step S9 will be executed. At this time, the interference of sequences of distantly related species will not be removed.

2. The method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology according to claim 1, characterized in that: The host sample includes a body fluid sample or a tissue sample.

3. The method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology according to claim 1, characterized in that: The nucleic acid sequence includes DNA sequence and RNA sequence.

4. The method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology according to claim 1, characterized in that: During the second filtering, a different alignment algorithm than that used in the first filtering needs to be used for nucleic acid sequence filtering.

5. The method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology according to claim 1, characterized in that: In step S4, if there is a species that may have homology with the sequence of the designated analysis species during the analysis process and affect the analysis results, the reference genome sequence of the species affecting the analysis is obtained, and the genomes of the designated analysis species and the species affecting the analysis are mixed to construct a species classification database.

6. The method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology according to claim 1, characterized in that: In step S6, the low-complexity nucleic acid sequence refers to a nucleic acid sequence in which the sequencing reads have an interval with more than two consecutive identical bases exceeding 5%.

7. The method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology according to claim 1, characterized in that: In step S6, different software is needed from that used in step S5 when aligning to the species reference genome to remove low-complexity nucleic acid sequences from the accurate species nucleic acid sequence.

8. The method for identifying host exogenous nucleic acid sequences based on high-throughput sequencing technology according to claim 1, characterized in that: In step S8, the distantly related species refers to species that are different at the family level when the paired-end sequences are annotated for species annotation using the NT library.

Citation Information

Patent Citations

  • Method and device for detecting microorganisms in blood

    CN105525033A

  • A Metagenomic Data Analysis Method and System

    CN108334750B

  • Method and device for microbiological analysis of host sample

    CN111009286A

  • Pathogenic microorganism analysis and identification system and application thereof

    CN111462821A

  • Metagenomics-based pathogen analysis method, analysis device, equipment and storage medium

    CN111951895A