Virus species identification method, identification system, equipment and medium
By constructing a reference database of viral genome sequences and comparing them using specific algorithms, the complexity and inaccuracy of viral species identification in the prior art are solved, and efficient and accurate viral species identification is achieved.
Patent Information
- Application Number
- CN202510386517.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art has problems such as insufficient resolution, complex sample preparation, insufficient antibody specificity, possible cross-reactions, lack of new or unknown virus sequence information, and genetic variation affecting detection accuracy in viral species identification, resulting in complex and inaccurate identification.
By obtaining the genomic sequences of non-redundant virus species and subspecies, the original reference database is constructed, and the second or third generation sequencing data are compared using the BWT algorithm and the BLAST algorithm, the bases of the homologous sequence are replaced, and the first and second reference databases are constructed to achieve accurate identification of viral species.
Accurate identification of viral species is achieved, and a variety of viral species can be identified, including new and unknown viruses, improving identification accuracy and speed, simplifying the process, and reducing the failure rate.
Smart Images

Figure CN119920328A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of species identification, and in particular to a virus species identification method, identification system, equipment and medium. Background Art
[0002] Accurate identification of viral species is crucial for disease prevention, control, and treatment. However, the large variety of viruses, rapid genetic variation, and high host specificity and environmental adaptability of many viruses make viral species identification a complex and arduous task.
[0003] Currently, virus species identification mainly relies on traditional microscopic imaging techniques, immune-serological detection and molecular biological detection methods. Although these methods are currently commonly used methods, their limitations are also very prominent.
[0004] Although microscopic imaging technology has the advantage of intuitively displaying the morphology of viruses and the infection process in virological research, its application is subject to multiple limitations. First, the morphological characteristics of viruses are extremely small and diverse, and traditional optical microscopes often have difficulty achieving sufficient resolution to clearly identify all types of viruses. Although electron microscopes can provide higher resolution, the sample preparation process is complicated and has high requirements for sample integrity, which can easily introduce artifacts or destroy the natural state of the virus during the preparation process. In addition, the dynamic process of viruses in cells is difficult to fully capture through static microscopic images, which limits the application of microscopic imaging technology in the study of viral life cycles.
[0005] Immuno-serological testing relies on the specific binding between antibodies and antigens, but this specificity is not absolute. The preparation process of antibodies may be affected by many factors, such as the purity of the antigen, the design of the immunization program, etc., resulting in insufficient antibody specificity or non-specific cross-reactions. This means that antibodies may not only bind to the target virus, but may also react with other microorganisms or host components in the sample, resulting in false positive results. In addition, for emerging or mutated viruses, immuno-serological testing may not be carried out in a timely and effective manner due to the lack of corresponding antibodies. The preparation and validation of antibodies also takes time, which is particularly disadvantageous in emerging epidemics.
[0006] Molecular biology detection methods, with their high sensitivity and high specificity, play an important role in virus detection. However, the application of this method also faces challenges. Molecular biology detection requires the design of primers based on known viral sequence information, which means that for new or unknown viruses, it may not be implemented or the effect may be poor due to the lack of sequence information. In addition, genetic variation of the virus may cause the match between the primer and the viral sequence to decrease, affecting the accuracy of the detection. Even for known viruses, if the viral sequence mutates significantly, the primers may need to be redesigned, which increases the detection time. In addition, molecular biology detection may also be affected by factors such as inhibitors, contamination or degradation in the sample, resulting in detection failure or inaccurate results.
[0007] Since microscopic imaging technology is limited by the morphological characteristics of the virus and sample preparation; immune-serological testing may be affected by antibody specificity and cross-reaction; molecular biological detection methods, such as PCR, although highly sensitive, usually require known virus sequence information as the basis for primer design, and have blind spots in the detection of new or unknown viruses, there is an urgent need to develop a virus species identification method, identification system, equipment and medium. Summary of the invention
[0008] In view of the above problems, the present invention provides a virus species identification method, identification system, device and medium.
[0009] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0010] In a first aspect, the present invention provides a method for identifying virus species, comprising:
[0011] Obtaining an original reference database, wherein the original reference database includes non-redundant genome sequences of virus species and representative functional fragment sequences of non-redundant virus subspecies;
[0012] All sequences in the original reference database are interrupted to obtain a number of first fragments, and each first fragment is compared with the original reference database using the BWT algorithm. When there is more than one sequence in the first fragment, all bases in the sequence segment in the original reference database are replaced with N to obtain a first reference database;
[0013] The original data from the second-generation or third-generation sequencing is aligned to the first reference database using the BWT algorithm to obtain unique aligned reads, and species abundance statistics are performed on the unique aligned reads, and the species with the highest aligned abundance is used as the identified species;
[0014] Sliding with a step length of a third length, all sequences in the original reference database are interrupted to obtain a plurality of second fragments, and each second fragment is compared with the original reference database using the BLAST algorithm. If the similarity of the second fragment and the sequence in the original reference database reaches a first threshold, and the length of the aligned sub-fragment is not less than the third length at more than 1 position, all bases of the aligned sequence segment in the original reference database are replaced with N to obtain a second reference database;
[0015] The spliced genome data were aligned to the second reference database using the BLAST algorithm, and the species to which the sequence with the highest bit-score value belonged was identified as the species.
[0016] In a preferred embodiment, the step of breaking all the sequences in the original reference database to obtain a plurality of first segments specifically comprises: sliding with a step length of the first length, breaking all the sequences in the original reference database according to a second length to obtain a plurality of first segments;
[0017] The step of interrupting all the sequences in the original reference database specifically includes: interrupting all the sequences in the original reference database according to a fourth length to obtain a plurality of second segments.
[0018] In a preferred embodiment, the first length is 10 bp, the second length is 50 bp, the third length is 50 bp, the fourth length is 500 bp, and the first threshold value ranges from [98%, 100%].
[0019] In a preferred embodiment, the identification method further comprises:
[0020] Obtain genetic data of the virus species to be identified;
[0021] Determine the data type of the genetic data of the virus species to be identified, wherein the data type includes the spliced genome data, the original data from the second-generation sequencing, and the original data from the third-generation sequencing.
[0022] In a second aspect, the present invention provides a virus species identification system, comprising:
[0023] An acquisition module, used to acquire an original reference database, wherein the original reference database includes non-redundant genome sequences of virus species and representative functional fragment sequences of non-redundant virus subspecies;
[0024] A first construction module is used to construct a first reference database, specifically for breaking all sequences in the original reference database to obtain a number of first fragments, using the BWT algorithm to compare each first fragment with the original reference database, and if there are more than one sequence in the first fragment, all bases of the sequence segment in the original reference database that is compared are replaced by N to obtain the first reference database;
[0025] The first identification module is used to use the BWT algorithm to align the original offline data of the second-generation or third-generation sequencing to the first reference database, obtain the reads on the unique alignment, perform species abundance statistics on the reads on the unique alignment, and use the species with the highest alignment abundance as the identified species;
[0026] A second construction module is used to construct a second reference database, specifically used to slide with a step length of a third length, break all sequences in the original reference database to obtain a plurality of second fragments, use a BLAST algorithm to compare each second fragment with the original reference database, if the sequence alignment similarity between the second fragment and the original reference database reaches a first threshold, and the length of the aligned sub-fragment is not less than the third length at more than 1 position, then all bases of the aligned sequence segment in the original reference database are replaced with N to obtain a second reference database;
[0027] The second identification module is used to compare the spliced genome data to the second reference database using the BLAST algorithm, and the species to which the sequence with the highest bit-score value belongs is used as the identified species.
[0028] In a third aspect, the present invention provides an electronic device comprising: a memory; one or more processors; one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods for identifying virus species according to the first aspect and its preferred embodiments.
[0029] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium comprises instructions, and when the instructions are executed on a computer, the computer executes a virus species identification method as described in any one of the first aspect and its preferred embodiments.
[0030] The present invention discloses a virus species identification method, identification system, equipment and medium. The original reference database is composed of non-redundant genome sequences of virus species and subspecies. The sequence that replaces the homologous sequence is designed as a reference database and a processing method is specifically designed. Then, the BWA or BLAST algorithm is used to compare and analyze the second-generation and third-generation original offline data or the spliced genome data, and the final species result is determined based on the comparison result. With such a design, the virus species or subtype level can be achieved, and accurate identification can also be achieved between strains with extremely high genome sequence similarity. When outputting the results, the strain information with the highest identification rate is directly given. The implementation principle is different from the prior art, and the species that can be identified exceed the prior art, and the identification accuracy is high. The species of new viruses or unknown viruses can also be detected. The present invention realizes simple identification, accurate results and short time consumption of virus identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0032] Figure 1 A schematic diagram of a method for identifying virus species;
[0033] Figure 2 A specific flow chart of a method for identifying a virus species;
[0034] Figure 3 A schematic diagram of a virus species identification system;
[0035] Figure 4 This is a structural framework diagram of a virus species identification system;
[0036] Figure 5 is a schematic diagram of an electronic device;
[0037] Among them, 10, acquisition module, 20, first construction module, 30, first identification module, 40, second construction module, 50, second identification module, 60, output module, 70, test data acquisition module, 80, type determination module. DETAILED DESCRIPTION
[0038] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0039] The existing commonly used methods for virus species identification are mainly microscopic imaging technology, immuno-serological testing and molecular biology detection methods, but these methods have very prominent limitations. Microscopic imaging technology is limited by the morphological characteristics of the virus and sample preparation; immuno-serological testing may be affected by antibody specificity and cross-reaction; molecular biology detection methods, such as PCR, although highly sensitive, usually require known virus sequence information as the basis for primer design, and have blind spots for the detection of new or unknown viruses.
[0040] To this end, the present invention provides a virus species identification method, identification system, equipment and medium. The present invention replaces the non-specific region by comparison, so as to solve the problem of low specificity of the comparison results based on the original offline data of short read length or the shorter genome splicing data. The present invention can achieve the virus species or subtype level, and can also realize accurate identification between different strains with extremely high genome sequence similarity, and directly give the strain information with the highest identification rate when outputting the results. The implementation principle is different from the prior art, and the species that can be identified exceed the prior art, and the identification accuracy is higher. For new viruses or unknown viruses, their species can also be detected, and the identification is simple. At the same time, the present invention can directly use the original sequencing data for identification, omitting the process of data splicing and assembly, and greatly shortening the analysis time.
[0041] The following embodiments of the present invention can be performed separately, and the embodiments can also be performed in combination with each other, and the embodiments of the present invention do not specifically limit this. In the embodiments of the present invention, "first", "second", etc. are used to describe various components, but these components should not be limited by these terms. These terms are only used to distinguish one component from another component. The "and / or" mentioned in the present invention refers to any and all combinations including one or more related listed items.
[0042] Below, the virus species identification method, identification system, equipment and medium and their technical effects are described.
[0043] Figure 1 A schematic diagram of a process for identifying a virus species provided in an embodiment, such as Figure 1 As shown, the method comprises the following steps:
[0044] Obtaining an original reference database, wherein the original reference database includes non-redundant genome sequences of virus species and representative functional fragment sequences of non-redundant virus subspecies;
[0045] All sequences in the original reference database are interrupted to obtain a number of first fragments, and each first fragment is compared with the original reference database using the BWT algorithm. If there are more than one sequence in the first fragment, all bases of the sequence segment in the original reference database are replaced with N to obtain a first reference database;
[0046] Using the BWT algorithm, the original data from the second-generation sequencing or the original data from the third-generation sequencing are aligned to the first reference database to obtain unique aligned reads, and the species abundance statistics are performed on the unique aligned reads, and the species with the highest aligned abundance is used as the identified species;
[0047] Sliding with a step length of a third length, all sequences in the original reference database are interrupted to obtain a plurality of second fragments, and each second fragment is compared with the original reference database using the BLAST algorithm. If the alignment similarity between the second fragment and a sequence in the original reference database exceeds (≥) a first threshold, and there is more than one position where the length of the aligned sub-fragment in the second fragment is not less than the third length, then all bases of the sequence segment in the original reference database whose alignment similarity is ≥ the first threshold are replaced with N to obtain a second reference database;
[0048] The spliced genome data were aligned to the second reference database using the BLAST algorithm, and the species to which the sequence with the highest bit-score value belonged was identified as the species.
[0049] It should be understood that the above method does not represent or imply that all steps must be executed in this order. A person skilled in the art can transform or change the execution order of the above steps based on the present invention. Some embodiments of the above method are illustrated below.
[0050] In one specific embodiment, see Figure 2 , the virus species identification method comprises the following steps:
[0051] S1. Obtain the original reference database;
[0052] We use the genome sequence of the virus species and the representative functional fragment sequence of the virus subspecies as the basic data for constructing the reference database, where the virus species is non-redundant and the virus subspecies is also non-redundant. Here, the genome sequence of the virus species is usually the complete gene sequence of the virus, while the sequence of the virus subspecies is usually the incomplete gene sequence of the virus. The representative functional fragment sequence of the virus subspecies is not limited in number. The virus subspecies has several functional fragment sequences, where the number of representative functional fragment sequences is specifically determined according to the virus subspecies. The representative functional fragment sequence is, for example, a fragment sequence that can characterize genetic characteristics and a fragment sequence that can characterize pathogenicity characteristics.
[0053] S2. constructing a first reference database based on the original reference database;
[0054] All sequences in the original reference database are broken into several first fragments according to the second length (preferably 50 bp) by sliding with a step length of the first length (preferably 10 bp), and a fastq file is generated by simulation. It can be understood that not all first fragments have the same length as the second length, and the last sequence segment of the sequence may be shorter than the second length after being broken, and the following second fragments can be understood in the same way.
[0055] The first fragment of each fastq file is aligned with the original reference database using the Burrows-Wheeler Transform (BWT) algorithm. For each first fragment, if the sequence alignment result between the first fragment and the original reference database is that there is more than one aligned sequence, all the bases of the aligned sequence segment in the original reference database are replaced with N. Through alignment and replacement, the original reference database after the replacement is used as the first reference database to obtain the first reference database, that is, after all the first fragments are aligned, all the bases of the aligned sequence segment in the original reference database are replaced with N when the alignment result is that there is more than one aligned sequence, and the replaced original reference database is used as the first reference database.
[0056] S3, constructing a second reference database based on the original reference database;
[0057] Slide with a step length of the third length (preferably 50 bp), break all sequences in the original reference database according to the fourth length (preferably 500 bp) to obtain several second fragments, and use the BLAST algorithm to compare each second fragment with the original reference database. If the similarity between the second fragment and the sequence in the original reference database is greater than or equal to 98% (that is, the similarity between the second fragment and a sequence segment in the original reference database is greater than or equal to 98%), and the length of the sub-fragment on the comparison (the sub-fragment can be understood as a partial fragment in the second fragment) is not less than the third length at more than 1 position, then replace all the bases of the sequence segment on the comparison in the original reference database with N to obtain a second reference database. That is, after all the fourth fragments are compared, the original reference database after the replacement operation of replacing all the bases of the sequence segments with a comparison similarity of ≥98% with N is used as the second reference database. In other embodiments, the first threshold value can also be other values, usually a value in the range of [98%, 100%]. It can be understood that the execution order between S2 and S3 is not limited.
[0058] S4. Obtaining genetic data of the virus species to be identified;
[0059] S5. Determine the data type of the genetic data of the virus species to be identified, where the data type includes the spliced genome data, the original data from the second-generation sequencing, and the original data from the third-generation sequencing.
[0060] If the data type is the original data from the second generation sequencing or the original data from the third generation sequencing, proceed to S6; if the data type is the spliced genome data, proceed to S7.
[0061] S6, using the BWT algorithm to align the original data (FASTA format) of the second-generation or third-generation sequencing to the first reference database, to obtain the reads on the unique alignment, to perform species abundance statistics on the reads on the unique alignment, and to use the species with the highest alignment abundance as the identified species, and to proceed to S8;
[0062] S7. Use the BLAST algorithm to align the assembled genome data (FASTQ format) to the second reference database, sort the BLAST results in reverse order according to the bit-score value, and use the species to which the sequence with the highest bit-score value belongs as the identified species for S8.
[0063] S8. Output the identified species.
[0064] like Figure 3 As shown, the present invention also provides a virus species identification system, comprising:
[0065] An acquisition module 10 is used to acquire an original reference database including genome sequences of non-redundant virus species and representative functional fragment sequences of non-redundant virus subspecies;
[0066] The first construction module 20 is used to construct a first reference database, specifically, to break all sequences in the original reference database into a plurality of first fragments, and to compare each first fragment with the original reference database using the BWT algorithm. If there are more than one sequence alignments between the first fragment and the sequence in the original reference database, all bases of the aligned sequence segments in the original reference database are replaced with N to obtain the first reference database;
[0067] The first identification module 30 is used to use the BWT algorithm to align the original offline data of the second-generation or third-generation sequencing to the first reference database, obtain the reads on the unique alignment, perform species abundance statistics on the reads on the unique alignment, and use the species with the highest alignment abundance as the identified species;
[0068] A second construction module 40 is used to construct a second reference database, specifically to slide with a step length of a third length, break all sequences in the original reference database to obtain a plurality of second fragments, and use a BLAST algorithm to compare each second fragment with the original reference database. If the sequence alignment similarity between the second fragment and the original reference database reaches a first threshold, and the length of the aligned sub-fragment is not less than the third length at more than one position, then all bases of the aligned sequence segment in the original reference database are replaced with N to obtain a second reference database;
[0069] The second identification module 50 is used to compare the spliced genome data to the second reference database using the BLAST algorithm, and the species to which the sequence with the highest bit-score value belongs is used as the identified species.
[0070] The method of breaking all the sequences in the original reference database into a plurality of first fragments specifically includes: sliding with a step length of the first length, breaking all the sequences in the original reference database into a second length to obtain a plurality of first fragments; the method of breaking all the sequences in the original reference database specifically includes: breaking all the sequences in the original reference database into a fourth length to obtain a plurality of second fragments.
[0071] The first length is 10 bp, the second length is 50 bp, the third length is 50 bp, the fourth length is 500 bp, and the first threshold value ranges from [98%, 100%].
[0072] See also Figure 4 , the identification system further comprises:
[0073] The test data acquisition module 70 is used to obtain the gene data of the virus species to be identified;
[0074] The type determination module 80 is used to determine the data type of the genetic data of the virus species to be identified, and the data type includes the spliced genome data, the original data from the second-generation sequencing, and the original data from the third-generation sequencing.
[0075] Here, the identification system also includes:
[0076] The output module 60 is used to output the identification results of the first identification module and the second identification module.
[0077] See also Figure 5 The present invention also provides an electronic device, comprising: a memory; one or more processors; one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any method according to a method for identifying a virus species.
[0078] The present invention also provides a computer-readable storage medium, which includes instructions. When the instructions are executed on a computer, the computer executes each step of a virus species identification method described in any of the above embodiments.
[0079] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0080] The present invention is based on the original reference database composed of non-redundant viral species and subspecies genome sequences, designs a sequence with masked homologous sequences as a reference database, and designs a method for comparing homologous sequences according to specific circumstances, uses BWA or BLAST algorithms to compare and analyze the second-generation and third-generation original offline data or spliced genome data, and determines the final species result based on the comparison results. In a specific embodiment, species identification of 15,327 viruses can be achieved.
[0081] Based on the above-designed method and system, it is possible to achieve virus species or subtype level, and accurate identification can be achieved between different strains with highly similar genomes. When outputting the results, the strain information with the highest identification rate is directly given.
[0082] Compared with microscopic imaging technology, the present invention has a different principle, and is therefore not limited by the morphological characteristics of the virus and sample preparation, does not have the problem of completely capturing the image, and has a high identification accuracy; compared with the immune-serological detection technology, the principle is different, the identification accuracy is higher and the identification is faster; compared with the molecular biology detection method, the detection principle is different, the present invention is easy to implement, has a higher detection accuracy, an extremely low failure rate, and a stable and fast identification speed.
[0083] The present invention not only encompasses species that can be detected by microscopic imaging technology, immune-serological detection technology, and molecular biological detection methods, and has high identification accuracy, but can also adapt to the identification of virus species that cannot be identified by microscopic imaging, immune-serological detection, and molecular biological methods, and can also detect the species of new viruses or unknown viruses. The present invention directly uses the original sequencing data for identification, which achieves simple identification, accurate results, and short time consumption for virus identification, while solving the problem of specificity of comparison results of short-read original offline data or shorter genome splicing data. The present invention is the first in the field to develop a standardized analysis process that can directly use reads data comparison strategies and is used for virus species identification. It can identify many species, can be accurate to the subspecies level, and has a high and fast identification rate.
[0084] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. The present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0085] The present invention is described with reference to flowcharts and / or block diagrams of methods, systems, and electronic devices according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0086] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0088] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0089] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for identifying virus species, characterized in that: include: Obtaining an original reference database, wherein the original reference database includes non-redundant genome sequences of virus species and representative functional fragment sequences of non-redundant virus subspecies; All sequences in the original reference database are interrupted to obtain a number of first fragments, and each first fragment is compared with the original reference database using the BWT algorithm. When there is more than one sequence in the first fragment, all bases in the sequence segment in the original reference database are replaced with N to obtain a first reference database; The original data from the second-generation or third-generation sequencing is aligned to the first reference database using the BWT algorithm to obtain unique aligned reads, and species abundance statistics are performed on the unique aligned reads, and the species with the highest aligned abundance is used as the identified species; Sliding with a step length of a third length, all genome sequences in the original reference database are interrupted to obtain a plurality of second fragments, and each second fragment is compared with the original reference database using a BLAST algorithm, if the similarity of the genome sequence comparison between the second fragment and the original reference database reaches a first threshold, and the length of the compared sub-fragment is not less than the third length at more than one position, then all bases of the compared sequence segment in the original reference database are replaced with N to obtain a second reference database; The spliced genome data were aligned to the second reference database using the BLAST algorithm, and the species to which the genome sequence with the highest bit-score value belonged was identified as the species.
2. A method for identifying virus species according to claim 1, characterized in that: The step of breaking all the sequences in the original reference database to obtain a plurality of first fragments specifically comprises: sliding with a step length of the first length, breaking all the genome sequences in the original reference database according to a second length to obtain a plurality of first fragments; The step of breaking all genome sequences in the original reference database specifically includes breaking all sequences in the original reference database according to a fourth length to obtain a plurality of second fragments.
3. A method for identifying virus species according to claim 2, characterized in that: The first length is 10 bp, the second length is 50 bp, the third length is 50 bp, the fourth length is 500 bp, and the first threshold value ranges from [98%, 100%].
4. A method for identifying virus species according to claim 1, characterized in that: The identification method further comprises: Obtain genetic data of the virus species to be identified; Determine the data type of the genetic data of the virus species to be identified, wherein the data type includes the spliced genome data, the original data from the second-generation sequencing, and the original data from the third-generation sequencing.
5. A virus species identification system, characterized in that: include: An acquisition module, used to acquire an original reference database, wherein the original reference database includes non-redundant genome sequences of virus species and representative functional fragment sequences of non-redundant virus subspecies; A first construction module is used to construct a first reference database, specifically for breaking all sequences in the original reference database to obtain a number of first fragments, using the BWT algorithm to compare each first fragment with the original reference database, and if there are more than one sequence in the first fragment, all bases of the sequence segment in the original reference database that is compared are replaced by N to obtain the first reference database; The first identification module is used to use the BWT algorithm to align the original offline data of the second-generation or third-generation sequencing to the first reference database, obtain the reads on the unique alignment, perform species abundance statistics on the reads on the unique alignment, and use the species with the highest alignment abundance as the identified species; A second construction module is used to construct a second reference database, specifically used to slide with a step length of a third length, break all sequences in the original reference database to obtain a plurality of second fragments, use a BLAST algorithm to compare each second fragment with the original reference database, if the sequence alignment similarity between the second fragment and the original reference database reaches a first threshold, and the length of the aligned sub-fragment is not less than the third length at more than 1 position, then all bases of the aligned sequence segment in the original reference database are replaced with N to obtain a second reference database; The second identification module is used to compare the spliced genome data to the second reference database using the BLAST algorithm, and the species to which the sequence with the highest bit-score value belongs is used as the identified species.
6. A virus species identification system as claimed in claim 5, characterized in that: The method of breaking all the sequences in the original reference database into a plurality of first fragments specifically includes: sliding with a step length of the first length, breaking all the sequences in the original reference database into a second length to obtain a plurality of first fragments; the method of breaking all the sequences in the original reference database specifically includes: breaking all the sequences in the original reference database into a fourth length to obtain a plurality of second fragments.
7. A virus species identification system as claimed in claim 6, characterized in that: The first length is 10 bp, the second length is 50 bp, the third length is 50 bp, the fourth length is 500 bp, and the first threshold value ranges from [98%, 100%].
8. A virus species identification system as claimed in claim 5, characterized in that: The identification system also includes: A test data acquisition module is used to obtain the genetic data of the virus species to be identified; The type determination module is used to determine the data type of the genetic data of the virus species to be identified, and the data type includes the spliced genome data, the original data from the second-generation sequencing, and the original data from the third-generation sequencing.
9. An electronic device, characterized in that: include: Memory; one or more processors; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods for identifying a virus species according to claims 1 to 4.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes instructions, and when the instructions are executed on a computer, the computer is caused to perform a virus species identification method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Virus infection detection and identification method based on metagenomics
CN105112569A
Method and device for detecting homologous sequences on basis of high-throughput sequencing
CN112513292A
Method and device for obtaining microbial species and related information through sequencing, computer readable storage medium and electronic equipment
CN114067911A
Methods for identification of organisms, assigning reads to organisms, and identification of genes in metagenomic sequences
US20150032711A1
Pathogen detection using next generation sequencing
US20180203976A1