Nucleic acid sequence database construction method, device, equipment and readable storage medium
By acquiring and deduplicating standard genome sequences, generating standard sequence indexes and alignment files, and constructing a nucleic acid sequence database, the contradiction between data coverage and retrieval efficiency in existing technologies is resolved, thereby improving the efficiency of high-throughput sequencing analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SANSURE BIOTECH INC
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-21
AI Technical Summary
Existing nucleic acid sequence databases have difficulty ensuring both comprehensive data coverage and efficient data retrieval during their construction, resulting in long data analysis times for high-throughput sequencing technologies.
By obtaining multiple standard genome sequences of the target pathogen, deduplication is performed based on classification and retrieval information to generate standard sequence index files and alignment files, and a nucleic acid sequence database is constructed to ensure comprehensive coverage of the target nucleic acid features and reduce the amount of data.
This approach improves data retrieval efficiency and shortens data analysis time for high-throughput sequencing technology without affecting data coverage.
Smart Images

Figure CN121350002B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for constructing a nucleic acid sequence database. Background Technology
[0002] With the continuous development of science and technology, high-throughput sequencing technology has become a core technical means for pathogen identification, such as targeted sequencing and metagenomic sequencing. Regardless of the type of high-throughput sequencing technology, the use of nucleic acid sequence databases is inevitable in the process of data analysis, pathogen feature comparison and accurate identification. Therefore, how to construct a high-quality, high-efficiency and comprehensive nucleic acid sequence database has become an urgent problem to be solved in this field.
[0003] Currently, in constructing nucleic acid sequence databases, the primary consideration is coverage. That is, the full and complete reference genome sequence is usually integrated into the nucleic acid sequence database to encompass the biological information of different pathogens as much as possible. However, due to the diversity of pathogen species, the nucleic acid sequence database is very large, which makes it easy to have long analysis time when performing data analysis based on high-throughput sequencing technology. Therefore, the current nucleic acid sequence databases are difficult to ensure both the comprehensiveness of data coverage and the efficiency of data retrieval. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for constructing a nucleic acid sequence database that ensures comprehensive data coverage while also taking into account the efficiency of data retrieval, in order to address the aforementioned technical problems.
[0005] Firstly, this application provides a method for constructing a nucleic acid sequence database, including:
[0006] Obtain multiple standard genome sequences of the target pathogen, wherein the standard genome sequences characterize the target nucleic acid features of the target pathogen;
[0007] Based on the classification and retrieval information of all standard genome sequences, duplicates of all standard genome sequences are removed to obtain multiple non-redundant genome sequences with different target nucleic acid features;
[0008] Generate a standard sequence index file containing classification retrieval information for each of the non-redundant genome sequences, and a standard sequence alignment file containing the targeted nucleic acid features of all the non-redundant genome sequences;
[0009] A nucleic acid sequence database of the target pathogen is constructed based on the standard sequence index file and the standard sequence alignment file.
[0010] In one embodiment, the classification retrieval information includes classification information and index information; the step of removing duplicates from all standard genome sequences based on the classification retrieval information of all standard genome sequences to obtain multiple non-redundant genome sequences with different target nucleic acid features includes:
[0011] Based on the classification information, candidate non-redundant genomic sequences for each taxonomic group of the target pathogen are screened from all the standard genomic sequences;
[0012] Based on the index information, all representative genome sequences are aligned and deduplicated to obtain non-redundant genome sequences with different target nucleic acid features.
[0013] In one embodiment, the step of screening candidate non-redundant genomic sequences for each taxonomic group of the target pathogen from all standard genomic sequences based on the classification information includes:
[0014] Based on the classification information, representative genome sequences for each taxonomic group of the target pathogen are selected from all the standard genome sequences;
[0015] Construct a standard sequence index table of the target pathogen based on all representative genome sequences;
[0016] Each representative genome sequence is segmented to obtain multiple genome sequence fragments to be aligned, which are adapted to the targeted sequencing read length and cover the targeted nucleic acid features of each representative genome sequence.
[0017] Based on the correspondence between each genome sequence fragment to be aligned and the standard sequence index table, redundant genome sequence fragments that cross taxonomic groups are removed from all genome sequence fragments to be aligned, thereby obtaining representative genome sequence fragments for each taxonomic group of the target pathogen.
[0018] All representative genomic sequence fragments were integrated into the candidate non-redundant genomic sequence.
[0019] In one embodiment, the target pathogen includes a viral pathogen; the step of removing redundant genomic sequence fragments across taxa from all genomic sequence fragments to be aligned, based on the correspondence between each genomic sequence fragment to be aligned and the standard sequence index table, to obtain a representative genomic sequence fragment for each taxa of the target pathogen, includes:
[0020] Based on the type of the viral pathogen, update the classification information of the viral pathogen in the standard sequence index table to obtain a target standard sequence index table adapted to the viral classification characteristics;
[0021] Multiple viral genome sequence fragments of the viral pathogen were identified from all the genome sequence fragments to be aligned;
[0022] Based on the correspondence between all viral genome sequence fragments to be compared and the target standard sequence index table, redundant genome sequence fragments that cross viral taxonomic groups are recursively removed from all viral genome sequence fragments to be compared in reverse order of the taxonomic lineage of the viral pathogen, so as to obtain the representative genome sequence fragment of each taxonomic group of the viral pathogen.
[0023] The process of integrating all representative genomic sequence fragments into the candidate non-redundant genomic sequence includes:
[0024] Based on the target nucleic acid feature distribution information of the viral pathogen, all representative genome sequence fragments are specifically spliced to obtain processed representative genome sequence fragments;
[0025] If all processed representative genomic sequence fragments pass quality verification, all processed representative genomic sequence fragments are integrated into a candidate non-redundant genomic sequence for the viral pathogen, wherein the target nucleic acid feature distribution information includes one or more of the target nucleic acid feature distribution location and the target nucleic acid feature variation interval.
[0026] In one embodiment, the step of aligning and deduplicating all representative genomic sequences according to the index information to obtain the plurality of non-redundant genomic sequences with different target nucleic acid features includes:
[0027] Based on the index information, all representative genome sequences are compared to multiple identical representative genome sequences to be deduplicated;
[0028] Identify target deduplication representative genome sequences that meet preset deduplication criteria from all representative genome sequences to be deduplicated, wherein the preset deduplication criteria are generated based on attribute difference information of targeted nucleic acid features;
[0029] The target deduplicated representative genome sequence is removed from all the representative genome sequences to obtain the non-redundant genome sequence.
[0030] In one embodiment, obtaining multiple standard genomic sequences of the target pathogen includes:
[0031] Based on the taxonomic information of the target pathogen, the corresponding initial genome sequence is searched from the metagenomic database;
[0032] If a reference genome sequence is detected in the initial genome sequence, the reference genome sequence is extracted from the reference database as the first candidate genome sequence.
[0033] If a reference genome sequence is detected as not existing in the initial genome sequence, a representative genome sequence is extracted from the initial nucleic acid sequence database as a second candidate genome sequence.
[0034] Based on the taxonomic lineage of the target pathogen, the first candidate genome sequence and the second candidate genome sequence are integrated into the plurality of standard genome sequences.
[0035] In one embodiment, integrating the first candidate genome sequence and the second candidate genome sequence into the plurality of standard genome sequences according to the taxonomic lineage of the target pathogen includes:
[0036] Based on the taxonomic lineage of the target pathogen, the first candidate genome sequence and the second candidate genome sequence are respectively located to their corresponding taxa;
[0037] According to the sequence redundancy rules of the taxonomic group, redundant genomic sequence fragments are removed from the first candidate genomic sequence and the second candidate genomic sequence to obtain multiple initial non-redundant genomic sequence fragments;
[0038] All the initial non-redundant genomic sequence fragments are converted into standard non-redundant genomic sequence fragments in a preset standard format;
[0039] Based on the attribute differences of the target nucleic acid features, multiple standard non-redundant genomic sequence fragments are integrated into the multiple standard genomic sequences.
[0040] Secondly, this application also provides a nucleic acid sequence database construction apparatus, comprising:
[0041] An acquisition module is used to acquire multiple standard genome sequences of a target pathogen, wherein the standard genome sequences characterize the target nucleic acid features of the target pathogen;
[0042] The deduplication module is used to remove duplicates from all standard genome sequences based on the classification retrieval information of all standard genome sequences, thereby obtaining multiple non-redundant genome sequences with different target nucleic acid features;
[0043] The generation module is used to generate a standard sequence index file containing classification retrieval information for each of the non-redundant genome sequences, and a standard sequence alignment file containing the target nucleic acid features of all the non-redundant genome sequences;
[0044] A construction module is used to construct a nucleic acid sequence database of the target pathogen based on the standard sequence index file and the standard sequence alignment file.
[0045] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0046] Multiple standard genome sequences of a target pathogen are obtained, wherein the standard genome sequences characterize the target nucleic acid features of the target pathogen; based on the classification retrieval information of all standard genome sequences, duplicates are removed from all standard genome sequences to obtain multiple non-redundant genome sequences with different target nucleic acid features; a standard sequence index file containing the classification retrieval information of each of the non-redundant genome sequences and a standard sequence alignment file containing the target nucleic acid features of all the non-redundant genome sequences are generated; based on the standard sequence index file and the standard sequence alignment file, a nucleic acid sequence database of the target pathogen is constructed.
[0047] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0048] Multiple standard genome sequences of a target pathogen are obtained, wherein the standard genome sequences characterize the target nucleic acid features of the target pathogen; based on the classification retrieval information of all standard genome sequences, duplicates are removed from all standard genome sequences to obtain multiple non-redundant genome sequences with different target nucleic acid features; a standard sequence index file containing the classification retrieval information of each of the non-redundant genome sequences and a standard sequence alignment file containing the target nucleic acid features of all the non-redundant genome sequences are generated; based on the standard sequence index file and the standard sequence alignment file, a nucleic acid sequence database of the target pathogen is constructed.
[0049] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0050] Multiple standard genome sequences of a target pathogen are obtained, wherein the standard genome sequences characterize the target nucleic acid features of the target pathogen; based on the classification retrieval information of all standard genome sequences, duplicates are removed from all standard genome sequences to obtain multiple non-redundant genome sequences with different target nucleic acid features; a standard sequence index file containing the classification retrieval information of each of the non-redundant genome sequences and a standard sequence alignment file containing the target nucleic acid features of all the non-redundant genome sequences are generated; based on the standard sequence index file and the standard sequence alignment file, a nucleic acid sequence database of the target pathogen is constructed.
[0051] The aforementioned method, apparatus, computer equipment, computer-readable storage medium, and computer program product for constructing a nucleic acid sequence database first obtain multiple standard genomic sequences characterizing the target nucleic acid features of the target pathogen, thus providing a data foundation for accurately covering the target nucleic acid features of the database to be constructed. Then, based on the classification and retrieval information of all standard genomic sequences, duplicates are removed from all standard genomic sequences to obtain multiple non-redundant genomic sequences with distinct target nucleic acid features, achieving the goal of redundancy removal without affecting the comprehensive coverage of the target nucleic acid features. Next, a standard sequence index file containing the classification and retrieval information of each non-redundant genomic sequence and a standard sequence alignment file containing the target nucleic acid features of all non-redundant genomic sequences are generated. Finally, based on the standard sequence index file and the standard sequence alignment file, a nucleic acid sequence database of the target pathogen is constructed. Due to the large number of nucleic acid sequences... The database construction process consistently uses the target nucleic acid characteristics of the pathogen as a benchmark. During the deduplication process, only standard genomic sequences with identical target nucleic acid characteristics are deduplicated. This allows for the reduction of the database size while preserving all target nucleic acid characteristics by generating standard sequence alignment files. Furthermore, since the nucleic acid sequence database is also built based on standard sequence index files, it provides a clear basis for high-throughput sequencing analysis, rather than directly relying on a full and complete reference genome sequence. Therefore, it overcomes the technical drawback of the large size of nucleic acid sequence databases due to the diversity of pathogen species, which can lead to long analysis times when using high-throughput sequencing technology. Thus, the constructed nucleic acid sequence database ensures comprehensive data coverage while maintaining efficient data retrieval. Attached Figure Description
[0052] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 This is a flowchart illustrating a method for constructing a nucleic acid sequence database in one embodiment;
[0054] Figure 2 A flowchart illustrating the method for constructing a nucleic acid sequence database in another embodiment;
[0055] Figure 3 This is a schematic diagram of the overall process of constructing a nucleic acid sequence database in another embodiment;
[0056] Figure 4 This is a structural block diagram of a nucleic acid sequence database construction apparatus in one embodiment;
[0057] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0059] Firstly, in the fields of bioinformatics and pathogen detection, high-throughput sequencing technology has become a core means of pathogen identification. Specifically, targeted sequencing relies on the design of target sequences and amplification primer sequences, requiring verification of sequence conservation and specificity from nucleic acid sequence databases. Similarly, metagenomic sequencing also requires database alignment for result confirmation. However, existing nucleic acid sequence databases all have certain limitations. For example, the coreNT library is extremely large and cannot perform efficient FM-indexing, relying solely on the slow BLAST method for alignment, resulting in long loading and alignment times. Furthermore, the coreNT library lacks the complete genomes of some species, containing only basic genomes. Because of the limited number of genome sequence fragments, the comprehensiveness of detection is restricted. Common methods employed include reference sequence-based compression, hierarchical classification coding, random access compression, and k-mer indexing and alignment optimization tools. On the other hand, other nucleic acid sequence databases prioritize coverage, typically integrating the full and complete reference genome sequence to encompass the biological information of different pathogens as comprehensively as possible. In summary, existing nucleic acid sequence databases often experience lengthy analysis times when using high-throughput sequencing technology. Therefore, there is an urgent need for a method to construct a nucleic acid sequence database that can ensure comprehensive data coverage while maintaining efficient data retrieval.
[0060] In one embodiment, such as Figure 1As shown, a method for constructing a nucleic acid sequence database is provided. This embodiment uses the application of this method to a terminal as an example. The terminal includes, but is not limited to, personal computers and laptops. The terminal is equipped with a nucleic acid sequence database construction device, which includes an acquisition module, a deduplication module, a generation module, and a construction module. The acquisition module is used to acquire multiple standard genomic sequences of the target pathogen, wherein the standard genomic sequences characterize the target nucleic acid features of the target pathogen. The deduplication module is used to remove duplicates from all standard genomic sequences according to the classification retrieval information of all standard genomic sequences, obtaining multiple non-redundant genomic sequences with different target nucleic acid features. The generation module is used to generate a standard sequence index file containing the classification retrieval information of each non-redundant genomic sequence, as well as a standard sequence index file containing all non-redundant genomic sequences. A standard sequence alignment file for the target nucleic acid features of redundant genome sequences; a construction module is used to construct a nucleic acid sequence database of the target pathogen based on the standard sequence index file and the standard sequence alignment file; through information interaction between the acquisition module, deduplication module, generation module and construction module, it is possible to reduce the amount of data in the nucleic acid sequence database while retaining all target nucleic acid features by generating a standard sequence alignment file. On the other hand, since the nucleic acid sequence database is also constructed based on the standard sequence index file, it can provide a clear basis for the high-throughput sequencing analysis process, rather than directly relying on the full and complete reference genome sequence to construct the nucleic acid sequence database. Therefore, the nucleic acid sequence database constructed based on the above can ensure the comprehensiveness of data coverage while taking into account the efficiency of data retrieval. It is understood that the nucleic acid sequence database construction system can also be deployed on a server, or on a system including terminals and servers, and implemented through the interaction between terminals and servers. In this embodiment, the method includes the following steps 102 to 108:
[0061] Step 102: Obtain multiple standard genome sequences of the target pathogen, wherein the standard genome sequences characterize the target nucleic acid features of the target pathogen.
[0062] It should be noted that the core object for constructing the target pathogen characterization nucleic acid sequence database can be one or more, specifically bacteria, viruses, or fungi. Among them, bacteria can specifically be Escherichia coli and Mycobacterium tuberculosis, viruses can be influenza A virus and influenza B virus, and fungi can be Aspergillus and Candida albicans. The standard genome sequence characterizes the target nucleic acid features of the target pathogen. The target nucleic acid features are used to distinguish pathogen-specific nucleic acid fragments. It can be understood that the target nucleic acid features can be used as a basis for pathogen identification. The target nucleic acid features can specifically include subtype / strain differential variation sites, functionally related nucleic acid features, and species-specific conserved sequences.
[0063] As an example, step 102 includes: querying all initial genome sequences of the target pathogen from a public database and using all initial genome sequences as standard genome sequences.
[0064] In one embodiment, obtaining multiple standard genome sequences of the target pathogen includes:
[0065] Based on the taxonomic information of the target pathogen, the corresponding initial genome sequence is searched in the metagenomic database. If a reference genome sequence is detected in the initial genome sequence, the reference genome sequence is extracted from the reference database as the first candidate genome sequence. If no reference genome sequence is detected in the initial genome sequence, a representative genome sequence is extracted from the initial nucleic acid sequence database as the second candidate genome sequence. Based on the taxonomic lineage of the target pathogen, the first candidate genome sequence and the second candidate genome sequence are integrated into multiple standard genome sequences.
[0066] It should be noted that, in order to avoid redundancy in standard genome sequences and ensure comprehensive coverage of target nucleic acid features, since the sources of genome sequences characterizing target nucleic acid features are diverse, candidate genome sequences covering core target features can be extracted through differentiated extraction methods during the process of obtaining genome sequences through different channels. Sequences matching the target taxonomic information are accurately filtered from metagenomic databases to ensure that the extracted candidate genome sequences can accurately reflect the target nucleic acid features. Redundancy is determined based on taxonomic lineage relationships to remove duplicate sequence fragments under the same taxonomic group.
[0067] It should be noted that taxonomic information represents the core information defining the classification of the target pathogen, specifically the Taxid. Taxonomic information can display taxonomic lineage relationships, which can be specifically Kingdom-Phylum-Class-Order-Family-Genus-Species-Subtype. Metagenomic databases represent public or private databases storing massive amounts of metagenomic sequencing data. Initial genome sequences represent the genome sequences of the target pathogen selected from metagenomic databases. Reference databases refer to databases composed of authoritatively verified, fully annotated, and standardized reference genome sequences. Initial nucleic acid sequence databases refer to databases containing genome sequences that have not undergone rigorous annotation but have broad coverage.
[0068] As an example, the initial genome sequence is found in the metagenomic database using the taxonomic information of the target pathogen as an index. If a reference genome sequence is detected in the initial genome sequence, the reference genome sequence is extracted from the reference database as the first candidate genome sequence. If a reference genome sequence is not detected in the initial genome sequence, a representative genome sequence is extracted from the initial nucleic acid sequence database as the second candidate genome sequence. Based on the taxonomic lineage of the target pathogen, the first candidate genome sequence and the second candidate genome sequence are integrated into multiple standard genome sequences.
[0069] In one feasible approach, during the standard genome sequence acquisition stage, the taxonomic lineage of pathogen species is first established based on the taxid node file, and a list of pathogen taxids from the metagenomic database is obtained. Further, for pathogen species with a reference genome sequence, the reference sequence is directly downloaded from NCBI (reference database), decompressed, and renamed. For pathogen species without a reference genome sequence, the longest genome sequence is extracted from the core_nt library (initial nucleic acid sequence database) as a representative genome sequence, thus obtaining the first candidate genome sequence and the second candidate genome sequence. Finally, based on the taxonomic lineage of the target pathogen, the first candidate genome sequence and the second candidate genome sequence are integrated into the standard genome sequence.
[0070] In one embodiment, based on the taxonomic lineage of the target pathogen, the first candidate genome sequence and the second candidate genome sequence are integrated into multiple standard genome sequences, including:
[0071] Based on the taxonomic lineage of the target pathogen, the first and second candidate genome sequences are located to their respective taxa. According to the sequence redundancy rules of the taxa, redundant genome sequence fragments are removed from the first and second candidate genome sequences to obtain multiple initial non-redundant genome sequence fragments. All multiple initial non-redundant genome sequence fragments are converted into standard non-redundant genome sequence fragments in a preset standard format. According to the attribute difference information of the target nucleic acid features, the multiple standard non-redundant genome sequence fragments are integrated into multiple standard genome sequences.
[0072] It should be noted that in the event of a conflict between the first and second candidate genome sequences, the first candidate genome sequence is prioritized and retained. That is, the first candidate genome sequence is larger than the second candidate genome sequence. This is because the first candidate genome sequence is derived from a reference database, which can further improve the accuracy of the standard genome sequence. Converting multiple initial non-redundant genome sequence fragments into genome sequences in a preset standard format can improve the efficiency of subsequent redundancy removal processing of candidate genome sequences. Attribute difference information characterizes the essential differences and functional correlations of the initial non-redundant genome sequence fragments in terms of target nucleic acid features, which may specifically include differences in feature type, taxonomic characteristics, functional correlation, and sequence structure.
[0073] As an example, using the taxonomic lineage of the target pathogen as an index, the first candidate genome sequence is located to its corresponding taxonomic group, and the second candidate genome sequence is located to its corresponding taxonomic group; according to the sequence redundancy rules of the taxonomic groups, redundant genome sequence fragments are removed from the first and second candidate genome sequences to obtain multiple initial non-redundant genome sequence fragments; all multiple initial non-redundant genome sequence fragments are converted into standard non-redundant genome sequence fragments in a preset standard format; according to the attribute difference information of the target nucleic acid features, the multiple standard non-redundant genome sequence fragments are integrated into multiple standard genome sequences.
[0074] In one feasible approach, a list of pathogenic microorganism taxonomic units is first obtained. Pathogen classification is then analyzed based on a taxonomic database, and taxonomic units containing major pathogen types are selected to ensure coverage of bacteria, viruses, fungi, and parasites. Specifically, the `taxid_nodes` and `taxid_name` files are read, and regular expressions are used to remove sequences containing identifiers such as "synthetic," "vector," "plasmid," and "cloning." That is, artificially constructed and non-pathogen-derived genome sequences are removed from all genome sequences of the target pathogen found in the metagenomic database to obtain the initial genome sequence. The specific regular expression can be as follows:
[0075] re.compile(r(?:^|[-_~.])(?:vector|artificial|unidentified|plasmid|recombinant|synthetic|cloning|clone)(?:$|[-_~.])', (re.IGNORECASE); then download the representative genome sequence of the corresponding target pathogen, prioritizing downloads from reference databases such as reference, refSeq, and genbank. Specific download commands can be obtained by editing a program script, for example, using a Python script to execute download commands such as axel-q-n5-k, with a maximum of 10 retries, each with a 10-second interval; after downloading, use gunzip to decompress the file; for eukaryotes, multiple chromosome sequences are merged into a single continuous sequence through sequence splicing; further, for pathogen species in the database, subspecies or varieties are used as the main representatives, and representative genome sequences are determined based on taxonomic relationships. The longest sequence is selected as the representative for other taxonomic units, and the selection logic is implemented through sequence length comparison; further, a species name annotation database is established, removing genome sequences containing artificially synthesized or cloning vector identifiers to ensure data purity, implemented using regular expression matching and filtering logic. Specifically, 1) construct a sequence index file, format package The index is calculated and written to the sequence file by parsing the sequence file, including tax unit identifiers, sequence identifiers, file offsets, and sequence lengths. Python file operations are used to read `core_nt_normalization.fasta` to generate the index file `coreNT.index`, recording the sequence information for each `taxid`, for example, in the format `taxid;ncbiID`. 2) A fast location file for the index is created to map tax units to index positions. A data structure is used to store the mapping relationship. An `index.index` file is created to store key-value pairs of `taxid` and file offsets, for example, `taxid`. 3) Special processing is applied to virus sequences. Subtype or variant sequences are merged based on taxonomic guidance. Mapping logic is used to implement classification merging, for example, mapping `taxid` to a parent `taxid` and updating the index. 4) For other species, sub-species tax units are merged to the species level to reduce redundant classifications. A tree structure is constructed by parsing the `taxid_nodes` file and recursively merged, removing redundant `taxid`s. Specific steps are as follows:
[0076] Based on the taxid in the mNGS file, vertebrate and plant sequences (except humans) are removed using list operations for filtering; non-pathogenic sequences are actively removed using filtering logic, such as excluding specific categories from the taxid list; artificial sequences are removed to ensure sequence quality by using name matching to remove identifiers such as 'synthetic', implemented using regular expressions; human sequences are retained for quality control and validation in subsequent processing; finally, the genome sequences are standardized to ensure a uniform format for all genome sequences, facilitating subsequent processing. The specific steps are as follows:
[0077] The process involves several steps: standardizing sequence formatting (e.g., converting to uppercase); standardizing sequence identifier formatting to ensure consistency and readability (e.g., using `taxid.fasta` for naming); handling special characters and non-standard bases by replacing invalid characters with regular expressions; connecting sequences using delimiters (e.g., polyN sequences) to form contiguous sequences; and finally, performing performance optimizations, employing multi-threading and memory optimization strategies to improve processing efficiency. The specific steps are as follows:
[0078] Multi-threaded processing is configured to improve efficiency. The number of threads is set via command-line parameters (default 64 threads) using the Python threading module. Threshold control is employed to prevent memory overflow. A loop threshold (500000000000 / seqLength) is set to process data in batches. Processing progress and performance data are recorded and written to a log file, with the format including timestamps and performance metrics. Temporary files are managed to ensure orderly processing. Intermediate files are stored in a working directory, and temporary files are cleaned up periodically.
[0079] In this way, by systematically integrating multi-source data, the initial genome sequence corresponding to the target pathogen can be obtained. Then, based on the specificity of the pathogen species, the first filtering is performed to simplify the amount of data and standardize the obtained standard non-redundant genome sequence fragments. This achieves comprehensive database protection in metagenomic sequencing scenarios, solves the risk of missed detection, and ensures the efficiency of the processing through performance optimization.
[0080] Step 104: Based on the classification and retrieval information of all standard genome sequences, duplicates are removed from all standard genome sequences to obtain multiple non-redundant genome sequences with different target nucleic acid characteristics.
[0081] It should be noted that the classification retrieval information represents the identification information of the taxonomic affiliation of the standard genome sequence, specifically including classification information, attribute information, and index information. Among them, the classification information includes the taxonomic hierarchy identifier of the target pathogen, which is used to identify the taxonomic group corresponding to the standard genome sequence; the attribute information includes the identifier of the genome sequence and the target nucleic acid feature type; and the index information may include the standard genome sequence.
[0082] It should be noted that by establishing a sequence indexing system, the classification and retrieval information for each standard genome sequence can be predetermined. For example, an index file can be created for each taxonomic unit, recording information such as sequence identifier, start position, and length, thereby quickly constructing the retrieval data structure. The specific steps are as follows: 1) Create a sequence index file, the format of which includes taxonomic unit identifier, sequence identifier, file offset, and sequence length. Calculate and write the index by parsing the sequence file. Use Python file operations to read core_nt_normalization.fasta to generate the index file coreNT.index, recording the sequence information for each taxid, for example, in the format... 1) Taxid;ncbiID; 2) Establish a fast location file for the index, realize the mapping from taxonomic units to index positions, use data structures to store the mapping relationship, create an index.index file, and store key-value pairs of taxid and file offset, such as taxid; 3) Perform special processing on virus sequences, merge subtype or variant sequences based on taxonomic guidance, and use mapping logic to realize classification merging, such as mapping taxid to parent taxid and updating the index; 4) For other pathogen species, merge subspecies taxonomic units to the species level to reduce redundant classifications, build a tree structure by parsing the taxid_nodes file and recursively merge, and remove redundant taxids.
[0083] As an example, step 104 includes: selecting multiple non-redundant genomic sequences with different target nucleic acid features from all standard genomic sequences by comparing classification retrieval information of different standard genomic sequences.
[0084] Step 106: Generate a standard sequence index file containing classification retrieval information for each non-redundant genome sequence, and a standard sequence alignment file containing the target nucleic acid features of all non-redundant genome sequences.
[0085] It should be noted that the standard sequence index file is a catalog file for retrieving non-redundant genomic sequences. The mapping relationship established in the standard sequence index file facilitates the improvement of retrieval efficiency in the subsequent use of nucleic acid sequence databases. The standard sequence alignment file is a total file that integrates the target nucleic acid features of all non-redundant genomic sequences, which can clearly show the differences in target nucleic acid features between different sequences.
[0086] As an example, step 106 includes: generating a standard sequence index file containing the classification retrieval information of each non-redundant genome sequence using the classification retrieval information of all standard genome sequences, and generating a standard sequence alignment file containing the targeted nucleic acid features of all non-redundant genome sequences by integrating all non-redundant genome sequences.
[0087] Step 108: Construct a nucleic acid sequence database of the target pathogen based on the standard sequence index file and the standard sequence alignment file.
[0088] As an example, standard sequence index files and standard sequence alignment files are packaged into a nucleic acid sequence database of the target pathogen.
[0089] In one feasible approach, standard sequence index files and standard sequence alignment files can be used as core data components, along with a structured storage directory, a data format standardization module, and retrieval interface adaptation rules, to encapsulate a nucleic acid sequence database for the target pathogen. Specifically, the database's file organization structure (hierarchical storage) needs to be clearly defined first, and then an association mapping needs to be established between the index files and alignment files (e.g., to achieve fast jumps via sequence IDs). At the same time, it should be compatible with industry-standard database formats to ensure that the constructed nucleic acid sequence database can be directly connected to high-throughput sequencing data analysis tools.
[0090] The aforementioned method, apparatus, computer equipment, computer-readable storage medium, and computer program product for constructing a nucleic acid sequence database first obtain multiple standard genomic sequences characterizing the target nucleic acid features of the target pathogen, thus providing a data foundation for accurately covering the target nucleic acid features of the database to be constructed. Then, based on the classification and retrieval information of all standard genomic sequences, duplicates are removed from all standard genomic sequences to obtain multiple non-redundant genomic sequences with distinct target nucleic acid features, achieving the goal of redundancy removal without affecting the comprehensive coverage of the target nucleic acid features. Next, a standard sequence index file containing the classification and retrieval information of each non-redundant genomic sequence and a standard sequence alignment file containing the target nucleic acid features of all non-redundant genomic sequences are generated. Finally, based on the standard sequence index file and the standard sequence alignment file, a nucleic acid sequence database of the target pathogen is constructed. Due to the large number of nucleic acid sequences... The database construction process consistently uses the target nucleic acid characteristics of the pathogen as a benchmark. During the deduplication process, only standard genomic sequences with identical target nucleic acid characteristics are deduplicated. This allows for the reduction of the database size while preserving all target nucleic acid characteristics by generating standard sequence alignment files. Furthermore, since the nucleic acid sequence database is also built based on standard sequence index files, it provides a clear basis for high-throughput sequencing analysis, rather than directly relying on a full and complete reference genome sequence. Therefore, it overcomes the technical drawback of the large size of nucleic acid sequence databases due to the diversity of pathogen species, which can lead to long analysis times when using high-throughput sequencing technology. Thus, the constructed nucleic acid sequence database ensures comprehensive data coverage while maintaining efficient data retrieval.
[0091] In one embodiment, such as Figure 2 As shown, the classification retrieval information includes classification information and index information; based on the classification retrieval information of all standard genome sequences, duplicates are removed from all standard genome sequences, resulting in multiple non-redundant genome sequences with different target nucleic acid features, including:
[0092] Step 202: Based on the classification information, candidate non-redundant genomic sequences for each taxonomic group of the target pathogen are screened from all standard genomic sequences;
[0093] It should be noted that, due to sequence redundancy between different taxa, such as the shared conserved regions of target pathogens of different species within the same genus, if the screening step of grouping by taxonomic information is skipped and all standard genome sequences are directly compared and deduplicated globally, it is easy to misjudge such conserved characteristic sequences between taxa as redundant sequences and remove them. This would cause the database to lose the specific target nucleic acid features of different taxa, thus affecting the comprehensiveness of the constructed nucleic acid sequence database. Therefore, it is advisable to first perform preliminary deduplication on all standard genome sequences through taxonomic information to obtain candidate non-redundant genome sequences for each taxa.
[0094] As an example, step 202 includes: performing preliminary deduplication on all standard genome sequences according to classification information to obtain candidate non-redundant genome sequences for each taxonomic group of the target pathogen.
[0095] In one embodiment, candidate non-redundant genomic sequences for each taxonomic group of the target pathogen are screened from all standard genomic sequences based on classification information, including:
[0096] Based on classification information, representative genome sequences for each taxonomic group of the target pathogen are screened from all standard genome sequences. A standard sequence index table for the target pathogen is constructed based on all representative genome sequences. Each representative genome sequence is segmented to obtain multiple genome sequence fragments to be aligned, which are adapted to the targeted sequencing read length and cover the target nucleic acid features of each representative genome sequence. Based on the correspondence between each genome sequence fragment to be aligned and the standard sequence index table, redundant genome sequence fragments that cross taxonomic groups are removed from all genome sequence fragments to be aligned, resulting in representative genome sequence fragments for each taxonomic group of the target pathogen. All representative genome sequence fragments are integrated into candidate non-redundant genome sequences.
[0097] It should be noted that, firstly, all standard genome sequences can be initially screened based on classification information. Then, the core information of all representative genome sequences (taxonomic group identifier, sequence ID, target feature location, and storage path, etc.) is integrated into a structured index table, which serves as a reference directory for subsequent fragment alignment. That is, a standard sequence index table of the target pathogen is obtained, which is used to achieve rapid location and association. By targeting the sequencing read length, each representative genome sequence is divided into short genome sequence fragments to be aligned, ensuring that each genome sequence fragment can both fit the length of the sequencing data and completely cover the target nucleic acid features of the sequence. Then, by using the correspondence between the genome sequence fragments to be aligned and the standard sequence index table, redundant genome sequence fragments across taxonomic groups are eliminated. Finally, representative fragments of the same taxonomic group are spliced or associated according to their original genome positions to form candidate non-redundant genome sequences that cover all target features of the taxonomic group and have no cross-taxonomic redundancy.
[0098] As an example, based on classification information, the longest standard genome sequence of each taxonomic group of the target pathogen is selected from all standard genome sequences as the representative genome sequence; a standard sequence index table of the target pathogen is constructed based on the representative genome sequence; each representative genome sequence is divided into multiple genome sequence fragments to be aligned according to sequence segmentation rules that are adapted to the targeted sequencing read length and cover the target nucleic acid features of each representative genome sequence; based on the correspondence between each genome sequence fragment to be aligned and the standard sequence index table, redundant genome sequence fragments that cross taxonomic groups are removed from all genome sequence fragments to be aligned to obtain the representative genome sequence fragments of each taxonomic group of the target pathogen; all representative genome sequence fragments are spliced together to form a candidate non-redundant genome sequence.
[0099] In one feasible approach, the specific steps for deduplication to obtain candidate non-redundant genomic sequences can be as follows: 1) Set the k-mer length to 80 bp and the step size to 1 bp, ensuring that the 75 bp sequencing read information is lossless. Use the argparse module via command-line parameters to set the seqLength parameter to 80; 2) Build a sequence alignment index for representative genomes. Use the efficient alignment tool bwa-mem2 to build the index, with the command bwa-mem2.avx512bw The index uses multi-threading optimization, with the number of threads set via the `threadNum` parameter (default 64 threads); 3) The sequence is split into k-mer fragments of a specified length. K-mers are generated using a sliding window method to achieve sequence segmentation. For example, for each sequence, starting from position 0, 80bp k-mers are generated with a step size of 1bp and written to a FASTQ format file with the format `@ncbiID:position+`; 4) K-mer alignment is performed using a sequence alignment tool with the parameter set to the maximum k-mer length (-k75). The alignment results are parsed to identify unaligned k-mers. The command is `bwa-mem2.avx512bw mem -t threadNum -k seqLength`. The output lines are parsed to check if a "4" header indicates an unaligned sequence; 5) Identifying identical redundant fragments: The alignment output is parsed to check if the k-mer is not present in the index and is not in the recorded k-mer set. If it is not present, the genome sequence fragment is retained, membership is checked using a Python set data structure, and position information is recorded in a dictionary, for example, stored as `ncbiID`. [position_list].
[0100] Step 204: Based on the index information, all representative genome sequences are aligned and deduplicated to obtain multiple non-redundant genome sequences with different target nucleic acid features.
[0101] It should be noted that, firstly, based on the filtering of classification information, by defining the boundaries of taxonomic groups, the deduplication scope is limited to the same taxonomic group, effectively avoiding conserved sequences between different taxonomic groups (which are misjudged as redundant and removed), ensuring that the database fully retains the specific target nucleic acid features of each taxonomic group, providing an accurate basis for subsequent precise classification and identification of pathogens; then, based on the comparison and deduplication of index information, for sequences with completely repeated features within the same taxonomic group, redundancy is accurately removed by the differences in the attributes of the target nucleic acid features, significantly reducing the amount of data in the database and avoiding storage and retrieval burdens; at the same time, the retained non-redundant sequences have different target features, which can be directly used for rapid comparison of high-throughput sequencing data, reducing matching ambiguities caused by repetitive sequences and improving the efficiency and accuracy of pathogen identification.
[0102] In one embodiment, the target pathogen includes a viral pathogen; based on the correspondence between each genome sequence fragment to be aligned and a standard sequence index table, redundant genome sequence fragments crossing taxa are removed from all genome sequence fragments to be aligned, resulting in a representative genome sequence fragment for each taxa of the target pathogen, including:
[0103] Based on the type of viral pathogen, the classification information of the viral pathogen in the standard sequence index table is updated to obtain a target standard sequence index table adapted to the viral classification characteristics; multiple viral genome sequence fragments to be compared are identified from all the genome sequence fragments to be compared; based on the correspondence between all the viral genome sequence fragments to be compared and the target standard sequence index table, redundant genome sequence fragments crossing viral taxonomic groups are recursively removed from all the viral genome sequence fragments to be compared in reverse order of the taxonomic lineage of the viral pathogen, to obtain representative genome sequence fragments for each taxonomic group of the viral pathogen; all representative genome sequence fragments are integrated into candidate non-redundant genome sequences, including:
[0104] Based on the distribution information of the target nucleic acid characteristics of the viral pathogen, all representative genome sequence fragments are specifically spliced to obtain processed representative genome sequence fragments;
[0105] If all processed representative genomic sequence fragments pass quality verification, all processed representative genomic sequence fragments are integrated into a candidate non-redundant genomic sequence for the viral pathogen. The target nucleic acid feature distribution information includes one or more of the target nucleic acid feature distribution locations and target nucleic acid feature variation intervals.
[0106] It should be noted that, given the specific characteristics of the genome sequence of a viral pathogen, specific screening rules will be set on top of the aforementioned redundancy removal mechanism to ensure that the candidate non-redundant genome sequences of the pathogen obtained after processing meet the requirements. The target nucleic acid feature distribution information is used to describe the specific distribution of viral target nucleic acid features in the genome sequence. Specifically, the target nucleic acid feature distribution information may include one or more of the following: the target nucleic acid feature distribution location and the target nucleic acid feature variation range. The target nucleic acid feature distribution location describes the specific location of the viral target nucleic acid feature in the genome sequence, and the target nucleic acid feature variation range describes the range of variation of the viral target nucleic acid feature in the genome sequence.
[0107] As an example, the classification information of viral pathogens in the standard sequence index table is classified and merged according to the type of viral pathogen to obtain a target standard sequence index table adapted to the viral classification characteristics. Based on the alignment relationship between all viral genome sequence fragments to be aligned and the target standard sequence index table, redundant genome sequence fragments that cross viral taxonomic groups are recursively removed from all viral genome sequence fragments to be aligned in reverse order of the taxonomic lineage of viral pathogens to obtain representative genome sequence fragments for each taxonomic group of viral pathogens. Based on the distribution location and variation range of the target nucleic acid features of viral pathogens, the representative genome sequence fragments are specifically spliced to obtain processed representative genome sequence fragments. The processed representative genome sequence fragments are quality verified, and if the quality verification is passed, all processed representative genome sequence fragments are spliced together to form a candidate non-redundant genome sequence of the viral pathogen.
[0108] Specifically, based on the distribution location and variation range of the target nucleic acid characteristics of the viral pathogen, all representative genomic sequence fragments were specifically spliced to obtain processed representative genomic sequence fragments, including:
[0109] Based on the distribution location and variation range of the target nucleic acid characteristics of the viral pathogen, each representative genomic sequence fragment is assigned a corresponding fragment splicing weight. Based on these weights, all representative genomic sequence fragments are sorted to obtain a sorting result. Based on the sorting result, target representative genomic sequence fragments are extracted from each of the representative genomic sequence fragments. All target representative genomic sequence fragments are then spliced together to form the processed representative genomic sequence fragment.
[0110] It should be noted that the functional attributes of the target nucleic acid feature distribution locations and the biological value of the target nucleic acid feature variation regions differ significantly for different viral pathogens. Therefore, in the specific splicing process, differentiated weights can be set based on the functional attributes and biological values of different viral pathogens to prioritize the integrity of core target features and the accuracy of variation information. For example, in one feasible approach, assuming the viral pathogen is hepatitis B virus (HBV), since the conserved functional regions of HBV are the core of viral replication and detection target design, the weights corresponding to the conserved functional regions are set at a higher level, and the variation regions are assigned corresponding weights according to their biological value. Among these, the fragment splicing weight represents the priority of the genomic sequence fragment in specific splicing.
[0111] As an example, a target nucleic acid feature mapping relationship is generated based on the distribution location and variation range of the target nucleic acid feature of the viral pathogen. All representative genomic sequence fragments are assigned splicing weights according to this mapping relationship. The target nucleic acid feature mapping relationship characterizes the correspondence between the functional attributes of the target nucleic acid feature distribution location and the biological value of the target nucleic acid feature variation range. All representative genomic sequence fragments are sorted in descending order according to their splicing weights to obtain the sorting results. Fragments with the same weight can be further sorted based on sequencing quality and coverage. High, medium, and low priority target representative genomic sequence fragments are extracted hierarchically according to the sorting results. First, high-priority fragments are spliced to obtain the first target representative genomic sequence fragment. Then, medium-priority fragments are integrated to obtain the second target representative genomic sequence fragment. Finally, a third target representative genomic sequence fragment is used to supplement the sequence fragments. The first, second, and third target representative genomic sequence fragments are then sequentially spliced to form the processed representative genomic sequence fragment.
[0112] In one feasible approach, assuming the viral pathogen is A, and all representative genomic sequence fragments include a1, a2, and a3, firstly, fragment splicing weights are assigned to each of the representative genomic sequence fragments a1, a2, and a3. Specifically, the distribution location of the target feature and the target nucleic acid feature variation interval are labeled for each representative genomic fragment. Weights are assigned based on the functional attributes and biological value of the target nucleic acid features: a1's target feature distribution location is a conserved functional region, with a location weight of 0.6; the target nucleic acid feature variation interval is a characteristic variation site, with a variation interval weight of 0.3; and the fragment splicing weight is 0.51 (0.6 × 0.7 + 0.3 × 0.3). a2's target feature distribution location is a cross-functional region junction, with a location weight of 0.5; the variation interval type is a common variation interval, with a variation interval weight of 0.1; and the fragment splicing weight is 0.38 (0.5 × 0.7 + 0.1 × 0.3). a3's target feature distribution location is a non-conserved non-functional region, with a location weight of 0.2; there is no target nucleic acid feature variation interval; and the variation interval is not specified. The inter-fragment weight is 0, and the fragment splicing weight is 0.14 (0.2×0.7+0×0.3). Then, the fragments are sorted in descending order of splicing weight to obtain a1>a2>a3. High-priority a1, medium-priority a2, and low-priority a3 target representative genome sequence fragments are extracted hierarchically. The first target representative genome sequence fragment is spliced with a1 as the anchor point without modifying the target feature sequence of a1 and without gaps at the anchor point. Then, the common variants of a2 are integrated. If there is no conflict with the first target representative genome sequence fragment, they are retained. The connection fragments between the variant regions and the core backbone are optimized through local alignment to obtain the second target representative genome sequence fragment. Finally, the second target representative genome sequence fragment is completed with a3. The fragment gap is allowed to be less than or equal to 50bp and filled with N. The completed sequence fragment is compared with the high-weight region, and low homology fragments are removed. The third target representative genome sequence fragment is spliced to obtain the final representative genome sequence fragment. Finally, the first, second, and third target representative genome sequence fragments are spliced together to form the processed representative genome sequence fragment.
[0113] In one feasible approach, the specific steps for compression processing of viral pathogen-specific data can be as follows:
[0114] First, identify the viral taxonomic units, analyze the viral classification based on the taxonomic database, and use the taxid_nodes file to determine the lineage relationship. Virus taxids begin with 10239. 1) For specific viruses, such as influenza viruses, merge subtype sequences to implement classification merging logic. For example, influenza A is merged into the HXNX subtype, and influenza B is merged into the taxid_nodes file. 11520, Update the index using a mapping table; 2) For other viruses, merge subspecies taxa into species level to maintain classification consistency, by resolving lineages and recursively processing subclasses: Then, perform redundancy removal according to the reverse lineage, that is, start redundancy removal from the lowest end of the taxonomic lineage to ensure comprehensiveness: 1) Obtain the list of secondary taxa, establish lineage relationships, and use the taxid_nodes file to resolve parent-child relationships; 2) Perform redundancy removal starting from the lowest level of classification, recursively process subclasses, and prioritize units with lower taxid levels; 3) When performing redundancy removal, integrate all downstream sequence data and build a temporary index file; 4) After building the index, delete redundant sequences, keep only the index sequences themselves, and update the core database; 5) Process lineage levels sequentially (such as level 3, level 2, or level 1) to ensure systematic processing, using loop logic; Then, perform specific k-mer processing on all representative genomic sequence fragments, targeting To address the small size of the viral genome, k-mer processing parameters were optimized: 1) For shorter sequences (e.g., less than 25M in length), the processing logic was optimized, potentially skipping the complete index construction and directly using the k-mer set for processing; 2) k-mer alignment parameters were adjusted to accommodate high variation characteristics, such as reducing the k-mer length or adjusting the step size, but in this embodiment, 80bp and 1bp were maintained to ensure consistency; 3) Special attention was paid to variant regions to ensure complete information preservation, and variant sites were analyzed through alignment results; Finally, quality control and verification were performed to ensure the integrity and accuracy of the information after redundancy removal from the viral sequence: 1) Sequence integrity checks were performed to ensure no important sequences were lost, and the sequence coverage before and after redundancy removal was compared; 2) The preservation of variant sites was checked, and alignment tools were used to verify whether k-mer correctly identified variants; 3) The alignment efficiency of sequencing reads was verified to ensure that the 75bp read information was intact, and tests were conducted using simulated sequencing data.
[0115] Thus, in this embodiment, by employing a virus-specific classification and merging strategy and a reverse lineage redundancy removal method, the system accurately adapts to the complex classification and rapid mutation characteristics of viruses, efficiently preserving the specific targeting features of each taxonomic group. This significantly improves the quality of the viral nucleic acid sequence database, strongly supporting diverse identification needs from broad-spectrum screening to precise typing. It solves the redundancy removal challenge brought about by the high variability and complex classification relationships of viral sequences, achieving efficient volume compression while ensuring the integrity of viral sequence information, making it particularly suitable for viral pathogen detection scenarios.
[0116] In one feasible approach, refer to Figure 3 , Figure 3 To illustrate the overall workflow of constructing a nucleic acid sequence database, the process begins with acquiring the taxonomic and phylogenetic information of the target pathogen. Based on this information, the corresponding initial genome sequence is searched from a metagenomic database. The source of the initial genome sequence is then determined. If a reference genome sequence is detected, it is downloaded and processed. If no reference genome sequence is found, the longest reference genome sequence is extracted, resulting in multiple standard genome sequences for the target pathogen. Finally, a BWA index-K75 is built to determine the classification and retrieval information for all standard genome sequences. This classification and retrieval information includes classification information and retrieval details. Information; then, based on classification information, candidate non-redundant genomic sequences for each taxonomic group of the target pathogen are screened from all standard genomic sequences; based on index information, all representative genomic sequences are aligned and deduplicated to obtain multiple non-redundant genomic sequences with different target nucleic acid characteristics, that is, redundant genomic sequences are discarded when they are aligned, and k-mer sequences are checked and recorded when they are not aligned, and finally integrated to obtain non-redundant genomic sequences; during this process, the standard sequence index table is updated and cleared based on the set threshold, and the standard sequence index file and standard sequence alignment file are finally output; thus, temporary files are deleted, and a nucleic acid sequence database is constructed based on the standard sequence index file and standard sequence alignment file.
[0117] Because the construction of the nucleic acid sequence database is always based on the target nucleic acid characteristics of the pathogen, and the deduplication process only targets standard genome sequences with the same target nucleic acid characteristics, it is possible to reduce the amount of data in the nucleic acid sequence database while retaining all target nucleic acid characteristics by generating standard sequence alignment files. On the other hand, since the nucleic acid sequence database is also built based on standard sequence index files, it can provide a clear basis for high-throughput sequencing analysis, rather than directly relying on the full and complete reference genome sequence to build the nucleic acid sequence database. Therefore, it overcomes the technical defect that the large size of the nucleic acid sequence database due to the diversity of pathogen species makes it easy to have long analysis time when using high-throughput sequencing technology. Therefore, the constructed nucleic acid sequence database can ensure the comprehensiveness of data coverage while taking into account the efficiency of data retrieval.
[0118] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0119] Based on the same inventive concept, this application also provides a nucleic acid sequence database construction apparatus for implementing the nucleic acid sequence database construction method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the nucleic acid sequence database construction apparatus provided below can be found in the limitations of the nucleic acid sequence database construction method described above, and will not be repeated here.
[0120] In one exemplary embodiment, such as Figure 4 As shown, a nucleic acid sequence database construction device is provided, including: an acquisition module 301, a deduplication module 302, a generation module 303, and a construction module 304, wherein:
[0121] The acquisition module 301 is used to acquire multiple standard genome sequences of the target pathogen, wherein the standard genome sequences characterize the target nucleic acid features of the target pathogen;
[0122] The deduplication module 302 is used to remove duplicates from all standard genome sequences based on the classification retrieval information of all standard genome sequences, and obtain multiple non-redundant genome sequences with different target nucleic acid features.
[0123] The generation module 303 is used to generate a standard sequence index file containing classification retrieval information for each non-redundant genome sequence, and a standard sequence alignment file containing the target nucleic acid features of all non-redundant genome sequences;
[0124] Module 304 is used to construct a nucleic acid sequence database of the target pathogen based on the standard sequence index file and the standard sequence alignment file.
[0125] In one embodiment, the classification retrieval information includes classification information and index information; the deduplication module 302 is further used for:
[0126] Based on classification information, candidate non-redundant genome sequences for each taxonomic group of the target pathogen are screened from all standard genome sequences; based on index information, all representative genome sequences are aligned and deduplicated to obtain multiple non-redundant genome sequences with different target nucleic acid characteristics.
[0127] In one embodiment, the deduplication module 302 is further configured to:
[0128] Based on classification information, representative genome sequences for each taxonomic group of the target pathogen are screened from all standard genome sequences. A standard sequence index table for the target pathogen is constructed based on all representative genome sequences. Each representative genome sequence is segmented to obtain multiple genome sequence fragments to be aligned, which are adapted to the targeted sequencing read length and cover the target nucleic acid features of each representative genome sequence. Based on the correspondence between each genome sequence fragment to be aligned and the standard sequence index table, redundant genome sequence fragments that cross taxonomic groups are removed from all genome sequence fragments to be aligned, resulting in representative genome sequence fragments for each taxonomic group of the target pathogen. All representative genome sequence fragments are integrated into candidate non-redundant genome sequences.
[0129] In one embodiment, the target pathogen includes a viral pathogen; the deduplication module 302 includes:
[0130] The rejection unit is used for:
[0131] Based on the type of viral pathogen, update the classification information of the viral pathogen in the standard sequence index table to obtain a target standard sequence index table adapted to the viral classification characteristics; identify multiple viral genome sequence fragments to be compared from all the genome sequence fragments to be compared; based on the correspondence between all the viral genome sequence fragments to be compared and the target standard sequence index table, recursively remove redundant genome sequence fragments that cross viral taxa from all the viral genome sequence fragments to be compared in reverse order of the taxonomic lineage of the viral pathogen to obtain the representative genome sequence fragment of each taxa of the viral pathogen;
[0132] Integration unit, the integration unit is used for:
[0133] Based on the distribution information of the target nucleic acid features of the viral pathogen, all representative genomic sequence fragments are specifically spliced to obtain processed representative genomic sequence fragments. If all processed representative genomic sequence fragments pass quality verification, all processed representative genomic sequence fragments are integrated into a candidate non-redundant genomic sequence of the viral pathogen. The distribution information of the target nucleic acid features includes one or more of the target nucleic acid feature distribution locations and target nucleic acid feature variation intervals.
[0134] In one embodiment, the deduplication module 302 is further configured to:
[0135] Based on the index information, all representative genome sequences are aligned to identify multiple identical representative genome sequences to be deduplicated; target deduplicated representative genome sequences that meet the preset deduplication criteria are identified from all representative genome sequences to be deduplicated, wherein the preset deduplication criteria are generated based on the attribute difference information of the target nucleic acid features; the target deduplicated representative genome sequences are removed from all representative genome sequences to obtain non-redundant genome sequences.
[0136] In one embodiment, the acquisition module 301 is further configured to:
[0137] Based on the taxonomic information of the target pathogen, the corresponding initial genome sequence is searched in the metagenomic database. If a reference genome sequence is detected in the initial genome sequence, the reference genome sequence is extracted from the reference database as the first candidate genome sequence. If no reference genome sequence is detected in the initial genome sequence, a representative genome sequence is extracted from the initial nucleic acid sequence database as the second candidate genome sequence. Based on the taxonomic lineage of the target pathogen, the first candidate genome sequence and the second candidate genome sequence are integrated into multiple standard genome sequences.
[0138] In one embodiment, the acquisition module 301 is further configured to:
[0139] Based on the taxonomic lineage of the target pathogen, the first and second candidate genome sequences are located to their respective taxa. According to the sequence redundancy rules of the taxa, redundant genome sequence fragments are removed from the first and second candidate genome sequences to obtain multiple initial non-redundant genome sequence fragments. All multiple initial non-redundant genome sequence fragments are converted into standard non-redundant genome sequence fragments in a preset standard format. According to the attribute difference information of the target nucleic acid features, the multiple standard non-redundant genome sequence fragments are integrated into multiple standard genome sequences.
[0140] Each module in the aforementioned nucleic acid sequence database construction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0141] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for constructing a nucleic acid sequence database. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0142] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0143] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0144] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0145] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0146] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0147] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for constructing a nucleic acid sequence database, characterized in that, The method includes: Obtain multiple standard genome sequences of the target pathogen, wherein the standard genome sequences characterize the target nucleic acid features of the target pathogen; Based on the classification and retrieval information of all standard genome sequences, duplicates of all standard genome sequences are removed to obtain multiple non-redundant genome sequences with different target nucleic acid features; Generate a standard sequence index file containing classification retrieval information for each of the non-redundant genome sequences, and a standard sequence alignment file containing the targeted nucleic acid features of all the non-redundant genome sequences; A nucleic acid sequence database of the target pathogen is constructed based on the standard sequence index file and the standard sequence alignment file, wherein the classification retrieval information includes classification information and index information; based on the classification retrieval information of all standard genome sequences, duplicates are removed from all standard genome sequences to obtain multiple non-redundant genome sequences with different target nucleic acid characteristics, including: Based on the classification information, candidate non-redundant genomic sequences for each taxonomic group of the target pathogen are screened from all standard genomic sequences; based on the index information, all representative genomic sequences are aligned and deduplicated to obtain multiple non-redundant genomic sequences with distinct target nucleic acid features, wherein the step of screening candidate non-redundant genomic sequences for each taxonomic group of the target pathogen from all standard genomic sequences based on the classification information includes: Based on the classification information, representative genome sequences for each taxonomic group of the target pathogen are selected from all standard genome sequences; a standard sequence index table for the target pathogen is constructed based on all representative genome sequences; each representative genome sequence is segmented to obtain multiple genome sequence fragments to be aligned, which are adapted to the targeted sequencing read length and cover the targeted nucleic acid features of each representative genome sequence; based on the correspondence between each genome sequence fragment to be aligned and the standard sequence index table, redundant genome sequence fragments that cross taxonomic groups are removed from all genome sequence fragments to be aligned, resulting in representative genome sequence fragments for each taxonomic group of the target pathogen; all representative genome sequence fragments are integrated into the candidate non-redundant genome sequence.
2. The method according to claim 1, characterized in that, The target pathogen includes viral pathogens; the step of removing redundant genomic sequence fragments across taxa from all genomic sequence fragments to be aligned, based on the correspondence between each genomic sequence fragment to be aligned and the standard sequence index table, to obtain a representative genomic sequence fragment for each taxa of the target pathogen, includes: Based on the type of the viral pathogen, update the classification information of the viral pathogen in the standard sequence index table to obtain a target standard sequence index table adapted to the viral classification characteristics; Multiple viral genome sequence fragments of the viral pathogen were identified from all the genome sequence fragments to be aligned; Based on the correspondence between all viral genome sequence fragments to be compared and the target standard sequence index table, redundant genome sequence fragments that cross viral taxonomic groups are recursively removed from all viral genome sequence fragments to be compared in reverse order of the taxonomic lineage of the viral pathogen, so as to obtain the representative genome sequence fragment of each taxonomic group of the viral pathogen. The process of integrating all representative genomic sequence fragments into the candidate non-redundant genomic sequence includes: Based on the target nucleic acid feature distribution information of the viral pathogen, all representative genome sequence fragments are specifically spliced to obtain processed representative genome sequence fragments; If all processed representative genomic sequence fragments pass quality verification, all processed representative genomic sequence fragments are integrated into a candidate non-redundant genomic sequence for the viral pathogen, wherein the target nucleic acid feature distribution information includes one or more of the target nucleic acid feature distribution location and the target nucleic acid feature variation interval.
3. The method according to claim 1, characterized in that, The step of aligning and deduplicating all representative genomic sequences according to the index information to obtain multiple non-redundant genomic sequences with distinct target nucleic acid features includes: Based on the index information, all representative genome sequences are compared to multiple identical representative genome sequences to be deduplicated; Identify target deduplication representative genome sequences that meet preset deduplication criteria from all representative genome sequences to be deduplicated, wherein the preset deduplication criteria are generated based on attribute difference information of targeted nucleic acid features; The target deduplicated representative genome sequence is removed from all the representative genome sequences to obtain the non-redundant genome sequence.
4. The method according to claim 1, characterized in that, The acquisition of multiple standard genome sequences of the target pathogen includes: Based on the taxonomic information of the target pathogen, the corresponding initial genome sequence is searched from the metagenomic database; If a reference genome sequence is detected in the initial genome sequence, the reference genome sequence is extracted from the reference database as the first candidate genome sequence. If a reference genome sequence is detected as not existing in the initial genome sequence, a representative genome sequence is extracted from the initial nucleic acid sequence database as a second candidate genome sequence. Based on the taxonomic lineage of the target pathogen, the first candidate genome sequence and the second candidate genome sequence are integrated into the plurality of standard genome sequences.
5. The method according to claim 4, characterized in that, The step of integrating the first candidate genome sequence and the second candidate genome sequence into the plurality of standard genome sequences according to the taxonomic lineage of the target pathogen includes: Based on the taxonomic lineage of the target pathogen, the first candidate genome sequence and the second candidate genome sequence are respectively located to their corresponding taxa; According to the sequence redundancy rules of the taxonomic group, redundant genomic sequence fragments are removed from the first candidate genomic sequence and the second candidate genomic sequence to obtain multiple initial non-redundant genomic sequence fragments; All the initial non-redundant genomic sequence fragments are converted into standard non-redundant genomic sequence fragments in a preset standard format; Based on the attribute differences of the target nucleic acid features, multiple standard non-redundant genomic sequence fragments are integrated into the multiple standard genomic sequences.
6. A nucleic acid sequence database construction device, characterized in that, The device includes: An acquisition module is used to acquire multiple standard genome sequences of a target pathogen, wherein the standard genome sequences characterize the target nucleic acid features of the target pathogen; The deduplication module is used to remove duplicates from all standard genome sequences based on the classification retrieval information of all standard genome sequences, thereby obtaining multiple non-redundant genome sequences with different target nucleic acid features; The generation module is used to generate a standard sequence index file containing classification retrieval information for each of the non-redundant genome sequences, and a standard sequence alignment file containing the target nucleic acid features of all the non-redundant genome sequences; The construction module is used to construct a nucleic acid sequence database of the target pathogen based on the standard sequence index file and the standard sequence alignment file, wherein the classification retrieval information includes classification information and index information; the deduplication module is further used for: Based on the classification information, candidate non-redundant genomic sequences for each taxonomic group of the target pathogen are screened from all standard genomic sequences; based on the index information, all representative genomic sequences are aligned and deduplicated to obtain multiple non-redundant genomic sequences with distinct target nucleic acid features, wherein the deduplication module is further used for: Based on the classification information, representative genome sequences for each taxonomic group of the target pathogen are selected from all standard genome sequences; a standard sequence index table for the target pathogen is constructed based on all representative genome sequences; each representative genome sequence is segmented to obtain multiple genome sequence fragments to be aligned, which are adapted to the targeted sequencing read length and cover the targeted nucleic acid features of each representative genome sequence; based on the correspondence between each genome sequence fragment to be aligned and the standard sequence index table, redundant genome sequence fragments that cross taxonomic groups are removed from all genome sequence fragments to be aligned, resulting in representative genome sequence fragments for each taxonomic group of the target pathogen; all representative genome sequence fragments are integrated into the candidate non-redundant genome sequence.
7. The apparatus according to claim 6, characterized in that, The target pathogen includes viral pathogens; the deduplication module includes: The rejection unit is used for: Based on the type of the viral pathogen, the classification information of the viral pathogen in the standard sequence index table is updated to obtain a target standard sequence index table adapted to the viral classification characteristics; multiple viral genome sequence fragments to be compared are identified from all the genome sequence fragments to be compared; based on the correspondence between all the viral genome sequence fragments to be compared and the target standard sequence index table, redundant genome sequence fragments that cross viral taxonomic groups are recursively removed from all the viral genome sequence fragments to be compared in reverse order of the taxonomic lineage of the viral pathogen to obtain a representative genome sequence fragment for each taxonomic group of the viral pathogen; Integration unit, the integration unit is used for: Based on the target nucleic acid feature distribution information of the viral pathogen, all representative genomic sequence fragments are specifically spliced to obtain processed representative genomic sequence fragments; if all processed representative genomic sequence fragments pass quality verification, all processed representative genomic sequence fragments are integrated into a candidate non-redundant genomic sequence of the viral pathogen, wherein the target nucleic acid feature distribution information includes one or more of the target nucleic acid feature distribution location and the target nucleic acid feature variation interval.
8. The apparatus according to claim 6, characterized in that, The deduplication module is also used for: Based on the index information, all representative genome sequences are aligned to multiple identical representative genome sequences to be deduplicated; a target deduplicated representative genome sequence that meets a preset deduplication standard is identified from all representative genome sequences to be deduplicated, wherein the preset deduplication standard is generated based on attribute difference information of targeted nucleic acid features; the target deduplicated representative genome sequence is removed from all representative genome sequences to obtain the non-redundant genome sequence.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Construction method of pathogenic microorganism metagenome database
CN118197436A
Construction method and application of pathogenic microorganism genome database
CN119993288A