A method, system, terminal and storage medium for identifying metagenome plasmids
By combining the improved Transformer model and the random forest model, the problem of low plasmid recognition accuracy and computational efficiency in metagenomic data is solved, and higher recognition accuracy and stability are achieved, which is suitable for contig analysis in the metagenome.
Patent Information
- Application Number
- CN202510101408.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-22
AI Technical Summary
In the existing metagenomic data, there is a high similarity between plasmids and chromosomal contigs, resulting in low recognition accuracy and calculation efficiency, and the inability to accurately identify plasmid contigs.
Plasmid recognition was performed using a method of combining the improved Transformer model with the random forest model. The improved Transformer model captures complex dependencies through the embedding layer, attention layer, and fully connected layer, and the random forest model provides a basis for decision fusion through genomic marker feature learning.
It improves the accuracy and stability of plasmid recognition, and can effectively process contig data in the metagenome, and is especially suitable for short-segment metagenomic data.
Smart Images

Figure CN119541645B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics DNA data mining, and in particular to a metagenomic plasmid recognition method, system, terminal and storage medium. Background Art
[0002] Metagenomics is an emerging technology for studying microbial communities. It uses high-throughput sequencing technology to obtain DNA information from environmental samples, which can comprehensively analyze the species composition, functional characteristics and evolutionary relationships of microbial communities. This technology has become the main method for microbiome research and has been widely used in agricultural ecology, environmental protection, medical health and other fields.
[0003] In complex metagenomic data, plasmids, as independent DNA molecules of bacteria, play a key role in the spread of antibiotic resistance and horizontal gene transfer. However, due to the high similarity between plasmids and chromosome contigs, the complexity and diversity of metagenomic data and the expansion of sample size, the existing plasmid recognition methods have low recognition accuracy and computational efficiency, and are unable to accurately identify plasmid contigs.
[0004] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0005] The main purpose of the present invention is to provide a metagenomic plasmid identification method, system, terminal and computer-readable storage medium, aiming to solve the problem in the prior art that due to the high similarity between plasmids and chromosome overlap groups, the complexity and diversity of metagenomic data and the expansion of sample size, the existing plasmid identification methods have low recognition accuracy and computational efficiency and cannot accurately identify plasmid overlap groups.
[0006] To achieve the above object, the present invention provides a method for identifying a metagenomic plasmid, the method comprising the following steps:
[0007] Obtaining a target genome contig, encoding the target genome contig according to a gene prediction tool to obtain an input feature vector, and aligning the target genome contig based on an alignment tool and a pre-constructed alignment library to obtain a genome feature;
[0008] Inputting the input feature vector into an improved Transformer model, performing plasmid recognition through an embedding layer, an attention layer, and a fully connected layer in the improved Transformer model, and outputting a first classification score, wherein the attention layer includes a gated attention unit and a hybrid block attention unit;
[0009] Inputting the genomic features into a random forest model, and obtaining a second classification score according to the voting results of each decision tree in the random forest model;
[0010] According to the classification model based on the attention mechanism, the first classification score and the second classification score are aggregated respectively to obtain a first matrix and a second matrix, and a plasmid recognition score is obtained according to the first matrix and the second matrix.
[0011] Optionally, encoding the target genome contigs according to a gene prediction tool to obtain an input feature vector specifically includes:
[0012] Predicting proteins in the target genome contigs according to a gene prediction tool to obtain a plurality of predicted proteins, and aligning all of the predicted proteins with a preset reference plasmid protein cluster to obtain a plurality of alignment results;
[0013] It is determined whether each comparison result is higher than a preset threshold value, and a corresponding plurality of determination results are obtained, and an input feature vector is generated according to all the determination results.
[0014] Optionally, the alignment tool and the pre-constructed alignment library are used to align the target genome contigs to obtain genome features, specifically including:
[0015] Pre-acquire the protein sequences in the data set library, and construct an alignment library according to the protein sequences in the data set library;
[0016] According to the target tool, gene prediction is performed on the target genome contigs to obtain a plurality of gene prediction proteins, and according to the comparison tool, all the gene prediction proteins are compared with the comparison library to obtain genome features.
[0017] Optionally, inputting the input feature vector into an improved Transformer model, performing plasmid recognition through an embedding layer, an attention layer, and a fully connected layer in the improved Transformer model, and outputting a first classification score specifically includes:
[0018] Inputting the input feature vector into the improved Transformer model, mapping the input feature vector into a vector representation according to the word embedding in the embedding layer, and adding position information to the vector representation according to the position embedding in the embedding layer to obtain an embedded vector;
[0019] Input the embedding vector into the hybrid block attention unit to divide the embedding vector to obtain multiple blocks, and calculate the local attention and global attention corresponding to each block respectively, and use the gated attention unit to calculate the attention representation of each block according to the local attention and global attention of each block, and splice all the attention representations to obtain the attention result;
[0020] The attention result is input into the fully connected layer for mapping to obtain the first classification score.
[0021] Optionally, inputting the genome feature into a random forest model, and obtaining a second classification score according to the voting result of each decision tree in the random forest model, specifically includes:
[0022] Presetting the parameters of the random forest model according to preset parameters;
[0023] Inputting the genome feature into a random forest model, predicting the genome feature according to each decision tree in the random forest model, and obtaining multiple classification results;
[0024] According to the classification results of all decision trees, a second classification score is generated.
[0025] Optionally, the first classification score and the second classification score are aggregated according to a classification model based on an attention mechanism to obtain a first matrix and a second matrix, and a plasmid recognition score is obtained according to the first matrix and the second matrix, specifically including:
[0026] Aggregating the first classification score and the second classification score according to the attention mechanism-based classification model to obtain the first matrix;
[0027] Obtaining the occurrence frequency of chromosome markers and the occurrence frequency of plasmid markers, and calculating the second matrix based on the classification model based on the attention mechanism;
[0028] The first matrix and the second matrix are multiplied row by row, and column average and normalization are performed to obtain a plasmid recognition score.
[0029] Optionally, the step of inputting the input feature vector into an improved Transformer model, performing plasmid recognition through an embedding layer, an attention layer, and a fully connected layer in the improved Transformer model, and outputting a first classification score, further comprises:
[0030] Based on the plasmid database and the chromosome database, a first preliminary data set is screened to obtain a preliminary data set; the first preliminary data set is screened to obtain a preliminary data set;
[0031] Dividing the preliminary data set into a test set and an original training set;
[0032] Performing a first preprocessing and a second preprocessing on the original training set to obtain a first training set and a second training set;
[0033] The improved Transformer model is trained according to the first training set, the random forest model is trained according to the second training set, and the improved Transformer model and the random forest model are tested according to the test set.
[0034] In addition, to achieve the above object, the present invention also provides a metagenomic plasmid recognition system, wherein the metagenomic plasmid recognition system comprises:
[0035] A feature acquisition module is used to obtain a target genome contig, encode the target genome contig according to a gene prediction tool to obtain an input feature vector, and align the target genome contig based on an alignment tool and a pre-built alignment library to obtain genome features;
[0036] A first classification score acquisition module, used for inputting the input feature vector into an improved Transformer model, performing plasmid recognition through an embedding layer, an attention layer, and a fully connected layer in the improved Transformer model, and outputting a first classification score, wherein the attention layer includes a gated attention unit and a hybrid block attention unit;
[0037] A second classification score acquisition module, used to input the genome feature into a random forest model, and obtain a second classification score according to the voting result of each decision tree in the random forest model;
[0038] A result generation module is used to aggregate the first classification score and the second classification score according to a classification model based on an attention mechanism to obtain a first matrix and a second matrix, and obtain a plasmid recognition score according to the first matrix and the second matrix.
[0039] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a metagenomic plasmid recognition program stored in the memory and executable on the processor, and when the metagenomic plasmid recognition program is executed by the processor, the steps of the metagenomic plasmid recognition method as described above are implemented.
[0040] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a metagenomic plasmid recognition program, and when the metagenomic plasmid recognition program is executed by a processor, the steps of the metagenomic plasmid recognition method as described above are implemented.
[0041] In the present invention, a target genome contig is obtained, the target genome contig is encoded according to a gene prediction tool to obtain an input feature vector, the target genome contig is compared based on an alignment tool and a pre-constructed alignment library to obtain a genome feature; the input feature vector is input into an improved Transformer model, plasmid recognition is performed through an embedding layer, an attention layer and a fully connected layer in the improved Transformer model, and a first classification score is output, wherein the attention layer includes a gated attention unit and a hybrid block attention unit; the genome feature is input into a random forest model, and a second classification score is obtained according to the voting result of each decision tree in the random forest model; according to a classification model based on an attention mechanism, the first classification score and the second classification score are respectively aggregated to obtain a first matrix and a second matrix, and a plasmid recognition score is obtained according to the first matrix and the second matrix. The present invention classifies plasmids by combining an improved Transformer model with a random forest model, wherein the improved Transformer model is optimized for a traditional Transformer structure to capture complex dependencies in overlapping groups of different lengths, so as to improve processing efficiency and performance, and is particularly suitable for overlapping group analysis in a metagenome; at the same time, the random forest model provides a basis for subsequent decision fusion by learning genome marker features, thereby improving the accuracy and stability of the overall classification; the classification results of the improved Transformer model are fused with the classification results of the random forest model, and an attention mechanism is used to achieve weighted fusion of classifier outputs. By combining these two methods, the improved Transformer's ability to understand complex sequences and the random forest's accurate learning of genome features can be fully utilized, thereby effectively improving the accuracy and stability of plasmid recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flow chart of a preferred embodiment of the method for identifying a metagenomic plasmid of the present invention;
[0043] Figure 2 is a schematic diagram of protein cluster generation in the metagenomic plasmid identification method of the present invention;
[0044] Figure 3 is a schematic diagram of the attention mechanism in the metagenomic plasmid identification method of the present invention;
[0045] Figure 4 It is a schematic diagram of the classification effect of the overlapping group test set in the metagenomic plasmid identification method of the present invention;
[0046] Figure 5 It is a schematic diagram of the classification effect of the short fragment contig test set in the metagenomic plasmid recognition method of the present invention;
[0047] Figure 6 It is a structural diagram of a preferred embodiment of the metagenomic plasmid recognition system of the present invention;
[0048] Figure 7 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0050] Metagenomics is an emerging technology for studying microbial communities. It uses high-throughput sequencing technology to obtain DNA information from environmental samples, which can comprehensively analyze the species composition, functional characteristics and evolutionary relationships of microbial communities. This technology has become the main method for microbiome research and has been widely used in agricultural ecology, environmental protection, medical health and other fields. In complex metagenomes, plasmids, as independent DNA molecules of bacteria, play a key role in the spread of antibiotic resistance and horizontal gene transfer. However, due to the high similarity between plasmids and chromosome contigs, the complexity and diversity of metagenomic data and the expansion of sample size, the recognition accuracy and computational efficiency of existing plasmid recognition methods are low, and plasmid contigs cannot be accurately identified. In addition, current recognition methods have insufficient performance when processing metagenomic data sets, especially for short-fragment metagenomic data. It is difficult to fully utilize the multi-dimensional feature information of plasmid contigs using a single machine learning or deep learning model, and accurate results cannot be obtained.
[0051] In response to one or more of the above problems, the present invention obtains a target genome contig, encodes the target genome contig according to a gene prediction tool to obtain an input feature vector, aligns the target genome contig based on an alignment tool and a pre-constructed alignment library to obtain a genome feature; inputs the input feature vector into an improved Transformer model, performs plasmid recognition through an embedding layer, an attention layer, and a fully connected layer in the improved Transformer model, and outputs a first classification score, wherein the attention layer includes a gated attention unit and a hybrid block attention unit; inputs the genome feature into a random forest model, and obtains a second classification score according to the voting results of each decision tree in the random forest model; according to a classification model based on an attention mechanism, aggregates the first classification score and the second classification score respectively to obtain a first matrix and a second matrix, and obtains a plasmid recognition score according to the first matrix and the second matrix.
[0052] The metagenome plasmid identification method described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the metagenomic plasmid identification method comprises the following steps:
[0053] Step S10, obtaining a target genome contig, encoding the target genome contig according to a gene prediction tool to obtain an input feature vector, and aligning the target genome contig based on an alignment tool and a pre-constructed alignment library to obtain genome features.
[0054] It should be noted that, in the present invention, the target genome contigs to be processed are processed separately, so that the obtained results can be processed and identified by the improved Transformer model and the random forest model.
[0055] Furthermore, encoding the target genome contigs according to the gene prediction tool to obtain an input feature vector specifically includes:
[0056] Predicting proteins in the target genome contigs according to a gene prediction tool to obtain a plurality of predicted proteins, and aligning all of the predicted proteins with a preset reference plasmid protein cluster to obtain a plurality of alignment results;
[0057] It is determined whether each comparison result is higher than a preset threshold value, and a corresponding plurality of determination results are obtained, and an input feature vector is generated according to all the determination results.
[0058] Specifically, natural language processing technology has achieved rapid development and has gradually expanded to the field of bioinformatics. Natural language models, especially those based on deep learning, have shown great potential in multiple bioinformatics tasks such as protein classification, genome or protein embedding representation, and molecular or protein interaction prediction; these models can effectively capture the complex correlations between elements (i.e., tokens) in sequences, thereby alleviating the common long-term dependency loss problem in biological sequence analysis. This ability makes natural language models an important tool for processing and understanding biological data. However, unlike natural language, biological sequences face unique challenges in the modeling process, such as how to determine the best vocabulary (i.e., token set). In bioinformatics, the elements for constructing token sets can include proteins, amino acids, or k-mers, each of which has its own advantages and disadvantages and will have an important impact on the performance of the model.
[0059] In the present invention, protein clusters (PC) are selected as the token set to make full use of the correlation between proteins to improve the prediction ability of the model. Protein clusters are selected as the token set because they can cover protein clusters from plasmids.
[0060] Therefore, the present invention uses a gene prediction tool (Prodigal, a protein coding gene prediction software tool for bacterial and archaeal genomes) to predict all protein sequences from the reference plasmid, and then uses the DIAMOND (sequence aligner for protein and translated DNA search) tool to perform an all-to-all alignment of all proteins. Based on the alignment results, a corresponding graph is constructed, in which the nodes represent proteins and the edges represent pairs with an e-value lower than 10 -5 Then, the Markov Clustering (MCL) algorithm is applied to cluster these proteins into protein clusters (PCs). Finally, only clusters containing at least two proteins are retained, thereby obtaining a total of a large number of PCs as reference protein clusters for comparison, i.e., reference plasmid protein clusters, and setting a corresponding index value for each protein cluster. The specific processing process of the genomic contigs in the present invention is as follows: Figure 2 As shown, for the genomic contigs, corresponding processing can be performed to obtain protein clusters corresponding to the corresponding proteins.
[0061] Specifically, for the obtained target genome contigs, the present invention first uses the Prodigal tool, i.e., a gene prediction tool, to predict proteins in the target genome contigs, and then uses the DIAMOND tool to compare these predicted proteins with proteins in the reference plasmid protein cluster, thereby obtaining a protein cluster with the smallest e value between the predicted protein and the reference plasmid protein cluster, and calculating the similarity value, i.e., identity value, between the predicted protein and the matching protein in the protein cluster, and the comparison result in the present invention is the corresponding similarity value. Based on the comparison result, if the comparison result is higher than a preset threshold, it is considered that the protein belongs to a cluster in the reference plasmid protein cluster, and the index value of the protein cluster on the comparison is obtained. The corresponding protein in the aligned target genome contigs will be marked as the index value of the PC to which it belongs, thereby converting it into a coded form; for similarity values lower than the preset threshold, an unknown tag (token ID is 1) is used to represent the corresponding predicted protein; and since the corresponding input feature vector has a fixed length, when the number of corresponding predicted proteins is insufficient to fill all input feature vectors, the corresponding vacancy is marked with a mask, i.e., token ID is 0, to fill this part of the content. Afterwards, the input feature vector corresponding to the target genome contig can be obtained, wherein each input feature vector has a corresponding index value, one or more of 0 and 1. Through this method, the present invention can convert the sequence into an input that can be processed by the model, and then use the deep learning model to analyze and predict it.
[0062] Furthermore, the target genome contigs are aligned based on an alignment tool and a pre-constructed alignment library to obtain genome features, specifically including:
[0063] Pre-acquire the protein sequences in the data set library, and construct an alignment library according to the protein sequences in the data set library;
[0064] According to the target tool, gene prediction is performed on the target genome contigs to obtain a plurality of gene prediction proteins, and according to the comparison tool, all the gene prediction proteins are compared with the comparison library to obtain genome features.
[0065] Specifically, in the present invention, genome features are obtained by comparing libraries.
[0066] The process of making the alignment library begins with a predetermined number of protein sequences collected from multiple sources, including genomes, metagenomes, and data from public databases. First, these sequences are clustered using the alignment tool MMseqs2 (a fast and efficient protein alignment and clustering tool specifically designed for large-scale protein data processing), and similar sequences are grouped into a cluster. Subsequently, multiple sequence alignments are performed on each cluster to extract conserved regions and variation patterns, and a feature model is generated using statistical methods (such as hidden Markov models, HMMs) to describe the key features within the cluster. To improve efficiency, the feature model is de-redundanted and multiple non-redundant models are generated. Afterwards, the classification specificity of these feature models is calculated using the reference genome, specific marker proteins are screened out, and the corresponding marker models are retained. The reference genome is a data set used for calibration and testing, which is composed of data from multiple databases. During the screening process, markers with Pielou specificity ≥ 0.4 or maximum SPM value ≥ 0.75 are retained, and different screening criteria are set for chromosome and plasmid markers, and finally the corresponding markers are screened out. In addition, these marker models were functionally and taxonomically annotated by docking with databases such as Pfam and KEGG Orthology, providing rich biological information and classification basis.
[0067] Among them, when calculating the classification specificity of these feature models, the corresponding SPM (Specificity Profile Measure) is introduced as a sensitive statistical indicator for quantitatively estimating the spatial expression pattern of genes in different tissues. SPM is a metric used to quantify the specificity of gene expression in a specific tissue. Its theoretical range is from 0 to 1. The closer the SPM value is to 1, the higher the expression specificity of the gene in the specific tissue.
[0068] The corresponding SPM values can be used to determine the specificity level, which includes CC, CP, CV, PC, PP, PV, VC, VP and VV; wherein each specificity level is obtained by the corresponding chromosome SPM, plasmid SPM and virus SPM.
[0069] According to the corresponding specificity level, the marker is assigned to one of nine levels. Different levels indicate the performance of the marker in the chromosome and plasmid genome. The present invention uses a total of 21 features for marker-based classification, that is, according to the following 21 features, the corresponding genomic features can be obtained. The 21 features include strand_switch_rate: the proportion of genes located on a different chain from the upstream gene; coding_density: the length of all protein coding regions (in base pairs) divided by the total sequence length; no_rbs_freq: the proportion of genes without detectable RBS (ribosome binding site) motifs; sd_canonical_rbs_freq: the proportion of genes predicted to have a canonical Shine-Dalgarno RBS motif; tatata_rbs_freq: the proportion of genes predicted to have a TATATA =Proportion of genes with RBS motifs; cc_marker_freq: number of genes assigned to CC-specific category divided by total number of genes; cp_marker_freq: number of genes assigned to CP-specific category divided by total number of genes; cv_marker_freq: number of genes assigned to CV-specific category divided by total number of genes; pc_marker_freq: number of genes assigned to PC-specific category divided by total number of genes; pp_marker_freq: number of genes assigned to PP-specific category divided by total number of genes; pv_marker_freq: number of genes assigned to PV-specific category divided by total number of genes; vc_marker_freq: number of genes assigned to VC-specific category divided by total number of genes; vp_marker_freq: number of genes assigned to VP-specific category divided by total number of genes; c_marker_freq: total frequency of chromosome markers (CC + CP + CV); p_marker_freq: total frequency of plasmid markers (PC + PP + PV); v_marker_freq: total frequency of viral markers (VC + VP + VV); median_c_spm: median of chromosomal SPM among all annotated genes; median_p_spm: median of plasmid SPM among all annotated genes; v_vs_c_score_logistic: composite score is applied to the range [0-1] using sigmoid function; v_vs_p_score_logistic: composite score is applied to the range [0-1] using sigmoid function; p_vs_c_score_logistic: composite score is applied to the range [0-1] using sigmoid function.
[0070] After obtaining the corresponding alignment library, functional annotation is first performed for the target genome contigs, and gene prediction is performed on the target genome contigs using the prodigal tool, and multiple corresponding gene prediction proteins are generated. Next, MMseqs2 is used to compare the gene prediction proteins with the marker protein models in the alignment library, providing detailed location information, functional annotations, and classification-specific marker information for each gene, so that classification-related marker information can be extracted to complete functional annotation. After that, classification is performed. Based on the position information, functional annotations, and classification-specific marker information generated in the functional annotation stage, a series of classification features are calculated to obtain the corresponding genomic features. The acquisition of these features can be divided into two categories: one is the features directly obtained through MMseqs2 alignment, including the number of various markers, SPM scores, single-copy conserved genes, and marker genes; the other is the features obtained through calculation, which can be further divided into three categories: (1) sequence structure features, such as the coding density obtained by dividing the length of the coding region by the total length, and the chain direction change rate calculated by counting the number of changes in the direction of adjacent gene chains; (2) RBS (ribosome binding site) features, which are obtained by counting the frequency of occurrence of each RBS type predicted by prodigal in the total genes; (3) composite features, including the marker frequency obtained by dividing the number of markers by the total number of genes, the median feature obtained by taking the median of the SPM value, and the scoring feature obtained by calculating the SPM difference of different types of markers and performing logistic transformation. Finally, these features were input into the random forest model as genomic features, and the second classification score was obtained by comprehensively evaluating the classification specificity of chromosomes and plasmids.
[0071] Step S20: input the input feature vector into the improved Transformer model, perform plasmid recognition through the embedding layer, attention layer and fully connected layer in the improved Transformer model, and output a first classification score, wherein the attention layer includes a gated attention unit and a hybrid block attention unit.
[0072] Specifically, in the present invention, the obtained input feature vector is processed by an improved Transformer model, which can process long sequence data while maintaining efficient computing performance. It successfully reduces the computational complexity from quadratic complexity to near linear complexity without significantly reducing the quality of the model by introducing the Gated Attention Unit (GAU) and the Mixed ChunkAttention Unit. This improvement enables the model to show superior adaptability and scalability when processing large-scale data and long sequence tasks. In the present invention, the improved Transformer model is mainly used for plasmid recognition tasks, that is, to determine whether the input genomic contigs are plasmids. By optimizing the processing power and computing efficiency of long sequence data, the model can efficiently extract features from large-scale genetic data and achieve high-precision plasmid classification.
[0073] The corresponding gated attention unit combines the attention mechanism with the gated linear unit (GLU) to form a unified layer structure. This structure not only reduces the complexity of the calculation, but also introduces a more efficient attention gating mechanism. Compared with the traditional MHSA, the attention calculation of GAU is simpler, which can maintain high attention accuracy while reducing the computational cost. The hybrid block attention unit divides the input sequence into multiple non-overlapping chunks. The model performs local secondary attention calculations in each chunk and uses global linear attention between chunks. This method can not only capture local contextual information, but also efficiently handle long-distance dependencies, significantly reducing computational overhead.
[0074] Furthermore, the input feature vector is input into the improved Transformer model, plasmid recognition is performed through the embedding layer, the attention layer and the fully connected layer in the improved Transformer model, and the first classification score is output, which specifically includes:
[0075] Inputting the input feature vector into the improved Transformer model, mapping the input feature vector into a vector representation according to the word embedding in the embedding layer, and adding position information to the vector representation according to the position embedding in the embedding layer to obtain an embedded vector;
[0076] Input the embedding vector into the hybrid block attention unit to divide the embedding vector to obtain multiple blocks, and calculate the local attention and global attention corresponding to each block respectively, and use the gated attention unit to calculate the attention representation of each block according to the local attention and global attention of each block, and splice all the attention representations to obtain the attention result;
[0077] The attention result is input into the fully connected layer for mapping to obtain the first classification score.
[0078] Specifically, in the present invention, the improved Transformer model includes three main components: an embedding layer that converts the encoding vector into a digital matrix, an attention layer that uses gated attention units and hybrid block attention units to learn to tag relevant information, and a fully connected layer as the final prediction classifier to finally obtain the corresponding output.
[0079] Among them, the function of the embedding layer is to convert the input feature vector into an embedding vector with position information, and provide input for the subsequent attention layer. The input is the input feature vector after the overlapping group conversion. In order to effectively capture the relative position relationship between the tags in the input sequence, the embedding layer of the improved Transformer model consists of two parts: word embedding and position embedding. First, word embedding maps the tag corresponding to each protein in the input feature vector to a vector representation of a fixed dimension. Then, position embedding is used to further add position information to the vector representation of these tags to obtain an embedded vector. This process can retain the relative position information of the tags during sequence processing, especially in the context of long sequences. In position embedding, in addition to the traditional absolute position embedding, rotational position embedding is also used to achieve periodic rotation of the position, making the model more robust when dealing with long-distance dependencies. After this process, the input feature vector X is converted into an embedding vector containing position information , the specific embedding process is shown in the following formula:
[0080] ;
[0081] Among them, WordEmbed represents the word embedding operation, and PositionEmbed represents the position embedding operation.
[0082] The role of the attention layer is to capture the dependencies between tags based on the embedding vector, and efficiently fuse local and global attention, and is responsible for processing the embedding vector output by the embedding layer. It is input into the hybrid block attention unit and the gated attention unit for further processing.
[0083] In the attention layer, the hybrid block attention unit achieves efficient attention calculation through a block mechanism, further improving the processing capability of long sequences. In the attention layer, the hybrid block attention unit divides the input embedding vector into several chunks of fixed length, and performs local attention calculation and global attention calculation in each chunk to capture the dependencies between tags in the chunk. The advantage of this hybrid block attention unit mechanism is that it can effectively reduce the computational complexity without sacrificing the ability to capture global and local information.
[0084] When the input embedding vector is divided into g non-overlapping blocks, for each block g, the local attention is calculated, which is expressed as:
[0085] ;
[0086] in, represents the local attention of the i-th block, represents the query vector of the ith block, represents the key vector of the ith block, represents the value vector of the i-th block, relu (rectified linear unit) is the activation function, T represents transpose, and b represents bias;
[0087] For g non-overlapping blocks, the global attention of each block is calculated, where for each block i, the corresponding global attention calculation process is expressed as:
[0088] ;
[0089] in, The i-th block , represents the global query of the ith block, represents the global key of the h-th block, It represents the global value of the h-th block, that is, when calculating the global attention of the ith block, it is necessary to calculate the cumulative result of the product of the global key and the global value of all blocks before the ith block.
[0090] GAU is one of the core modules of the model, which aims to efficiently capture the relationship between tokens in the input sequence. Unlike the multi-head self-attention mechanism in the Transformer model, GAU combines local quadratic attention and global linear attention for each block output by the hybrid block attention unit to improve computational efficiency. Specifically, GAU introduces a gating mechanism to fuse the outputs of local attention and global attention to obtain the corresponding attention representation. The gating mechanism can dynamically adjust the weight of attention according to the characteristics of the input data, thereby enhancing the performance of the model on complex data. Among them, GAU calculates the attention representation of each block based on the local attention and global attention of each block, which is specifically expressed as:
[0091] ;
[0092] in, represents the attention representation of the i-th block, represents the gating vector of the ith block, which is used to control the weights of local attention and global attention, is the output transformation matrix, Represents element-wise multiplication.
[0093] Afterwards, for multiple attention representations, GAU correspondingly concatenates them to obtain the final attention result .
[0094] After the fully connected layer obtains the output of the attention layer, the attention result is first compressed by adaptive average pooling, then standardized by layer normalization, and finally directly mapped to a scalar output by a linear layer to obtain the first classification score for the final classification prediction. This design not only maintains the simplicity of the model structure, but also can effectively process long sequence data.
[0095] Furthermore, before the improved Transformer model is used, it must be trained accordingly. During the training process, the binary cross-entropy loss function (BCE) is used as the objective function, and the parameters of the model are continuously adjusted through the backpropagation algorithm to minimize the prediction error.
[0096] Step S30: input the genome feature into a random forest model, and obtain a second classification score according to the voting result of each decision tree in the random forest model.
[0097] Specifically, the obtained genomic features are processed using a random forest model to obtain the second classification score. Random forests can effectively improve the accuracy, robustness and noise resistance of the model by integrating multiple relatively independent decision trees. Compared with a single decision tree, random forests have significant advantages in processing high-dimensional data and preventing overfitting, and are particularly suitable for data sets with high diversity and complex structures.
[0098] Furthermore, the step of inputting the genome feature into a random forest model and obtaining a second classification score according to the voting result of each decision tree in the random forest model specifically includes:
[0099] Presetting the parameters of the random forest model according to preset parameters;
[0100] Inputting the genome feature into a random forest model, predicting the genome feature according to each decision tree in the random forest model, and obtaining multiple classification results;
[0101] According to the classification results of all decision trees, a second classification score is generated.
[0102] Specifically, in the present invention, the random forest model is used to use genomic features as input variables to predict whether the overlapping group belongs to a plasmid or a chromosome. The random forest introduces corresponding features, which come from the input. By inputting these features into the random forest, the model can perform efficient classification and improve the accuracy and generalization ability of classification through an integrated learning mechanism.
[0103] It should be noted that random forest is an algorithm based on ensemble learning. Its core idea is to improve the accuracy and stability of classification by building multiple relatively independent decision trees and using ensemble methods. Each decision tree makes independent predictions on the input sample, and the final result is determined by voting (classification task) or averaging (regression task) of all trees. Specifically, random forest increases the robustness of the model through a strategy called "Bagging". The basic principle of Bagging is to generate multiple different training sets for training multiple decision tree models by sampling the original training data set multiple times with replacement. Each decision tree will build a classification or regression model based on these different training subsets. In addition, when constructing each node, random forest will randomly select features for splitting. This randomness of feature selection enhances the diversity of the model and further reduces the risk of overfitting. A significant advantage of the random forest model is its feature importance evaluation function. By calculating the information gain or the reduction in the Gini index brought by each feature when the tree node is split, the importance of each feature can be quantified, including helping to understand model decisions and feature selection and optimization.
[0104] In the present invention, parameter optimization is performed in advance to set the parameters of the random forest model, that is, the number of trees (n_estimators) is set to 100, indicating that the model contains 100 decision trees. Increasing the number of trees can improve the stability of the model, but requires more computing resources; the maximum depth (max_depth) is set to 10 to prevent over-growth of the tree from causing overfitting. Controlling the depth of the tree can ensure the good generalization ability of the model on new data; the minimum number of sample splits (min_samples_split) is set to 2 to control the minimum number of samples required for node splitting, ensuring that each node split has at least enough data support; the minimum number of sample leaf nodes (min_samples_leaf) is set to 1, allowing leaf nodes to contain a minimum number of samples, so that the tree can capture more data details, although it may increase the risk of overfitting; the maximum number of features (max_features) is set to "sqrt", indicating that the square root of all features are used in each split. This setting improves the diversity of the model and reduces the impact of interdependence between features.
[0105] After setting the corresponding parameters, the genome feature is input into the random forest model, and the genome feature is predicted according to each decision tree in the random forest model to obtain a second classification score.
[0106] Furthermore, the input feature vector is input into the improved Transformer model, plasmid recognition is performed through the embedding layer, the attention layer and the fully connected layer in the improved Transformer model, and the first classification score is output, which also includes:
[0107] Based on the plasmid database and chromosome database, a preliminary data set was obtained by screening;
[0108] Dividing the preliminary data set into a test set and an original training set;
[0109] Performing a first preprocessing and a second preprocessing on the original training set to obtain a first training set and a second training set;
[0110] The improved Transformer model is trained according to the first training set, the random forest model is trained according to the second training set, and the improved Transformer model and the random forest model are tested according to the test set.
[0111] Specifically, in the present invention, a preliminary data set is obtained by screening the data in the plasmid database and the chromosome database. In one embodiment of the present invention, the PLSDB plasmid database is used, and after deleting plasmids shorter than 1K or longer than 350K, 33,125 plasmids are retained; all bacterial and archaeal genomes with "complete" and "representative" labels are obtained from NCBI (a database website), and a total of 3,775 bacterial genomes and 232 archaeal genomes are obtained. Then, the keywords "plasmid", "mitochondria" and "chloroplasts" are used to filter out non-chromosomal sequences in these genomes, and finally, 4,005 sequences are obtained as reference chromosomes; 33,125 plasmids and 4,005 chromosomes are aggregated to obtain the first preliminary data set; wherein NCBI is the corresponding chromosome database. 20% of the obtained preliminary data set can be divided as a test set, and the rest as an original training set.
[0112] The original training set is subjected to the first preprocessing and the second preprocessing respectively to obtain the first training set and the second training set, wherein the first preprocessing is to obtain the input feature vector and the corresponding plasmid classification of each sample in the original training set, so as to use the first training set to train the improved Transformer model; the second preprocessing is to obtain the genomic features and plasmid classification corresponding to each sample, and generate multiple different training subsets for the obtained genomic features through the bootstrap sampling method (Bootstrap Sampling), each subset contains samples sampled with replacement from the genomic features obtained from the original training set. In this way, each decision tree is trained on a different data subset, thereby improving the diversity of the model and its robustness to noise.
[0113] When the number of training times or the training effect meets the requirements, the training is stopped and the test set is used for testing. When the test results meet the requirements, the corresponding model is used for practical applications.
[0114] Step S40: According to the classification model based on the attention mechanism, the first classification score and the second classification score are respectively aggregated to obtain a first matrix and a second matrix, and a plasmid recognition score is obtained according to the first matrix and the second matrix.
[0115] Specifically, in the present invention, in the task of plasmid identification, a hybrid strategy is adopted, combining the improved Transformer model with the random forest model. Different from the database query method, the improved Transformer model classifies the input overlapping groups; while the random forest model classifier classifies by the presence of specific protein markers, which are highly informative in the classification task. By combining these two methods, the method of the present invention can make full use of the advantages of the two models, thereby improving the accuracy of classification.
[0116] Furthermore, the first classification score and the second classification score are aggregated according to the classification model based on the attention mechanism to obtain a first matrix and a second matrix, and a plasmid recognition score is obtained according to the first matrix and the second matrix, specifically including:
[0117] Aggregating the first classification score and the second classification score according to the attention mechanism-based classification model to obtain the first matrix;
[0118] Obtaining the occurrence frequency of chromosome markers and the occurrence frequency of plasmid markers, and calculating the second matrix based on the classification model based on the attention mechanism;
[0119] The first matrix and the second matrix are multiplied row by row, and column average and normalization are performed to obtain a plasmid recognition score.
[0120] Specifically, in order to further improve the classification performance, the output results of the two classifiers are aggregated into a unified classification result. Since the two classifiers use different and complementary classification methods, by aggregating their outputs, the advantages of the two methods can be more fully utilized, thereby providing more accurate classification results. Specifically, this aggregation is achieved through a classification model based on an attention mechanism. The working principle of the attention mechanism is that when most of the genes in the input sequence can be assigned to specific markers, the weight of the marker branch will increase; on the contrary, when the marker information is relatively scarce, the model will rely more on the sequence classification branch based on the improved Transformer model. In this way, it is possible to rely more on the marker classifier when the marker information is abundant, and rely more on the sequence classifier when the marker information is insufficient, thereby optimizing the overall classification performance.
[0121] The classification model based on the attention mechanism weights the output of each classifier according to the frequency of chromosome and plasmid markers in the input sequence, as shown in the following formula:
[0122] ;
[0123] ;
[0124] Among them, Neural net scores are the first classification scores of the improved Transformer model, Marker-based scores are the second classification scores of the random forest model based on markers, Score matrix is the score matrix composed of two groups of classification results, that is, the first matrix, Attention matrix is the matrix of the attention mechanism, that is, the second matrix, T represents transposition, W1 and W2 represent the weights of the classification model based on the attention mechanism, Cmarker freq is the frequency of chromosome markers, and Pmarker freq is the frequency of plasmid markers, where the frequency of chromosome markers and the frequency of plasmid markers can be obtained from genome features. The first matrix and the second matrix are multiplied row by row, and the final plasmid recognition score is obtained through column average and softmax.
[0125] Furthermore, for the classification model based on the attention mechanism, after training the corresponding improved Transformer model and random forest model, the original training sets of the improved Transformer model and the random forest model are processed to obtain the first training classification score and the second training classification score, all the first training classification scores and the second training classification scores are used as input for training the classification model based on the attention mechanism, and the classification results of the corresponding original training sets are used as corresponding labels. The classification model based on the attention mechanism is trained accordingly, and the weight of the classification model based on the attention mechanism can be obtained.
[0126] Furthermore, if Figure 3 As shown in the figure, when the attention mechanism aggregates the outputs of different branches, as the proportion of genes assigned to labels in the input sequence increases, the contribution of the label branch also increases. This branch aggregation strategy ensures that when most genes in the input sequence can be assigned to useful labels, the classification results of the model can be more accurate; and when the label information is insufficient, the neural network branch can still effectively provide classification information, ensuring the stability and accuracy of the classification.
[0127] In one embodiment of the present invention, an experiment was conducted to download a plasmid dataset from the PLSDB plasmid database. After deleting plasmids shorter than 1K or longer than 350K, 33,125 plasmids were retained. All "complete" and "representative" bacterial and archaeal genomes were downloaded from NCBI, and a total of 3,775 bacterial genomes and 232 archaeal genomes were obtained. Then, the keywords "plasmid", "mitochondria" and "chloroplasts" were used to filter non-chromosomal sequences in these genomes. Finally, 4,005 sequences were obtained as reference chromosomes. 20% of the dataset was divided as a test set, totaling 6,618 plasmids and 800 chromosomes, and the chromosomes were randomly cut to generate a balanced test set containing 6,618 plasmids and 6,618 chromosome fragments. The following performance indicators were used to evaluate the classification effect of the model: Precision, Recall, F1-Score and Accuracy. These indicators measure the prediction performance of the model from different perspectives.
[0128] The present invention is compared with three popular plasmid identification methods, namely PlasClass, Platon and PPR-Meta, where PlasClass, Platon and PPR-Meta are three different identification tools. Figure 4 As shown, it shows that compared with other methods, the present invention has achieved excellent results in multiple indicators. A key factor in achieving higher recognition performance in the present invention is that it combines the advantages of neural networks and labeled classifiers, and effectively integrates the outputs of different classifiers through the attention mechanism. Using this mechanism, the present invention can rely on the powerful modeling ability of neural networks when overlapping group label information is scarce, and fully utilize the specificity of labeled classifiers when label information is abundant, thereby demonstrating excellent classification capabilities in various types of complex data.
[0129] In order to evaluate the effect of each method in short plasmid identification, plasmid contigs with lengths ranging from 1000 to 3999 base pairs were screened from the test set, and the classification performance indicators of different methods on these short plasmids were compared. Figure 5 The specific performance of each method in terms of Recall, F1-score and Accuracy is shown. Figure 5 It can be seen from the analysis that the model of the present invention is significantly better than other models in all indicators, which shows that the model of the present invention has strong generalization ability and recognition ability for positive samples, and is suitable for classification tasks of short fragment overlap groups.
[0130] The present invention obtains a target genome contig, encodes the target genome contig according to a gene prediction tool to obtain an input feature vector, aligns the target genome contig based on an alignment tool and a pre-constructed alignment library to obtain a genome feature; inputs the input feature vector into an improved Transformer model, performs plasmid recognition through an embedding layer, an attention layer and a fully connected layer in the improved Transformer model, and outputs a first classification score, wherein the attention layer includes a gated attention unit and a hybrid block attention unit; inputs the genome feature into a random forest model, and obtains a second classification score according to the voting result of each decision tree in the random forest model; and according to a classification model based on an attention mechanism, aggregates the first classification score and the second classification score respectively to obtain a first matrix and a second matrix, and obtains a plasmid recognition score according to the first matrix and the second matrix. The present invention classifies plasmids by combining an improved Transformer model with a random forest model, wherein the improved Transformer model is optimized for a traditional Transformer structure to capture complex dependencies in overlapping groups of different lengths, so as to improve processing efficiency and performance, and is particularly suitable for overlapping group analysis in a metagenome; at the same time, the random forest model provides a basis for subsequent decision fusion by learning genome marker features, thereby improving the accuracy and stability of the overall classification; the classification results of the improved Transformer model are fused with the classification results of the random forest model, and an attention mechanism is used to achieve weighted fusion of classifier outputs. By combining these two methods, the improved Transformer's ability to understand complex sequences and the random forest's accurate learning of genome features can be fully utilized, thereby effectively improving the accuracy and stability of plasmid recognition.
[0131] Furthermore, if Figure 6 As shown, based on the above-mentioned metagenomic plasmid identification method, the present invention also provides a metagenomic plasmid identification system, wherein the metagenomic plasmid identification system comprises:
[0132] A feature acquisition module 61 is used to acquire a target genome contig, encode the target genome contig according to a gene prediction tool to obtain an input feature vector, and align the target genome contig based on an alignment tool and a pre-built alignment library to obtain a genome feature;
[0133] A first classification score acquisition module 62 is used to input the input feature vector into the improved Transformer model, perform plasmid recognition through the embedding layer, attention layer and fully connected layer in the improved Transformer model, and output a first classification score, wherein the attention layer includes a gated attention unit and a hybrid block attention unit;
[0134] A second classification score acquisition module 63, used to input the genome feature into the random forest model, and obtain a second classification score according to the voting result of each decision tree in the random forest model;
[0135] The result generating module 64 is used to aggregate the first classification score and the second classification score respectively according to the classification model based on the attention mechanism to obtain a first matrix and a second matrix, and obtain a plasmid recognition score according to the first matrix and the second matrix.
[0136] Furthermore, if Figure 7 As shown, based on the above-mentioned metagenomic plasmid identification method and system, the present invention also provides a terminal accordingly, and the terminal includes a processor 10, a memory 20 and a display 30. Figure 7 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0137] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed in the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a metagenomic plasmid recognition program 40 is stored on the memory 20, and the metagenomic plasmid recognition program 40 can be executed by the processor 10, thereby realizing the metagenomic plasmid recognition method in the present invention.
[0138] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the metagenomic plasmid identification method.
[0139] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface.
[0140] In one embodiment, when the processor 10 executes the metagenomic plasmid identification program 40 in the memory 20 , the steps of the above metagenomic plasmid identification method are implemented.
[0141] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a metagenomic plasmid recognition program, and when the metagenomic plasmid recognition program is executed by a processor, the steps of the metagenomic plasmid recognition method as described above are implemented.
[0142] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or terminal including the element.
[0143] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.
[0144] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A method for identifying a metagenomic plasmid, characterized in that: The metagenomic plasmid identification method comprises: Obtaining a target genome contig, encoding the target genome contig according to a gene prediction tool to obtain an input feature vector, and aligning the target genome contig based on an alignment tool and a pre-constructed alignment library to obtain a genome feature; Inputting the input feature vector into an improved Transformer model, performing plasmid recognition through an embedding layer, an attention layer, and a fully connected layer in the improved Transformer model, and outputting a first classification score, wherein the attention layer includes a gated attention unit and a hybrid block attention unit; Inputting the genomic features into a random forest model, and obtaining a second classification score according to the voting results of each decision tree in the random forest model; According to the classification model based on the attention mechanism, the first classification score and the second classification score are respectively aggregated to obtain a first matrix and a second matrix, and a plasmid recognition score is obtained according to the first matrix and the second matrix; The encoding of the target genome contigs according to the gene prediction tool to obtain an input feature vector specifically includes: Predicting proteins in the target genome contigs according to a gene prediction tool to obtain a plurality of predicted proteins, and aligning all of the predicted proteins with a preset reference plasmid protein cluster to obtain a plurality of alignment results; Determine whether each comparison result is higher than a preset threshold, obtain corresponding multiple judgment results, and generate an input feature vector according to all the judgment results; The target genome contigs are aligned based on an alignment tool and a pre-constructed alignment library to obtain genome features, specifically including: Pre-acquire the protein sequences in the data set library, and construct an alignment library according to the protein sequences in the data set library; Perform gene prediction on the target genome contig according to the target tool to obtain multiple gene prediction proteins, and compare all the gene prediction proteins with the comparison library according to the comparison tool to obtain genome features; The method further comprises: aggregating the first classification score and the second classification score according to the classification model based on the attention mechanism to obtain a first matrix and a second matrix, and obtaining a plasmid recognition score according to the first matrix and the second matrix. Aggregating the first classification score and the second classification score according to the attention mechanism-based classification model to obtain the first matrix; Obtaining the occurrence frequency of chromosome markers and the occurrence frequency of plasmid markers, and calculating the second matrix based on the classification model based on the attention mechanism; The first matrix and the second matrix are multiplied row by row, and column average and normalization are performed to obtain a plasmid recognition score.
2. The method for identifying a metagenomic plasmid according to claim 1, characterized in that: The step of inputting the input feature vector into the improved Transformer model, performing plasmid recognition through the embedding layer, the attention layer and the fully connected layer in the improved Transformer model, and outputting a first classification score specifically includes: Inputting the input feature vector into the improved Transformer model, mapping the input feature vector into a vector representation according to the word embedding in the embedding layer, and adding position information to the vector representation according to the position embedding in the embedding layer to obtain an embedded vector; Input the embedding vector into the hybrid block attention unit to divide the embedding vector to obtain multiple blocks, and calculate the local attention and global attention corresponding to each block respectively, and use the gated attention unit to calculate the attention representation of each block according to the local attention and global attention of each block, and splice all the attention representations to obtain the attention result; The attention result is input into the fully connected layer for mapping to obtain the first classification score.
3. The method for identifying a metagenomic plasmid according to claim 1, characterized in that: The step of inputting the genome feature into the random forest model and obtaining a second classification score according to the voting result of each decision tree in the random forest model specifically includes: Presetting the parameters of the random forest model according to preset parameters; Inputting the genome feature into a random forest model, predicting the genome feature according to each decision tree in the random forest model, and obtaining multiple classification results; According to the classification results of all decision trees, a second classification score is generated.
4. The method for identifying a metagenomic plasmid according to claim 1, characterized in that: The step of inputting the input feature vector into the improved Transformer model, performing plasmid recognition through an embedding layer, an attention layer, and a fully connected layer in the improved Transformer model, and outputting a first classification score, further includes: Based on the plasmid database and the chromosome database, a first preliminary data set is screened to obtain a preliminary data set; the first preliminary data set is screened to obtain a preliminary data set; Dividing the preliminary data set into a test set and an original training set; Performing a first preprocessing and a second preprocessing on the original training set to obtain a first training set and a second training set; The improved Transformer model is trained according to the first training set, the random forest model is trained according to the second training set, and the improved Transformer model and the random forest model are tested according to the test set.
5. A metagenomic plasmid recognition system, characterized in that: The metagenomic plasmid recognition system comprises: A feature acquisition module is used to obtain a target genome contig, encode the target genome contig according to a gene prediction tool to obtain an input feature vector, and align the target genome contig based on an alignment tool and a pre-built alignment library to obtain genome features; A first classification score acquisition module, used for inputting the input feature vector into an improved Transformer model, performing plasmid recognition through an embedding layer, an attention layer, and a fully connected layer in the improved Transformer model, and outputting a first classification score, wherein the attention layer includes a gated attention unit and a hybrid block attention unit; A second classification score acquisition module, used to input the genome feature into a random forest model, and obtain a second classification score according to the voting result of each decision tree in the random forest model; A result generation module, used for respectively aggregating the first classification score and the second classification score according to a classification model based on an attention mechanism to obtain a first matrix and a second matrix, and obtaining a plasmid recognition score according to the first matrix and the second matrix; The encoding of the target genome contigs according to the gene prediction tool to obtain an input feature vector specifically includes: Predicting proteins in the target genome contigs according to a gene prediction tool to obtain a plurality of predicted proteins, and aligning all of the predicted proteins with a preset reference plasmid protein cluster to obtain a plurality of alignment results; Determine whether each comparison result is higher than a preset threshold, obtain corresponding multiple judgment results, and generate an input feature vector according to all the judgment results; The target genome contigs are aligned based on an alignment tool and a pre-constructed alignment library to obtain genome features, specifically including: Pre-acquire the protein sequences in the data set library, and construct an alignment library according to the protein sequences in the data set library; Perform gene prediction on the target genome contig according to the target tool to obtain multiple gene prediction proteins, and compare all the gene prediction proteins with the comparison library according to the comparison tool to obtain genome features; The method further comprises: aggregating the first classification score and the second classification score according to the classification model based on the attention mechanism to obtain a first matrix and a second matrix, and obtaining a plasmid recognition score according to the first matrix and the second matrix. Aggregating the first classification score and the second classification score according to the attention mechanism-based classification model to obtain the first matrix; Obtaining the occurrence frequency of chromosome markers and the occurrence frequency of plasmid markers, and calculating the second matrix based on the classification model based on the attention mechanism; The first matrix and the second matrix are multiplied row by row, and column average and normalization are performed to obtain a plasmid recognition score.
6. A terminal, characterized in that: The terminal comprises: a memory, a processor, and a metagenomic plasmid identification program stored in the memory and executable on the processor, wherein the metagenomic plasmid identification program, when executed by the processor, implements the steps of the metagenomic plasmid identification method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a metagenomic plasmid identification program, and when the metagenomic plasmid identification program is executed by a processor, the steps of the metagenomic plasmid identification method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Anomaly-Transform photovoltaic power generation anomaly detection method
CN117972419A
Method for identifying drug-resistant pathogenic microorganisms based on deep learning technology
CN118155720A