A method for metagenomic species reconstruction based on pre-training and deep clustering
Patent Information
- Application Number
- CN202211069609.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-08-31
AI Technical Summary
[0005]鉴于上述现有技术的不足,本发明的目的在于提供一种基于预训练和深度聚类的宏基因组物种重建方法,旨在解决因宏基因组DNA片段数据长度较短、特征较少的复杂性以及物种间相对丰度不均匀等特性带来的宏基因组DNA片段聚类与拼接难的问题
[0039]有益效果:本发明提供了一种基于预训练和深度聚类的宏基因组物种重建方法。基于宏基因组DNA片段的深度LSTM自编码器联合聚类方法,设计了基于图卷积神经网络联合Focal Loss损失函数的词嵌入特征提取模型以及基于LSTM自编码器联合改进的FCM算法的深度聚类模型。本发明将所选取的宏基因组数据集中每条DNA片段用滑窗法以k-mer的形式编码,应用两层GCN网络作为词嵌入模型训练高维特征,将宏基因组中没有比对标签的片段用训练后得到的词嵌入结果表示,作为后续深度聚类的数据样本。本发明构建了一种深度LSTM自编码器联合聚类算法模型,对上述处理后的数据集进行聚类分箱,将深度学习与聚类结合在一起,重构误差与聚类误差同步优化,相比于其他算法,可以进一步提升二者性能,同时计算量也较小。最后,利用拼接软件对聚类后的片段按照类别进行组装,并用check-m工具计算组装后的重叠群(contigs)的完整度与污染度。在用户使用时,只需要针对所选取的数据集的大小及序列长度对整个模型的参数进行调整,重新运行模型即可得到聚类结果,大大提高了准确度与便利性。经过真实数据集的验证以及与现有的其他多种方法的聚类结果进行比较,结果证实,本发明能够得到更加优秀的聚类结果。不仅如此,在发现未知物种的层面上,本发明所发现的未知物种相较于其他方法而言,完整度更高,污染度更低。
Smart Images

Figure CN115579068B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics analysis, and in particular to a metagenomic species reconstruction method based on pre-training and deep clustering. Background Technology
[0002] Microorganisms, including bacteria, fungi, and some small protozoa, are tiny organisms and constitute the largest, most numerous, and most widely distributed group of organisms on Earth. The interactions between microorganisms are crucial to human health and disease. For a long time, research on microorganisms has primarily focused on culturable single species. In fact, culturable microorganisms in the environment represent a very small fraction of the total number of microorganisms in nature; studies have confirmed that 1 gram of soil contains approximately 103 culturable microorganisms. 7 Of the 100 bacterial species, only about 1% to 10% can be cultured, and the majority of microbial species and functions have not yet been discovered.
[0003] Today, next-generation sequencing (NGS), or high-throughput sequencing, can analyze the genomes of all microorganisms in environmental samples in detail, not just those that have been cultured beforehand. Metagenomics, utilizing next-generation sequencing technology, can obtain most of the genetic material in the environment without laboratory culturing. Microbial community clustering analysis is one of the most challenging tasks in metagenomic research. Unlike traditional sequencing methods, the raw data obtained by second-generation sequencing technology is generally a large number of short DNA fragments from various microorganisms with uneven relative abundance among species. Since the metagenomic DNA fragments of many species in the environment are intertwined, efficient clustering and accurate binning are key steps in analyzing detailed information and corresponding functions of each species. Moreover, each DNA fragment contains a small number of bases and features, and the complexity of the sequence and the relative imbalance among species increase the difficulty of subsequent clustering and splicing. The complexity of DNA fragments seriously hinders the accurate study and efficient application of metagenomic sequences, increasing the difficulty of constructing and optimizing clustering models. Therefore, how to efficiently and accurately cluster metagenomic DNA fragments is a key problem that needs to be solved. Most research methods first assemble the raw DNA fragments and then analyze the resulting contigs. Contigs are characterized by their longer length and greater number of features, and their dataset complexity and imbalance are lower than those of DNA fragments. Currently, although some researchers have proposed clustering and splicing algorithms directly targeting metagenomic DNA fragments extracted from the environment, their performance still needs further improvement.
[0004] Therefore, existing technologies still need improvement. Summary of the Invention
[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a metagenomic species reconstruction method based on pre-training and deep clustering, which aims to solve the problem of difficult clustering and splicing of metagenomic DNA fragments caused by the short length, limited features, and uneven relative abundance among species of metagenomic DNA fragment data.
[0006] The technical solution of the present invention is as follows:
[0007] This invention provides a metagenomic species reconstruction method based on pre-training and deep clustering, comprising the following steps:
[0008] The first step, feature extraction, includes:
[0009] Raw metagenomic datasets of different environmental microorganisms were extracted and preprocessed. The raw datasets included DNA sequence features of different species.
[0010] Build a word embedding model for the preprocessed dataset;
[0011] Construct the model error function;
[0012] The word embedding model is trained and its parameters are adjusted using the constructed model error function;
[0013] Save the output feature vector matrix;
[0014] The second step, deep clustering, includes:
[0015] A deep LSTM autoencoder joint clustering model is constructed. The LSTM autoencoder consists of two parts: an encoder and a decoder. The encoder learns features from the input feature vector time series data, and the decoder reconstructs the data features using the current hidden layer state and network parameters.
[0016] The clustering loss function of the model is constructed by combining an LSTM autoencoder with an FCM algorithm clustering model.
[0017] Input the metagenomic dataset of the microorganism to be tested, and calculate the overall loss error of the model using the model reconstruction loss function and the clustering loss function;
[0018] Calculate and analyze the clustering performance metrics of the model;
[0019] Adjust the model parameters to obtain the optimal clustering performance of the model;
[0020] The unknown species in the dataset to be tested are clustered and assembled, and the integrity and contamination of the assembled contiguous groups are analyzed by the software, and the clustering results of the unknown species are output.
[0021] The metagenomic species reconstruction method based on pre-training and deep clustering, wherein the preprocessing of the original dataset specifically includes the following steps:
[0022] a) Download the metagenomic sequence dataset of the microbial community;
[0023] b) Based on the quality value information of each DNA fragment stored in the dataset, optimize the data using quality control software tools and filter out low-quality sequences;
[0024] c) Replace the N base in the remaining sequence after filtering in step b) with one of A, G, C, or T;
[0025] d) Use the BLAST tool to compare the tags of all metagenomic sequences processed in step c), and distinguish different samples based on primers and index sequences;
[0026] e) Based on the BLAST alignment results, filter out DNA fragment pairs that do not belong to the same species, and use the filtered sequences as input to the word embedding model.
[0027] The metagenomic species reconstruction method based on pre-training and deep clustering, wherein the low-quality sequence includes sequences containing more than 5 consecutive N bases.
[0028] The metagenomic species reconstruction method based on pre-training and deep clustering, wherein the construction of the word embedding model specifically includes the following steps:
[0029] a) The pre-processed DNA fragment sequence was cut using the sliding window method to convert the sequence into overlapping fixed-length k-mer sequences;
[0030] b) Construct a topology graph from the segmented sequence database;
[0031] c) Establish the edge between two k-mer nodes using the co-occurrence information of the k-mer sequence;
[0032] d) Input the constructed topology graph into a two-layer graph convolutional neural network to construct a word embedding model.
[0033] The metagenomic species reconstruction method based on pre-training and deep clustering, wherein the step of constructing the model error function uses Focal Loss as the error function of the feature extraction model.
[0034] The metagenomic species reconstruction method based on pre-training and deep clustering includes an encoder and a decoder, each comprising two LSTM layers and two fully connected layers, with a clustering layer constructed between the encoder and decoder.
[0035] The metagenomic species reconstruction method based on pre-training and deep clustering is described in which the neurons of the LSTM are composed of several recursively connected memory blocks, each memory block including one memory unit and three logical units, the logical units including a forget gate, an input gate and an output gate.
[0036] The metagenomic species reconstruction method based on pre-training and deep clustering, wherein the clustering performance index of the model calculated and analyzed in the step includes precision, recall and adjusted Land coefficient.
[0037] A storage medium storing one or more programs that can be executed by one or more processors to implement the steps in the metagenomic species reconstruction method based on pre-training and deep clustering as described above.
[0038] A terminal device includes a processor adapted to implement various instructions; and a storage medium adapted to store a plurality of instructions, said instructions being loaded by the processor and executed as described above in the metagenomic species reconstruction method based on pre-training and deep clustering.
[0039] Beneficial Effects: This invention provides a metagenomic species reconstruction method based on pre-training and deep clustering. A deep LSTM autoencoder joint clustering method based on metagenomic DNA fragments is proposed, employing a word embedding feature extraction model based on a graph convolutional neural network combined with a Focal Loss loss function, and a deep clustering model based on an improved FCM algorithm using LSTM autoencoders. This invention encodes each DNA fragment in the selected metagenomic dataset using a sliding window method in k-mer form, applies a two-layer GCN network as a word embedding model to train high-dimensional features, and represents fragments without alignment labels in the metagenomic dataset using the word embedding results obtained after training, serving as data samples for subsequent deep clustering. This invention constructs a deep LSTM autoencoder joint clustering algorithm model to cluster and bin the processed dataset, combining deep learning with clustering, simultaneously optimizing reconstruction and clustering errors. Compared to other algorithms, this further improves the performance of both while reducing computational cost. Finally, the clustered fragments are assembled according to categories using splicing software, and the completeness and contamination of the assembled contigs are calculated using the check-m tool. When using this method, users only need to adjust the parameters of the entire model according to the size of the selected dataset and the sequence length, and then rerun the model to obtain the clustering results, greatly improving accuracy and convenience. Validation with real datasets and comparison with the clustering results of other existing methods confirms that this invention can obtain superior clustering results. Furthermore, in discovering unknown species, the unknown species discovered by this invention have higher completeness and lower contamination compared to other methods. Attached Figure Description
[0040] Figure 1 This is a flowchart of the overall model of a metagenomic species reconstruction method based on pre-training and deep clustering in an embodiment of the present invention.
[0041] Figure 2 This is a schematic diagram of the DNA-FLGCN model structure in an embodiment of the present invention.
[0042] Figure 3 This is a flowchart illustrating the process of constructing a word embedding feature vector matrix using the sliding window method in an embodiment of the present invention.
[0043] Figure 4 This is a schematic diagram of the LSTM network neuron structure in an embodiment of the present invention.
[0044] Figure 5 This is a schematic diagram of the processing method of the LSTM autoencoder decoder in an embodiment of the present invention (the semantic vector is only used as the initial state for computation).
[0045] Figure 6This is a schematic diagram of the processing method of the LSTM autoencoder decoder in an embodiment of the present invention (the semantic vector participates in the operation at all time points).
[0046] Figure 7 This is a schematic diagram of the joint clustering model structure of the deep LSTM autoencoder in an embodiment of the present invention.
[0047] Figure 8 This is a comparison of the SRR492190 test clustering results under different methods in the embodiments of the present invention.
[0048] Figure 9 This invention provides a comparison of the integrity of contiguous groups after clustering and splicing of SRR492190 unknown species fragments using different methods in this embodiment.
[0049] Figure 10 This invention presents a comparison of contamination levels of contigs after clustering and splicing SRR492190 unknown species fragments using different methods in this embodiment. Detailed Implementation
[0050] This invention provides a metagenomic species reconstruction method based on pre-training and deep clustering. To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0051] This invention provides a metagenomic species reconstruction method based on pre-training and deep clustering, comprising the following steps:
[0052] S100, Feature extraction steps;
[0053] S200, deep clustering steps.
[0054] Because many metagenomic sequences have low similarity to sequences in the NCBI official nucleic acid sequence database, it is impossible to align these sequences with tags. To date, most methods are based on sequence fragments with alignment tags, without detailed analysis of sequences without alignment tags. This invention's deep LSTM autoencoder-based joint clustering method for metagenomic DNA fragments differs significantly from current methods. It designs a word embedding feature extraction model based on a graph convolutional neural network combined with a Focal Loss loss function, and a deep clustering model based on an improved FCM algorithm using an LSTM autoencoder. Figure 1The diagram shows the overall model flowchart of the metagenomic species reconstruction method based on pre-training and deep clustering described in this embodiment of the invention. This embodiment directly extracts raw metagenomic datasets (reads) obtained from high-throughput sequencing technology in the environment, conducting detailed research on such datasets with few features, short lengths, and highly uneven relative abundance among species. Features of the dataset are extracted using a pre-constructed word embedding model to obtain a feature vector matrix. A deep autoencoder joint clustering model, which performs clustering while training, is constructed to cluster the dataset and conduct detailed analysis, effectively reducing the model's time complexity while improving clustering performance. This embodiment mainly analyzes sequences with unknown labels in the selected metagenomic dataset, representing them with feature vectors obtained after training the word embedding model. These vectors are input into the pre-constructed deep clustering model and parameter tuning is performed to obtain better clustering performance indicators. The clustered results are then spliced together, and the integrity and contamination of the spliced contigs are analyzed.
[0055] In some implementations, S100 specifically includes the following steps:
[0056] S101. Extract raw metagenomic datasets of different environmental microorganisms and preprocess the raw datasets, wherein the raw datasets include DNA sequence features of different species;
[0057] S102. Construct a word embedding model for the preprocessed dataset;
[0058] S103. Construct the model error function;
[0059] S104. Train the word embedding model and adjust its parameters using the constructed model error function;
[0060] S105. Save the output feature vector matrix.
[0061] In some embodiments, step S101, which involves preprocessing the original dataset, specifically includes the following steps:
[0062] a) Download the metagenomic sequence dataset of the microbial community;
[0063] b) Based on the quality value information of each DNA fragment stored in the dataset, optimize the data using quality control software tools and filter out low-quality sequences;
[0064] c) Replace the N base in the remaining sequence after filtering in step b) with one of A, G, C, or T;
[0065] d) Use the BLAST tool to compare the tags of all metagenomic sequences processed in step c), and distinguish different samples based on primers and index sequences;
[0066] e) Based on the BLAST alignment results, filter out DNA fragment pairs that do not belong to the same species, and use the filtered sequences as input to the word embedding model.
[0067] Metagenomics contains a large number of DNA sequences from different species. Therefore, extracting the characteristics of each DNA sequence and classifying it according to its species is an important step in metagenomic fragment sequence analysis. The sequence characteristics of DNA refer to the arrangement of the four nucleotides A, T, G, and C. If two DNA sequences have similar combinations, the two DNA sequences are highly similar and are more likely to originate from the same species.
[0068] Specifically, in this embodiment of the invention, metagenomic sequence datasets of different microbial communities, such as the human gut microbiota metagenomic sequence dataset, are downloaded from the NCBI website. As a specific example, the SRR492190 dataset is used. The paired-end sequence data obtained by high-throughput sequencing technology are stored in FASTQ or FASTA format. Taking FASTQ format as an example, the downloaded high-throughput sequencing sequences are divided into two files, fq1 and fq2. The sequences in fq1 are referred to as read1, and the sequences in fq2 are referred to as read2. The corresponding read1 and read2 sequences are two inverse complementary strands, each fragment being 100 bp in length.
[0069] Specifically, in this embodiment of the invention, optimization is performed using the quality control software FastP based on the quality value information of each DNA fragment (read) stored in the file. Low-quality sequences (e.g., sequences with a large number of consecutive N bases) are filtered out using FastP and other data quality control software. More specifically, this embodiment of the invention stipulates that when a sequence contains more than 5 consecutive N bases, the sequence is also considered a low-quality sequence and filtered out.
[0070] Specifically, the base N in the remaining filtered sequences is randomly replaced with any one of A, G, C, or T. If the downloaded high-throughput sequencing file contains the base N, it means that the position of the base N in that sequence has not been clearly sequenced. Therefore, in this embodiment of the invention, it is randomly replaced with any one of A, G, C, or T.
[0071] Specifically, the BLAST tool is used to align the tags of all processed metagenomic sequences, and different samples are distinguished based on primers and index sequences. Since the DNA fragments in the fq1 and fq2 files are paired-end sequencing results, each pair of reverse complementary DNA fragments should belong to the same species. The BLAST alignment results are checked, and DNA fragment pairs that do not belong to the same species are filtered out. The filtered sequences in the fq1 file are then used as input to the word embedding model.
[0072] In some implementations, step S102, constructing the word embedding model, specifically includes the following steps:
[0073] a) The pre-processed DNA fragment sequence was cut using the sliding window method to convert the sequence into overlapping fixed-length k-mer sequences;
[0074] b) Construct a topology graph from the segmented sequence database;
[0075] c) Establish the edge between two k-mer nodes using the co-occurrence information of the k-mer sequence;
[0076] d) Input the constructed topology graph into a two-layer graph convolutional neural network to construct a word embedding model.
[0077] This invention establishes a word embedding model based on DNA fragments (reads) (DNA-FLGCN), which is an improvement and innovation on the basis of Text GCN.
[0078] Specifically, the sliding window method is used to segment the sequences in the preprocessed fq1 file, transforming them into overlapping fixed-length k-mer sequences. Assuming each sequence has a length of n, there are (n-k+1) windows in the sequence. For example, ATGCCGTAT is transformed into five 5-mer sequences: {ATGCC,TGCCG,GCCGT,CCGTA,CGTAT}.
[0079] The DNA-FLGCN model structure proposed in this embodiment of the invention is shown in the figure below. Figure 2 As shown.
[0080] Specifically, the segmented sequence database is constructed into a topological graph, whose nodes consist of the entire document and all k-mer sequences, explicitly modeling global word co-occurrence to facilitate subsequent training with a graph convolutional neural network (GCN).
[0081] Specifically, edges between two k-mer nodes are established using the co-occurrence information of the k-mer sequences. The edge between each k-mer node and a document node is established using the frequency of that k-mer and the frequency of the document containing that k-mer. The weights of the edges between two k-mer nodes are defined as follows:
[0082]
[0083] Where TF-IDF(i,j) represents the weight of the edge between the document node and each k-mer sequence node, and PMI(i,j) is the weight between two k-mer sequence nodes. When PMI(i,j) is positive, it indicates that sequence i and sequence j have a strong semantic association; when PMI(i,j) is negative, it indicates that words i and j have a low semantic association.
[0084] Specifically, the constructed topology graph is input into a two-layer Graph Convolutional Neural Network (GCN) to build a word embedding model. The output size of the first-layer GCN is the dimension of the word embedding vectors after training, and the output of the second-layer GCN is the total number of label categories. In the constructed word embedding model, the test accuracy increases with the increase of the window size, but a window that is too small cannot fully generate global word co-occurrence information, while a window that is too large may add edges between k-mer sequences with low correlation. At the same time, the word embedding dimension is also an important factor affecting the entire model. A low-dimensional word embedding dimension cannot effectively propagate complete label information to the constructed graph, while a word embedding dimension that is too high may reduce classification performance and consume more training time. Preferably, in this embodiment of the invention, the window size is set to 10 and the word embedding dimension is set to 10. Model Selection The function is used as the activation function. The Dropout function is added to prevent the model from overfitting. The Dropout function is used to adjust the relationship between the model parameters and the number of samples. The parameter of the Dropout function is set to 0.5.
[0085] In some implementations, in step S103, Focal Loss is used as the error function of the classification model.
[0086] In real-world scenarios, most metagenomic datasets are extremely imbalanced, with significant differences in relative abundance among species. Using only traditional cross-entropy as the loss function is insufficient for accurate classification; therefore, further improvements to the loss function are necessary. Weights are assigned to the loss function based on the ease of sample differentiation. Samples with classification confidence scores close to 1 or 0 are considered easily differentiated, while the rest are considered difficult to differentiate. Considering various factors, this embodiment of the invention employs Focal Loss as the error function for the classification model.
[0087] Specifically, the formula for calculating the Focal Loss error function is as follows:
[0088] FL(p t )=-α t (1-p t ) γ log(p t (2)
[0089] In the formula, pt α represents the probability distribution of the current sample's category within the total dataset. In multi-class classification problems, this is the probability output after passing the predicted label obtained from training a graph convolutional neural network through a softmax layer. t α and γ are two hyperparameters. t Its main function is to solve the imbalance problem between positive and negative samples (α). t ∈[0,1]), the more classes the sample belongs to, the higher α becomes. t The smaller the value of γ, the lower the training loss for high-abundance species, allowing subsequent classification and optimization to focus on low-abundance and very low-abundance species. γ addresses the imbalance between easy and difficult samples, effectively reducing the loss from negative samples. The larger γ is, the lower the Focal Loss value can be for easier samples with higher probabilities. Preferably, a value of γ of 2 yields the best results.
[0090] The DNA-FLGCN method proposed in this invention encodes each k-mer sequence, that is, it uses a feature vector to represent a k-mer, and each element of the feature vector is a quantitative description of a certain feature of that k-mer. The choice of the k value in the k-mer affects subsequent genome clustering and assembly. Smaller k-mers will reduce the edges of the constructed graph, increase the chance of all k-mers overlapping, and reduce the memory required to store DNA fragments, but will face the risk of multiple vertices leading to a single k-mer sequence, and will lose some information in the dataset; larger k values are helpful for genome construction, but will lead to the disjointing of DNA fragments, and will generate a large number of short contigs during the assembly process.
[0091] Preferably, in this embodiment of the invention, the value of k is 13, such as... Figure 3 As shown, a word embedding feature vector matrix is constructed using the sliding window method, where L is the DNA fragment sequence length. Taking the SRR492190 dataset used in this embodiment as an example, the value of L is 100, and each sequence can be segmented into 88 13-mer sequences using the sliding window method. The length of the feature vector depends on the specific situation; the more feature elements, the more comprehensive the representation of the word. Preferably, the feature vector length is set to 10, and each k-mer sequence in the sequence can be represented by a corresponding 88×10 vector feature. The word vectors generated by this method have good semantic properties, and they are the result of the application of neural networks in the field of natural language processing. By using deep learning methods to obtain the distribution representation of words, it can be used for natural language processing tasks such as text classification, sentiment computing, and word construction. The advantage of this feature representation is that it can clearly show the similarity between different k-mer sequences, and the number of sequence features after processing is significantly increased, which is more conducive to subsequent autoencoder deep learning and clustering.
[0092] In some implementations, step S104, which involves training the word embedding model using the constructed model error function and adjusting its parameters, specifically includes: using the preprocessed dataset from the fq1 file as input to the word embedding model, randomly selecting 75% of the sequences as the training set and the remaining 25% as the test set, and inputting this into the constructed word embedding model. When training the word embedding model, because it is necessary to construct the relevance of each sequence in detail, all training set data must be input into the model at once in one iteration, and the precision, recall, and adjusted Rand Index (ARI) of the word embedding training model after each iteration are calculated using the test set. Taking the dataset SRR492190 used in this embodiment as an example, this invention preferably sets the learning rate lr_rate of the adaptive moment estimation (Adam) optimizer to 0.15 and the training iteration period to 600. For different datasets, better performance metrics need to be obtained by adjusting the adaptive learning rate and the training iteration period.
[0093] In some implementations, step S105 involves saving the feature vector matrix output by the word embedding model to provide a basis for the next step of detailed analysis of deep clustering of metagenomic DNA fragments.
[0094] In some implementations, S200 specifically includes the following steps:
[0095] S201. Construct a deep LSTM autoencoder joint clustering model. The LSTM autoencoder includes an encoder and a decoder. The encoder learns features from the input feature vector time series data, and the decoder reconstructs the data features using the current hidden layer state and network parameters.
[0096] S202. Construct the model clustering loss function by combining the LSTM autoencoder with the FCM algorithm clustering model;
[0097] S203. Input the metagenomic dataset of the microorganism to be tested, and calculate the overall loss error of the model using the model reconstruction loss function and the clustering loss function.
[0098] S204. Calculate and analyze the clustering performance indicators of the model;
[0099] S205. Adjust the model parameters to obtain the optimal clustering performance of the model;
[0100] S206. Cluster and assemble the unknown species in the dataset to be tested, and use software to analyze the integrity and contamination of the assembled contiguous groups, and output the clustering results of the unknown species.
[0101] In some implementations, in step S201, the encoder and decoder each include two LSTM layers and two fully connected layers, with a clustering layer constructed between the encoder and decoder.
[0102] In some specific implementations, LSTM neurons consist of several recursively connected memory blocks. Specifically, each memory block includes a memory unit and three logic units, namely a forget gate, an input gate, and an output gate.
[0103] The neuronal structure of Long Short-Term Memory (LSTM) networks is as follows: Figure 4 As shown, it consists of a series of recursively connected memory blocks. Each memory block includes a memory unit and three logical units, which control the transmission of information. These three logical units are named the forget gate, input gate, and output gate. Each light-colored circle in the diagram represents an operation performed on a vector, and each rectangle represents a neural network layer. The characters inside the rectangle represent the activation function used by the corresponding neural network. The lines running through the units represent the hidden states of neurons, and σ represents the sigmoid function. As x approaches negative infinity, f(x) approaches 0, becoming a nonlinear function acting on the neuron. t This represents the neuron's memory after time t, encompassing the neural network's "summary" of all input information up to time t+1. t-1 This represents the neuron's memory from the previous moment. The memory unit represents the neuron's state memory; the input gate and output gate are used to receive and output parameters, respectively. The input gate uses the tanh function to extract valid information, and... This means that the sigmoid function is used to control which memories are placed into the cell state, using i t This means that each component is rated, and the higher the rating, the more memory input is added to the LSTM memory unit. The output gate combines the current input value and the previous output value into a vector, extracts information using the sigmoid function, then compresses the current memory unit to the interval (-1, 1) using the tanh function, and finally multiplies the processed unit state with the result of the sigmoid function to obtain the output of the LSTM neural network at time t. The forget gate controls whether to retain the historical information stored in the current hidden layer node. The LSTM neural network decides which part of the previous memory to forget based on the new input and the previous output. t and the output h of the previous step t-1 Integrate into a single vector, multiply it point-to-point with the current unit state through a sigmoid layer, and use o tThis means that the sigmoid function compresses the input into the (0,1) interval. Therefore, if a component in the integrated vector becomes 0 after passing through the sigmoid layer, the corresponding component will also become 0 after the corresponding unit state is multiplied bitwise, meaning that the information on that component is "forgotten." Conversely, it means that the complete memory is retained. t This is the output at the current moment. Long Short-Term Memory (LSTM) networks can store important information for a long time and can dynamically adjust according to the input, where f... t This is the output vector of the sigmoid neural layer. Long Short-Term Memory (LSTM) networks have the ability to read, reset, and update historical information. To minimize training error, gradient descent can be used to modify the network weights at each training iteration.
[0104] In some implementations, in step S201, the LSTM autoencoder includes two parts: an encoder and a decoder. The encoder performs feature learning on the input feature vector time series data, and the decoder reconstructs the data features using the current hidden layer state and network parameters.
[0105] The Seq2seq model belongs to the LSTM autoencoder architecture. Its main idea is to utilize two recurrent neural networks (RNNs), one as an encoder and the other as a decoder. The encoder is responsible for reducing the dimensionality of the input sequence and outputting a vector of a specified length, while the decoder generates a specified sequence based on the semantic vector. The decoder can process data in two ways, such as... Figure 5 and Figure 6 As shown in the figure, h0, h1, h2, h3 are the encoder's state vectors, h0', h1', h2', h3' are the decoder's state vectors, x1, x2, x3, x4 are the input vectors, and y1, y2, y3 are the output vectors. As can be seen from the figure, one approach is that the semantic vector C only participates in the computation as the initial state, and the subsequent computations of the decoder are independent of C; the other approach is that the semantic vector C participates in the computation of all time points in the sequence.
[0106] The advantage of LSTM autoencoders lies in their ability to solve the gradient vanishing and gradient exploding problems that occur during the training of DNA fragments, and to effectively handle variable-length sequences without affecting model training. This model allows specifying the effective sequence length for each sample in each batch. Within the effective length, the state and output values of a sequence remain unchanged, but the state values of the portion exceeding the effective length will not change, while the output values will become zero vectors. The calculation of output and state values depends not only on the current input value but also on the state value from the previous time step. Internally, a mask matrix is used to mark the effective and invalid parts. Invalid parts do not cause parameter updates during backpropagation and do not affect the subsequent calculation of the error function.
[0107] Assume the input data sequence is X i X i ={x i1 ,x i2 ,L,x in}, for each sequence X i At this point, the hidden layer state vector corresponding to the i-th column of t∈{1,2,L,n} is:
[0108]
[0109] In the formula Let be the output state vector of the i-th coding unit at time (t-1). Let X be the input vector, and W and R be m×d and m×m sparse weight matrices, respectively; the function k(·) is usually set as the tanh activation function, which will... i Each column vector in the code is input into the encoder, and the output is:
[0110]
[0111] In the formula Let be the output of the i-th coding unit at time t, and α be the parameters of the encoder. Also using the tanh activation function. Pooling is necessary during the training of high-dimensional data, primarily for dimensionality reduction. Pooling methods include random pooling, mean pooling, and max pooling. This paper uses max pooling for training, which, compared to other methods, better preserves the original features of the data and reduces the computational complexity of the model. After pooling, the data is input into the decoder, and the reconstructed result is:
[0112]
[0113]
[0114] in, To reconstruct the data, This represents the output of the i-th coding unit at time t. This is the hidden state vector of the decoder. And ρ(·) is usually set as a tanh function, thus derived from the formula The reconstruction error was calculated, where, To reconstruct the data, This is the original data.
[0115] In some implementations, the deep clustering model constructed in this invention includes an encoder and a decoder, each comprising two layers of Long Short-Term Memory (LSTM) network and two fully connected layers. A clustering layer is constructed between the encoder and decoder to achieve learning-while-clustering functionality. This improves the clustering performance of the model while reducing time complexity. Furthermore, the number of layers and parameters of the model can be appropriately modified according to different datasets.
[0116] In some implementations, in step S202, the clustering loss function of the model is constructed by using an LSTM autoencoder combined with an FCM algorithm clustering model.
[0117] In current research on feature learning and clustering of datasets, most methods involve deep learning of sequence features before clustering or binning, with cross-optimization of reconstruction and clustering losses. However, the clustering results are not very stable, and the model has high time complexity. Utilizing unsupervised clustering algorithms to jointly optimize deep neural networks has become an active research area. Exploratory methods combining deep learning include discriminative and soft clustering, K-means clustering, subspace clustering, Student's-t distribution, graph clustering, and Kullback-Leibler (KL) divergent clustering, among others. Addressing the difficulty of clustering short metagenomic fragments mentioned above, this invention proposes an improved FCM algorithm clustering model jointly implemented with an LSTM autoencoder. This model combines deep learning and clustering, simultaneously optimizing reconstruction and clustering errors, further improving their performance while reducing computational cost.
[0118] This invention represents data in a feature space generated by an autoencoder network, extracts the hidden layer output vector, utilizes a clustering layer to map data features from a high-dimensional space to a low-dimensional space, iteratively optimizes the clustering objective, and calculates the Student's-t distribution to measure the similarity between the hidden layer and the cluster centers. The N input data in the model are represented as X = {x1, ..., x...} i ,…,x N},x i ∈R d The output of the hidden layer of the autoencoder is defined as Z = {z1, ..., z}. i ,…,z N},z i ∈R c , where d and c are the dimensions of the input layer and the hidden feature layer, respectively. Let z... i The input is fed into the constructed clustering layer to achieve clustering based on the Student's-t distribution, as shown in the following formula:
[0119]
[0120] Specifically, the above formula defines the output probability distribution q of the clustering layer after the Student's-t distribution. ij Where μ = [μ1, ..., μ j ,…,μ c [ ] represents the weights of the clustering layer after training, c is the number of categories, and j is the number of clusters in the clustering layer. th Neuron. q ij For x i The membership degree of the j-th cluster, μ is the cluster center or represents the Student's-t distribution, and α is a hyperparameter, usually set to 1.
[0121] Traditional fuzzy C-means clustering algorithms suffer from a "uniformity effect" when dealing with imbalanced datasets, tending to generate clusters of similar size, resulting in poor clustering performance and low precision and recall. Assuming the dataset contains only two classes, the FCM objective function can be adjusted as follows:
[0122]
[0123] Where n represents the total number of samples, n1 and n2 represent the number of samples in each class, and v1 and v2 represent the cluster centers of the two classes. Since the input data x i x j Since it is a constant, Minimize L (where L is a constant) FCM Equivalent to maximizing 2n1n2||v1-v2|| 2 If ||v1-v2|| 2 If n1 is a constant, then n1n2 takes the maximum value if and only if n1 = n2 = n / 2, i.e., L FCM The optimal solution is obtained. At this point, the fuzzy C-means produces a uniform effect.
[0124] To address the aforementioned issues, this invention combines a deep LSTM autoencoder with an improved FCS algorithm, as shown in the model structure diagram below. Figure 7 As shown. Let Z = {z1,…,z} i ,…,z N},z i ∈R c Let be the vector representation of the hidden layer of the autoencoder, where c is the feature dimension of the hidden layer and N is the amount of input data. The resulting hidden layer output vector is passed through a self-constructed fuzzy clustering layer, with Z as input, providing the corresponding fuzzy membership matrix q. ij :
[0125]
[0126] In the formula, m is the fuzzy coefficient, which is generally taken as 2, β is a hyperparameter used to balance intra-class distance and inter-class distance, and μ=[μ1,…,μ j ,…,μ c ] represents the weights of the clustering layer after training, c represents the number of categories, and j represents the j-th neuron in the clustering layer. Defined as the average value of the hidden layer vectors, i.e. z i This represents the low-dimensional feature vector of the data. Similar to the traditional FCS algorithm, the membership formula is derived from the loss function. The clustering loss function of the entire deep clustering model is:
[0127]
[0128] Wherein, η j As shown in formula (9), it is unified as To ensure that the fuzzy membership matrix obtained during clustering from the hidden layer output vector is positive, where m is the fuzzy coefficient. It is a fuzzy membership matrix. Defined as the average value of the hidden layer vectors, i.e. The low-dimensional feature vector z i Inter-class distance Defined as d1, the intra-class distance ||z i -μ j || 2 Let d2 be the membership degree. Generally, in the early stages of training, the size of d2 tends to be too large. However, as training progresses, the size of d2 gradually decreases, while the size of d1 gradually increases, failing to positively influence d2. If d2 - d1 < 0, then the membership degree q is zero. ij A value less than 0 is inconsistent with the constraint objective and cannot mitigate the impact of the uniformity effect on clustering. Therefore, to ensure that d1 and d2 maintain the same proportion, d1 is... Normalization can clearly reflect the sensitivity of the distance difference between d1 and d2.
[0129] Expected target p based on Student's-t distribution ij Defined by the following formula, it not only increases q ij The variance is calculated, and the result is normalized relative to the j-th cluster center:
[0130]
[0131] Wherein, membership degree q ij The clustering layer output feature vector z i Relevance, membership degree q ij The probability distribution of the output feature vectors of the constructed clustering layer after passing through the softmax function is obtained. Based on the calculated p... ij and qij Calculate the KL divergence error, and make p by calculating the KL divergence. ij and q ij To get closer to the target while minimizing the loss function, the KL divergence error function is as follows:
[0132]
[0133] In this embodiment of the invention, K-means clustering is first used to initialize the cluster centers using the hidden layer vectors obtained after the first training iteration. Starting from the second iteration, the model updates the cluster centers using the clustering algorithm constructed above. Regarding the selection of the number of clusters, conventional algorithms such as DBI are only suitable for clustering contigs with many features. Since metagenomic DNA fragments are short and have few features, these algorithms cannot determine the specific number of clusters. Therefore, when performing clustering, the range of the number of clusters for unknown species sequences needs to be roughly determined based on the species distribution of known-labeled sequences in the selected metagenomic dataset.
[0134] This invention employs the Adam algorithm optimizer to update network parameters. Compared to the SGD algorithm and other optimization algorithms, the Adam optimizer is less prone to getting trapped in local optima, faster, more effective in learning, and attempts to correct some problems in other optimization techniques, such as learning rate vanishing or large fluctuations in the loss function caused by high variance parameter updates.
[0135] In some implementations, in step S203, a subset of DNA fragments with known metagenomic tags is used as a test set to test the clustering performance of the constructed model.
[0136] This embodiment of the invention uses the SRR492190 metagenomic dataset as an example. High-abundance species sequences are used as the test set for clustering. This test set contains 36,778 sequences belonging to 10 species, with a sequence ratio of 9:1 between the highest and lowest abundance species, still conforming to an imbalanced dataset. The test dataset is segmented into several 13-mer sequences using a sliding window method. These sequences are represented by the feature vector matrix output by the aforementioned word embedding model and input into the constructed deep clustering model. Since there are tens of millions of permutations and combinations of 13-mer sequences, there may be cases where the constructed word embedding feature vector matrix does not contain a certain 13-mer sequence from the test set. This embodiment of the invention supplements such 13-mer sequences with an all-zero vector matrix. Taking the SRR492190 test set as an example, each DNA fragment is 100 bp in length, and each sequence, after being represented by a word embedding feature vector, has a dimension of 88 × 10. The deep clustering model constructed in this embodiment of the invention achieves simultaneous training and clustering during the clustering process. Therefore, the overall loss error function of the model is the sum of the reconstruction error and the clustering error loss, as shown in the following formula:
[0137]
[0138] In the formula, L1 is the intra-class distance index, L2 is the inter-class distance index, L3 is the reconstruction error function, and δ and γ are two hyperparameters (δ,γ∈(0,1)). The values of the two hyperparameters are adjusted appropriately according to the clustering performance index so that the model obtains the optimal clustering performance.
[0139] In some implementations, the clustering performance metrics in step S204 include precision, recall, and adjusted RAND coefficient.
[0140] Traditional classification algorithms typically use classification accuracy to evaluate the results, which is calculated by dividing the number of correctly classified samples by the total number of samples. However, for imbalanced data, accuracy is not suitable for detailed evaluation and analysis. For example, in an imbalanced dataset containing two classes, where the majority class comprises 90% of the total samples, a 90% classification accuracy can be achieved by correctly classifying only all samples of the majority class. However, this evaluation metric is meaningless because the other 10% of minority class sequences are not correctly classified into their respective clusters. Therefore, this invention utilizes precision, recall, and the adjusted Rand Index (ARI) as performance metrics for clustering imbalanced datasets, all calculated from the confusion matrix.
[0141] The confusion matrix effectively reflects the degree of overlap between instance category classification and clustering results, as shown in Table 1. Positive classes are the minority class, and negative classes are the majority class. TP (True Positive) and FN (False Negative) represent the number of positive samples assigned to the positive and negative classes, respectively, while FP (False Positive) and TN (True Negative) represent the number of negative samples assigned to the positive and negative classes, respectively.
[0142] Table 1. Specific Representations of the Confusion Matrix
[0143]
[0144] Specifically, Generally, ε = 1. Only when both recall and precision are high can the F-value represent good clustering performance. Therefore, the F-value can evaluate the clustering performance of a clustering algorithm for minority classes. The F-value is defined as:
[0145]
[0146] Specifically, the degree of overlap is measured by the Adjusted Rand Index (ARI), defined as:
[0147]
[0148] In multi-category evaluation metrics, considering factors such as inconsistencies in sample sequence length, differences between the number of bins and the number of species with true labels, and imbalances among true species, this embodiment of the invention performs multi-faceted analysis of multi-label clustering results. Specific steps include:
[0149] a) For cases with inconsistent sequence lengths, the three performance metrics are calculated as the overall evaluation metrics for the final multi-label clustering results. Preferably, when adjusting the Adjusted Rand Index (ARI) for comparison of clustering algorithms, the value range is [-1, 1], and the values of the precision and recall performance metrics range are [0, 1], with values closer to 1 indicating better clustering results.
[0150] b) Analyzing inter-species clustering: When clustering multiple species within the same genus, it is difficult to accurately cluster them due to the similarity of sequence information among the species. Therefore, using precision and recall to analyze clustering performance is more effective in reflecting the quality of clustering. Higher precision and recall in intra-genus analysis indicates that this deep clustering algorithm model is more sensitive to different species within the same genus.
[0151] c) When analyzing the clustering performance of sequences among species with significant imbalance, combine all the above performance indicators for unified analysis to jointly measure the clustering effect and avoid inaccurate evaluations due to incomplete analysis.
[0152] In some implementations, in step S205, the constructed test set is input into the deep LSTM autoencoder joint clustering model, and the model parameters are reasonably adjusted based on the performance indicators obtained from multiple iterations of training and clustering.
[0153] In some specific implementations, for the test dataset in the SRR492190 metagenomics, the parameters of the deep clustering model are set as follows: the encoder and decoder each have two layers of Long Short-Term Memory (LSTM) networks and two fully connected layers. The input dimension of each sequence is 88×10, and the output dimensions of the two LSTM networks are 88×20 and 88×5, respectively. The output dimensions of each sequence after passing through the LSTM network are merged into a one-dimensional vector, and then passed through two fully connected layers with output dimensions of 200 and 64, respectively. This vector is then passed through the clustering layer constructed above, so that the hidden layer output dimension of each sequence is consistent with the set number of clusters. This facilitates the calculation of the probability of each sequence belonging to each species during the clustering process. The softmax function is used to obtain the species with the maximum probability, and the sequence is clustered into the class of that species. The encoder and decoder have a symmetrical structure, and the dimension setting of each layer of the decoder is consistent with that of the encoder, so that the output dimension of the final autoencoder is consistent with the input dimension, which is still 88×10. The hyperparameters in the overall loss error function of the model are δ = 0.5 and γ = 0.9. The iteration period is set to 500 epochs and the number of clusters is set to 10. In each iteration, 100 sequences are randomly selected from the dataset for training. After a whole iteration cycle, the hidden layer feature vector matrix is updated, deep clustering is performed, and the overall error function of the model is calculated. When the overall loss error of the model has not decreased after 10 consecutive iteration cycles, the clustering results are output, and the clustering performance index of the model is calculated and analyzed.
[0154] In some implementations, step S206, which involves clustering and assembling the unknown species in the dataset to be tested, and using software to analyze the integrity and contamination of the assembled contiguous groups, and outputting the clustering results of the unknown species, specifically includes the following steps:
[0155] a) Represent the unknown species sequences in the dataset using the feature vector matrix output by the constructed word embedding model, and input it into the constructed deep LSTM autoencoder joint clustering model for deep clustering;
[0156] b) After performing deep clustering on the unknown species, assemble them according to the clustering results using Spades assembly software;
[0157] c) Use check-m software to analyze the assembled contiguous groups and calculate the contamination and integrity.
[0158] The overall model flowchart of this invention embodiment is as follows: Figure 1As shown, the unknown species sequence portions of the selected dataset are input into the constructed deep LSTM autoencoder joint clustering model. Since the species distribution of unknown-labeled sequences in metagenomic datasets is similar to that of known-labeled sequences, and metagenomic DNA fragments are short and have few features, it is difficult to determine the approximate range of the number of clusters using conventional algorithms. Therefore, it is necessary to determine the number of species in the unknown-labeled sequences based on the distribution of known species sequences.
[0159] In some specific implementations, for the test dataset in the SRR492190 metagenomics, there are 12 high-abundance species (species with more than 1000 sequences) known tag sequences, accounting for approximately 97% of the total number of sequences in the dataset. Other species are considered to be of very low abundance. Therefore, the number of clusters of unknown species is determined to be approximately in the range of [10, 15]. In the final embodiment of this invention, the number of clusters of unknown species sequences is set to 10 based on this dataset. The unknown species sequence was represented by the feature vector matrix of the word embedding model mentioned above and input into the deep LSTM autoencoder joint clustering model. Iterative parameter tuning and adjustment of the number of network layers were performed. Finally, when the encoder and decoder each included one layer of long short-term memory network (LSTM) and three layers of fully connected network, and the autoencoder structure was set to a symmetric structure with 10*5*250*500*100*10*100*500*250*5*10 hidden layers, the learning rate was 0.001, the iteration period was 200 times, and the batch size was set to 256, the overall clustering performance of the model was optimal.
[0160] After performing deep clustering on unknown species, the Spades assembly software is used to assemble them according to the clustering results. The check-m software is then used to analyze the assembled contigs and calculate their contamination and integrity. The lower the contamination and the higher the integrity, the better the deep clustering model is.
[0161] This invention also provides a storage medium storing one or more programs that can be executed by one or more processors to implement the steps in the metagenomic species reconstruction method based on pre-training and deep clustering as described above.
[0162] This invention also provides a terminal device, including a processor adapted to implement various instructions; and a storage medium adapted to store multiple instructions, the instructions being adapted to be loaded by the processor and executed as described above in the metagenomic species reconstruction method based on pre-training and deep clustering.
[0163] In some implementations, the logical instructions in the storage medium can be implemented as software functional units and sold or used as independent products.
[0164] The storage medium, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of this disclosure. The processor executes functional applications and data processing by running the software programs, instructions, or modules stored in the storage medium, thereby implementing the methods in the above embodiments.
[0165] Storage media may include a stored program area and a stored data area. The stored program area may store the operating system and at least one application program required for a given function; the stored data area may store data created based on the use of the terminal device. Furthermore, storage media may include high-speed random access memory (RAM) and non-volatile memory. Examples include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks; these can also be transient storage media.
[0166] The following specific embodiments further illustrate the metagenomic species reconstruction method based on pre-training and deep clustering of the present invention:
[0167] Example 1
[0168] The Human Gut Microbiome Genome Research Project (MetaHIT) is a sub-project under the EU Seventh Framework Programme, aiming to study the human gut microbiota and understand its function and impact on human health. In this example, a relevant metagenomic dataset—the Preborninfant gut metagenome-Carrol stool sample—was extracted from articles published in the Human Gut Microbiome Genome Research Project, using the SRR492190 dataset as an example.
[0169] After quality control processing of the entire dataset and tag alignment using BLAST software, 1,020,842 DNA fragments were found to lack corresponding tags, with 510,421 fragments each for read1 and read2. The main purpose of this embodiment is to discover unknown species, and the distribution of unknown species sequences should be consistent with the distribution of tagged species sequences. The tagged sequences (a total of 1,523,952 DNA fragments, with 761,976 each for read1 and read2) were filtered to remove extremely low abundance species. In this embodiment, for the SRR492190 dataset, species with fewer than 1000 sequences were defined as extremely low abundance species. After processing, 12 species with a total of 1,477,442 DNA fragments remained, with the highest abundance species comprising 1,206,372 DNA fragments, approximately 81.65% of the total. Therefore, this metagenomic dataset still exhibits extreme imbalance after preprocessing.
[0170] The clustering performance of this embodiment is compared and analyzed with that of several methods, including MBMC (2016), Metaprob (2016), MetaBin (2019), Sparse Coding (2020), Kexue Li, et al. (2020), and Metaprob2 (2021). Among them, Metaprob (2016) and Metaprob2 (2021) solved the problem of variable distribution of k-mers and imbalanced metagenomic fragment clustering; MBMC (2016) uses Markov chains to represent each genome, and Markov chains describe the stochastic process of transitioning from one state to another in the state space; MetaBin (2019) is a method for comparing metrics based on variable-length sequences of metagenomics; Kexue Li, et al. (2020) uses statistical information obtained from multiple samples of the dataset to reduce underclustering problems; Sparse Coding (2020) uses sparse coding technology and elastic network regularization to realize potential genome recovery, and discovers and recovers unknown species in the genome through the sparsity and non-negativity constraints of k-mer sequence counting. Compared to previous studies, this embodiment combines deep learning with clustering, and the simultaneous training and clustering approach effectively reduces time complexity. The training set portion of the selected dataset is input into the constructed word embedding model for feature extraction. The test set is represented by the feature vector matrix output by the word embedding model, and further clustered and analyzed using a deep LSTM autoencoder and a joint clustering model, resulting in superior clustering performance metrics, such as... Figure 8 As shown in Table 2.
[0171] Table 2 Comparison of SRR492190 clustering results under different methods
[0172]
[0173]
[0174] Directly performing K-means clustering on the data represented by the word embedding feature vector matrix yielded a precision of 83.0% and a recall of 76.9%, with an adjusted Rand Index (ARI) of 80.2%. Inputting the data into a pre-built deep clustering model further improved the optimal clustering performance, achieving an adjusted Rand Index (ARI) of 91.3%, a precision of 94.0%, and a recall of 89.4%. Clustering the test dataset directly through several other comparative experiments showed poor clustering performance compared to this embodiment, with all performance metrics failing to reach 80%. This may be because the selected test dataset was too small; other methods could not obtain accurate feature vectors from such limited datasets, thus failing to achieve better clustering. In contrast, this embodiment first extracts features from the training set portion of the dataset, training the feature vector matrix to a certain level. Using this matrix to represent the test set of the dataset yields better clustering performance. Furthermore, the dataset selected in this embodiment is an imbalanced dataset. By combining a word embedding feature extraction model with a deep clustering model, the loss error function and the update method of cluster centers in the clustering layer are further improved, resulting in superior clustering results. Unknown species are represented using the trained feature vector matrix and clustered using the deep clustering model constructed in this embodiment. The results are then concatenated, and the completeness and contamination degree of the resulting contigs are calculated. For the SRR492190 dataset, there are 1,020,842 unknown-label sequences, with read1 and read2 each containing 510,421 sequences. Since read1 and read2 must belong to the same species, only deep clustering of read1 is needed. The species distribution of read2 sequences is consistent with the deep clustering results of read1. Because the known-label portion of the selected metagenomics is a highly imbalanced dataset, with the most abundant species sequences accounting for 81.65% of the total, the distribution of unknown-species sequences should be roughly similar, also an imbalanced dataset. After deep clustering, as shown... Figure 9 , Figure 10 And as shown in Table 3:
[0175] Table 3 Comparison of sequence integrity and contamination level after clustering and splicing of SRR492190 unknown species fragments using different methods.
[0176]
[0177]
[0178] This embodiment identified one unknown species (completeness > 50%, contamination < 6%), with a completeness of 74.23% and a contamination of 4.88%. Comparative experiments using other methods revealed one species in each of these methods. The contigs assembled using the Sparse Coding (2020) method had a contamination of 4.40%, but a low completeness of only 67.68%. Compared to other methods, the contigs assembled using the model constructed in this embodiment for deep clustering showed a clear advantage in overall performance, exhibiting higher completeness and relatively lower contamination.
[0179] In summary, this invention provides a metagenomic species reconstruction method based on pre-training and deep clustering. A joint clustering method based on deep LSTM autoencoders of metagenomic DNA fragments is proposed, employing a word embedding feature extraction model based on graph convolutional neural networks combined with Focal Loss, and a deep clustering model based on an improved FCM algorithm using LSTM autoencoders. This invention encodes each DNA fragment in the selected metagenomic dataset using a sliding window method in k-mer form, applies a two-layer GCN network as a word embedding model to train high-dimensional features, and represents the word embedding results obtained after training the above model for fragments without alignment labels in the metagenomic dataset as data samples for subsequent deep clustering. This invention constructs a joint clustering algorithm model based on deep LSTM autoencoders to perform clustering and binning on the processed dataset, combining deep learning and clustering, and simultaneously optimizing reconstruction and clustering errors. Compared to other algorithms, this further improves the performance of both while reducing computational cost. Finally, the clustered fragments are assembled according to categories using splicing software, and the completeness and contamination of the assembled contigs are calculated using the check-m tool. When using this method, users only need to adjust the parameters of the entire model according to the size of the selected dataset and the sequence length, and then rerun the model to obtain the clustering results, greatly improving accuracy and convenience. Validation with real datasets and comparison with the clustering results of other existing methods confirms that this invention can obtain superior clustering results. Furthermore, in discovering unknown species, the unknown species discovered by this invention have higher completeness and lower contamination compared to other methods.
[0180] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A metagenomic species reconstruction method based on pre-training and deep clustering, characterized in that, Including the following steps: The first step, feature extraction, includes: Raw metagenomic datasets of different environmental microorganisms were extracted and preprocessed. The raw datasets included DNA sequence features of different species. Build a word embedding model for the preprocessed dataset; Construct the model error function; The word embedding model is trained and its parameters are adjusted using the constructed model error function; Save the output feature vector matrix; The second step, deep clustering, includes: A deep LSTM autoencoder joint clustering model is constructed. The LSTM autoencoder consists of two parts: an encoder and a decoder. The encoder learns features from the input feature vector time series data, and the decoder reconstructs the data features using the current hidden layer state and network parameters. The clustering loss function of the model is constructed by combining an LSTM autoencoder with an FCM algorithm clustering model. Input the metagenomic dataset of the microorganism to be tested, and calculate the overall loss error of the model using the reconstruction error function and the clustering loss function; the formula for calculating the overall loss error is: , In the formula, It is an intra-class distance metric. As an inter-class distance metric, To reconstruct the error function, , For two hyperparameters ( ); calculate and analyze the clustering performance metrics of the model; Adjust the model parameters to obtain the optimal clustering performance of the model; The unknown species in the dataset to be tested are clustered and assembled, and the integrity and contamination of the assembled contiguous groups are analyzed by the software, and the clustering results of the unknown species are output. The preprocessing of the original dataset specifically includes the following steps: a) Download the metagenomic sequence dataset of the microbial community; b) Based on the quality value information of each DNA fragment stored in the dataset, optimize the data using quality control software tools and filter out low-quality sequences; c) Replace the N base in the remaining sequence after filtering in step b) with one of A, G, C, or T; d) Use the BLAST tool to compare the tags of all metagenomic sequences processed in step c), and distinguish different samples based on primers and index sequences; e) Based on the BLAST alignment results, filter out DNA fragment pairs that do not belong to the same species, and use the filtered sequences as input to the word embedding model; The low-quality sequence contains more than 5 consecutive N bases; The specific steps for constructing the word embedding model are as follows: a) The pre-processed DNA fragment sequence was cut using the sliding window method to convert the sequence into overlapping fixed-length k-mer sequences; b) Construct a topology graph from the segmented sequence database; c) Establish the edge between two k-mer nodes using the co-occurrence information of the k-mer sequence; d) Input the constructed topology graph into a two-layer graph convolutional neural network to build a word embedding model; In the first step, the step of constructing the model error function uses Focal Loss as the model error function; The encoder and decoder are each composed of two LSTM layers and two fully connected layers, with a clustering layer built between the encoder and decoder. The neurons of the LSTM are composed of several recursively connected memory blocks. Each memory block consists of one memory unit and three logic unit processes. The logic unit consists of a forget gate, an input gate, and an output gate. In the step of calculating and analyzing the clustering performance index of the model, the clustering performance index consists of precision, recall, and adjusted RAND coefficient.
2. A storage medium, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the steps in the metagenomic species reconstruction method based on pre-training and deep clustering as described in claim 1.
3. A terminal device, characterized in that, It includes a processor adapted to implement various instructions; and a storage medium adapted to store multiple instructions, said instructions being adapted to be loaded by the processor and executed in the metagenomic species reconstruction method based on pre-training and deep clustering as described in claim 1.
Citation Information
Patent Citations
Biological sequence feature extraction method based on word embedding and auto-encoder fusion
CN113392929A
Image clustering method and system based on depth subspace fuzzy clustering
CN114821142A