System generation tree construction method based on metric learning

Through the method based on metric learning, the construction accuracy and adaptability of the phylogenetic tree are optimized, and the problems of low computing efficiency and strong limitations of model assumptions in the existing technology are solved, and efficient and accurate construction of phylogenetic tree is achieved.

CN120067895AActive Publication Date: 2025-05-30YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510533729.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing phylogenetic tree construction methods have shortcomings in terms of computing efficiency, construction accuracy and model adaptability, especially when processing large-scale data, the model hypothesis is highly limited, and the error propagation problem is prominent.

Method used

Using a phylogenetic tree construction method based on metric learning, we use the phylogenetic tree construction method to collect nucleic acid sequences and corresponding evolution tree data, perform feature extraction and model training, design triple networks and triple loss functions, optimize the similarity metrics between data points, and build a phylogenetic tree using the generated similarity matrix.

Benefits of technology

It improves the construction accuracy of the phylogenetic tree, is compatible with multimodal data, improves the adaptability and practicality of the model, and overcomes the problems of computational complexity and model hypothesis limitations in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067895A_ABST
    Figure CN120067895A_ABST
Patent Text Reader

Abstract

The invention relates to a systematic occurrence tree construction method based on metric learning. The problems that in the prior art, system generation tree construction is high in calculation complexity, low in efficiency and poor in accuracy are solved. The method comprises the following steps: S1, collecting a public data set containing a nucleic acid sequence and a corresponding constructed evolutionary tree; s2, performing feature extraction on the collected data set; s3, designing and constructing a measurement learning model, and optimizing similarity measurement between data points; s4, training the metric learning model by using the training set and generating a similarity matrix; s5, based on the generated similarity matrix, constructing a tree structure by using a system generation tree construction algorithm, and performing result verification; and S6, performing application analysis. The method has the advantages that the problems of low distance measurement precision, high model hypothesis limitation, high calculation complexity and the like in a traditional method are solved, and efficient construction of the system generation tree under large-scale biological data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics and computer science technology, and particularly relates to a method for constructing a phylogenetic tree based on metric learning. Background Art

[0002] Existing methods for constructing phylogenetic trees usually rely on molecular sequence alignment and evolutionary models, such as Multiple Sequence Alignment (MSA) and Maximum Likelihood (ML) methods. Multiple Sequence Alignment finds the similarity between sequences by globally or locally aligning multiple DNA, RNA, or protein sequences and constructs an evolutionary tree; Maximum Likelihood searches for an evolutionary tree that maximizes the likelihood of the observed data based on a specific evolutionary model; Bayesian Inference uses Bayesian statistical methods to infer the evolutionary tree distribution based on sequence data. Although existing methods are widely used in phylogenetic tree construction, there are still significant limitations: high computational complexity, low efficiency of multiple sequence alignment and model inference methods in processing large-scale data; limitations of model assumptions, existing models may not accurately reflect the actual evolutionary process; error propagation problems, alignment errors will have a significant impact on subsequent tree construction.

[0003] In summary, the current methods for constructing phylogenetic trees have deficiencies in terms of computational efficiency, construction accuracy, and model adaptability. Summary of the Invention

[0004] The object of the present invention is to provide a method for constructing a phylogenetic tree based on metric learning for the above problems.

[0005] To achieve the above object, the present invention adopts the following technical solutions: A method for constructing a phylogenetic tree based on metric learning, the method includes the following steps: S1. Collect a publicly available dataset containing nucleic acid sequences and the corresponding constructed phylogenetic trees; S2. Extract features from the collected dataset; S3. Design and construct a metric learning model to optimize the similarity metric between data points; S4. Use the training set to train the metric learning model and generate a similarity matrix; S5. Based on the generated similarity matrix, use a phylogenetic tree construction algorithm to construct a tree structure and verify the results; S6. Application analysis.

[0006] In the above method for constructing a phylogenetic tree based on metric learning, step S1 specifically includes the following steps: S11. Download a dataset containing nucleic acid sequences and their corresponding phylogenetic trees from a publicly available nucleic acid evolution database; S12. In the dataset, each set of data contains a nucleic acid sequence set and a corresponding phylogenetic tree structure, and each set of data is respectively labeled as a group of units. There are n groups of units in the dataset; S13. Randomly divide the n groups of data units collected into a training set and a test set.

[0007] In the above method for constructing a phylogenetic tree based on metric learning, step S2 specifically includes the following steps: S21. For nucleic acid sequence data, use a sliding window method to extract k-mer features; S22. For phenotypic or ecological data, select a statistical method for feature extraction; S23. Standardize all the extracted features; S24. Integrate data features of different types.

[0008] In the above method for constructing a phylogenetic tree based on metric learning, step S3 specifically includes the following steps: S31. Select a suitable metric learning model structure: According to the goal, adopt a triplet network as the basic architecture of the metric learning model; S32. Design a loss function: According to the characteristics of the triplet network, adopt a triplet loss function as the objective function of the model. The calculation formula of the triplet loss function is: , where, represents the distance between the anchor point and the positive sample, represents the distance between the anchor point and the negative sample, and α is a hyperparameter representing the minimum threshold of the distance difference; S33. Process the training samples and construct triplets: Construct the data in the training set in the form of triplets, ensuring that each triplet contains an anchor point, a positive sample, and a negative sample; S34. Model architecture design and training details; S35. Regularization and hyperparameter tuning: To prevent overfitting, use L2 regularization to constrain the weights of the network. At the same time, adjust the hyperparameters through a cross-validation method.

[0009] In the above method for constructing a phylogenetic tree based on metric learning, the training of the metric learning model in step S4 specifically includes the following steps: S41. Dataset division and preparation: Randomly divide the data samples in the training set into subsets for training and validation; S42. Select an optimization algorithm: Adopt an appropriate optimization algorithm to adjust the parameters of the model; S43. Set the learning rate and batch size: Select an appropriate learning rate to control the step size of each parameter update; S44. Model training and loss function calculation: During the training process, each time a triple sample is obtained from the training set and input into the triple network, calculate the similarity measure of the network output, calculate the distances between the anchor and the positive sample, and between the anchor and the negative sample, and use the triple loss function to calculate the loss value; S45. Monitor the training process and overfitting: During the training process, regularly evaluate the performance of the model on the validation set, and calculate metrics such as loss value and accuracy; S46. Model tuning and hyperparameter optimization: After the initial training is completed, further optimize the hyperparameters through methods such as cross-validation and grid search; S47. Model evaluation and stability check: After the training process ends, use the validation set to conduct a final evaluation of the model, and check whether the similarity determination between different category samples is accurate.

[0010] The specific steps for generating the similarity matrix in step S4 are as follows: S480. Use the trained metric learning model: Apply the metric learning model to the test set and calculate the similarity between each sample.

[0011] S481. Calculate the similarity or distance between samples: For each pair of samples in the test set, use the trained metric learning model to calculate the similarity measure between them; S482. Construct the similarity matrix: According to the similarity values calculated in step S52, construct a symmetric matrix, and the dimension of this matrix is the number of samples in the test set , and each element of the matrix represents the similarity between sample and sample ; S483. Check the symmetry and sparsity of the matrix; S484. Handle outliers and noise; S485. Visualize the similarity matrix.

[0012] In step S5, the specific steps for constructing the tree structure are as follows: S51. Select the phylogenetic tree construction algorithm: Select the Neighbor-Joining method as the phylogenetic tree construction algorithm; S52. Use the similarity matrix as input: Use the similarity matrix generated in step S4 as input and pass it to the selected phylogenetic tree construction algorithm; each element in the matrix represents the similarity or distance between samples, and the algorithm will gradually merge samples or clusters based on this similarity / distance information to form a tree structure. S53. Construct the tree structure: After the similarity matrix is input, the phylogenetic tree construction algorithm will gradually construct the branches of the tree according to the similarity or distance. S54. Post - processing and optimization of the tree: By introducing a branch length correction method, optimize the branch lengths of the tree to make them conform to the actual evolutionary distances. The branch length correction uses the least - squares method. S55. Verify the rationality and consistency of the tree: By comparing with known biological prior knowledge or standard data sets, verify the rationality of the constructed phylogenetic tree. Use the RF distance of the tree, that is, the Robinson - Foulds distance, as an evaluation index to measure the matching degree between the constructed tree and the known evolutionary relationships. S56. Visualize the phylogenetic tree: Through an evolutionary tree visualization tool, convert the phylogenetic tree into a graphical form to intuitively display information such as the topological structure, branch lengths, and node confidence levels of the tree. In step S51, the specific steps of the phylogenetic tree construction algorithm are as follows: S511. Calculate the distance matrix: First, calculate the evolutionary distances between samples according to the similarity matrix generated in step S4. S512. Construct the initial tree structure: According to the calculated distance matrix, select the two samples or clusters with the smallest distance as the initial branches of the tree, then create a new node by merging these two clusters, update the distance matrix, and continue to select the clusters with the smallest distance for merging. S513. Gradually merge clusters: The algorithm gradually constructs the topological structure of the tree by repeatedly selecting and merging the clusters with the smallest distance. S514. Calculate the branch lengths: After the tree is constructed, use the information in the distance matrix to calculate the lengths of each branch of the tree.

[0013] In step S55, the RF distance of the tree is an index used to measure the similarity or difference between two phylogenetic trees. By comparing the topological structures of the two trees, it quantifies their differences. The RF distance is calculated based on the number of branches, that is, cuts, in the tree, reflecting the structural differences between the two trees. Given two phylogenetic trees, the steps for calculating the RF distance are as follows: S551. Tree cutting: Each edge that divides the tree into two parts. In the tree, the tree is divided into two sub - trees by removing an edge. S552. Calculate the difference: Calculate the comparison between the cutting sets of two trees to determine whether each cut exists in both trees. If a certain cut exists only in one of the trees, it means that the two trees are different in this cut. S553. Calculate the RF distance: The RF distance is the number of differences between the cutting sets of two trees; the calculation formula is: , where and are the cutting sets of tree and tree respectively, and represents the number of cuts that are in but not in

[0014] In step S5, the result verification includes the following steps: S570. Verify using known biological prior knowledge: By comparing the topological structure and branch relationships of the phylogenetic tree, check whether it conforms to the existing biological consensus. If the constructed tree is highly consistent with the known species relationships, it indicates that this method has high reliability in constructing the evolutionary tree. S571. Calculate the consistency index of the tree: Quantitatively evaluate the accuracy of the construction result by calculating the tree consistency index Tree Consistency Index, i.e., TCI. S572. Compare with traditional phylogenetic tree construction methods: Compare the phylogenetic tree constructed by this method with the evolutionary trees constructed by traditional algorithms such as distance matrix or neighbor-joining method. S573. Cross-validation and multiple validations: Divide the dataset into multiple subsets, alternately use different subsets as the training set and the test set, construct multiple phylogenetic trees, and compare the structural stability and consistency of the trees under different test sets. S574. Verify the biological rationality of the results: In the study of species evolutionary relationships, the constructed phylogenetic tree can be compared with the actual species genome or phenotypic data to check whether the branches in the tree conform to the known species evolutionary laws. S575. Result visualization and interpretation: Through visualizing the constructed phylogenetic tree, display the tree structure and node information, and at the same time provide visualization tools and analysis methods so that users can intuitively understand the branch relationships and similarities / differences between species.

[0015] Step S6 includes the following steps: S61. Species evolution and gene function analysis: Apply the constructed phylogenetic tree to the study of species evolutionary relationships, predict the evolutionary positions of unknown species, and explore the functional evolution of genes or proteins in different species. S62, Ecology and Species Conservation: Using phylogenetic trees to analyze the ecological relationships between species, providing a scientific basis for species conservation and environmental management, and at the same time providing feedback for optimizing evolutionary models.

[0016] Compared with the prior art, the advantages of the present invention are as follows: 1. Improve construction accuracy: By introducing a metric learning model, especially the triplet network and the triplet loss function, this solution optimizes the similarity metric between biological features. Compared with traditional methods, metric learning can better distinguish similar and different biological samples, thus improving the construction accuracy of phylogenetic trees; 2. Compatible with multi-modal data: This method not only supports nucleic acid sequence data, but also can process phenotypic or ecological data. Through feature extraction, normalization and fusion, different types of data can be effectively integrated, improving the comprehensive performance and adaptability of the model, which is of great significance for complex biological problems involving multiple data sources; 3. Flexible phylogenetic tree construction method: By generating a similarity matrix and applying the neighbor-joining algorithm, the construction of phylogenetic trees is both accurate and efficient. The subsequent branch length correction and tree optimization methods make the tree structure more biologically interpretable, thereby improving the practicality of the model; 4. Wide application: In addition to the study of species evolutionary relationships, this solution can also be used in various biological fields such as the functional study of genes and proteins. The constructed phylogenetic tree can not only reflect the evolutionary relationships between species, but also provide a reference for predicting the evolutionary positions of unknown species; In summary: Using a metric learning model to deeply represent and optimize biological features, generating a more reliable similarity matrix, and then combining with an advanced tree construction algorithm to finally generate a phylogenetic tree with high biological significance. By introducing metric learning technology, problems such as low distance measurement accuracy, strong limitations of model assumptions, and high computational complexity in traditional methods are overcome, realizing the efficient construction of phylogenetic trees under large-scale biological data. It has important application value in the field of bioinformatics, especially showing broad application prospects in analyzing species evolutionary relationships, exploring functional genes, and studying complex biological networks. Brief Description of the Drawings

[0017] Figure 1 is the method flow chart of the present invention; Figure 2 is the schematic diagram of the triplet neural network architecture in the present invention; Figure 3 is the result graph of the RF distance between the phylogenetic trees constructed by different methods in the present invention and the standard tree. Detailed Embodiments

[0018] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0019] As Figures 1 - 3 shown, a method for constructing a phylogenetic tree based on metric learning, the method includes the following steps: S1. Collect a publicly available dataset containing nucleic acid sequences and corresponding constructed phylogenetic trees; S2. Extract features from the collected dataset; S3. Design and construct a metric learning model to optimize the similarity metric between data points; S4. Use the training set to train the metric learning model and generate a similarity matrix; S5. Based on the generated similarity matrix, use a phylogenetic tree construction algorithm to construct a tree structure and verify the results; S6. Apply analysis.

[0020] Step S1 specifically includes the following steps: S11. Download a dataset containing nucleic acid sequences and their corresponding phylogenetic trees from a publicly available nucleic acid evolution database; S12. In the dataset, each group of data contains a nucleic acid sequence set and a corresponding phylogenetic tree structure, and each group of data is respectively labeled as a group of units. There are n groups of units in the dataset; The nucleic acid sequence is a DNA or RNA sequence, and all sequences are preprocessed into standard FASTA files to ensure uniform format. The phylogenetic tree is stored in the standardized Newick format.

[0021] In this example, the dataset contains at least n = 100 groups of unit data, and each group of units includes multiple nucleic acid sequences and corresponding phylogenetic tree structures.

[0022] S13. Randomly divide the collected n groups of data units into a training set and a test set.

[0023] Among them, the training set contains at least groups of data units, and the test set contains at least groups of data units. The nucleic acid sequences and phylogenetic tree structures of each group of data maintain a paired relationship to ensure data consistency during training and testing.

[0024] Step S2 specifically includes the following steps: S21. For nucleic acid sequence data, use a sliding window method to extract k-mer features; Specifically, by setting a window size \(k\), the nucleic acid sequence is segmented into subsequences of length \(k\), and the occurrence frequency of each \(k\)-mer in the sequence is counted. These frequency values form the feature vector of each sequence, representing the local structural features of the sequence. If the dataset contains sequences of multiple different lengths, \(k\)-mer features can be extracted separately for each length to ensure the comprehensiveness of the features.

[0025] S22. For phenotypic or ecological data, select statistical methods for feature extraction; Common methods include using basic statistics such as mean, variance, skewness, kurtosis, or reducing the dimension through principal component analysis to map high-dimensional data to a low-dimensional space and extract the most representative features. Such features can usually effectively capture the distribution characteristics of phenotypic or ecological data.

[0026] S23. Standardize all the extracted features; For the extracted feature vectors, use standardization methods such as z-score standardization to ensure that all features have the same scale. This step helps to eliminate the influence caused by different feature dimensions, enabling the subsequent metric learning model to process various features more accurately.

[0027] S24. Integrate different types of data features.

[0028] If the dataset contains both sequence data and phenotypic or ecological data, the extracted features need to be fused. Through methods such as concatenation or weighted averaging, the features from different sources are combined into a unified feature vector as the input for the subsequent metric learning model. This step helps to enhance the performance of the model when dealing with multi-modal data.

[0029] Step S3 specifically includes the following steps: S31. Select an appropriate metric learning model structure: According to the objective, adopt a triplet network as the basic architecture of the metric learning model; The triplet network learns the similarity and difference between samples by inputting three samples, namely the anchor, the positive sample, and the negative sample, so as to optimize the distance metric of the model. The core idea of the triplet network is to make the distance between the anchor and the positive sample as small as possible, and the distance between the anchor and the negative sample as large as possible.

[0030] S32. Design a loss function: According to the characteristics of the triplet network, adopt a triplet loss function as the objective function of the model. The calculation formula of the triplet loss function is: , where, represents the distance between the anchor and the positive sample, Denote the distance between the anchor point and the negative sample, and α is a hyperparameter representing the minimum threshold of the distance difference; The objective of the loss function is to make the distance between the anchor point and the positive sample less than the distance between the anchor point and the negative sample, while ensuring that the gap between the two is greater than the threshold α, thereby improving the model's ability to distinguish between similar and different samples.

[0031] S33. Process the training samples and construct triplets: Construct the data in the training set in the form of triplets, ensuring that each triplet contains an anchor point, a positive sample, and a negative sample; Specifically, the anchor point and the positive sample should belong to similar categories or have similar biometric features, while the negative sample belongs to a different category or has significantly different biometric features. In this way, the model can learn the distance relationship between similar and different samples during training and optimize the similarity metric.

[0032] S34. Model architecture design and training details; The triplet network usually consists of neural networks with shared weights, and the specific network structure can be selected according to the characteristics of the dataset. For example, use a convolutional neural network (CNN) to process sequence data or a multi-layer perceptron (MLP) to process phenotypic data. When constructing the network, select appropriate activation functions such as ReLU or Sigmoid, optimization methods such as Adam or SGD, and ensure that the network can converge effectively during training.

[0033] S35. Regularization and hyperparameter tuning: To prevent overfitting, use L2 regularization to constrain the weights of the network. At the same time, adjust the hyperparameters through the cross-validation method. Such as the α value in the triplet loss function, the learning rate of the network, and the batch size, to ensure that the model can achieve the best performance.

[0034] The specific training of the metric learning model in step S4 includes the following steps: S41. Dataset partitioning and preparation: Randomly partition the data samples in the training set into subsets for training and validation; The training set is used to optimize the model parameters, and the validation set is used to evaluate the generalization ability of the model during training to prevent overfitting. Ensure that the triplets (anchor point, positive sample, negative sample) in each subset are consistent to ensure the reliability of the data during the training and validation processes.

[0035] S42. Select an optimization algorithm: Use an appropriate optimization algorithm to adjust the parameters of the model. Commonly used optimization algorithms include the Adam optimizer and Stochastic Gradient Descent (SGD). In the present invention, the Adam optimizer is used because it has good convergence and can automatically adjust the learning rate to adapt to the update speeds of different parameters. During the training process, the optimization algorithm gradually adjusts the model weights according to the gradient information calculated by the Triplet Loss function to minimize the value of the loss function.

[0036] S43. Set the learning rate and batch size: Select an appropriate learning rate to control the step size of each parameter update. In this example, it is set to 1e-3. A too large learning rate will cause the training process to be unstable, while a too small learning rate will result in an overly slow convergence speed. In addition, the batch size is also an important hyperparameter affecting the training process. In this example, it is set to 64. An appropriate batch size can balance the training time and the model's convergence speed.

[0037] S44. Model training and loss function calculation: During the training process, each time a triplet sample is obtained from the training set and input into the triplet network, the similarity metric output by the network is calculated. By calculating the distances between the anchor and the positive sample, and the anchor and the negative sample, the loss value is calculated using the Triplet Loss function. The model will adjust the parameters according to this loss value, continuously optimizing the similarity metric, thereby improving the ability to distinguish between similar and different samples.

[0038] S45. Monitor the training process and overfitting: During the training process, regularly evaluate the performance of the model on the validation set, and calculate metrics such as the loss value and accuracy. If it is found that the loss on the training set continues to decrease while the loss on the validation set levels off or increases, overfitting may occur. At this time, early stopping techniques can be used to prevent overfitting, that is, stop the training when the performance on the validation set no longer improves, or use methods such as Dropout and L2 regularization to further control the complexity of the model.

[0039] S46. Model tuning and hyperparameter optimization: After the initial training is completed, further optimize the hyperparameters through methods such as cross-validation and grid search. For example, adjust parameters such as the threshold in the Triplet Loss, the learning rate, and the batch size, and try different network architectures and activation functions to find the configuration that best suits this dataset. Through these tuning means, the performance of the model on the training set and the validation set can be improved.

[0040] S47. Model evaluation and stability check: After the training process ends, use the validation set to finally evaluate the model and check whether the similarity determination between different category samples is accurate.

[0041] Meanwhile, the stability of the model can be checked by repeating the training multiple times to ensure its consistent performance during different initializations and training processes.

[0042] In step S4, generating the similarity matrix specifically includes the following steps: S480. Use the trained metric learning model: Apply the metric learning model to the test set and calculate the similarity between each pair of samples.

[0043] In step S4, the well-trained metric learning model already has the ability to distinguish the similarity and difference of biometric features. The input of the metric learning model is the data in the test set, and the metric learning model will output the similarity or distance between each pair of samples.

[0044] S481. Calculate the similarity or distance between samples: For each pair of samples in the test set, use the trained metric learning model to calculate the similarity metric between them. Generally, the metric learning model will output a numerical value representing the distance or similarity between two samples. The smaller the distance, the more similar the samples are, and the larger the distance, the more different the samples are. Each element in the similarity matrix represents the similarity between two samples in the test set.

[0045] S482. Construct the similarity matrix: According to the similarity values calculated in step S52, construct a symmetric matrix whose dimension is the number of samples in the test set , and each element of the matrix represents the similarity between sample and sample ; for each pair of samples , The value range of is usually from 0 to 1, or it is normalized according to the output of the model. The diagonal elements

[0046] of the matrix are usually 1, indicating that the sample has the highest similarity to itself. Since the similarity matrix is symmetric, i.e., , it is necessary to check whether the matrix satisfies this condition. At the same time, if the number of samples is large, the matrix may be very sparse. In this case, a sparse matrix storage method can be adopted to improve the calculation efficiency and avoid excessive memory consumption.

[0047] S484. Handle outliers and noise; After the similarity matrix is generated, there may be some outliers or noise, such as the similarity calculation of some samples being too high or too low; these outliers can be processed by setting thresholds or using smoothing techniques such as Gaussian filtering to ensure the quality of the similarity matrix.

[0048] For example, if the similarity of a pair of samples is much higher or lower than the normal range, it can be adjusted to a reasonable value to avoid affecting the accuracy of the subsequent phylogenetic tree construction of the system.

[0049] S485. Visualize the similarity matrix.

[0050] For the convenience of subsequent analysis and verification, the generated similarity matrix is visualized, and a heat map is used to display the structure of the matrix. By observing the change of color depth, the similarity distribution between samples can be intuitively observed, and further analysis can be carried out on which samples have strong similarity and which samples have large differences.

[0051] In step S5, constructing the tree structure specifically includes the following steps: S51. Select the phylogenetic tree construction algorithm: Select the Neighbor-Joining method as the phylogenetic tree construction algorithm; The Neighbor-Joining method is a commonly used distance-based tree construction method. It constructs a phylogenetic tree by gradually merging the most similar samples or clusters and is suitable for large-scale sequence datasets.

[0052] The basic idea of the Neighbor-Joining method is that in each step of the merging process, select the two clusters with the smallest distance as the new parent cluster until all samples or clusters are merged into a complete tree.

[0053] S52. Use the similarity matrix as input: Use the similarity matrix generated in step S4 as input and pass it to the selected phylogenetic tree construction algorithm; each element in the matrix represents the similarity or distance between samples, and the algorithm will gradually merge samples or clusters according to this similarity / distance information to form a tree structure; For the Neighbor-Joining method, the "minimum distance" strategy is usually adopted, and the samples or clusters with the smallest distance are selected for merging in each step of the merging process.

[0054] S53. Construct the tree structure: After the similarity matrix is input, the phylogenetic tree construction algorithm will gradually construct the branches of the tree according to the similarity or distance; Each branch represents a group of similar samples, and the branch length of the tree reflects the similarity or evolutionary distance between samples. Shorter branches indicate higher similarity between samples, while longer branches indicate larger differences between samples.

[0055] S54. Post-processing and optimization of the tree: The initially constructed phylogenetic tree may require further post-processing and optimization, especially the correction of branch lengths, to improve the accuracy and biological significance of the tree. By introducing a branch length correction method, optimize the branch lengths of the tree to make them conform to the actual evolutionary distance, and the branch length correction uses the least squares method for correction; The least squares method is a commonly used method for branch length correction. This method adjusts the branch lengths of the tree by minimizing the branch length differences of all nodes in the phylogenetic tree. Specifically, the least squares method calculates the error between the actual distances in the tree and the distances predicted according to the model, with the goal of minimizing these errors by adjusting the branch lengths. This method is applicable to dealing with branch length biases caused by sequencing errors or preliminary tree construction methods (such as the NJ algorithm).

[0056] S55. Verify the rationality and consistency of the tree: By comparing with known biological prior knowledge or standard data sets, verify the rationality of the constructed phylogenetic tree. Use the RF distance of the tree, that is, the Robinson-Foulds distance, as an evaluation index to measure the matching degree between the constructed tree and the known evolutionary relationships; If the evaluation results show that the tree structure is reasonable and conforms to known biological laws, it indicates that the constructed tree has good accuracy and reliability.

[0057] S56. Visualize the phylogenetic tree: For the convenience of subsequent analysis and application, the phylogenetic tree is usually visualized. Through an evolutionary tree visualization tool, such as FigTree, convert the phylogenetic tree into a graphical form to intuitively display information such as the topological structure of the tree, branch lengths, and node confidence levels.

[0058] The visualization results can help researchers more clearly understand the evolutionary relationships between species and provide strong support for further biological research.

[0059] In step S51, the specific steps of the phylogenetic tree construction algorithm are as follows: S511. Calculate the distance matrix: First, calculate the evolutionary distances between samples according to the similarity matrix generated in step S4; Since the similarity matrix represents the similarity or resemblance between samples, the neighbor-joining algorithm converts these similarities into a distance matrix, usually using the distance d(i, j) = 1 - S(i, j) to calculate, where S(i, j) is the similarity value between sample i and sample j.

[0060] S512. Construct the initial tree structure: According to the calculated distance matrix, select the two samples or clusters with the smallest distance as the initial branches of the tree, then create new nodes by merging these two clusters, update the distance matrix, and continue to select the clusters with the smallest distance for merging; S513. Gradually merge clusters: The algorithm gradually constructs the topological structure of the tree by repeatedly selecting the clusters with the smallest distance and merging them; At each merge, update the distance matrix to ensure that the distances between each new cluster and the remaining clusters are recalculated.

[0061] S514. Calculate the branch length: After the tree is constructed, use the information in the distance matrix to calculate the length of each branch of the tree.

[0062] The length of the branch reflects the evolutionary distance between samples. A shorter branch indicates a high similarity between samples, while a longer branch indicates a greater difference.

[0063] By using the neighbor-joining algorithm, a phylogenetic tree can be constructed quickly and effectively, especially suitable for large-scale biological datasets. In practical applications, the neighbor-joining algorithm has a low computational complexity and can provide a relatively reasonable tree structure.

[0064] In step S55, the RF distance of the tree is an index used to measure the similarity or difference between two phylogenetic trees. By comparing the topological structures of the two trees, it quantifies their degree of difference. The RF distance is calculated based on the number of branches, that is, the number of cuts, in the tree, reflecting the structural difference between the two trees. Given two phylogenetic trees, the RF distance calculation steps are as follows: S551. Tree cutting: Each edge that divides the tree into two parts. In the tree, removing an edge divides the tree into two subtrees; S552. Calculate the difference: Calculate the comparison between the cutting sets of the two trees, and determine whether each cut exists in both trees. If a cut exists only in one of the trees, it means that the two trees are different at this cut; S553. Calculation of the RF distance: The RF distance is the number of differences between the cutting sets of the two trees; the calculation formula is: , where and are the cutting sets of tree and tree respectively represents the number of cuts that are in but not in

[0065] The value range of the RF distance is from 0 to the maximum value. The maximum value is the number of cutting differences when the two trees are completely different. A value of 0 means that the two trees are exactly the same and the topological structures of the trees are exactly the same. The RF distance is used to compare the similarity of two trees and is often used to evaluate the difference between the constructed result of a phylogenetic tree and a known tree or a tree constructed by other methods. For example, in the evaluation of algorithm performance, the RF distance can quantify the accuracy of phylogenetic trees constructed by different methods.

[0066] In step S5, the result verification includes the following steps: S570. Verification using known biological prior knowledge: By comparing the topological structure and branch relationships of the phylogenetic tree, check whether it conforms to the existing biological consensus. If the constructed tree is highly consistent with the known species relationships, it indicates that this method has high reliability in constructing the evolutionary tree; To evaluate the construction result of the phylogenetic tree, it is necessary to compare the constructed phylogenetic tree with the existing biological prior knowledge, which can be derived from known species taxonomic information, traditional phylogenetic research results, or evolutionary trees in public databases.

[0067] S571. Calculate the consistency index of the tree: Quantitatively evaluate the accuracy of the construction result, and calculate the Tree Consistency Index (TCI) of the tree; TCI is used to measure the similarity between the phylogenetic tree and the known evolutionary tree. The closer the value is to 1, the higher the consistency between the constructed tree and the known tree. In addition, other evaluation criteria such as the F1 score can be adopted to further verify the accuracy of the tree structure by comparing the prediction results of the phylogenetic tree with the actual evolutionary relationships.

[0068] S572. Compare with traditional phylogenetic tree construction methods: Compare the phylogenetic tree constructed by this method with the evolutionary trees constructed by traditional algorithms such as distance matrix or neighbor-joining method; By quantitatively comparing indicators such as the tree consistency, tree depth, and branch stability of the two, evaluate the advantages of this method in terms of accuracy, stability, etc. If the construction result of this method is superior to the traditional method in the evaluation indicators, it indicates that this method can better improve the accuracy and biological significance of the phylogenetic tree with the introduction of metric learning.

[0069] S573. Cross-validation and multiple validations: Divide the dataset into multiple subsets, and alternately use different subsets as the training set and test set to construct multiple phylogenetic trees, and compare the structural stability and consistency of the trees under different test sets; To further verify the robustness of the model, techniques such as cross-validation are used to verify the model, which helps to evaluate whether the constructed phylogenetic tree has good generalization ability and avoid tree structure deviation caused by data overfitting.

[0070] S574. Verify the biological rationality of the results: In the study of species evolutionary relationships, the constructed phylogenetic tree can be compared with the actual species genome or phenotypic data to check whether the branches in the tree conform to the known species evolutionary laws; If the branches of the tree are consistent with the known evolutionary relationships between species, it indicates that the constructed tree has strong biological significance.

[0071] S575, Result Visualization and Explanation: By visualizing the constructed phylogenetic tree, the structure and node information of the tree are displayed. Meanwhile, visualization tools and analysis methods are provided so that users can intuitively understand the branch relationships of the tree and the similarities / differences between species.

[0072] In addition, the nodes of the tree can be annotated to display relevant biological information (such as genomic, phenotypic, or functional data), which helps researchers further analyze the evolutionary processes of species or genes in the tree.

[0073] Step S6 includes the following steps: S61, Species Evolution and Gene Function Analysis: Apply the constructed phylogenetic tree to the study of species evolutionary relationships, predict the evolutionary positions of unknown species, and explore the functional evolution of genes or proteins in different species; S62, Ecology and Species Conservation: Use the phylogenetic tree to analyze the ecological relationships between species and provide a scientific basis for species conservation and environmental management, while providing feedback for optimizing the evolutionary model.

[0074] In summary, the principle of this embodiment is as follows: Through the metric learning model, optimize the similarity measurement between biological features, thereby improving the accuracy and biological significance of the phylogenetic tree; secondly, use the triplet network as the metric learning model, and optimize the similarity measurement between samples through the triplet loss function, so that the model can more accurately distinguish similar and different biological features; and use the neighbor-joining method as the main algorithm for constructing the phylogenetic tree, and combine the similarity matrix to construct the tree structure; at the same time, optimize the accuracy of the tree through post-processing steps such as branch length correction and Bootstrap test.

[0075] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar ways to substitute them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A phylogenetic tree construction method based on metric learning, characterized in that: The method comprises the following steps: S1. Collect public data sets containing nucleic acid sequences and corresponding constructed evolutionary trees, and randomly divide the collected data into training sets and test sets; S2, extract features from the collected data set and perform standardization to integrate data features; S3. Design and build a metric learning model, design a loss function, optimize the similarity measure between data points, process training samples and construct triplets, and adjust hyperparameters through cross-validation methods; S4, divide and prepare the data set, select the optimization algorithm for model training and loss function calculation, monitor the model training process and overfitting, tune the model and optimize the hyperparameters, evaluate the model and check the stability, and finally generate a similarity matrix based on the model; S5. Based on the generated similarity matrix, use the similarity matrix as input to construct the branches of the tree, optimize the branch length of the tree by introducing the branch length correction method, verify the rationality of the constructed phylogenetic tree and the degree of match between the known evolutionary relationships, and intuitively display the information of the tree through the evolutionary tree visualization tool; S6. Apply the constructed phylogenetic tree to the study of species evolutionary relationships and ecological species conservation.

2. A phylogenetic tree construction method based on metric learning according to claim 1, characterized in that: Step S1 specifically includes the following steps: S11. Downloading a data set containing nucleic acid sequences and their corresponding evolutionary trees from a public nucleic acid evolution database; S12. In the data set, each set of data includes a set of nucleic acid sequences and a corresponding evolutionary tree structure. Each set of data is marked as a group of units. There are n groups of units in the data set. S13. Randomly divide the collected n groups of data units into training sets and test sets.

3. A phylogenetic tree construction method based on metric learning according to claim 2, characterized in that: Step S2 specifically includes the following steps: S21. For nucleic acid sequence data, a sliding window method is used to extract k-mer features; S22. For phenotypic or ecological data, select statistical methods for feature extraction; S23, standardizing all the extracted features; S24. Integrate different types of data features.

4. A phylogenetic tree construction method based on metric learning according to claim 3, characterized in that: Step S3 specifically includes the following steps: S31. Select the appropriate metric learning model structure: According to the goal, the triplet network is used as the basic architecture of the metric learning model; S32. Design loss function: According to the characteristics of the triplet network, the triplet loss function is used as the objective function of the model. The calculation formula of the triplet loss function is: , in, represents the distance between the anchor point and the positive sample, represents the distance between the anchor point and the negative sample, is a hyperparameter, indicating the minimum threshold of the distance difference; S33, process training samples and construct triplets: construct the data in the training set into the form of triplets, ensuring that each triplet contains an anchor point, a positive sample and a negative sample; S34, model architecture design and training details; S35, Regularization and hyperparameter adjustment: In order to prevent overfitting, L2 regularization is used to constrain the weights of the network. At the same time, the hyperparameters are adjusted through cross-validation method.

5. A phylogenetic tree construction method based on metric learning according to claim 4, characterized in that: The metric learning model training described in step S4 specifically includes the following steps: S41. Dataset division and preparation: randomly divide the data samples in the training set into subsets for training and validation; S42, selecting an optimization algorithm: using an appropriate optimization algorithm to adjust the parameters of the model; S43, set learning rate and batch size: select an appropriate learning rate to control the step size of each parameter update; S44, model training and loss function calculation: During the training process, a triplet sample is obtained from the training set each time, input into the triplet network, and the similarity measure of the network output is calculated. The loss value is calculated using the triplet loss function by calculating the distance between the anchor point and the positive sample, and the distance between the anchor point and the negative sample; S45. Monitor the training process and overfitting: During the training process, regularly evaluate the performance of the model on the validation set and calculate indicators such as loss value and accuracy; S46. Model tuning and hyperparameter optimization: After the initial training is completed, the hyperparameters are further optimized through methods such as cross-validation and grid search; S47, Model evaluation and stability check: After the training process is completed, the model is finally evaluated using the validation set to check whether its similarity judgment between samples of different categories is accurate.

6. A phylogenetic tree construction method based on metric learning according to claim 5, characterized in that: The generation of the similarity matrix described in step S4 specifically includes the following steps: S480, using the trained metric learning model: applying the metric learning model to the test set, and calculating the similarity between each sample; S481. Calculate the similarity or distance between samples: for each pair of samples in the test set, calculate the similarity measure between them using the trained metric learning model; S482, construct a similarity matrix: According to the similarity values ​​calculated in step S52, construct a symmetric matrix, the dimension of which is the number of samples in the test set , each element of the matrix Representation sample With sample The similarity between S483, check the symmetry and sparsity of matrices; S484, processing outliers and noise; S485. Visualize the similarity matrix.

7. A phylogenetic tree construction method based on metric learning according to claim 6, characterized in that: In step S5, constructing the tree structure specifically includes the following steps: S51, select the phylogenetic tree construction algorithm: select the neighbor-joining method as the phylogenetic tree construction algorithm; S52, using the similarity matrix as input: passing the similarity matrix generated in step S4 as input to the selected phylogenetic tree construction algorithm; S53, constructing a tree structure: after the similarity matrix is ​​input, the phylogenetic tree construction algorithm will gradually construct the branches of the tree according to similarity or distance; S54, tree post-processing and optimization: optimize the branch length of the tree by introducing the branch length correction method; S55. Verify the rationality and consistency of the tree: Verify the rationality of the constructed phylogenetic tree by comparing it with known biological prior knowledge or standard data sets, and use the RF distance of the tree to measure the degree of match between the constructed tree and the known evolutionary relationship; S56. Visualize phylogenetic tree: Convert phylogenetic tree into graphical form through evolutionary tree visualization tool.

8. A phylogenetic tree construction method based on metric learning according to claim 7, characterized in that: In step S55, the RF distance of the tree is an indicator for measuring the similarity or difference between two phylogenetic trees. The difference between the two trees is quantified by comparing their topological structures. The RF distance is calculated based on the number of branches, i.e., cuts, in the tree, reflecting the structural difference between the two trees. Given two phylogenetic trees, the RF distance calculation steps are as follows: S551, Tree Cutting: Each edge that divides a tree into two parts, in a tree, splits the tree into two subtrees by removing an edge; S552, calculating differences: calculating the comparison between the cut sets of the two trees, and determining whether each cut exists in both trees. If a cut only exists in one of the trees, it means that the two trees are different in the cut. S553. Calculation of RF distance: RF distance is the number of differences between two tree cut sets; the calculation formula is: , in, and The trees And Tree The cutting set, express But The number of cuts not in .

9. A phylogenetic tree construction method based on metric learning according to claim 6, characterized in that: In step S5, the result verification includes the following steps: S570, Verify using known biological prior knowledge: By comparing the topological structure and branching relationship of the phylogenetic tree, check whether it conforms to the existing biological consensus. If the constructed tree is highly consistent with the known species relationship, it means that this method has high reliability in constructing evolutionary trees; S571. Compute tree consistency index: quantitatively evaluate the accuracy of the construction results and calculate the tree consistency index TreeConsistency Index, or TCI; S572. Comparison with traditional phylogenetic tree construction methods: Compare the phylogenetic tree constructed by this method with the traditional evolutionary tree constructed based on algorithms such as distance matrix or neighbor-joining method; S573, cross-validation and multiple validation: Divide the data set into multiple subsets, use different subsets as training sets and test sets in turn, construct multiple phylogenetic trees, and compare the structural stability and consistency of the trees under different test sets; S574. Verify the biological plausibility of the results: In the study of species evolutionary relationships, the constructed phylogenetic tree can be compared with the actual species genome or phenotypic data to check whether the branches in the tree conform to the known laws of species evolution; S575. Result visualization and interpretation: By visualizing the constructed phylogenetic tree, the tree structure and node information are displayed, and visualization tools and analysis methods are provided so that users can intuitively understand the branching relationship of the tree and the similarities / differences between species.

10. A phylogenetic tree construction method based on metric learning according to claim 9, characterized in that: Step S6 includes the following steps: S61. Species evolution and gene function analysis: Apply the constructed phylogenetic tree to the study of species evolutionary relationships, predict the evolutionary position of unknown species, and explore the functional evolution of genes or proteins in different species; S62. Ecology and species conservation: Use phylogenetic trees to analyze the ecological relationships between species and provide a scientific basis for species conservation and environmental management, while providing feedback for the optimization of evolutionary models.

Citation Information

Patent Citations

  • Method for constructing model for classifying nucleic acid sequences and application thereof

    CN112599196A

  • Pedigree tracing method based on whole genome re-sequencing SNP big data and deep learning

    CN118248210A

  • Development tree inference method and device based on biological big language model and medium

    CN119443274A

  • Mining method and system for synthetic biological functional element and storage medium

    CN119673283A

  • System and method for statistical mapping between genetic information and facial image data

    US20110206246A1