A phylogenetic tree construction method and system based on a nucleic acid sequence distance matrix

By processing and optimizing nucleic acid sequence data, an optimized distance matrix with minimal supermetric deviation and function is generated, which solves the problems of accuracy and large computational complexity in the construction of nucleic acid sequence distance matrices in the existing technology, and realizes efficient and accurate phylogenetic tree construction.

CN119721195BActive Publication Date: 2025-10-17SHENZHEN MSU-BIT UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510230572.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-10-17
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

In the prior art, when reconstructing the nucleic acid sequence distance matrix, there are problems of poor accuracy and large amount of calculation, which affect the construction of the phylogenetic tree.

Method used

By acquiring and processing the original nucleic acid sequence data, an initial distance matrix is ​​generated, and iterative optimization is performed using the global alignment algorithm and gradient optimization algorithm to generate an optimized distance matrix with the minimum supermetric deviation and function. Finally, a phylogenetic tree is constructed based on the optimized distance matrix.

Benefits of technology

It achieves efficient and accurate construction of phylogenetic trees in the case of incomplete or large data sets, maintaining the accuracy of biological properties and effective use of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721195B_ABST
    Figure CN119721195B_ABST
Patent Text Reader

Abstract

The application provides a phylogenetic tree construction method and system based on a nucleic acid sequence distance matrix, comprising the following steps: obtaining and processing original nucleic acid sequence data to obtain standard nucleic acid sequence data; extracting a plurality of nucleic acid sequences from the standard nucleic acid sequence data, obtaining a global alignment distance between any two nucleic acid sequences in the plurality of nucleic acid sequences, and generating an initial distance matrix of the standard nucleic acid sequence data; obtaining a hypermetric deviation sum function of the initial distance matrix; iteratively optimizing the initial distance matrix to obtain an optimized distance matrix with the minimum hypermetric deviation sum function; and constructing a phylogenetic tree based on the optimized distance matrix. The effective nucleic acid sequence distance matrix is obtained through a hybrid optimization technique, so that the phylogenetic tree can be accurately constructed under the condition that the initial data is incomplete or too large in scale.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of phylogenetic tree construction, and particularly relates to a phylogenetic tree construction method and system based on nucleic acid sequence distance matrix. BACKGROUND

[0002] In phylogenetic research and evolutionary biology, accurate construction of phylogenetic tree is the basis for understanding the evolutionary relationship between species or individuals. By quantifying genetic differences, gene sequences can be compared to obtain a nucleic acid sequence distance matrix as a raw parameter for constructing a phylogenetic tree, to assist in the construction of a phylogenetic tree. However, in real experimental processes, incomplete or inaccurate distance matrices often occur, for example, sequence data loss caused by sequencing reaction failure, calculation errors or missing data caused by calculation limitations in large-scale genome comparison, and quality control processes for removing unreliable sequences also affect the integrity of the distance matrix. Such incomplete or inaccurate distance matrix will undoubtedly affect the construction process of the phylogenetic tree.

[0003] In the prior art, the problem of missing data in the distance matrix is generally solved by pairwise deletion or mean estimation, but this solution often leads to loss of valuable information and introduces bias, thereby affecting the accuracy of the results. In order to improve the accuracy, detailed sequence alignment is performed, for example, using a global alignment algorithm such as the Needleman-Wunsch algorithm, which results in excessive computation and is not suitable for large data sets.

[0004] Therefore, the prior art still needs to be improved and developed. SUMMARY

[0005] In view of the deficiencies of the prior art described above, the purpose of the present application is to provide a phylogenetic tree construction method based on nucleic acid sequence distance matrix, which aims to solve the problem of poor accuracy and large amount of calculation in reconstructing a complete distance matrix, which affects the construction of a phylogenetic tree in the prior art.

[0006] The first aspect of the present application provides a phylogenetic tree construction method based on nucleic acid sequence distance matrix, comprising the steps of:

[0007] obtaining and processing original nucleic acid sequence data to obtain standard nucleic acid sequence data;

[0008] extracting a plurality of nucleic acid sequences from the standard nucleic acid sequence data, obtaining a global alignment distance between any two nucleic acid sequences in the plurality of nucleic acid sequences, and generating an initial distance matrix of the standard nucleic acid sequence data;

[0009] obtaining a hypermetric deviation and function of the initial distance matrix;

[0010] iteratively optimizing the initial distance matrix to obtain an optimized distance matrix with minimum hypermetric deviation and function;

[0011] constructing a phylogenetic tree based on the optimized distance matrix.

[0012] In an embodiment, raw nucleic acid sequence data is obtained and processed to obtain standard nucleic acid sequence data, specifically including:

[0013] obtaining nucleic acid sequences of a species of interest, wherein the nucleic acid sequences of the species of interest include raw mitochondrial DNA sequences of the species of interest;

[0014] deleting nucleic acid sequences containing non-standard nucleotide bases from the nucleic acid sequences of the species of interest to obtain first nucleic acid sequence data;

[0015] deleting nucleic acid sequences containing unknown nucleotides from the first nucleic acid sequence data to obtain second nucleic acid sequence data;

[0016] deleting nucleic acid sequences with a length less than a preset length from the second nucleic acid sequence data to obtain third nucleic acid sequence data;

[0017] converting all nucleotides in the third nucleic acid sequence data to uppercase and deleting spaces and line breaks to obtain the standard nucleic acid sequence data.

[0018] In an embodiment, a number of nucleic acid sequences are extracted from the standard nucleic acid sequence data, global alignment distances between any two nucleic acid sequences in the number of nucleic acid sequences are obtained, and an initial distance matrix of the standard nucleic acid sequence data is generated, specifically including:

[0019] extracting nucleic acid sequences constituting 10-30% elements of the initial distance matrix from the standard nucleic acid sequence data;

[0020] selecting two nucleic acid sequences from the extracted nucleic acid sequences, obtaining a corresponding matching score matrix, and generating a backtracking path of the matching score matrix;

[0021] obtaining a global alignment distance of the two nucleic acid sequences based on the backtracking path;

[0022] repeating the above steps to obtain global alignment distances between any two nucleic acid sequences in the extracted nucleic acid sequences;

[0023] generating an initial distance matrix for all nucleic acid sequences in the standard nucleic acid sequence data, wherein for two nucleic acid sequences for which a global alignment distance has been obtained, the global alignment distance is filled as a known initial distance in the initial distance matrix, and for two nucleic acid sequences for which a global alignment distance has not been obtained, -1 is filled as an unknown initial distance in the initial distance matrix.

[0024] In one embodiment, two nucleic acid sequences are selected from the extracted nucleic acid sequences, a corresponding matching score matrix is obtained, and a backtracking path of the matching score matrix is generated, specifically comprising:

[0025] determining the matching score, the mismatch penalty and the gap penalty between the two nucleic acid sequences;

[0026] for two nucleic acid sequences and with lengths of and respectively, a score matrix with a size of is obtained, wherein each element of the matching score matrix is:

[0027] for , the gap penalty;

[0028] for , the gap penalty;

[0029] for ,

[0030] ;

[0031] wherein represents the element of the matching score matrix in the i-th row and the j-th column, represents the i-th nucleic acid on the nucleic acid sequence , represents the j-th nucleic acid on the nucleic acid sequence , when , is the matching score, otherwise is the mismatch penalty; from the element of the score matrix

[0032] to the element , the backtracking path is obtained, and the backtracking rule comprises: when

[0033] the left upper diagonal direction is backtracked when ;

[0034] when , the upper direction is backtracked when

[0035] ;​ When the left trace is reached, trace back to the left.

[0036] In an embodiment, based on the trace path, a global alignment distance of the two nucleic acid sequences is obtained, specifically comprising:

[0037] Based on the trace path, an aligned sequence corresponding to the two nucleic acid sequences is determined;

[0038] Based on the total length of the aligned sequence and the number of mismatches and gaps in the aligned sequence, a global alignment distance of the two nucleic acid sequences is obtained.

[0039] In an embodiment, a hypermetric deviation sum function of the initial distance matrix is obtained, specifically comprising:

[0040] Three nucleic acid sequences are extracted from the standard nucleic acid sequence data;

[0041] Based on the initial distance matrix, three groups of initial distances between the three nucleic acid sequences are obtained;

[0042] Based on the relationship between the three groups of initial distances and the triangle inequality, a hypermetric deviation of the three nucleic acid sequences is obtained;

[0043] Repeat the above steps to obtain the hypermetric deviation of any three nucleic acid sequences in the standard nucleic acid sequence data and sum them up to obtain the hypermetric deviation sum function.

[0044] In an embodiment, the initial distance matrix is iteratively optimized to obtain an optimized distance matrix with the minimum hypermetric deviation sum function, specifically comprising:

[0045] The elements in the initial distance matrix are initialized, the known initial distances are retained, and the unknown initial distances are replaced with the average value of the known initial distances;

[0046] The first moment estimate and the second moment estimate of the initial distance matrix are initialized to 0;

[0047] The first iteration is performed to calculate the current gradient that reduces the hypermetric deviation sum function;

[0048] Based on the current gradient, the first moment estimate and the second moment estimate are updated, and the bias correction of the first moment estimate and the second moment estimate is obtained;

[0049] Based on the bias correction, the updated unknown initial distance is obtained, and the updated distance matrix is obtained;

[0050] Repeat the iteration step until the change of the hypermetric deviation sum function of the updated distance matrix is less than a preset threshold, and the optimized distance matrix is obtained.

[0051] In an embodiment, the computing the current gradient that reduces the hypermetric deviation sum function further comprises:

[0052] regularizing the current gradient by weight decay to prevent overfitting.

[0053] In an embodiment, the constructing the phylogenetic tree based on the optimized distance matrix specifically comprises:

[0054] determining the distance sum between each nucleic acid sequence and other nucleic acid sequences based on the optimized distance matrix;

[0055] obtaining a corrected distance between any two nucleic acid sequences based on the distance sum of each nucleic acid sequence;

[0056] taking each nucleic acid sequence as an initial node of the phylogenetic tree, aggregating two nucleic acid sequences with the minimum corrected distance into a higher node, obtaining a distance between the two nucleic acid sequences with the minimum corrected distance and the higher node as a branch length of the two nucleic acid sequences on the phylogenetic tree;

[0057] obtaining distances between the higher node and other nucleic acid sequences, deleting the two nucleic acid sequences with the minimum corrected distance and adding the higher node in the optimized distance matrix, updating the distances between the higher node and other nucleic acid sequences to obtain an aggregated distance matrix;

[0058] repeating the aggregation process on the aggregated distance matrix until all nucleic acid sequences are aggregated to complete the construction of the phylogenetic tree.

[0059] The second aspect of the present application discloses a phylogenetic tree construction system based on a nucleic acid sequence distance matrix, comprising:

[0060] a preprocessing module, the preprocessing module is used for obtaining and processing original nucleic acid sequence data to obtain standard nucleic acid sequence data;

[0061] an initialization module, the initialization module is used for extracting a plurality of nucleic acid sequences from the standard nucleic acid sequence data, obtaining a global alignment distance between any two nucleic acid sequences in the plurality of nucleic acid sequences, and generating an initial distance matrix of the standard nucleic acid sequence data;

[0062] a computing module, the computing module is used for obtaining a hypermetric deviation sum function of the initial distance matrix;

[0063] an optimization module, the optimization module is used for iteratively optimizing the initial distance matrix to obtain an optimized distance matrix with the minimum hypermetric deviation sum function;

[0064] a building module for building a phylogenetic tree based on the optimized distance matrix.

[0065] The present application provides a phylogenetic tree building method and system based on nucleic acid sequence distance matrix, comprising the steps of: obtaining and processing original nucleic acid sequence data to obtain standard nucleic acid sequence data; extracting a plurality of nucleic acid sequences from the standard nucleic acid sequence data, obtaining the global alignment distance between any two nucleic acid sequences in the plurality of nucleic acid sequences, and generating an initial distance matrix of the standard nucleic acid sequence data; obtaining the hypermetric deviation sum function of the initial distance matrix; iteratively optimizing the initial distance matrix to obtain an optimized distance matrix with the minimum hypermetric deviation sum function; and building a phylogenetic tree based on the optimized distance matrix. By using a hybrid optimization technique, an effective nucleic acid sequence distance matrix is obtained, so that a phylogenetic tree can be accurately built even when the initial data is incomplete or too large in scale. BRIEF DESCRIPTION OF DRAWINGS

[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0067] Figure 1 A step flow chart of the phylogenetic tree building method based on nucleic acid sequence distance matrix in an embodiment of the present application;

[0068] Figure 2 A running schematic diagram of the phylogenetic tree building system based on nucleic acid sequence distance matrix in an embodiment of the present application;

[0069] Figure 3 A matching score matrix of two nucleic acid sequences in an embodiment of the present application;

[0070] Figure 4 An initial distance matrix obtained in an embodiment of the present application;

[0071] Figure 5 An optimized distance matrix obtained in an embodiment of the present application;

[0072] Figure 6 A change of the hypermetric deviation sum function in the iteration process of the distance matrix in an embodiment of the present application;

[0073] Figure 7 An optimized distance matrix obtained by using the conventional Needleman-Wunsch algorithm in an embodiment of the present application;

[0074] Figure 8 A phylogenetic tree for visual display in an embodiment of the present application;

[0075] Figure 9 A schematic diagram of the intelligent terminal of the present application. DETAILED DESCRIPTION

[0076] The present application provides a phylogenetic tree construction method based on a nucleic acid sequence distance matrix, to solve the problem of poor accuracy and large amount of calculation in reconstructing a complete distance matrix in the prior art, which affects the construction of a phylogenetic tree. In order to make the purpose, technical scheme and effect of the present application more clear and definite, the present application will be further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0077] In phylogenetic research and evolutionary biology, a distance matrix can generally be obtained by comparing nucleic acid sequences (such as DNA sequences) of different species, and a phylogenetic tree can be constructed based on the distance matrix, so as to describe the evolutionary relationship between different species, deduce their common ancestors and evolutionary paths, and reflect the evolutionary time or degree of change of each species through the branch length of the node of different species on the phylogenetic tree.

[0078] However, in actual application, the distance matrix between nucleic acid sequences often cannot be accurately obtained, and missing data or data errors often occur. This situation is often caused by the limitation of computing power in reality and sequencing reaction failure, which cannot be well solved under the limitation of the prior art. Therefore, when constructing a phylogenetic tree using a distance matrix of nucleic acid sequences, it is often necessary to optimize and reconstruct the distance matrix to ensure the smooth construction of the phylogenetic tree.

[0079] In the prior art, sequence alignment algorithms such as Needleman-Wunsch algorithm are generally used to optimize the distance matrix. Although these algorithms are effective for small data sets, the amount of calculation grows quadratically relative to the growth of data volume, so the amount of calculation of such algorithms will become very large for large-scale analysis. At the same time, other existing optimization distance matrix schemes often have great limitations, for example, deleting pairs of sequences lacking data to reduce the amount of calculation and perfecting the distance matrix, which is easy to lose key information related to evolutionary relationship; and using average value substitution or other attribution methods may introduce bias, thereby distorting the analysis results of the phylogenetic tree. Further, using a single algorithm for calculation has a huge requirement for the amount of calculation, and has too high a requirement for time and resources, and cannot maintain the basic biological properties of hypermetricity and distance additivity when constructing a phylogenetic tree, resulting in that the finally constructed phylogenetic tree is not reasonable in biology. Finally, the prior art has low utilization rate of data, and cannot effectively expand with the size of the genomic data set, resulting in insufficient universality.

[0080] The present application discloses a method and system for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix, and introduces a hybrid optimization method. First, the distance of a specific subset of nucleic acid sequence data is directly estimated using a sequence alignment algorithm, and then the distance of the remaining nucleic acid sequence data is estimated using a gradient-based optimization algorithm. This method fully utilizes the accuracy of direct calculation for manageable data subsets and the efficiency of optimization algorithm for the remaining data, thereby achieving more complete and accurate distance matrix optimization with less computational resources, and assisting in more accurate and efficient construction of a phylogenetic tree, especially when dealing with incomplete data or large genomic data.

[0081] Specifically, as shown in Figure 1 The method for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix comprises the following steps:

[0082] S100, obtaining and processing raw nucleic acid sequence data to obtain standard nucleic acid sequence data.

[0083] Raw nucleic acid sequence data often has various defects. By pre-processing the raw nucleic acid sequence data, high-quality and effective nucleic acid sequences are ensured for subsequent construction of distance matrix and phylogenetic tree, so as to ensure the reliability and accuracy of the final results. In an embodiment of the present application, mitochondrial DNA sequences are used as nucleic acid sequences, and standard DNA sequences are obtained after preprocessing for subsequent analysis and processing.

[0084] Specifically, obtaining and processing raw nucleic acid sequence data to obtain standard nucleic acid sequence data specifically comprises: obtaining nucleic acid sequences of a species of interest, wherein the nucleic acid sequences of the species of interest include raw mitochondrial DNA sequences of the species of interest. Specifically, mitochondrial DNA sequences of the species of interest are read from an input file in FASTA format as raw nucleic acid sequence data, and one mitochondrial DNA sequence corresponds to one species, i.e. the number of nucleic acid sequences in the raw nucleic acid sequence data corresponds to the number of species on the subsequently constructed phylogenetic tree.

[0085] After obtaining the raw nucleic acid sequence data, first, the nucleic acid sequences containing non-standard nucleotide bases in the nucleic acid sequences of the species of interest are deleted to obtain first nucleic acid sequence data. When mitochondrial DNA sequences are selected as nucleic acid sequences, standard nucleotide bases include adenine (A), cytosine (C), guanine (G) and thymine (T), and characters other than A, C, G and T in the nucleic acid sequence are identified as containing invalid characters and should be deleted.

[0086] After obtaining the first nucleic acid sequence data, nucleic acid sequences containing unknown nucleotides in the first nucleic acid sequence data are deleted to obtain second nucleic acid sequence data. Specifically, by detecting sequences with missing or ambiguous data in the first nucleic acid sequence data, such as unknown nucleotides represented by "N" or nucleotide sequences represented by spaces or other characters, they are separated from other complete nucleic acid sequences and deleted or the missing / ambiguous characters are replaced accordingly.

[0087] After obtaining the second nucleic acid sequence data, nucleic acid sequences with a length less than a preset length in the second nucleic acid sequence data are deleted to obtain third nucleic acid sequence data. By setting minimum sequence length and other standards for quality filtering, sequences that cannot provide meaningful information because they are too short or for other reasons are deleted, thereby improving the overall quality of the data set.

[0088] After obtaining the third nucleic acid sequence data, all nucleotides in the third nucleic acid sequence data are converted to uppercase and spaces and line breaks are deleted to obtain the standard nucleic acid sequence data. Specifically, by converting all nucleotides in the third nucleic acid sequence data to uppercase and deleting extra spaces or line breaks in the sequence, the format of the mitochondrial DNA sequence is standardized. The processed mitochondrial DNA sequence is finally written in FASTA format to a new output file for subsequent distance matrix and phylogenetic tree construction, improving the reliability and accuracy of the subsequent processing process.

[0089] Further, as shown in Figure 1 The phylogenetic tree construction method based on nucleic acid sequence distance matrix according to the present application further comprises, after step S100:

[0090] S200, extracting a plurality of nucleic acid sequences from the standard nucleic acid sequence data, obtaining a global alignment distance between any two nucleic acid sequences in the plurality of nucleic acid sequences, and generating an initial distance matrix of the standard nucleic acid sequence data.

[0091] Specifically, S100 ensures that the input nucleic acid sequence is in a standard format, which can be used for subsequent processing using bioinformatics tools. Alternatively, in step S200, the Needleman-Wunsch algorithm is used for global sequence alignment of a specific number of nucleic acid sequences in the standard nucleic acid sequence to obtain the corresponding distance matrix.

[0092] Specifically, step S200 comprises:

[0093] S210, extracting nucleic acid sequences constituting 10-30% of the elements of the initial distance matrix from the standard nucleic acid sequence data;

[0094] S220, selecting two nucleic acid sequences from the extracted nucleic acid sequences, obtaining a corresponding matching score matrix, and generating a backtracking path of the matching score matrix;

[0095] S230, obtaining a global alignment distance of the two nucleic acid sequences based on the backtracking path;

[0096] S240, repeating the above steps to obtain a global alignment distance between any two nucleic acid sequences in the extracted nucleic acid sequences;

[0097] S250, generating an initial distance matrix for all nucleic acid sequences in the standard nucleic acid sequence data, wherein for two nucleic acid sequences for which a global alignment distance has been obtained, the global alignment distance is filled into the initial distance matrix as a known initial distance, and for two nucleic acid sequences for which a global alignment distance has not been obtained, -1 is filled into the initial distance matrix as an unknown initial distance.

[0098] Wherein, the number of elements in the final expected initial distance matrix is determined according to the number of extracted nucleic acid sequences, when the number of species of interest is small, i.e. the number of nucleic acid sequences in the standard nucleic acid sequence data is small, the number of elements in the initial distance matrix is small, and nucleic acid sequences corresponding to a larger proportion (such as 30%) of the number of elements can be selected for subsequent processing; and when the number of species of interest is large, i.e. the number of nucleic acid sequences in the standard nucleic acid sequence data is large, the number of elements in the initial distance matrix is large, and nucleic acid sequences corresponding to a smaller proportion (such as 10%) of the number of elements can be selected for subsequent processing to reduce the amount of calculation.

[0099] Further, step S220 specifically includes:

[0100] determining the matching score, the mismatch penalty and the gap penalty between the two nucleic acid sequences, wherein the matching score is a positive score corresponding to the same nucleic acid at the same position of the two nucleic acid sequences; the mismatch penalty and the gap penalty are negative scores and equal, wherein the mismatch penalty corresponds to different nucleic acids at the same position of the two nucleic acid sequences, and the gap penalty corresponds to the misalignment between the positions of the same nucleic acid of the two nucleic acid sequences, and the matching score, the mismatch penalty and the gap penalty are used to reflect the matching and non-matching relationship between the biological characteristics of the nucleic acid sequences. Alternatively, the matching score is +5 points, and the mismatch penalty and the gap penalty are -4 points.

[0101] After determining the matching score, the mismatch penalty and the gap penalty, for two nucleic acid sequences with lengths of and , a score matrix with a size of is obtained. ​​, where the matching score matrix The elements of are:

[0102] for , open field penalty;

[0103] for , open field penalty;

[0104] for ,

[0105] ;

[0106] in Represents the matching score matrix Middle Rank Elements of the column, Indicates nucleic acid sequence Previous nucleic acids, Indicates nucleic acid sequence Previous nucleic acid, when hour, is the matching score, otherwise is the mismatch penalty. Through the above dynamic steps, it can be ensured that the two nucleic acid sequences obtained in the end and The matching score matrix between them accurately reflects the sequence alignment relationship between the two, so as to facilitate the subsequent selection of the best backtracking path and form the optimal alignment sequence.

[0107] Get the matching score matrix Then, from the score matrix Elements Back to element , get the backtracking path, that is, from the matching score matrix The bottom right corner element is traced back to the top left corner element. The backtracking rules are:

[0108] when Time-sharing, backtracking to the upper left diagonal direction, at this time the nucleic acid sequence Previous nucleic acids and nucleic acid sequences Previous nucleic acids correspond;

[0109] when When tracing back upward, the nucleic acid sequence Previous nucleic acids corresponds to a gap on the nucleic acid sequence ;

[0110] When , backtrack to the left, at this time the gap on the nucleic acid sequence corresponds to the first nucleic acid on the nucleic acid sequence .

[0111] By repeating the above steps, eventually backtrack to the element , and get the alignment sequence corresponding to the backtracking path. Through the above method, all feasible backtracking paths can be obtained, and by limiting the backtracking path, an optimal alignment sequence can be found.

[0112] Further, step S230 specifically comprises:

[0113] Based on the backtracking path, determine the alignment sequence corresponding to the two nucleic acid sequences;

[0114] Based on the total length of the alignment sequence and the number of mismatches and gaps in the alignment sequence, obtain the global alignment distance of the two nucleic acid sequences.

[0115] Specifically, in combination with the backtracking rule corresponding to each step of the backtracking path, the two nucleic acid sequences are matched and aligned to obtain the corresponding alignment sequence. Alternatively, the optimization effect of the alignment sequence is evaluated by the element in the lower right corner of the score matrix , i.e. . Through the element , the alignment sequence with the highest score can be obtained under the definition of the present application, so as to find the optimal alignment sequence corresponding to the two nucleic acid sequences. After obtaining the optimal alignment sequence, the distance between the two nucleic acid sequences is evaluated by the proportion of mismatches and gaps in the alignment sequence, i.e. the global alignment distance between the two nucleic acid sequences

[0116]

[0117] , wherein the number of mismatches and gaps is the total number of mismatched nucleic acids and matched gap nucleic acids in the optimal alignment sequence, and the length of the alignment sequence is the total number of matched nucleic acids, mismatched nucleic acids and matched gap nucleic acids. It can be seen that the global alignment distance between the two nucleic acid sequences is 0 (the same nucleic acid sequence) to 1 (two completely different nucleic acid sequences).

[0118] ​​Thereafter, steps S220 and S230 are repeated to obtain the global alignment distance between any two of the extracted nucleic acid sequences and fill in the initial distance matrix, wherein each element of the initial distance matrix corresponds to the initial distance between two nucleic acid sequences. Thus, in the initial distance matrix, for the same nucleic acid sequence, the corresponding initial distance is 0; for two nucleic acid sequences for which the global alignment distance has been obtained, the corresponding initial distance is the global alignment distance obtained by steps S220 and S230; for two nucleic acid sequences for which the global alignment distance has not been obtained, i.e. one of the 70%-90% of nucleic acid sequences not selected and any other nucleic acid sequence, the initial distance is an unknown initial distance and is preset to -1. In this way, a distance matrix with partial accurate distance data is obtained, providing a basis for the subsequent optimization process, which not only ensures the progress of the subsequent optimization process, but also avoids excessive computational load and effectively reduces the demand for resources.

[0119] Further, as shown in Figure 1 , the phylogenetic tree construction method based on the nucleic acid sequence distance matrix of the present application further comprises, after step S200:

[0120] S300, obtaining the hypermetric deviation sum function of the initial distance matrix.

[0121] In order to evaluate and improve the quality of the initial distance matrix obtained in step S200, the present application uses the hypermetric deviation sum function (Delta sum function) to quantify the degree to which the distance matrix satisfies the hypermetric property. Specifically, we assume that the triangle formed by the distance between any two of the three species is an isosceles right triangle in the ideal state, i.e. any two species are expected to have similar genetic distances from the third species. Therefore, the isosceles right triangle is the best state, and the hypermetric property is to evaluate whether the distances from all nodes to the root are equal, to evaluate whether the evolutionary rate is uniform. The distance between any three nucleic acid sequences in the standard nucleic acid sequence data is measured to obtain the hypermetric deviation sum function of the corresponding three nucleic acid sequences, and the optimization of the hypermetric deviation sum function can achieve the optimization of the initial distance matrix, thereby facilitating the further construction of the phylogenetic tree based on the distance matrix.

[0122] Specifically, in step S300, three nucleic acid sequences are first extracted from the standard nucleic acid sequence data. Alternatively, three nucleic acid sequences are extracted as an example to avoid repeated extraction, i.e. to ensure that and , wherein takes a value of 1 to , takes a value of i +1 to , and the value of j +1 to , is the total number of nucleic acid sequences in the standard nucleic acid sequence data.

[0123] Then, for the extracted three nucleic acid sequences, based on the initial distance matrix, three groups of initial distances between the three nucleic acid sequences are obtained. Wherein, for the initial distance matrix obtained from step S200 D , the three groups of initial distances between the three nucleic acid sequences are respectively . .

[0124] After obtaining the three groups of initial distances, based on the relationship between the three groups of initial distances and the triangle inequality, the hypermetric deviation of the three nucleic acid sequences is obtained. Specifically, first, the three groups of initial distances are arranged in descending order according to length; and the three groups of initial distances are compared with the triangle inequality as the three sides of a triangle, and the hypermetric deviation of the three nucleic acid sequences is determined by whether the three groups of initial distances conform to the triangle inequality. Specifically, the triangle inequality is , wherein can be any three groups of initial distances ;

[0125] When the condition occurs between the three groups of initial distances, it means that the three groups of initial distances completely violate the triangle inequality, and the hypermetric deviation of the three nucleic acid sequences is:

[0126]

[0127] wherein, is the hypermetric deviation of the three nucleic acid sequences , , is a constant;

[0128] When the three groups of initial distances conform to the triangle relationship, the cosine values of the three internal angles of the triangle composed of the three groups of initial distances are calculated:

[0129]

[0130] wherein, corresponds to the longest side, corresponds to the second longest side, and corresponds to the shortest side;

[0131] If all the calculated cosine values fall within the effective range , then the sizes of the three internal angles are calculated by the inverse cosine function, and the hypermetric deviation of the three nucleic acid sequences is:

[0132]

[0133] If the cosine value obtained is not within the effective range , set the hypermetric deviation of the three nucleic acid sequences as .

[0134] Repeat the steps of extracting three nucleic acid sequences from the standard nucleic acid sequence data and calculating the hypermetric deviation corresponding to the three nucleic acid sequences, obtain the hypermetric deviation of any three nucleic acid sequences in the standard nucleic acid sequence data and sum them up, and obtain the hypermetric deviation sum function. Specifically, for the obtained hypermetric deviation , the final hypermetric deviation sum function is:

[0135] ,

[0136] wherein D corresponds to the initial distance matrix, corresponds to any three nucleic acid sequences in the standard nucleic acid sequence data, and the total number of nucleic acid sequences in the standard nucleic acid sequence data.

[0137] Further, as shown in Figure 1 , the phylogenetic tree construction method based on the nucleic acid sequence distance matrix of the present application further comprises the following steps after step S300:

[0138] S400, iteratively optimizing the initial distance matrix to obtain an optimized distance matrix with the minimum hypermetric deviation sum function.

[0139] In the present application, the purpose of optimizing the initial distance matrix is to adjust the missing distance in the initial distance matrix, i.e., the value of the unknown initial distance, to minimize the total hypermetric deviation of the distance matrix, thereby enhancing the hypermetric property of the matrix, wherein the optimization process can be represented as:

[0140]

[0141] wherein, corresponds to the optimized optimized distance matrix, corresponds to the hypermetric deviation sum function of the initial distance matrix . In order to ensure that the optimized distance matrix remains valid and can accurately represent the distance between nucleic acid sequences, the element corresponding to the distance between any two nucleic acid sequences in the optimized distance matrix is subject to the following conditions:

[0142] Non-negativity: for all , ;

[0143] Zero diagonal: for all , ;

[0144] Symmetry: For all , .

[0145] Under the above constraints, the present invention uses the Adam optimization algorithm to iteratively update the initial distance in the initial distance matrix. The optimization update process can be summarized as follows:

[0146] Initializing the elements in the initial distance matrix, retaining the known initial distances, and replacing the unknown initial distances with the average value of the known initial distances;

[0147] Initializing the first-order moment estimate and the second-order moment estimate of the initial distance matrix to 0;

[0148] Performing a first iteration, calculating a current gradient that reduces the ultrametric deviation sum function;

[0149] Based on the current gradient, updating the first-order moment estimate and the second-order moment estimate, and obtaining a bias correction for the first-order moment estimate and the second-order moment estimate;

[0150] Based on the deviation correction, an updated unknown initial distance is obtained to obtain an updated distance matrix.

[0151] Specifically, from iteration number t=0 to T (maximum iteration number), first replace the element corresponding to the unknown initial distance in the initial distance matrix (-1 when t=0) with the average value of the known initial distance, and set the initialized first-order moment estimate and second-order moment estimates Then the gradient of the optimization step is calculated using the finite difference method, and the current gradient of the optimization step is approximated by evaluating the change caused by a small change in the distance. Specifically, for the unknown initial distance in the upper half of the initial distance matrix (i.e., the elements in the upper right triangle of the initial distance matrix that maintains symmetry, where , the elements in the lower left triangle can be obtained by symmetry), in the The current gradient at iteration Approximately:

[0152]

[0153] in It is The distance matrix at the iteration, is a small constant (e.g. ), exist a matrix with 1 at the position and 0 elsewhere, is the initial distance, a matrix with 1 at the position and 0 elsewhere, to calculate how the function should be changed in order to reduce the deviation of the hypermetric .

[0154] Specifically, in approximating the current gradient, the current gradient is also regularized by weight decay to prevent overfitting. Specifically, the regularization process is:

[0155]

[0156] wherein, is the weight decay coefficient (for example, take ), through this regularization term, the penalty for too large distance change can be imposed, so as to ensure the consistency and reality of the current gradient.

[0157] After approximating the current gradient, in the th iteration, update the first moment estimate and the second moment estimate:

[0158]

[0159]

[0160] wherein, and are the first moment estimate and the second moment estimate of the th element updated in the th iteration, and are the decay rates (corresponding to the momentum parameter and the mean square gradient parameter respectively), which can be taken as and .

[0161] Based on the updated first moment estimate and the second moment estimate, the deviation correction of the first moment estimate and the second moment estimate is calculated:

[0162]

[0163] wherein, and are the deviation corrections of the first moment estimate and the second moment estimate in the th iteration.

[0164] Based on the deviation correction, the unknown initial distance is updated as:

[0165]

[0166] wherein the learning rate , the bias correction term is a small constant to avoid division by zero. It is noted that the non-negativity and symmetry of the distance matrix are always guaranteed in the process of updating the unknown initial distance, i.e.

[0167]

[0168] After one iteration to obtain the updated unknown initial distance, the hypermetric deviation sum function of the current distance matrix can be calculated by step S300 .

[0169] The above iteration and updating steps are repeated until the change of the hypermetric deviation sum function of the updated distance matrix is less than a preset threshold, and the optimized distance matrix is obtained. The end condition can be expressed as:

[0170]

[0171] The preset threshold can be set as needed, for example, set to 0.1.

[0172] The hypermetric deviation sum function of the distance matrix is smaller, which means that the distance matrix has better hypermetric property, and the distance matrix corresponding to the final end condition hypermetric deviation sum function is the final optimized distance matrix.

[0173] Further, the original complete matrix obtained by the Needleman-wunsh algorithm and the root mean square error (RMSE) can be used to evaluate the relative deviation of the optimized distance matrix:

[0174]

[0175] wherein is an element of the original complete matrix, is an element of the optimized distance matrix.

[0176] The optimized distance matrix is obtained by optimizing the hypermetric deviation sum function of the distance matrix, so as to ensure that the distance matrix used to construct the phylogenetic tree has accurate and complete data, thereby ensuring efficient and accurate construction of the phylogenetic tree.

[0177] Further, as shown in Figure 1 , the phylogenetic tree construction method based on the nucleic acid sequence distance matrix of the present application further comprises, after step S400:

[0178] ​​S500, constructing a phylogenetic tree based on the optimized distance matrix.

[0179] Specifically, each element in the optimized distance matrix is extracted as the distance between each nucleic acid sequence in the standard nucleic acid sequence data, and a phylogenetic tree is constructed using the Neighbor-Joining (NJ) algorithm. In summary, step S500 specifically includes:

[0180] Based on the optimized distance matrix, the distance sum between each nucleic acid sequence and other nucleic acid sequences is determined.

[0181] Based on the distance sum of each nucleic acid sequence, the corrected distance between any two nucleic acid sequences is obtained.

[0182] Taking each nucleic acid sequence as an initial node of the phylogenetic tree, the two nucleic acid sequences with the smallest corrected distance are aggregated into a higher node, and the distance between the two nucleic acid sequences with the smallest corrected distance and the higher node is obtained as the branch length of the two nucleic acid sequences on the phylogenetic tree.

[0183] The distance between the higher node and other nucleic acid sequences is obtained, the two nucleic acid sequences with the smallest corrected distance are deleted in the optimized distance matrix and the higher node is added, and the distance between the higher node and other nucleic acid sequences is updated to obtain an aggregated distance matrix.

[0184] The aggregation process is repeated for the aggregated distance matrix until all nucleic acid sequences are aggregated, and the phylogenetic tree is constructed.

[0185] Specifically, the present application uses the NJ algorithm to construct a phylogenetic tree by gradually aggregating the nearest two nodes, while minimizing the total branch length (total distance) of the tree to reflect the evolutionary relationship between species. In an embodiment of the present application, each nucleic acid sequence in the standard nucleic acid sequence data is aggregated as a node of the phylogenetic tree, and the distance between each nucleic acid sequence in the distance matrix is used to construct the phylogenetic tree, wherein each nucleic acid sequence corresponds to a species of interest, so that a phylogenetic tree about the evolutionary relationship between species of interest can be constructed.

[0186] In one embodiment, the sum of the distances between each nucleic acid sequence and other nucleic acid sequences is first calculated:

[0187]

[0188] wherein, is the sum of the distances between the nucleic acid sequence and other nucleic acid sequences, N is the total number of nucleic acid sequences in the standard nucleic acid sequence data, is the nucleic acid sequence between the other nucleic acid sequences in the optimized distance matrix.

[0189] Then, based on the distance of each nucleic acid sequence, the corrected distance between a pair of nodes on the phylogenetic tree is calculated:

[0190]

[0191] wherein, d is the corrected distance, and d is the distance between the pair of nucleic acid sequences in the conventional NJ algorithm, and d is the distance between the pair of nucleic acid sequences in the optimized distance matrix. The corrected distance obtained in the present application can accurately determine the pair of nodes with the shortest distance on the phylogenetic tree. After finding the pair of nodes with the minimum corrected distance d, the distance between the nucleic acid sequences and the new node (i.e. the upper node) is:

[0192] The distance obtained is the branch length between the node and the upper node on the phylogenetic tree, to ensure the additivity of the finally constructed phylogenetic tree.

[0193]

[0194] The distance obtained is the branch length between the node and the upper node on the phylogenetic tree, to ensure the additivity of the finally constructed phylogenetic tree.

[0195] After determining the new node obtained by the new aggregation, the distance between the node and other nodes is:

[0196]

[0197] wherein, d corresponds to a nucleic acid sequence other than the nucleic acid sequences and in the standard nucleic acid sequence data. In this way, the elements corresponding to the distance between the nucleic acid sequences and other nucleic acid sequences in the optimized distance matrix are removed, and the corresponding elements corresponding to the distance between the node and other nucleic acid sequences are added to form a new optimized distance matrix, and enter the next aggregation process.

[0198] ​​​​​​​​​​​​​​​​​​​​The final phylogenetic tree is obtained by repeatedly aggregating the nodes with the minimum correction distance until all the nucleic acid sequences in the standard nucleic acid sequence data are aggregated. Optionally, the present application uses a drawing tool to visually construct the phylogenetic tree, thereby improving the clarity and presentation effect, and better displaying the evolutionary relationship between the species of interest.

[0199] The phylogenetic tree construction method based on the nucleic acid sequence distance matrix of the present application uses a hybrid optimization technique to effectively estimate the DNA distance matrix. The sequence alignment algorithm (Needleman-Wunsch algorithm) and the advanced optimization method (Adma optimization algorithm) are seamlessly integrated. By directly calculating a specific proportion (10%-30%) of elements (corresponding to the distance between a pair of nucleic acid sequences) in the distance matrix in the data set, and estimating the remaining proportion (70%-90%) of elements in the distance matrix through the optimization process, the phylogenetic tree is efficiently and accurately constructed. Compared with the phylogenetic tree construction methods in the prior art, the phylogenetic tree construction method based on the nucleic acid sequence distance matrix of the present application can significantly reduce the computational requirements, and has the following advantages: preserving key genetic data without introducing bias, improving the accuracy of phylogenetic tree analysis; reducing the dependence on computationally intensive algorithms, suitable for large-scale genomic research; ensuring basic biological constraints such as ultrametricity and additivity, forming a phylogenetic tree with biological rationality; expanding the applicability in different data analysis according to the size of the genomic data set; and reconstructing the complete distance matrix from partial distance information, achieving more comprehensive and accurate analysis.

[0200] Further, based on the phylogenetic tree construction method based on the nucleic acid sequence distance matrix, a phylogenetic tree construction system based on the nucleic acid sequence distance matrix is disclosed to realize the corresponding steps, such as Figure 2 As shown, the system comprises:

[0201] A preprocessing module is configured to obtain and process the original nucleic acid sequence data to obtain standard nucleic acid sequence data.

[0202] An initialization module is configured to extract a plurality of nucleic acid sequences from the standard nucleic acid sequence data, obtain the global alignment distance between any two nucleic acid sequences in the plurality of nucleic acid sequences, and generate an initial distance matrix of the standard nucleic acid sequence data.

[0203] A calculation module is configured to obtain the ultrametric bias and function of the initial distance matrix.

[0204] An optimization module is configured to iteratively optimize the initial distance matrix to obtain an optimized distance matrix with the minimum ultrametric bias and function.

[0205] A construction module is used to construct a phylogenetic tree based on the optimized distance matrix.

[0206] Specifically, if Figure 2 As shown, the preprocessing module first performs a rapid test on the sequence validity of the input nucleic acid sequence data, including deleting nucleic acid sequences containing non-standard nucleotide bases, deleting nucleic acid sequences containing unknown nucleotides, and deleting nucleic acid sequences with a length less than a preset length; then the data is formatted and normalized into FASTA format; finally, the nucleic acid sequence data in the output file is tested for missing data and quality screening to ensure the reliability and accuracy of subsequent processing. Furthermore, after receiving the data output by the preprocessing module, the initialization module extracts part of the nucleic acid sequence from the data and calculates the initial distance using the Needleman-Wunsch algorithm, thereby obtaining an initial distance matrix with approximately 10%-30% of the known initial distance. The calculation module then calculates the supermetric deviation and function using the initial distance matrix output by the initialization module, and evaluates the supermetricity of the initial distance matrix based on the calculation results. The optimization module optimizes the supermetric deviation and function obtained by the calculation module, minimizes the supermetric deviation and function using the Adam optimization algorithm, thereby retaining the known distances in the distance matrix and updating the unknown distances to obtain an updated distance matrix, and checks whether the change in the supermetric deviation and function between the updated distance matrix and the current distance matrix is ​​less than a preset threshold. When the change in the supermetric deviation and function is greater than the preset threshold, the calculation module recalculates the supermetric deviation and function based on the updated distance matrix, and then the optimization module repeats the optimization steps until the change in the supermetric deviation and function is less than the preset threshold, thereby obtaining an optimized distance matrix. Finally, the construction module constructs a phylogenetic tree after obtaining a complete optimized distance matrix, and performs a visual display based on the constructed phylogenetic tree data to obtain a final phylogenetic tree.

[0207] Based on the above-mentioned method and system for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix, the present invention also provides an intelligent terminal, such as Figure 9 As shown, the intelligent terminal includes at least one processor, a display screen, and a memory, and may also include a communication interface and a bus. The processor, display screen, memory, and communication interface can communicate with each other via the bus. The display screen is configured to display a user guidance interface preset in an initial setup mode. The communication interface can transmit information. The processor can call logic instructions in the memory to execute the phylogenetic tree construction method based on a nucleic acid sequence distance matrix described in the present invention.

[0208] The following describes the application scenarios and results of the method for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix according to the present invention with reference to specific examples:

[0209] In the embodiment, the mitochondrial DNA sequences of 10 different monkeys constitute the original nucleic acid sequence data, which ensures that the mitochondrial DNA sequences of the 10 monkeys only contain adenine (A), cytosine (C), guanine (G) and thymine (T) four nucleotides, and are complete nucleic acid sequences, the length meets the minimum length requirement, all the nucleotide symbols in the mitochondrial DNA sequences are unified as capital letters, and the extra spaces and line breaks are deleted. Finally, the standard nucleic acid sequence data in FASTA format is output.

[0210] The initial distance matrix composed of 10 mitochondrial DNA sequences includes a total of 100 distance elements (10x10), 30% of which are extracted, and the global alignment distance is obtained by selecting two mitochondrial DNA sequences corresponding to each element.

[0211] Taking the nucleic acid sequence =ACGTACGT and the nucleic acid sequence =ACGTCGTT as an example, the corresponding global alignment distance is obtained, wherein the matching score is set to +5, the mismatch penalty is set to -4, and the gap penalty is set to -4. For the nucleic acid sequence and , a 9x9 matching score matrix is constructed, and the elements in the matrix satisfy the following rules:

[0212] ;

[0213] The obtained matching score matrix is shown in Figure 3 , and based on the matching score matrix, the alignment sequence corresponding to the optimal backtracking path is obtained ACGTACG-T, and the alignment sequence CGT-CGTT, wherein the horizontal bar corresponds to the gap in the matching process. It can be seen that the number of matches in the alignment sequence is 7, the number of mismatches is 0, the number of gaps is 2, and the total length is 9. Therefore, the global alignment distance between the nucleic acid sequences and is:

[0214]

[0215] Repeat the above steps to calculate the global alignment distance between any two nucleic acid sequences in the extracted nucleic acid sequences, and fill the global alignment distance as the initial distance into the initial distance matrix , as shown in Figure 4 , the initial distance matrix The elements in the initial distance matrix are the initial distances between the extracted nucleic acid sequences Seq1-Seq10 and the nucleic acid sequences Seq1-Seq10, wherein the initial distance between the same nucleic acid sequences is 0, the initial distance between the extracted nucleic acid sequences is the global alignment distance calculated by the above steps, and the initial distance between the remaining nucleic acid sequences is set to -1.

[0216] Then the unknown initial distances (-1) in the initial distance matrix are optimized to obtain an optimized distance matrix, wherein the unknown initial distances are first replaced by the average of the known initial distances, and then the iterative optimization of the distance matrix is realized by using the ultrametric deviation sum function of the present application, wherein the distance matrix that minimizes the ultrametric deviation sum function is the optimized distance matrix. Specifically, as shown in Figure 5 the optimized distance matrix obtained after 1200 iterations is shown, wherein the elements correspond to the distances between any two nucleic acid sequences among the nucleic acid sequences Seq1-Seq10, as shown in Figure 6 the ultrametric deviation sum function values of the distance matrix obtained after every 100 iterations in the iterative optimization process can be seen. The ultrametric deviation sum function of the optimized distance matrix obtained after 1200 iterations is reduced from 218.2 of the initial distance matrix to 2.297. As shown in Figure 7 the distance matrix obtained by using the conventional Needleman-Wunsch algorithm, and the ultrametric deviation sum function thereof is 3.646. The phylogenetic tree construction method based on the nucleic acid sequence distance matrix of the present application has better ultrametricity, higher integrity, and is more in line with biological rationality than the traditional algorithm, and can more accurately reflect the evolutionary relationship, thereby more efficiently and accurately constructing the phylogenetic tree.

[0217] In this embodiment, the NJ algorithm is used to construct the phylogenetic tree based on the above obtained optimized distance matrix, and the visualization is shown as the phylogenetic tree shown in Figure 8 , thereby showing the evolutionary relationship between the 10 monkeys.

[0218] In summary, the present application provides a phylogenetic tree construction method and system based on nucleic acid sequence distance matrix, comprising the steps of: obtaining and processing original nucleic acid sequence data to obtain standard nucleic acid sequence data; extracting a plurality of nucleic acid sequences from the standard nucleic acid sequence data, obtaining the global alignment distance between any two nucleic acid sequences in the plurality of nucleic acid sequences, and generating an initial distance matrix of the standard nucleic acid sequence data; obtaining the hypermetric deviation sum function of the initial distance matrix; iteratively optimizing the initial distance matrix to obtain an optimized distance matrix with the minimum hypermetric deviation sum function; and constructing a phylogenetic tree based on the optimized distance matrix. By using a hybrid optimization technique to obtain an effective nucleic acid sequence distance matrix, the phylogenetic tree can be accurately constructed even when the initial data is incomplete or too large in scale.

[0219] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand; the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not deviate from the spirit and scope of the corresponding technical solutions, and should be included in the protection scope of the present application.

Claims

1. A method for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix, characterized in that: Including steps: Acquiring and processing raw nucleic acid sequence data to obtain standard nucleic acid sequence data; extracting a plurality of nucleic acid sequences from the standard nucleic acid sequence data, obtaining a global alignment distance between any two nucleic acid sequences among the plurality of nucleic acid sequences, and generating an initial distance matrix for the standard nucleic acid sequence data; Obtaining a supermetric deviation and function of the initial distance matrix; Iteratively optimizing the initial distance matrix to obtain an optimized distance matrix with the minimum supermetric deviation and function; constructing a phylogenetic tree based on the optimized distance matrix; The step of obtaining the supermetric deviation and function of the initial distance matrix specifically includes: Extracting three nucleic acid sequences from the standard nucleic acid sequence data; Based on the initial distance matrix, obtaining three sets of initial distances between the three nucleic acid sequences; Obtaining the supermetric deviations of the three nucleic acid sequences based on the relationship between the three sets of initial distances and the triangle inequality; Repeat the above steps to obtain the supermetric deviations of any three nucleic acid sequences in the standard nucleic acid sequence data and sum them to obtain the supermetric deviation sum function. The supermetric deviation sum function is: , in D Corresponding to the initial distance matrix, Corresponding to any three nucleic acid sequences in the standard nucleic acid sequence data, is the supermetric deviation, is the total number of nucleic acid sequences in the standard nucleic acid sequence data.

2. The method for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix according to claim 1, wherein: Obtaining and processing raw nucleic acid sequence data to obtain standard nucleic acid sequence data, specifically including: Obtaining a nucleic acid sequence of a species of interest, wherein the nucleic acid sequence of the species of interest includes an original mitochondrial DNA sequence of the species of interest; Deleting the nucleic acid sequence containing non-standard nucleotide bases in the nucleic acid sequence of the species of interest to obtain first nucleic acid sequence data; deleting the nucleic acid sequence containing unknown nucleotides in the first nucleic acid sequence data to obtain second nucleic acid sequence data; Deleting nucleic acid sequences whose length is less than a preset length in the second nucleic acid sequence data to obtain third nucleic acid sequence data; All nucleotides in the third nucleic acid sequence data are converted to uppercase and spaces and line breaks are deleted to obtain the standard nucleic acid sequence data.

3. The method for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix according to claim 1, wherein: Extracting a plurality of nucleic acid sequences from the standard nucleic acid sequence data, obtaining a global alignment distance between any two nucleic acid sequences in the plurality of nucleic acid sequences, and generating an initial distance matrix for the standard nucleic acid sequence data specifically includes: Extracting nucleic acid sequences that constitute 10%-30% of the elements of the initial distance matrix from the standard nucleic acid sequence data; Selecting two nucleic acid sequences from the extracted nucleic acid sequences, obtaining corresponding matching score matrices, and generating a backtracking path of the matching score matrices; Based on the backtracking path, obtaining a global alignment distance between the two nucleic acid sequences; Repeat the above steps to obtain the global alignment distance between any two nucleic acid sequences in the extracted nucleic acid sequences; An initial distance matrix is ​​generated for all nucleic acid sequences in the standard nucleic acid sequence data, wherein for two nucleic acid sequences for which a global alignment distance has been obtained, the global alignment distance is filled in the initial distance matrix as a known initial distance; and for two nucleic acid sequences for which a global alignment distance has not been obtained, -1 is filled in the initial distance matrix as an unknown initial distance.

4. The method for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix according to claim 3, wherein: Selecting two nucleic acid sequences from the extracted nucleic acid sequences, obtaining corresponding matching score matrices, and generating a backtracking path of the matching score matrices, specifically including: Determining a match score, mismatch penalty, and gap penalty between two nucleic acid sequences; For the lengths and Two nucleic acid sequences and , get the size The matching score matrix , where the matching score matrix The elements of are: for , open field penalty; for , open field penalty; for , ; in Represents the matching score matrix Middle Rank Elements of the column, Indicates nucleic acid sequence Previous nucleic acids, Indicates nucleic acid sequence Previous nucleic acid, when hour, is the matching score, otherwise Penalty points for mismatches; From the matching score matrix Elements Back to element , the backtracking path is obtained, and the backtracking rules include: when Time sharing, retrace to the upper left diagonal direction; when When, trace back upward; when , backtrack to the left.

5. The method for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix according to claim 3, wherein: Based on the backtracking path, obtaining the global alignment distance of the two nucleic acid sequences specifically includes: Based on the backtracking path, determining the alignment sequence corresponding to the two nucleic acid sequences; Based on the total length of the aligned sequences and the number of mismatches and gaps in the aligned sequences, a global alignment distance of the two nucleic acid sequences is obtained.

6. The method for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix according to claim 3, wherein: Iteratively optimizing the initial distance matrix to obtain an optimized distance matrix with the minimum supermetric deviation and function, specifically including: Initializing the elements in the initial distance matrix, retaining the known initial distances, and replacing the unknown initial distances with the average value of the known initial distances; Initializing the first-order moment estimate and the second-order moment estimate of the initial distance matrix to 0; Performing a first iteration, calculating a current gradient that reduces the ultrametric deviation sum function; Based on the current gradient, updating the first-order moment estimate and the second-order moment estimate, and obtaining a bias correction for the first-order moment estimate and the second-order moment estimate; Based on the deviation correction, obtaining an updated unknown initial distance to obtain an updated distance matrix; The iterative steps are repeated until the supermetric deviation and function change of the updated distance matrix is ​​less than a preset threshold, thereby obtaining the optimized distance matrix.

7. The method for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix according to claim 6, wherein: After calculating the current gradient that reduces the supermetric deviation and function, the method further includes: The current gradient is regularized by weight decay to prevent overfitting.

8. The method for constructing a phylogenetic tree based on a nucleic acid sequence distance matrix according to claim 1, wherein: Based on the optimized distance matrix, a phylogenetic tree is constructed, specifically comprising: Determining the sum of the distances between each nucleic acid sequence and other nucleic acid sequences based on the optimized distance matrix; Based on the sum of the distances of the nucleic acid sequences, obtaining the corrected distance between any two nucleic acid sequences; Taking each nucleic acid sequence as an initial node of the phylogenetic tree, aggregating the two nucleic acid sequences with the smallest corrected distances into a parent node, and obtaining the distance between the two nucleic acid sequences with the smallest corrected distances and the parent node as the branch length of the two nucleic acid sequences with the smallest corrected distances on the phylogenetic tree; Obtaining the distance between the parent node and other nucleic acid sequences, deleting the two nucleic acid sequences with the smallest corrected distances in the optimized distance matrix and adding the parent node, and updating the distance between the parent node and other nucleic acid sequences to obtain a post-aggregation distance matrix; The aggregation process is repeated for the aggregated distance matrix until all nucleic acid sequences are aggregated, thereby completing the construction of the phylogenetic tree.

9. A phylogenetic tree construction system based on nucleic acid sequence distance matrix, characterized in that: include: A preprocessing module, which is used to obtain and process raw nucleic acid sequence data to obtain standard nucleic acid sequence data; an initialization module, the initialization module being used to extract a plurality of nucleic acid sequences from the standard nucleic acid sequence data, obtain a global alignment distance between any two nucleic acid sequences among the plurality of nucleic acid sequences, and generate an initial distance matrix for the standard nucleic acid sequence data; A calculation module, the calculation module is used to obtain the supermetric deviation and function of the initial distance matrix; An optimization module, configured to iteratively optimize the initial distance matrix to obtain an optimized distance matrix with a minimum supermetric deviation and function; A construction module, wherein the construction module is used to construct a phylogenetic tree based on the optimized distance matrix; The calculation module obtains the supermetric deviation and function of the initial distance matrix, specifically including: Extracting three nucleic acid sequences from the standard nucleic acid sequence data; Based on the initial distance matrix, obtaining three sets of initial distances between the three nucleic acid sequences; Obtaining the supermetric deviations of the three nucleic acid sequences based on the relationship between the three sets of initial distances and the triangle inequality; Repeat the above steps to obtain the supermetric deviations of any three nucleic acid sequences in the standard nucleic acid sequence data and sum them to obtain the supermetric deviation sum function. The supermetric deviation sum function is: , in D Corresponding to the initial distance matrix, Corresponding to any three nucleic acid sequences in the standard nucleic acid sequence data, is the supermetric deviation, is the total number of nucleic acid sequences in the standard nucleic acid sequence data.

Citation Information

Patent Citations

  • Syncmer-based evolutionary distance estimation and phylogenetic tree construction method

    CN118824374A