Method for biological species homology analysis based on protein sequence data
By generating probability vectors of amino acid frequency, physicochemical properties, and position, and combining them with dimensionality reduction techniques, the problem of high cost and low efficiency in existing biological species homology analysis has been solved, enabling rapid and accurate homology analysis and supporting epidemic prevention and control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 36TH RES INST OF CETC
- Filing Date
- 2021-11-19
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods for analyzing the homology of biological species are costly and inefficient, making it impossible to quickly and effectively identify species and analyze the homology of diseases caused by viruses, thus increasing the difficulty of epidemic prevention and control.
By obtaining the protein sequences of biological species, generating amino acid frequency information vectors, average physicochemical property vectors, and position probability vectors, and combining dimensionality reduction techniques, the distance between protein sequences can be calculated, enabling rapid and accurate homology analysis.
It enables homology analysis of biological species based on protein sequence data, reducing time, manpower, and financial costs, improving analysis efficiency, and helping medical professionals to quickly adopt targeted solutions.
Smart Images

Figure CN116153409B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biological species homology analysis technology, and in particular to a method for biological species homology analysis based on protein sequence data. Background Technology
[0002] In recent years, various diseases caused by viruses have gradually increased. Under such circumstances, it is crucial to conduct rapid and effective homology analysis and species identification of biological species that cause outbreaks. Timely analysis of homology and species classification among biological species is of great significance for medical workers to develop corresponding drugs and antibodies and for relevant government departments to respond to epidemic prevention and control.
[0003] Currently, the identification and homology analysis of biological species often involve conducting biomedical experiments and extracting relevant comparative features, which requires significant time, manpower, material resources, and financial costs. During the pandemic, prolonged duration often means a wider spread of the disease, greater difficulty in prevention and control, and more severe economic and social losses. Existing testing methods are not only costly but also inefficient.
[0004] Therefore, there is a lack of existing technologies for analyzing the homology of biological species based on protein sequence data. Summary of the Invention
[0005] In view of the above analysis, the embodiments of the present invention aim to provide a method for analyzing the homology of biological species based on protein sequence data, so as to solve the problems of high cost and low efficiency of existing detection methods.
[0006] On one hand, embodiments of the present invention provide a method for analyzing the homology of biological species based on protein sequence data, including:
[0007] Obtain the protein sequences of X groups of biological species, and generate an amino acid frequency information vector and an average amino acid physicochemical property vector for the X groups of protein sequences based on the frequency of each amino acid and the physicochemical properties of each amino acid in the protein sequences of each group.
[0008] Based on the positional information of amino acids 1 to K in the protein sequence, X groups of amino acid position probability vectors are generated; when K≥2, the dimensionality of each group of amino acid position probability vectors is reduced to obtain the dimensionality-reduced amino acid position probability vector; the k-character amino acid is k specified consecutive amino acids, where 1≤k≤K;
[0009] Based on the amino acid frequency information vector, the average value vector of amino acid physicochemical properties, and the reduced amino acid position probability vector, the distance between each pair of protein sequences is calculated, and the homology of the protein sequences is analyzed based on the magnitude of the distance.
[0010] Further, based on the positional information of amino acids 1 to K in the protein sequence, an X-group amino acid position probability vector is generated, including:
[0011] For each group of protein sequences, the following operation is performed to obtain the probability vector of amino acid positions for group X:
[0012] The protein sequence is sorted from 1, and the sorting number corresponding to the first amino acid in the k-word amino acid sequence is used as the position information value of the k-word amino acid.
[0013] Calculate the various k-word amino acids in sequence The sum of positional information values in the protein sequence Where i is the i-th type of k-character amino acid, 1≤i≤20 K ;
[0014] Through various k-amino acids Sum of location information values The ratio of the ... Obtain the probability vector D of the position of the k-word amino acid k Where 1≤k≤K;
[0015] The position probability vector D of amino acids from 1 to K 1 ~D K The amino acid position probability vector V' is formed by splicing these together. d .
[0016] Furthermore, the amino acid position probability vector V d ', meaning:
[0017] V' d =(D 1 D 2 ...D k …, D K )
[0018]
[0019]
[0020]
[0021]
[0022] Where k is the number of consecutive amino acids in the k-word amino acid sequence, 1≤k≤K; D k Let k be the probability vector of the amino acid positions. The positional information of the amino acid in the i-th k-th character represents the proportion of its content. For the i-th k-word amino acid The sum of positional information values appearing in the protein sequence, where N is the total number of amino acids in the protein sequence.
[0023] Furthermore, the amino acid position probability vector V' d Let M1 be an M1-dimensional vector, where M1 = 20 + 20 2 +…20 k +…20 K ;
[0024] When K≥2, the dimensionality reduction of the amino acid position probability vectors in each group is performed to obtain the dimensionality-reduced amino acid position probability vectors, including:
[0025] The amino acid position probability vector is zero-mean normalized to obtain the measurement matrix X′;
[0026] The covariance matrix S of the measurement matrix X′ is decomposed into M1 eigenvalues, which are then arranged in descending order. The eigenvectors corresponding to the first M eigenvalues are used to form an eigenvector matrix. Obtaining the eigenvector matrix The corresponding amino acid position probability vector V d ;
[0027] V d Let be the M-dimensional amino acid position probability vector obtained after dimensionality reduction.
[0028] Further, the generation of the amino acid frequency information vector of X groups of protein sequences includes:
[0029] For each group of protein sequences, the following operations are performed to obtain the amino acid frequency information vector of group X protein sequences:
[0030] The frequency of each amino acid in the protein sequence is counted, and the amino acid frequency information vector is obtained by calculating the ratio of the frequency of each amino acid to the total number of amino acids in the protein sequence; the amino acid frequency information vector V f , expressed as:
[0031] V f = (f1, f2, ..., f i …, f 20 )
[0032]
[0033]
[0034] Among them, f i amino acids Frequency information, n i amino acids The number of times it appears, where N is the total number of amino acids in the protein sequence. It is the i-th amino acid in a 1-word amino acid series.
[0035] Further, the generation of the average vector of amino acid physicochemical properties of X groups of protein sequences includes:
[0036] The following operations are performed on each group of protein sequences to obtain the average vector of amino acid physicochemical properties of group X protein sequences:
[0037] J physicochemical property parameter values of various 1-word amino acids were selected. Based on the maximum and minimum values of the physicochemical property parameter values of various 1-word amino acids, the physicochemical property parameter values of each amino acid were standardized to obtain the standardized physicochemical property parameters of each amino acid.
[0038] Based on the standardized physicochemical properties of various amino acids and their frequency of occurrence, the average values of each physicochemical property are calculated to obtain the average value vector of amino acid physicochemical properties; the average value vector V of amino acid physicochemical properties is... p , expressed as:
[0039]
[0040]
[0041]
[0042] in, To standardize physical property data, P ji For the i-th type of amino acid The value of the j-th physicochemical property parameter, P ab For the bth type of amino acid The value of the a-th physicochemical property parameter, f represents the average value of each physicochemical property in the protein sequence. i For the i-th type of amino acid Frequency information, 1≤j≤J.
[0043] Furthermore, based on the amino acid frequency information vector, the average value vector of amino acid physicochemical properties, and the amino acid position information vector, a numerical representation vector for different protein sequences is constructed. This numerical representation vector is expressed as:
[0044] V = (V f V d V p )
[0045] Among them, Vf V is the amino acid frequency information vector. d V is the amino acid position information vector. p This is the vector of average values of the physicochemical properties of amino acids.
[0046] Further, based on the numerical representation vectors of the different protein sequences, the distance d(S,T) between every two sets of protein sequences S and T is calculated. The distance d(S,T) between the two sets of protein sequences is expressed as:
[0047]
[0048] Among them, V S [q] and V T [q] represents the q-th element in the numerical representation vectors of protein sequence S and protein sequence T, respectively, 1≤q≤Q, Q=20+M+8, and M is the amino acid position probability vector V. d The dimension of.
[0049] Furthermore, when the distance between a group of protein sequences of unknown biological species and a group of protein sequences of known biological species is less than a distance threshold d... th If so, then the unknown biological species is homologous to the known biological species.
[0050] Furthermore, when the distance between the protein sequences of a certain group of unknown biological species and the protein sequences of all known biological species is greater than a distance threshold d... th In this case, the biological species with the closest homology to the unknown biological species is determined based on the shortest distance between the protein sequence of the unknown biological species and the protein sequences of all known biological species.
[0051] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0052] 1. This invention combines the frequency information of amino acid occurrence, the average value information of amino acid physicochemical properties, and the probability information of K-type amino acid positions to comprehensively and accurately analyze protein sequences. By comparing the distance between two protein sequences, protein homology analysis can be performed more accurately.
[0053] 2. This invention uses a species homology comparison analysis method based on protein sequence data to quickly classify the genetic information of species, which is beneficial for relevant medical workers to take targeted measures.
[0054] 3. The method and system for analyzing the homology of biological species and identifying species based on biological protein sequence data greatly reduces the time required for experiments compared with traditional methods, saving human, material and financial costs.
[0055] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0056] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.
[0057] Figure 1 This is a flowchart illustrating a biological species homology analysis method based on protein sequence data, as shown in one embodiment of this application.
[0058] Figure 2 A schematic diagram of the hardware structure of an electronic device for performing the biological species homology analysis method based on protein sequence data provided in the embodiments of this application. Detailed Implementation
[0059] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0060] like Figure 1 As shown, a specific embodiment of the present invention discloses a method for analyzing the homology of biological species based on protein sequence data, comprising:
[0061] S10. Obtain the protein sequences of X groups of biological species, and generate an amino acid frequency information vector and an average amino acid physicochemical property vector for the X groups of protein sequences based on the frequency of occurrence of various amino acids and the physicochemical properties of each amino acid in the protein sequences of each group; specifically, the protein sequences of the X groups of biological species include: protein sequences of unknown biological species and protein sequences of known biological species.
[0062] Specifically, with the emergence and rapid development of new technologies such as big data, artificial intelligence, and transfer learning, bioinformatics has also entered a period of rapid development. Analyzing the homology between species and identifying species categories through the combination of bioinformatics and existing new technologies is characterized by "data-driven, rapid, and accurate" approaches. Whether it's a virus or other organism, the main functions are accomplished by protein and gene sequences. Protein synthesis is controlled by genes, meaning that proteins are the dominant expression of genetic information. Therefore, homology analysis and species identification based on biological sequence data are very helpful in studying the evolutionary relationships of genetic information in different species and classifying species, and are also key to quickly identifying the source of an epidemic.
[0063] Specifically, a protein sequence is a sequence composed of different amino acids arranged in a specific order. There are 20 types of amino acids. The types, number, and order of amino acids in protein sequences with different functions are also different. Ultimately, to achieve activity and function, it may need to go through other processes such as rotation and folding, but all of these are achieved on the basis of producing a protein sequence.
[0064] When performing homology analysis of biological species based on protein sequence data, the raw protein sequence data of the biological species is first extracted using various protein sequence extraction methods. The raw protein sequence data is then preprocessed to remove outliers. Optionally, the protein sequence extraction includes, but is not limited to, methods such as biological gene transcription, chemical methods, or electromagnetic methods; the preprocessing of the raw data includes, but is not limited to, specified data extraction, data cleaning, and data feature transformation.
[0065] Specifically, the content and types of various amino acids differ in different types of protein sequences; there are a total of 20 types of amino acids. Each represents one of 20 amino acids, where 1 ≤ i ≤ 20.
[0066] Specifically, the generation of the amino acid frequency information vector of the X groups of protein sequences includes:
[0067] For each group of protein sequences, the following operations are performed to obtain the amino acid frequency information vector of group X protein sequences: The frequency of each 1-word amino acid in the protein sequence is counted, and the amino acid frequency information vector is obtained by the ratio of the frequency of each amino acid to the total number of amino acids in the protein sequence; the amino acid frequency information vector V... f , expressed as:
[0068] V f =(f1, f2, ..., f i …, f 20 )
[0069]
[0070]
[0071] Among them, f i amino acids Frequency information, n i amino acids The number of times it appears, where N is the total number of amino acids in the protein sequence. It is the i-th amino acid in a 1-word amino acid series.
[0072] Specifically, different amino acids have a variety of different physicochemical properties. Selecting the common physicochemical properties among them, an average vector is generated. Physicochemical properties refer to physical and chemical properties. The physical and chemical properties of different amino acids are certain and are known information. There are many physicochemical properties. Optionally, this embodiment uses eight physicochemical properties of amino acids, including hydrophobicity, molecular weight, solubility, specific rotation ([a]D(H2O)), specific optical rotation ([a]D(HCl)), isoelectric point, ionization state of the carboxyl group of the amino acid in aqueous solution (pk1(-COOH)) and ionization state of the amino group of the amino acid in aqueous solution (pk2(-NH3)).
[0073] Specifically, the generation of the average vector of amino acid physicochemical properties of X groups of protein sequences includes:
[0074] The following operations are performed on each group of protein sequences to obtain the average vector of amino acid physicochemical properties of group X protein sequences:
[0075] Select J kinds of physicochemical property parameter values for various 1-word amino acids, and standardize the physicochemical property parameter values of each amino acid according to the maximum and minimum values of the physicochemical property parameter values of each amino acid to obtain the standardized physicochemical property parameters of each amino acid; optionally, J = 8.
[0076] Based on the standardized physicochemical properties of various amino acids and their frequency of occurrence, the average values of each physicochemical property are calculated to obtain the average value vector of amino acid physicochemical properties; the average value vector V of amino acid physicochemical properties is... p , expressed as:
[0077]
[0078]
[0079]
[0080] in, To standardize physical property data, P ji For the i-th type of amino acid The value of the j-th physicochemical property parameter, P ab For the bth type of amino acid The value of the a-th physicochemical property parameter, f represents the average value of each physicochemical property in the protein sequence. i For the i-th type of amino acid Frequency information, 1≤j≤J.
[0081] S20. Based on the position information of amino acids 1 to K in the protein sequence, generate X groups of amino acid position probability vectors; when K≥2, reduce the dimensionality of each group of amino acid position probability vectors to obtain the dimensionality-reduced amino acid position probability vectors; the k-character amino acid is k specified consecutive amino acids, where 1≤k≤K;
[0082] Specifically, the value of K can be freely chosen based on the protein sequence length and computing power. When K=1, it indicates that only the case of 20 amino acids appearing alone is analyzed; when K=2, it indicates that the case of two amino acid combinations appearing simultaneously is analyzed, for example, 400 amino acid combinations such as II, IV, VI, and IL; and so on, depending on the value of K, it can analyze the case of different amino acid combinations appearing simultaneously. K We will analyze this situation. The larger the K value, the greater the computing power required. Therefore, the K value can be selected based on the actual application platform.
[0083] Specifically, based on the positional information of amino acids 1 to K in the protein sequence, an X-group amino acid position probability vector is generated, including:
[0084] For each group of protein sequences, the following operation is performed to obtain the probability vector of amino acid positions for group X:
[0085] The protein sequence is sorted from 1, and the sorting number corresponding to the first amino acid in the k-word amino acid sequence is used as the position information value of the k-word amino acid.
[0086] Calculate the various k-word amino acids in sequence The sum of positional information values in the protein sequence Where i is the i-th type of k-character amino acid, 1≤i≤20 K ;
[0087] Through various k-amino acids Sum of location information values The ratio of the ... Obtain the probability vector D of the position of the k-word amino acid k Where 1≤k≤K;
[0088] The position probability vector D of amino acids from 1 to K 1 ~D K The amino acid position probability vector V' is formed by splicing these together. d .
[0089] More specifically, the amino acid position probability vector V' d , expressed as:
[0090] V d '=(D 1 D 2 ...D k …, D K )
[0091]
[0092]
[0093]
[0094]
[0095] Where k is the number of consecutive amino acids in the k-word amino acid sequence, 1≤k≤K; D k Let k be the probability vector of the amino acid positions. The positional information of the amino acid in the i-th k-th character represents the proportion of its content. For the i-th k-word amino acid The sum of positional information values appearing in the protein sequence, where N is the total number of amino acids in the protein sequence.
[0096] Specifically, when K≥2, the amino acid position probability vector V' d The dimension is M1 = 20 + 20 2 +…20 k +…20 K The selection of different K values is to find patterns in protein sequence similarity analysis after collecting a large amount of data on the arrangement and combination of amino acids in the protein sequence. However, a large amount of data will increase the workload of data analysis to a certain extent. More importantly, there may be correlations between many data points, which increases the complexity of problem analysis. Therefore, in the analysis process, high-dimensional data can be preprocessed by dimensionality reduction to retain the most important features and remove noise and unimportant features, thereby improving the purpose of data processing. This can save a lot of time and cost in our engineering practice within a certain range of information loss.
[0097] Before dimensionality reduction, the amino acid position probability vector V' dLet M1 be an M1-dimensional vector, where M1 = 20 + 20 2 +…20 k +…20 K Specifically, when K≥2, the dimensionality of the amino acid position probability vectors in each group is reduced to obtain the dimensionality-reduced amino acid position probability vectors, including:
[0098] The amino acid position probability vector is zero-mean normalized to obtain the measurement matrix X′;
[0099] The covariance matrix S of the measurement matrix X′ is decomposed into M1 eigenvalues, which are then arranged in descending order. The eigenvectors corresponding to the first M eigenvalues are used to form an eigenvector matrix. Obtaining the eigenvector matrix The corresponding amino acid position probability vector V d Optionally, M = 72;
[0100] V d This is the M-dimensional amino acid position probability vector obtained after dimensionality reduction, which is the final 72-dimensional amino acid position probability vector obtained after dimensionality reduction.
[0101] Specifically, let's take the example of obtaining a 72-dimensional amino acid position probability vector with K=2 as an example:
[0102] When K=2, then M1=420, which is the amino acid position probability vector V'. d 420-dimensional vector
[0103] The amino acid position probability vectors are used to construct a 1*420 amino acid position probability matrix X, where... 1≤m≤420;
[0104] The amino acid position probability matrix was zero-mean normalized to obtain a 1*420 measurement matrix.
[0105] The covariance matrix S of the 1*420 measurement matrix X′ is decomposed into M1 eigenvalues, which are then arranged in descending order. The eigenvectors corresponding to the first M eigenvalues are used to form an eigenvector matrix. Obtaining the eigenvector matrix The corresponding amino acid position probability vector V d Optionally, M = 72; where the covariance matrix S is expressed as:
[0106]
[0107] The 420 values on the diagonal of the covariance matrix S are eigenvalues of the covariance matrix S, i.e., z m,m(1≤m≤420) represents the eigenvalues of the covariance matrix S. The 420 eigenvalues are arranged in descending order, and the proportions of the 72 amino acid position information corresponding to the first 72 eigenvalues are taken to form the M-dimensional amino acid position probability vector V obtained after dimensionality reduction. d This refers to the 72-dimensional amino acid position probability vector obtained after dimensionality reduction.
[0108] S30. Based on the amino acid frequency information vector, the average vector of amino acid physicochemical properties, and the reduced amino acid position probability vector, calculate the distance between each pair of protein sequences, and analyze the homology of the protein sequences based on the magnitude of the distance. Optionally, the distance between the two pairs of protein sequences can be calculated using methods such as Euclidean distance, Manhattan distance, Chebyshev distance, or genetic distance.
[0109] Specifically, based on the amino acid frequency information vector, the average value vector of amino acid physicochemical properties, and the amino acid position information vector, a numerical representation vector for different protein sequences is constructed, wherein the numerical representation vector is expressed as:
[0110] V = (V f V d V p )
[0111] Among them, V f V is the amino acid frequency information vector. d V is the amino acid position information vector. p This is the vector of average values of the physicochemical properties of amino acids.
[0112] Specifically, based on the numerical representation vectors of the different protein sequences, the distance d(S,T) between each pair of protein sequences S and T is calculated. The distance d(S,T) between the two pairs of protein sequences is expressed as:
[0113]
[0114] Among them, V S [q] and V T [q] represents the q-th element in the numerical representation vectors of protein sequence S and protein sequence T, respectively, 1≤q≤Q, Q=20+M+8, and M is the amino acid position probability vector V. d The dimension of.
[0115] Specifically, when the distance between a certain group of protein sequences and at least two other groups of protein sequences from other species is not clearly distinguishable, the next step is to process them by dividing the position probability vectors D of amino acids 1 to K. 1 ~D K The amino acid position probability vector V' is formed by splicing these together.d A second dimensionality reduction is performed, and the position probability vector of the first 20 dimensions, the amino acid frequency information vector, and the average value vector of amino acid physicochemical properties are selected to form the second protein sequence digital representation vector. Then, the distance between protein sequences is calculated.
[0116] Specifically, the protein sequences of biological species in group X include protein sequences of unknown biological species in group x1 and protein sequences of known biological species in group x2, where X = x1 + x2; the distance between each protein sequence in group x1 of unknown biological species and all protein sequences in group x2 of known biological species is calculated, and the homology of the protein sequences of unknown biological species is analyzed by the distance between each protein sequence of unknown biological species and the protein sequences of all known biological species;
[0117] More specifically, when the distance between a group of protein sequences of an unknown biological species and a group of protein sequences of a known biological species is less than a distance threshold d. th If so, then the unknown biological species is homologous to the known biological species.
[0118] Specifically, when the distance between the protein sequences of a certain group of unknown biological species and the protein sequences of all known biological species is greater than a distance threshold d. th In this case, the species with the closest homology to the unknown species is determined based on the shortest distance between the protein sequences of the unknown species and the protein sequences of all known species. That is, the known species corresponding to the shortest distance is the species with the closest homology to the unknown species.
[0119] When multiple groups of protein sequences from unknown biological species are homologous to, or have the closest homology to, protein sequences from known biological species in the same group, they can be grouped together for easier subsequent analysis.
[0120] Compared with existing technologies, the biological species homology analysis method based on protein sequence data proposed in this invention firstly, by combining the frequency information of amino acid occurrences, the average value information of amino acid physicochemical properties, and the probability information of K-shaped amino acid positions, it can comprehensively and accurately analyze protein sequences. By comparing the distance between two protein sequences, it can more accurately analyze protein homology. Secondly, through the species homology comparison analysis method based on protein sequence data, this invention can quickly classify the genetic information of species, which is beneficial for relevant medical workers to take targeted measures. Finally, the biological species homology analysis and species identification method and system based on biological protein sequence data greatly reduces the time required for experiments compared with traditional methods, saving human, material, and financial costs.
[0121] See Figure 2 Another embodiment of the present invention also provides an electronic device for performing the biological species homology analysis method based on protein sequence data described in the above embodiments. The electronic device includes:
[0122] One or more processors 710 and memory 720, Figure 2 Take the 710 processor as an example.
[0123] The electronic device for performing bio-species homology analysis methods based on protein sequence data may further include: an input device 730 and an output device 740.
[0124] The processor 710, memory 720, input device 730, and output device 740 can be connected via a bus or other means. Figure 2 Taking the example of a connection between China and Israel via a bus.
[0125] The memory 720, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules (units) corresponding to the biological species homology analysis method based on protein sequence data in the embodiments of the present invention. The processor 710 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 720, thereby implementing the icon display method of the above-described method embodiments.
[0126] The memory 720 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store information such as the number of reminders from the acquired applications. Furthermore, the memory 720 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 720 may optionally include memory remotely located relative to the processor 710, and these remote memories can be connected to the processing device operating the list items via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0127] The input device 730 can receive input digital or character information, as well as key signal inputs related to user settings and function control of the bio-species homology analysis device based on protein sequence data. The output device 740 may include a display device such as a screen.
[0128] The one or more modules are stored in the memory 720, and when executed by the one or more processors 710, they perform the biological species homology analysis method based on protein sequence data in any of the above method embodiments.
[0129] The above-described product can perform the methods provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of the present invention.
[0130] The electronic devices of the embodiments of the present invention may exist in various forms, including but not limited to:
[0131] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include: smartphones (e.g., iPhones), multimedia phones, feature phones, and low-end phones, etc.
[0132] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include PDAs, MIDs, and UMPCs, such as the iPad.
[0133] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes: audio and video players (e.g., iPods), handheld game consoles, e-book readers, as well as smart toys and portable car navigation devices.
[0134] (4) Server: A device that provides computing services. The components of a server include a processor, hard disk, memory, system bus, etc. Servers are similar to general computer architectures, but because they need to provide highly reliable services, they have higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.
[0135] (5) Other electronic devices with reminder recording function.
[0136] The device embodiments described above are merely illustrative. The units (modules) described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0137] This invention provides a non-transitory computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are executed by an electronic device, the electronic device performs the biological species homology analysis method based on protein sequence data in any of the above method embodiments.
[0138] This invention provides a computer program product, wherein the computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, wherein when the program instructions are executed by an electronic device, the electronic device performs the biological species homology analysis method based on protein sequence data in any of the above method embodiments.
[0139] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the prior art, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0140] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the principles or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for analyzing the homology of biological species based on protein sequence data, characterized in that, include: The protein sequences of X groups of biological species are obtained. Based on the frequency of occurrence of various amino acids and the physicochemical properties of each amino acid in the protein sequences of each group, an amino acid frequency information vector and an average amino acid physicochemical property vector are generated for the X groups of protein sequences. The protein sequences of the X groups of biological species include protein sequences of unknown biological species and protein sequences of known biological species. The value of K is selected based on the protein sequence length and computing power; based on the positional information of amino acids 1 to K in the protein sequence, an X-group amino acid position probability vector is generated. include: For each group of protein sequences, the following operation is performed to obtain the probability vector of amino acid positions for group X: The protein sequence is sorted from 1, and the position information value of the first amino acid in the k-word amino acid sequence is used as the position information value of the k-word amino acid; the k-word amino acid sequence consists of k specified consecutive amino acids, where 1≤k≤K; Calculate the various k-word amino acids in sequence The sum of positional information values in the protein sequence Where i represents the i-th type of k-word amino acid, ; Through various k-amino acids Sum of location information values The ratio of the ... The probability vector of the position of amino acid k is obtained. ; The position probability vector of amino acids from 1 to K Concatenate to form the amino acid position probability vector of this group of amino acids ,include: The amino acid position probability vector , expressed as: Where k is the number of consecutive amino acids in the k-word amino acid, 1≤k≤K; Let k be the probability vector of the amino acid positions. The positional information of the amino acid in the i-th k-th character represents the proportion of its content. For the i-th k-th amino acid The sum of positional information values appearing in the protein sequence, where N is the total number of amino acids in the protein sequence; When K≥2, the dimensionality of the amino acid position probability vectors in each group is reduced to obtain the dimensionality-reduced amino acid position probability vectors, including: The amino acid position probability vector for A dimensional vector, where, ; The amino acid position probability vector is zero-mean normalized to obtain the measurement matrix. ; For the measurement matrix covariance matrix Eigenvalue decomposition is performed to obtain the covariance matrix. of The eigenvalues are denoted as M, arranged in descending order. The eigenvectors corresponding to the first M eigenvalues are used to form the eigenvector matrix. ; Obtain the eigenvector matrix The corresponding amino acid position probability vector ; The result after dimensionality reduction A dimensional probability vector of amino acid positions; Based on the amino acid frequency information vector, the average value vector of amino acid physicochemical properties, and the reduced amino acid position probability vector, the distance between each pair of protein sequences is calculated, and the homology of the protein sequences is analyzed based on the magnitude of the distance. Calculate the distance between any two sets of protein sequences, including: When the distance between a certain group of protein sequences and at least two other groups of protein sequences is not clearly distinguishable, the amino acid position probability vectors corresponding to each group of protein sequences are analyzed separately. A second dimensionality reduction is performed, and the distance between the protein sequence and the other two groups of protein sequences is calculated again.
2. The method for analyzing the homology of biological species based on protein sequence data according to claim 1, characterized in that, The generated amino acid frequency information vector of the X group of protein sequences includes: For each group of protein sequences, the following operations are performed to obtain the amino acid frequency information vector of group X protein sequences: The frequency of each amino acid in the protein sequence is counted, and the amino acid frequency information vector is obtained by calculating the ratio of the frequency of each amino acid to the total number of amino acids in the protein sequence; the amino acid frequency information vector , expressed as: in, amino acids Frequency information amino acids Number of times it appears This refers to the total number of amino acids in a protein sequence. The first amino acid in the 1st word A type of amino acid.
3. The method for analyzing the homology of biological species based on protein sequence data according to claim 2, characterized in that, The vector of average amino acid physicochemical properties of the X group of protein sequences includes: The following operations are performed on each group of protein sequences to obtain the average vector of amino acid physicochemical properties of group X protein sequences: J physicochemical property parameter values of various 1-word amino acids were selected. Based on the maximum and minimum values of the physicochemical property parameter values of various 1-word amino acids, the physicochemical property parameter values of each amino acid were standardized to obtain the standardized physicochemical property parameters of each amino acid. Based on the standardized physicochemical properties of various amino acids and their frequency of occurrence, the average values of each physicochemical property are calculated to obtain an average vector of amino acid physicochemical properties; the average vector of amino acid physicochemical properties... , expressed as: in, To standardize physical property data, For the i-th type of amino acid The value of the j-th physicochemical property parameter, For the bth type of amino acid The value of the a-th physicochemical property parameter, This represents the average of the physicochemical properties in the protein sequence. For the i-th type of amino acid Frequency information .
4. The method for analyzing the homology of biological species based on protein sequence data according to any one of claims 1 to 3, characterized in that, Based on the amino acid frequency information vector, the average value vector of amino acid physicochemical properties, and the amino acid position information vector, a numerical representation vector for different protein sequences is constructed. This numerical representation vector is expressed as: in, This is a vector representing the frequency information of amino acids. This is a vector containing amino acid position information. This is the vector of average values of the physicochemical properties of amino acids.
5. The method for analyzing the homology of biological species based on protein sequence data according to claim 4, characterized in that, Based on the numerical representation vectors of the different protein sequences, the distance between each pair of protein sequences S and T is calculated. The distance between the two sets of protein sequences , expressed as: Among them, and The numerical representation vectors of protein sequences S and T are respectively the first and second most significant vectors. Each corresponding element , , The amino acid position probability vector The dimension of.
6. The method for analyzing the homology of biological species based on protein sequence data according to claim 5, characterized in that, When the distance between a group of protein sequences of an unknown biological species and a group of protein sequences of a known biological species is less than a distance threshold... If so, then the unknown biological species is homologous to the known biological species.
7. The method for analyzing the homology of biological species based on protein sequence data according to claim 5, characterized in that, When the distance between the protein sequences of a group of unknown biological species and the protein sequences of all known biological species is greater than a distance threshold... In this case, the biological species with the closest homology to the unknown biological species is determined based on the shortest distance between the protein sequence of the unknown biological species and the protein sequences of all known biological species.