A protein structure prediction method, device, platform and storage medium

By constructing an initial three-dimensional structure model in protein structure prediction and filling in the missing parts, combined with molecular mechanics minimization optimization, the problem of inability to fully match the target sequence in homology modeling is solved, achieving more accurate and stable protein structure prediction, and supporting secondary structure prediction and visualization.

CN112530517BActive Publication Date: 2025-09-26KANGMAXIN (SHANGHAI) INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910880279.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-09-18
Publication Date
2025-09-26
Estimated Expiration
2039-09-18

AI Technical Summary

Technical Problem

The existing homology modeling method is used to predict protein structure, which affects the completeness and accuracy of the prediction because the target sequence cannot completely find the template sequence.

Method used

By extracting the target sequence from the protein file to be tested, matching it with the protein database of known structures, constructing an initial three-dimensional structure model, and combining the unmatched sequence fragments with the adjacent matched sequence fragments into a sub-target sequence, the matching subsequences and their structures are searched in the database, the missing parts are filled, and the structure is optimized by combining molecular mechanics minimization.

Benefits of technology

It improves the accuracy and completeness of protein structure prediction, optimizes the stability of structural models, provides secondary structure prediction capabilities, and supports visual display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112530517B_ABST
    Figure CN112530517B_ABST
Patent Text Reader

Abstract

The present invention discloses a protein structure prediction method, comprising: extracting a target sequence from a protein file to be tested; matching the target sequence against a protein database of known structures to find a matching sequence; obtaining a matching structure of the matching sequence based on the matching sequence; constructing an initial three-dimensional structural model of the target sequence based on the matching sequence and its matching structure; combining unmatched sequence segments of the target sequence with a portion of adjacent matched sequence segments to form a sub-target sequence; and searching for matching subsequences and their structures of the sub-target sequence in a protein database of known structures; and filling in missing portions of the initial three-dimensional structural model based on the found matching subsequences and their structures to obtain the three-dimensional structure of the protein file to be tested. The protein structure prediction method of the present invention can more accurately predict the structure of a protein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of protein structure prediction, and in particular to a protein structure prediction method, device, platform and storage medium. Background Art

[0002] Studying the three-dimensional structure of proteins can explain protein function at the molecular level. For example, protein complexes are often central to many cellular metabolic processes. Describing the interactions and quaternary structure of the components of a protein complex at the molecular level helps researchers understand how protein complexes function within their metabolic environment and provides insights into how these complexes regulate this process. In recent years, the number of protein complexes in the Protein Crystallographic Database (PDB) has grown rapidly due to the continuous development of structure determination methods based on electron microscopy. However, the speed at which experimental methods can determine the three-dimensional structure of protein complexes still cannot keep up with the demands of high-throughput screening of protein-protein complexes. Therefore, predicting protein structure through computational prediction methods is an effective approach to address this problem.

[0003] Currently, protein prediction methods include ab initoprediction, threading modeling, and homology modeling, depending on the specific situation. Compared to the first two prediction methods, homology modeling is the most mature method currently developed and applied. During evolution, homologous proteins share similar sequences and structures. Based on this principle, homology modeling searches for amino acid sequences of known structures (template sequences: models) in a database and compares the similarity between the target protein and the template sequence to select the optimal modeling template. The residues of the target protein sequence are then mapped to the template protein structure to complete the three-dimensional structure construction. During the homology modeling process, it is possible that the target sequence cannot completely find the template sequence, that is, the similarity between the two is low. This problem affects the completeness and accuracy of the target protein structure prediction. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides a protein structure prediction method, device, platform and storage medium; specifically, the technical solution of the present invention is as follows:

[0005] In a first aspect, the present invention provides a protein structure prediction method, comprising: extracting a target sequence from a protein file to be tested; matching the target sequence in a protein database of known structures to find a matching sequence; obtaining a matching structure of the matching sequence based on the matching sequence; constructing an initial three-dimensional structure model of the target sequence based on the matching sequence and its matching structure; combining unmatched sequence fragments of the target sequence with a portion of adjacent matched sequence fragments to form a sub-target sequence; and searching for a matching subsequence and its structure of the sub-target sequence in the protein database of known structures; filling in the missing parts in the initial three-dimensional structure model based on the found matching subsequence and its structure to obtain the three-dimensional structure of the protein file to be tested.

[0006] Preferably, the combining of the unmatched sequence fragments of the target sequence with a part of the adjacent matched sequence fragments into a sub-target sequence; and searching for the matching subsequence and structure of the sub-target sequence in the protein database of known structure specifically includes: obtaining the sequence fragments that match the target sequence and the matching sequence as the matched sequence fragments; taking the sequence fragments on the target sequence that do not match the matching sequence as the unmatched sequence fragments; when the fragment length of the unmatched sequence fragment is greater than a preset length value, determining the cut site in the matched sequence fragment, and combining the unmatched sequence fragment with the adjacent cut matched sequence fragment to form a sub-target sequence; when the fragment length of the unmatched sequence fragment is less than or equal to the preset length value, truncating the unmatched sequence fragment as the sub-target sequence; matching the sub-target sequence in the protein database of known structure to find the matching subsequence and structure.

[0007] Preferably, the step of determining the cleavage site in the matched sequence fragment, when the length of the unmatched sequence fragment is greater than a preset length value, determining the cleavage site in the matched sequence fragment, and combining the unmatched sequence fragment with the adjacent matched sequence fragment after cleavage to form a sub-target sequence specifically includes: if the starting site of the unmatched sequence fragment is set to M e1 , the termination site is M s2 ; Set the adjacent matched sequence fragment of the unmatched sequence fragment as the first matched sequence fragment and / or the second matched sequence fragment; the first matched sequence fragment has M1 amino acids, and its starting site is M s1 , the termination site is M e1 The second matching sequence fragment has M2 amino acids, and its starting site is M s2 , the termination site is M e2 ;

[0008] When M s2 -Me1 >15: If the matching sequence segment adjacent to the unmatched sequence segment is only the first matching sequence segment, the cleavage site T1 of the first matching sequence is intercepted to the termination site M of the unmatched sequence segment. S2 The sequence fragment between is the sub-target sequence; if the matching sequence fragment adjacent to the unmatched sequence fragment has only the second matching sequence fragment, the starting site M of the unmatched sequence fragment is intercepted. e1 The sequence fragment between the cleavage site T1 of the first matching sequence and the cleavage site T2 of the second matching sequence is taken as the sub-target sequence; if the matching sequence fragments adjacent to the unmatched sequence fragment are the first matching sequence fragment and the second sequence fragment, the sequence fragment between the cleavage site T1 of the first matching sequence and the cleavage site T2 of the second matching sequence is taken as the sub-target sequence; wherein:

[0009] If M1 / 2>10, select the cleavage site T1 of the first matching sequence as M e1 -9;

[0010] If M1 / 2≦10, select the cleavage site T1 of the first matching sequence as (M e1 -M1 / 2+1); wherein, M1 / 2 is rounded up or down;

[0011] If M2 / 2>10, select the cleavage site T2 of the second matching sequence as M s2 +9;

[0012] If M2 / 2≦10, select the cleavage site T2 of the second matching sequence as (M s2 +M2 / 2-1); wherein, the M2 / 2 is rounded up or down.

[0013] When the length of the unmatched sequence segment is less than or equal to the preset length value, intercepting the unmatched sequence segment as a sub-target sequence specifically includes: when M s2 -M e1 ≦15, intercept the unmatched sequence fragment M e1 To M s2 As a sub-target sequence fragment.

[0014] Preferably, constructing the initial three-dimensional structural model of the target sequence based on the matching sequence and its matching structure specifically includes: obtaining the target sequence and the matched sequence fragments in the matching sequence; constructing the initial three-dimensional structural model of the target sequence using the matching structure of the matching sequence as a template; in the initial three-dimensional structural model, the structure of the matched sequence fragment of the target sequence adopts the structure of the corresponding matched sequence fragment in the matching sequence; and treating the structure of the unmatched sequence fragment in the target sequence as a missing structure.

[0015] Preferably, when no matching sequence or matching structure is found in the protein database of known structures, matching sequence fragments of the target sequence are searched; from the matching sequence fragments found, the best matching sequence fragment is selected as the target sequence fragment; based on the positional relationship of the target sequence fragment on the target sequence, the cleavage site is determined on the target sequence to obtain a sub-target sequence; based on the sub-target sequence, a template is searched in the protein database of known structures to construct the initial three-dimensional structure of the target sequence.

[0016] Preferably, determining the cleavage site on the target sequence according to the positional relationship of the target sequence fragment on the target sequence, and obtaining the sub-target sequence specifically includes: assuming that the target sequence has N amino acids, its starting site is Ns and its ending site is Ne; the target sequence fragment has Q amino acids, its starting site is Qs and its ending site is Qe;

[0017] When Qs-Ns>15: If Q / 2>10, the cleavage site of the target sequence is selected as Qs+9, and the fragment from Ns to (Qs+9) is intercepted as the sub-target sequence; if Q / 2≦10, the cleavage site of the target sequence is selected as (Qs+Q / 2-1), and the fragment from Ns to (Qs+Q / 2-1) is intercepted as the sub-target sequence; wherein Q / 2 is rounded up or down;

[0018] When Qs-Ns≦15, the cleavage site of the target sequence is selected as Qs, and the fragment from Ns to Qs is intercepted as the sub-target sequence;

[0019] When Ne-Qe>15: If Q / 2>10, the cleavage site of the target sequence is selected as Qe-9, and the fragment from (Qe-9) to Ne is intercepted as the sub-target sequence; if Q / 2≦10, the cleavage site of the target sequence is selected as (Qe-Q / 2+1), and the fragment from (Qe-Q / 2+1) to Ne is intercepted as the sub-target sequence; wherein Q / 2 is rounded up or down;

[0020] When Ne-Qe≦15, the unmatched sequence fragments Qe to Ne are intercepted as the sub-target sequence.

[0021] Preferably, the protein structure prediction method of the present invention further comprises: reconstructing the side chains of the three-dimensional structure of the protein to be tested; and minimizing the energy of the three-dimensional structure model of the protein to be tested using molecular mechanics.

[0022] Preferably, the protein structure prediction method of the present invention further comprises: predicting the secondary structure of the protein to be tested based on the three-dimensional structure of the protein to be tested and combining the information of the three-dimensional spatial structure and secondary structure of each protein sequence in the protein database of known structures.

[0023] In a second aspect, the present invention discloses a protein structure prediction device, comprising: a sequence extraction module for extracting a target sequence from a protein file to be tested; a matching search module for matching the target sequence in a protein database of known structures to find a matching sequence; and obtaining a matching structure of the matching sequence based on the matching sequence; a model construction module for constructing an initial three-dimensional structure model of the target sequence based on the matching sequence and its matching structure; a combination module for combining unmatched sequence fragments of the target sequence with a part of adjacent matched sequence fragments to form a sub-target sequence; and searching for a matching subsequence and its structure of the sub-target sequence in the protein database of known structures through the matching search module; a filling module for filling the missing parts in the initial three-dimensional structure model based on the found matching subsequence and its structure to obtain the three-dimensional structure of the protein file to be tested.

[0024] Preferably, the combination module includes: an acquisition submodule, used to acquire sequence fragments that match the target sequence and the matching sequence as matched sequence fragments; and to treat sequence fragments on the target sequence that do not match the matching sequence as unmatched sequence fragments; a determination submodule, used to determine the cut site in the matched sequence fragment when the fragment length of the unmatched sequence fragment is greater than a preset length value, and to combine the unmatched sequence fragment with the adjacent cut matched sequence fragment to form a sub-target sequence; and also used to intercept the unmatched sequence fragment as a sub-target sequence when the fragment length of the unmatched sequence fragment is less than or equal to the preset length value; the match search module matches the sub-target sequence in the protein database of known structure to find a matching subsequence and its structure.

[0025] Preferably, the determination submodule includes a selection unit and a cut-off unit; wherein: if the starting position of the unmatched sequence fragment is set to M e1 , the termination site is M s2 ; Set the adjacent matched sequence fragment of the unmatched sequence fragment as the first matched sequence fragment and / or the second matched sequence fragment; the first matched sequence fragment has M1 amino acids, and its starting site is M s1 , the termination site is M e1 The second matching sequence fragment has M2 amino acids, and its starting site is M s2 , the termination site is M e2 ;

[0026] When M s2 -M e1 >15: If the matching sequence segment adjacent to the unmatched sequence segment is only the first matching sequence segment, the interception unit intercepts the cutting site T1 of the first matching sequence to the termination site M of the unmatched sequence segment. S2 The sequence segment between is the sub-target sequence; if the matching sequence segment adjacent to the unmatched sequence segment has only the second matching sequence segment, the interception unit intercepts the starting site M of the unmatched sequence segment. e1 to the cleavage site T2 of the second matching sequence as a sub-target sequence; if the matching sequence segments adjacent to the unmatched sequence segment are the first matching sequence segment and the second sequence segment, the clipping unit clips the sequence segment between the cleavage site T1 of the first matching sequence and the cleavage site T2 of the second matching sequence as a sub-target sequence; wherein:

[0027] If M1 / 2>10, the selection unit selects the cleavage site T1 of the first matching sequence as M e1 -9;

[0028] If M1 / 2≦10, the selection unit selects the cleavage site T1 of the first matching sequence as (M e1 -M1 / 2+1); wherein, M1 / 2 is rounded up or down;

[0029] If M2 / 2>10, the selection unit selects the cutting site T2 of the second matching sequence as M s2 +9;

[0030] If M2 / 2≦10, the selection unit selects the cleavage site T2 of the second matching sequence as (M s2 +M2 / 2-1); wherein, M2 / 2 is rounded up or down;

[0031] When M s2 -M e1 ≦15: the interception unit intercepts the unmatched sequence segment M s2 To M e1 As a sub-target sequence fragment.

[0032] Preferably, the protein structure prediction device further comprises: a side chain construction module for reconstructing the side chains of the three-dimensional structure of the protein to be tested; and a structure optimization module for minimizing the energy of the three-dimensional structure model of the protein to be tested using molecular mechanics.

[0033] Preferably, the protein structure prediction device further comprises: a secondary structure prediction module for predicting the secondary structure of the protein to be tested based on the three-dimensional structure of the protein to be tested and in combination with the information on the three-dimensional spatial structure and secondary structure of each protein sequence in the protein database of known structures.

[0034] In a third aspect, the present invention discloses a storage medium storing a plurality of instructions, wherein the plurality of instructions are executed by one or more processors to implement the steps of any one of the protein structure prediction methods of the present invention.

[0035] In a fourth aspect, the present invention discloses a protein structure prediction platform, comprising a protein structure prediction device according to any one of the present inventions. The protein structure prediction platform is constructed on a server and is equipped with an online visualization program for protein three-dimensional structure and secondary structure to visualize the structure of the protein.

[0036] The present invention includes at least one of the following technical effects:

[0037] (1) In the protein structure prediction technology of the present invention, the target sequence of the protein to be tested is first extracted, and then the matching sequence of the target sequence is searched for, and then the structural template of the matching sequence is selected to establish the initial three-dimensional structure. In the modeling process, the missing parts of the model are processed accordingly. Specifically, the unmatched sequence fragments of the target sequence are combined with a part of the adjacent matched sequence fragments to form a sub-target sequence; then, the matching subsequences and their structures of the sub-target sequence are searched in the protein database of known structures to fill the missing parts of the initial three-dimensional structure of the protein to be tested, thereby making the prediction results more accurate.

[0038] (2) The protein structure prediction technology of the present invention provides a method for determining the cleavage site of the matching sequence. Different determination methods are adopted according to the different lengths of the unmatched sequences and the different lengths of the matched sequences, so that the new sequence fragments (sub-target sequences) obtained by interception are more reasonable and more likely to be matched with similar sequence structures, thereby filling in the corresponding missing parts and improving the prediction accuracy.

[0039] (3) In the protein structure prediction technology of the present invention, based on the three-dimensional structure of the protein to be tested after the gaps are filled, the structural model is also subjected to side chain reconstruction and the energy of the model is minimized using molecular mechanics, thereby optimizing the predicted structural model and making the structural model more stable.

[0040] (4) The present invention can also utilize information in existing databases, including information on secondary structure and spatial structure, to predict the secondary structure of a protein from its target sequence, thereby making the visual display of the protein more structured.

[0041] (5) The protein structure prediction platform of the present invention includes the protein structure prediction device of the present invention. The platform is set on a server and is provided with a visualization program for the secondary structure and three-dimensional structure of the protein. The user can log in to the platform on any smart device to perform protein structure prediction operations. After the protein structure prediction platform obtains the protein to be tested input by the user, the protein structure prediction device in the platform can use the protein structure prediction method of the present invention to predict the three-dimensional structure of the protein to be tested input by the user, and visualize the final prediction results for the user to view conveniently and quickly. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 A flowchart of an embodiment of a protein structure prediction method of the present invention;

[0044] Figure 2 Schematic diagram of the extraction process for extracting the target sequence of the protein to be tested in the present invention;

[0045] Figure 3 A flowchart of another embodiment of a protein structure prediction method of the present invention;

[0046] Figure 4a A schematic diagram of a determination of a sub-goal sequence;

[0047] Figure 4b Schematic diagram of another determination of sub-goal sequence;

[0048] Figure 4c Schematic diagram of another determination of sub-goal sequence;

[0049] Figure 4d determining a schematic diagram for the subtarget sequence used to find the template structure;

[0050] Figure 5 A schematic flow chart of another embodiment of a protein structure prediction method of the present invention;

[0051] Figure 6 Schematic diagram of the three-dimensional structure prediction results of the protein to be tested;

[0052] Figure 7 Schematic diagram of the secondary structure prediction results of the protein to be tested;

[0053] Figure 8 This is a structural block diagram of an embodiment of a protein structure prediction device of the present invention;

[0054] Figure 9 This is a structural block diagram of another embodiment of a protein structure prediction device of the present invention. DETAILED DESCRIPTION

[0055] In the following description, specific details such as specific system structures and technologies are provided for illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present application with unnecessary details.

[0056] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections.

[0057] To simplify the drawings, only the parts relevant to the present invention are schematically depicted in each figure; they do not represent the actual structure of the product. Furthermore, to simplify the drawings and facilitate understanding, in some figures, only one component with the same structure or function is schematically depicted or labeled. As used herein, "one" refers not only to "only one" but also to "more than one."

[0058] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0059] In addition, in the description of the present application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the specific embodiments of the present invention will be described below with reference to the accompanying drawings. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings and other embodiments can be obtained based on these drawings without inventive work.

[0061] Figure 1 The present invention provides a flowchart of an embodiment of a method for predicting protein structure, comprising:

[0062] S101, extracting the target sequence from the protein file to be tested;

[0063] Specifically, there are twenty basic amino acids in proteins. The English names of amino acids in the protein file are in the form of three-letter abbreviations. The corresponding one-letter abbreviations need to be converted based on the atom ATOM beginning of the standard residue in the protein file (PDB). Then, the corresponding amino acid name is obtained according to the sequence number, and the chain identifier is obtained as the amino acid sequence contained in different chains.

[0064] It should be noted that repeated sequences in a single protein file need to be deduplicated. The atomic coordinates of standard residues mixed with non-standard residues need to be filtered. In addition, for residue sequence numbers and residue insertion codes, the residue insertion codes need to be added to the sequence numbers. Then a sequence for each chain will be obtained, and then deduplication will be performed. The extracted sequences are the same but the chains or file names are different, so they are added together. You will get the NR library (protein library). The NR database is a merger of several protein sequence libraries. Then install blast (basic comparison search tool) and put the extracted sequence library in to format the sequence library. The flowchart for extracting protein sequences (target sequences) is as follows: Figure 2 shown.

[0065] The Non-Redundant Protein Sequence Database (NR) is a non-redundant protein database. For all known or potential coding sequences in GenBank, EMBL, DDBJ, and PDB, the NR records provide the corresponding amino acid sequence (inferred from known or potential reading frames) and the sequence number in the specialized protein database. The NR database acts as a cross-index based on nucleic acid sequences, linking nucleic acid and protein data.

[0066] S102, matching the target sequence with a protein database of known structures to find a matching sequence;

[0067] Concrete, after obtaining the target sequence, then call blastp again to search the similar sequence that coupling and target sequence mate, to choose corresponding structural template.In this step, can select PDB (Protein Data Bank) as the protein database of known structure, the target sequence that extracts in the inventive method comprises aminoacid sequence or nucleotide sequence, if target sequence is aminoacid sequence, then can call blastp and directly search and compare, and return retrieval result.If target sequence is nucleotide sequence, then call blastx, earlier this nucleotide sequence is translated into protein sequence, then search and mate, and return retrieval result.

[0068] blast, short for Basic Local Alignment Search Tool, is a tool for searching for sequences based on local alignments. Blast works by first creating a database with the target sequence (this database is called a database, and each sequence in it is called a subject). Then, the database is searched with the sequence to be searched (called a query). Each query is then aligned with each subject in the database, yielding a complete set of alignments.

[0069] Blast is an integrated program package that implements five possible sequence alignment methods by calling different alignment modules:

[0070] blastp: Compare protein sequences with protein libraries and directly compare the homology of protein sequences.

[0071] blastx: compares nucleic acid sequences to protein libraries. First, the nucleic acid sequence is translated into a protein sequence (which can be translated into 6 possible protein sequences according to the phase), and then compared with the protein library.

[0072] blastn: Alignment of nucleic acid sequences to nucleic acid libraries, directly comparing the homology of nucleic acid sequences.

[0073] tblastn: Compare protein sequences to nucleic acid libraries, translate the nucleic acids in the library into protein sequences, and then compare them.

[0074] tblastx: compares nucleic acid sequences to nucleic acid libraries at the protein level, translates both the library and the sequence to be queried into protein sequences, and then compares the protein sequences.

[0075] S103, obtaining a matching structure of the matching sequence according to the matching sequence;

[0076] Specifically, after obtaining the final matching result, the best matching result is selected as the matching sequence, and then the structure of the matching sequence is selected as the model template.

[0077] S104, constructing an initial three-dimensional structure model of the target sequence based on the matching sequence and its matching structure;

[0078] Specifically, the matching sequence and its corresponding matching structure are obtained in the previous step. Then, the matching structure can be used as a structural template for the target sequence to construct an initial three-dimensional structural model of the target sequence. The target sequence is then aligned with the initial three-dimensional structural model one by one, and the matching sequence is aligned with the matching structure one by one. Then, a modeling engine (such as promod3) is used to generate a protein model.

[0079] S105, combining the unmatched sequence fragment of the target sequence with a portion of the adjacent matched sequence fragment to form a sub-target sequence; and searching for a matching subsequence and its structure of the sub-target sequence in a protein database of known structures;

[0080] Specifically, during the modeling process, missing parts in the structural model need to be filled in to complete its three-dimensional structure. For unmatched sequence fragments, some of the matched sequence fragments can be combined to form a new sequence fragment (sub-target sequence), and then blastp is called to find the structural template corresponding to the new sequence fragment, thereby filling the missing parts of the initial three-dimensional structural model of the target sequence.

[0081] S106, filling in the missing parts of the initial three-dimensional structure model according to the found matching subsequences and their structures, and obtaining the three-dimensional structure of the protein file to be tested.

[0082] Specifically, after the missing parts in the initial three-dimensional structure model are filled in using the above method, the three-dimensional structure of the protein file to be tested is obtained.

[0083] Another embodiment of the protein structure prediction method of the present invention is as follows Figure 3 As shown, this embodiment adds a detailed description of step S105 (S205--S209) on the basis of the previous embodiment. The structure prediction method of this embodiment specifically includes:

[0084] S201, extracting the target sequence from the protein file to be tested;

[0085] S202, matching the target sequence with a protein database of known structures to find a matching sequence;

[0086] S203, acquiring a matching structure of the matching sequence according to the matching sequence;

[0087] S204, constructing an initial three-dimensional structural model of the target sequence based on the matching sequence and its matching structure;

[0088] S205, obtaining sequence segments that match between the target sequence and the matching sequence as matched sequence segments; and obtaining sequence segments that do not match between the target sequence and the matching sequence as unmatched sequence segments;

[0089] S206, determining whether the length of the unmatched sequence segment is greater than a preset length value, if so, proceeding to step S207, otherwise, proceeding to step S208;

[0090] S207, determining the cleavage site in the matched sequence fragment, combining the unmatched sequence fragment with the adjacent matched sequence fragment after cleavage to form a sub-target sequence; proceeding to step S209;

[0091] S208, intercepting the unmatched sequence fragment as a sub-target sequence;

[0092] S209, matching the sub-target sequence with a protein database of known structures to find a matching sub-sequence and its structure;

[0093] S210 , filling in the missing parts in the initial three-dimensional structure model according to the found matching subsequences and their structures, to obtain the three-dimensional structure of the protein file to be tested.

[0094] For unmatched sequences whose sequence fragment length is less than or equal to the preset length, since the unmatched sequence fragment is short, the unmatched sequence fragment can be directly intercepted as a sub-target sequence, and then the sub-target sequence is searched in a protein database with known structures to obtain the corresponding template structure to fill the missing part corresponding to the unmatched sequence fragment on the initial three-dimensional structure of the target sequence.

[0095] For unmatched sequence fragments whose length is greater than the preset length value, a new sequence fragment (sub-target sequence) is formed by combining part of the adjacent matched sequence fragments, and then the new sequence fragment is matched and searched in the protein database of known structures to see if a matching sequence can be found, thereby obtaining the structural template of the new sequence fragment to fill the gap in the original three-dimensional structure model of the target sequence.

[0096] Selecting a portion of the adjacent matched sequence fragment will affect the composition of the new sequence fragment (sub-target sequence), and thus affect the subsequent matching search results. Therefore, how to select and determine the best cut site in the matched sequence fragment is particularly important. Preferably, in step S207, when the length of the unmatched sequence fragment is greater than a preset length value, determining the cut site in the matched sequence fragment and combining the unmatched sequence fragment with the adjacent cut matched sequence fragment to form the sub-target sequence specifically includes:

[0097] If the starting position of the unmatched sequence fragment is set to M e1 , the termination site is M s2 ; Set the adjacent matched sequence fragment of the unmatched sequence fragment as the first matched sequence fragment and / or the second matched sequence fragment; the first matched sequence fragment has M1 amino acids, and its starting site is M s1 , the termination site is M e1The second matching sequence fragment has M2 amino acids, and its starting site is M s2 , the termination site is M e2 ;

[0098] 1. When M s2 -M e1 >3 PM:

[0099] (1) Figure 4a As shown, if the matching sequence segment adjacent to the unmatched sequence segment is only the first matching sequence segment, the cleavage site T1 of the first matching sequence is intercepted to the termination site M of the unmatched sequence segment. S2 The sequence segments between are sub-target sequences;

[0100] (2) Figure 4b As shown, if the matching sequence segment adjacent to the unmatched sequence segment has only the second matching sequence segment, the starting site M of the unmatched sequence segment is intercepted. e1 The sequence fragment between the cleavage site T2 of the second matching sequence is used as the sub-target sequence;

[0101] (3) Figure 4c As shown, if the matching sequence segments adjacent to the unmatched sequence segment are the first matching sequence segment and the second matching sequence segment, the sequence segment between the cleavage site T1 of the first matching sequence and the cleavage site T2 of the second matching sequence is intercepted as the sub-target sequence;

[0102] Where: If M1 / 2>10, select the cleavage site T1 of the first matching sequence as M e1 -9;

[0103] If M1 / 2≦10, select the cleavage site T1 of the first matching sequence as (M e1 -M1 / 2+1); wherein, M1 / 2 is rounded up or down;

[0104] If M2 / 2>10, select the cleavage site T2 of the second matching sequence as M s2 +9;

[0105] If M2 / 2≦10, select the cleavage site T2 of the second matching sequence as (M s2 +M2 / 2-1); wherein, the M2 / 2 is rounded up or down.

[0106] 2. In step S208, intercepting the unmatched sequence fragment as a sub-target sequence specifically includes:

[0107] When M s2 -M e1 ≦15, intercept the unmatched sequence fragment Ms2 To M e1 as target sequence fragments.

[0108] In any of the above method embodiments, the target sequence is matched in a pre-stored protein structure database, and finding the matching sequence specifically includes: matching and searching the target sequence in the pre-stored protein structure database, and selecting the sequence with the highest similarity from the similar sequences found as the matching sequence; preferably, the best matching sequence is selected from the similar sequences found based on factors such as similarity and the length of the matching sequence fragment.

[0109] Specifically, if a protein of unknown structure shares sufficient sequence similarity with a protein of known structure, an approximate three-dimensional model can be constructed based on the principle of similarity. If a portion of the target protein sequence is similar to a domain region of a known protein, the target protein can be assumed to share the same domain or functional region. Homology modeling is the most reliable method for protein structure prediction.

[0110] The main idea of ​​the homology modeling method is: for a protein of unknown structure, find a homologous protein with a known structure, and use the structure of this protein as a template to build a structural model for the protein of unknown structure. Generally, if a sequence identity of 25% or more with the target sequence can be found (better, if the sequence identity of the two exceeds 30%), the homology modeling method can be used to predict the protein structure. The homology comparison of proteins is often carried out with the help of sequence alignment, and the evolutionary relationship between proteins can be discovered through sequence alignment. In terms of protein structure analysis, sequence conservation patterns or mutation patterns can be discovered through sequence alignment. These sequence patterns contain very useful three-dimensional structural information. The homology modeling method can predict the structure of 10-30% of proteins.

[0111] In general, after extracting the target sequence, similarity data is obtained by calling blastp (comparing protein sequences in a protein database) to obtain a result file. When evaluating and scoring, the Score value is the result of the scoring. The longer the matching fragment and the higher the similarity, the larger the Scoer value. Then we can select the sequence with the highest similarity, that is, the sequence with the highest Scoer value as the best matching sequence. Preferably, in addition to referring to its similarity, other reference factors can also be combined to select the best matching sequence. For example, considering the similarity to the target sequence, the number and length of the alignment gaps, etc., a sequence with high similarity to the target sequence and a small number and length of the alignment gaps can be selected as the template sequence. Of course, other reference factors can also be set, such as in conjunction with template resolution, etc.

[0112] After obtaining the matching sequence and matching structure, the initial three-dimensional model of the target sequence is constructed. For the sequence fragments that are not matched, the corresponding part of the structure is treated as missing data when constructing the initial three-dimensional model; if missing data occurs, the missing data corresponding to the target sequence, target structure, template sequence, and template structure are also removed. If there is data with no comparison results, it is necessary to generate a protein model based on promod3 (modeling engine). The specific complete process diagram is as follows Figure 5 shown.

[0113] Specifically, to generate a protein model using promod3 (a modeling engine), you first need to install all the dependencies required by promod3. For details, see the promod3 documentation. You need to pass in the input target sequence and structure as well as the matching sequence and structure alignment. Then the modeling steps are:

[0114] 1. Build an initial 3D structural model from the template structure.

[0115] 2. Perform loop modeling and fill missing gaps. Specifically, there are two methods for loop modeling, one is to calculate the loop prediction from scratch based on conformational search or conformational counting in a given environment, due to scoring or energy function guidance. Different protein representations, energy function terms and optimization or enumeration algorithms are used. The other is a database method for loop prediction, which includes finding the main chain segments of the two stem regions suitable for the loop, and performing a search for such fragments through databases of many known protein structures, rather than just simulating homologs of proteins. In this embodiment, for the missing gaps in the initial three-dimensional structure model, the present invention is used to combine unmatched sequence fragments with adjacent partially matched sequence fragments to form new sequence fragments, and then search for matches in the known structure protein database, search for matching sequence fragments, and then obtain the corresponding template method, and the corresponding template structure obtained is used to fill the corresponding missing gap part in the initial three-dimensional structure model. The search and filling are repeated in this way until a complete three-dimensional structure model is obtained.

[0116] 3. Side chain reconstruction. Specifically, because amino acid side chains in proteins are connected to α-carbon atoms via σ bonds, which are relatively flexible, the side chain conformation is difficult to determine. Side chain reconstruction involves using the SCWRL method to find the conformation of a specific amino acid side chain in a specific protein. By reconstructing the side chains of the three-dimensional structure of the protein being tested, the three-dimensional structural model of the protein being tested becomes more stable.

[0117] 4. Use molecular mechanics to minimize the energy of the final model. Specifically, using molecular mechanics to minimize the energy of the three-dimensional structural model of the protein to be tested can adjust the distances between atoms in the structure, stabilize the model structure, and thus optimize the three-dimensional structure of the protein to be tested.

[0118] Preferably, blastp finds only a small portion of similar sequences, and the rest cannot be found. Figure 4d As shown in FIG, assuming that the protein or its fragment for which structure prediction is required has N amino acids, the start site is Ns, and the end site is Ne. The fragment of the best blastp alignment result has Q amino acids, the start site is Qs, and the end site is Qe.

[0119] If Qs-Ns>15:

[0120] If Q / 2>10, the fragment from Ns to (Qs+9) is cut (cleavage site T=Qs+9), and blastp is called to search for the template.

[0121] If Q / 2<=10, the fragment from Ns to (Qs+Q / 2-1) is cut (cutting site T=Qs+Q / 2-1), and blastp is called to search for the template.

[0122] If Qs-Ns<=15:

[0123] The fragment from Ns to Qs was cut and the corresponding structure was searched in the protein database of known structure.

[0124] If Ne-Qe>15:

[0125] If Q / 2>10, the fragment from (Qe-9) to Ne is cut, and blastp is called to search for the template.

[0126] If Q / 2<=10, the fragment from (Qe-Q / 2+1) to Ne is cut, and blastp is called to search for the template.

[0127] If Ne-Qe<=15:

[0128] The fragment from Qe to Ne was extracted and the corresponding structure was retrieved from the protein database with known structure.

[0129] That is to say, if a matching sequence of the target sequence is searched for, if no matching sequence and corresponding structural template are found, for example, only some matching sequence fragments are found, then the fragment with the best alignment result is selected from the found matching sequence fragments as the target sequence fragment (for example, the sequence fragment with the longest matching fragment is selected as the fragment with the best alignment result), and based on the positional relationship of the target sequence fragment on the target sequence, the cleavage site is determined on the target sequence to obtain a sub-target sequence; based on the sub-target sequence, a template is searched in a protein database of known structures to construct an initial three-dimensional structure of the target sequence.

[0130] In addition, in proQed3 modeling, most of the filling is done for the missing parts in the middle, but if there is only one missing part (unmatched part) at both ends of the target sequence, it will not be filled. In this case, you can first remove the adjacent sequences including the three-dimensional structure, and then fill in the two missing parts.

[0131] When using promod3 for modeling, promod3 cannot predict some new amino acids such as U during the modeling process. The present invention needs to handle such cases separately, converting U into C through sequence recognition method and then performing structure prediction.

[0132] The promed3 modeling engine predicts three-dimensional structures by homology modeling. Based on this, this embodiment fully utilizes information in existing databases, including information on secondary structure and spatial structure, to predict the secondary structure from the protein sequence, making the protein visualization more structured.

[0133] When the sequence is continuously modeled to generate a complete three-dimensional structure model, we will get Figure 6 The final result, followed by the prediction of the secondary structure model, will be Figure 7 The final result. After the prediction result is obtained, a protein (PDB) file is returned.

[0134] In addition, if no similar sequence to the target sequence is found in the protein database of known structure (no comparison results, or no known sequence containing an equivalent sequence fragment is found), you can also make full use of the information in the existing database, including information on secondary structure and spatial structure. First, predict the secondary structure from the protein sequence, and then predict the spatial structure of the protein based on the secondary structure, or use ab initio prediction method for prediction.

[0135] Based on the same technical concept, the present invention discloses a protein structure prediction device, which can use any of the above protein structure prediction methods to predict the three-dimensional structure of a protein. Specifically, an embodiment of the protein structure prediction device of the present invention is as follows: Figure 8 As shown, including:

[0136] The sequence extraction module 100 is used to extract the target sequence from the protein file to be tested; specifically, the amino acid sequence is first extracted from the protein file. The atomic coordinates of the standard residue atoms mixed with non-standard residues need to be filtered. In addition, for the residue sequence numbers mixed with residue insertion codes, the residue insertion codes need to be added to the sequence numbers. Then, a sequence for each chain is obtained, and then deduplication is performed to remove duplicate sequences. The extracted sequences are the same but the chains or file names are different, so they are added together to obtain the nr library (protein library). The nr database is a merger of several protein sequence libraries. Then, blast (basic comparison search tool) is installed and the extracted sequence library is put into it to format the sequence library.

[0137] The matching search module 200 is used to match the target sequence against a database of proteins with known structures to find matching sequences. Based on the matching sequences, the matching structure of the matching sequences is then determined. Specifically, after obtaining the target sequence, blastp is used to search for similar sequences that match the target sequence to select a corresponding structural template. For example, if a matching sequence with a similarity of 25% or greater (preferably 30% or greater) is obtained, an initial three-dimensional structural model of the target sequence can be constructed using homology modeling.

[0138] A model building module 300 is used to build an initial three-dimensional structure model of the target sequence based on the matching sequence and its matching structure;

[0139] The combination module 400 is used to combine unmatched sequence segments of the target sequence with a portion of adjacent matched sequence segments to form a sub-target sequence. The matching search module then searches for matching subsequences and structures of the sub-target sequence in a protein database of known structures. Specifically, during the modeling process, missing portions of the structural model need to be filled to complete its three-dimensional structure. For unmatched sequence segments, a new sequence segment (sub-target sequence) can be formed by combining some of the matched sequence segments. blastp is then called to search for the structural template corresponding to this new sequence segment, thereby filling in the missing portions of the initial three-dimensional structural model of the target sequence.

[0140] Filling module 500 is used to fill in missing portions of the initial three-dimensional structure model based on the found matching subsequences and their structures, thereby obtaining the three-dimensional structure of the protein file to be tested. Specifically, after all missing portions of the initial three-dimensional structure model are filled in using the above-described method, the three-dimensional structure of the protein file to be tested is obtained.

[0141] Another embodiment of the prediction device of the present invention is as follows Figure 9 As shown, based on the above-mentioned prediction device embodiment, the combination module 400 includes:

[0142] The acquisition submodule 410 is used to acquire the sequence segments that match between the target sequence and the matching sequence as matched sequence segments; and to acquire the sequence segments that do not match between the target sequence and the matching sequence as unmatched sequence segments;

[0143] The determination submodule 420 is used to determine the cleavage site in the matched sequence fragment when the fragment length of the unmatched sequence fragment is greater than a preset length value, and to combine the unmatched sequence fragment with the adjacent matched sequence fragment after cleavage to form a sub-target sequence; it is also used to determine the cleavage site in the unmatched sequence fragment when the fragment length of the unmatched sequence fragment is less than or equal to the preset length value, and to intercept a part of the unmatched sequence fragment as the sub-target sequence; the matching search module matches the sub-target sequence in a protein database of known structures to find a matching subsequence and its structure.

[0144] For unmatched sequence fragments, since they are treated as missing structures in the initial three-dimensional structure, the missing parts in the initial three-dimensional structure need to be filled. For unmatched sequence fragments, different processing methods can be adopted according to the length of the fragments. Specifically, for unmatched sequence fragments whose length is less than or equal to the preset length, the cutting site can be determined on the unmatched sequence fragment, and a part of the unmatched sequence fragment can be intercepted as a sub-target sequence. The sub-target sequence is then searched in a protein database of known structures to obtain the corresponding template structure to fill the missing part corresponding to the unmatched sequence fragment in the initial three-dimensional structure of the target sequence. For the case where the length of the unmatched sequence fragment is greater than the preset length value, a new sequence fragment (sub-target sequence) can be formed by combining adjacent partially matched sequence fragments, and then the new sequence fragment is matched and searched in a protein database of known structures to see if a matching sequence can be found, thereby obtaining a structural template for the new sequence fragment to fill the missing part in the original three-dimensional structure model of the target sequence.

[0145] So, how to better determine the cutting site? Specifically, the determination submodule 420 includes a selection unit 420 and a cut unit 422; wherein: if the starting site of the unmatched sequence fragment is set to M e1 , the termination site is M s2 ; Set the adjacent matched sequence fragment of the unmatched sequence fragment as the first matched sequence fragment and / or the second matched sequence fragment; the first matched sequence fragment has M1 amino acids, and its starting site is M s1 , the termination site is M e1 The second matching sequence fragment has M2 amino acids, and its starting site is M s2 , the termination site is M e2 ;

[0146] 1. When Ms2 -M e1 >3 PM:

[0147] (1) If the matching sequence segment adjacent to the unmatched sequence segment is only the first matching sequence segment, the interception unit intercepts the cutting site T1 of the first matching sequence to the termination site M of the unmatched sequence segment. S2 The sequence segments between are sub-target sequences;

[0148] (2) If the matching sequence segment adjacent to the unmatched sequence segment has only the second matching sequence segment, the interception unit intercepts the starting position M of the unmatched sequence segment. e1 The sequence fragment between the cleavage site T2 of the second matching sequence is used as the sub-target sequence;

[0149] (3) If the matching sequence segments adjacent to the unmatched sequence segment are the first matching sequence segment and the second matching sequence segment, the clipping unit clips the sequence segment between the cleavage site T1 of the first matching sequence and the cleavage site T2 of the second matching sequence as the sub-target sequence;

[0150] Wherein: If M1 / 2>10, the selection unit selects the cutting site T1 of the first matching sequence as M e1 -9;

[0151] If M1 / 2≦10, the selection unit selects the cleavage site T1 of the first matching sequence as (M e1 -M1 / 2+1); wherein, M1 / 2 is rounded up or down;

[0152] If M2 / 2>10, the selection unit selects the cutting site T2 of the second matching sequence as M s2 +9;

[0153] If M2 / 2≦10, select the cleavage site T2 of the second matching sequence as (M s2 +M2 / 2-1); wherein, M2 / 2 is rounded up or down;

[0154] 3. When M s2 -M e1 ≤3:00 PM:

[0155] The interception unit intercepts the unmatched sequence fragment M s2 To M e1 as target sequence fragments.

[0156] Another embodiment of the prediction device of the present invention, based on any of the above device embodiments, further comprises:

[0157] A side chain building module 600 is used to reconstruct the side chains of the three-dimensional structure of the protein to be tested;

[0158] The structure optimization module 700 is used to minimize the energy of the three-dimensional structure model of the protein to be tested using molecular mechanics.

[0159] Based on any of the above prediction device embodiments, the protein structure prediction device further includes:

[0160] The secondary structure prediction module is used to predict the secondary structure of the protein to be tested based on the three-dimensional structure of the protein to be tested and the information on the three-dimensional spatial structure and secondary structure of each protein sequence in the protein database with known structures.

[0161] Specifically, for example, in this embodiment, the modeling engine promed3 can be used to predict the three-dimensional structure through homology modeling (including remote homology). On this basis, full use can be made of the information in the known structural protein database, including information on secondary structure and spatial structure, to predict the secondary structure from the protein sequence, making the protein visualization more structured.

[0162] In addition, in promed3 modeling, most of the filling is done for the missing parts in the middle, but if there is only one missing part (unmatched part) at both ends of the target sequence, it will not be filled. In this case, you can first remove the adjacent sequences including the three-dimensional structure, and then fill in the two missing parts.

[0163] During the protein modeling process, promod3 cannot predict some new amino acids such as U. The present invention needs to handle such cases separately, converting U into C through sequence recognition method before performing structure prediction.

[0164] Similarly, the present invention also discloses a storage medium storing a plurality of instructions, which are executed by one or more processors to implement the steps of the protein structure prediction method of any of the above-mentioned prediction method embodiments of the present invention.

[0165] For example, the storage medium of this embodiment stores multiple instructions to implement the following protein structure prediction method steps:

[0166] Extract the target sequence from the protein file to be tested;

[0167] Matching the target sequence with a protein database of known structures to find a matching sequence;

[0168] According to the matching sequence, obtaining a matching structure of the matching sequence;

[0169] Based on the matching sequence and its matching structure, constructing an initial three-dimensional structural model of the target sequence;

[0170] Combining the unmatched sequence fragment of the target sequence with a portion of the adjacent matched sequence fragment to form a sub-target sequence; and searching for a matching subsequence and its structure of the sub-target sequence in the protein database of known structure;

[0171] According to the found matching subsequences and their structures, the missing parts in the initial three-dimensional structure model are filled to obtain the three-dimensional structure of the protein file to be tested.

[0172] Of course, the storage medium storing the implementation instructions of the above-mentioned other protein structure prediction method embodiments also belongs to the storage medium of the present invention, and will not be described in detail here.

[0173] Finally, the present invention discloses a protein structure prediction platform, including a protein structure prediction device in any prediction device embodiment of the present invention. The protein structure prediction platform is constructed on a server and is installed with an online visualization program for protein three-dimensional structure and / or secondary structure to visualize the structure of the protein.

[0174] The protein structure prediction platform includes the protein structure prediction device of the present invention, which corresponds to the various steps of the protein structure prediction method of the present invention. The prediction platform is established on a server, preferably on a cloud server, and the prediction platform is also installed with an online visualization program for protein three-dimensional structure or secondary structure. Users can predict the structure of proteins through the protein structure prediction platform, and the prediction platform can also display the secondary structure or three-dimensional structure of proteins to users.

[0175] The protein structure prediction method of the present invention corresponds to the protein structure prediction device. The technical details of the embodiments of the protein structure prediction method of the present invention are also applicable to the protein structure prediction device of the present invention. To reduce repetition, they are not repeated here.

[0176] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0177] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0179] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0180] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A protein structure prediction method, characterized in that: include: Extract the target sequence from the protein file to be tested; Matching the target sequence with a protein database of known structures to find a matching sequence; According to the matching sequence, obtaining a matching structure of the matching sequence; Based on the matching sequence and its matching structure, constructing an initial three-dimensional structural model of the target sequence; Combining the unmatched sequence fragment of the target sequence with a portion of the adjacent matched sequence fragment to form a sub-target sequence; and searching for a matching subsequence and its structure of the sub-target sequence in the protein database of known structure; Filling in the missing parts of the initial three-dimensional structure model according to the found matching subsequences and their structures to obtain the three-dimensional structure of the protein file to be tested; The unmatched sequence segments of the target sequence are combined with a portion of adjacent matched sequence segments to form a sub-target sequence; Searching for a matching subsequence and structure of the subtarget sequence in the protein database of known structure specifically includes: obtaining sequence segments that match between the target sequence and the matching sequence as matched sequence segments; and obtaining sequence segments that do not match between the target sequence and the matching sequence as unmatched sequence segments; When the length of the unmatched sequence fragment is greater than a preset length value, determining a cut site in the matched sequence fragment, and combining the unmatched sequence fragment with an adjacent cut matched sequence fragment to form a sub-target sequence; When the length of the unmatched sequence segment is less than or equal to the preset length value, intercepting the unmatched sequence segment as a sub-target sequence; The sub-target sequence is matched against the protein database of known structure to find a matching sub-sequence and its structure.

2. A protein structure prediction method according to claim 1, characterized in that: When the length of the unmatched sequence fragment is greater than a preset length value, determining the cut site in the matched sequence fragment, and combining the unmatched sequence fragment with the adjacent cut matched sequence fragment to form a sub-target sequence specifically includes: If the starting position of the unmatched sequence fragment is set to M e1 , the termination site is M s2 ; Set the adjacent matched sequence fragment of the unmatched sequence fragment as the first matched sequence fragment and / or the second matched sequence fragment; the first matched sequence fragment has M1 amino acids, and its starting site is M s1 , the termination site is M e1 The second matching sequence fragment has M2 amino acids, and its starting site is M s2 , the termination site is M e2 ; When M s2 -M e1 >3 PM: If the matching sequence segment adjacent to the unmatched sequence segment is only the first matching sequence segment, the cleavage site T1 of the first matching sequence is intercepted to the termination site M of the unmatched sequence segment. s2 The sequence segments between are sub-target sequences; If the matching sequence segment adjacent to the unmatched sequence segment has only the second matching sequence segment, the starting position M of the unmatched sequence segment is intercepted. e1 The sequence fragment between the cleavage site T2 of the second matching sequence is used as the sub-target sequence; If the matching sequence segments adjacent to the unmatched sequence segment are the first matching sequence segment and the second matching sequence segment, the sequence segment between the cleavage site T1 of the first matching sequence and the cleavage site T2 of the second matching sequence is intercepted as the sub-target sequence; in: If M1 / 2>10, select the cleavage site T1 of the first matching sequence as M e1 -9; If M1 / 2≦10, select the cleavage site T1 of the first matching sequence as (M e1 -M1 / 2+1); wherein, M1 / 2 is rounded up or down; If M2 / 2>10, select the cleavage site T2 of the second matching sequence as M s2 +9; If M2 / 2≦10, select the cleavage site T2 of the second matching sequence as (M s2 +M2 / 2-1); wherein, M2 / 2 is rounded up or down; When the length of the unmatched sequence segment is less than or equal to the preset length value, intercepting the unmatched sequence segment as a sub-target sequence specifically includes: When M s2 -M e1 ≦15, intercept the unmatched sequence fragment M e1 To M s2 As a sub-target sequence fragment.

3. A protein structure prediction method according to claim 1, characterized in that: The step of constructing an initial three-dimensional structural model of the target sequence based on the matching sequence and its matching structure specifically includes: Obtaining matched sequence segments between the target sequence and the matching sequence; An initial three-dimensional structural model of the target sequence is constructed using the matching structure of the matching sequence as a template; in the initial three-dimensional structural model, the structure of the matched sequence fragment of the target sequence adopts the structure of the corresponding matched sequence fragment in the matching sequence; and the structure of the unmatched sequence fragment in the target sequence is processed as a missing structure.

4. A protein structure prediction method according to claim 1, characterized in that: Also includes: When no matching sequence or matching structure is found in the protein database of known structure, searching for a matching sequence fragment of the target sequence; Selecting the best matching sequence segment from the found matching sequence segments as the target sequence segment; Determine the cleavage site on the target sequence according to the positional relationship of the target sequence fragments on the target sequence to obtain a sub-target sequence; According to the sub-target sequence, a template is searched in a protein database of known structures to construct an initial three-dimensional structure of the target sequence.

5. A protein structure prediction method according to claim 4, characterized in that: Determining the cleavage site on the target sequence according to the positional relationship of the target sequence fragments on the target sequence, and obtaining the sub-target sequence specifically includes: If the target sequence is set to have N amino acids, its starting site is Ns and its ending site is Ne; the target sequence fragment is set to have Q amino acids, its starting site is Qs and its ending site is Qe; When Qs-Ns>15: If Q / 2>10, the cleavage site of the target sequence is selected as Qs+9, and the fragment from Ns to (Qs+9) is intercepted as the sub-target sequence; If Q / 2≦10, the cleavage site of the target sequence is selected as (Qs+Q / 2-1), and the fragment from Ns to (Qs+Q / 2-1) is intercepted as the sub-target sequence; wherein Q / 2 is rounded up or down; When Qs-Ns≦15, the cleavage site of the target sequence is selected as Qs, and the fragment from Ns to Qs is intercepted as the sub-target sequence; When Ne-Qe>15: If Q / 2>10, the cleavage site of the target sequence is selected as Qe-9, and the fragment from (Qe-9) to Ne is intercepted as the sub-target sequence; If Q / 2≦10, the cleavage site of the target sequence is selected as (Qe-Q / 2+1), and the fragment from (Qe-Q / 2+1) to Ne is intercepted as the sub-target sequence; wherein Q / 2 is rounded up or down; When Ne-Qe≦15, the cleavage site of the target sequence is selected as Qe, and the fragment from Qe to Ne is intercepted as the sub-target sequence.

6. A protein structure prediction method according to any one of claims 1 to 5, characterized in that: Also includes: Reconstructing the side chains of the three-dimensional structure of the protein to be tested; Molecular mechanics is used to minimize the energy of the three-dimensional structural model of the protein to be tested.

7. A protein structure prediction method according to claim 1, characterized in that: Also includes: The secondary structure of the protein to be tested is predicted based on the three-dimensional structure of the protein to be tested and combined with the information on the three-dimensional spatial structure and secondary structure of each protein sequence in the protein database with known structures.

8. A protein structure prediction device, characterized in that: include: Sequence extraction module, used to extract target sequences from the protein file to be tested; A matching search module is used to match the target sequence in a protein database with known structures to find a matching sequence; and obtain a matching structure of the matching sequence based on the matching sequence; A model building module, used to build an initial three-dimensional structural model of the target sequence based on the matching sequence and its matching structure; A combining module is configured to combine the unmatched sequence segments of the target sequence with a portion of the adjacent matched sequence segments to form a sub-target sequence; and to search for a matching subsequence and its structure of the sub-target sequence in the protein database of known structures through the matching search module; A filling module is used to fill in the missing parts in the initial three-dimensional structure model according to the found matching subsequences and their structures, so as to obtain the three-dimensional structure of the protein file to be tested; The combined module comprises: an acquisition submodule, configured to acquire sequence segments that match between the target sequence and the matching sequence as matched sequence segments; and to acquire sequence segments that do not match between the target sequence and the matching sequence as unmatched sequence segments; a determination submodule, configured to, when the length of the unmatched sequence fragment is greater than a preset length value, determine a cut site in the matched sequence fragment, and combine the unmatched sequence fragment with an adjacent cut matched sequence fragment to form a sub-target sequence; and, when the length of the unmatched sequence fragment is less than or equal to the preset length value, intercept the unmatched sequence fragment as a sub-target sequence; The matching search module matches the sub-target sequence in the protein database of known structures to find matching sub-sequences and their structures.

9. A protein structure prediction device according to claim 8, characterized in that: The determination submodule includes a selection unit and an interception unit; wherein: If the starting position of the unmatched sequence fragment is set to M e1 , the termination site is M s2 ; Set the adjacent matched sequence fragment of the unmatched sequence fragment as the first matched sequence fragment and / or the second matched sequence fragment; the first matched sequence fragment has M1 amino acids, and its starting site is M s1 , the termination site is M e1 The second matching sequence fragment has M2 amino acids, and its starting site is M s2 , the termination site is M e2 ; When M s2 -M e1 >3 PM: If the matching sequence segment adjacent to the unmatched sequence segment is only the first matching sequence segment, the interception unit intercepts the cutting site T1 of the first matching sequence to the termination site M of the unmatched sequence segment. s2 The sequence segments between are sub-target sequences; If the matching sequence segment adjacent to the unmatched sequence segment has only the second matching sequence segment, the interception unit intercepts the starting position M of the unmatched sequence segment. e1 The sequence fragment between the cleavage site T2 of the second matching sequence is used as the sub-target sequence; If the matching sequence segments adjacent to the unmatched sequence segment are the first matching sequence segment and the second matching sequence segment, the clipping unit clips the sequence segment between the cleavage site T1 of the first matching sequence and the cleavage site T2 of the second matching sequence as the sub-target sequence; in: If M1 / 2>10, the selection unit selects the cleavage site T1 of the first matching sequence as M e1 -9; If M1 / 2≦10, the selection unit selects the cleavage site T1 of the first matching sequence as (M e1 -M1 / 2+1); wherein, M1 / 2 is rounded up or down; If M2 / 2>10, the selection unit selects the cutting site T2 of the second matching sequence as M s2 +9; If M2 / 2≦10, the selection unit selects the cleavage site T2 of the second matching sequence as (M s2 +M2 / 2-1); wherein, M2 / 2 is rounded up or down; When M s2 -M e1 ≦15: the interception unit intercepts the unmatched sequence segment M s2 To M e1 As a sub-target sequence fragment.

10. The protein structure prediction device according to claim 8, characterized in that: Also includes: A side chain building block for reconstructing the side chain of the three-dimensional structure of the protein to be tested; The structure optimization module is used to minimize the energy of the three-dimensional structure model of the protein to be tested by using molecular mechanics.

11. A protein structure prediction device according to any one of claims 8 to 10, characterized in that: Also includes: The secondary structure prediction module is used to predict the secondary structure of the protein to be tested based on the three-dimensional structure of the protein to be tested and the information on the three-dimensional spatial structure and secondary structure of each protein sequence in the protein database with known structures.

12. A storage medium, characterized in that: The storage medium stores a plurality of instructions, and the plurality of instructions are executed by one or more processors to implement the steps of the protein structure prediction method according to any one of claims 1 to 7.

13. A protein structure prediction platform, characterized in that: The protein structure prediction device comprises the protein structure prediction device according to any one of claims 8 to 11, wherein the protein structure prediction platform is constructed on a server and is installed with an online visualization program for protein three-dimensional structure and secondary structure to visualize the structure of the protein.

Citation Information

Patent Citations

  • Prediction method for protein three-dimensional structure

    CN101294970A

  • Protein three-dimensional structure predicting method and predicting cloud platform build by same

    CN109300501A