Multi-modal deep learning intelligent design method based on protein sequence and structure

By using a multimodal deep learning intelligent design method, structurally compatible protein sequences are generated, which solves the problem of sequence-structure inconsistency in protein design in existing technologies, improves the accuracy of functional site prediction and structural stability, and optimizes the design process.

CN121213645AActive Publication Date: 2025-12-26SHUIMU BIOSCIENCES LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511757937.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2025-12-26
Estimated Expiration
2045-11-27

AI Technical Summary

Technical Problem

Existing technologies lack effective constraint mechanisms in protein sequence and structure design, making it impossible to effectively cope with structural changes and sequence variations. This results in inconsistent generated sequences with the target structure, low accuracy in predicting functional sites, and affects the quality of protein design and its ability to adapt to complex biological functions.

Method used

By constructing a multimodal deep learning intelligent design method, we can obtain multimodal collaborative feature node groups, generate an embedding constraint set, screen structurally compatible sequence fragments, and combine geometric offset analysis to optimize the residue combination at the splicing interface and generate the target protein structure design sequence.

Benefits of technology

It improves the accuracy of protein functional site optimization and structural stability prediction, ensures the overall consistency and functionality of the target protein structural sequence, and optimizes the operability and reliability of the protein design process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121213645A_ABST
    Figure CN121213645A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a multi-modal deep learning intelligent design method based on a protein sequence and structure, which comprises the following steps of: acquiring a protein three-dimensional structure and sequence characteristics, constructing a collaborative node with a gradient and an included angle synchronously changing, generating an embedding constraint according to the collaborative node, and then combining candidate fragments with hydrophobicity and direction matching, and screening a compatible path by evaluating geometric offset and conformation stability, and finally splicing and correcting high-matching fragments to generate a target protein design sequence. According to the method, by analyzing the multi-modal collaborative feature node group, collaborative modeling of the sequence and the structure is achieved, the accuracy of functional site optimization and stability prediction is improved, then compatibility sequence fragments are screened according to the gradient and included angle synchronous change, stability and accuracy of the compatibility sequence fragments are ensured through geometric offset analysis, finally, an adaptive fragment list is spliced and corrected, and the accuracy of functional site optimization and stability prediction is improved. The conformation deviation is reduced, the overall consistency and functionality of the target protein sequence are ensured, and the operability and reliability of the design process are optimized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a multi-modal deep learning intelligent design method based on protein sequence and structure. BACKGROUND

[0002] The technical field of artificial intelligence includes algorithm systems, data modeling mechanisms, and knowledge expression and processing methods that simulate human cognitive behavior to achieve autonomous learning and reasoning. The core content of this technical field includes using deep learning, graph neural networks, natural language processing, and other methods to model high-dimensional features of structured or unstructured data and generalize reasoning, thereby playing an intelligent decision-making role in image recognition, speech recognition, automatic reasoning, biological information processing, and other scenarios. Artificial intelligence technology is widely used in cross-disciplines such as smart medicine, intelligent manufacturing, and biological computing, and its development has driven the formation of key directions such as multi-modal perception fusion, human-machine collaborative systems, and knowledge-based intelligent behavior modeling. In particular, after integration with biological information science, a series of intelligent generation design methods based on biological sequence and structure have been derived.

[0003] Among them, the multi-modal deep learning intelligent design method based on protein sequence and structure refers to an intelligent generation design method that integrates protein primary sequence information and three-dimensional spatial structure data to build a joint modeling system based on deep neural networks. For protein function site optimization, structure stability prediction, and novel conformation design, relying on residue interaction matrices, spatial topological coordinate tensors, and amino acid evolution conservation scores as main input basis, local sequence fragment context embedding features are extracted through convolutional neural networks, and graph attention mechanisms are combined to encode and process structure adjacency graphs, forming a unified representation space across modalities. Then, based on the variational auto-encoding framework, conditional generation modeling is realized, and a multi-objective loss function is used to constrain the mapping consistency between generated sequences and target structures, thereby completing the collaborative modeling process between protein sequence and structure design.

[0004] Existing technologies lack effective constraint mechanisms in protein sequence and structure design, and cannot effectively deal with structural changes and sequence variations, resulting in inconsistencies between generated sequences and target structures. Traditional methods rely on single sequence-structure mapping and ignore the impact of local structural changes on overall stability. The lack of effective integration of multi-modal data results in low accuracy in function site prediction, affecting the overall quality of protein design and the ability to adapt to complex biological functions. SUMMARY

[0005] To solve the technical problems existing in the prior art, the embodiments of the present application provide a multi-modal deep learning intelligent design method based on protein sequence and structure. The technical solution is as follows:

[0006] A multi-modal deep learning intelligent design method based on protein sequence and structure, comprising the following steps:

[0007] S1: Obtain the 3Di structure string of the target protein, collect the three-dimensional coordinates, rotation angles and direction vectors of the main chain nodes, extract the index number, hydrophobicity value and charge polarity of each residue in the target protein sequence, construct a pairing relationship based on the gradient amplitude and angle change characteristics, screen the point position combination, and generate a multi-modal collaborative feature node group;

[0008] S2: Based on the multi-modal collaborative feature node group, extract the main chain torsion angle change value, residue potential difference value and spatial density change rate corresponding to the node, analyze the relative variation interval between the three types of values, screen the point set with the variation amplitude in the continuous gradient region, calculate the spatial mean and standard deviation, and generate an embedding constraint condition set;

[0009] S3: According to the window three-dimensional coordinate mean and main chain direction vector in the embedding constraint condition set, call the candidate residue group with the same length, compare the hydrophobicity value and spatial direction vector angle of each residue, and generate a candidate residue sequence fragment group;

[0010] S4: Based on the density distribution, main chain connection angle and spatial angle of the target embedding position, calculate the geometric offset of the three types of parameters, screen the sequence fragments compatible with the structure, and generate an adaptive sequence fragment list.

[0011] As a further scheme of the present application, the multi-modal collaborative feature node group includes collaborative feature points, associated domain gradient values and sequence domain angle change rates, the embedding constraint condition set specifically refers to the reference window coordinate mean, target main chain direction vector and torsion angle change threshold, the candidate residue sequence fragment group specifically refers to the hydrophobicity matched amino acid combination, consistent number after verification and potential connection path, and the adaptive sequence fragment list includes high-score compatibility sequences, geometric offset evaluation data and conformational stability fitting values.

[0012] As a further scheme of the present application, the acquisition step of the multi-modal collaborative feature node group is:

[0013] S101: Obtain the 3Di structure string of the target protein, collect the three-dimensional coordinate data, rotation angle value and direction vector coordinate information of the main chain nodes in the structure, extract the index number of each corresponding residue in the target protein sequence according to the spatial distribution order of each main chain node in the structure expression domain, and simultaneously call the hydrophobicity value and charge polarity value information of the corresponding residue in the database based on the index number, to generate a residue conformation attribute group;

[0014] S102: Based on the index number of the residues in the residue conformation attribute group, the corresponding gradient value and the rate of change of the angle data are called in the sequence and structure expression domain respectively, the difference interval of the gradient amplitude value and the rate of change of the angle value under each number is calculated, the gradient angle joint trajectory sequence is established according to the index number sequence of the residues, and the joint trajectory variation interval value is obtained.

[0015] S103: According to the joint trajectory variation interval value, the synchronous variation trend of the gradient amplitude value and the rate of change of the angle value is judged by adjacent points, the point combination of synchronous growth or synchronous attenuation of the gradient is screened out, and the three-dimensional coordinates, the hydrophobic value and the charge polarity value in the corresponding point combination are combined to form the multi-modal node attribute, and the multi-modal collaborative feature node group is generated.

[0016] As a further scheme of the application, the obtaining step of the embedding constraint condition set is:

[0017] S201: Based on the multi-modal collaborative feature node group, the torsion angle change value data of the main chain residues corresponding to the node is extracted, the potential value of the residues corresponding to the node is matched, the potential value difference between adjacent residues is compared according to the sequence index, the spatial density change rate of the corresponding node is collected in combination with the three-dimensional coordinate range, the joint variation sequence is established according to the node number of the three values, and the joint structure variation parameter set is generated.

[0018] S202: According to the distribution variation interval of the three types of values in the joint structure variation parameter set in the node index sequence, the start and end positions of the continuous gradient variation region are judged, the node index combination in which each parameter is simultaneously in the continuous unidirectional variation interval is screened, the whole node number set in the corresponding combination is obtained, and the continuous variation node set is established.

[0019] S203: The three-dimensional coordinate set and the direction vector value of the nodes contained in the continuous variation node set are called, the spatial coordinate mean value and the direction vector standard deviation are calculated respectively, the node number group corresponding to the minimum standard deviation is selected as the structure embedding reference condition in all standard deviation results, and the embedding constraint condition set is established in combination with the mean value and the vector data.

[0020] As a further scheme of the application, the obtaining step of the residue sequence fragment group is:

[0021] S301: According to the window three-dimensional coordinate mean value and the main chain direction vector value in the embedding constraint condition set, a candidate residue group of the same number is called, the hydrophobic value corresponding to each residue in the candidate group is extracted respectively, the angle value between the direction vector and the window direction vector is calculated, the corresponding matching matrix of the hydrophobic value difference and the angle value is obtained, and the hydrophobic angle matching data set is generated.

[0022] S302: Based on the hydrophobic angle matching data set, the path combination with the angle value and the hydrophobicity difference value located in the angle threshold interval and the hydrophobicity threshold interval is screened, the continuity of the residue number sequence in the screened path is judged, and the index difference value of the residue number sequence is judged to verify the sequence consistency, and a continuous sequence consistency path group is obtained;

[0023] S303: The direction vector sequence and the residue hydrophobicity gradient of all path combinations in the continuous sequence consistency path group are called, the trend direction of the direction vector direction change rate and the hydrophobicity change interval is extracted, the path combination with the consistent direction change trend and the hydrophobic gradient direction is integrated, and a candidate residue sequence fragment group is established.

[0024] As a further scheme of the application, the adaptive sequence fragment list acquisition step is:

[0025] S401: Based on the number information of each sequence fragment in the candidate residue sequence fragment group, the three-dimensional coordinate density value, the main chain connection angle value and the space angle value corresponding to the target embedded position are called, and the coordinate difference, the angle difference and the angle offset of each fragment are calculated, the total offset value of each fragment under the three types of parameters is obtained, and a three-parameter geometric offset set is generated;

[0026] S402: According to the number sequence of each path group in the three-parameter geometric offset set, the local conformation and stability value of the corresponding residue are extracted, the fitting accuracy value and the stability value are combined into a cooperative change sequence in a fixed range, and it is judged whether the change trend is within the fluctuation rate change threshold range, the number path satisfying the trend cooperation condition is obtained, and a cooperative conformation trend sequence is obtained.

[0027] S403: The path number in the cooperative conformation trend sequence is called, the offset combination value corresponding to the path number in the three-parameter geometric offset set is searched, the sequence path with the total offset value less than the geometric compatibility threshold is screened, and the screened path is numbered and integrated to establish an adaptive sequence fragment list.

[0028] As a further scheme of the application, the method further comprises:

[0029] S5: According to the high matching sequence fragment in the adaptive sequence fragment list, the residue side chain charge and the contact surface projection area at the interface are spliced, the rotation angle deviation of the connection part is identified, the direction continuity range and the boundary are corrected, and the adjusted residue combination is integrated into the target protein structure design sequence.

[0030] The target protein structure design sequence is specifically a complete splicing sequence, an optimized corrected interface conformation and a final residue three-dimensional coordinate set.

[0031] As a further scheme of the present application, the obtaining step of the target protein structure design sequence is:

[0032] S501: According to the numbering order of each high-matching sequence fragment in the adaptive sequence fragment list, the corresponding node number in the original structure coordinates is spliced, the side chain charge data and contact surface projection area value of the splicing interface are called, the electronic distribution and physical contact range information of the residue contact area are extracted, and the fragment splicing interface parameter set is obtained;

[0033] S502: Based on the contact surface information of each interface residue in the fragment splicing interface parameter set, the matching of the electronic repulsion region and the spatial coordination region in the residue docking combination is analyzed, the main chain rotation angle value of each interface connection part is detected, and the angle deviation value of the angle and the original main chain direction is calculated, and the connection angle offset value table is established;

[0034] S503: The connection points with offset values meeting the direction continuity range in the connection angle offset value table are called, the local conformation data of the key connection points are extracted and direction boundary difference correction processing is performed, the direction data of the corrected connection nodes is combined with the original sequence fragment combination, and the target protein structure design sequence is generated.

[0035] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:

[0036] In the present application, based on the generation and analysis of the multi-modal collaborative feature node group, the collaborative modeling of protein sequence and structure is realized, and the accuracy of protein function site optimization and structure stability prediction is improved. Through the synchronous feature analysis of gradient amplitude and angle change, sequence fragments with structural compatibility are selected, and geometric offset analysis ensures the stability and accuracy of the sequence. The splicing strategy of the adaptive sequence fragment list effectively reduces the conformation deviation and directional inconsistency problem, ensures the overall consistency and functionality of the target protein structure sequence, and optimizes the operability and reliability in the protein design process. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 The method flowchart of the present application is shown in the figure;

[0038] Figure 2 The acquisition flowchart of the multi-modal collaborative feature node group of the present application is shown in the figure;

[0039] Figure 3 The acquisition flowchart of the embedded constraint condition set of the present application is shown in the figure;

[0040] Figure 4 The acquisition flowchart of the candidate residue sequence fragment group of the present application is shown in the figure;

[0041] Figure 5The flowchart for obtaining the adaptive sequence fragment list of the application;

[0042] Figure 6 The flowchart for obtaining the target protein structure design sequence of the application. DETAILED DESCRIPTION

[0043] The technical solutions in the application will be described below with reference to the drawings.

[0044] In the embodiments of the application, the words such as “example”, “for example” and the like are used to represent an example, illustration or description. Any embodiment or design scheme described as “example” in the application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word “example” is intended to present the concept in a specific manner. In addition, in the embodiments of the application, the meaning expressed by “and / or” can be both, or can be one of the two.

[0045] In the embodiments of the application, “image” and “picture” can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized. “Of”, “corresponding” and “corresponding” can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized.

[0046] In the embodiments of the application, the subscript such as W1 may be written in the form of non-subscript such as W1 at times, and the meanings expressed are consistent when the distinction is not emphasized.

[0047] To make the technical problems, technical solutions and advantages to be solved by the application more clear, the following will be described in detail with reference to the drawings and specific embodiments.

[0048] Please refer to Figure 1 The application provides a technical solution: a multi-modal deep learning intelligent design method based on protein sequence and structure, including the following steps:

[0049] S1: obtaining a 3Di structure string of a target protein, collecting three-dimensional coordinates, rotation angles and direction vectors of backbone nodes, extracting index numbers, hydrophobicity values and charge polarity of each residue in the target protein sequence, calling gradient values and angle change rates in the sequence and structure expression domain corresponding to the numbers, constructing a pairing relationship based on gradient amplitude and angle synchronous change characteristics, screening point position combinations with corresponding trends, and generating a multi-modal collaborative feature node group;

[0050] S2: Based on the multi-modal collaborative feature node group, the main chain torsion angle change value, the residue potential difference value and the spatial density change rate corresponding to the node are extracted, the point set with the variation amplitude in the continuous gradient region is screened by analyzing the relative variation interval between the three types of values, the spatial mean value of the point set and the standard deviation of the direction vector are calculated, the standard deviation minimum set is called as the subsequent structure embedding reference condition, and the embedding constraint condition set is generated;

[0051] S3: According to the window three-dimensional coordinate mean value and the main chain direction vector in the embedding constraint condition set, a candidate residue group with the same length is called, the hydrophobicity value of each residue in the candidate group and the spatial direction vector of the window node are compared, the path with the matching interval of the included angle and the hydrophobicity difference value is combined, the path continuity is judged and the consistency of the amino acid number sequence is verified, the fragment path meeting the continuous structure direction characteristics and the hydrophobic gradient trend is integrated, and the candidate residue sequence fragment group is generated;

[0052]

[0053] S5: According to the high matching sequence fragment in the adaptive sequence fragment list, the fragments are spliced according to the node number sequence corresponding to the original structure coordinates, the residue side chain charge data and the contact surface projection area value at the splicing interface are analyzed, the residue docking combination mode at the fragment interface is analyzed, the rotation angle deviation of the connecting part is identified, the continuous structure sequence is integrated by limiting the direction continuity range and correcting the conformation of the key connecting point, and the target protein structure design sequence is generated.

[0054] The multi-modal collaborative feature node group includes collaborative feature points, associated domain gradient values and sequence domain included angle change rates, the embedding constraint condition set specifically refers to the reference window coordinate mean value, the target main chain direction vector and the torsion angle change threshold, the candidate residue sequence fragment group specifically refers to the hydrophobicity matching amino acid combination, the consistent number after verification and the potential connection path, the adaptive sequence fragment list includes the high score compatibility sequence, the geometric offset evaluation data and the conformation stability fitting value, and the target protein structure design sequence specifically refers to the complete splicing sequence, the optimized corrected interface conformation and the final residue three-dimensional coordinate set.

[0055] Please refer to Figure 2 , the acquisition step of S1 is:

[0056] ​S101: Obtain the 3Di structure string of the target protein, collect the three-dimensional coordinate data, rotation angle value and direction vector coordinate information of the main chain nodes in the structure, extract the index number of each corresponding residue in the target protein sequence according to the spatial distribution order of each main chain node in the structure expression domain, and call the hydrophobicity value and charge polarity value information of the corresponding residue in the database based on the index number to generate the residue conformation attribute group;

[0057] According to the obtained 3Di structure string representation of the target protein, human serum albumin (PDB ID: 1AO6), specifically the LGEKH segment with index numbers 201 to 205, first collect the three-dimensional coordinate data of each main chain node (Cα) in the structure of the segment, and extract the coordinates of leucine (L) at index 201 as (20.1, 15.3, 22.8) Å, the coordinates of glycine (G) at index 202 as (21.5, 18.9, 23.4) Å, the coordinates of glutamic acid (E) at index 203 as (24.9, 18.5, 25.1) Å, the coordinates of lysine (K) at index 204 as (25.3, 22.1, 26.3) Å, and the coordinates of histidine (H) at index 205 as (28.8, 22.5, 27.9) Å, and collect the main chain dihedral angle rotation angle values, for example, the φ angle of glutamic acid at index 203 is -65.7°, and the ψ angle is -42.1°, and then calculate the direction vectors between the main chain Cα atoms, such as the direction vector between indexes 201 and 202 (0.35, 0.90, 0.15) is obtained by subtracting and normalizing the coordinates. According to this spatial distribution order, the index numbers 201, 202, 203, 204, 205 of each corresponding residue in the target protein sequence are extracted, and then the hydrophobicity values of the corresponding numbered residues in the database are called based on the Kyte-Doolittle hydrophobicity scale, L201 is 3.8, G202 is -0.4, E203 is -3.5, K204 is -3.9, and H205 is -3.2. At the same time, the charge polarity value information under the pH 7.4 environment is called, L201 is neutral (0), G202 is neutral (0), E203 is negative charge (-1), K204 is positive charge (+1), and H205 is positive charge (+1). The three-dimensional coordinates, rotation angles, direction vectors, index numbers, hydrophobicity values and charge polarity values extracted above are structured and integrated to generate the residue conformation attribute group.

[0058] S102: Based on the index number of the residue in the residue conformation attribute group, call the corresponding gradient value and angle change rate data in the sequence and structure expression domain respectively, calculate the difference interval of the gradient amplitude value and the angle change rate value under each number, establish the gradient angle joint trajectory sequence according to the index number order of the residue, and obtain the joint trajectory variation interval value;

[0059] Based on the residue conformation attribute group, the hydrophobic gradient value of each residue is calculated in the sequence expression domain using the index number of each residue The gradient value here is defined as the absolute value of the difference between the hydrophobic values of adjacent residues, that is For example, the gradient value between indices 201 and 202 is calculated as Similarly, the gradient value between indices 202 and 203 is calculated as , In the structure expression domain, first calculate the angle between the vectors of Cα atoms i-1 to i and i to i+1 , then define the angle change rate as the absolute value of the difference between adjacent angles , set , , , , then the angle change rate is calculated as , is calculated as To unify the dimensions, the angle value is normalized by dividing by the maximum possible change range of 180° to obtain , Next, calculate the difference interval between the gradient amplitude value and the normalized angle change rate value , the difference is , According to the residue index number order (202, 203), establish the gradient angle joint trajectory sequence [3.078, 0.361], and obtain the joint trajectory variation interval value.

[0060] S103: According to the joint trajectory variation interval value, judge the synchronous variation trend of the gradient amplitude value and the angle change rate value, filter out the point combination of synchronous growth or synchronous decay of the gradient, and combine the three-dimensional coordinates, hydrophobicity value, and charge polarity value in the corresponding point combination to form a multi-modal node attribute, and generate a multi-modal collaborative feature node group;

[0061] According to the obtained joint trajectory variation interval value [0.361, 3.078], judge the synchronous variation trend of the gradient amplitude value and the angle change rate value , the judgment standard of the synchronous variation trend is: if and , it is determined as synchronous growth, if and , it is determined as synchronous decay, and a synchronization tolerance is set , the tolerance value is determined by statistical analysis of 1000 alpha-helix structures in the protein database PDB, and the difference value is taken as The difference value range that reaches 90% consistency of change trend is determined as 0.5, and when , it is determined to be synchronous. Taking the above data as an example, from index 201 to 202, from 4.2 to 3.1, from an initial value (for example ) to 0.022, which meets the synchronous decay trend, from index 202 to 203, from 3.1 to 0.4, from 0.022 to 0.039, which is a non-synchronous change. Therefore, the point combination (201, 202) with a synchronous decay trend is screened out, and the three-dimensional coordinates (20.1, 15.3, 22.8) and (21.5, 18.9, 23.4) corresponding to index 201 and 202 in the point combination are combined to form a multi-modal node attribute, and a multi-modal collaborative feature node group is generated.

[0062] Please refer to Figure 3 , the acquisition step of S2 is:

[0063] S201: Based on the multi-modal collaborative feature node group, the torsion angle change value data of the main chain residues corresponding to the node is extracted, the potential value of the residues corresponding to the node is matched, and the difference between the potential values of adjacent residues is compared with the sequence index, the spatial density change rate of the corresponding node is collected by combining the three-dimensional coordinate range, and the joint variation sequence is established according to the node number of the three values. Generate a joint structure variation parameter set;

[0064] Extract the torsion angle change value data of the main chain residues corresponding to the node (201, 202) in the multi-modal collaborative feature node group, specifically extract the ψ angle (-45.1°) of index 201 and the φ angle (-60.2°) of index 202, and the torsion angle change value is the difference between them , and match the potential value of the residues corresponding to the node. Here, the potential value is derived from the calculation of the electrostatic potential energy of the residues in the solvent environment, and the potential value of L201 is set to -0.05 unit charge, and the potential value of G202 is set to +0.02 unit charge. The difference between the potential values of adjacent residues is compared with the sequence index, and , by calculating the number of non-hydrogen atoms within a radius of 5Å around the Cα atom, the spatial density of the corresponding node is collected. The density of L201 is 18 atoms, and the density of G202 is 12 atoms. The spatial density change rate is ​, the joint variation sequence of torsion angle change value, potential difference value, and spatial density change rate is established according to the node number, that is, nodes (201, 202) correspond to (15.1°, 0.07, 0.333), and a joint structural variation parameter set is generated.

[0065] S202: According to the distribution variation interval of the three types of values in the node index sequence in the joint structural variation parameter set, the start and end positions of the continuous gradient variation region are judged, the node index combination in which each parameter is simultaneously in a continuous unidirectional variation interval is screened, the node number set in the corresponding combination is obtained, and a continuous variation node set is established.

[0066] According to the joint structural variation parameter set, the sequence (201-205) extended to 5 nodes is analyzed, and the following parameter table is obtained:

[0067] Table 1: Residue structural variation parameter table

[0068] Node index Torsion angle change value (°) Potential difference value Spatial density change rate 201 12.5 0.07 0.33 202 15.1 0.82 0.15 203 18.3 0.95 0.11 204 22.0 1.10 0.08 205 19.8 0.21 0.13

[0069] As shown in Table 1, by observing the distribution variation interval of the three types of values in the node index sequence, the start and end positions of the continuous gradient variation region are judged, where the continuous gradient variation region is defined as at least three nodes in which the three parameters simultaneously present monotonic increase or monotonic decrease. Analysis of the data shows that from node 202 to 204, the torsion angle change value sequence is (15.1, 18.3, 22.0), which presents monotonic increasing, the potential difference value sequence is (0.82, 0.95, 1.10), which presents monotonic increasing, and the spatial density change rate sequence is (0.15, 0.11, 0.08), which presents monotonic decreasing, so this interval does not meet the condition of simultaneous unidirectional variation of the three parameters. The judgment standard is adjusted to allow one of the parameters to fluctuate reversely within a specified tolerance (for example, 15%). Under this standard, the spatial density change rate of nodes 202-204 decreases, while the other two increase, which does not meet the condition. It is assumed that there is another segment (310-312) whose three parameters are (14.2, 16.8, 19.5), (0.4, 0.6, 0.8), and (0.25, 0.21, 0.18), respectively. Among them, the first two increase, and the third decreases. If the condition is relaxed to be consistent with the trend of two main parameters, then the combination is screened, which is a hypothetical increasing set, the node number set {201, 202, 203} in the corresponding combination is obtained, and a continuous variation node set is established.

[0070] S203: Call the three-dimensional coordinate set and the direction vector value of the nodes contained in the continuous variable node set, respectively calculate the spatial coordinate mean and the standard deviation of the direction vector, select the node number group corresponding to the minimum standard deviation in all standard deviation results as the structure embedding reference condition, and establish the embedding constraint condition set combining the mean and vector data;

[0071] Call the three-dimensional coordinate set and the direction vector value of the nodes contained in the continuous variable node set {201, 202, 203}, the coordinates are (20.1, 15.3, 22.8), (21.5, 18.9, 23.4), (24.9, 18.5, 25.1), and the direction vectors are , , First, calculate the spatial coordinate mean , which is calculated by summing the x, y, and z coordinates of all nodes and dividing by the number of nodes 3, that is , then calculate the standard deviation of the direction vector , first calculate the vector mean , then calculate the difference vector of each vector and the mean vector, then calculate the variance of each component of the difference vector, and finally take the square root of the sum of the variances of each component to get the standard deviation, for example, the standard deviation of the x component , and similarly , , the total standard deviation is , among all the standard deviation results, assuming that another node set {310, 311, 312} has a standard deviation of 0.31, since , select the node number group {310, 311, 312} with the smallest standard deviation as the reference condition for subsequent structure embedding, and combine the coordinate mean (35.4, 41.2, 19.8) and the direction vector mean (0.6, 0.7, -0.4) calculated to establish the embedding constraint condition set.

[0072] Please refer to Figure 4 , the acquisition steps of S3 are:

[0073] S301: According to the window three-dimensional coordinate mean and the main chain direction vector value in the embedding constraint condition set, call the same number of candidate residue groups, respectively extract the hydrophobicity value corresponding to each residue in the candidate group, and calculate the included angle value between the direction vector and the window direction vector, obtain the corresponding matching matrix of the hydrophobicity difference and the included angle value, and generate the hydrophobicity-included angle matching data set;

[0074] According to the embedding constraint condition set, the window length is 3 residues, the window three-dimensional coordinate mean is (35.4, 41.2, 19.8), and the main chain direction vector mean is , a database containing 1000 tripeptide fragment candidate residue groups is called, one candidate group (alanine-valine-isoleucine, AVI) is extracted, and the hydrophobicity values of each residue in the group are extracted, A is 1.8, V is 4.2, I is 4.5, the hydrophobicity difference between the corresponding positions of the original fragment (assuming LGE, hydrophobicity value is 3.8, -0.4, -3.5) is calculated, for example, the difference between the first A and L , the difference between the second V and G , the difference between the third I and E At the same time, the angle between the direction vector (for example, defined by its Cα-Cβ vector) of each residue in the candidate group and the average direction vector of the window main chain is calculated , assuming that the direction vector of A is , then the angle between it and is The hydrophobicity difference vector and the angle value vector of all candidate groups are matched correspondingly to generate a hydrophobic angle matching data set.

[0075] S302: Based on the hydrophobic angle matching data set, the path combination whose angle value and hydrophobicity difference are located in the angle threshold interval and the hydrophobicity threshold interval is screened, the continuity of the residue number sequence in the screened path is judged, and the index difference of the residue number sorting result is judged to verify the sequence consistency, and a continuous sequence consistency path group is obtained;

[0076] Based on the hydrophobic angle matching data set, the angle threshold interval and the hydrophobicity threshold interval are set to screen the path combination, the angle threshold is set to be less than 30°, and the hydrophobicity threshold is set to be less than 3.0. The setting of these two thresholds is based on statistical analysis of natural protein homologous substitution data, and the parameter range covering 85% of conservative substitution instances is selected. For the candidate group AVI, the hydrophobicity difference vector is [2.0, 4.6, 8.0], and both 4.6 and 8.0 are greater than , so this path combination is excluded. Another candidate group (phenylalanine-glycine-glutamic acid, FGE) is selected, and its hydrophobicity difference vector is [|2.8-3.8|=1.0, |-0.4-(-0.4)|=0, |-3.5-(-3.5)|=0], all values are less than 3.0, and its angle value vector is [25°, 15°, 28°], all values are less than 30°, so this path combination is retained. Next, the continuity of the residue number sequence in the screened path is judged. In the candidate fragment database, the FGE fragment is derived from the continuous positions of protein (PDBID: 2XW4) with indexes 55, 56, 57. By judging the index difference of the residue number sorting result, and , and the sequence consistency is verified to obtain a continuous sequence consistency path group.

[0077] S303: Call the direction vector sequence and the residue hydrophobic gradient of all path combinations in the continuous sequence consistency path group, extract the trend direction of the direction vector direction change rate and the hydrophobic change interval, integrate the path combinations with the consistent direction change trend and the hydrophobic gradient direction, and establish a candidate residue sequence fragment group;

[0078] Call the direction vector sequence and the residue hydrophobic gradient of the FGE path combination in the continuous sequence consistency path group. The direction vector sequence of this path is , the hydrophobic value sequence is , and the direction vector direction change rate, i.e., the sequence of the included angle between adjacent vectors, is calculated , the trend direction in the hydrophobic change interval is calculated. The hydrophobicity changes from 2.8 to -0.4 and then to -3.5, showing a monotonic decreasing trend. This trend indicates that the fragment is transitioning from a hydrophobic core to a hydrophilic surface. At this time, the change trend of the direction vector should show gradual divergence to adapt to the solvent environment, i.e., the included angle between the vectors should be greater than , assuming that the calculated , satisfies , the direction change trend is consistent with the hydrophobic gradient direction, and this FGE path combination is integrated to establish a candidate residue sequence fragment group.

[0079] Please refer to Figure 5 , the acquisition step of S4 is:

[0080] S401: Based on the number information of each sequence fragment in the candidate residue sequence fragment group, call the three-dimensional coordinate density value, the main chain connection angle value, and the space included angle value corresponding to the target embedded position, and perform item-by-item calculation of the coordinate difference, the angle difference, and the included angle offset of the structure parameters corresponding to each fragment to obtain the total offset value of each fragment under the three types of parameters, and generate a three-parameter geometric offset set.

[0081] The FGE sequence fragment in the candidate residue sequence fragment group is selected, and the average atomic density value of the target position is set to 15 atoms / 100ų, the main chain connection angle with the previous residue is 110°, the main chain connection angle with the next residue is 115°, and the space angle between the average Cα-Cβ vector of the fragment itself and the main chain direction is 35° according to the three-dimensional coordinate density value, the main chain connection angle value and the space angle value corresponding to the target embedded position (i.e. the original LGE fragment position) of the FGE sequence fragment in the original protein. Then the structure parameters of the FGE fragment itself are calculated, and the average atomic density is set to 18 atoms / 100ų, the main chain connection angles are 112° and 113° respectively, and the space angle is 42°. The coordinate difference, angle difference and angle offset are calculated one by one, the density difference , the angle difference , and the angle offset are calculated. In order to obtain the total offset value, the weights of the three types of parameters are set, and the weight is obtained according to the large-scale mutagenesis experiment data through multivariate regression analysis to determine the contribution of each parameter to the structure stability, and the weight is set. Then the total offset value is obtained. The calculation is performed for all candidate fragments to generate a three-parameter geometric offset set.

[0082] S402: According to the numbering sequence of each path group in the three-parameter geometric offset set, the local conformation and stability value of the corresponding residue are extracted, and the fitting accuracy value and the stability value are combined to form a coordinated change sequence in a fixed range. It is judged whether the change trend is within the fluctuation rate change threshold range, the numbering path satisfying the trend coordination condition is obtained, and a coordinated conformation trend sequence is obtained.

[0083] According to the three-parameter geometric offset set, the numbering sequence of the FGE path group is extracted, and the local conformation and stability value are called. The conformation is determined by the Lagrange graph area, and the FGE is located in the allowed area. The stability value is evaluated by using the statistical energy score based on the distance potential. The lower the score, the more stable. The stability score of the FGE is set to -25.8, and the offset value of another candidate fragment WYE is 3.520, and the stability score is -28.4. The fitting accuracy value is defined as the reciprocal of the total offset value, that is , the fitting accuracy of the FGE is , and the fitting accuracy of the WYE is The fitting accuracy value and the stability value of a series of candidate fragments are combined to form a coordinated change sequence, and it is judged whether the change trend is within the fluctuation rate change threshold range. The threshold is determined by analyzing the relationship between sequence variation and stability change in the natural protein family, and is set to the linear regression slope When the fitting accuracy is improved (the offset value is reduced), the stability score should be reduced (more stable) accordingly. Linear fitting is performed on the two points (0.243, -25.8) and (0.284, -28.4), and the slope is Since the stability score is lower the better, the negative slope represents a positive correlation, which meets the trend synergy condition. The number paths FGE and WYE that meet the condition are obtained, and the conformation trend sequence is obtained.

[0084] S403: Call the path number in the conformation trend sequence, retrieve the corresponding offset combination value in the three-parameter geometric offset set, filter the sequence paths with a total offset value less than the geometric compatibility threshold, and number and integrate the filtered paths to establish an adaptive sequence fragment list.

[0085] Call the path numbers FGE and WYE in the obtained conformation trend sequence, retrieve the corresponding offset combination values in the three-parameter geometric offset set, which are 4.115 and 3.520 respectively, and set a geometric compatibility threshold The threshold is determined by a retrospective analysis of 500 verified successful protein design cases, calculating the geometric offset value of the replacement fragment, and selecting the upper limit of the offset value that covers 95% of the successful cases. The threshold is set to Filter the sequence paths with a total offset value less than the threshold. Comparison shows that the offset value of FGE is 4.115, which is greater than 4.0 and does not meet the condition, and the offset value of WYE is 3.520, which is less than 4.0 and meets the condition. Therefore, the WYE path is numbered and integrated to establish an adaptive sequence fragment list.

[0086] Please refer to Figure 6 The acquisition step of S5 is:

[0087] S501: According to the numbering order of each high-matching sequence fragment in the adaptive sequence fragment list, splice it according to the corresponding node number in the original structure coordinates, call the side chain charge data and contact surface projection area value of the splicing interface residues, extract the electronic distribution and physical contact range information of the residue contact area, and obtain the fragment splicing interface parameter set.

[0088] According to the adaptive sequence fragment list, the highest matching sequence fragment WYE is selected, and the corresponding node number order (201, 202, 203) in the original structure coordinates is spliced to replace the original LGE fragment. Next, the side chain charge data and contact surface projection area value of the residues at the splicing interface are called. Interface 1 is the connection between the original residue 200 and the first residue W (tryptophan) of the WYE fragment, and interface 2 is the connection between the last residue E (glutamic acid) of the WYE fragment and the original residue 204. Assuming that residue 200 is aspartic acid (D, negative), residue 204 is lysine (K, positive), W is neutral, and E is negative, the side chain charges of D and W at interface 1 are -1 and 0, respectively, and the side chain charges of E and K at interface 2 are -1 and +1, respectively. By calculating the projection of the side chain atoms of the two residues on the contact plane, the contact surface projection area is obtained. Assuming that the area value of interface 1 is 45.2 Å2 and that of interface 2 is 58.6 Å2, these information together constitute the fragment splicing interface parameter set, as shown in the following table.

[0089] Table 2: Fragment splicing interface parameter table

[0090] Interface number Connected residue pair Side chain charge pair Contact surface projected area (A2) 1 D200-W201 (-1,0) 45.2 2 E203-K204 (-1,+1) 58.6

[0091] As shown in Table 2, the table lists the key physicochemical parameters of the two interface positions after splicing.

[0092] S502: Based on the contact surface information of each interface residue in the fragment splicing interface parameter set, analyze the matching of the electron repulsion region and the spatial coordination region in the residue docking combination, detect the main chain rotation angle value of each interface connection position, and calculate the angle deviation value of the angle from the original main chain direction to establish a connection angle deviation value table;

[0093] Analyze the contact surface information of each interface residue in the fragment splicing interface parameter set. At interface 2, the negative charge of residue E and the positive charge of residue K form a favorable salt bridge, which belongs to spatial coordination region matching. At interface 1, the negative charge of D and the neutral bulky side chain of W have no obvious charge repulsion, but the van der Waals radius needs to be checked to avoid spatial steric conflict. Through atomic coordinate calculation, it is found that the atomic distance is not less than 0.8 times the sum of their van der Waals radii, and there is no serious collision. Subsequently, the main chain rotation angle value of each interface connection position is detected, especially the peptide bond plane angle ω. Assuming that the ω angle measurement value of D200-W201' at interface 1 is 178.5°, and the ω angle of E203'-K204 at interface 2 is -177.0°, the angle deviation value of the angle from the ideal trans peptide bond (180°) or the original main chain direction is calculated. The deviation of interface 1 is , and the deviation of interface 2 is These deviation values and the deviations of φ and ψ angles are recorded together to establish a connection angle deviation value table.

[0094] Table 3: Table of connecting angle offset values

[0095] Interface number ω angle deviation (°) φ angle deviation (°) ψ angle deviation (°) 1 1.5 5.2 8.1 2 3.0 6.8 4.5

[0096] As shown in Table 3, the table quantifies the disturbance caused by the splicing operation to the main chain geometric conformation.

[0097] S503: Call the connecting points whose offset values in the connecting angle offset value table meet the direction continuity range, extract the local conformation data of the key connecting points and perform direction boundary difference correction processing, combine and integrate the direction data of the corrected connecting nodes with the original sequence fragments to generate the target protein structure design sequence;

[0098] Call the offset values recorded in the connecting angle offset value table and compare them with the preset direction continuity range, which is determined by the statistical distribution of the main chain angles in the high-resolution crystal structure database, set the ω angle deviation to be less than 5°, and the φ / ψ angle deviation to be less than 10°, as shown in Table 3, all offset values of interface 1 and interface 2 are within the respective threshold range, so it is determined that both connecting points meet the direction continuity requirement and do not need to perform boundary correction, if it is assumed that the φ angle deviation of interface 1 is 12.7°, which exceeds the threshold of 10°, then the direction boundary difference correction processing needs to be performed for the connecting point, the specific execution process of the correction processing is as follows: first, take the ψ angle of residue D200 at interface 1 and the φ angle of residue W201' as adjustable variables, and perform combined fine-tuning within a ±5° range with a step size of 1°, generating a series of candidate conformations, for each candidate conformation, recalculate the deviation value of the φ angle of W201' from the original main chain direction, and discard the conformations whose deviation values are still greater than 10°, for the conformations that meet the angle deviation requirement, further calculate their local comprehensive energy score, which is composed of two parts, the weights are obtained by regression analysis according to the contribution of each energy term to the structure stability in large-scale molecular dynamics simulation, set the weight of van der Waals term to 0.6 and the weight of electrostatic term to 0.4, the first is the van der Waals collision penalty term, which is calculated by the distance between the side chain atoms of D200 and W201', when the distance is less than 0.9 times the sum of the van der Waals radii, a energy penalty value is given which is inversely proportional to the square of the distance difference, for example, the initial collision penalty value is 15.3, the second is the electrostatic interaction energy, which is calculated based on the local charge distribution of the side chain atoms of D200 and W201', the initial value is -5.8, the initial comprehensive energy score is After traversing all candidate conformations, the conformation that makes the local integrated energy score lowest and satisfies the angle deviation constraint is selected as the correction result. For example, when the ψ angle of D200 is adjusted by +2° and the φ angle of W201' is adjusted by -3°, the φ angle deviation of W201' is reduced to 9.5°, the van der Waals collision penalty value is reduced to 2.1 due to the increase of the interatomic distance, and the electrostatic interaction energy is changed to -6.2 due to the optimization of the relative position. At this time, the new integrated energy score is The conformation is adopted because its integrated energy score is the lowest. The final angle data of the corrected connection node (D200-W201') is updated to the structure coordinates, and is combined and integrated with the WYE sequence fragment and the remaining part of the protein to generate the target protein structure design sequence.

[0099] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which shall be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A multimodal deep learning intelligent design method based on protein sequence and structure, characterized in that, Includes the following steps: S1: Obtain the 3Di structure string of the target protein, collect the three-dimensional coordinates, rotation angle and direction vector of the main chain node, extract the index number, hydrophobicity value and charge polarity of each residue in the target protein sequence, construct the pairing relationship based on the gradient amplitude and angle change characteristics, screen the site combination, and generate multimodal collaborative feature node group. S2: Based on the multimodal collaborative feature node group, extract the main chain twist angle change value, residue potential difference value and spatial density change rate corresponding to the node. By analyzing the relative variation range between the three types of values, screen out the set of points with variation amplitude in the continuous gradient region, calculate the spatial mean and standard deviation, and generate the embedded constraint condition set. S3: Based on the mean of the three-dimensional coordinates of the window in the embedded constraint set and the main chain direction vector, call the candidate residue group of the same length, compare the hydrophobicity value of each residue with the angle between the spatial direction vector, and generate a candidate residue sequence fragment group. S4: Based on the density distribution of the candidate residue sequence fragment group and the target embedding position, the main chain connection angle and the spatial angle, calculate the geometric offset of three types of parameters, screen structurally compatible sequence fragments, and generate an adaptive sequence fragment list.

2. The multimodal deep learning intelligent design method based on protein sequence and structure according to claim 1, characterized in that: The multimodal collaborative feature node group includes collaborative feature points, associated structural domain gradient values ​​and sequence domain angle change rates. The embedding constraint set specifically includes the reference window coordinate mean, target main chain direction vector and twist angle change threshold. The candidate residue sequence fragment group specifically refers to hydrophobic matching amino acid combinations, verified consistency numbers and potential connection paths. The adaptive sequence fragment list includes high-scoring compatibility sequences, geometric offset evaluation data and conformational stability fitting values.

3. The multimodal deep learning intelligent design method based on protein sequence and structure according to claim 1, characterized in that, The steps for obtaining the multimodal collaborative feature node group are as follows: S101: Obtain the 3Di structure string of the target protein, collect the three-dimensional coordinate data, rotation angle value and direction vector coordinate information of the main chain nodes in the structure, extract the index number of each corresponding residue in the target protein sequence according to the spatial distribution order of each main chain node in the structural expression domain, and at the same time, call the hydrophobicity value and charge polarity value information of the corresponding residue in the database based on the index number to generate the residue conformation attribute group. S102: Based on the index number of the residue in the residue conformation attribute group, call the corresponding gradient value and angle change rate data in the sequence and structure expression domains respectively, calculate the difference range between the gradient amplitude value and the angle change rate value under each number, establish the gradient angle joint trajectory sequence according to the residue index number order, and obtain the joint trajectory variation range value. S103: Based on the joint trajectory variation interval value, the adjacent points are judged according to the synchronous variation trend of the gradient amplitude value and the angle change rate value, and the point combination with synchronous gradient growth or synchronous decay is selected. The three-dimensional coordinates, hydrophobicity value and charge polarity value in the corresponding point combination are combined to form multimodal node attributes, and multimodal collaborative feature node group is generated.

4. The multimodal deep learning intelligent design method based on protein sequence and structure according to claim 1, characterized in that, The steps for obtaining the embedded constraint set are as follows: S201: Based on the multimodal collaborative feature node group, extract the twist angle change value data of the main chain residue corresponding to the node, match the potential value of the residue corresponding to the node, compare the potential value difference between adjacent residues by sequence index, collect the spatial density change rate of the corresponding node in combination with the three-dimensional coordinate range, establish a joint variation sequence based on the node number of the three values, and generate a joint structural variation parameter set. S202: Based on the distribution and change range of the three types of values ​​in the node index sequence in the set of joint structure variation parameters, determine the start and end positions of the continuous gradient change zone, screen out the node index combinations in which all parameters are simultaneously in the continuous unidirectional change zone, obtain the set of all node numbers in the corresponding combination, and establish a continuously changing node set. S203: Call the three-dimensional coordinate set and direction vector values ​​of the nodes contained in the continuously changing node set, calculate the mean of the spatial coordinates and the standard deviation of the direction vector respectively, select the node number group corresponding to the minimum standard deviation from all standard deviation results as the structural embedding reference condition, and establish the embedding constraint condition set by combining its mean and vector data.

5. The multimodal deep learning intelligent design method based on protein sequence and structure according to claim 1, characterized in that, The steps for obtaining the residue sequence fragment group are as follows: S301: Based on the mean of the three-dimensional coordinates of the window and the value of the main chain direction vector in the embedded constraint condition set, call the same number of candidate residue groups, extract the hydrophobicity value corresponding to each residue in the candidate group, calculate the angle between its direction vector and the window direction vector, obtain the corresponding matching matrix of hydrophobicity difference and angle value, and generate a hydrophobic angle matching dataset. S302: Based on the hydrophobic angle matching dataset, filter path combinations where the angle value and the hydrophobicity difference are simultaneously located in the angle threshold range and the hydrophobicity threshold range, determine the continuity of the residue number sequence in the filtered path, and perform index difference judgment on the residue number sorting result to verify sequence consistency, thereby obtaining a continuous sequence consistent path group. S303: Call the direction vector sequence and residue hydrophobic gradient of all path combinations in the continuous sequence consistency path group, extract the trend direction of the direction vector change rate and the hydrophobic change interval, integrate the path combinations that are consistent with the direction of the hydrophobic gradient, and establish a candidate residue sequence fragment group.

6. The multimodal deep learning intelligent design method based on protein sequence and structure according to claim 1, characterized in that, The steps for obtaining the adaptive sequence fragment list are as follows: S401: Based on the numbering information of each group of sequence fragments in the candidate residue sequence fragment group, call the three-dimensional coordinate density value, main chain connection angle value and spatial angle value corresponding to its target embedding position, and calculate the coordinate difference, angle difference and angle offset item by item with the structural parameters corresponding to each group of fragments to obtain the total offset value of each fragment under the three types of parameters, and generate a three-parameter geometric offset set. S402: Based on the numbering sequence of each path group in the three-parameter geometric offset set, extract the local conformation and stability values ​​of the corresponding residues, combine the fitting accuracy value and stability value into a cooperative change sequence within a fixed range, determine whether its change trend is within the fluctuation rate change threshold range, obtain the numbered path that meets the trend coordination condition, and obtain the cooperative conformation trend sequence. S403: Call the path number in the cooperative conformation trend sequence, retrieve the corresponding offset combination value in the three-parameter geometric offset set, filter the sequence paths whose total offset value is less than the geometric compatibility threshold, and integrate the filtered paths by number to establish an adaptive sequence fragment list.

7. The multimodal deep learning intelligent design method based on protein sequence and structure according to claim 1, characterized in that, The method further includes: S5: Based on the highly matched sequence fragments in the adaptive sequence fragment list, the side chain charge of the residues at the splicing interface and the projected area of ​​the contact surface are used to identify the rotation angle deviation of the connection site. By limiting the range of directional continuity and boundary correction, the adjusted residue combination is integrated into the target protein structure design sequence. The target protein structure design sequence specifically includes a complete splicing sequence, an optimized and corrected interface conformation, and a final set of three-dimensional residue coordinates.

8. The multimodal deep learning intelligent design method based on protein sequence and structure according to claim 7, characterized in that, The steps for obtaining the target protein structure design sequence are as follows: S501: According to the numbering order of each highly matched sequence fragment in the adaptive sequence fragment list, splice them according to the node number corresponding to them in the original structural coordinates, call the side chain charge data and contact surface projection area value of the residues at the splicing interface, extract the electronic distribution and physical contact range information of the residue contact area, and obtain the fragment splicing interface parameter set; S502: Based on the contact surface information of each interface residue in the fragment splicing interface parameter set, analyze the matching of the electron repulsion region and the spatial coordination region in the residue docking combination, detect the main chain rotation angle value of each interface connection part, calculate the angle deviation value between the angle and the original main chain direction, and establish a connection angle offset value table. S503: Call the connection points whose offset values ​​in the connection angle offset value table meet the directional continuity range, extract the local conformation data of the key connection points and perform directional boundary difference correction processing, combine and integrate the directional data of the corrected connection nodes with the original sequence fragments to generate the target protein structure design sequence.

Citation Information

Patent Citations

  • Protein structure prediction method and system based on deep learning

    CN112233723A

  • Multi-modal protein design method, device and system and storage medium thereof

    CN119943206A

  • Multi-dimensional screening method for protein design based on lexicographical order optimization strategy

    CN120280002A

  • Efficient protein stability prediction method for selective state space modeling

    CN120496640A

  • Protein inverse folding method based on multi-modal pretrained large model

    US20250246261A1