A method for predicting sequence structure of biological macromolecule

By introducing geometric position encoding and multi-scale attention broadcasting mechanism into the diffusion model, the calculation of amino acid positional relationships is optimized, solving the problem that long-chain amino acid interactions are difficult to capture in traditional methods, and improving the accuracy and efficiency of biomacromolecule structure prediction.

CN120853678BActive Publication Date: 2025-11-25INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511368692.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-11-25
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Traditional attention mechanisms struggle to accurately capture interactions in long-chain amino acid sequences, leading to decreased accuracy in predicting the structure of biological macromolecules.

Method used

By introducing a geometric position encoding gating mechanism and a multi-scale attention broadcasting mechanism (PAB), the calculation of amino acid positional relationships is optimized through dynamic geometric position encoding and attention weight adjustment, thereby improving the model's ability to identify long-chain amino acid interactions.

Benefits of technology

It improves the accuracy and efficiency of predicting the sequence structure of biological macromolecules, reduces the amount of computation, enhances the model's ability to perceive amino acid position information, and reduces structure prediction errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853678B_ABST
    Figure CN120853678B_ABST
Patent Text Reader

Abstract

The application discloses a biological macromolecule sequence structure prediction method, and relates to the technical field of computers. In the diffusion process, a geometric position coding gating mechanism is introduced, and a multi-scale attention broadcast mechanism is used to broadcast the attention weights corresponding to different attention types of the target information processed in different broadcast range modes, and the attention weights in different time steps in the broadcast range are reused, so that the same attention weight is used in the steps of the broadcast range, the calculation amount is reduced, the model efficiency performance is improved, the perception ability of the model to the amino acid position information in the protein sequence can be effectively improved, the model can fully consider the position relationship between amino acids when calculating the attention weight, the model can better identify the long-chain amino acid interaction, the structure prediction error caused by the missing position information is reduced, and the prediction precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, and in particular to a biological macromolecule sequence structure prediction method. BACKGROUND

[0002] Current traditional protein structure determination methods, such as X-ray crystal diffraction, nuclear magnetic resonance spectroscopy, etc., can obtain high-precision structure information, but have the disadvantages of long experimental cycle, high cost, complex operation, etc., and are difficult to meet the demand of large-scale protein structure determination. For large-scale protein structure determination, the artificial intelligence (AI) software AlphaFold3 provided by the related technology successfully predicts the structures and interactions of almost all biological macromolecules (proteins, DNA, RNA, ligands, etc.) with unprecedented accuracy. This ability to accurately predict and analyze the complex interactions of proteins and other molecules with computers helps bring new insights into disease pathways, genomics, therapeutic targets, protein engineering, and synthetic biology. However, the original diffusion model used in AlphaFold3 has the problem of decreasing prediction accuracy as the length of the amino acid increases in the iterative denoising process, because the traditional attention mechanism cannot maintain the interaction of long-chain amino acids. When calculating the attention weight, the traditional attention mechanism cannot accurately capture the dependency relationship between long-chain amino acids as the sequence length increases, resulting in a decrease in the overall structure prediction accuracy of long sequence proteins. SUMMARY

[0003] The present application provides a biological macromolecule sequence structure prediction method to at least solve the problem of decreasing prediction accuracy caused by the difficulty of the attention mechanism in accurately capturing the dependency relationship between long-chain amino acids when calculating the attention weight in the iterative denoising process of the diffusion model in the related technology.

[0004] The present application provides a biological macromolecule sequence structure prediction method, comprising:

[0005] Encoding the geometric position information of the amino acids in the biological macromolecule sequence, determining the dynamic geometric position encoding of the relative spatial relationship between a plurality of residue pairs, and generating the geometric position encoding adjusted by the gate signal by adjusting the weight of the dynamic geometric position encoding;

[0006] Combining the geometric position encoding with the attention mechanism to determine the attention weight that adjusts the position relationship between the amino acids, and adjusting the broadcast range of the attention weight according to the attention type of the target information processed;

[0007] Based on the geometric position encoding, the attention weight, and the broadcast range of the attention weight, correcting the residue position, and updating the geometric position encoding according to the corrected residue position;

[0008] determine a predicted biological macromolecular sequence structure according to the updated geometric position encoding.

[0009] The application introduces a geometric position encoding gating mechanism in the diffusion process, uses a multi-scale attention broadcast mechanism (PAB) to broadcast the attention weights corresponding to different attention types of the target information processed in different broadcasting ranges, and reuses the attention weights in different time steps in the broadcasting range, so that the same attention weight is used in the step of broadcasting range, reduces the calculation amount, improves the model efficiency performance, and can effectively improve the model's perception of the position information of amino acids in the protein sequence, so that the model can fully consider the position relationship between amino acids when calculating the attention weight, can help the model better identify long-chain amino acid interactions, reduce structure prediction errors caused by missing position information, and improve the prediction accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0011] Figure 1 The principle diagram of the original diffusion model adopted in AlphaFold3 is difficult to maintain long-chain amino acid interaction;

[0012] Figure 2 The application environment diagram of the biological macromolecular sequence structure prediction method in one embodiment of the present application;

[0013] Figure 3 The logic diagram of broadcasting the attention weights corresponding to different attention types by introducing a geometric position encoding gating mechanism in the diffusion process in one embodiment of the present application;

[0014] Figure 4 The flowchart of the biological macromolecular sequence structure prediction method in one embodiment of the present application;

[0015] Figure 5 The structural block diagram of the biological macromolecular sequence structure prediction device in one embodiment of the present application;

[0016] Figure 6 The internal structure diagram of the computer device in one embodiment of the present application. DETAILED DESCRIPTION

[0017] With reference to the drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0018] It should be noted that in the description of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0019] In order to enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0020] Every plant, animal and human cell in nature contains hundreds of millions of molecular machines. These molecular machines are composed of proteins, DNA, RNA and other ligand molecules. It is these small machines composed of biological macromolecules that maintain the operation and continuation of life. In essence, life is built on the structural support at the molecular level and the interaction between molecules. Proteins are the main bearers of life activities, and their functions are closely related to their three-dimensional structures. Understanding the three-dimensional structure of proteins helps to reveal their action mechanisms in biological bodies, such as the catalytic function of enzymes and the immune recognition function of antibodies. Traditional protein structure determination methods, such as X-ray crystal diffraction and nuclear magnetic resonance spectroscopy, can obtain high-precision structure information, but have the disadvantages of long experimental period, high cost and complex operation, which are difficult to meet the needs of large-scale protein structure determination.

[0021] On May 9, 2024, researchers from DeepMind and Isomorphic Labs developed a new artificial intelligence (AI) software called AlphaFold3, which successfully predicted the structure and interaction of almost all biological macromolecules (proteins, DNA, RNA, ligands, etc.) with unprecedented accuracy.

[0022] This ability to precisely predict and analyze complex protein interactions with other molecules using computers can help bring new insights into disease pathways, genomics, therapeutic targets, protein engineering, and synthetic biology. More importantly, AlphaFold3 opens up exciting possibilities for drug development, which could revolutionize traditional drug development models.

[0023] The original diffusion model used in AlphaFold3 has a problem of decreasing prediction accuracy as the length of the amino acid increases in the iterative denoising process, as shown in the accompanying Figure 1 As shown in the accompanying

[0024] To address the problem of decreasing prediction accuracy of long-chain amino acids in the iterative denoising process of the AlphaFold3 model, the biological macromolecule sequence structure prediction method provided by the present application can be applied in the application environment as shown in the accompanying Figure 1 The terminal 102 and the server 104 communicate through a network. The terminal 102 can input biological macromolecule sequences, ligands, and covalent bond information into the server 104, and the server 104 processes and predicts the structure of the input biological macromolecule sequence, ligand, and covalent bond information and outputs it. The terminal 102 can be, but is not limited to, various personal computers, notebook computers, smartphones, tablet computers, and portable wearable devices, and the server 104 can be implemented with a standalone server or a server cluster composed of multiple servers.

[0025] As shown in the accompanying Figure 3As shown, the present application introduces a geometric position encoding gating mechanism in the diffusion process, uses a multi-scale attention broadcast mechanism (PAB) to adopt different broadcast range methods for attention weight calculation of different change frequencies, and the broadcast range refers to the reuse operation of attention weight in different time steps, which realizes the use of the same attention weight in the broadcast range step. This mechanism not only effectively improves the model's perception ability of the position information of amino acids in the protein sequence, thereby improving the accuracy of structure prediction, but also reduces the computational complexity and improves the efficiency of the model. The geometric position encoding gating mechanism encodes the geometric position information of amino acids and combines it with the attention mechanism, so that the model can fully consider the positional relationship between amino acids when calculating the attention weight. This mechanism can help the model better identify long-chain amino acid interactions and reduce structure prediction errors caused by missing position information.

[0026] Among them, the multi-scale attention broadcast mechanism (PAB) reduces redundant attention calculation by using the same attention weight in the broadcast range step, achieving acceleration. The attention difference of different time steps presents a U-shaped pattern, and the attention difference of the middle step is small; different types of attention have different stability and difference. Based on this, PAB proposes a pyramid attention broadcast, which sets different broadcast ranges according to the stability and difference of different attention, avoiding redundant attention calculation.

[0027] As shown in Figure 4 The embodiment of the present application provides a biological macromolecule sequence structure prediction method, which comprises the following steps:

[0028] Step S1, encode the geometric position information of amino acids in the biological macromolecule sequence, determine the dynamic geometric position encoding of the relative spatial relationship between a plurality of residue pairs, and generate the gated geometric position encoding by adjusting the weight of the dynamic geometric position encoding through the gating signal;

[0029] Step S2, combine the geometric position encoding with the attention mechanism to determine the attention weight for adjusting the position relationship between amino acids, and adjust the broadcast range of the attention weight according to the attention type of the target information processed;

[0030] Step S3, correct the residue position based on the geometric position encoding, the attention weight and the broadcast range of the attention weight, and update the geometric position encoding according to the corrected residue position;

[0031] Step S4, determine the predicted biological macromolecule sequence structure according to the updated geometric position encoding.

[0032] The application introduces a geometric position coding gating mechanism in the diffusion process, adopts different broadcasting range ways to broadcast the attention weights corresponding to different attention types of the processed target information by using a multi-scale attention broadcast mechanism (PAB), and the attention weights in the broadcasting range are reused in different time steps, so that the same attention weight is used in the broadcasting range step, the calculation amount is reduced, the model efficiency performance is improved, and the model's perception ability of the position information of amino acids in the protein sequence is effectively improved, so that the model can fully consider the position relationship between amino acids when calculating the attention weight, can help the model better identify long-chain amino acid interactions, reduce structure prediction errors caused by missing position information, and improve the prediction accuracy.

[0033] Moreover, the application can broadcast the attention weights corresponding to different attention types of the processed target information in different broadcasting range ways, accurately calculate the attention weights for long-chain amino acids by fully considering long-chain amino acid interactions, and effectively improve the calculation resources of the attention weights for long-chain amino acids, thereby improving the prediction accuracy and efficiency of the sequence structure of biological macromolecules as a whole.

[0034] The spatial geometric information and sequence interaction are integrated, and the structure is gradually refined by the diffusion model to realize accurate prediction of the sequence structure of biological macromolecules.

[0035] In this embodiment, the geometric position information of amino acids in the sequence of biological macromolecules is coded, the dynamic geometric position coding of the relative spatial relationship between a plurality of residue pairs is determined, the gated adjusted geometric position coding is generated by adjusting the weight of the dynamic geometric position coding through a gating signal, including:

[0036] The input biological macromolecule sequence, ligand and covalent bond information are received, and the initial residue feature representation of the biological macromolecule sequence is generated;

[0037] The starting three-dimensional coordinate point of the biological macromolecule sequence is generated as the starting point of the dynamic geometric position coding;

[0038] Based on the starting three-dimensional coordinate point, the three-dimensional coordinates of the residues are mapped to high-dimensional features through linear transformation, and the coding vector representing the relative spatial relationship between the residues is generated as the dynamic geometric position coding between a plurality of residue pairs by combining the noise disturbance process of the diffusion model;

[0039] One or more gating signals are generated, and the gating signal is used to adjust the weight of the dynamic geometric position coding;

[0040] The gating signal is applied to the dynamic geometric position coding to generate the gated adjusted geometric position coding.

[0041] wherein initial three-dimensional coordinates are generated for each residue by a conformer generation module. The generation of the gating signal aims to suppress noise interference in the diffusion denoising process and to strengthen the spatial relationship between key residues. The gating signal is used to adjust the weight of the dynamic geometric position encoding to selectively control its influence on subsequent feature updates or attention mechanisms; wherein the geometric position encoding includes but is not limited to: the discrete encoding of the relative distance between the center atoms (e.g. Cα atoms) of the residue pairs; the encoding of the relative orientation of the local coordinate system between the residue pairs, such as by rotation matrix or quaternion representation; and the encoding of the relative angle or torsion angle (e.g. dihedral angle) formed by key atoms (e.g. Cα, C, N atoms) between residue pairs.

[0042] wherein the initial input is received, the initial structure representation is generated, and the geometric relationship between the residues is dynamically calculated in combination with the diffusion noise, and the weight of the geometric information is adjusted by the gating mechanism to provide accurate and controlled geometric information for the subsequent attention mechanism.

[0043] In this embodiment, the determination of the attention weight for adjusting the positional relationship between the amino acids by combining the geometric position encoding with the attention mechanism includes:

[0044] The geometric position encoding is injected as a logarithmic bias term into the attention score calculation of the attention mechanism, and the attention score between the amino acid residues is adjusted as the attention weight for adjusting the positional relationship between the amino acids.

[0045] wherein by injecting the geometric encoding as a logarithmic bias term into the attention score calculation, the attention between the amino acid residues is directly adjusted, thereby strengthening the model's perception of local spatially proximal atom / residue interactions.

[0046] In this embodiment, the broadcast range of the attention weight is adjusted according to the attention type of the target information being processed, including:

[0047] determining the change frequency or influence change rate of the target information being processed in the iterative denoising process;

[0048] determining the attention type corresponding to the target information according to the change frequency or influence change rate of the target information, and setting the broadcast ratio of the attention type corresponding to the target information;

[0049] controlling the broadcast range of the attention weight corresponding to the target information according to the broadcast ratio.

[0050] wherein by evaluating the change frequency or influence change rate of the target information being processed in the iterative denoising process, the attention type and the broadcast ratio are adaptively determined, thereby optimizing the information propagation efficiency and ensuring that different types of information are reasonably focused in the model.

[0051] In the embodiment, the attention type corresponding to the target information is determined according to the change frequency or the influence change rate of the target information, and the broadcast ratio of the attention type corresponding to the target information is set, including:

[0052] In response to the target information being atomic local structure information, the instantaneous change of hydrogen bond or van der Waals force is captured, and the target information is classified as spatial attention feature information;

[0053] In response to the target information being sequence evolution information, the target information is classified as time attention feature information;

[0054] In response to the target information being ligand-protein binding mode in the global structure generation stage, the target information is classified as cross-modal attention feature information;

[0055] The broadcast ratio of the spatial attention feature information is set to be less than the broadcast ratio of the time attention feature information, and the broadcast ratio of the time attention feature information is set to be less than the broadcast ratio of the cross-modal attention feature information.

[0056] According to the nature of the processing information (atomic local structure, sequence evolution, and cross-modal binding), the attention is divided into three categories of space, time, and cross-modal, and an increasing broadcast ratio is set to match the demand of different information types for the propagation range.

[0057] In the embodiment, the broadcast range of the attention weight corresponding to the target information is controlled according to the broadcast ratio, including:

[0058] The total diffusion step number of the attention weight corresponding to the spatial attention feature information when performing single broadcast is set to be less than the total diffusion step number of the attention weight corresponding to the time attention feature information when performing single broadcast, and the total diffusion step number of the attention weight corresponding to the time attention feature information when performing single broadcast is set to be less than the total diffusion step number of the attention weight corresponding to the cross-modal attention feature information when performing single broadcast;

[0059] In response to determining whether the current time step is within the broadcast range at each time step within the total diffusion step number when performing single broadcast on the target attention weight;

[0060] In response to the current time step being within the broadcast range, the next time step is controlled to reuse the target attention weight of the previous time step;

[0061] In response to the current time step not being within the broadcast range, the attention weight of the next time step is recalculated according to the attention mechanism, and the attention weight of the next time step is saved for broadcast within the total diffusion step number of the next single broadcast.

[0062] Wherein, by setting the total diffusion step number of single broadcast of different attention types, and judging whether it is in the broadcast range at each time step, it is determined whether to reuse the existing weight or to recalculate, in order to balance the calculation efficiency and accuracy, and avoid unnecessary repeated calculation.

[0063] As shown in Figure 3 , the time steps of broadcasting and propagating the attention weights corresponding to the spatial attention feature information, the temporal attention feature information and the cross-modal attention feature information include:

[0064] Let the current time step be X t , the remaining time steps within the total diffusion step number of single broadcast are X t-1 , X t-2 , X t-3 , ……X t-n ; Wherein t-n=1, n is an integer;

[0065] According to the order of the amino acid sequence of the generated biological macromolecular sequence structure, the spatial attention feature information sequentially broadcasts the corresponding attention weights from the current time step X t , the temporal attention feature information sequentially broadcasts the corresponding attention weights from the current time step X t , and the cross-modal attention feature information respectively sequentially broadcasts the corresponding attention weights from the current time step X t .

[0066] Since the total diffusion step number of single broadcast of the attention weights corresponding to the spatial attention feature information is less than that of the temporal attention feature information, and the total diffusion step number of single broadcast of the attention weights corresponding to the temporal attention feature information is less than that of the cross-modal attention feature information, in Figure 3 , it is exemplarily shown that the attention weights corresponding to the spatial attention feature information are broadcast from the current time step X t to the time step X t-1 , the attention weights corresponding to the temporal attention feature information are broadcast from the current time step X t to the time step X t-2 , and the attention weights corresponding to the cross-modal attention feature information are broadcast from the current time step X t to the time step X t-3 .

[0067] In this embodiment, the broadcast range of the attention weights corresponding to the target information according to the broadcast ratio also includes:

[0068] In the process of broadcasting the attention weights corresponding to the target information, the calculation resource load rate is obtained;

[0069] The broadcast ratio of the spatial attention feature information, the time attention feature information, and the cross-modal attention feature information is adjusted according to the computing resource load rate to adjust the broadcast range.

[0070] Among them, when broadcasting complex target information, the broadcast frequency of the corresponding attention feature information is automatically increased; and when broadcasting simple target information, the broadcast ratio is reduced, thereby further improving the efficiency while ensuring the quality.

[0071] In the embodiment, when broadcasting the spatial attention feature information, the time attention feature information, and the cross-modal attention feature information, the method further comprises:

[0072] obtaining the spatial attention feature information, the time attention feature information, and the cross-modal attention feature information to be broadcast at the current time step;

[0073] setting the spatial attention feature information, the time attention feature information, and the cross-modal attention feature information to be broadcast in a group broadcast mode or a hierarchical broadcast mode.

[0074] Among them, in addition to a single broadcast mechanism, different types of broadcast mechanisms, such as a group broadcast mode and a hierarchical broadcast mode, are tried to tap greater acceleration potential.

[0075] In the embodiment, in response to determining whether the current time step is within the broadcast range at each time step within the total diffusion step number when broadcasting the target attention weight a single time, the method comprises:

[0076] At the current time step, obtaining the current diffusion step of the target attention weight, obtaining the total diffusion step number of the target attention weight to be broadcast and propagated, and determining the real-time diffusion step position according to the current diffusion step and the total diffusion step number.

[0077] calculating a U-shaped stability curve according to the real-time diffusion step position;

[0078] If the next diffusion step is within the total diffusion step number, the U-shaped stability curve is greater than a first threshold value, and the attention stability is greater than an attention stability threshold value, it is determined that the current time step is within the broadcast range, otherwise it is determined that the current time step is not within the broadcast range.

[0079] Among them, the method of calculating the U-shaped stability curve based on the real-time diffusion step position is introduced, and the comprehensive judgment is combined with the attention stability threshold value to more intelligently and robustly determine whether to reuse the attention weight, so as to ensure broadcasting in the stage where the attention weight is stable.

[0080] In the embodiment, determining the real-time diffusion step position according to the current diffusion step and the total diffusion step number comprises:

[0081] obtaining the current diffusion step and the total diffusion steps of the single round;

[0082] determining the real-time diffusion step position according to the quotient of the current diffusion step and the total diffusion steps of the single round.

[0083] Specifically, the real-time diffusion step position is calculated by normalized_step = step / total_steps, where normalized_step is the real-time diffusion step position, step is the current diffusion step, and total_steps is the total diffusion steps of the single round.

[0084] where normalized_current_diffusion_progress is the normalized current diffusion progress, which provides a unified input for the subsequent calculation of the U-shaped stability curve.

[0085] In this embodiment, the calculation of the U-shaped stability curve according to the real-time diffusion step position comprises:

[0086] obtaining the absolute value of the difference between the real-time diffusion step position and a first threshold value;

[0087] calculating the product of the absolute value and a second threshold value;

[0088] determining the U-shaped stability curve according to the difference between a third threshold value and the product.

[0089] Specifically, the U-shaped stability curve is calculated by stability_curve = 1.0-2.0 * abs(normalized_step-0.5), where stability_curve is the U-shaped stability curve, and abs() is the absolute value function.

[0090] where the first threshold value is preferably 0.5, the second threshold value is preferably 2.0, and the third threshold value is preferably 1.0, which quantitatively evaluates the stability of the attention weight at the current diffusion step and serves as a key basis for determining whether to perform broadcasting.

[0091] In this embodiment, the control of the broadcast range of the attention weight corresponding to the target information according to the broadcast ratio comprises:

[0092] obtaining the total diffusion steps of the single round and the broadcast ratio;

[0093] obtaining a first product according to the total diffusion steps of the single round and the broadcast ratio;

[0094] rounding the first product to determine the broadcast range of the attention weight corresponding to the target information.

[0095] Specifically, the broadcast range of the attention weight corresponding to the target information is calculated by broadcast_range = int(total_steps*broadcast_ratio), where broadcast_range is the broadcast range, int() is the integer function, total_steps is the total number of single diffusion steps, and broadcast_ratio is the broadcast ratio.

[0096] Wherein, based on the total number of diffusion steps and the broadcast ratio, the effective propagation distance of the attention weight in the diffusion process is quantitatively determined.

[0097] In this embodiment, setting the broadcast ratio of the target information corresponding to the attention type further includes:

[0098] Obtain the target attention type corresponding to the target information, and obtain the confidence score of broadcasting the attention weight corresponding to the target information in the target attention type;

[0099] In response to the confidence score being less than the confidence preset score threshold, the broadcast ratio of the target information corresponding to the target attention type is increased.

[0100] Wherein, when the broadcast confidence score of the target attention weight is lower than the confidence preset score threshold (preferably 0.7), the broadcast ratio is increased to increase the information update frequency, so as to prompt the model to reevaluate and update more frequently when the model is uncertain about the current attention weight.

[0101] In this embodiment, based on the geometric position encoding, the attention weight and the broadcast range of the attention weight, the residue position is corrected, and the geometric position encoding is updated according to the corrected residue position, which includes:

[0102] In the iterative denoising process, the initial three-dimensional coordinates of the atoms corresponding to the amino acid sequence in the geometric position encoding are obtained;

[0103] In each iteration, the initial three-dimensional coordinates and the amino acid sequence pairing information are input into the diffusion model, and the diffusion model combines the geometric position encoding and the attention weight to identify and predict the noise vector;

[0104] Subtract the noise vector from the initial three-dimensional coordinates to correct the residue position, and update the geometric position encoding based on the corrected three-dimensional coordinates.

[0105] In each iteration, new noise is introduced to explore various biological macromolecular sequence structures until the target three-dimensional coordinates of the corresponding amino acid sequence atoms are obtained. By continuously removing existing noise and introducing new noise through the iterative denoising process, the diversity of biological macromolecular sequence structures can be increased, the best structure can be found, the target three-dimensional coordinates of the corresponding amino acid sequence can be obtained, and the accuracy of structure prediction can be improved.

[0106] In this embodiment, determining the predicted biological macromolecular sequence structure according to the updated geometric position code comprises:

[0107] Using the distance constraint network, the target three-dimensional coordinates are optimized based on the predicted distance matrix to conform to the molecular force field rules, and the optimized target three-dimensional coordinates are obtained.

[0108] According to the optimized target three-dimensional coordinates, a predicted biological macromolecular sequence structure is generated.

[0109] In this embodiment, by introducing the distance constraint network, the target three-dimensional coordinates are optimized based on the predicted distance matrix to ensure that the predicted structure conforms to the molecular mechanics rules and physical constraints, thereby improving the quality and biological rationality of the predicted structure.

[0110] In the above biological macromolecular sequence structure prediction method, the present application introduces a geometric position coding gating mechanism in the diffusion process, and uses a multi-scale attention broadcast mechanism (PAB) to broadcast the attention weights corresponding to different attention types of the target information processed in different broadcasting ranges. The attention weights in different time steps in the broadcasting range are reused, the same attention weight is used within the broadcasting range step, the calculation amount is reduced, the model efficiency performance is improved, and the model's perception ability of the amino acid position information in the protein sequence is effectively improved. When calculating the attention weight, the model can fully consider the position relationship between amino acids, can help the model better identify long-chain amino acid interactions, reduce structure prediction errors caused by missing position information, and improve prediction accuracy.

[0111] The application implements a two-level architecture for the diffusion module of the AlphaFold3 model: first processing atomic coordinates, then processing Token representation, and finally regressing atomic coordinates. The PAB can establish an attention bridge across levels during this process. For example, during atomic-level processing, local attention is used to capture chemical details such as bond lengths and bond angles; during Token-level processing, global attention is used to integrate sequence evolution information, and stable attention results (such as interaction patterns of conserved domains) are broadcast to the atomic level to guide local structure optimization. This cross-level interaction can enhance the synergy of features at different abstraction levels and improve the coherence of structure prediction. Experiments show that the AlphaFold3 model with PAB has a 5-10 point increase in predicted residue pLDDT scores.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software and the necessary general hardware platform, of course, it can also be implemented by hardware, but in many cases the former is a better implementation.

[0113] In one embodiment, as shown in Figure 5 A biological macromolecule sequence structure prediction device 10 is provided, comprising a geometric position encoding module 1, a broadcast range management module 2, a correction module 3, and a biological macromolecule generation module 4.

[0114] The geometric position encoding module 1 is used to encode the geometric position information of amino acids in the biological macromolecule sequence, determine the dynamic geometric position encoding of the relative spatial relationship between multiple residue pairs, and generate the gated adjusted geometric position encoding by adjusting the weight of the dynamic geometric position encoding through the gating signal.

[0115] The broadcast range management module 2 is used to determine the attention weight for adjusting the position relationship between amino acids by combining the geometric position encoding with the attention mechanism, and adjust the broadcast range of the attention weight according to the attention type of the target information processed.

[0116] The correction module 3 is used to correct the residue position based on the geometric position encoding, the attention weight, and the broadcast range of the attention weight, and update the geometric position encoding according to the corrected residue position.

[0117] The biological macromolecule generation module 4 is used to determine the predicted biological macromolecule sequence structure according to the updated geometric position encoding.

[0118] In this embodiment, the geometric position information of amino acids in the biological macromolecule sequence is encoded, the dynamic geometric position encoding of the relative spatial relationship between multiple residue pairs is determined, and the gated adjusted geometric position encoding is generated by adjusting the weight of the dynamic geometric position encoding through the gating signal, comprising:

[0119] receiving inputted biological macromolecule sequence, ligand and covalent bond information, generating initial residue feature representation of biological macromolecule sequence;

[0120] generating starting three-dimensional coordinate points of biological macromolecule sequence as starting point of dynamic geometric position coding;

[0121] based on the starting three-dimensional coordinate points, mapping the three-dimensional coordinates of the residues into high-dimensional features through linear transformation, and generating coding vectors representing the relative spatial relationship between the residue pairs as dynamic geometric position coding between multiple residue pairs by combining the noise disturbance process of the diffusion model;

[0122] generating one or more gating signals, the gating signals being used to adjust the weight of the dynamic geometric position coding;

[0123] applying the gating signal to the dynamic geometric position coding to generate a gated and adjusted geometric position coding.

[0124] In this embodiment, the geometric position coding is combined with the attention mechanism to determine the attention weight for adjusting the position relationship between amino acids, which includes:

[0125] injecting the geometric position coding into the attention score calculation of the attention mechanism as a logarithmic bias term, and adjusting the attention score between the amino acid residue pairs as the attention weight for adjusting the position relationship between the amino acids.

[0126] In this embodiment, the broadcast range of the attention weight is adjusted according to the attention type of the target information being processed, which includes:

[0127] determining the change frequency or influence change rate of the target information being processed in the iterative denoising process;

[0128] determining the attention type corresponding to the target information according to the change frequency or influence change rate of the target information, and setting the broadcast ratio of the attention type corresponding to the target information;

[0129] controlling the broadcast range of the attention weight corresponding to the target information according to the broadcast ratio.

[0130] In this embodiment, the broadcast ratio of the attention type corresponding to the target information is set according to the change frequency or influence change rate of the target information, which includes:

[0131] in response to the target information being atomic-level local structure information being processed, capturing the instantaneous change of hydrogen bond or van der Waals force, the target information is classified as spatial attention feature information;

[0132] in response to the target information being sequence evolution information being processed, the target information is classified as time attention feature information;

[0133] in response to the target information being used to generate a ligand-protein binding mode of a processing global structure, the target information is classified as cross-modal attention feature information;

[0134] a broadcast ratio of the spatial attention feature information is set to be less than a broadcast ratio of the temporal attention feature information, and the broadcast ratio of the temporal attention feature information is set to be less than a broadcast ratio of the cross-modal attention feature information.

[0135] In the embodiment, the broadcast range of the attention weight corresponding to the target information is controlled according to the broadcast ratio, including:

[0136] a total diffusion step number when the attention weight corresponding to the spatial attention feature information is broadcasted once is set to be less than a total diffusion step number when the attention weight corresponding to the temporal attention feature information is broadcasted once, and the total diffusion step number when the attention weight corresponding to the temporal attention feature information is broadcasted once is set to be less than a total diffusion step number when the attention weight corresponding to the cross-modal attention feature information is broadcasted once;

[0137] in response to determining whether the current time step is within the broadcast range at each time step within the total diffusion step number when the target attention weight is broadcasted once;

[0138] in response to the current time step being within the broadcast range, the target attention weight of the next time step is controlled to reuse the target attention weight of the previous time step;

[0139] in response to the current time step not being within the broadcast range, the attention weight of the next time step is recalculated according to the attention mechanism, and the attention weight of the next time step is saved for broadcasting within the total diffusion step number of the next single broadcast.

[0140] In the embodiment, in response to determining whether the current time step is within the broadcast range at each time step within the total diffusion step number when the target attention weight is broadcasted once, including:

[0141] at the current time step, a current diffusion step of the target attention weight is obtained, a total diffusion step number of a single broadcast propagation of the target attention weight is obtained, and a real-time diffusion step position is determined according to the current diffusion step and the total diffusion step number of the single broadcast;

[0142] a U-shaped stability curve is calculated according to the real-time diffusion step position;

[0143] if the next diffusion step is within the total diffusion step number of the single broadcast, the U-shaped stability curve is greater than a first threshold value, and the attention stability is greater than an attention stability threshold value, it is determined that the current time step is within the broadcast range, otherwise it is determined that the current time step is not within the broadcast range.

[0144] In the embodiment, determining the real-time diffusion step position according to the current diffusion step and the total number of single diffusion steps comprises:

[0145] obtaining the current diffusion step and the total number of single diffusion steps;

[0146] determining the real-time diffusion step position according to the quotient of the current diffusion step and the total number of single diffusion steps.

[0147] In the embodiment, calculating the U-shaped stability curve according to the real-time diffusion step position comprises:

[0148] obtaining the absolute value of the difference between the real-time diffusion step position and the first threshold value;

[0149] calculating the product of the absolute value and the second threshold value;

[0150] determining the U-shaped stability curve according to the difference between the third threshold value and the product.

[0151] In the embodiment, controlling the broadcast range of the attention weight corresponding to the target information according to the broadcast ratio comprises:

[0152] obtaining the total number of single diffusion steps and the broadcast ratio;

[0153] obtaining a first product according to the total number of single diffusion steps and the broadcast ratio;

[0154] rounding the first product to determine the broadcast range of the attention weight corresponding to the target information.

[0155] In the embodiment, setting the broadcast ratio of the attention type corresponding to the target information further comprises:

[0156] obtaining the target attention type corresponding to the target information, and obtaining the confidence score of broadcasting the attention weight corresponding to the target information in the target attention type;

[0157] in response to the confidence score being less than the preset confidence score threshold, increasing the broadcast ratio of the target attention type corresponding to the target information.

[0158] In the embodiment, correcting the residue position based on the geometric position encoding, the attention weight, and the broadcast range of the attention weight, and updating the geometric position encoding according to the corrected residue position comprises:

[0159] In the iterative denoising process, obtaining the initial three-dimensional coordinates of the atoms corresponding to the amino acid sequence in the geometric position encoding;

[0160] In each iteration, input the initial three-dimensional coordinates and the amino acid sequence pairing information into the diffusion model, and the diffusion model combines the geometric position encoding and the attention weight to identify and predict the noise vector;

[0161] Subtract the noise vector from the initial three-dimensional coordinates to correct the residue position, and update the geometric position code based on the corrected three-dimensional coordinates.

[0162] In this embodiment, determining the predicted biological macromolecular sequence structure according to the updated geometric position code comprises:

[0163] Using the distance constraint network, the target three-dimensional coordinates are optimized based on the predicted distance matrix to conform to the molecular force field rules, and the optimized target three-dimensional coordinates are obtained.

[0164] According to the optimized target three-dimensional coordinates, a predicted biological macromolecular sequence structure is generated.

[0165] In the above biological macromolecular sequence structure prediction device, the present application introduces a geometric position code gating mechanism in the diffusion process, uses a multi-scale attention broadcast mechanism (PAB) to broadcast the attention weights corresponding to different attention types of the target information processed in different broadcasting ranges, and reuses the attention weights in different time steps in the broadcasting range, so that the same attention weight is used in the step of broadcasting range, reducing the calculation amount and improving the model efficiency performance. It can also effectively improve the model's perception of amino acid position information in the protein sequence, so that the model can fully consider the position relationship between amino acids when calculating the attention weight, help the model better identify long-chain amino acid interactions, reduce structure prediction errors caused by missing position information, and improve the prediction accuracy.

[0166] The features of the embodiments corresponding to the biological macromolecular sequence structure prediction device can be referred to the related descriptions of the embodiments corresponding to the biological macromolecular sequence structure prediction method, which will not be repeated here.

[0167] Embodiments of the present application also provide an electronic device comprising a memory and a processor, the memory storing a computer program, and the processor being configured to run the computer program to perform the steps in any of the above biological macromolecular sequence structure prediction method embodiments.

[0168] In one embodiment, the electronic device can be a server, and its internal structure diagram can be as shown in Figure 6As shown. The electronic device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the electronic device is used to store biological macromolecular sequence structure prediction data. The network interface of the electronic device is used to communicate with the external terminal through the network connection. The computer program is executed by the processor to implement a biological macromolecular sequence structure prediction method.

[0169] Embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above biological macromolecular sequence structure prediction method embodiments when running:

[0170] The geometric position information of the amino acids in the biological macromolecular sequence is encoded, the dynamic geometric position encoding of the relative spatial relationship between the plurality of residue pairs is determined, the weight of the dynamic geometric position encoding is adjusted through the gating signal to generate the geometric position encoding adjusted by the gating, and the geometric position encoding adjusted by the gating is obtained.

[0171] The attention weight adjusting the position relationship between the amino acids is determined by combining the geometric position encoding with the attention mechanism, and the broadcast range of the attention weight is adjusted according to the attention type of the target information processed.

[0172] The residue position is corrected based on the geometric position encoding, the attention weight and the broadcast range of the attention weight, and the geometric position encoding is updated according to the corrected residue position.

[0173] The predicted biological macromolecular sequence structure is determined according to the updated geometric position encoding.

[0174] In an exemplary embodiment, the above computer readable storage medium can include but is not limited to: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0175] Embodiments of the present application also provide a computer program product, the above computer program product includes a computer program, and the computer program is executed by the processor to implement the steps in any of the above biological macromolecular sequence structure prediction method embodiments:

[0176] The geometric position information of the amino acids in the biological macromolecule sequence is encoded, dynamic geometric position encoding of relative spatial relationships between a plurality of residue pairs is determined, and the weight of the dynamic geometric position encoding is adjusted through a gating signal to generate gated adjusted geometric position encoding;

[0177] The geometric position encoding is combined with an attention mechanism to determine attention weights that adjust the positional relationships between the amino acids, and the broadcast range of the attention weights is adjusted according to the attention type of the target information being processed;

[0178] The residue positions are corrected based on the geometric position encoding, the attention weights, and the broadcast range of the attention weights, and the geometric position encoding is updated according to the corrected residue positions;

[0179] The predicted biological macromolecule sequence structure is determined according to the updated geometric position encoding.

[0180] Embodiments of the present application also provide another computer program product, comprising a non-volatile computer readable storage medium, the non-volatile computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps in any of the above biological macromolecule sequence structure prediction method embodiments:

[0181] The geometric position information of the amino acids in the biological macromolecule sequence is encoded, dynamic geometric position encoding of relative spatial relationships between a plurality of residue pairs is determined, and the weight of the dynamic geometric position encoding is adjusted through a gating signal to generate gated adjusted geometric position encoding;

[0182] The geometric position encoding is combined with an attention mechanism to determine attention weights that adjust the positional relationships between the amino acids, and the broadcast range of the attention weights is adjusted according to the attention type of the target information being processed;

[0183] The residue positions are corrected based on the geometric position encoding, the attention weights, and the broadcast range of the attention weights, and the geometric position encoding is updated according to the corrected residue positions;

[0184] The predicted biological macromolecule sequence structure is determined according to the updated geometric position encoding.

[0185] Those skilled in the art will further appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in general terms above as being generally described in the above description. Whether such functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0186] The above describes in detail a biological macromolecule sequence structure prediction method and an electronic device provided by the present application. The principles and implementation manners of the present application are described by using specific examples, and the above description of the examples is only applicable to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A method for predicting the sequence structure of biological macromolecules, characterized in that, include: The geometric position information of amino acids in the sequence of biological macromolecules is encoded to determine the dynamic geometric position encoding of the relative spatial relationship between multiple residue pairs. The weight of the dynamic geometric position encoding is adjusted by a gating signal to generate a gated geometric position encoding. The geometric position encoding is combined with the attention mechanism to determine the attention weights that adjust the positional relationships between amino acids, and the broadcast range of the attention weights is adjusted according to the attention type of the target information being processed. The residue positions are corrected based on the geometric position encoding, the attention weight, and the broadcast range of the attention weight, and the geometric position encoding is updated according to the corrected residue positions; The predicted biomolecular sequence structure was determined based on the updated geometric position coding; The process of encoding the geometric position information of amino acids in a biomacromolecule sequence, determining the dynamic geometric position encoding of the relative spatial relationships between multiple residue pairs, and generating a gated geometric position encoding by adjusting the weights of the dynamic geometric position encoding through a gating signal includes: Receive input biomacromolecule sequences, ligands, and covalent bond information, and generate an initial residue feature representation of the biomacromolecule sequence; The starting three-dimensional coordinates of the biomacromolecule sequence are generated as the starting point of the dynamic geometric position encoding; Based on the initial three-dimensional coordinate point, the three-dimensional coordinates of the residues are mapped to high-dimensional features through linear transformation. Combined with the noise perturbation process of the diffusion model, an encoding vector characterizing the relative spatial relationship between residue pairs is generated as the dynamic geometric position encoding between multiple residue pairs. Generate one or more gating signals, which are used to adjust the weights of the dynamic geometric position encoding; The gating signal is applied to the dynamic geometric position code to generate a gating-adjusted geometric position code; The step of combining the geometric position encoding with the attention mechanism to determine the attention weights for adjusting the positional relationships between amino acids includes: The geometric position encoding is injected as a logarithmic bias term into the attention score calculation of the attention mechanism, and the attention score between amino acid residue pairs is used as the attention weight to adjust the positional relationship between amino acids.

2. The method for predicting the sequence structure of biological macromolecules according to claim 1, characterized in that, The broadcast range for adjusting the attention weight based on the attention type of the processed target information includes: Determine the frequency of change or rate of change of the target information being processed during the iterative denoising process; Determine the attention type corresponding to the target information based on the frequency of change or the rate of change of influence of the target information, and set the broadcast ratio of the attention type corresponding to the target information. The broadcast range of the attention weight corresponding to the target information is controlled according to the broadcast ratio.

3. The method for predicting the sequence structure of biological macromolecules according to claim 2, characterized in that, The step of determining the attention type corresponding to the target information based on the change frequency or influence change rate of the target information, and setting the broadcast ratio of the attention type corresponding to the target information, includes: When the target information is to process atomic-level local structural information and capture instantaneous changes in hydrogen bonds or van der Waals forces, the target information is classified as spatial attention feature information. When the target information is processing sequence evolution information, the target information is classified as temporal attention feature information; When the target information is a ligand-protein binding mode in the global structure generation stage, the target information is classified as cross-modal attention feature information; The broadcast ratio of the spatial attention feature information is set to be less than the broadcast ratio of the temporal attention feature information, and the broadcast ratio of the temporal attention feature information is less than the broadcast ratio of the cross-modal attention feature information.

4. The method for predicting the sequence structure of biological macromolecules according to claim 3, characterized in that, The broadcast range for controlling the attention weight corresponding to the target information based on the broadcast ratio includes: The total number of diffusion steps when the attention weight corresponding to the spatial attention feature information is broadcast once is less than the total number of diffusion steps when the attention weight corresponding to the temporal attention feature information is broadcast once, and the total number of diffusion steps when the attention weight corresponding to the temporal attention feature information is broadcast once is less than the total number of diffusion steps when the attention weight corresponding to the cross-modal attention feature information is broadcast once. In response to a single broadcast of the target attention weights, determine at each time step within the total number of diffusion steps whether the current time step is within the broadcast range; In response to the current time step being within the broadcast range, control the reuse of the target attention weight of the previous time step in the next time step; In response to the current time step not being within the broadcast range, the attention weight for the next time step is recalculated according to the attention mechanism, and the attention weight for the next time step is saved for broadcasting within the total number of diffusion steps of the next single broadcast.

5. The method for predicting the sequence structure of biological macromolecules according to claim 4, characterized in that, The response to determining whether the current time step is within the broadcast range at each time step within the total number of diffusion steps during a single broadcast of the target attention weights includes: At the current time step, obtain the current diffusion step of the target attention weight, obtain the total number of single diffusion steps for the target attention weight to be broadcast, and determine the real-time diffusion step position based on the current diffusion step and the total number of single diffusion steps. Calculate the U-shaped stability curve based on the location of the real-time diffusion step; If the following conditions are met: the next diffusion step is within the total number of diffusion steps in a single run, the U-shaped stability curve is greater than the first threshold, and the attention stability is greater than the attention stability threshold, then the current time step is determined to be within the broadcast range; otherwise, the current time step is determined not to be within the broadcast range.

6. The method for predicting the sequence structure of biological macromolecules according to claim 5, characterized in that, Determining the real-time diffusion step location based on the current diffusion step and the total number of single diffusion steps includes: Obtain the current diffusion step and the total number of diffusion steps in a single run; The real-time diffusion step position is determined based on the quotient of the current diffusion step and the total number of diffusion steps in a single run; The calculation of the U-shaped stability curve based on the location of the real-time diffusion step includes: Obtain the absolute value of the difference between the location of the real-time diffusion step and the first threshold; Calculate the product of the absolute value and the second threshold; The U-shaped stability curve is determined based on the difference between the third threshold and the product. The step of controlling the broadcast range of the attention weight corresponding to the target information according to the broadcast ratio includes: Obtain the total number of diffusion steps in a single run and the broadcast ratio; The first product is obtained based on the total number of diffusion steps in a single run and the broadcast ratio; The first product is rounded down to determine the broadcast range of the attention weight corresponding to the target information.

7. The method for predicting the sequence structure of biological macromolecules according to claim 2, characterized in that, The setting of the broadcast ratio for the attention type corresponding to the target information also includes: Obtain the target attention type corresponding to the target information, and obtain the confidence score for broadcasting the attention weight corresponding to the target information when the target attention type is specified. In response to the confidence score being less than a preset confidence score threshold, the broadcast ratio of the target attention type corresponding to the target information is increased.

8. The method for predicting the sequence structure of biological macromolecules according to claim 1, characterized in that, The step of correcting the residue position based on the geometric position encoding, the attention weight, and the broadcast range of the attention weight, and updating the geometric position encoding according to the corrected residue position, includes: During the iterative denoising process, the initial three-dimensional coordinates of the corresponding amino acid sequence atoms in the geometric position encoding are obtained; In each iteration, the initial three-dimensional coordinates and amino acid sequence pairing information are input into the diffusion model, which keeps the broadcast range of the attention weights unchanged. The model combines the geometric position encoding and the attention weights to identify and predict noise vectors. The noise vector is subtracted from the initial three-dimensional coordinates to correct the residue position, and the geometric position code is updated based on the corrected three-dimensional coordinates.

Citation Information

Patent Citations

  • Chemical molecular structure identification method and device based on multi-stage sequence, and medium

    CN118097665A

  • Protein residue motion coordinate prediction method, equipment, medium and program product

    CN120656557A