Prediction control program, information processing apparatus, and prediction control method

The predictive control program adjusts intermediate features of structural prediction models to align with actual density maps, enabling the prediction of diverse three-dimensional structures, addressing the limitations of existing models that output only one typical structure.

JP2025127934APending Publication Date: 2025-09-02FUJITSU LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024024940
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-21
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Existing machine learning models for predicting protein structures, such as AlphaFold2, are limited to outputting only one typical atomic structure and cannot predict diverse three-dimensional structures.

Method used

A predictive control program and method that adjust the intermediate features of a structural prediction model to minimize the difference between the predicted structure and actual measured density maps, allowing for the prediction of various three-dimensional structures.

Benefits of technology

Enables the prediction of a variety of three-dimensional structures by refining the model's intermediate features to align with actual density maps, improving the accuracy and diversity of predicted structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025127934000001_ABST
    Figure 2025127934000001_ABST
Patent Text Reader

Abstract

To provide a prediction control program, an information processing apparatus, and a prediction control method for predicting atomic structure models of various three-dimensional structures.SOLUTION: There is provided a prediction control program for a structure prediction model 11 that predicts a three-dimensional structure of an organic compound from sequence information of the organic compound, the prediction control program causing a computer to execute a process of changing an intermediate feature value of the structure prediction model 11 so that a difference between a three-dimensional density map m1 corresponding to a predicted structure output as a prediction result from the structure prediction model 11 and a three-dimensional density map m0 different from the three-dimensional density map m1 and actually measured is decreased.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a predictive control program, an information processing device, and a predictive control method. [Background technology]

[0002] Techniques for predicting atomic structures from amino acid sequences have been disclosed (see Non-Patent Documents 1 and 2). For example, machine learning models such as AlphaFold2, OpenFold, and RoseTTAFold output one typical three-dimensional structure from among the atomic structures of proteins (all-atom structure model) for an input amino acid sequence.

[0003] Figure 8 is a reference diagram showing the AlphaFold2 machine learning model. As shown in Figure 8, when an InputSequence indicating an amino acid sequence is input, the AlphaFold2 machine learning model outputs one 3D structure as a typical protein structure.

[0004] In drug discovery, prediction of protein atomic structures is an important elemental technology, and for application to drug discovery, prediction of diverse 3D structures other than typical 3D structures is required. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2000-229994 [Non-patent literature]

[0006] [Non-Patent Document 1] Highly accurate protein structure prediction with AlphaFold Bibliographic information: https: / / www.nature.com / articles / s41586-021-03819-2 [Non-patent document 2] Yosuke Oyama, Akihiro Tabuchi, Atsushi Tokuhisa, "Accelerating AlphaFold2 Inference of Protein Three-Dimensional Structure on the Supercomputer Fugaku", FlexScience '23: Proceedings of the 13th Workshop on AI and Scientific Computing at Scale using Flexible ComputingAugust 11 ,2023, Pages 1-9, (https: / / doi.org / 10.1145 / 3589013.3596674) Summary of the Invention [Problem to be solved by the invention]

[0007] However, the machine learning models shown in the prior art only output one typical atomic structure, and are therefore unable to predict diverse three-dimensional structures.

[0008] In one aspect, the present invention aims to provide a predictive control program, an information processing device, and a predictive control method that are capable of predicting a variety of three-dimensional structures. [Means for solving the problem]

[0009] A prediction control program according to one aspect is a prediction control program for a structural prediction model that predicts the three-dimensional structure of an organic compound from sequence information of the organic compound, and causes a computer to execute a process of changing intermediate features of the structural prediction model so as to reduce the difference between first limiting information corresponding to the predicted structure output as a prediction result from the structural prediction model and second limiting information that differs from the first limiting information. [Effects of the Invention]

[0010] A variety of three-dimensional structures can be predicted. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram illustrating an image of predictive control according to the first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a functional configuration of the information processing apparatus according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing the spread or existence probability of atoms. [Figure 4] FIG. 4 is a diagram showing an example of the display of the output all-atom structure model. [Figure 5] FIG. 5 is a diagram showing an example of display of coordinates of the output all-atom structure model. [Figure 6] FIG. 6 is a diagram illustrating an example of a flowchart of the predictive control process according to the first embodiment. [Figure 7] FIG. 7 is a diagram illustrating an example of a hardware configuration. [Figure 8] Figure 8 is a reference diagram showing the AlphaFold2 machine learning model. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the predictive control program, information processing device, and predictive control method according to the present application will be described with reference to the accompanying drawings. Each embodiment merely illustrates one example or aspect, and does not limit the range of values, functions, or usage scenarios. Furthermore, each embodiment can be appropriately combined within the scope of not causing contradictions in the processing content. [Example]

[0013] (Image of predictive control) First, an image of predictive control according to Example 1 will be described with reference to FIG. 1. FIG. 1 is a diagram illustrating an image of predictive control according to the example. The predictive control process shown in FIG. 1 predicts an atomic structure by changing intermediate feature quantities of a structure prediction model so as to conform to information that limits the target atomic structure. In this example, the atomic structure to be predicted may be referred to as an all-atom structure model. Furthermore, hereinafter, the information that limits the all-atom structure model may be referred to as limiting information.

[0014] Hereinafter, a protein will be taken as an example of the target, but the target is not necessarily limited to a protein. For example, the target may be an organic compound other than a protein, such as a polymer material. Furthermore, a three-dimensional density map that limits the all-atom structural model of a protein will be taken as an example of information that limits the all-atom structural model of the target (limiting information), but the present invention is not limited to this.

[0015] 1, the prediction control process has a structure prediction model 11. The structure prediction model 11 may be, for example, a machine learning model such as AlphaFold2, OpenFold, or RoseTTAFold. In the embodiment, AlphaFold2 is applied to the structure prediction model 11, but the structure prediction model 11 applied in the embodiment does not need to be limited to AlphaFold2.

[0016] The prediction control process inputs sequence information of an organic compound to the structure prediction model 11. For example, the prediction control process inputs amino acid sequence information as an example of the sequence information of an organic compound to the structure prediction model 11 (S1). Then, the structure prediction model 11 extracts intermediate features from the amino acid sequence information, inputs the extracted intermediate features to a layer, and outputs an all-atom structure model (three-dimensional structure) of the protein as a prediction result (S2). Note that, for convenience of explanation, the layer of the structure prediction model 11 is described as a single layer, but multiple layers may also be used.

[0017] The prediction control process converts the all-atom structure model output as a prediction result into a three-dimensional density map that defines the all-atom structure model in a differentiable form (S3).

[0018] Then, the prediction control process calculates the difference between the 3D density map corresponding to the predicted all-atom structural model and the actually measured 3D density map (S4). The actually measured 3D density map is, for example, a 3D density map reconstructed from an EM (Electron Microscope) image taken by an electron microscope such as a cryo-electron microscope, and is a 3D density map that defines the all-atom structural model of the protein, but is not limited to this, and may be any 3D voxel data shape.

[0019] The predictive control process then performs backpropagation of the difference to calculate intermediate features that minimize the difference (S5). The predictive control process then updates the intermediate features of the structure prediction model 11 to the calculated intermediate features (S6). That is, the predictive control process changes the intermediate features of the structure prediction model 11 so as to reduce the difference between the 3D density map corresponding to the all-atom structure model of the predicted result and the actually measured 3D density map. In other words, the predictive control process changes the intermediate features of the structure prediction model 11 according to constraints imposed by the actually measured 3D density map.

[0020] Thereafter, the structure prediction model 11 predicts an all-atom structure model of the protein using the updated intermediate features. The prediction control process repeats steps S2 to S6, making it possible to predict a variety of three-dimensional structures.

[0021] If the structure prediction model 11 has multiple layers, the prediction control process only needs to update the intermediate features input to any of the multiple layers. The prediction control process updates the intermediate features of the structure prediction model 11, but does not update the learning parameters of the structure prediction model 11. This is to avoid forgetting knowledge held by the structure prediction model 11 by not destroying the existing learning parameters due to re-learning of the structure prediction model 11.

[0022] (Functional configuration of information processing device) 2 is a diagram illustrating an example of a functional configuration of an information processing device according to Example 1. The information processing device 1 illustrated in FIG. 2 is an example of a computer that executes a predictive control process. As illustrated in FIG. 2, the information processing device 1 includes a control unit 10 and a storage unit 20.

[0023] The storage unit 20 has a Protein Data Bank (PDB) file 21, a three-dimensional electron microscope data bank (EMDB) data 22, and EM data 23.

[0024] The PDB file 21 is a file that stores information on the three-dimensional structure of a protein. The PDB file 21 includes, for example, information on the all-atom structural model of the three monomers that make up the protein, as well as information on which chain each atom belongs to. The PDB file 21 can be obtained from the PDB on the Web.

[0025] The EMDB data 22 is a file that stores, as voxel data, a 3D density map of proteins reconstructed from 2D EM images captured by an electron microscope such as a cryo-electron microscope. The EMDB data 22 may be obtained, for example, from the EMDB on the Web.

[0026] The EM data 23 is a three-dimensional density map calculated from the all-atom structure model predicted by the predictive control process. The EM data 23 is stored in the storage unit 20 by the output unit 16, which will be described later.

[0027] The control unit 10 includes a plurality of structure prediction models 11 , a preprocessing unit 12 , a conversion unit 13 , a difference calculation unit 14 , an update unit 15 , and an output unit 16 .

[0028] The structure prediction models 11 predict an all-atom structure model of a protein from amino acid sequence information. Each structure prediction model 11 is applied to one monomer. When the amino acid sequence information is a multimer, each structure prediction model 11 predicts an all-atom structure model of the monomers that make up the protein for each chain. In Example 1, a case where the target protein is a trimer will be described.

[0029] The preprocessing unit 12 performs preprocessing for predictive control.

[0030] As an example of the first preprocessing step, the preprocessing unit 12 uses the PDB file 21 to determine three rigid transformations that fit each chain of the all-atom structural model of each of the three predicted monomers (trimers) through point cloud alignment. This is to determine rigid transformations that correspond to the individual chains. The PDB file 21 contains information about the all-atom structural models of the three monomers that make up the protein, as well as information about which chain each atom belongs to. Therefore, by using the PDB file 21, the preprocessing unit 12 can perform point cloud alignment for each chain of the all-atom structural model of each predicted trimer, and determine three rigid transformations that fit each chain. The preprocessing unit 12 then applies the determined rigid transformations to each chain of the all-atom structural model of each of the three predicted monomers, combining them into a single multimer.

[0031] As an example of the second preprocessing, the preprocessing unit 12 determines a rigid body transformation that best fits the combined multimer and the 3D density map of the target in the EMDB data 22. That is, the preprocessing unit 12 determines a rigid body transformation used to align the all-atom structural model with the constraining 3D density map. The equation for aligning is expressed, for example, by Equation (1). Note that R c , t c is a rigid transformation that indicates rotation and translation. x ca , x´ ca indicates the atomic coordinates of the original atom a and the atomic coordinates of atom a after rigid body transformation.

number

[0032] For example, in the second preprocessing, the preprocessing unit 12 performs a first step and a second step. In the first step, the preprocessing unit 12 performs center alignment (translation) and principal component axis alignment (rotation) on the combined multimers using a limited three-dimensional density map to obtain a rough t c and R c In the second stage, the preprocessing unit 12 determines R c , t c With x as a variable, ca is fixed as a constant, and the alignment is fine-tuned to obtain t c and R c In other words, the second preprocessing does not use information about which chain each atom of each monomer originally belonged to, but instead determines how to translate and rotate a rigid body, a single polymer formed by combining each monomer, so that it best matches the restricting 3D density map. This allows the preprocessing unit 12 to align the all-atom structural model with the restricting 3D density map.

[0033] After the preprocessing by the preprocessing unit 12 is completed, the control unit 10 forward propagates the layers of the structure prediction model 11 for each chain. Note that the functional unit that performs the forward propagation is not limited to the control unit 10, but may be the structure prediction model 11 or the like. Alternatively, the forward propagation may be performed by a user. The equation for forward propagating the layers is expressed by, for example, equation (2). Note that, r ci indicates the intermediate feature for each chain. i is the index for each residue. x ca indicates the atomic coordinates of the predicted atomic structure. When forward propagating through layers, the intermediate feature values ​​of the structure prediction model 11 are not changed.

number

[0034] The second preprocessing was divided into two stages, Stage 1 and Stage 2, to determine the rigid body transformation that best fits the combined multimers and the limiting 3D density map. However, the second preprocessing is not limited to this, and a genetic algorithm may be used to determine the rigid body transformation that best fits the combined multimers and the limiting 3D density map.

[0035] The conversion unit 13 converts the atomic structure output as a prediction result from the structure prediction model 11 into a three-dimensional density map (EM) that defines the entire atomic structure model by a differentiable operation.

[0036] For example, the conversion unit 13 applies the rigid body transformation obtained by the preprocessing unit 12 to the all-atom structure model of the prediction result to obtain the atomic structure of one multimer (trimer). That is, the conversion unit 13 obtains the all-atom structure model of one multimer (trimer) from the all-atom structure model of the prediction result using formula (1).

[0037] Then, the conversion unit 13 converts the predicted all-atom structure model of the trimer into a three-dimensional density map that defines the all-atom structure model using an interpolation formula that satisfies the density conservation law. The interpolation formula that satisfies the density conservation law is expressed, for example, by the following formulas (3), (4), and (5).

[0038] d used in Eq. (3) N indicates the length of one side of the voxel. Ngk , x´ cak indicates the k-component of the coordinates of the vertices of the voxel and the k-component of the atomic coordinates, respectively. Ngcak is d N indicates the distance between atom a and the vertex g of the voxel normalized by

number

[0039] u used in Eq. (4) Ngcak is the calculation result of equation (3), d N indicates the distance between atom a and the vertex g of the voxel normalized by p Ngcakindicates the probability of atom a existing at vertex g of a voxel.

number

[0040] n used in Eq. (5) a indicates the atomic number of atom a. Ngcak indicates the probability of atom a existing at vertex g of the voxel, which is the calculation result of equation (4). Ng pred indicates the 3D density map to be calculated.

number

[0041] Here, the distance u between atom a and the vertex g of the voxel shown in equation (3) Ngcak and the probability of atom a existing at vertex g of the voxel, P Ngcak The relationship between these is shown in Figure 3. Figure 3 is a diagram showing the spread or existence probability of atoms. The x-axis of the graph shown in Figure 3 is a group of values ​​obtained from an equation in which the absolute value of the numerator shown on the right side of equation (3) is removed. In other words, the x-axis is a group of values ​​normalized to show how close the atom is to the vertex of the voxel. The y-axis is the existence probability p of atom a at vertex g of the voxel shown in equation (4). Ngcak This graph shows that the closer the atom is to the vertex of the voxel (the closer the x-axis value is to 0), the higher the probability of the atom's existence, and the farther the atom is from the vertex of the voxel (the farther the x-axis value is from 0), the lower the probability of the atom's existence. If the x-axis value is greater than "2" or less than "-2", the probability of the atom's existence on the y-axis is "0". Therefore, the conversion unit 13 calculates the probability of atom a's existence at vertex g of the voxel, Pp, which is calculated from equations (3) and (4). Ngcak By using this, atoms can be converted into a smooth image (3D density map) with a blurred appearance (concept of a filter). In other words, the conversion unit 13 realizes a conversion process that integrates the concept of atomic spread and the concept of a filter into one.

[0042] In this way, the conversion unit 13 converts the atomic structure of the trimer of the predicted result into a three-dimensional density map ρ Ng pred These formulas (3), (4) and (5) are differentiable transformation functions that transform the all-atom structure model into a three-dimensional density map that defines the all-atom structure model.

[0043] The difference calculation unit 14 calculates the difference between the 3D density map corresponding to the predicted all-atom structural model of the trimer and the 3D density map of the target in the EMDB data 22. For example, the difference calculation unit 14 calculates the difference between the 3D density map that defines the predicted all-atom structural model and the 3D density map of the target by calculating the cross-correlation. The equation for calculating the cross-correlation is expressed, for example, by the following equation (6).

[0044] ρ used in Eq. (6) Ng pred denotes the three-dimensional density map calculated by equations (3), (4) and (5). Ng targ indicates the 3D density map of the target. N indicates the cross-correlation value of the 3D density map. g and N indicate the vertex of the voxel and the division number of the voxel, respectively.

number

[0045] In this way, the difference calculation unit 14 uses formula (6) to calculate the difference between the 3D density map of atom a of the trimer in the prediction result and the 3D density map of the target. This formula (6) is a differentiable formula. Note that the formula for calculating the difference has been described as a formula for calculating cross-correlation, but is not limited to this. The formula for calculating the difference may also use the L2 norm or the L1 norm.

[0046] The difference calculation unit 14 also uses the calculated difference to calculate an objective function that indicates constraints on the predicted trimer all-atom structural model. Additionally, the difference calculation unit 14 adds constraints that the protein should have to the calculated objective function to calculate a final objective function. For example, the difference calculation unit 14 calculates the final objective function using the following equation (7):

[0047] L used in Eq. (7) N is the value calculated by equation (6). bondlength , L bondangle are examples of objective functions that maintain the distance and angle of peptide bonds, respectively. N indicates the number of voxel divisions. L total is the final objective function.

number

[0048] The first term on the right side of equation (7) is the main objective function. This first term smoothly reduces the difference by using a technique of mixing coarse and fine region division.

[0049] In this way, the difference calculation unit 14 uses equation (7) to calculate an objective function that indicates constraints on the atomic structure of the trimer of the predicted result. This equation (7) is a differentiable equation. Note that while the added constraints have been described as, for example, the distance and angle of peptide bonds, they are not limited to this and may also be excluded volume effects, disulfide bond distances, hydrogen bond energy, etc. In short, the difference calculation unit 14 simply adds constraints that the protein must have as options to the main objective function.

[0050] Then, the control unit 10 calculates the back propagation of the difference. Note that the back propagation of the difference is not limited to the control unit 10, but may be performed by the difference calculation unit 14 or the like. For example, the control unit 10 calculates the back propagation of the difference using the following equations (8) to (13). Equation (8) is obtained by multiplying L shown in equation (7) by total L N Equation (9) is derived with respect to L shown in equation (6). N ρ Ngpred Equation (10) is differentiated with respect to ρ shown in equation (5). Ng pred ρ Ngcak Equation (11) is differentiated with respect to ρ shown in equation (4). Ngcak U Ngcak Equation (12) is differentiated with respect to u shown in equation (3). Ngcak x´ cak Then, equation (13) is derived from L shown in equation (6). N x´ cak is differentiated with respect to

[0051]

number

[0052]

number

[0053]

number

[0054]

number

[0055]

number

[0056]

number

[0057] In addition, the control unit 10 calculates L for each chain by backpropagating the difference. totalThe gradient is calculated based on the intermediate feature amount of the above. Note that the calculation of the gradient is not limited to being performed by the control unit 10, and may be performed by the difference calculation unit 14 when backpropagation of the difference is performed by the difference calculation unit 14. Alternatively, the calculation of the gradient may be performed by the structure prediction model 11 or the conversion unit 13.

[0058] The update unit 15 calculates the intermediate feature for each chain by the following formula (14) using the gradient calculated by backpropagation of the difference. ci indicates the intermediate feature for each chain. i is the index indicating each residue. Also, δx ca / δr ci is the atomic coordinate x shown in formula (2) ca The intermediate feature value r ci is the derivative with respect to

number

[0059] Furthermore, the update unit 15 updates the intermediate feature values ​​for each chain to the structure prediction model 11 corresponding to each chain. That is, the update unit 15 changes the intermediate feature values ​​of the structure prediction model 11 so as to reduce the difference between the all-atom structure model of the prediction result and the all-atom structure model of the target. In other words, the update unit 15 changes the intermediate feature values ​​of the structure prediction model 11 according to the constraints imposed by the limiting 3D density map.

[0060] The output unit 16 stores the three-dimensional density map (EM) converted by the conversion unit 13 in the storage unit 20 as EM data 23.

[0061] (Example of output atomic structure) 4 is a diagram showing a display example of the output all-atom structure model. The image shown in FIG. 4 shows the atomic coordinates x' of the all-atom structure model output from the structure prediction model 11. ca (See formula (1)) is displayed on the screen. The atomic coordinates x' of this all-atom structure model are caare the atomic coordinates after rigid transformation to each chain. The user can use the atomic coordinates of the all-atom structure model displayed on the screen to fit (adjust) the 3D density map.

[0062] In Example 1, the limiting information that limits the all-atom structure model is described as a three-dimensional density map (EM). However, the limiting information that limits the all-atom structure model is not limited to a three-dimensional density map (EM) and may be a three-dimensional density map obtained by X-ray structural analysis. Furthermore, the limiting information that limits the all-atom structure model may be the coordinates of the all-atom structure model.

[0063] Here, a display example of an output all-atom structure model when the limiting information limiting the all-atom structure model is the coordinates of the all-atom structure model will be described with reference to FIG. 5. FIG. 5 is a diagram showing a display example of the coordinates of the output all-atom structure model. The image shown in FIG. 5 displays the coordinates of the all-atom structure model output from the structure prediction model 11 on the screen (solid lines). The atomic coordinates of this all-atom structure model are the coordinates after rigid body transformation into each chain. In addition, the image shown in FIG. 5 displays the coordinates of the limiting information on the screen (dashed lines). Note that when the limiting information limiting the all-atom structure model is the coordinates of the atomic structure, the limiting information of the target can be obtained, for example, from the PDB file 21. The user can use this to align the coordinates (solid lines) of the all-atom structure model output from the structure prediction model 11 displayed on the screen with the coordinates (dashed lines) of the limiting information.

[0064] Also, Figure 5 shows only the carbon atoms (Cα) present in the main chain of each amino acid. The number of Cα atoms is about one-tenth of the total number of atoms. The atomic coordinates and line structures of the all-atom structure model output from the structure prediction model 11 contain position information for all atoms, not just the displayed Cα. On the other hand, the coordinates and line structures of the target's limited information are obtained from the PDB file 21, and contain position information for atoms other than Cα, but only Cα is used as the limited information. In other words, even if only about one-tenth of all atoms are available as atomic structures for limited information, the user can still reconstruct the position information for all atoms by using the all-atom structure model output from the structure prediction model 11 as a hint.

[0065] (Flowchart of predictive control process) Here, a flowchart of the predictive control process performed by the information processing device 1 will be described with reference to Fig. 6. Fig. 6 is a diagram illustrating an example of a flowchart of the predictive control process according to the first embodiment.

[0066] 6, the information processing device 1 predicts the structure of a monomer for each chain from the amino acid sequence (step S11). For example, the structure prediction model 11 predicts an all-atom structure model of a monomer constituting a protein for each chain from the amino acid sequence.

[0067] As a preprocessing step, the information processing device 1 uses the PDB file 21 to determine three rigid body transformations that fit each chain of the predicted three monomer (trimer) structure (step S12). For example, the preprocessing unit 12 executes a first preprocessing step. Then, the information processing device 1 applies the rigid body transformations determined for each chain to combine them into one multimer (step S13).

[0068] Then, as a preprocessing step, the information processing device 1 performs a rigid body transformation R that fits the three-dimensional density map (EM) to limit the multimer. c , t c (Step S14) For example, the preprocessing unit 12 executes a second preprocessing.

[0069] Then, the information processing device 1 performs forward propagation through each layer for each chain (step S15). For example, the information processing device 1 executes equation (2). Here, the information processing device 1 calculates the intermediate feature r ci Do not change.

[0070] Then, the information processing device 1 applies the rigid body transformation obtained in the preprocessing of S12 and S14 to obtain an all-atom structural model of the trimer (step S16). For example, the information processing device 1 executes equation (1).

[0071] Then, the information processing device 1 converts the atomic structure of the trimer into a three-dimensional density map (step S17). For example, the information processing device 1 executes the formulas (3), (4), and (5).

[0072] Then, the information processing device 1 obtains the difference between the converted three-dimensional density map and the three-dimensional density map (EM) of the target (step S18). For example, the information processing device 1 executes equation (6).

[0073] Then, the information processing device 1 calculates the objective function using the obtained difference (step S19). For example, the information processing device 1 executes equation (7).

[0074] Then, the information processing device 1 calculates back propagation (S19 → S18 → S17 → S16 → S15) and obtains intermediate features (step S20). For example, the information processing device 1 executes equations (8) to (14) to calculate back propagation and obtain intermediate features. Then, the information processing device 1 updates the structure prediction model 1 with the obtained intermediate features (step S21).

[0075] Then, the information processing device 1 determines whether the structure prediction by the structure prediction model 11 has converged (step S22). If it is determined that the structure prediction has not converged (step S22; No), the information processing device 1 proceeds to step S15 to perform the next structure prediction.

[0076] On the other hand, if it is determined that the structure prediction has converged (step S22; Yes), the information processing device 1 ends the prediction control process.

[0077] As described above, the information processing device 1 according to this embodiment changes the intermediate feature values ​​of the structure prediction model 11 so as to reduce the difference between a 3D density map obtained by converting a predicted structure output as a prediction result from the structure prediction model 11 using an interpolation formula that satisfies the density conservation law and a 3D density map that is different from the 3D density map. Therefore, the information processing device 1 according to this embodiment can predict a variety of 3D structures. Furthermore, by changing the intermediate feature values ​​of the structure prediction model 11 so as to reduce the difference between the 3D density map obtained by converting the predicted structure and the correct 3D density map, the information processing device 1 can predict a variety of 3D structures in a realistic amount of time. [Example]

[0078] In the information processing device 1 according to the first embodiment, the conversion unit 13 converts the atomic structure output as a prediction result from the structure prediction model 11 into a three-dimensional density map (EM) that defines the atomic structure using an interpolation formula that satisfies the density conservation law as a differentiable form. However, the conversion unit 13 is not limited to the interpolation formula that satisfies the density conservation law, and may convert the entire atomic structure model into a three-dimensional density map (EM) that defines the entire atomic structure using, for example, an approximation formula based on a spherically symmetric Gaussian distribution.

[0079] When the molecular shape is the limiting information for defining an all-atom structure model, a three-dimensional electron density map of the molecule can be used, which can be calculated accurately and based on physics using given atomic positions and atomic scattering factors.

[0080] On the other hand, because the 3D electron density map has many ups and downs, there is a possibility that a search using limited information may become stuck at a local solution. Therefore, by approximating the functional form of the atomic scattering factor and smoothing the 3D electron density map converted from the approximated functional form in advance, it is possible to improve the smoothness of the search while maintaining the accuracy of the molecular shape.

[0081] The atomic scattering factor f(q) is expressed by four Gaussian functions and a constant term as shown in equation (15). The atomic scattering factor f(q) can be approximated by four or fewer Gaussian functions as shown in equation (16). f represents the number of Gaussian functions used for approximation. a, b, and c are fitting parameters, and new fitting may be performed.

number

number

[0082] As an example of a method for smoothing a 3D electron density map (creating a smooth image), a low-pass filter in wavenumber space (cutoff wavenumber = f c When using a low-pass filter in wavenumber space, the target resolution (degree of smoothing) is R [Å], the molecular size parameter is L [Å], and the number of voxels per dimension N can be given by the following equation (17): N=2L / R=Lf c ...Equation (17) The molecular size parameter L can be estimated with high accuracy by using an accurately calculated three-dimensional electron density map and the molecular surface defined therefrom.

[0083] In this way, the conversion unit 13 converts the predicted all-atom structure model of the trimer into a three-dimensional density map that defines the all-atom structure model using equations (16) and (17). This equation (16) is a conversion function from the all-atom structure model to the three-dimensional density map that defines the all-atom structure model, and is a differentiable conversion function.

[0084] As a result, the information processing device 1 according to the second embodiment changes the intermediate feature amount of the structure prediction model 11 so as to reduce the difference between a 3D density map obtained by converting a predicted structure output as a prediction result from the structure prediction model 11 using an approximation formula based on a spherically symmetric Gaussian distribution and a 3D density map different from the 3D density map. Therefore, the information processing device 1 according to the second embodiment can predict a variety of 3D structures. The information processing device 1 can predict a variety of 3D structures in a realistic time by changing the intermediate feature amount of the structure prediction model 11 so as to reduce the difference between the 3D density map obtained by converting the predicted structure and the correct 3D density map. [Example]

[0085] Although the embodiments of the disclosed device have been described above, the present invention may be embodied in various different forms other than the above-described embodiments. Therefore, other embodiments included in the present invention will be described below.

[0086] The information including the processing procedures, control procedures, specific names, various data and parameters shown in the documents and drawings of the first and second embodiments may be changed arbitrarily unless otherwise specified.

[0087] Furthermore, the specific form of distribution or integration of the components of each device is not limited to that shown in the drawings. That is, all or some of the components may be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions of each device may be realized by a CPU and a program analyzed and executed by the CPU, or may be realized as hardware using wired logic. Furthermore, for all or any part of the processing functions of each device, a computing device such as a GPU (Graphics Processing Unit) or a TPU (Tensor Processing Unit) may be used instead of or in combination with the CPU.

[0088] The various processes described in the first and second embodiments can be realized by executing a prepared program on a computer such as a personal computer or a workstation. Alternatively, a so-called supercomputer or an HPC (High Performance Computing) machine may be used as the computer that executes the various processes described in the first and second embodiments. An example of a computer that executes a predictive control program having the same functions as those in the first and second embodiments will be described below with reference to FIG. 7.

[0089] Fig. 7 is a diagram showing an example of a hardware configuration. As shown in Fig. 7, a computer 100 has an operation unit 110a, a speaker 110b, a camera 110c, a display 120, and a communication unit 130. The computer 100 also has a CPU 150, a ROM 160, an HDD 170, and a RAM 180. These units 110 to 180 are connected via a bus 140.

[0090] 7, the HDD 170 stores a prediction control program 170a that performs the same functions as the structure prediction model 11, preprocessing unit 12, conversion unit 13, difference calculation unit 14, update unit 15, and output unit 16 (in other words, control unit 10) shown in the first embodiment. This control program 170a may be integrated or separated, similar to the components of the structure prediction model 11, preprocessing unit 12, conversion unit 13, difference calculation unit 14, update unit 15, and output unit 16 shown in FIG. 2. In other words, the HDD 170 does not necessarily have to store all of the data shown in the first embodiment, as long as the data used for processing is stored in the HDD 170.

[0091] Under such an environment, the CPU 150 reads the predictive control program 170a from the HDD 170 and loads it into the RAM 180. As a result, the predictive control program 170a functions as a predictive control process 180a, as shown in FIG. 7. The predictive control process 180a loads various data read from the HDD 170 into an area of ​​the storage area of ​​the RAM 180 allocated to the predictive control process 180a, and executes various processes using the loaded data. For example, examples of processes executed by the predictive control process 180a may include the processes shown in FIG. 6. Note that the CPU 150 does not necessarily need to operate all of the processing units shown in the first embodiment, as long as the processing units corresponding to the processes to be executed are virtually implemented.

[0092] The predictive control program 170a does not necessarily have to be stored in the HDD 170 or the ROM 160 from the beginning. For example, the predictive control program 170a may be stored in a portable physical medium, such as a flexible disk (FD, CD-ROM, DVD disk, magneto-optical disk, or IC card) inserted into the computer 100. The computer 100 may then acquire and execute the predictive control program 170a from such a portable physical medium. Alternatively, the predictive control program 170a may be stored in another computer or server device connected to the computer 100 via a public line, the Internet, a LAN, a WAN, or the like. The predictive control program 170a stored in this manner may be downloaded to the computer 100 and then executed. [Explanation of symbols]

[0093] 1. Information processing equipment 10 Control Unit 11 Structural prediction model 12 Pretreatment section 13 Conversion unit 14 Difference calculation part 15 Update section 16 Output section 20 Memory section 21 PDB files 22 EMDB data 23 EM data

Claims

1. A prediction control program for a structure prediction model that predicts a three-dimensional structure of an organic compound from sequence information of the organic compound, The intermediate feature amount of the structure prediction model is changed so that a difference between first limiting information corresponding to a predicted structure output as a prediction result from the structure prediction model and second limiting information different from the first limiting information becomes small. A predictive control program that causes a computer to execute the process.

2. The changing process changes intermediate feature quantities of the structure prediction model using gradients obtained by backpropagation of the difference.

2. The predictive control program according to claim 1.

3. the changing process converts the predicted structure into the first limited information using either an interpolation formula satisfying the law of density conservation or an approximation formula based on a spherically symmetric Gaussian distribution; Calculating the difference between the first limited information and the second limited information 2. The predictive control program according to claim 1.

4. The process of calculating the difference uses any one of cross-correlation, L2 norm, and L1 norm to calculate the difference.

4. The predictive control program according to claim 3.

5. The process of calculating the difference further includes calculating an objective function in which a second objective function indicating a constraint on the organic compound is added to a first objective function indicating a constraint on the predicted structure using the difference.

5. The predictive control program according to claim 4.

6. The process of calculating the difference uses a mixture of coarse and fine region divisions as the first objective function.

6. The predictive control program according to claim 5.

7. performing preprocessing to determine a rigid body transformation that fits the predicted structure and the second constraint information; The process of changing the rigid body transformation is performed by applying the rigid body transformation to a predicted structure that is newly output as a prediction result from the structure prediction model.

2. The predictive control program according to claim 1.

8. The process of performing the preprocessing further includes, when the three-dimensional structure to be predicted is a multimer, determining a rigid transformation of the predicted structure into a chain by point cloud registration; The process of changing the structure involves applying the rigid transformation to a predicted structure that is newly output as a prediction result from the structure prediction model.

8. The predictive control program according to claim 7.

9. The first and second limiting information are any one of a three-dimensional density map derived from an electron microscope, a three-dimensional density map derived from X-ray analysis, and atomic coordinates of an all-atom structural model.

2. The predictive control program according to claim 1.

10. An information processing device that controls a structure prediction model that predicts a three-dimensional structure of an organic compound from sequence information of the organic compound, a control unit that changes intermediate feature quantities of the structure prediction model so as to reduce a difference between first limiting information corresponding to a predicted structure output as a prediction result from the structure prediction model and second limiting information different from the first limiting information. An information processing device comprising:

11. A method for controlling a structure prediction model that predicts a three-dimensional structure of an organic compound from sequence information of the organic compound, comprising: The intermediate feature amount of the structure prediction model is changed so that a difference between first limiting information corresponding to a predicted structure output as a prediction result from the structure prediction model and second limiting information different from the first limiting information becomes small. A predictive control method in which processing is performed by a computer.

Citation Information

Patent Citations

  • JP1145358901A

  • Method and apparatus for estimating three-dimensional structure of protein

    JP2000229994A