Information processing program, information processing method, and information processing device
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2026-08-13
Smart Images

Figure JP2025003929_13082026_PF_FP_ABST
Abstract
Description
Information processing program, information processing method, and information processing device
[0001] The present invention relates to an information processing program, an information processing method, and an information processing apparatus.
[0002] In fields such as drug discovery and food, the analysis of protein structure and structural changes is crucial. That is, proteins change structure in response to their environment. One method for analyzing proteins that change structure in various ways is the use of cryo-electron microscopy (Cryo-EM). This method uses Cryo-EM to project images of the protein's diverse structures from various angles, allowing researchers to determine the protein's three-dimensional structure in response to its environment. The protein's three-dimensional structure can be represented, for example, by its three-dimensional density structure. Based on the protein's three-dimensional structure in response to its environment, structural changes in the protein can be predicted.
[0003] One technique that can be used to predict structural changes in proteins is the String method. In the String method, several fixed points and several movable intermediate points are set in a latent space, and a path search is performed to optimize these intermediate points. The points along the searched path correspond to the structure of the change process.
[0004] Regarding technologies related to protein structure analysis, for example, a predictive system capable of predicting the reaction rate between molecules that form resins has been proposed. A simulation device that can improve computational accuracy and reduce computation time in simulations predicting the dynamic behavior of systems containing biomacromolecules has also been proposed. An information processing device for searching for highly accurate reaction pathways of molecules has also been proposed. Methods for identifying binding partners in proteins and other macromolecules have also been proposed. Furthermore, a method has been proposed that automatically generates plausible structural change pathways using experimental data from Cryo-EM directly, via a theoretical isochromatic latent space learned by a deep autoencoder model.
[0005] Japanese Patent Publication No. 2024-78744, Japanese Patent Publication No. 2013-84255, International Publication No. 2024 / 038864, U.S. Patent Application Publication No. 2005 / 0010368, Specification
[0006] Atsushi Tokuhisa, Kimihiro Yamazaki, Yuichiro Wada, Mutsuyo Wada, Takashi Katoh, Akira Nakagawa, Yoshinobu Akinaga, Yoko Sasakura, Yasushi Okuno, "PaStEL: Generator of Pathways with Structural Change on Pseudo Free-Energy Landscape from CryoEM Images", ChemRxiv, 20 March 2024, Version 1, DOI: 10.26434 / chemrxiv-2024-8t70s
[0007] When using the String method to predict protein structural changes, a bias may occur in the positions of the intermediate points moved during optimization, resulting in paths passing through these intermediate points not being appropriate. For example, a bias in intermediate points can cause some distances between adjacent intermediate points to become extremely long. If the distances between intermediate points are too long, the path connecting these points may pass through regions corresponding to structures that are highly unlikely to exist as part of the protein structure. Pathways passing through such regions are inappropriate as representing protein structural changes. This problem is not limited to protein structural changes; it also occurs when predicting structural changes of various substances whose structures are represented by ultra-high-dimensional data, such as three-dimensional density structures.
[0008] In one respect, this project aims to enable accurate prediction of structural changes in materials.
[0009] One proposal provides an information processing program that causes a computer to perform the following steps: The computer sets one or more second points on a path connecting multiple first points corresponding to each of multiple first structures of the object in a latent space corresponding to the object's possible structures. The computer sets one or more third points on subpaths obtained by dividing the path at the positions of the second points. Based on the information regarding the positions of the second and third points, the computer calculates an objective function relating to whether the object becomes the structure corresponding to the second and third points during the deformation process from one of the first structures to the other. The computer then moves the second points in the direction that optimizes the objective function.
[0010] According to one embodiment, it becomes possible to accurately predict structural changes in a substance. The above and other objects, features and advantages of the present invention will become apparent from the following description in conjunction with the accompanying drawings illustrating preferred embodiments as examples of the present invention.
[0011] This figure shows an example of an information processing method according to the first embodiment. This figure shows an example of a system configuration according to the second embodiment. This figure shows an example of server hardware. This figure shows an example of a protein structural change. This figure shows an example of a method for analyzing protein shape changes. This figure shows an example of a case where the optimized intermediate point position is inappropriate. This figure shows an example of minimum energy path search with intermediate points not subject to optimization added. This figure shows an example of a function that the server has for analyzing protein shape changes. This flowchart shows an example of a procedure for analyzing protein shape changes. This figure shows an example of intermediate point setting processing in the second embodiment. This flowchart shows an example of a path search process procedure in the second embodiment. This figure shows an example of intermediate point setting processing in the third embodiment. This flowchart shows an example of a path search process procedure in the third embodiment. This figure shows an example of intermediate point setting processing in the fourth embodiment. This flowchart shows an example of a path search process procedure in the fourth embodiment. This figure shows an example of intermediate point setting processing in the fifth embodiment. This flowchart shows an example of a path search process procedure in the fifth embodiment. This figure shows an example of intermediate point setting processing in the sixth embodiment. This flowchart shows an example of a path search process procedure in the sixth embodiment. This figure shows an example of intermediate point setting processing in the seventh embodiment. This is a flowchart showing an example of the procedure for the route search process in the seventh embodiment. This is a diagram showing an example of the intermediate point setting process in the eighth embodiment. This is a flowchart showing an example of the procedure for the route search process in the eighth embodiment.
[0012] The following description of this embodiment will be made with reference to the drawings. Note that each embodiment can be implemented by combining multiple embodiments within a reasonable range. [First Embodiment] The first embodiment is an information processing method for enabling appropriate prediction of structural changes of a substance to be analyzed.
[0013] Figure 1 is a diagram showing an example of an information processing method according to the first embodiment. Figure 1 shows an information processing device 10 for implementing the information processing method according to the first embodiment. The information processing device 10 can implement the information processing method according to the first embodiment by, for example, executing a predetermined information processing program.
[0014] The information processing device 10 includes a storage unit 11 and a processing unit 12. The storage unit 11 is, for example, a memory or storage device of the information processing device 10. The processing unit 12 is, for example, a processor of the information processing device 10. The information processing device 10 may have multiple processors. Some of the multiple processes performed by the information processing device 10 may be executed on different processors.
[0015] The memory unit 11 stores an autoencoder 1 that can convert the structure of the substance being analyzed (hereinafter sometimes referred to as "the target") into the position (coordinates) of a point in the latent space 3, and convert that point into the structure corresponding to that point based on information such as the coordinates of the point in the latent space 3. The autoencoder 1 is, for example, a trained model obtained by machine learning. For example, a neural network is used for the autoencoder 1.
[0016] The latent space 3 corresponds to the possible structures of the object. For example, the latent space 3 represents the position of a point associated with the possible structures of the substance being analyzed. The structure of a substance changes depending on the environment in which it is placed. The structure of a substance under specific conditions can be determined based on experimental or observational results. For example, if multiple (e.g., two) structures of a substance (first structures 2a, 2b) are known, the autoencoder 1 can be used to obtain points in the latent space 3 corresponding to those structures 2a, 2b (first points 3a, 3b).
[0017] The processing unit 12 enables accurate prediction of structural changes in a substance. For example, the processing unit 12 predicts the process of structural change when a substance changes from a first structure 2a to a first structure 2b using the following procedure.
[0018] The processing unit 12 sets one or more second points 5a, 5b on a path 4 that connects two first points 3a, 3b corresponding to the first structures 2a, 2b, respectively, within the latent space 3. Next, the processing unit 12 sets one or more third points 6a, 5f on the partial paths 4a, 4c obtained by dividing the path 4 at the positions of the second points 5a, 5b.
[0019] The processing unit 12 calculates an objective function (calculates the objective function and obtains a value) based on information regarding the positions of the second points 5a and 5b and the third points 6a to 6f. The objective function is a function that determines whether the deformation process of the material from one of the first structures 2a and 2b to the other results in structures 2c to 2j corresponding to the second points 5a and 5b and the third points 6a to 6f. The objective function is, for example, the MaxFluxObjective described later.
[0020] The processing unit 12 moves the second points 5a and 5b to positions where the value of the objective function is predicted to change in a predetermined direction through optimization. For example, the processing unit 12 moves the second points 5a and 5b in the direction in which the objective function is optimized (minimized or maximized). The positions to which the second points 5a and 5b are moved can be determined by an optimization method such as the gradient method.
[0021] Furthermore, the processing unit 12 moves the third points 6a to 6f onto the path 4, which has been transformed to pass through the second points 5a and 5b after the movement. The path 4 is represented, for example, by a smooth curve that passes through the second points 5a and 5b and connects the first points 3a and 3b. The third points 6a to 6f are moved to positions that are equally spaced on the partial path 4a to 4c of the destination.
[0022] The processing unit 12 then repeats the process of calculating the value of the objective function, moving the second points 5a and 5b, and moving the third points 6a to 6f until the convergence condition is satisfied. The convergence condition is, for example, that the change in the value of the objective function due to the movement of the second points 5a and 5b falls below a threshold. Another convergence condition is, for example, that the number of executions of the process reaches a predetermined number. In addition, other arbitrary conditions may be set as the convergence condition.
[0023] Using the autoencoder 1, structures 2c to 2j corresponding to the positions of the second points 5a, 5b or the third points 6a to 6f when the convergence condition is met can be obtained. These structures 2c to 2j can be presumed to be generated during the structural change process between the first structures 2a and 2b.
[0024] By predicting the structural change process using this procedure, the path 4 representing the change process is prevented from passing through regions where the material structure is unlikely to occur. In other words, the objects whose positions are moved by optimization are limited to the second points 5a and 5b. Therefore, before the convergence condition is met, for example, the subpath between the second point 5a and the second point 5b may pass through regions where the material structure is unlikely to occur. However, third points 6c and 6d are set on the subpath between the second point 5a and the second point 5b. If these third points 6c and 6d are in an inappropriate position, the value of the objective function will not be the target value (e.g., the minimum value). Therefore, by repeatedly optimizing the positions of the second points 5a and 5b, the path 4 is modified so that the positions of the third points 6c and 6d are also in appropriate positions. As a result, accurate prediction of the structural change of the material becomes possible.
[0025] The processing unit 12 can also optimize the number of third points 6a to 6f at each iteration until the convergence condition is satisfied. For example, if the convergence condition is not satisfied, the processing unit 12 determines the number of third points 6a to 6f to set according to the length of the subpath 4a to 4c based on the moved second points 5a and 5b before calculating the value of the next objective function. Then, the processing unit 12 adds or deletes third points 6a to 6f on the subpath 4a to 4c according to the determined number.
[0026] As a result, if the movement of the second points 5a and 5b increases the length of any of the subpaths 4a to 4c, the number of third points on that subpath increases. Conversely, if the movement of the second points 5a and 5b decreases the length of any of the subpaths 4a to 4c, the number of third points on that subpath decreases. This maintains an appropriate number of third points on each subpath 4a to 4c, improving the search accuracy of path 4.
[0027] When the convergence condition is satisfied, the processing unit 12 can also increase the number of the second points 5a and 5b and repeat the optimization process until the convergence condition is satisfied again. For example, when the convergence condition is satisfied, the processing unit 12 changes a third point, for which the amount of change in the value of the objective function when the position is shifted by a predetermined amount satisfies a predetermined condition, to the second point. Then, until the convergence condition is satisfied again, the processing unit 12 repeats the process of calculating the value of the objective function, the process of moving the second point (including the second point changed from the third point), and the process of moving the third point. Thereby, it becomes possible to obtain an appropriate path 4 with high precision.
[0028] When the convergence condition is satisfied, the processing unit 12 can also change the second points 5a and 5b to the first point and change the third points 6a to 6f to the second point and repeat the optimization process until the convergence condition is satisfied again. In this case, the processing unit 12 newly sets the third point for a partial path obtained by dividing the path 4 at the positions of the newly set first point and second point. Then, until the convergence condition is satisfied again, the processing unit 12 repeats the process of calculating the value of the objective function, the process of moving the newly installed second point, and the process of moving the third point.
[0029] The first point changed from the second points 5a and 5b does not move even if the optimization process is repeated. Therefore, by changing the second points 5a and 5b to the first point, the range of path search becomes narrower. As a result, it becomes possible to obtain an appropriate path 4 with high precision by limiting the search range compared to when the convergence condition was first satisfied.
[0030] The processing unit 12 can also change at least a part of the second points 5a and 5b and the third points 6a to 6f to the first point using the numerical values used when calculating the value of the objective function. For example, in the process of calculating the value of the objective function, the processing unit 12 calculates an evaluation value indicating whether a corresponding structure is formed for each of the second points 5a and 5b and each of the third points 6a to 6f using the information regarding the position. The processing unit 12 sets the sum of the evaluation values of each of the second points 5a and 5b and each of the third points 6a to 6f as the value of the objective function.
[0031] In the optimization process using such an objective function, when the convergence condition is satisfied, the processing unit 12 converts the second point or the third point whose evaluation value satisfies a predetermined condition into the first point. Further, the processing unit 12 converts the third point whose evaluation value does not satisfy the predetermined condition into the second point. The predetermined condition is, for example, that the evaluation value is greater than or equal to a threshold value, or that the evaluation value is less than or equal to a threshold value. After the conversion process, the processing unit 12 sets the third point in the partial path obtained by dividing the path 4 at the positions of the newly set first point and second point. Then, until the convergence condition is satisfied again, the processing unit 12 repeats the process of calculating the value of the objective function, the process of moving the second point, and the process of moving the third point.
[0032] Thereby, the region to be intensively searched can be determined based on the evaluation value, and an appropriate path within the corresponding region can be searched in detail. As a result, an appropriate prediction result can be obtained according to the purpose of analyzing the structural change of the substance.
[0033] Further, after satisfying the convergence condition, the processing unit 12 can also reset the second point and the third point at equal intervals on the path 4 at that time. For example, when the convergence condition is satisfied, the processing unit 12 deletes the existing second points 5a, 5b and the third points 6a to 6f from the path 4, and sets one or more new second points at equal intervals on the path 4. Further, the processing unit 12 sets one or more new third points at equal intervals in the partial path obtained by dividing at the position of the new second point on the path 4.
[0034] Thereby, even when the interval between the second points 5a, 5b becomes too wide or too narrow in the process of optimization, the optimization process is continued by resetting at equal intervals. As a result, an appropriate path 4 can be efficiently obtained.
[0035] [Second Embodiment] The second embodiment is a computer system that realizes highly accurate prediction of protein structural changes.
[0036] Figure 2 shows an example of the system configuration of the second embodiment. For example, a server 100 that performs highly accurate structural change prediction of proteins is connected to the network 20. The server 100 is a computer on which protein structural change prediction software is implemented. A terminal device 30 is further connected to the network 20. The terminal device 30 is a computer used by a user who wants to have the server 100 predict the structural changes of a protein.
[0037] Figure 3 shows an example of server hardware. The entire server 100 is controlled by a processor 101. The processor 101 is connected to memory 102 and several peripheral devices via a bus 109.
[0038] The server 100 may be a multiprocessor system having multiple processors. A collection of multiple processors in a multiprocessor system can be called a processor 101. The processor 101 may also be called a processor circuitry. Each of the multiple processors can execute some or all of the processes executed by the server 100. When there are multiple related processes, two or more of those processes may be executed by different processors.
[0039] The processor 101 is, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or a DSP (Digital Signal Processor). At least some of the functions that the processor 101 implements by executing a program may be implemented by electronic circuits such as an ASIC (Application Specific Integrated Circuit) or a PLD (Programmable Logic Device).
[0040] Memory 102 is used as the main memory of the server 100. Memory 102 temporarily stores at least a portion of the OS (Operating System) program and application programs that are to be executed by the processor 101. Memory 102 also stores various data used for processing by the processor 101. As memory 102, a volatile semiconductor memory device such as RAM (Random Access Memory) is used.
[0041] Peripheral devices connected to the bus 109 include a storage device 103, a graphics controller 104, an input interface 105, an optical drive device 106, a device connection interface 107, and a network interface 108.
[0042] The storage device 103 electrically or magnetically writes and reads data from its built-in recording medium. The storage device 103 is used as an auxiliary storage device for the server 100. The storage device 103 stores the OS program, application programs, and various data. For example, the storage device 103 can be an HDD (Hard Disk Drive) or an SSD (Solid State Drive).
[0043] The graphics controller 104 is an arithmetic unit that performs image processing. The graphics controller 104 is, for example, a GPU (Graphics Processing Unit). A monitor 21 is connected to the graphics controller 104. The graphics controller 104 displays images on the screen of the monitor 21 according to instructions from the processor 101. The monitor 21 can be a display device using organic EL (Electro Luminescence) or a liquid crystal display device. If a GPU is used as the graphics controller 104, the graphics controller 104 can also perform complex numerical calculations such as matrix calculations.
[0044] The input interface 105 is connected to a keyboard 22 and a mouse 23. The input interface 105 transmits signals from the keyboard 22 and mouse 23 to the processor 101. Note that the mouse 23 is just one example of a pointing device, and other pointing devices can also be used. Other pointing devices include touch panels, tablets, touchpads, and trackballs.
[0045] The optical drive device 106 uses laser light or the like to read data recorded on the optical disc 24 or write data to the optical disc 24. The optical disc 24 is a portable recording medium on which data is recorded in a way that makes it readable by the reflection of light. Examples of optical discs 24 include DVD (Digital Versatile Disc), DVD-RAM, CD-ROM (Compact Disc Read Only Memory), and CD-R (Recordable) / RW (ReWritable).
[0046] The device connection interface 107 is a communication interface for connecting peripheral devices to the server 100. For example, a memory device 25 and a memory reader / writer 26 can be connected to the device connection interface 107. The memory device 25 is a recording medium equipped with a communication function with the device connection interface 107. The memory reader / writer 26 is a device that writes data to or reads data from the memory card 27. The memory card 27 is a card-type recording medium.
[0047] The network interface 108 is connected to the network 20. The network interface 108 transmits and receives data to and from other computers or communication devices via the network 20. The network interface 108 is a wired communication interface, for example, connected by cable to a wired communication device such as a switch or router. Alternatively, the network interface 108 may be a wireless communication interface, connected by radio waves to a wireless communication device such as a base station or access point.
[0048] The server 100 can implement the processing functions of the second embodiment using the hardware described above. The information processing device 10 shown in the first embodiment can also be implemented using hardware similar to that of the server 100 shown in Figure 3.
[0049] The server 100 implements the processing functions of the second embodiment by executing a program recorded on, for example, a computer-readable recording medium. The program describing the processing content to be executed by the server 100 can be recorded on various recording media. For example, the program to be executed by the server 100 can be stored in the storage device 103. The processor 101 loads at least a portion of the program in the storage device 103 into the memory 102 and executes the program. Alternatively, the program to be executed by the server 100 can be recorded on a portable recording medium such as an optical disc 24, a memory device 25, or a memory card 27. The program stored on the portable recording medium becomes executable after being installed in the storage device 103, for example, under control from the processor 101. The processor 101 can also directly read and execute the program from the portable recording medium.
[0050] Next, we will explain the difficulty of predicting protein structural changes. Figure 4 shows an example of protein structural changes. Protein structural analysis and prediction of structural changes are performed in the following four steps, for example: [Step 1] Projection of diverse protein structures from various orientations using Cryo-EM [Step 2] Estimation of projection orientation [Step 3] Reconstruction of three-dimensional structure from images [Step 4] Prediction of protein shape changes based on the reconstructed three-dimensional structure For example, suppose that when the three-dimensional structure of a protein is constructed, four structures 41 to 44 are obtained. For example, the protein of structure 41 changes to structure 42 or structure 43 due to environmental changes. The protein of structure 42 changes to structure 43 or structure 44 due to environmental changes. The protein of structure 43 changes to structure 44 due to environmental changes.
[0051] Structures 41-44, obtained from projection results using Cryo-EM, are realistically plausible protein structures. When a protein's structure changes, the resulting structure is also a realistically plausible structure. Therefore, if the process of protein structural change can be accurately predicted, possible structures of proteins other than those experimentally confirmed can be identified. When a protein takes on a specific structure, that structure may have the potential to perform a useful function in the human body. For this reason, accurately predicting the process of protein structural change is important in various industries that handle proteins, such as drug discovery and food.
[0052] Figure 5 shows an example of a method for analyzing changes in protein shape. For example, the server 100 can use PaStEL (see Non-Patent Literature 1) to automatically calculate changes in protein structures 41-44 obtained from projection results by Cryo-EM.
[0053] The protein structures 41-44 are represented by ultra-high-dimensional data. Changes in protein shape can be analyzed using a special autoencoder (e.g., CryoTWIN) to use a lower-dimensional latent space 40 (e.g., an 8-dimensional space) instead of the space representing the actual shape (ultra-high-dimensional space). Each of the protein structures 41-44 corresponds to a point in the latent space 40.
[0054] For example, server 100 uses an AI (Artificial Intelligence) technology called DeepTWIN to learn the protein structure using cryo-EM images, and from that learning result, the probability distribution P of the protein structure in the latent space 40. ψ The server 100 seeks the probability distribution P after training. ψ This is used to automatically calculate a reasonable continuous deformation. For example, the positions of intermediate points set between the start and end points in the latent space are optimized, and a continuous deformation represented by a path passing through the optimized intermediate points is considered a reasonable continuous deformation. Specifically, a path in the latent space 40 that shows a reasonable continuous deformation is a path that minimizes the following MaxFluxObjective.
[0055] When the interval between the starting point and the ending point is divided into K (K is a natural number) segments, k = {0, 1, 2, ···, K}. The point where k = 0 is the starting point, and the point where k = K is the ending point. z k is the coordinate within the latent space 40 of the k-th intermediate point. "||z k+1 - z k ||" is the distance (L2 norm) between the (k + 1)-th point (intermediate point or ending point) and the k-th point (starting point or intermediate point). Equation (1) is the sum of the evaluation values ((1 / P ψ )||z k+1 - z k ||) corresponding to each of "k = 0, 1, 2, ···, K - 1". The evaluation value indicates the degree of contribution of each intermediate point to the value of MaxFluxObjective. The evaluation value becomes smaller as the probability distribution P ψ at the intermediate point is larger, and becomes smaller as the distance to the next intermediate point is shorter.
[0056] For example, based on the probability distribution P ψ within the latent space 40, the existence probability of the protein shape at each position within the latent space 40 can be represented by the contour map 40a. The two structures 41, 43 obtained based on the Cryo-EM image often become positions with high existence probability. In this case, the existence probability decreases as the distance from each of the structures 41, 43 increases.
[0057] When the protein transforms from structure 41 to structure 43, the trajectory of the position within the latent space 40 corresponding to the structure during the transformation is represented by the path 49 from the starting point corresponding to structure 41 to the ending point corresponding to structure 43. This path 49 is expected to pass through positions where the existence probability is higher than the surroundings.
[0058] According to Equation (1), the larger the existence probability "P ψ (z k )" of the intermediate point on the path 49, the smaller the MaxFluxObjective. Also, the length of the divided partial path "||z k+1 - z kThe shorter the ||, the smaller the MaxFluxObjective. The path 49 connecting the intermediate points that minimize the MaxFluxObjective represents the transformation process of the protein structure. For example, when a protein transforms from structure 41 to structure 43, it is estimated that during the transformation, it will become structures 45 and 46, which correspond to points on path 49.
[0059] If the path 49 from structure 41 to structure 43 within the latent space 40 can be correctly determined in this way, the deformation process from structure 41 to structure 43 will become clear. However, even if the positions of the intermediate points on the path 49 are optimized using MaxFluxObjective, it may not result in an appropriate path 49.
[0060] Figure 6 shows an example of a case where the optimized intermediate point positions are inappropriate. In the optimization of MaxFluxObjective, the start and end points are set as fixed points 51 and 52. Multiple intermediate points 53 to 57 are set on a path 50 connecting the two fixed points 51 and 52. The initial state of the path 50 is, for example, a straight line. Intermediate points 53 to 57 are the objects of differentiation of MaxFluxObjective. By differentiating MaxFluxObjective with respect to each of the intermediate points 53 to 57, for example by gradient descent, intermediate points 53 to 57 are moved in the direction in which the value of MaxFluxObjective decreases.
[0061] The optimization process ends when the value of MaxFluxObjective can no longer decrease after repeatedly moving between intermediate points 53 and 57. The curve that smoothly connects intermediate points 53 to 57 and fixed points 51 and 52 after the optimization process is the optimized path 50. This algorithm for finding the optimal path 50 is called the String method.
[0062] When using the String method to find the minimum value of MaxFluxObjective, the minimum energy path is searched by assuming that locations with lower probability of existence in the probability distribution have higher energy. When the String method is applied to search for the minimum energy path in a specific case, a bias in the intermediate points may occur.
[0063] For example, in latent space 59, a region 59e with higher energy than its surroundings is widely distributed between the fixed points 59a and 59b, which serve as the starting and ending points. In this case, when the minimum energy path is searched using the String method, the intermediate points move to positions that avoid the high-energy region 59e. If the high-energy region 59e is widespread, the intermediate points may concentrate around either of the fixed points 59a or 59b. If such a bias in the distribution of intermediate points occurs, the paths connecting the intermediate points will pass through the high-energy region 59e.
[0064] For example, if intermediate points 59c and 59d are adjacent, the optimized path will be one that passes near the shortest line connecting these intermediate points 59c and 59d. In this case, the optimized path will pass through a high-energy region 59e. In other words, the path obtained when there is a bias in the intermediate points represents a deformation that includes a structure that is unlikely to actually exist, and is therefore an inappropriate path.
[0065] One reason for the bias in the distribution of intermediate points is that the String method does not consider the energy (probability distribution) in the region between adjacent intermediate points. Furthermore, even if the number of intermediate points is increased, there is no guarantee that the uneven distribution of intermediate points around fixed points 59a and 59b can be avoided.
[0066] Therefore, the server 100 performs pathfinding using multiple first intermediate points for optimizing the positions between fixed points, as well as second intermediate points set between adjacent first intermediate points that are not subject to position optimization. In doing so, the server 100 uses both the first and second intermediate points as variables in the objective function (MaxFluxObjective) used for optimization. This makes it possible to predict the optimal path within the latent space representing the shape of the protein. Hereafter, the first and second intermediate points, together, may simply be referred to as intermediate points.
[0067] Figure 7 shows an example of minimum energy path search with the addition of intermediate points not subject to optimization. Multiple first intermediate points 62a to 62c, which are subject to calculation of appropriate positions through optimization, are set on the path 60 connecting fixed points 61a and 61b. Second intermediate points 63a to 63p are set in each of the multiple subpaths divided by the first intermediate points 62a to 62c.
[0068] In the example shown in Figure 7, two second intermediate points 63a to 63b are set between the starting fixed point 61a and the first first intermediate point 62a. Three second intermediate points 63c to 63e are set between the first first intermediate point 62a and the second first intermediate point 62b. Eight second intermediate points 63f to 63m are set between the second first intermediate point 62b and the third first intermediate point 62c. Three second intermediate points 63n to 63p are set between the third first intermediate point 62c and the ending fixed point 61b.
[0069] There are three first intermediate points 62a to 62c, and sixteen second intermediate points 63a to 63p. Therefore, the total number of intermediate points is 19. When calculating the objective function MaxFluxObjective, the value is calculated when the distance between fixed points 61a and 61b is divided into 20 parts (K=20) by the first intermediate points 62a to 62c and the second intermediate points 63a to 63p. The first intermediate points 62a to 62c are then moved to minimize the objective function. When the first intermediate points 62a to 62c are moved, the path 60 is modified to pass through the moved first intermediate points 62a to 62c. As a result of the modification of the path 60, the second intermediate points 63a to 63p set on the path 60 also move.
[0070] The second intermediate points 63a to 63p are moved to positions that divide the partial path between the fixed point and the first intermediate point, or the partial path between the two first intermediate points, into equal lengths. For example, the two second intermediate points 63a and 63b between the fixed point 61a and the first intermediate point 62a are placed at two points that divide the path between the fixed point 61a and the first intermediate point 62a after optimization into three parts.
[0071] Thus, in the path search to optimize the first intermediate points 62a to 62c, the objective function used for optimization is calculated based on the positions of the first intermediate points 62a to 62c and the second intermediate points 63a to 63p. As a result, if the probability of existence of the structure corresponding to the positions of the second intermediate points 63a to 63p is low (high energy), the value of the objective function will also be high. Consequently, the positions of the first intermediate points 62a to 62c are optimized so that the second intermediate points 63a to 63p move to a lower energy position.
[0072] For example, suppose there is a high-energy region 64c between two fixed points 64a and 64b in the latent space 64. If the first intermediate point is located near the two fixed points 64a and 64b, the second intermediate point will be located within the high-energy region 64c. This will increase the energy value of the objective function. As a result, the placement of adjacent first intermediate points across the high-energy region 64c is suppressed, and the first intermediate points are positioned so as to create a path that avoids the high-energy region 64c.
[0073] By suppressing the uneven distribution of the first intermediate point in this way, optimization yields an appropriate pathway within the latent space 64 representing protein deformation. Obtaining an appropriate pathway enables accurate prediction of the structure during the protein deformation process.
[0074] Figure 8 shows an example of the functions a server has for analyzing changes in protein shape. Server 100 has a storage unit 110, an analysis request receiving unit 120, a latent space transformation unit 130, and a path search unit 140.
[0075] The memory unit 110 stores protein structure data 111. Protein structure data 111 is data showing the structure of a protein under various environmental conditions, obtained based on cryo-EM images. Protein structure data 111 is, for example, data representing the three-dimensional density structure of a protein.
[0076] The analysis request receiving unit 120 receives an analysis request for protein shape change from the terminal device 30. The analysis request includes information such as the protein to be analyzed, the starting structure, and the ending structure. When the analysis request receiving unit 120 receives an analysis request from the terminal device 30, it transmits the contents of the received analysis request to the latent space transformation unit 130. When the analysis request receiving unit 120 receives the analysis results from the latent space transformation unit 130, it visualizes the analysis results and transmits them to the terminal device 30. The analysis results include structural data representing the three-dimensional density structure of the deformation process. The analysis request receiving unit 120 generates a three-dimensional model showing the three-dimensional structure of the protein based on the structural data shown in the analysis results, and transmits data representing that three-dimensional model to the terminal device 30.
[0077] The latent space conversion unit 130 is an autoencoder that converts from a three-dimensional density structure to coordinates of points in latent space and from coordinates of points in latent space back to a three-dimensional density structure. For example, based on an analysis request, the latent space conversion unit 130 obtains structural data showing the structure of the starting point and structural data showing the structure of the endpoint of the protein to be analyzed from the storage unit 110. The latent space conversion unit 130 then converts the obtained structural data into coordinates of points in latent space, for example, using the encoder of CryoTWIN. The latent space conversion unit 130 transmits the coordinates of the starting point and the coordinates of the endpoint in latent space to the path search unit 140. When the latent space conversion unit 130 receives coordinates of points along the path showing the shape changes of the protein from the path search unit 140, the CryoTWIN decoder generates structural data showing the structure of the protein corresponding to those coordinates. The latent space conversion unit 130 transmits the generated structural data as an analysis result to the analysis request reception unit 120.
[0078] The path search unit 140 obtains the coordinates of the start and end points in latent space from the latent space transformation unit 130, and then searches for an appropriate path that shows the deformation process when the shape of the protein is deformed from the start point to the end point. The path search unit 140 then transmits the coordinates of the points on the predicted path to the latent space transformation unit 130.
[0079] The functions of each element shown in Figure 9 can be realized, for example, by having the processor 101 execute the program module corresponding to that element. Next, the procedure for analyzing protein shape changes will be explained in detail.
[0080] Figure 9 is a flowchart showing an example of the procedure for analyzing changes in protein shape. The process shown in Figure 9 will be described below according to the step numbers. [Step S101] When the analysis request receiving unit 120 receives an analysis request, it transmits the contents of the analysis request to the latent space conversion unit 130. The latent space conversion unit 130 obtains the structural data of the start and end points of the protein indicated in the analysis request from the storage unit 110.
[0081] [Step S102] The latent space transformation unit 130 determines the positions of the start point and the end point in the latent space. For example, the latent space transformation unit 130 uses the coordinates in the latent space obtained by encoding the structural data of the start point as the position of the start point. The latent space transformation unit 130 also uses the coordinates in the latent space obtained by encoding the structural data of the end point as the position of the end point. The latent space transformation unit 130 transmits the determined positions of the start point and the end point to the path search unit 140.
[0082] [Step S103] The path search unit 140 performs a path search process to search for a path in the latent space that represents the shape change from the starting point to the ending point. Details of the path search process will be described later (see Figure 11). The path search unit 140 transmits the coordinates on the path obtained in the path search process to the latent space transformation unit 130.
[0083] [Step S104] The latent space transformation unit 130 decodes the coordinates in the latent space obtained as a result of the path search and converts them into structural data that shows the structure of the protein. The latent space transformation unit 130 transmits the analysis results, including the structural data generated by the transformation, to the analysis request reception unit 120.
[0084] [Step S105] The analysis request receiving unit 120 visualizes the three-dimensional structure shown in the protein structure data. For example, the analysis request receiving unit 120 generates three-dimensional model data showing the three-dimensional structure of the protein and transmits the three-dimensional model data to the terminal device 30.
[0085] In this way, the user can observe how the protein's shape deforms on the terminal device 30. The user can also use the terminal device 30 to analyze the protein's structure during deformation and verify the usefulness of that protein structure.
[0086] Next, the path search process will be explained in detail. Figure 10 shows an example of the intermediate point setting process in the second embodiment. For example, the path search unit 140 connects the fixed points 71a and 71b corresponding to the start point and end point with, for example, a straight path 70 as the initial state. The path search unit 140 sets first intermediate points 72a to 72c, which are subject to position calculation by optimization, at equal intervals on the path 70. In the example of Figure 10, three first intermediate points 72a to 72c are set, so the path 70 is divided into four sub-paths with the first intermediate points 72a to 72c as the boundaries.
[0087] The path search unit 140 sets second intermediate points 73a to 73l at equal intervals on each partial path. For example, on the partial path connecting the fixed point 71a and the first intermediate point 72a with a straight line, second intermediate points 73a to 73c are set at equal intervals. On the partial path connecting the first intermediate point 72a and the first intermediate point 72b with a straight line, second intermediate points 73d to 73f are set at equal intervals. On the partial path connecting the first intermediate point 72b and the first intermediate point 72c with a straight line, second intermediate points 73g to 73i are set at equal intervals. On the partial path connecting the first intermediate point 72c and the fixed point 71b with a straight line, second intermediate points 73j to 73l are set at equal intervals.
[0088] The objective function (MaxFluxObjective) is calculated using the positions of the first midpoints 72a–72c and the second midpoints 73j–73l as variables, with the value (e.g., energy value) being calculated. However, the second midpoints 73a–73l are excluded from the calculation of their destination positions through optimization (e.g., calculations using differentiation). When the positions of the first midpoints 72a–72c are moved through optimization, the path 70 becomes a smooth curve passing through the positions of the first midpoints 72a–72c. When the path 70 is changed, the second midpoints 73a–73l move to positions that equally divide each subpath of the changed path 70.
[0089] Figure 11 is a flowchart showing an example of the procedure for the pathfinding process in the second embodiment. The process shown in Figure 11 will be described below in accordance with the step numbers. [Step S111] The pathfinding unit 140 sets the first intermediate points to be optimized at equal intervals between the fixed points corresponding to the start point and the end point.
[0090] [Step S112] The path search unit 140 sets second intermediate points, which are not to be optimized, at equal intervals on each sub-path divided by the first intermediate point. [Step S113] The path search unit 140 calculates the objective function based on the positions of the first intermediate point and the second intermediate point. This gives, for example, a value of MaxFluxObjective.
[0091] [Step S114] The pathfinding unit 140 determines whether the convergence condition has been met. For example, the pathfinding unit 140 determines that the convergence condition has been met if the difference between the value of the objective function calculated immediately before and the value of the objective function calculated before that is less than or equal to a threshold. If the convergence condition has been met, the pathfinding unit 140 terminates the pathfinding process. If the convergence condition has not been met, the pathfinding unit 140 proceeds to step S115.
[0092] [Step S115] The pathfinding unit 140 changes the position of the first intermediate points in a direction that minimizes the value of the objective function. For example, the pathfinding unit 140 differentiates the objective function at each position of the first intermediate points and calculates the optimized position of each first intermediate point using a minimization algorithm such as the gradient method. The pathfinding unit 140 moves the first intermediate points to the calculated positions.
[0093] [Step S116] The path search unit 140 moves the position of the second intermediate point based on the position of the first intermediate point after optimization. For example, the path search unit 140 generates a smooth curve that passes through the first intermediate point and connects the two fixed points, and uses the generated curve as the path after optimization. The path search unit 140 then moves the second intermediate point onto the path after optimization. After that, the path search unit 140 proceeds to step S113.
[0094] In this way, by introducing a second intermediate point that is not subject to optimization between the first intermediate points, and by calculating a value that reflects the position of the second intermediate point in the objective function calculation, the uneven distribution of the first intermediate point in the latent space is suppressed. As a result, an appropriate pathway that shows the shape deformation process of the protein is obtained.
[0095] [Third Embodiment] In the third embodiment, when the length of a partial path divided at the first intermediate point is changed due to the movement of the first intermediate point, the number of second intermediate points on the partial path is increased or decreased according to that change in length.
[0096] Figure 12 shows an example of the intermediate point setting process in the third embodiment. For example, in the initial state, the first intermediate points 72a to 72c and the second intermediate points 73a to 73l are set, similar to the second embodiment. In the initial state, the path 70 is divided at equal intervals by the first intermediate points 72a to 72c, and the length of each divided section of the path is the same. Therefore, the number of second intermediate points set in each section of the path is also the same, three in each case.
[0097] If the first intermediate points 72a to 72c are moved to appropriate positions during the pathfinding process, the lengths of the sub-paths divided by the first intermediate points 72a to 72c become uneven. In the example in Figure 12, the distance between the fixed point 71a and the first intermediate point 72a decreases, and the sub-path 74a between them becomes shorter. Therefore, two second intermediate points 73b and 73c are removed from the sub-path 74a between the fixed point 71a and the first intermediate point 72a. The distance between the first intermediate point 72a and the first intermediate point 72b increases, and the sub-path 74b between them becomes longer. Therefore, three second intermediate points 73m to 73o are added to the sub-path 74b between the first intermediate point 72a and the first intermediate point 72b. The distance between the first intermediate point 72b and the first intermediate point 72c increases, and the sub-path 74c between them becomes longer. Therefore, a second intermediate point 73p is added to the partial path 74c between the first intermediate point 72b and the first intermediate point 72c. The distance between the first intermediate point 72c and the fixed point 71b is reduced, and the partial path 74d between them is shortened. Therefore, two second intermediate points 73k and 73l are removed from the partial path 74d between the first intermediate point 72c and the fixed point 71b.
[0098] Figure 13 is a flowchart showing an example of the procedure for the pathfinding process in the third embodiment. The process shown in Figure 13 will be described below in accordance with the step numbers. [Step S211] The pathfinding unit 140 sets the first intermediate points to be optimized at equal intervals between the fixed points corresponding to the start point and the end point.
[0099] [Step S212] The path search unit 140 sets a number of second intermediate points on each sub-path divided by the first intermediate point, the number of intermediate points corresponding to the length of the sub-path. For example, the path search unit 140 uses the quotient obtained by dividing the length of the sub-path by a predetermined reference value as the number of second intermediate points to set on that sub-path.
[0100] [Step S213] The pathfinding unit 140 calculates the objective function based on the positions of the first intermediate point and the second intermediate point. [Step S214] The pathfinding unit 140 determines whether the convergence condition has been met. If the convergence condition has been met, the pathfinding unit 140 terminates the pathfinding process. If the convergence condition has not been met, the pathfinding unit 140 proceeds to step S215.
[0101] [Step S215] The pathfinding unit 140 moves the first intermediate point in a direction that minimizes the value of the objective function. [Step S216] Based on the position of the first intermediate point after optimization, the pathfinding unit 140 increases or decreases the number of second intermediate points for each subpath and determines the position of the second intermediate points for each subpath. For example, the pathfinding unit 140 generates a smooth curve that passes through the first intermediate point and connects two fixed points, and uses the generated curve as the path after optimization. Next, the pathfinding unit 140 calculates the length of each subpath obtained by dividing the path at the first intermediate point and determines the number of second intermediate points to set for each subpath. Then, the pathfinding unit 140 sets a number of second intermediate points at equal intervals on each subpath corresponding to the length of that subpath. After that, the pathfinding unit 140 proceeds to step S213.
[0102] In this way, the number of second intermediate points set in a subpathway can be dynamically changed according to the length of the subpathway. This prevents the subpathway between two adjacent first intermediate points from passing through high-energy regions (regions with a low probability of protein structure existence), even if the distance between them becomes long.
[0103] [Fourth Embodiment] The fourth embodiment involves changing the second midpoint, which is located at a position where the influence on the objective function satisfies a certain criterion, to the first midpoint.
[0104] Figure 14 shows an example of the intermediate point setting process in the fourth embodiment. Multiple first intermediate points 72a to 72c and multiple second intermediate points 73a to 73l are set on a path 70 connecting fixed points 71a and 71b. During the path search process, the variance of the second intermediate points 73a to 73l is calculated. The variance is the amount of change in the value of the objective function when the position of the second intermediate point is shifted by a predetermined amount. At this time, any second intermediate point whose variance is greater than a predetermined threshold is changed to a first intermediate point 72d.
[0105] Figure 15 is a flowchart showing an example of the procedure for the pathfinding process in the fourth embodiment. The process shown in Figure 15 will be described below in accordance with the step numbers. [Step S311] The pathfinding unit 140 sets the first intermediate points to be optimized at equal intervals between the fixed points corresponding to the start point and the end point.
[0106] [Step S312] The path search unit 140 sets a number of second intermediate points on each sub-path divided by the first intermediate point, corresponding to the length of the sub-path. [Step S313] The path search unit 140 sets the loop variable i to an initial value of "1".
[0107] [Step S314] The pathfinding unit 140 determines whether the loop variable i is less than or equal to n (where n is a predetermined natural number). If i is less than or equal to n (i ≤ n), the pathfinding unit 140 proceeds to step S315. If i is greater than n (i > n), the pathfinding unit 140 terminates the pathfinding process.
[0108] [Step S315] The pathfinding unit 140 calculates the objective function based on the positions of the first intermediate point and the second intermediate point. [Step S316] The pathfinding unit 140 determines whether the convergence condition for the i-th optimization process has been satisfied. The convergence condition can be made stricter (the threshold value can be made smaller) as the value of i increases, for example. This makes it possible to gradually improve the accuracy of the pathfinding. If the convergence condition is satisfied, the pathfinding unit 140 proceeds to step S319 of the pathfinding process. If the convergence condition is not satisfied, the pathfinding unit 140 proceeds to step S317 of the process.
[0109] [Step S317] The pathfinding unit 140 moves the first intermediate point in a direction that minimizes the value of the objective function. [Step S318] The pathfinding unit 140 increases or decreases the number of second intermediate points for each sub-path according to the distance between the first intermediate points after the move through optimization, and determines the position of the second intermediate point for each sub-path. After that, the pathfinding unit 140 proceeds to step S213.
[0110] [Step S319] The pathfinding unit 140 calculates the variance for each of the second intermediate points. The pathfinding unit 140 converts the second intermediate points whose variance is greater than or equal to a predetermined value into first intermediate points.
[0111] [Step S320] The path search unit 140 adds 1 to the loop variable i (i = i + 1). Then, the path search unit 140 proceeds to step S314. In this way, the second intermediate point, which has a large influence on the objective function, is changed to the first intermediate point, and the position of the first intermediate point is optimized. This makes it possible to search for a path that represents a structural change of the protein with high accuracy.
[0112] [Fifth Embodiment] In the fifth embodiment, after the convergence conditions are satisfied, the first intermediate point is converted to a fixed point, and the optimization process is performed again.
[0113] Figure 16 shows an example of the intermediate point setting process in the fifth embodiment. Multiple first intermediate points 72a to 72c and multiple second intermediate points 73a to 73g are set on the path 70 connecting the fixed points 71a and 71b. When the convergence condition is met, the first intermediate points 72a to 72c are converted to fixed points 71c to 71e. Also, the second intermediate points 73a to 73g are converted to first intermediate points 72d to 72j. As a result of the conversion of the second intermediate points 73a to 73g to first intermediate points 72d to 72j, the path 70 is divided into partial paths with the fixed points 71c to 71e and the first intermediate points 72d to 72j as boundaries. Then, new second intermediate points 73h to 73y are set on the partial paths.
[0114] Figure 17 is a flowchart showing an example of the procedure for the pathfinding process in the fifth embodiment. The process shown in Figure 17 will be described below in accordance with the step numbers. [Step S411] The pathfinding unit 140 sets the first intermediate points to be optimized at equal intervals between the fixed points corresponding to the start point and the end point.
[0115] [Step S412] The pathfinding unit 140 sets the loop variable i to an initial value of "1". [Step S413] The pathfinding unit 140 determines whether the loop variable i is less than or equal to n. If i is less than or equal to n (i ≤ n), the pathfinding unit 140 proceeds to step S414. If i is greater than n (i > n), the pathfinding unit 140 terminates the pathfinding process.
[0116] [Step S414] The pathfinding unit 140 sets a number of second intermediate points on each sub-path divided by the first intermediate point, corresponding to the length of the sub-path. [Step S415] The pathfinding unit 140 calculates the objective function based on the positions of the first intermediate point and the second intermediate point.
[0117] [Step S416] The pathfinding unit 140 determines whether the convergence condition for the i-th optimization process has been satisfied. If the convergence condition is satisfied, the pathfinding unit 140 proceeds to step S419. If the convergence condition is not satisfied, the pathfinding unit 140 proceeds to step S417.
[0118] [Step S417] The pathfinding unit 140 moves the first intermediate point in a direction that minimizes the value of the objective function. [Step S418] The pathfinding unit 140 increases or decreases the number of second intermediate points for each sub-path according to the distance between the first intermediate points after the move through optimization, and determines the position of the second intermediate point for each sub-path. After that, the pathfinding unit 140 proceeds to step S415.
[0119] [Step S419] The pathfinding unit 140 converts the first intermediate point into a fixed point. [Step S420] The pathfinding unit 140 converts the second intermediate point into the first intermediate point. [Step S421] The pathfinding unit 140 adds 1 to the loop variable i (i = i + 1). After that, the pathfinding unit 140 proceeds to step S413.
[0120] In this way, each time the convergence condition is satisfied, the first intermediate point is converted to a fixed point and the second intermediate point is converted to the first intermediate point, and the process of optimizing the position of the first intermediate point is performed again. By repeating this process a predetermined number of times, it is possible to start with a wide interval between the first intermediate points to perform a rough path search, and then gradually perform a more detailed path search. As a result, highly accurate path searches can be performed efficiently.
[0121] [Sixth Embodiment] The sixth embodiment is an evaluation value ((1 / P) that indicates the degree of contribution to MaxFluxObjective. ψ )||z k+1 -z k This method converts the first or second midpoint with a high value (||) into a fixed point, and converts the second midpoint with a low evaluation value into the first midpoint.
[0122] Figure 18 shows an example of the intermediate point setting process in the sixth embodiment. Multiple first intermediate points 72a to 72c and multiple second intermediate points 73a to 73g are set on the path 70 connecting the fixed points 71a and 71b. When the convergence condition is met, an evaluation value is calculated for each of the first intermediate points 72a to 72c and the second intermediate points 73a to 73g. Intermediate points with high evaluation values (second intermediate point 73b, second intermediate point 73d, first intermediate point 72c) are converted to fixed points 71c to 71e. Second intermediate points 73a, 73c, 73e, 73f, and 73g with low evaluation values are converted to first intermediate points 72d to 72h. As the second intermediate points 73a, 73c, 73e, 73f, and 73g are transformed into the first intermediate points 72d to 72h, the path 70 is divided into sub-paths bounded by the fixed points 71c to 71e and the first intermediate points 72d to 72h. Then, new second intermediate points 73h to 73q are set on the sub-paths.
[0123] Figure 19 is a flowchart showing an example of the procedure for the pathfinding process in the sixth embodiment. The processes in steps S511 to S518 and S521 in Figure 19 are the same as steps S411 to S418 and S421 in the flowchart of the fifth embodiment shown in Figure 17. The processes in steps S519 and S520, which differ from those in the fifth embodiment, will be described below.
[0124] [Step S519] When the convergence condition is met, the path search unit 140 evaluates the value "(1 / P)" ψ )||z k+1 -z k Convert the first or second midpoint where the value of "||" is high (exceeds the threshold) into a fixed point.
[0125] [Step S520] The route search unit 140 evaluates the value "(1 / P)" ψ )||z k+1 -z k The second intermediate point with a low value (below the threshold) is converted to the first intermediate point. In this way, if the path passes through a region that has a high influence on the objective function, the number of first intermediate points in that region can be increased, improving the accuracy of the path calculation in that region. For example, if the high-energy region (low probability of protein structure existence) is limited, the search can be focused on finding appropriate paths in the low-energy region (high probability of protein structure existence).
[0126] [Seventh Embodiment] The seventh embodiment is an evaluation value ((1 / P) that indicates the degree of contribution to MaxFluxObjective. ψ )||z k+1 -z k This method converts the first or second midpoint with a low value (||) into a fixed point, and converts the second midpoint with a high evaluation value into the first midpoint.
[0127] Figure 20 shows an example of the intermediate point setting process in the seventh embodiment. Multiple first intermediate points 72a to 72c and multiple second intermediate points 73a to 73g are set on a path 70 connecting fixed points 71a and 71b. When the convergence condition is met, evaluation values are calculated for each of the first intermediate points 72a to 72c and the second intermediate points 73a to 73g. Intermediate points with low evaluation values (second intermediate point 73a, first intermediate points 72a and 72b, second intermediate points 73c, 73e, 73f, and 73g) are converted to fixed points 71c to 71i. Second intermediate points 73b and 73d with high evaluation values are converted to first intermediate points 72d and 72e. As the second intermediate points 73b and 73d are transformed into the first intermediate points 72d and 72e, the path 70 is divided into sub-paths bounded by the fixed points 71c to 71i and the first intermediate points 72c, 72d, and 72e. Then, new second intermediate points 73h to 73n are set on the sub-path between the fixed points and the first intermediate points, or on the sub-path between the two first intermediate points.
[0128] Figure 21 is a flowchart showing an example of the procedure for the pathfinding process in the seventh embodiment. The processes in steps S611 to S618 and S621 in Figure 21 are the same as steps S411 to S418 and S421 in the flowchart of the fifth embodiment shown in Figure 17. The processes in steps S619 and S620, which differ from those in the fifth embodiment, will be described below.
[0129] [Step S619] When the convergence condition is met, the path search unit 140 evaluates the value "(1 / P)" ψ )||z k+1 -z k Convert the first or second midpoint with a low value (below the threshold) of || into a fixed point.
[0130] [Step S620] The route search unit 140 calculates the evaluation value "(1 / P) ψ )||z k+1 -z kThe second intermediate point with a high value (exceeding the threshold) is converted to the first intermediate point. In this way, if the path passes through a region that has little influence on the objective function, the number of first intermediate points in that region can be increased, improving the accuracy of the path calculation in that region. For example, if the low-energy region (where the probability of finding the protein structure is high) is limited, the search can be focused on finding an appropriate path in the high-energy region (where the probability of finding the protein structure is low).
[0131] [Eighth Embodiment] In the eighth embodiment, the first intermediate point and the second intermediate point are reset to equal intervals each time the convergence condition is met, and the position of the first intermediate point is repeatedly optimized. That is, even if the first intermediate point is placed at equal intervals in the initial state, the interval of the first intermediate point becomes uneven after the optimization process. If there are places where the interval of the first intermediate point is too wide, the accuracy of the path at those places may decrease. Also, if there are places where the interval of the first intermediate point is too narrow, the computational efficiency at those places may decrease. Therefore, in the eighth embodiment, the rearrangement of the first intermediate point and the second intermediate point is repeated a predetermined number of times so that the interval is uniform, thereby suppressing partial decreases in accuracy and achieving a predetermined level of efficiency.
[0132] Figure 22 shows an example of the intermediate point setting process in the eighth embodiment. Multiple first intermediate points 72a to 72c and multiple second intermediate points 73a to 73g are set on the path 70 connecting the fixed points 71a and 71b. When the convergence condition is met, the first intermediate points and second intermediate points are reset to equal intervals on the converged path 70.
[0133] In the example shown in Figure 22, three first intermediate points 72a to 72c are set at equal intervals between fixed points 71a and 71b. Second intermediate points 73a and 73b are set at equal intervals between fixed point 71a and first intermediate point 72a. Second intermediate points 73c and 73d are set at equal intervals between first intermediate point 72a and first intermediate point 72b. Second intermediate points 73e and 73f are set at equal intervals between first intermediate point 72b and first intermediate point 72c. Second intermediate points 73g and 73h are set at equal intervals between first intermediate point 72c and fixed point 71b.
[0134] Figure 23 is a flowchart showing an example of the procedure for the pathfinding process in the eighth embodiment. The processes in steps S711 to S718 and S720 in Figure 23 are the same as steps S411 to S418 and S421 in the flowchart of the fifth embodiment shown in Figure 17. The process in step S719, which differs from that of the fifth embodiment, will be described below.
[0135] [Step S719] When the convergence condition is met, the path search unit 140 resets the first intermediate point and the second intermediate point on the path at equal intervals. For example, when the convergence condition is met, the path search unit 140 generates a smooth curve that passes through the first intermediate point and connects the two fixed points, and this becomes a path representing the structure of the protein. Next, the path search unit 140 deletes the existing first intermediate point and the second intermediate point. Furthermore, the path search unit 140 divides the generated path into a predetermined number of subpathways and sets the first intermediate point at the boundary of the subpathways. Then, the path search unit 140 sets the second subpathways at the points that divide each subpathway into a predetermined number of equal parts.
[0136] In this way, each time the convergence condition is met, the first and second intermediate points are reset to equal intervals. This allows for efficient pathfinding with accuracy free from location bias.
[0137] [Other Embodiments] The second to eighth embodiments showed examples of analyzing the process of shape change of proteins, but if the structure can be represented as a point in latent space, the process of shape change of other substances can be analyzed in a similar manner.
[0138] The above merely illustrates the principle of the present invention. Furthermore, numerous modifications and changes are possible for those skilled in the art, and the present invention is not limited to the exact configurations and applications shown and described above. All corresponding modifications and equivalents are considered to be within the scope of the present invention as defined by the appended claims and their equivalents.
[0139] 1 Autoencoder 2a-2j Structure 3 Latent space 3a, 3b First point 4 Path 4a-4c Partial path 5a, 5b Second point 6a-6f Third point 10 Information processing device 11 Storage unit 12 Processing unit
Claims
1. An information processing program that causes a computer to perform the following steps:
1. In a latent space corresponding to the possible structures of the object, set one or more second points on a path connecting a plurality of first points corresponding to each of a plurality of first structures of the object; set one or more third points on a subpath obtained by dividing the path at the positions of the second points; calculate an objective function based on information regarding the positions of the second and third points, relating to whether the deformation process of the object from one of the first structures to the other results in structures corresponding to the second and third points; and move the second points in a direction that optimizes the objective function.
2. The information processing program according to claim 1, which causes the computer to further perform the process of moving the third point onto the path which has been transformed to pass through the second point after the move.
3. The information processing program according to claim 2, which causes a computer to further perform the process of calculating the objective function, moving the second point, and moving the third point, until the convergence condition is satisfied.
4. The information processing program according to claim 3, which, if the convergence condition is not satisfied, causes the computer to perform a further process before the next calculation of the objective function, which involves determining the number of third points to set according to the length of the partial path based on the second point after movement, and adding or deleting the third points on the partial path according to the determined number.
5. The information processing program according to claim 3, wherein, if the convergence condition is satisfied, the program causes the computer to further perform the following steps: change the third point to the second point if the amount of change in the objective function when its position is shifted by a predetermined amount satisfies a predetermined condition, and repeat the process of calculating the objective function, moving the second point, and moving the third point until the convergence condition is satisfied again.
6. The information processing program according to claim 3, wherein, if the convergence condition is satisfied, the program causes the computer to further perform the following operations: changing the second point to the first point, changing the third point to the second point, dividing the path at the positions of the first point and the second point to obtain the third point in the resulting partial path, and repeating the process of calculating the objective function, moving the second point, and moving the third point until the convergence condition is satisfied again.
7. The information processing program according to claim 3, wherein in the process of calculating the objective function, for each of the second points and each of the third points, an evaluation value indicating whether a corresponding structure is formed is calculated using positional information, the sum of the evaluation values for each of the second points and each of the third points is taken as the value of the objective function, and if the convergence condition is satisfied, the second point or third point whose evaluation value satisfies a predetermined condition is converted to the first point, the third point whose evaluation value does not satisfy the predetermined condition is converted to the second point, the third point is set in the sub-path obtained by dividing the path at the positions of the first point and the second point, and the process of calculating the objective function, the process of moving the second point and the process of moving the third point are repeated until the convergence condition is satisfied again.
8. The information processing program according to claim 3, which, if the convergence condition is satisfied, causes the computer to further perform the following processes: delete the existing second point and third point from the path, set one or more second points at equal intervals on the path, and set one or more third points at equal intervals on the partial path obtained by dividing the path at the positions of the second points.
9. An information processing method performed by a computer, comprising: setting one or more second points on a path connecting a plurality of first points corresponding to each of a plurality of first structures of an object in a latent space corresponding to the possible structures of the object; setting one or more third points on a subpath obtained by dividing the path at the positions of the second points; calculating an objective function relating to whether the structure corresponding to the second and third points is formed during the deformation process of the object from one of the first structures to the other, based on information regarding the positions of the second and third points; and moving the second points in a direction in which the objective function is optimized.
10. An information processing device having: one or more second points on a path connecting a plurality of first points corresponding to each of a plurality of first structures of the object in a latent space corresponding to the possible structures of the object; one or more third points on a subpath obtained by dividing the path at the positions of the second points; an objective function relating to whether the deformation process of the object from one of the first structures to the other results in structures corresponding to the second and third points, based on information regarding the positions of the second and third points; and a processing unit that moves the second points in a direction in which the objective function is optimized.