Information processing apparatus, information processing system, and information processing method

By using the structure prediction and eigenvalue prediction units of the information processing device and system, the problem of low efficiency in compound design and evaluation has been solved, and efficient and accurate compound design and evaluation have been achieved.

CN121336264APending Publication Date: 2026-01-13SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480039721.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-21
Filing Date
2024-06-07
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing technologies require significant time, effort, and cost to design and evaluate compounds, making it difficult to achieve high-throughput cyclic design and evaluation of desired compounds.

Method used

An information processing device and system are used to predict the three-dimensional structure and characteristic values ​​of a compound based on its sequence information through a structure prediction unit and a characteristic value prediction unit, and then present the information to the user through a presentation unit, thereby achieving efficient compound design.

Benefits of technology

It improves the efficiency and accuracy of compound design, reduces time and cost, and enables high-throughput cyclic design and evaluation of desired compounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121336264A_ABST
    Figure CN121336264A_ABST
Patent Text Reader

Abstract

An information processing apparatus according to an embodiment of the present disclosure includes: a structure prediction unit that predicts a stereostructure of a candidate compound based on sequence information of the candidate compound and obtains stereostructure information of the candidate compound; a feature value prediction unit that predicts a feature value of the candidate compound on the basis of the sequence information and the stereostructure information on the candidate compound, and obtains feature value information on the candidate compound; and a presentation unit that presents the feature value information and the stereostructure information on the candidate compound to a user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to an information processing apparatus, an information processing system, and an information processing method. BACKGROUND

[0002] In fields such as drug discovery, biological material development, and food development, high-throughput cycles are required to design and evaluate new compounds, i.e., desired compounds, having desired characteristics and functions. However, when designing desired compounds, trial and error needs to be repeated to actually synthesize and analyze the compounds, which requires a large amount of time, effort, and cost.

[0003] Therefore, with the development of image recognition and natural language processing deep learning techniques in recent years, machine learning has been proposed to solve the above problems in the field of bioinformatics. For example, PTL 1 proposes a method of predicting the three-dimensional structure of a protein using a neural network with an amino acid sequence as input.

[0004] PRIOR ART DOCUMENTS

[0005] PATENT LITERATURE

[0006] PTL 1: US 2021 / 0166779 A

[0007] NON-PATENT LITERATURE

[0008] NPL 1: John Jumper et al., "Highly Accurate Protein Structure Prediction with AlphaFold", Nature, 26 August 2021.

[0009] NPL 2: Minkyung Beak et al., "Accurate Prediction of Protein Structures and Interactions using a Three-track Neural Network", Science, 15 July 2021 (First Release).

[0010] NPL 3: Hannes Stark et al., "EquiBind: Geometric Deep Learning for Drug Binding Structure Prediction", Proceedings of the 39th International Conference on Machine Learning, 4 June 2022. SUMMARY

[0011] TECHNICAL PROBLEM

[0012] However, there is still a need for a technology capable of efficiently producing a desired compound by high-throughput cyclic design and evaluation of a required compound.

[0013] Therefore, the present disclosure proposes an information processing apparatus, an information processing system, and an information processing method capable of efficiently producing a desired compound.

[0014] SOLUTION TO PROBLEM

[0015] The information processing apparatus according to an aspect of the present disclosure includes a structure prediction unit configured to predict a three-dimensional structure of a candidate compound based on sequence information about the candidate compound, and obtain three-dimensional structure information about the candidate compound; a feature value prediction unit configured to predict a feature value of the candidate compound based on the sequence information and the three-dimensional structure information about the candidate compound, and obtain feature value information about the candidate compound; and a presentation section configured to present the feature value information and the three-dimensional structure information about the candidate compound to a user.

[0016] The information processing system according to an aspect of the present disclosure includes a structure prediction unit configured to predict a three-dimensional structure of a candidate compound based on sequence information about the candidate compound, and obtain three-dimensional structure information about the candidate compound; a feature value prediction unit configured to predict a feature value of the candidate compound based on the sequence information and the three-dimensional structure information about the candidate compound, and obtain feature value information about the candidate compound; and a presentation section configured to present the feature value information and the three-dimensional structure information about the candidate compound to a user.

[0017] An information processing method according to an aspect of the present disclosure includes: predicting, by an information processing apparatus, a stereo structure of a candidate compound based on sequence information about the candidate compound and obtaining stereo structure information about the candidate compound, predicting a characteristic value of the candidate compound based on the sequence information and the stereo structure information about the candidate compound, and obtaining characteristic value information about the candidate compound; and presenting the characteristic value information and the stereo structure information about the candidate compound to a user. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 is a diagram illustrating a configuration example of an information processing system according to a first embodiment.

[0019] Figure 2 is a diagram illustrating a configuration example of an information processing apparatus according to the first embodiment.

[0020] Figure 3 is a diagram illustrating an example of displaying stereo structure information and characteristic value information about a compound according to the first embodiment.

[0021] Figure 4 is a flowchart illustrating a process in an example of a compound design process according to the first embodiment.

[0022] Figure 5 is a diagram for describing a structure / sequence prediction AI (Stage 1) according to the first embodiment.

[0023] Figure 6 is a diagram for describing a structure / sequence prediction AI (Stage 2) according to the first embodiment.

[0024] Figure 7 is a diagram for describing a characteristic value prediction AI (Stage 1) according to the first embodiment.

[0025] Figure 8 is a diagram for describing a characteristic value prediction AI (Stage 2) according to the first embodiment.

[0026] Figure 9 is a diagram for describing a characteristic value prediction AI (Stage 3) according to the first embodiment.

[0027] Figure 10 is a diagram illustrating a configuration example of an information processing apparatus according to a second embodiment.

[0028] Figure 11 is a flowchart illustrating a process in an example of a compound design process according to the second embodiment.

[0029] Figure 12is a diagram for describing the interactive prediction AI (Stage 1) according to the second embodiment.

[0030] Figure 13 is a diagram for describing the interactive prediction AI (Stage 2) according to the second embodiment.

[0031] Figure 14 is a diagram showing a configuration example of an information processing apparatus according to the third embodiment.

[0032] Figure 15 is a diagram for describing the calculation of the prediction reliability according to the third embodiment.

[0033] Figure 16 is a diagram showing a hardware configuration example of a computer. DETAILED DESCRIPTION

[0034] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Note that the embodiments are not limited to systems, apparatuses, methods, and the like according to the present disclosure. In the following embodiments, substantially the same elements are denoted by the same reference numerals, and overlapping descriptions will be omitted. The following embodiments include examples, modifications, and the like.

[0035] One or more of the following embodiments can each be realized independently. On the other hand, at least a part of each of the following multiple embodiments can be realized in combination with at least a part of another embodiment as appropriate. These multiple embodiments can include novel features different from each other. Therefore, each of the embodiments can serve different purposes or solve different problems and can achieve different effects. Note that the effects in the embodiments are merely examples and are not limited thereto, and other effects can be provided.

[0036] The present disclosure will be described in the following order.

[0037] 1. First Embodiment

[0038] 1-1. Configuration Example of Information Processing System

[0039] 1-2. Configuration Example of Information Processing Apparatus

[0040] 1-3. Example of Compound Design Process

[0041] 1-4. Learning Method of Neural Network

[0042] 2. Second Embodiment

[0043] 2-1. Configuration Example of Information Processing Apparatus

[0044] 2-2. Example of Compound Design Process

[0045] 2-3. Learning method of neural network

[0046] 3. Third embodiment

[0047] 3-1. Configuration example of information processing apparatus

[0048] 3-2. Learning method of neural network

[0049] 4. Operation and effects

[0050] 5. Other embodiments

[0051] 6. Hardware configuration example

[0052] 7. Supplementary explanation

[0053] 1. First embodiment

[0054] 1-1. Configuration example of information processing system

[0055] A configuration example of an information processing system 1 according to a first embodiment will be described with reference to Figure 1 A configuration example of an information processing system 1 according to a first embodiment will be described with reference to Figure 1 is a diagram illustrating a configuration example of the information processing system 1 according to the first embodiment. The information processing system 1 according to the first embodiment functions as a compound design system.

[0056] As shown in Figure 1 The information processing system 1 of the first embodiment has an information processing apparatus 10, a sequence structure information DB 20, and a feature information DB 30. The information processing apparatus 10, the sequence structure information DB 20, and the feature information DB 30 are configured to be able to communicate (transmit and receive) various information via a network 40.

[0057] The information processing apparatus 10 is an apparatus for designing a compound such as an organic compound, which is used by a user such as a designer of the designed compound. For example, the information processing apparatus 10 performs a process of designing a compound using machine learning such as a neural network (for example, various machine learning models). For example, the information processing apparatus 10 generates various machine learning models for designing a compound. At the time of generating the machine learning models, the information processing apparatus 10 uses various types of information stored in the sequence / structure information DB 20 and the feature information DB 30. The process of designing a compound and the generation of the machine learning models will be described in detail below.

[0058] The information processing apparatus 10 described above can be any one of a local PC, a physical server, a cloud server, or the like, for example. As a local PC, a desktop personal computer (PC), a notebook PC, a tablet device, or a smart phone can be used, for example. The information processing apparatus 10 is capable of accessing the sequence / structure information DB 20, the feature information DB 30, and the like via the network 40.

[0059] The sequence / structure information DB 20 records and manages sequence information and stereostructure information of each of various compounds. The sequence / structure information DB 20 can be, for example, a locally deployed database or a cloud-based database. As the sequence / structure information DB 20, a database of the Worldwide Protein Database (wwPDB), for example, can be used. The sequence information and the stereostructure information on the sequence / structure information DB 20 are used as template information of sequences and stereostructures of compounds.

[0060] The sequence information on a compound is information on a sequence of the compound. The sequence information on a compound includes, for example, an alphabetical string representing an order. The stereostructure information on a compound is information on a stereostructure of the compound (e.g., an inherent stereostructure and a function of the compound). The sequence information and the stereostructure information on a compound are paired information for one kind of compound.

[0061] The feature information DB 30 records and manages feature information of each of various compounds. Similar to the sequence / structure information DB 20, the feature information DB 30 can be, for example, a locally deployed database or a cloud-based database. The feature information of the feature information DB 30 is used as template information of compound features.

[0062] The feature information on a compound is information on a feature of the compound. The feature information on a compound includes, for example, feature value information and experimental condition information. The feature value information is information on a feature value of a compound. The feature value information includes, for example, various feature values corresponding to the compound. Examples of the various feature values include a denaturation temperature, a solubility, and a pH feature. The experimental condition information is information on an experimental condition of production of the compound. The experimental condition of compound manufacturing is a condition of an experiment of compound manufacturing, and includes, for example, various condition values related to the experiment. Examples of the various condition values include a temperature and a humidity during the experiment, an experimental step, a droplet amount, and a possibility of contamination by a foreign matter.

[0063] The network 40 can be wireless and / or wired, and can be a dedicated line and / or a public line. The network 40 can be any network that allows the information processing apparatus 10, the sequence / structure information DB 20, and the feature information DB 30 to communicate with each other.

[0064] The network 40 can be a communication network such as a LAN (Local Area Network), a WAN (Wide Area Network), a cellular network, a fixed telephone network, a regional IP (Internet Protocol) network, or the Internet, for example. The network 40 can include a core network. The core network is an EPC (Evolved Packet Core), for example, or a 5GC (5G Core Network). The network 40 can include a data network other than the core network. The data network can be a service network of a telecommunications carrier, for example, an IMS (IP Multimedia Subsystem) network. The data network can be a private network such as an intranet. The network 40 can be implemented as an SDN (Software Defined Network).

[0065] Note that, in the embodiment of the information processing apparatus 10, Figure 1 In the embodiment of the information processing apparatus 10, Figure 1 In the embodiment of the information processing apparatus 10, there is one sequence / structure information DB 20 and one feature information DB 30, but there can be multiple DBs for both.

[0066] The generation of the machine learning model is performed by the information processing apparatus 10, but can be performed by an apparatus different from the information processing apparatus 10, for example. In this case, the machine learning model generated by the other device is stored in a storage device such as a server, for example, and used by one or more information processing apparatuses 10.

[0067] Embodiments of the compound

[0068] Here, the embodiment of the compound is a protein. A protein is formed of several tens of types of amino acids connected in a line in different orders, for example. The characteristics (functions) of a protein are exhibited not only by the amino acid sequence (primary structure) but also by forming a specific stereo structure.

[0069] The amino acid sequence is usually a sequence of several tens to several hundreds of amino acid residues. When these amino acid residues are written by a rational formula or the like, the rational formulas are very redundant. Therefore, in order to write the amino acid sequence simply, a method in which the type of the amino acid residue is written by a single letter of the alphabet is used. For example, a serine residue is written by the letter "S" and a glutamine residue is written by the letter "Q". Furthermore, each of the 20 types of amino acid residues is written by a single letter of the alphabet. Such a string of letters is sequence information, for example. Of course, the sequence information can include any other information on the amino acid sequence.

[0070] The stereo structure information is information including, for example, the stereo structure inherent to a protein and the function. The stereo structure information includes at least one of the protein structure or the protein function.

[0071] Protein structure is information about the structure of a protein. For example, protein structure includes a sequence of coordinates containing the three-dimensional coordinates of each of the atoms, molecules, bonds, functional groups, etc., that make up the protein. This sequence of three-dimensional coordinates can be called volumetric data. Of course, protein structure can include any information about the structure of a protein.

[0072] Protein function includes at least one of, for example, hydrophilicity or rigidity. Some proteins exhibit localized hydrophilicity within a portion of their structure. Other proteins exhibit localized rigidity (resistance to bending). Functional markers representing such hydrophilicity or rigidity are included in protein function, for example. A functional marker is a numerical value, for example, representing the range of three-dimensional coordinates indicating hydrophilicity or rigidity and the level of hydrophilicity or rigidity. Conversely, a functional marker may include numerical values ​​representing the range of three-dimensional coordinates indicating hydrophobicity or non-rigidity. Furthermore, when a protein locally possesses a Y-shaped structure, the arms of the "Y" may express a function of capturing viruses. Functional markers representing this immune function may be included in protein function. Of course, protein function may include any information about the function of the protein.

[0073] 1-2. Configuration Examples of Information Processing Devices

[0074] Reference Figure 2 and Figure 3 An example of the configuration of the information processing device 10 according to the first embodiment will be described. Figure 2 This is a diagram illustrating an example of the structure of the information processing apparatus 10 according to the first embodiment. Figure 3 This is a diagram illustrating an embodiment according to the first embodiment, showing information about the three-dimensional structure and characteristic values ​​of the compound.

[0075] like Figure 2 As shown, the information processing device 10 of the first embodiment includes a structure prediction unit 11, an eigenvalue prediction unit 12, a selection unit 13, and a sequence prediction unit 14.

[0076] The structure prediction unit 11 receives sequence information (compound sequence information A1) of candidate compounds (candidate compounds) for design, predicts the stereostructure of the compound from the sequence based on the sequence information, obtains stereostructure information about the predicted stereostructure, and sends the stereostructure information to the feature value prediction unit 12. For example, the structure prediction unit 11 receives a sequence in which the molecules of the compound are represented by strings as input, and predicts the stereostructure using a neural network (e.g., a machine learning model). Then, for example, the structure prediction unit 11 sends the stereostructure information representing the predicted stereostructure of the compound as a sequence of atomic stereo coordinates to the feature value prediction unit 12.

[0077] The neural network serving as the structure prediction unit 11 can utilize methods already published in papers or similar publications. For example, as a neural network configuration, AlphaFold, disclosed in NPL 1, or RoseTTAFold, disclosed in NPL 2, can be used. However, the neural network configuration is not limited to these.

[0078] It should be noted that when the candidate compound for design is a protein, the sequence information is the amino acid sequence, but is not limited to this. That is, the sequence information is not limited to amino acids and can also be a string representing the molecular formula of nucleic acids, polymers, etc. As sequence information, the sequence of a known compound can be used as a template, or a new string prepared by the designer can be used.

[0079] The eigenvalue prediction unit 12 predicts various eigenvalues ​​of the compound based on its sequence and stereostructure information, obtains eigenvalue information about the predicted eigenvalues, and sends the eigenvalue information to the selection unit 13. For example, the eigenvalue prediction unit 12 receives sequence and stereostructure information about the compound as input and uses a neural network (e.g., a machine learning model) to predict the chemical or physical eigenvalues ​​of the compound.

[0080] The neural network of the eigenvalue prediction unit 12 is, for example, a multi-head neural network with multiple output layers that predicts multiple eigenvalues. However, the configuration of the neural network is not limited to this. For example, denaturation temperature, solubility, pH characteristics, biosynthetic cost (molecular synthesis cost), and torsional stiffness can be used as eigenvalues. For example, active sites indicating which part of the compound is the active center can also be used as eigenvalues.

[0081] The selection unit 13 includes a presentation unit 13a, an input unit 13b, and a processing unit 13c. The selection unit 13 presents the predicted stereostructure information and characteristic value information of the compound to the designer through the presentation unit 13a, receives input operations from the designer through the input unit 13b, and selects the compound being presented by the processing unit 13c as the desired compound or the next candidate compound in response to the received input operations (requests) from the designer.

[0082] The presentation unit 13a visualizes and presents the predicted stereostructure and characteristic value information of the compound to the designer. The presentation unit 13a can be any type of display unit, printer, etc. As a display unit, for example, a liquid crystal display, an organic EL (electroluminescent) display, or a head-mounted display can be used. Multiple display units can be provided depending on the application.

[0083] Input unit 13b receives input operations from users such as designers. Input unit 13b can be any of, for example, a keyboard and mouse, operation keys, a touch panel, etc. When using a touch panel, the user performs various operations by touching the touch panel with a finger or stylus. Input unit 13b can be a voice input device (e.g., a microphone) that receives input operations via the user's voice.

[0084] In response to a user's input operation to the input unit 13b, the processing unit 13c determines the compound (candidate compound) presented by the presentation unit 13a as the desired compound. In response to a user's input operation to the input unit 13b, the processing unit 13c changes the stereoscopic structure information of the compound (candidate compound) presented by the presentation unit 13a, and determines the compound whose stereoscopic structure has been changed as the next candidate compound.

[0085] like Figure 3 As shown, the presentation unit 13a, for example, is a display unit that visualizes the three-dimensional structure information predicted by the structure prediction unit 11 as a three-dimensional structure image A2, and visualizes the feature value information predicted by the feature value prediction unit 12 as a table A3. The three-dimensional structure image A2 and the table A3 are then presented to the designer, who is the user. The three-dimensional structure image A2 represents the three-dimensional structure, and the table A3 represents various feature values. This allows the designer to visually evaluate candidate compounds on the information processing device 10 based on the predicted three-dimensional structure information and various feature values.

[0086] For example, when a designer views the stereoscopic structure image A2 and table A3 presented by the presentation unit 13a and determines that the candidate compound being displayed is the desired compound, the designer operates the input unit 13b (such as the keyboard 131 and mouse 132) to select the candidate compound being displayed as the desired compound. This allows the designer to obtain information such as stereoscopic structure information and sequence information about the desired compound.

[0087] For example, when the designer determines that the displayed candidate compound is not the desired compound, the designer operates the input unit 13b to edit the stereostructure on the display screen of the presentation unit 13a. While viewing the predicted stereostructure, the designer can operate the input unit 13b to deform the stereostructure, for example, by bending a portion of the stereostructure. At this time, the processing unit 13c updates the stereo coordinate sequence of the atoms according to the edited stereostructure. When the designer completes the editing of the stereostructure, the processing unit 13c sends information about the next candidate stereostructure representing the updated stereo coordinate sequence of the atoms to the sequence prediction unit 14.

[0088] return Figure 2The sequence prediction unit 14 predicts the sequence of a compound based on the stereoscopic structural information of the compound selected by the selection unit 13 (the next candidate compound), obtains sequence information about the predicted sequence, and sends the sequence information to the structure prediction unit 11. For example, the sequence prediction unit 14 converts the stereoscopic structural information of the compound selected by the selection unit 13 into sequence information about the compound. For example, the sequence prediction unit 14 receives the stereoscopic coordinate sequence of atoms as input and uses a neural network (e.g., a machine learning model) to convert the stereoscopic coordinate sequence of atoms into the molecular sequence of the compound.

[0089] The neural network serving as sequence prediction unit 14 can, for example, be a graph transformer or graph autoencoder that represents the sequence of three-dimensional coordinates of atoms as an inter-atom graph and converts the inter-atom graph into a sequence. However, the neural network configuration is not limited to this.

[0090] The sequence information predicted by sequence prediction unit 14 is again input to structure prediction unit 11 and eigenvalue prediction unit 12, and selection unit 13 presents to the designer whether the results edited by the designer (e.g., stereostructure information and eigenvalue information) are expected. As described above, for example, when the candidate compound is the desired compound, the designer operates input unit 13b to select the candidate compound as the desired compound. This allows the designer to obtain information such as stereostructure information and sequence information about the desired compound.

[0091] 3D structural images

[0092] Here, the aforementioned stereoscopic structure image A2 can be, for example, a point cloud image, a polygon image, a mesh image, a surface image, a slice image, a stereoscopic view, etc. The specific display format of the stereoscopic structure image A2 is not limited.

[0093] A point cloud image is an image in which data is represented by a set of points. For example, each atom in a protein is represented as a point, and the protein is displayed as a point cloud image. Specifically, the positions of points in the point cloud image are calculated based on the three-dimensional coordinates of the atoms contained in the three-dimensional structure, and the point cloud image is generated. Note that not only atoms, but also molecules, functional groups, functional markers, or the main chain and side chains of a protein can be represented by points, and the protein is displayed as a point cloud image. Alternatively, points may be displayed in different colors depending on the type of atom or functional marker. It should be noted that a point cloud may be referred to simply as a point cloud.

[0094] A polygonal image is an image in which data is represented by polygons. For example, the local shape of a protein is represented by triangles or quadrilaterals. A mesh image is an image in which data is represented by multiple polygons. For example, the shape of a protein is represented by a shape obtained by connecting triangles and quadrilaterals. A mesh image can be thought of as a collection of polygonal images. A surface image is an image in which data is represented by smooth, curved surfaces. For example, the shape of a protein is represented by a smooth, curved surface.

[0095] A slice image is an image representing a cross-section of a protein. For example, a cross-section at a predetermined location in a point cloud image is displayed as a slice image. Alternatively, a cross-section of a polygonal image, a mesh image, or a surface image can be displayed. A three-view drawing is an image representing the shape of a protein when viewed from three directions. For example, a three-view drawing may include views viewed from any direction, such as a front view, a top view, a bottom view, a right side view, a left side view, and a rear view, where a predetermined surface of the protein is the front view.

[0096] By displaying the three-dimensional structural image A2 in these display formats, users can intuitively identify protein structures, etc. The slice images allow users to easily identify the internal structure of proteins (structures not visible from the outside). It should be noted that, for example, users can appropriately change the display format, the position of the cross-section in the slice image, the orientation of the three-dimensional image, etc., via the input unit 13b.

[0097] Editing of three-dimensional structures

[0098] The input operations for editing the aforementioned three-dimensional structures include at least one of, for example, editing protein structures or editing protein functions. Note that to implement these editing operations, any graphical user interface (GUI) can be configured, such as various windows, buttons, checkboxes, labels, and input fields.

[0099] For example, changing the arrangement of atoms is equivalent to editing the protein structure, and operations corresponding to this editing can be performed through dragging operations on the points representing atoms in the 3D structure image A2. Besides changing the arrangement of atoms, editing such as placing new atoms, deleting atoms, selecting atoms, and changing the type of atoms (α-carbon, β-carbon, oxygen, nitrogen, etc.) is possible. Operations corresponding to these editing operations can be performed through clicking operations, dragging operations, etc., on the points representing atoms in the 3D structure image A2. At this time, the input information corresponding to the editing, such as "delete atoms" or "the new type of atom is carbon," is input to the processing unit 13c. Alternatively, similar editing is possible for molecules, functional groups, the main chain and side chains of proteins, etc. In this case, editing such as molecular deformation may be possible.

[0100] Furthermore, atoms can be placed together in a desired area. That is, not only can atoms be precisely positioned at a single point, but also, for example, a method can be used in which a desired area is dragged and atoms are placed on the entire area simultaneously. Similarly, all atoms within an area can be selected, moved, deleted, etc., at once.

[0101] The bonding relationships between atoms can be edited. For example, two atoms can be specified by clicking, a screen for selecting the bond type can be displayed by right-clicking, and the desired type (hydrogen bond, etc.) can be selected using checkboxes, etc. Furthermore, users can specify only the protein's backbone (rough shape), and the detailed arrangement of atoms can be automatically determined based on the specified backbone, etc.

[0102] The assignment of functional markers is, for example, the editing of protein function, and may involve the local assignment of functional markers representing functions such as hydrophilicity, hydrophobicity, rigidity, or non-rigidity. For example, by dragging to select the desired region in the three-dimensional structure image A2, and then using an input operation such as selecting the functional label to be assigned, an operation corresponding to the editing can be performed. At this time, the input information corresponding to the editing, such as "the new function of the functional marker is hydrophilic, and the coordinate range is X=10~20, Y=10~30, Z=20~40," is input to the processing unit 13c.

[0103] For example, once a functional marker is assigned, the arrangement of atoms, etc., is automatically determined based on the assigned functional marker. For instance, when the functional marker "hydrophilic" is assigned to a region, the arrangement of atoms, etc., within that region is automatically determined, thus giving the protein the function of "hydrophilicity" in that region. This allows for the assignment of function even when a user wants to assign a protein a desired function but does not know how to arrange the atoms, etc.

[0104] 1-3. Examples of compound design processes

[0105] Reference Figure 4 Examples describing the compound design process according to the first embodiment. Figure 4 This is a flowchart illustrating a process in an embodiment of the compound design process according to the first embodiment.

[0106] like Figure 4As shown, in step S11, the structure prediction unit 11 predicts the stereostructure of the compound based on the input sequence information about the compound, and obtains stereostructure information about the compound. In step S12, the feature value prediction unit 12 predicts the feature values ​​of the compound based on the input sequence information and the predicted stereostructure information about the compound, and obtains feature value information about the compound. In step S13, the presentation unit 13a presents the stereostructure and feature values ​​of the compound to the designer based on the obtained stereostructure information and feature value information. In step S14, in response to the designer's input operation on the input unit 13b, the processing unit 13c determines whether the compound being displayed (candidate compound) is the desired compound.

[0107] For example, when the compound being displayed is not the desired compound, the designer operates the input unit 13b to change the sequence information to be input to the structure prediction unit 11 to other sequence information, or to edit the stereostructure of the compound being displayed. When the compound being displayed is the desired compound, the designer operates the input unit 13b to notify the processing unit 13c that the compound being displayed is the desired compound.

[0108] In step S14, when the processing unit 13c determines, in response to the designer's input operation on the input unit 13b, that the compound being displayed is not the desired compound ("No" in step S14), in step S15, the processing unit 13c, in response to the designer's input operation on the input unit 13b, changes the sequence information of the compound to be input to the structure prediction unit 11 to other sequence information or updates the stereostructure of the compound being displayed, thereby determining the next candidate compound.

[0109] On the other hand, in step S14, when the processing unit 13c determines that the compound being displayed is the desired compound in response to the designer's input operation on the input unit 13b ("Yes" in step S14), in step S16, the processing unit 13c determines that the compound being displayed is the desired compound.

[0110] 1-4. Learning methods of neural networks

[0111] See Figures 5 to 9 A learning method for a neural network according to a first embodiment is described. Figure 5 and Figure 6 It is a diagram used to describe the structure / sequence prediction artificial intelligence (AI) according to the first embodiment. Figures 7 to 9 It is a graph used to describe AI prediction based on the feature values ​​of the first embodiment.

[0112] The structure / sequence prediction AI includes, for example, a structure prediction unit 11 and a sequence prediction unit 14, and the eigenvalue prediction AI includes, for example, an eigenvalue prediction unit 12 (see...). Figure 2 These learning processes will be described sequentially. Note that learning a neural network refers to, for example, adjusting parameters (such as weights and biases) so that the output layer produces the expected result (correct answer) as humans would.

[0113] Structure / Sequence Prediction AI

[0114] The learning process for AI based on structure / sequence prediction is divided into two phases (Phase 1 and Phase 2). See also Figure 5 Describe phase 1, and see also Figure 6 Description phase 2. Figure 5 and Figure 6 Each of the predictive models in the model is a machine learning model (AI model). These machine learning models are implemented using machine learning methods such as neural networks.

[0115] First of all, Figure 5 In stage 1 shown, the structure prediction model 11a and the sequence prediction model 14a collect training data from the sequence / structure information DB 20, in which compound sequence information B1 and compound structure information (stereoscopic structure information of the compound) B2 are paired, and perform parallel asynchronous learning.

[0116] exist Figure 5 In the embodiments, structure prediction model 11a predicts the stereostructure of the compound from compound sequence information B1 and outputs predicted structure information B3, which is stereostructure information about the predicted stereostructure. Sequence prediction model 14a predicts the sequence based on compound structure information B2 and outputs predicted sequence information B4 about the predicted sequence.

[0117] For example, structure prediction model 11a performs learning to minimize the error between compound structure information B2 on sequence / structure information DB 20 and predicted structure information B3 (error minimization 11b). For example, sequence prediction model 14a performs learning to minimize the error between compound sequence information B1 on sequence / structure information DB 20 and predicted sequence information B4 (error minimization 14b). After the two learning sessions have converged, learning is performed in the subsequent second half (stage 2). The convergence of learning is determined, for example, by the designer and in response to the designer's input operation decision on input unit 13b.

[0118] Next, in Figure 6In stage 2 shown, only compound sequence information B1 is used to train the learning model. First, compound sequence information B1 regarding sequence / structure information DB 20 is input to structure prediction model 11a, and structure prediction model 11a predicts the stereostructure based on compound sequence information B1. Predicted structure information B3 regarding the predicted stereostructure is input to sequence prediction model 14a, and sequence prediction model 14a predicts the sequence based on predicted structure information B3. The error between predicted sequence information B4 regarding the predicted sequence and compound sequence information B1 on sequence / structure information DB 20 is evaluated (error minimization 14c). The error value is used to update the parameters of structure prediction model 11a and sequence prediction model 14a through backpropagation.

[0119] By repeating this error minimization cycle 14c, a wealth of sequence information about unknown structures existing in the world can be used for learning. Consistent learning can be achieved by simultaneously updating the parameters of the sequence prediction model 14a and the structure prediction model 11a end-to-end, so that the structure prediction results are translated into the original sequence.

[0120] End-to-end learning, for example in deep learning, involves learning the weights of all layers in a neural network, from the input layer (end) to the output layer (end), all at once for a single problem or task. That is, end-to-end machine learning generates a structure prediction model 11a and a sequence prediction model 14a.

[0121] Here, for example, in Figure 2 The structure prediction unit 11 shown includes the structure prediction model 11a trained above, and the structure prediction model 11a can be used to predict the stereostructure from sequence information (compound sequence information A1) about the compound. For example, Figure 2 The sequence prediction unit 14 described herein includes the trained sequence prediction model 14a described above, and the sequence prediction model 14a can be used to predict sequences from information about the stereostructure of the compound.

[0122] Eigenvalue prediction AI

[0123] The learning process for eigenvalue prediction AI is divided into three stages (Stage 1, Stage 2, and Stage 3). See also... Figure 7 Describe phase 1, see Figure 8 Describe phase 2, and see also Figure 9 Description phase 3. Figure 7 , Figure 8 and Figure 9 Each of the prediction models in the dataset is a machine learning model.

[0124] First of all, Figure 7In stage 1 shown, training data for learning the feature values ​​of the compound is selected through machine learning. The information processing device 10 has an experimental reliability prediction model 15. For example, the feature value prediction unit 12 of the information processing device 10 may include the experimental reliability prediction model 15.

[0125] Experimental reliability prediction model 15 collects feature value information C1 about the feature values ​​of the compound and experimental condition information C2 as information about the experimental conditions when obtaining the feature values ​​from feature information DB 30 about the characteristics of the compound, and performs learning. Figure 7 In the embodiment, the experimental reliability prediction model 15 predicts the experimental reliability C3 based on the feature value information C1 and the experimental condition information C2, and outputs the predicted experimental reliability C3.

[0126] For example, the experimental reliability prediction model 15 performs learning by performing error minimization 15a to regress the reliability score C4 based on the compound's eigenvalue information C1 and experimental condition information C2, and outputting experimental reliability C3. After learning convergence, the experimental reliability prediction model 15 can evaluate the validity of the experiment as experimental reliability C3 based on the eigenvalue information C1 and experimental condition information C2 on the feature information DB 30.

[0127] Information regarding experimental conditions, temperature and humidity during the experiment, experimental procedures, droplet volume, and the possibility of foreign contamination can be used. The characteristic values ​​accumulated by compounds through experimental analysis can vary significantly due to subtle changes in experimental conditions. Therefore, evaluators, as experts in chemical synthesis, assess the validity of characteristic values ​​by referring to experimental conditions, and define the degree of evaluation as a reliability score C4.

[0128] The information processing device 10 includes an evaluation unit 16. The evaluation unit 16 includes a presentation unit 16a, an input unit 16b, and a processing unit 16c. The evaluation unit 16 presents characteristic value information C1 and experimental condition information C2 to the designer via the presentation unit 16a, receives input operations from the designer via the input unit 16b, and determines a reliability score C4 in response to the received input operations (requests) from the designer via the processing unit 16c. For example, the more reliable the experimental conditions, the higher the reliability score C4 determined by the designer.

[0129] Presentation unit 16a corresponds to presentation unit 13a, input unit 16b corresponds to input unit 13b, and processing unit 16c corresponds to processing unit 13c. That is, evaluation unit 16 corresponds to selection unit 13, but processing unit 16c performs the processing of determining reliability score C4 in addition to the processing of processing unit 13c.

[0130] Processing unit 16c determines a reliability score C4 in response to input operations performed by the designer on input unit 16b. For example, the designer operates input unit 16b while referencing experimental conditions to evaluate the validity of the feature value and determine the reliability score C4.

[0131] exist Figure 8 In stage 2 shown, the eigenvalue prediction model 12a collects compound sequence information C5, compound structural information (stereoscopic structural information about the compound) C6, and various eigenvalues ​​C7 from the feature information DB 30, and performs learning. Figure 8 In the embodiment, the eigenvalue prediction model 12a predicts eigenvalues ​​based on compound sequence information C5 and compound structure information C6, and outputs eigenvalue information C8 about the predicted eigenvalues.

[0132] Eigenvalue prediction model 12a is implemented, for example, by a multi-head neural network with multiple output layers to predict various types of eigenvalues. As various eigenvalues, for example, active site C8a, denaturation temperature C8b, solubility C8c, pH characteristic C8d, and biosynthetic cost C8e in the sequence can be used.

[0133] For example, eigenvalue prediction model 12a predicts various eigenvalues ​​of a compound based on compound sequence information C5 and compound structure information C6, and performs learning while comparing eigenvalue information C8 regarding the predicted eigenvalues ​​with various eigenvalues ​​C7 from the training data collected from eigenvalue information DB 30. At this time, experimental reliability prediction model 15, trained in the first half of stage 1, is used in parallel.

[0134] For example, the experimental reliability prediction model 15 evaluates whether the various feature values ​​C7 used as training data are valid as experimental reliability C3, and the feature value prediction model 12a performs weighted error minimization 12b during learning, which gives more weight to data with higher experimental reliability C3. This enables more effective learning from feature value information C8 that varies greatly from experimental analysis.

[0135] exist Figure 9 In stage 3 shown, the experimental condition information C2 and various eigenvalues ​​C7 are known, but data for compounds with unknown stereostructures are added, and the learning parameters are further tuned (fine-tuned).

[0136] In this adjustment, the predicted structural information B3 regarding the stereostructure predicted by the structural prediction model 11a is input into the eigenvalue prediction model 12a. Then, the error between the eigenvalue information C8 regarding the predicted eigenvalues ​​and the correct eigenvalues ​​C7 is backpropagated to the structural prediction model 11a and the eigenvalue prediction model 12a, thereby updating the learning parameters of the structural prediction model 11a and the eigenvalue prediction model 12a. This allows the accuracy of the structural prediction model 11a and the eigenvalue prediction model 12a to be further improved using information about compounds with unknown stereostructures.

[0137] Here, for example, Figure 2 The eigenvalue prediction unit 12 shown includes the eigenvalue prediction model 12a trained above, and the eigenvalue prediction model 12a can be used to predict various eigenvalues ​​of the compound (including the active site) based on the compound's sequence information (compound sequence information A1) and information about its stereostructure.

[0138] By using the information processing system 1 according to the first embodiment described above, i.e., the compound design system, the designer, as a user, can visually and interactively design desired compounds (e.g., new compounds that achieve the desired structure and desired characteristics). Furthermore, the efficiency of designing desired compounds (e.g., the efficiency of developing new compounds) can be improved.

[0139] 2. Second Implementation Method

[0140] 2-1. Example of Information Processing Device Configuration

[0141] Reference Figure 10 A configuration example of the information processing apparatus 10A according to the second embodiment is described. Figure 10 This is a diagram showing an example of the configuration of the information processing apparatus 10A according to the second embodiment.

[0142] like Figure 10 As shown, the information processing device 10A of the second embodiment includes, in addition to the structure prediction unit (first structure prediction unit) 11, feature value prediction unit 12, selection unit 13 and sequence prediction unit 14 involved in the first embodiment, a structure prediction unit (second structure prediction unit) 21, an active site prediction unit 22, a conformation prediction unit 23 and a stability prediction unit 24.

[0143] The structure prediction unit 21 predicts the stereostructure of the second compound based on the sequence information of the second compound (second compound sequence information A1b), and sends the stereostructure information regarding the predicted stereostructure to the active site prediction unit 22. For example, the structure prediction unit 21 can use the same prediction method as the structure prediction unit 11 according to the first embodiment. Note that the structure prediction unit 11 according to the first embodiment receives the sequence information regarding the first compound (first compound sequence information A1a). For example, proteins, glycoproteins, or nucleic acids can be used as the first and second compounds.

[0144] The active site prediction unit 22 predicts the active site of the second compound based on its sequence and stereostructure information, and sends the active site information to the conformation prediction unit 23. The active site is one of the different feature values. Therefore, the active site prediction unit 22 can use the same prediction method as the feature value prediction unit 12 according to the first embodiment.

[0145] The conformation prediction unit 23 predicts the conformation (e.g., docking conformation) of the complex of the first and second compounds based on the stereostructure and active site information of the first compound and the stereostructure and active site information of the second compound, and sends the conformation information of the predicted conformation of the complex to the stability prediction unit 24 and the selection unit 13. For example, the conformation prediction unit 23 uses a neural network (e.g., a machine learning model) to predict the conformation of the complex formed by combining the first compound (e.g., a candidate compound) and the second compound (e.g., a protein).

[0146] The neural network serving as the conformation prediction unit 23 can utilize methods already disclosed in papers, etc. For example, EquiBind, disclosed in NPL 3, can be used as a configuration for the neural network. However, the configuration of the neural network is not limited to this.

[0147] The stability prediction unit 24 predicts the stability of the complex based on conformational information about the complex and sends stability information (e.g., docking score) about the stability of the complex to the selection unit 13. For example, the stability prediction unit 24 uses a neural network (e.g., a machine learning model) to predict the stability of the complex of the first compound and the second compound as a docking score.

[0148] The neural network serving as the stability prediction unit 24 can, for example, be a graphical neural network that represents the first and second compounds in a graphical model and predicts the stability of the complex. However, the configuration of the neural network is not limited to this.

[0149] According to this configuration, the presentation unit 13a, in addition to visualizing the stereostructure information of the first compound predicted by the structure prediction unit 11 and the eigenvalue information predicted by the eigenvalue prediction unit 12, also visualizes the conformational information of the complex predicted by the conformation prediction unit 23 and the stability information of the complex predicted by the stability prediction unit 24, and presents the visualized information to the user. This allows the designer to visually evaluate candidate compounds on the information processing device 10A based on the conformational and stability information of the compound, in addition to the predicted stereostructure and eigenvalue information of the first compound. Thereafter, the designer's process in compound design is the same as in the first embodiment.

[0150] Graphical Model

[0151] Here, the graphical model is a graph representing probability dependencies. Specifically, the graphical model consists of multiple nodes and multiple edges. Nodes are connected to each other by edges, which are schematically represented as circles, and edges are represented in many cases as lines connecting the nodes. For example, the length of the edge connecting two nodes is determined based on the magnitude of a certain probability associated with them. When the probability is high, the edge is relatively short, and when the probability is low, the edge is relatively long.

[0152] For example, a graphical model can be created by treating atoms as nodes and the probability of atoms combining with each other as edges. For instance, when the probability of the first and second atoms combining is high, the node representing the first atom and the node representing the second atom are connected by a short edge. Conversely, when the probability of combining is low, the nodes are connected by long edges. Note that the probability of atoms combining with each other depends on the distance between the atoms.

[0153] For example, when the distance between atoms is short, the probability of atoms combining with each other is high. Conversely, when the distance is long, the probability of atoms combining with each other is low. That is, a graphical model can be created by treating the distance between atoms as edges. In this case, when the distance between atoms is long, nodes are connected by long edges. This also means that the probability of atoms combining with each other is low. Conversely, when the distance between atoms is short, nodes are connected by short edges. This also means that the probability of atoms combining with each other is high.

[0154] Atoms can be connected via edges only when the distance is less than a predetermined threshold (e.g., 10 angstroms). Such atom pairs with a distance shorter than the threshold (considered as contact) are called contact atom pairs. Functional labels can be embedded in nodes and edges. That is, node features and edge features can be generated based on functional labels. Furthermore, the specific method for generating the graphical model is unrestricted.

[0155] 2-2. Examples of compound design treatment

[0156] ReferenceFigure 11 Examples of compound design treatment according to the second embodiment are described. Figure 11 It is a flowchart illustrating the procedures in an embodiment of the compound design process according to the second embodiment.

[0157] like Figure 11 As shown, steps S11 to S16 according to the second embodiment are substantially the same as steps S11 to S16 according to the first embodiment, but the amount of information presented in step S13a according to the second embodiment (e.g., three-dimensional structure, feature value, conformation and docking score) is greater than the amount of information presented in step S13 according to the first embodiment (e.g., three-dimensional structure and feature value).

[0158] In step S21, the structure prediction unit 21 predicts the stereostructure of the second compound based on the input sequence information of the second compound, and obtains stereostructure information of the second compound about the predicted stereostructure.

[0159] In step S22, the active site prediction unit 22 predicts the active site of the second compound based on the input sequence information of the second compound and the predicted stereostructure information of the second compound, and obtains information about the active site of the second compound (an example of feature value information).

[0160] In step S23, the conformation prediction unit 23 predicts the conformation of the complex of the first compound and the second compound based on the stereostructure information and active site information of the first compound and the obtained stereostructure information and active site information of the second compound, and obtains conformation information about the predicted complex.

[0161] In step S24, the stability prediction unit 24 predicts the stability (docking score) of the complex based on the obtained conformational information about the complex, obtains stability information about the predicted stability of the complex, and transfers the processing to step S13a.

[0162] In step S13a, in addition to the stereostructure and eigenvalues ​​of the first compound, the presentation unit 13a presents the conformation and stability (docking score) of the complex to the designer based on the stereostructure and eigenvalue information of the first compound, as well as the conformation and stability information of the complex. This allows the designer to identify the conformation and stability of the complex in addition to the predicted stereostructure and eigenvalues ​​of the first compound.

[0163] 2-3. Learning methods of neural networks

[0164] See Figure 12 and Figure 13 A learning method for a neural network according to a second embodiment is described.Figure 12 and Figure 13 This is a diagram used to describe the interactive prediction AI according to the second embodiment. For example, the interactive prediction AI includes a conformation prediction unit 23 and a stability prediction unit 24 (see...). Figure 10 The learning process will be described in sequence.

[0165] Interactive Predictive AI

[0166] The learning of AI through interactive prediction is divided into two phases (Phase 1 and Phase 2). See also... Figure 12 Describe phase 1, and see also Figure 13 Description phase 2. Figure 12 and Figure 13 Each of the prediction models in the dataset is a machine learning model.

[0167] First of all, Figure 12 In stage 1 shown, the conformation prediction model 23a collects complex structure information D1 and sequence information (first compound sequence information D2a and second compound sequence information D2b), stereostructure information (first compound structure information D3a and second compound structure information D3b), and active site information (first compound active site information D4a and second compound active site information D4b) of the complex information DB 50 regarding the compound, and performs learning. Furthermore, the complex information DB 50 is configured to communicate with the information processing device 10A via the network 40.

[0168] exist Figure 12 In the embodiments described, the conformation prediction model 23a predicts the conformation of the complex based on the sequence information D2a of the first compound, the sequence information D2b of the second compound, the structural information D3a of the first compound, the structural information D3b of the second compound, the active site information D4a of the first compound, and the active site information D4b of the second compound, and outputs conformational information D5 related to the predicted conformation. The conformation prediction model 23a can predict the conformation of the compound based at least on the structural information D3a of the first compound, the structural information D3b of the second compound, the active site information D4a of the first compound, and the active site information D4b of the second compound.

[0169] For example, conformation prediction model 23a is trained to minimize the error between complex structure information D1 and conformation information D5 on complex information DB 50 (error minimization 23b). After this training has converged, training is performed in the subsequent second half (stage 2). The convergence of training is determined, for example, by the designer and in response to the designer's input operation decision on input unit 13b.

[0170] Next, in Figure 13In stage 2 shown, stability prediction model 24a performs learning to predict the stability of the predicted conformation of the complex. Stability prediction model 24a learns to predict docking score D7 based on conformational information D5, first compound structure information D3a, second compound structure information D3b, and binding affinity information D6, which are input from conformational prediction model 23a. Figure 13 In the embodiments described, the stability prediction model 24a predicts the docking score D7 of the stability of the complex based on the conformational information D5, the first compound structure information D3a, and the second compound structure information D3b, and outputs the predicted docking score D7.

[0171] For example, stability prediction model 24a evaluates the error between binding affinity information D6 and the docking score D7 output from stability prediction model 24a (error minimization 23c), and updates the parameters in stability prediction model 24a and conformation prediction model 23a end-to-end through backpropagation. That is, conformation prediction model 23a and stability prediction model 24a are generated through end-to-end machine learning. This allows stability prediction model 24a to be tuned to further improve the prediction accuracy of the final docking score D7. Binding affinity information D6 can be, for example, binding affinity such as IC50 (50% inhibition concentration), dissociation constant, and binding energy.

[0172] Here, for example, Figure 10 The conformation prediction unit 23 shown includes the above-mentioned trained conformation prediction model 23a. Using the conformation prediction model 23a, the conformation prediction unit 23 can predict the conformation of the complex of the first compound and the second compound based on the stereostructure information and active site information of the first compound and the stereostructure information and active site information of the second compound.

[0173] For example, Figure 10 The stability prediction unit 24 shown includes the trained stability prediction model 24a described above, and using the stability prediction model 24a, the stability prediction unit 24 can predict the stability of the complex (docking score D7) based on the conformational information D5 of the complex.

[0174] By using the information processing system 1 according to the second embodiment described above, i.e., the compound design system, the same effect as in the first embodiment can be obtained. For example, a designer as a user can visually and interactively design the desired compound while considering the desired compound and the complex conformation information D5, stability information, etc., of the compound used for compound formation.

[0175] 3. Third Implementation Method

[0176] 3-1. Example of Information Processing Device Configuration

[0177] Reference Figure 14 A configuration example of the information processing apparatus 10B according to the third embodiment is described. Figure 14 This is a diagram showing an example of the configuration of the information processing apparatus 10B according to the third embodiment.

[0178] like Figure 14 As shown, the information processing apparatus 10B according to the third embodiment is similar to the information processing apparatus 10A according to the second embodiment, but updates the neural networks (e.g., structure prediction model 11a, eigenvalue prediction model 12a, sequence prediction model 14a, etc.) of the corresponding blocks (e.g., structure prediction unit 11, eigenvalue prediction unit 12, sequence prediction unit 14a, and structure prediction unit 21).

[0179] In the third embodiment, data available to the designer is used to update the learning parameters of the neural network in each block. For example, when the designer obtains new information about the sequence and stereostructure of a compound, this information is used to perform additional learning on one or both blocks (e.g., structure prediction model 11a) of structure prediction unit 11 and structure prediction unit 21. At this time, the block to be updated can be selected based on the available additional information.

[0180] To demonstrate to the designer whether additional learning is required for each block, the presentation unit 13a can present the predicted reliability for each block. Figure 14 In the embodiments described above, the predicted reliability is represented as "Reliability: __". The calculation of this predicted reliability will be described below.

[0181] 3-2. Learning methods of neural networks

[0182] See Figure 15 A learning method for a neural network according to a third embodiment is described. Figure 15 This is a diagram used to describe the calculation of the predicted reliability according to the third embodiment.

[0183] like Figure 15 As shown, the prediction unit 100 according to the third embodiment also predicts the prediction reliability as the reliability of the prediction itself. For example, the prediction unit 100 corresponds to any one of the structure prediction unit 11, the eigenvalue prediction unit 12, the sequence prediction unit 14, and the structure prediction unit 21. It should be noted that the prediction unit 100 may correspond to any one of, for example, the active site prediction unit 22, the conformation prediction unit 23, and the stability prediction unit 24.

[0184] exist Figure 15In the embodiments, it is assumed that the neural network (e.g., a machine learning model) of the prediction unit 100 has multiple outputs. The neural network corresponds to, for example, any of the structure prediction model 11a, the eigenvalue prediction model 12a, and the sequence prediction model 14a, or to any of the conformation prediction model 23a and the stability prediction model 24a.

[0185] When information is input into the neural network of prediction unit 100, a prediction result is output. That is, output information E2 is generated in response to input information E1. The learning parameters of the neural network are updated to minimize the error between the output information E2, which is the prediction result, and the correct answer information E3. At this time, the error itself is regressed by another output (the output of prediction error E4).

[0186] Subsequently, the converged neural network learns to predict unknown input information and simultaneously outputs the expected prediction error E4 in that case. The prediction error E4 is used as a measure of prediction reliability, indicating the dependability of the prediction.

[0187] With this configuration, by presenting the predicted reliability to the designer through the presentation unit 13a, etc., the designer can be encouraged to add additional data. Furthermore, in both the first and second embodiments, the designer can identify the validity of the output of the prediction unit 100.

[0188] By using the information processing system 1 according to the third embodiment described above, i.e., the compound design system, effects similar to those in the first and second embodiments can be obtained. Furthermore, additional learning can be implemented for any one or all blocks. Moreover, by presenting the predictive reliability of each block, the designer can be notified whether additional learning is needed for each block.

[0189] 4. Functions and Effects

[0190] As described above, any one of the information processing apparatuses 10, 10A, and 10B according to the above embodiments includes: a structure prediction unit 11 configured to predict the stereostructure of a candidate compound based on sequence information about the candidate compound, and to obtain stereostructure information about the candidate compound; a feature value prediction unit 12 configured to predict feature values ​​of the candidate compound based on the sequence information and stereostructure information of the candidate compound, and to obtain feature value information about the candidate compound; and a presentation unit 13a configured to present the feature value information and stereostructure information about the candidate compound to a user. This enables users, such as designers, to easily identify the feature value information and stereostructure information about the candidate compound, thereby achieving high throughput in the design and evaluation of desired compounds and enabling efficient production of desired compounds.

[0191] Any of the information processing apparatuses 10, 10A, and 10B according to the above embodiments may further include: an input unit 13b configured to receive input operations from a user; and a processing unit 13c configured to determine a candidate compound as the next candidate compound in response to an input operation from the user on the input unit 13b. Thus, the user can appropriately specify the next candidate, achieving high processing capacity for the desired compound design evaluation cycle (design evaluation cycle).

[0192] The processing unit 13c can change the stereostructure information about the candidate compound in response to user input on the input unit 13b. This allows the user to appropriately change the stereostructure information about the candidate compound, thereby achieving high throughput in the design evaluation cycle.

[0193] Any of the information processing apparatuses 10, 10A, and 10B according to the above embodiments may further include: a sequence prediction unit 14 configured to predict the sequence of a next candidate compound based on stereostructural information about the next candidate compound, and to obtain sequence information about the next candidate compound; wherein, the structure prediction unit 11 can predict the stereostructure of the next candidate compound based on the sequence information about the next candidate compound, and to obtain stereostructural information about the next candidate compound; the feature value prediction unit 12 can predict the feature values ​​of the next candidate compound based on the sequence information and stereostructural information about the next candidate compound, and to obtain feature value information about the next candidate compound; and the presentation unit 13a can present the feature value information and stereostructural information about the next candidate compound to a user. This ensures high throughput in the design evaluation cycle, thereby enabling the efficient production of desired compounds.

[0194] The structure prediction unit 11 can use a structure prediction model 11a generated through machine learning to predict the stereostructure information of candidate compounds, based on the sequence and stereostructure information of each of the various compounds. Similarly, the sequence prediction unit 14 can use a sequence prediction model 14a generated through machine learning to predict the sequence information of the next candidate compound, based on the sequence and stereostructure information of each of the various compounds. This allows for the acquisition of stereostructure information of candidate compounds and sequence information of the next candidate compound with high accuracy.

[0195] The structure prediction model 11a and sequence prediction model 14a can be generated through end-to-end machine learning. Therefore, even when the stereostructure information of a candidate compound is edited by the user, consistency between the edited stereostructure information and the sequence information of the next candidate compound can be maintained.

[0196] The sequence prediction unit 14 can determine the prediction reliability of the sequence information of the next candidate compound, and the presentation unit 13a can present the prediction reliability to the user. This allows the user to identify the prediction reliability of the sequence prediction unit 14. For example, when the prediction reliability is lower than the expected value, the user can operate the input unit 13b to instruct the sequence prediction model 14a of the sequence prediction unit 14 to perform relearning.

[0197] The structure prediction unit 11 can use a structure prediction model 11a generated by machine learning to predict the stereostructure of candidate compounds. The machine learning is based on the sequence information and stereostructure information of each of the various compounds. This allows for the acquisition of stereostructure information about candidate compounds with high accuracy.

[0198] The structure prediction unit 11 can determine the prediction reliability of the stereostructure information of the candidate compound, and the presentation unit 13a can present the prediction reliability to the user. This allows the user to identify the prediction reliability of the structure prediction unit 11. For example, when the prediction reliability is lower than the expected value, the user can operate the input unit 13b to instruct the structure prediction model 11a of the structure prediction unit 11 to perform relearning.

[0199] The eigenvalue prediction unit 12 can use an eigenvalue prediction model 12a generated by machine learning to predict the eigenvalues ​​of candidate compounds. The machine learning is based on the sequence information, stereostructure information, and eigenvalue information (e.g., various eigenvalues ​​C7) of each of the various compounds. This allows eigenvalue information about candidate compounds to be obtained with high accuracy.

[0200] The eigenvalue prediction unit 12 can use an experimental reliability prediction model 15 generated through machine learning to predict the experimental reliability C3 of the experiment determining the candidate compounds. The machine learning is based on eigenvalue information C1 and experimental condition information C2 for each of the various compounds, and, in addition to sequence information, stereostructure information, and eigenvalue information for each of the various compounds, can also generate the eigenvalue prediction model 12a based on the experimental reliability C3. This allows for the acquisition of eigenvalue information about the candidate compounds with high accuracy.

[0201] In addition to the eigenvalue information and experimental condition information C2 for each of the various compounds, an experimental reliability prediction model 15 can be generated based on machine learning using a reliability score C4, which is evaluated by the user, regarding the experimental reliability of each of the various compounds. This allows for obtaining experimental reliability C3 with high accuracy.

[0202] The stereostructure information for each of the various compounds can include stereostructure information about the candidate compounds predicted by the structure prediction unit 11. Therefore, by using information about compounds with unknown stereostructures, the prediction accuracy of the eigenvalue prediction model 12a can be improved.

[0203] The eigenvalue prediction unit 12 can determine the prediction reliability of the eigenvalues ​​of the candidate compounds, and the presentation unit 13a can present the prediction reliability to the user. This allows the user to identify the prediction reliability of the eigenvalue prediction unit 12. For example, when the prediction reliability is lower than the expected value, the user can operate the input unit 13b to instruct the eigenvalue prediction model 12a of the eigenvalue prediction unit 12 to perform relearning.

[0204] Any one of the information processing apparatuses 10A and 10B according to the above embodiments may further include: a structure prediction unit 21 configured to predict the stereostructure of the compound for complex formation based on sequence information of the compound for complex formation of the candidate compound, and obtain stereostructure information of the compound for complex formation; an active site prediction unit 22 configured to predict the active site of the compound for complex formation based on sequence information and stereostructure information of the compound for complex formation, and obtain active site information of the compound for complex formation; a conformation prediction unit 23 configured to predict the conformation of the compound of the candidate compound and the compound for complex formation based on the active site information and stereostructure information of the compound for complex formation, as well as the stereostructure information and eigenvalue information of the candidate compound, and obtain conformation information D5 of the compound; and a stability prediction unit 24 configured to predict the stability of the compound based on the conformation information D5 of the compound, and obtain stability information of the compound, wherein, in addition to the stereostructure information and eigenvalue information of the candidate compound, the presentation unit 13a may present the stability information of the compound to the user. This allows users, such as designers, to easily identify information about the stability of the complex, in addition to characteristic and stereostructural information about the candidate compound, thereby enabling high-throughput design evaluation cycles and efficient production of desired compounds.

[0205] The conformation prediction unit 23 can use a conformation prediction model 23a generated by machine learning to predict the conformation of the complex. The machine learning is based on sequence information, stereostructure information, and active site information for each of multiple compounds in each of the various compounds, and stereostructure information for each of the various complexes. Similarly, the stability prediction unit 24 can use a stability prediction model 24a generated by machine learning to predict the stability of the complex. The machine learning is based on stereostructure information and conformational information D5 for each of multiple compounds in each of the various complexes. This allows for the acquisition of morphological information D5 and stability information about the compounds with high accuracy.

[0206] Conformation prediction model 23a and stability prediction model 24a can be generated through end-to-end machine learning. This can improve the prediction accuracy of stability prediction model 24a.

[0207] The conformation prediction unit 23 or the stability prediction unit 24 can determine the predictive reliability of the conformation or stability of the complex, and the presentation unit 13a can present the predictive reliability to the user. This allows the user to identify the predictive reliability of the conformation prediction unit 23 or the stability prediction unit 24. For example, when the predictive reliability is lower than the expected value, the user can operate the input unit 13b to instruct the conformation prediction model 23a of the conformation prediction unit 23 or the stability prediction model 24a of the stability prediction unit 24 to perform relearning.

[0208] 5. Other implementation methods

[0209] The processing described in the above embodiments (or modifications) can be implemented in various other embodiments (modifications) besides those described above. For example, in the processes described in the above embodiments, all or some of the processes described as being executed automatically can be executed manually, or all or some of the processes described as being executed manually can be executed automatically using known methods. Furthermore, unless otherwise stated, the processing steps, specific names, and information including various data and parameters described in this specification and the accompanying drawings may be changed as needed. For example, the various types of information shown in the accompanying drawings are not limited to the information shown.

[0210] Each component of the apparatus shown in the accompanying drawings is conceptual in function and does not necessarily have to be physically implemented as shown in the drawings. That is, the specific embodiments of the dispersion and integration of the various devices are not limited to those shown in the accompanying drawings, and all or part of each device may be functionally or physically dispersed or integrated in any unit, depending on various loads, usage conditions, etc.

[0211] The above-described embodiments (or variations) can be appropriately combined within the scope of not contradicting each other. Furthermore, the effects described in this specification are merely examples and are not limiting; other effects may be obtained.

[0212] In the above embodiments (or modifications), a system means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are in the same housing. Therefore, multiple devices housed in a separate housing and connected via network 40, and a single device including multiple modules housed in a single housing, are both systems.

[0213] In the above embodiments (or variations), a cloud computing configuration may be employed, wherein a function is shared and processed by multiple devices via network 40. Furthermore, each step described in the above processing procedure (e.g., the flowchart) may be executed by a single device or may be shared and executed by multiple devices. Additionally, when a single step includes multiple processes, the multiple processes included in that single step may be executed by a single device or may be shared and executed by multiple devices.

[0214] 6. Examples of Hardware Configuration

[0215] For example, by having such Figure 16 The computer 500 configured as shown can implement the information processing apparatus 10, 10A and 10B according to the first to third embodiments described above. Figure 16 This is a diagram illustrating an embodiment of the hardware configuration of computer 500.

[0216] like Figure 16 As shown, the computer 500 includes a CPU (Central Processing Unit) 510, a ROM (Read-Only Memory) 520, and a RAM (Random Access Memory) 530.

[0217] CPU 510, ROM 520, and RAM 530 are interconnected via bus 540. Input / output interface 550 is further connected to bus 540. Input unit 560, output unit 570, recording unit 580, communication unit 590, and driver 600 are connected to input / output interface 550.

[0218] Input unit 560 includes, for example, a keyboard and mouse, a microphone, or an imaging element. Output unit 570 includes, for example, a display or a speaker. Recording unit 580 includes, for example, a hard disk or non-volatile memory. Communication unit 590 includes, for example, a network interface. Driver 600 drives a removable recording medium 610 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0219] In the computer 500 implemented as described above, the CPU 510 performs the series of processes described above, for example, by loading a program recorded in the recording unit 580 into the RAM 530 via the input / output interface 550 and the bus 540 and executing the program.

[0220] That is, when the computer 500 is used as any one of the information processing apparatuses 10, 10A and 10B according to the first to third embodiments, the CPU 510 of the computer 500 executes the program loaded into the RAM 530 to implement the function of each unit of the information processing apparatuses 10, 10A and 10B according to the first to third embodiments.

[0221] For example, a program executed by the CPU 510 may be provided by recording on a removable recording medium 610 such as an encapsulation medium, or via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0222] That is, by attaching the removable recording medium 610 to the driver 600, the program can be installed in the recording unit 580 via the input / output interface 550, or the program can be received by the communication unit 590 and installed in the recording unit 580 via a wired or wireless transmission medium. Alternatively, the program can be pre-installed in the ROM 520 or the recording unit 580.

[0223] Note that a program may be a program in which processing is performed sequentially in the order described in this specification, or a program in which processing is performed in parallel or at necessary time intervals (such as when invoked).

[0224] 7. Supplementary Explanation

[0225] It should be noted that this technology may also have the following configurations.

[0226] (1) An information processing apparatus, comprising: a structure prediction unit configured to predict the stereostructure of a candidate compound based on sequence information about the candidate compound, and to obtain stereostructure information about the candidate compound; an eigenvalue prediction unit configured to predict eigenvalues ​​of the candidate compound based on the sequence information and stereostructure information about the candidate compound, and to obtain eigenvalue information about the candidate compound; and a presentation unit configured to present the eigenvalue information and stereostructure information about the candidate compound to a user.

[0227] (2) The information processing apparatus according to (1) further includes: an input unit configured to receive input operations from a user; and a processing unit configured to determine a candidate compound as the next candidate compound in response to the user's input operation on the input unit.

[0228] (3) The information processing apparatus according to (2), wherein the processing unit changes the stereostructure information of the candidate compound in response to the user’s input operation on the input section.

[0229] (4) The information processing apparatus according to (2) or (3) further includes: a sequence prediction unit configured to predict the sequence of a next candidate compound based on the stereostructure information of the next candidate compound, and to obtain sequence information of the next candidate compound, wherein the structure prediction unit predicts the stereostructure of the next candidate compound based on the sequence information of the next candidate compound, and to obtain stereostructure information of the next candidate compound, the feature value prediction unit predicts the feature value of the next candidate compound based on the sequence information and stereostructure information of the next candidate compound, and to obtain feature value information of the next candidate compound, and the presentation unit presents the feature value information and stereostructure information of the next candidate compound to the user.

[0230] (5) The information processing apparatus according to (4), wherein the structure prediction unit uses a structure prediction model generated by machine learning to predict the stereostructure information of the candidate compound, the machine learning being based on the sequence information and stereostructure information of each of the various compounds, and the sequence prediction unit uses a sequence prediction model generated by machine learning to predict the sequence information of the next candidate compound, the machine learning being based on the sequence information and stereostructure information of each of the various compounds.

[0231] (6) The information processing device according to (5), wherein a structural prediction model and a sequence prediction model are generated by end-to-end machine learning.

[0232] (7) An information processing apparatus according to any one of (4) to (6), wherein the sequence prediction unit determines the prediction reliability of the sequence information of the next candidate compound, and the presentation unit presents the prediction reliability to the user.

[0233] (8) An information processing apparatus according to any one of (1) to (7), wherein the structure prediction unit uses a structure prediction model generated by machine learning to predict the stereostructure of candidate compounds, the machine learning being based on sequence information and stereostructure information of each of the various compounds.

[0234] (9) An information processing apparatus according to any one of (1) to (8), wherein the structure prediction unit determines the prediction reliability of the stereostructure information of the candidate compound, and the presentation unit presents the prediction reliability to the user.

[0235] (10) An information processing apparatus according to any one of (1) to (9), wherein the eigenvalue prediction unit uses an eigenvalue prediction model generated by machine learning to predict the eigenvalues ​​of candidate compounds, the machine learning being based on sequence information, stereostructure information and eigenvalue information of each of the various compounds.

[0236] (11) The information processing apparatus according to (10), wherein the feature value prediction unit uses an experimental reliability prediction model generated by machine learning to predict the experimental reliability of the experiment for determining the candidate compound, the machine learning being based on the feature value information and experimental condition information of each of the various compounds, and generating the feature value prediction model based on the experimental reliability in addition to the sequence information, stereostructure information and feature value information of each of the various compounds.

[0237] (12) The information processing apparatus according to (11), wherein the experimental reliability prediction model is generated by machine learning based on reliability scores of the experimental reliability of each of the various compounds, in addition to the characteristic value information and experimental condition information of each of the various compounds.

[0238] (13) An information processing apparatus according to any one of (10) to (12), wherein the stereostructure information of each of the various compounds includes stereostructure information about the candidate compounds predicted by the structure prediction unit.

[0239] (14) An information processing apparatus according to any one of (1) to (13), wherein the eigenvalue prediction unit determines the prediction reliability of the eigenvalues ​​of the candidate compound, and the presentation unit presents the prediction reliability to the user.

[0240] (15) The information processing apparatus according to any one of (1) to (14) further includes: a structure prediction unit configured to predict the stereostructure of the compound for forming the complex based on sequence information of the compound for forming the candidate compound, and to obtain stereostructure information of the compound for forming the complex; an active site prediction unit configured to predict the active site of the compound for forming the complex based on the sequence information and stereostructure information of the compound for forming the complex, and to obtain active site information of the compound for forming the complex; a conformation prediction unit configured to predict the conformation of the complex of the candidate compound and the compound for forming the complex based on the active site information and stereostructure information of the compound for forming the complex, and on the stereostructure information and eigenvalue information of the candidate compound, and to obtain conformation information of the complex; and a stability prediction unit configured to predict the stability of the complex based on the conformation information of the complex, and to obtain stability information of the complex, wherein the presentation unit presents the stability information of the complex to the user in addition to the stereostructure information and eigenvalue information of the candidate compound.

[0241] (16) The information processing apparatus according to (15), wherein the conformation prediction unit uses a conformation prediction model generated by machine learning to predict the conformation of the compound, the machine learning being based on sequence information, stereostructure information and active site information for each of the plurality of compounds in each of the various complexes, and stereostructure information for each of the various compounds, and the stability prediction unit uses a stability prediction model generated by machine learning to predict the stability of the complex, the machine learning being based on stereostructure information and conformation information for each of the plurality of compounds in each of the various complexes.

[0242] (17) The information processing apparatus according to (16), wherein the conformation prediction model and the stability prediction model are generated by end-to-end machine learning. (18)

[0244] The information processing apparatus according to any one of (15) to (17) wherein the conformation prediction unit or the stability prediction unit determines the prediction reliability of the conformation or stability of the predicted complex, and the presentation unit presents the prediction reliability to the user.

[0245] (19) An information processing system, comprising: a structure prediction unit configured to predict the stereostructure of a candidate compound based on sequence information about the candidate compound, and to obtain stereostructure information about the candidate compound; an eigenvalue prediction unit configured to predict eigenvalues ​​of the candidate compound based on the sequence information and stereostructure information about the candidate compound, and to obtain eigenvalue information about the candidate compound; and a presentation unit configured to present the eigenvalue information and stereostructure information about the candidate compound to a user.

[0246] (20) An information processing method, comprising: using an information processing device, predicting the stereostructure of a candidate compound based on sequence information of the candidate compound and obtaining stereostructure information of the candidate compound; predicting the characteristic value of the candidate compound based on the sequence information and stereostructure information of the candidate compound and obtaining characteristic value information of the candidate compound; and presenting the characteristic value information and stereostructure information of the candidate compound to a user.

[0247] (21) An information processing apparatus comprising any one or all of a plurality of constituent elements included in an information processing apparatus according to any one of (1) to (18).

[0248] (22) An information processing method using any one or all of a plurality of constituent elements included in an information processing apparatus according to any one of (1) to (18).

[0249] Symbol Explanation

[0250] 1. Information Processing System

[0251] 10. Information processing device

[0252] 10A Information Processing Device

[0253] 10B Information Processing Device

[0254] 11 Structural Prediction Unit

[0255] 11a Structural Prediction Model

[0256] 11b Error Minimization

[0257] 12 Eigenvalue Prediction Units

[0258] 12a Eigenvalue Prediction Model

[0259] 12b Error Minimization

[0260] 13 Selection Unit

[0261] 13a Presentation Section

[0262] 13b Input Section

[0263] 131 Keyboard

[0264] 132 mouse

[0265] 13c processing unit

[0266] 14 Sequence Prediction Units

[0267] 14a Sequence Prediction Model

[0268] 14b Error Minimization

[0269] 14c Error Minimization

[0270] 15 Experimental Reliability Prediction Model

[0271] 15a Error Minimization

[0272] 16 Evaluation Units

[0273] 16a Presentation Section

[0274] 16b Input Section

[0275] 16c processing unit

[0276] 21 Structural Prediction Unit

[0277] 22 Active Site Prediction Unit

[0278] 23 Conformation Prediction Units

[0279] 23a Conformation Prediction Model

[0280] 23b Error Minimization

[0281] 23c Error Minimization

[0282] 24 Stability Prediction Units

[0283] 24a Stability Prediction Model

[0284] 20 Sequence / Structure Information DB

[0285] 30 Feature Information DB

[0286] 40 Network

[0287] 50 Complex Information DB

[0288] 100 prediction units

[0289] A1a First compound sequence information

[0290] A1b Second compound sequence information

[0291] A2 3D structural image

[0292] A3 Table

[0293] B1 Compound Sequence Information

[0294] B2 Compound Structure Information

[0295] B3 Predictive structural information

[0296] B4 Predicted Sequence Information

[0297] C1 eigenvalue information

[0298] C2 Experimental Conditions Information

[0299] C3 Experimental Reliability

[0300] C4 Reliability Score

[0301] C5 compound sequence information

[0302] C6 compound structural information

[0303] C7 various eigenvalues

[0304] C8 eigenvalue information

[0305] C8a active site

[0306] C8b denaturation temperature

[0307] C8c solubility

[0308] C8d pH characteristics

[0309] C8e biosynthesis cost

[0310] D1 Complex Structural Information

[0311] D2a First compound sequence information

[0312] D2b Second compound sequence information

[0313] D3a First compound structure information

[0314] D3b Second compound structural information

[0315] D4a First compound active site information

[0316] Information on the active site of the second compound, D4b.

[0317] D5 Construct Information

[0318] D6 Combines affinity information

[0319] D7 Docking Score

[0320] E1 Input Information

[0321] E2 Output Information

[0322] E3 Correct Answer Information

[0323] E4 Prediction error.

Claims

1. An information processing apparatus, comprising: The structure prediction unit is configured to predict the stereostructure of the candidate compound based on the sequence information of the candidate compound, and to obtain the stereostructure information of the candidate compound. The feature value prediction unit is configured to predict the feature values ​​of the candidate compound based on the sequence information and the stereostructure information of the candidate compound, and to obtain feature value information about the candidate compound. as well as The presentation unit is configured to present the feature value information and the stereostructure information of the candidate compound to the user.

2. The information processing apparatus according to claim 1, further comprising: The input section is configured to receive user input. as well as The processing unit is configured to determine the candidate compound as the next candidate compound in response to the user's input operation on the input section.

3. The information processing apparatus according to claim 2, wherein, The processing unit changes the stereostructure information of the candidate compound in response to the user's input operation on the input section.

4. The information processing apparatus according to claim 2, further comprising: A sequence prediction unit is configured to predict the sequence of the next candidate compound based on the stereostructure information of the next candidate compound, and to obtain the sequence information of the next candidate compound, wherein... Based on the sequence information of the next candidate compound, the structure prediction unit predicts the stereostructure of the next candidate compound and obtains stereostructure information about the next candidate compound. The feature value prediction unit predicts the feature values ​​of the next candidate compound based on the sequence information and the stereostructure information of the next candidate compound, and obtains feature value information about the next candidate compound. The presentation unit presents the feature value information and the stereostructure information of the next candidate compound to the user.

5. The information processing apparatus according to claim 4, wherein, The structure prediction unit uses a structure prediction model generated through machine learning to predict the stereostructure information of the candidate compounds. The machine learning is based on the sequence and stereostructure information of each of the various compounds. The sequence prediction unit uses a sequence prediction model generated by machine learning to predict the sequence information of the next candidate compound, the machine learning being based on the sequence information and stereostructure information of each of the various compounds.

6. The information processing apparatus according to claim 5, wherein, The structure prediction model and the sequence prediction model are generated through end-to-end machine learning.

7. The information processing apparatus according to claim 4, wherein, The sequence prediction unit determines the prediction reliability of the sequence information for predicting the next candidate compound, and The presentation unit presents the predicted reliability to the user.

8. The information processing apparatus according to claim 1, wherein, The structure prediction unit uses a structure prediction model generated by machine learning to predict the stereostructure of the candidate compounds, the machine learning being based on sequence information and stereostructure information for each of the various compounds.

9. The information processing apparatus according to claim 1, wherein, The structure prediction unit determines the prediction reliability of the stereostructure information of the candidate compound, and The presentation unit presents the predicted reliability to the user.

10. The information processing apparatus according to claim 1, wherein, The eigenvalue prediction unit uses an eigenvalue prediction model generated by machine learning to predict the eigenvalues ​​of the candidate compounds, the machine learning being based on sequence information, stereostructure information, and eigenvalue information for each of the various compounds.

11. The information processing apparatus according to claim 10, wherein, The eigenvalue prediction unit uses an experimental reliability prediction model generated through machine learning to predict the experimental reliability of the experiment for determining the candidate compounds. The machine learning is based on the eigenvalue information and experimental condition information for each of the various compounds. The eigenvalue prediction model is generated by machine learning based on the experimental reliability, in addition to the sequence information, stereostructure information, and eigenvalue information of each of the various compounds.

12. The information processing apparatus according to claim 11, wherein, The experimental reliability prediction model is generated by machine learning based on reliability scores for the experimental reliability of each of the various compounds, as evaluated by the user, in addition to the feature value information and experimental condition information for each of the various compounds.

13. The information processing apparatus according to claim 10, wherein, The stereostructure information for each of the various compounds includes stereostructure information about the candidate compounds predicted by the structure prediction unit.

14. The information processing apparatus according to claim 1, wherein, The eigenvalue prediction unit determines the prediction reliability of the eigenvalues ​​of the candidate compounds, and The presentation unit presents the predicted reliability to the user.

15. The information processing apparatus according to claim 1, further comprising: The structure prediction unit is configured to predict the stereostructure of the compound for complex formation based on the sequence information of the compound for complex formation of the candidate compound, and to obtain the stereostructure information of the compound for complex formation. The active site prediction unit is configured to predict the active site of the compound for complex formation based on the sequence information and stereostructure information of the compound for complex formation, and to obtain the active site information of the compound for complex formation. The conformation prediction unit is configured to predict the conformation of the candidate compound and the compound used for complex formation based on the active site information and stereostructure information of the compound used for complex formation, as well as the stereostructure information and eigenvalue information of the candidate compound, and to obtain conformation information about the complex. as well as A stability prediction unit is configured to predict the stability of the complex based on its conformational information and obtain stability information about the complex, wherein... In addition to the stereostructure information and feature value information of the candidate compound, the presentation unit also presents the stability information of the complex to the user.

16. The information processing apparatus according to claim 15, wherein, The conformation prediction unit uses a conformation prediction model generated by machine learning to predict the conformation of the complex. The machine learning is based on sequence information, stereoscopic information, and active site information for each of multiple compounds in each of the various complexes, and stereoscopic information for each of the various complexes. The stability prediction unit uses a stability prediction model generated by machine learning to predict the stability of the complex, the machine learning unit being based on the stereostructure and conformational information of each of the various complexes with respect to each of the plurality of compounds.

17. The information processing apparatus according to claim 16, wherein, The conformation prediction model and the stability prediction model are generated through end-to-end machine learning.

18. The information processing apparatus according to claim 15, wherein, The conformation prediction unit or the stability prediction unit determines the prediction reliability of the conformation or stability of the composite, and The presentation unit presents the predicted reliability to the user.

19. An information processing system, comprising: The structure prediction unit is configured to predict the stereostructure of the candidate compound based on the sequence information of the candidate compound, and to obtain the stereostructure information of the candidate compound. The feature value prediction unit is configured to predict the feature value of the candidate compound based on the sequence information and the stereostructure information of the candidate compound, and to obtain the feature value information of the candidate compound. as well as The presentation unit is configured to present the feature value information and the stereostructure information of the candidate compound to the user.

20. An information processing method, comprising: Through information processing devices, Predict the stereostructure of the candidate compound based on its sequence information, and obtain the stereostructure information of the candidate compound. Based on the sequence information and stereostructure information of the candidate compounds, the feature values ​​of the candidate compounds are predicted, and the feature value information of the candidate compounds is obtained. as well as The characteristic value information and the stereostructure information of the candidate compounds are presented to the user.

Citation Information

Patent Citations

  • Protein Structure Prediction from Amino Acid Sequences Using Self-Attention Neural Networks

    US20210166779A1