PROGRAM, INFORMATION PROCESSING APPARATUS, METHOD AND SYSTEM

By running a program using machine learning models on a computer, the acquisition and screening of biopolymer receptor activity information of compounds is simplified, and the problem of difficult to effectively screen compounds useful for specific receptors in the prior art is solved, achieving an efficient screening process.

JP7678924B1Active Publication Date: 2025-05-16ASAHI GRP HLDG LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024200325
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-05-16
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

It is difficult to obtain information on the activity of compounds on biopolymer receptors effectively, especially when screening for compounds useful for specific receptors.

Method used

By running a program on a computer, the program uses machine learning models to simulate the activity information of a compound as an explanatory variable and the activity information of a specific receptor as a target variable, in order to simplify the acquisition and screening of the activity information of the compound on the receptor.

Benefits of technology

It realizes simple simulation of the activity information of compounds on specific receptors, facilitates screening of compounds useful for specific receptors, and improves screening efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007678924000001_ABST
    Figure 0007678924000001_ABST
Patent Text Reader

Abstract

Easily simulate the activity information of compounds against specific receptors. [Solution] This is a program for operating a computer 20 including a processor 29 and memories 25 and 26. The memories 25 and 26 store a machine learning model in which information for identifying a compound and structural information of a compound identified by the information are used as explanatory variables, and activity information of the compound against a specific receptor is used as a target variable. The program causes the processor 29 to execute a first step of inputting information for identifying a compound to be screened and structural information of a compound identified by the information into the machine learning model, and acquiring activity information of the compound against a specific receptor as an inference result, and a second step of screening the compound based on the activity information acquired in the first step.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a program, an information processing device, a method, and a system. [Background technology]

[0002] Screening for compounds that exhibit significant activity against biological receptors is a complex task due to the enormous number of compounds involved, and therefore there is a demand for efficient screening techniques.

[0003] In relation to the above-mentioned technology, a technology is known in which a target biopolymer is specified and the three-dimensional structure of a compound to be predicted for binding is obtained, the three-dimensional structure of the biopolymer corresponding to the specification is obtained from a three-dimensional structure database that accumulates the three-dimensional structures of biopolymers, a predicted three-dimensional structure of a complex between the biopolymer and the compound is generated based on the obtained three-dimensional structure, the generated predicted three-dimensional structure is compared with a plurality of interaction patterns defined based on the statistics of the spatial arrangement distribution of ligand atoms located around the residues of the biopolymer, converted into a predicted three-dimensional structure vector representing the result of the comparison with the interaction pattern, and the predicted three-dimensional structure vector is discriminated using a machine learning algorithm, thereby predicting the binding between the three-dimensional structure of the biopolymer and the three-dimensional structure of the compound (Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2019-28879 A Summary of the Invention [Problem to be solved by the invention]

[0005] Although the technology described in Patent Document 1 can predict the binding affinity between the three-dimensional structure of a biopolymer and the three-dimensional structure of a compound, it is difficult to obtain activity information of the compound against a receptor that is part of the biopolymer.

[0006] Therefore, the present disclosure provides a technology that can easily simulate activity information of a compound against a specific receptor, thereby making it possible to easily screen for a compound that is useful against a specific receptor. [Means for solving the problem]

[0007] A program for operating a computer having a processor and a memory, the program screening a compound based on activity information of the compound against a receptor. The memory stores a machine learning model in which information for identifying a compound and structural information of a compound identified by the information are used as explanatory variables, and activity information of the compound against a specific receptor is used as a target variable. The program causes the processor to execute a first step of inputting information for identifying a compound to be screened and structural information of a compound identified by the information into the machine learning model, and acquiring activity information of the compound against a specific receptor as an inference result, and a second step of screening the compound based on the activity information acquired in the first step. Effect of the Invention

[0008] According to the present disclosure, activity information of a compound for a specific receptor can be easily simulated, thereby making it possible to easily screen for a compound that is useful for a specific receptor. [Brief description of the drawings]

[0009] [Figure 1] 1 is a diagram showing an overall configuration of a system according to an embodiment; [Diagram 2] FIG. 2 is a diagram illustrating a functional configuration of a terminal device according to an embodiment. [Diagram 3] FIG. 2 is a diagram illustrating a functional configuration of a server according to an embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of a data structure of an experiment library database according to an embodiment. [Diagram 5]FIG. 2 is a diagram illustrating an example of a data structure of teacher data according to an embodiment. [Figure 6] 11 is a flowchart illustrating an example of a processing flow in a system according to an embodiment. [Figure 7] 8 is a flowchart showing an example of a processing flow in a system according to an embodiment, which is a flowchart following the processing flow of FIG. 7. [Figure 8] FIG. 13 is a diagram showing the prediction accuracy of the machine learning model used in the experimental example. [Figure 9] FIG. 1 shows predicted results for compounds in experimental examples. [Figure 10] FIG. 2 is a block diagram showing the basic hardware configuration of a computer 90. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings. In all the drawings explaining the embodiment, the same reference numerals are given to common components, and repeated explanations are omitted. Note that the following embodiment does not unduly limit the contents of the present disclosure described in the claims. In addition, not all of the components shown in the embodiment are essential components of the present disclosure. In addition, each figure is a schematic diagram and is not necessarily illustrated strictly.

[0011] In the following description, a "processor" refers to one or more processors. The at least one processor is typically a microprocessor such as a CPU (Central Processing Unit), but may be another type of processor such as a GPU (Graphics Processing Unit). The at least one processor may be a single-core or multi-core.

[0012] Furthermore, the at least one processor may be a processor in the broad sense, such as a hardware circuit (for example, a field-programmable gate array (FPGA) or an application specific integrated circuit (ASIC)) that performs part or all of the processing.

[0013] In the following explanation, information that gives an output for an input may be described using an expression such as "xxx table", but this information may be data of any structure or a learning model such as a neural network that generates an output for an input. Therefore, the "xxx table" may be called "xxx information".

[0014] Furthermore, in the following description, the configuration of each table is an example, and one table may be divided into two or more tables, or all or part of two or more tables may be one table.

[0015] In addition, in the following explanation, the processing may be described with the "program" as the subject, but since the program is executed by a processor to perform a specified processing step by appropriately using a memory unit and / or an interface unit, etc., the subject of the processing may be the processor (or a device such as a controller having the processor).

[0016] The program may be installed in a device such as a computer, or may be, for example, in a program distribution server or a computer-readable (e.g., non-transitory) recording medium. In the following description, two or more programs may be realized as one program, or one program may be realized as two or more programs.

[0017] The functions performed by the components described herein may be implemented in circuitry or processing circuitry, including general purpose processors, application specific processors, integrated circuits, ASICs, CPUs, conventional circuits, and / or combinations thereof, programmed to perform the functions described. Processors include transistors and other circuits and are considered to be circuitry or processing circuitry. A processor may be a programmed processor that executes a program stored in a memory.

[0018] In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed in this specification or any hardware known to be programmed to realize or perform the described functions.

[0019] If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.

[0020] In the following description, identification numbers are used as identification information for various objects, but other types of identification information (for example, identifiers including alphabetic characters or symbols) may be used instead.

[0021] In addition, in the following description, when describing elements of the same type without distinguishing between them, reference signs (or common signs among the reference signs) may be used, and when describing elements of the same type with distinction between them, the identification numbers (or reference signs) of the elements may be used.

[0022] In the following description, the control lines and information lines are those that are considered necessary for the description, and not all control lines and information lines in the product are necessarily shown. All components may be connected to each other.

[0023] <0 System Overview>

[0024] Below, the outline of the system according to the present disclosure will be described, but the following description should not be construed in a limiting manner, and the contents of the present disclosure should be understood based on the disclosure of this specification and the ordinary technical knowledge and common sense of a person of ordinary skill in the art.

[0025] The system according to the present disclosure is a system for determining activity information of a compound for a specific receptor by computer simulation. In this specification, "computer simulation" refers to estimating activity information of a compound for a specific receptor on a computer, unlike a method for determining activity information of a compound for a specific receptor by experiment. Such a method is called in silico.

[0026] In an embodiment described later, as a method of computer simulation, which is an in silico analysis, a machine learning model is generated in advance, in which information for identifying a compound and structural information of a compound identified by this information are used as explanatory variables, and activity information of the compound against a specific receptor is used as a target variable, and information for identifying a compound to be screened and structural information of a compound identified by this information are input to this machine learning model, activity information of the compound against a specific receptor is obtained as an inference result, and screening of the compound is performed based on the obtained activity information. Note that creating a so-called rule base for the relationship between the structural information of a compound and the activity information of the compound and obtaining activity information of the compound based on this rule base is also included in one method of computer simulation, which is the in silico analysis described above, and there is no intention to exclude such a method in the system according to the present disclosure.

[0027] As a background of the system according to the present disclosure, a case of screening functional ingredients (compounds) used in foods (including food, beverages, and supplements) will be described as an example. Currently, there are foods containing a wide variety of functional ingredients. Foods containing functional ingredients that are useful to the human body are approved by the Consumer Affairs Agency as, for example, foods for specified health uses, and are labeled on the food after submitting a notification to the Consumer Affairs Agency as functional foods. However, there are a huge number of compounds (for example, more than tens of thousands of types) that are candidates for functional ingredients, and if the above-mentioned foods for specified health uses, etc. are newly developed, it is very difficult to simply screen for compounds with the desired functionality. Screening includes not only a method of selecting compounds by confirming activity in vitro, but also so-called virtual screening, which selects compounds based on data.

[0028] When a compound (hereinafter, appropriately referred to as a "ligand") binds to a receptor in a cell of a living body, the ligand has a specific activity on the receptor (the active concentration is equal to or higher than a certain level). Therefore, by estimating the binding between the ligand and the receptor by computer simulation, the receptor activity of the compound on the receptor can be estimated. The system according to the present disclosure has been realized based on such knowledge.

[0029] In general, a ligand broadly means a substance that specifically interacts with a protein, such as a substrate for an enzyme, a low molecular weight compound that interacts with a receptor protein, or a coenzyme or a regulatory factor. In some cases, the ligand may be interpreted as being limited to a substance that binds to a receptor present on a cell membrane or an intracellular receptor. However, in this specification, the term "ligand" is used in a broad sense to include a substance that specifically interacts with a protein, including a substrate for an enzyme, a coenzyme, a regulatory factor, a substance that binds to a receptor, etc. Thus, a ligand may be either a low molecular weight compound or a high molecular weight compound, or may mean a partial region of a compound.

[0030] There is already software that predicts whether a receptor will be activated (CzeekS: https: / / www.intage-healthcare.co.jp / service / data-science / insilico / czeeks / ). There is also software that simulates the binding free energy between receptors and compounds (MOE: https: / / www.molsis.co.jp / lifescience / moe / ).

[0031] Therefore, the inventor of the system according to the present disclosure simulated the receptor activity of the capsaicin receptor TRPV1 of candidate compounds, including known compounds having receptor activity against the capsaicin receptor TRPV1, using CzeekS and MOE, and the detection probability was low for MOE (details will be described later). In addition, as described above, CzeekS is software that predicts whether the receptor is activated or not, and it is not possible to know the concentration of the compound required for receptor activation.

[0032] In the system according to the present disclosure, as one method of computer simulation, which is an in silico analysis, a machine learning model is generated in advance, in which information for identifying a compound and structural information of the compound identified by this information are used as explanatory variables, and activity information of the compound against a specific receptor is used as a target variable. Information for identifying a compound to be screened and structural information of the compound identified by this information are input to this machine learning model, activity information of the compound against a specific receptor is obtained as an inference result, and compounds are screened based on the obtained activity information.

[0033] The machine learning model in the system according to the present disclosure is obtained by having the machine learning model perform machine learning based on teacher data in accordance with a model learning program.

[0034] The machine learning algorithm used in the system of the present disclosure may be any algorithm that constructs a known learning model such as a general regression model or classification model, and examples of such regression models include Gaussian process regression, linear regression, neural networks, and decision trees (such as LightGBM: Light Gradient Boosting Machine).

[0035] The learning model according to the system of the present disclosure is, for example, a parameterized composite function in which multiple functions are combined. The parameterized composite function is defined by a combination of multiple adjustable functions and parameters. The prediction model according to the present disclosure may be any parameterized composite function that satisfies the above requirements, but is assumed to be a multi-layer network model (hereinafter referred to as a multi-layer network). A prediction model using a multi-layer network has an input layer, an output layer, and at least one intermediate layer or hidden layer provided between the input layer and the output layer. The prediction model is expected to be used as a program module that is part of artificial intelligence software.

[0036] As the multi-layered network according to the present disclosure, for example, a deep neural network (DNN), which is a multi-layered neural network that is the subject of deep learning, may be used. As the DNN, for example, a convolution neural network (CNN) that targets images may be used.

[0037] In this disclosure, the terms "machine learning" and "deep learning" are used without any particular distinction between them, that is, the term "machine learning" includes the term "deep learning".

[0038] Therefore, the system according to the present disclosure can simulate the receptor activity of existing compounds (particularly compounds derived from food) against known (target) receptors without biological experiments, thereby making it possible to screen existing compounds based on their receptor activity.

[0039] In particular, since the compound concentration required for receptor activity can be obtained by simulation, quantitative evaluation is possible and compounds can be screened in order of concentration.

[0040] <One embodiment> <1 Overall system configuration> FIG. 1 is a diagram showing the overall configuration of an experiment result analysis system (hereinafter simply referred to as the "system") 1 of this embodiment. As shown in FIG. 1, the system 1 includes a plurality of terminal devices (terminal devices 10A and 10B are shown in FIG. 1. Hereinafter, they may be collectively referred to as "terminal devices 10"), a server 20, and an external server 40. The terminal devices 10, the server 20, and the external server 40 are connected to each other so as to be able to communicate with each other via a network 80. The network 80 is configured as a wired or wireless network. In this embodiment, the server 20 is a server having a function as a Web server (including a cloud server), and exchanges information with the terminal device 10 through Web pages. In addition, a Web page browser for browsing Web pages is installed in the terminal device 10, but a dedicated application for providing the services of the server 20 may be installed and configured to be able to be browsed through the dedicated application. Alternatively, the terminal device 10 and the server 20 may be connected through a so-called intranet, and the server 20 may be realized as an on-premise server. In other words, the server 20 may be a so-called cloud server, or may not be so.

[0041] The terminal device 10 is a device operated by a user who screens compounds (ligands) using the system 1 according to the present disclosure. The terminal device 10 is realized by a stationary PC (Personal Computer), a laptop PC, or the like, which are examples of information processing devices. In addition, the terminal device 10 may be a mobile terminal such as a tablet compatible with a mobile communication system or a smartphone, which are also examples of information processing devices.

[0042] The terminal device 10 is communicatively connected to the server 20 via a network 80. The terminal device 10 is connected to the network 80 by communicating with communication devices such as a wireless base station 81 compatible with communication standards such as 4G, 5G, and LTE (Long Term Evolution), and a wireless LAN router 82 compatible with wireless LAN (Local Area Network) standards such as IEEE (Institute of Electrical and Electronics Engineers) 802.11. As shown in FIG. 1, the terminal device 10 includes a communication IF (Interface) 12, an input device 13, an output device 14, a memory 15, a storage unit 16, and a processor 19.

[0043] The communication IF 12 is an interface for inputting and outputting signals so that the terminal device 10 can communicate with an external device. The input device 13 is an input device (e.g., a keyboard, a touch panel, a touch pad, a pointing device such as a mouse, etc.) for receiving an input operation from a user. The output device 14 is an output device (a display, a speaker, etc.) for presenting information to a user. The memory 15 is for temporarily storing a program and data processed by the program, etc., and is a volatile memory such as a DRAM (Dynamic Random Access Memory). The storage unit 16 is a storage device for saving data, and is, for example, a flash memory or a HDD (Hard Disc Drive). The processor 19 is hardware for executing an instruction set described in a program, and is composed of an arithmetic unit, a register, a peripheral circuit, etc.

[0044] The server 20 is managed by an administrator of the system 1 of this embodiment, and the stored contents are appropriately modified / added / deleted by users of the terminal devices 10 .

[0045] The server 20 is a computer that is an example of an information processing device connected to the network 80. The server 20 includes a communication IF 22, an input / output IF 23, a memory 25, a storage 26, and a processor 29.

[0046] The communication IF 22 is an interface for inputting and outputting signals so that the server 20 can communicate with an external device. The input / output IF 23 functions as an interface with an input device for receiving input operations from a user and an output device for presenting information to a user. The memory 25 is for temporarily storing programs and data processed by the programs, etc., and is a volatile memory such as a DRAM (Dynamic Random Access Memory). The storage 26 is a storage device for saving data, such as a flash memory or a HDD (Hard Disc Drive). The processor 29 is hardware for executing an instruction set described in a program, and is composed of an arithmetic unit, a register, a peripheral circuit, etc.

[0047] The external server 40 is a computer (information processing device) separate from the terminal device 10 and the server 20, and various databases that serve as sources of data acquired by the server 20 are stored in this external server 40. An example of the database in this embodiment is ChEMBL, which is a database of open data on biologically active small molecules such as pharmaceuticals and pharmaceutical candidate compounds. In addition, various libraries for machine learning models may be stored in this external server 40.

[0048] <1.1 Functional configuration of the terminal device 10> FIG. 2 is a block diagram showing a functional configuration of the terminal device 10 constituting the system 1 of the first embodiment. As shown in FIG. 2, the terminal device 10 includes a plurality of antennas (antenna 111, antenna 112), wireless communication units (first wireless communication unit 121, second wireless communication unit 122) corresponding to the respective antennas, an operation reception unit 130 (including a keyboard 131 and a mouse 132), a voice processing unit 140, a microphone 141, a speaker 142, a display 150, a storage unit 170, and a control unit 180. The terminal device 10 also has functions and configurations (for example, a battery for storing power, a power supply circuit for controlling the supply of power from the battery to each circuit, etc.) that are not particularly shown in FIG. 2. As shown in FIG. 2, each block included in the terminal device 10 is electrically connected by a bus or the like.

[0049] The antenna 111 emits a signal generated by the terminal device 10 as a radio wave. The antenna 111 also receives a radio wave from space and provides the received signal to the first wireless communication unit 121.

[0050] The antenna 112 radiates a signal generated by the terminal device 10 as a radio wave. The antenna 112 also receives the radio wave from space and provides the received signal to the second radio communication unit 122.

[0051] The first wireless communication unit 121 performs modulation / demodulation processing and the like for transmitting and receiving signals via the antenna 111 so that the terminal device 10 can communicate with other wireless devices. The second wireless communication unit 122 performs modulation / demodulation processing and the like for transmitting and receiving signals via the antenna 112 so that the terminal device 10 can communicate with other wireless devices. The first wireless communication unit 121 and the second wireless communication unit 122 are communication modules including a tuner, a Received Signal Strength Indicator (RSSI) calculation circuit, a Cyclic Redundancy Check (CRC) calculation circuit, a high-frequency circuit, and the like. The first wireless communication unit 121 and the second wireless communication unit 122 perform modulation / demodulation and frequency conversion of wireless signals transmitted and received by the terminal device 10, and provide the received signals to the control unit 180.

[0052] The operation reception unit 130 has a mechanism for receiving input operations from a user. Specifically, the operation reception unit 130 includes a keyboard 131 and a mouse 132. Note that the operation reception unit 130 may be configured as a touch screen that detects the user's contact position on the touch panel, for example, by using a capacitive touch panel.

[0053] The keyboard 131 accepts input operations by the user of the terminal device 10. The keyboard 131 is a device for inputting characters, and outputs input character information to the control unit 180 as an input signal.

[0054] The mouse 132 accepts input operations by the user of the terminal device 10. The mouse 132 is a pointing device for selecting an object displayed on the display 150, and outputs position information of an object selected on the screen and information indicating that a button has been pressed to the control unit 180 as input signals.

[0055] The audio processing unit 140 modulates and demodulates an audio signal. The audio processing unit 140 modulates a signal provided from the microphone 141 and provides the modulated signal to the control unit 180. The audio processing unit 140 also provides the audio signal to the speaker 142. The audio processing unit 140 is realized by, for example, a processor for audio processing. The microphone 141 accepts audio input and provides an audio signal corresponding to the audio input to the audio processing unit 140. The speaker 142 converts the audio signal provided from the audio processing unit 140 into audio and outputs the audio to the outside of the terminal device 10.

[0056] Display 150 displays data such as images, videos, and text under the control of control unit 180. Display 150 is realized by, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.

[0057] The storage unit 170 is configured with, for example, a flash memory, and stores data and programs used by the terminal device 10.

[0058] The control unit 180 reads a program stored in the storage unit 170 and executes instructions included in the program to control the operation of the terminal device 10. The control unit 180 is, for example, an application that is pre-installed in the terminal device 10. The control unit 180 operates according to the program to fulfill the functions of an input operation reception unit 1801, a transmission / reception unit 1802, a data processing unit 1803, and a notification control unit 1804.

[0059] The input operation receiving unit 1801 performs processing for receiving input operations by a user via an input device such as the keyboard 131 .

[0060] The transmitting / receiving unit 1802 performs processing for the terminal device 10 to transmit and receive data to and from an external device such as the server 20 in accordance with a communication protocol.

[0061] The data processing unit 1803 performs calculations on the data input by the terminal device 10 according to a program, and outputs the calculation results to a memory or the like.

[0062] The notification control unit 1804 performs processing to present information to the user. The notification control unit 1804 performs processing to display a display image on the display 150, processing to output sound to the speaker 142, and the like.

[0063] <1.2 Functional configuration of server 20> 3 is a diagram showing an example of a functional configuration of the server 20. As shown in FIG. 3, the server 20 fulfills the functions of a communication unit 201, a storage unit 202, and a control unit 203.

[0064] The communication unit 201 performs processing for the server 20 to communicate with external devices.

[0065] The storage unit 202 stores data and programs used by the server 20. The storage unit 202 stores an experiment library database (DB: DataBase) 2022, teacher data 2023, a machine learning model 2024, and the like.

[0066] The experimental library DB2022 is a database that stores experimental values ​​obtained for compounds (ligands) to be screened and activity concentrations of the compounds estimated by the system 1 according to this embodiment. In the system 1 according to this embodiment, the experimental library DB2022 is provided for each receptor, and if the system 1 screens compounds for a single receptor, a single experimental library DB2022 is stored in the storage unit 202, and if the system 1 screens compounds for a plurality of receptors, the same number of experimental library DB2022 as the number of corresponding receptors are stored in the storage unit 202.

[0067] In the system 1 according to the present embodiment, the activity concentration of the compound against the receptor is used as the activity information of the compound. Naturally, information other than the activity concentration, such as the activity information of the compound against the receptor, may also be used.

[0068] In the system 1 according to this embodiment, the experimental library DB2022 includes at least information for identifying a compound, structural information of the compound identified by this information, and the active concentration of the compound for a specific receptor. The information for identifying a compound and the structural information of the compound identified by this information are experimental or measured values, and the active concentration of the compound for a specific receptor is an estimated value by the system 1 according to this embodiment. Of the information stored in the experimental library DB2022, the information for identifying a compound and the structural information of the compound identified by this information are obtained from the above-mentioned external server 40, preferably ChEMBL. It goes without saying that this information may be obtained from an external database other than ChEMBL.

[0069] In the system 1 according to this embodiment, the teacher data 2023 also includes at least information for identifying a compound, structural information of the compound identified by this information, and the active concentration of the compound for a specific receptor, similar to the experimental library DB 2022. However, the active concentration of the compound in the teacher data 2023 is an experimental value. In the system 1 according to this embodiment, the teacher data 2023 is also provided on a receptor-by-receptor basis.

[0070] The machine learning model 2024 is a machine learning model that has been subjected to machine learning using the teacher data 2023. In the system 1 of this embodiment, the machine learning model 2024 stored in the storage unit 202 is a machine learning model that has been subjected to machine learning for a plurality of candidate machine learning models and selected based on their prediction accuracy, as described below. In the system 1 according to this embodiment, the machine learning model 2024 is also provided on a receptor-by-receptor basis.

[0071] The experiment library DB 2022 and the training data 2023 will be described in detail later.

[0072] The control unit 203 performs functions indicated by various modules, such as a reception control module 2031, a transmission control module 2032, a data acquisition module 2033, a model generation module 2034, a model evaluation and selection module 2035, a screening module 2036, and a presentation control module 2037, by the processor of the server 20 performing processing in accordance with the application program 2021 stored in the memory unit 202.

[0073] The reception control module 2031 controls the process in which the server 20 receives a signal from an external device in accordance with a communication protocol.

[0074] The transmission control module 2032 controls the process in which the server 20 transmits signals to external devices in accordance with a communication protocol.

[0075] As an example, the data acquisition module 2033 searches the external server 40 based on predetermined conditions, and acquires data that is the source of the experiment library DB 2022 and the teacher data 2023 from the external server 40. The data acquisition module 2033 then performs an operation to change the format of the data acquired from the external server 40, and performs processing suitable for a screening operation described later, after which it generates the experiment library DB 2022 and the teacher data 2023 and stores them in the memory unit 202.

[0076] Similarly, the data acquisition module 2033 acquires a machine learning algorithm for generating a machine learning model from a library of the external server 40. The machine learning algorithm acquired by the data acquisition module 2033 is provided to a model generation module 2034, which will be described later.

[0077] In the server 20 of this embodiment, the data acquisition module 2033 acquires multiple machine learning algorithms, and generates multiple machine learning models (trained models) as a result of machine learning performed by the model generation module 2034 described later using teacher data 2023. Then, the model evaluation and selection module 2035 calculates the prediction accuracy of the multiple generated candidate machine learning models, and based on this prediction accuracy, selects a machine learning model to be actually used for screening, and stores this machine learning model 2024 in the storage unit 202. In this specification, the multiple machine learning algorithms acquired by the data acquisition module 2033 from the library of the external server 40 are referred to as candidate machine learning models.

[0078] The model generation module 2034 uses the teacher data 2023 to perform machine learning using the multiple candidate machine learning models acquired by the data acquisition module 2033, thereby generating trained candidate machine learning models.

[0079] The applicable methods of machine learning algorithms vary depending on the data to be analyzed and its purpose. In machine learning, it is usually necessary to set multiple values ​​called hyperparameters, and the accuracy of the results changes depending on these settings. Generally, there is no "machine learning algorithm that can achieve the best accuracy for every problem," and it has been necessary for data scientists to have high levels of skill and experience to choose from the many machine learning algorithms available and how to tune the hyperparameters.

[0080] Recently, a technology called AutoML (Automated Machine Learning) has been proposed and realized to automate various tasks carried out in analysis using machine learning. Many automated machine learning tools optimize the settings of hyperparameters for machine learning algorithms, and also compare multiple machine learning algorithms to automatically find the algorithm that produces the most accurate results. (The above explanation is taken from https: / / www.nri.com / jp / knowledge / glossary / lst / sa / automl)

[0081] The model generation module 2034 of this embodiment uses an automated machine learning tool stored in the external server 40 to optimize the parameter setting values ​​used when machine learning is performed using the teacher data 2023 for multiple candidate machine learning models acquired by the data acquisition module 2033.

[0082] Naturally, the optimization process of parameter values ​​using an automated machine learning tool is optional in the system 1 according to the present disclosure, and it does not preclude a data scientist (the developer / operator of the server 20 in this embodiment) from optimizing parameter values ​​manually as in the past.

[0083] The model evaluation and selection module 2035 calculates the predictive accuracy of multiple trained candidate machine learning models generated by the model generation module 2034, and based on the calculated predictive accuracy, selects a machine learning model to be used on the server 20 (i.e., to be used for screening work), and stores it in the memory unit 202 as a machine learning model 2024.

[0084] The evaluation of candidate machine learning models and the selection of the machine learning model 2024 by the model evaluation and selection module 2035 may be appropriately selected from well-known procedures, but an example will be described below.

[0085] First, the model evaluation and selection module 2035 inputs information for identifying a compound and structural information of a compound identified by this information, which are included in the training data 2023, into each trained candidate machine learning model, and obtains the activity concentration of the compound, which is the inference result of these candidate machine learning models. Next, the model evaluation and selection module 2035 calculates a correlation coefficient between the activity concentration of the compound, which is the inference result, and the experimental value of the activity concentration of the compound included in the training data 2023. Then, the model evaluation and selection module 2035 regards the correlation coefficient calculated based on the candidate machine learning model as the prediction accuracy, adopts the candidate machine learning model that has provided the best correlation coefficient as the machine learning model to be used in the server 20, and stores it in the storage unit 202.

[0086] Alternatively, the model evaluation and selection module 2035 may calculate the prediction accuracy of the candidate machine learning model using a known index for calculating the prediction accuracy of the machine learning model. Examples of such known indexes include the mean absolute error (MAE) and the mean squared error (MSE). The predicted value of MAE, MSE, etc. may be the activity concentration of the compound, which is the inference result, and the actual measured value may be the experimental value of the activity concentration of the compound.

[0087] There are a number of known public indices, and when a machine learning model 2024 is selected using a number of known indices, the same trend of prediction accuracy is not necessarily obtained from each of the indices. In other words, a candidate machine learning model that is deemed to have the best prediction accuracy for one index may not be deemed to have the best prediction accuracy for another index. The candidate machine learning model in this embodiment is basically a regression analysis model, and the above-mentioned known indices are premised on a certain correlation between predicted values ​​and actual values, so it is unlikely that a contradictory relationship will occur in the prediction accuracy based on the multiple indices. However, when the same trend of prediction accuracy is not obtained from the multiple indices, the machine learning model 2024 may be selected by prioritizing one of the indices.

[0088] Alternatively, the candidate machine learning model with the highest predictive accuracy may be selected using the automated machine learning tool described above, and the selected candidate machine learning model may be stored in the storage unit 202 as the machine learning model 2024.

[0089] In addition, the model generation module 2034 may divide all data included in the teacher data 2023 into data for training (machine learning) and testing (verification), and may use the training teacher data 2023 to learn the candidate machine learning model, and the model evaluation and selection module 2035 may recalculate the prediction accuracy using the testing teacher data 2023 to check whether the expected prediction accuracy has been obtained.

[0090] The screening module 2036 uses the machine learning model 2024 selected by the model evaluation and selection module 2035 to screen compounds stored in the experimental library DB 2022. Specifically, the screening module 2036 inputs information for identifying compounds and structural information of compounds identified by this information, which are stored in the experimental library DB 2022, to the machine learning model 2024, and obtains the activity concentration of the child compound as an inference result. The screening module 2036 stores the obtained activity concentration of the compound in the experimental library DB 2022.

[0091] The screening module 2036 performs the above-mentioned operation for a plurality of compounds to obtain the activity concentrations of the compounds, and screens the compounds based on the obtained activity concentrations of the compounds. As already described, the experimental library DB 2022, the teacher data 2023, and the machine learning model 2024 are provided for each receptor, and therefore the screening operation by the screening module 2036 is a screening operation for each specific receptor.

[0092] There are no particular limitations on the specific content of the screening operation by the screening module 2036 based on the acquired compound activity concentration, and suitable examples include determining the compound with the highest compound activity concentration as the screening result, sorting the compound activity concentrations in ascending order and determining the compound in a predetermined rank, for example, the top 5, as the screening result.

[0093] The presentation control module 2038 generates a display control signal for displaying a predetermined screen on the display 150 of the terminal device 10, transmits it to the terminal device 10, and acquires an operation input input by a user operating the operation reception unit 130 of the terminal device 10. The presentation control module 2038 also generates a display control signal for displaying the contents of the experiment library DB 2022 and the teacher data 2023, or for displaying the input / output contents of the model generation module 2034, the model evaluation / selection module 2035, and the screening module 2036, transmits it to the terminal device 10, and causes the display 150 to display data, etc.

[0094] <2 Data Structure> 4 and 5 are diagrams showing the data structure of a database stored in the server 20. Note that Fig. 4 and Fig. 5 are merely examples and do not exclude data that is not listed. In addition, even data that is listed in the same table may be stored in separate storage areas in the storage unit 202.

[0095] The databases shown in Figures 4 and 5 are relational databases, which are used to manage data sets called tables in a tabular format that are structurally defined by rows and columns, by associating them with each other. In a database, a table is called a table, a column in a table is called a column, and a row in a table is called a record. In a relational database, it is possible to set relationships between tables and associate them.

[0096] Usually, a column that serves as a primary key for uniquely identifying a record is set in each table, but setting a primary key to a column is not essential. The control unit 203 of the server 20 can cause the processor 29 to add, delete, or update records in a specific table stored in the storage unit 202 according to various programs.

[0097] Fig. 4 is a diagram showing an example of the data structure of the experimental library DB2022. As shown in Fig. 4, each record of the experimental library DB2022 includes, for example, an item "SMILES notation", an item "compound name", an item "CAS number", an item "molecular weight", and an item "receptor activity concentration". Each item of the experimental library DB2022 is acquired by the data acquisition module 2033 from an external server 40 (ChEMBL as an example), and the data is converted as necessary before being stored by the data acquisition module 2033. The information stored in the experimental library DB2022 can be changed and updated as appropriate.

[0098] The item "SMILES notation" is structural information of a compound to be screened in the system 1 (more precisely, the server 20) of this embodiment. SMILES notation (simplified molecular input line entry system) is a notation method that does not ambiguize the structure, in which the chemical structure of a molecule is converted into a string of alphanumeric characters in ASCII code. Since the SMILES notation itself is a string, when it is actually input as an explanatory variable to the machine learning model 2024 (including a candidate machine learning model), the model generation module 2034 and the model evaluation and selection module 2035 convert the SMILES notation into a numerical value and input it to the machine learning model 2024, etc. The method of converting the SMILES notation into a numerical value is well known, and any method can be adopted, but the server 20 of this embodiment uses a method called Morgan Fingerprints.

[0099] Morgan Fingerprint is a kind of circular fingerprint that counts the number of substructures located at a certain distance from each atom. A (numerical) conversion library from SMILES notation to Morgan Fingerprints values ​​is, for example, available as an open library in the RDKit, a software for cheminformatics.

[0100] The item "Compound name" is the name of a compound to be screened in the system 1 of this embodiment. The item "Compound name" is used as a label or identification information of a compound in the experimental library DB2022 (and the teacher data 2023 described later), so there is no particular limitation on the naming method as long as the naming method of compounds is unified in one experimental library DB2022. In the experimental library DB2022 of this embodiment, as an example, the EC 50 (The concentration at which a drug, antibody, etc. shows 50% of the maximum response from the minimum value) The names of compounds included in the data are used.

[0101] The item "CAS number" is a globally standardized identification number used worldwide that is unique to each chemical substance and is managed by the Chemical Abstracts Service (CAS), a subsidiary of the American Chemical Society.

[0102] The item "molecular weight" is information indicating the molecular weight of the compound identified by the item "SMILES notation" or the item "compound name".

[0103] The item "receptor activity concentration" is information indicating the activity concentration of the target receptor for the compound identified by the item "SMILES notation" or the item "compound name." Note that since the experimental library DB2022 is for the compound to be screened, the cell of the item "receptor activity concentration" may be blank. If the item "receptor activity concentration" is blank, the screening module 2036 makes an estimate using the machine learning model 2024 and stores the estimated value.

[0104] Fig. 5 is a diagram showing the data structure of the teacher data 2023. The data structure of the teacher data 2023 shown in Fig. 5 is the same as the data structure of the experimental library DB 2022 shown in Fig. 4, but for the teacher data 2023, all cells of the item "receptor activity concentration" store experimental values.

[0105] <3 Example of operation> An example of the operation of the server 20 will be described below. Note that the order of operations of each step shown in the flowchart showing an example of the operation of the server 20 is not limited to that shown in the figure, and the order of operations can be changed as appropriate. It should be noted that not all steps are essential to the system 1 according to the present disclosure.

[0106] (Machine learning model generation) FIG. 6 is a flowchart for explaining an example of the operation of the server 20 when generating a machine learning model 2024 to be stored in the storage unit 202 of the server 20 of this embodiment.

[0107] First, in step S600, the control unit 203 of the server 20 acquires data to be stored in the experiment library DB 2022 from the external server 40. Specifically, for example, the control unit 203 searches the external server 40 under predetermined conditions using the data acquisition module 2033, acquires data that is the search result from the external server 40, performs predetermined data processing as necessary, and generates the experiment library DB 2022 based on the acquired data.

[0108] Next, in step S601, the control unit 203 divides the data acquired in step S600 into data for training the machine learning model 2024 and data for testing the machine learning model 2024. Specifically, for example, the control unit 203 divides the data acquired in step S600 by the model generation module 2034 into data for training the machine learning model 2024 and data for testing the machine learning model 2024. The division is basically performed based on the number of data.

[0109] Next, in step S602, the control unit 203 uses the data divided for training to execute a machine learning algorithm on the multiple candidate machine learning models with an automated machine learning tool, and optimizes parameters of each of the candidate machine learning models. Specifically, for example, the control unit 203 uses the data divided for training with the model generation module 2034 to execute a machine learning algorithm on the multiple candidate machine learning models with an automated machine learning tool, and optimizes parameters of each of the candidate machine learning models. The automated machine learning tool may be one provided in the external server 40.

[0110] Next, in step S603, the control unit 203 obtains the prediction accuracy of each of the candidate machine learning models after performing parameter optimization for the multiple candidate machine learning models, which are the execution results of the operation of step S602, and determines the machine learning model 2024 to be actually stored in the storage unit 202 of the server 20 of this embodiment based on the prediction accuracy. Specifically, for example, the control unit 203 obtains the prediction accuracy of each of the candidate machine learning models after performing parameter optimization for the multiple candidate machine learning models, which are the execution results of the operation of step S602, by the model evaluation and selection module 2035, and determines the machine learning model 2024 to be actually stored in the storage unit 202 of the server 20 of this embodiment based on the prediction accuracy. The determined machine learning model 2024 is stored in the storage unit 202 by the model evaluation and selection module 2035.

[0111] Then, in step S604, the control unit 203 inputs the divided data for testing into the candidate machine learning model selected in step S603, outputs a predicted value of the activity concentration of the compound as an inference result, calculates a correlation coefficient between the actual measured value of the activity concentration of the compound contained in the test data, and evaluates the validity of the selected candidate machine learning model (machine learning model 2024).

[0112] (Screening process) FIG. 7 is a flowchart for explaining an example of the operation of the server 20 when performing a compound screening process using the server 20 of this embodiment.

[0113] First, in step S700, the control unit 203 acquires data to be screened from the experiment library DB 2022 stored in the storage unit 202. Specifically, for example, the control unit 203 acquires data to be screened from the experiment library DB 2022 stored in the storage unit 202 by the screening module 2036.

[0114] Next, in step S701, the control unit 203 extracts compound identification information and compound structural information from the data acquired in step S700, and inputs them to the machine learning model 2024 stored in the storage unit 202. Specifically, for example, the control unit 203 extracts compound identification information and compound structural information from the data acquired in step S700 by the screening module 2036, and inputs them to the machine learning model 2024 stored in the storage unit 202.

[0115] Next, in step S702, the control unit 203 acquires the active concentration of the compound obtained as an inference result as a result of inputting the data to the machine learning model 2024 in step S701. Specifically, for example, the control unit 203 acquires the active concentration of the compound obtained as an inference result as a result of inputting the data to the machine learning model 2024 in step S701 by the screening module 2036.

[0116] Next, in step S703, the control unit 203 stores the inferred value of the activity concentration of the compound obtained in step S702 in the experiment library DB 2022. Specifically, for example, the control unit 203 stores the inferred value of the activity concentration of the compound obtained in step S702 in the experiment library DB 2022 by the screening module 2036.

[0117] Then, in step S704, the control unit 203 performs screening of the compound based on the inferred value of the active concentration of the compound stored in the experiment library DB 2022 in step S703. Specifically, for example, the control unit 203 performs screening of the compound by the screening module 2036 based on the inferred value of the active concentration of the compound stored in the experiment library DB 2022 in step S703.

[0118] <4. Effects of one embodiment> As described above in detail, according to the system 1 of the present embodiment, a machine learning model is prepared in which the structural information of a compound is used as an explanatory variable and the activity information of a compound against a specific receptor is used as a target variable, information for identifying a compound to be screened and the structural information of a compound identified by the information are input to the machine learning model, activity information of the compound against a specific receptor is obtained as an inference result, and screening of the compound is performed based on the obtained activity information, so that it is possible to simulate the receptor activity of an existing compound (particularly a compound derived from food) against a known (target) receptor without performing biological experiments. This allows existing compounds to be screened based on their receptor activity.

[0119] In other words, according to the system 1 of this embodiment, activity information of a compound for a specific receptor can be easily simulated, and thus compounds useful for a specific receptor can be easily screened.

[0120] In particular, since the compound concentration required for receptor activity can be obtained by simulation, quantitative evaluation is possible and compounds can be screened in order of concentration.

[0121] In addition, in system 1 of this embodiment, multiple candidate machine learning models are prepared, and the prediction accuracy of the candidate machine learning models is calculated using the estimated values ​​of receptor activity concentrations output from each candidate machine learning model and the known receptor activity concentrations, and the machine learning model to be actually used is selected based on this prediction accuracy, thereby further improving the accuracy of the predicted values ​​of receptor activity concentrations.

[0122] In addition, data for training a candidate machine learning model is divided into training data and test data, the candidate machine learning model is trained using the training data, the test data is input to a machine learning model selected based on its predictive accuracy, and the correlation coefficient between the inference value, which is the inference result, and the actual measured value of the active concentration of the receptor contained in the test data is calculated, thereby confirming that the selected machine learning model actually has effective predictive accuracy. This makes it possible to guarantee the predictive accuracy of the machine learning model. Experimental example

[0123] <5 Experimental example using capsaicin receptor TRPV1> Below, an experimental example will be described in which the capsaicin receptor TRPV1 is used as the receptor.

[0124] EC of TRPV1 registered in ChEMBL as training data 2023 50 From the data, compound names, structural information of the compounds (SMILES information), and active concentrations of the compounds were extracted.

[0125] Among the data extracted from ChEMBL, data with ambiguous information on the active concentration (activity information) was deleted. In other words, data with no specific concentration information, such as an active concentration of >100 μM, was deleted, and only data with a specific value, such as 100 μM, was used for further processing. As a result, the number of data obtained from ChEMBL as training data 2023 was 338.

[0126] The SMILES information obtained from ChEMBL was converted into Morgan Fingerprints values ​​using a library written in python (a programming language) available on RDkit.

[0127] The teacher data 2023 was split into 270 pieces of training data and 68 pieces of test data. The 270 pieces of training data generated by splitting were used to perform machine learning on 19 candidate machine learning models, each with a different algorithm. All 19 candidate machine learning models were written in Python. The actual machine learning was performed using pycaret, an automated machine learning tool. Machine learning was performed on the 19 machine learning models using the automated machine learning tool, and the parameters of each candidate machine learning model were optimized.

[0128] Next, using pycaret in the same way, the prediction accuracy of 19 candidate machine learning models was calculated using seven indices such as MAE and MST, and one candidate machine learning model was selected based on the calculated prediction accuracy, and this was used as the machine learning model to be used for actual screening. The machine learning model selected was the Bayesian linear regression model (Bayesian Ridge). In this experimental example, the Bayesian linear regression model provided the best prediction accuracy for all indices.

[0129] Next, 68 pieces of test data were used to estimate the receptor activity concentration for a Bayesian linear regression model, a trained machine learning model, and the correlation coefficient was calculated between the estimated value and the actual measured value of the receptor activity concentration contained in the test data.

[0130] A graph showing the correlation between the estimated (predicted) and actual measured values ​​of receptor activity concentrations when the training data and the test data were used is shown in Figure 8. The correlation coefficient for the training data was r = 0.90, and for the test data was r = 0.70, demonstrating that the selected machine learning model has a correlation coefficient that is suitable for practical use.

[0131] The accuracy of compound screening using in silico analysis, including the selected machine learning model, was verified using an experimental library in which active concentrations for specific receptors were already known.

[0132] The experimental library used for the verification was the experimental results of the active concentration of the capsaicin receptor TRPV1 (2000 compounds). Approximately 1% (20 compounds) of this experimental library was selected, and the number of compounds that could be screened using multiple methods including the selected machine learning model was calculated.

[0133] The multiple methods were: (1) a machine learning model selected in an experimental example (this is a machine learning model written in Python, so for convenience it is referred to as Python), (2) the method using the structure-activity relationship by CzeekS described above, (3) the method using the docking simulation by MOE described above, and (4) random selection from 20 compounds.

[0134] For (1) to (3), the compound similarity is calculated using the compound's partial structure data, and filtering is performed in advance to target only compounds with a similarity of 0.8 or more. In this experimental example, the Tanimoto coefficient was used as the compound similarity. The Tanimoto coefficient is an index obtained by converting the molecular structure of a compound into a fingerprint and calculating the similarity between fingerprints. Note that since there are multiple methods for calculating the compound similarity itself, it is not excluded to calculate the compound similarity using a method other than the Tanimoto coefficient.

[0135] The screening results for (1) to (4) are shown in Figure 9. According to the method (1) of this experimental example, the probability of detecting a compound can be improved by up to about 15 times compared to (4) random selection.

[0136] <6.1 Basic hardware configuration of computer> 10 is a block diagram showing the basic hardware configuration of a computer 90. The computer 90 includes at least a processor 901, a main storage device 902, an auxiliary storage device 903, and a communication IF 991 (interface). These are electrically connected to each other by a communication bus 921.

[0137] The processor 901 is hardware for executing an instruction set described in a program, and is composed of an arithmetic unit, a register, a peripheral circuit, and the like.

[0138] The main memory device 902 is for temporarily storing programs, data to be processed by the programs, etc. For example, it is a volatile memory such as a DRAM (Dynamic Random Access Memory).

[0139] The auxiliary storage device 903 is a storage device for saving data and programs, such as a flash memory, a hard disk drive (HDD), a magneto-optical disk, a CD-ROM, a DVD-ROM, or a semiconductor memory.

[0140] The communication IF 991 is an interface for inputting and outputting signals for communicating with other computers via a network using a wired or wireless communication standard. The network is composed of the Internet, a LAN, various mobile communication systems constructed by wireless base stations, etc. For example, the network includes 3G, 4G, 5G mobile communication systems, LTE (Long Term Evolution), wireless networks that can connect to the Internet via a specified access point (e.g., Wi-Fi (registered trademark)), etc. In the case of wireless connection, communication protocols include, for example, Z-Wave (registered trademark), ZigBee (registered trademark), Bluetooth (registered trademark), etc. In the case of wired connection, the network also includes a network that is directly connected by a USB (Universal Serial Bus) cable or the like.

[0141] It should be noted that the computer 90 can be virtually realized by distributing all or part of each hardware configuration among multiple computers 90 and connecting them together via a network. In this way, the computer 90 is a concept that includes not only a computer 90 housed in a single housing or case, but also a virtualized computer system.

[0142] <6.2 Basic functional configuration of computer 90> A description will now be given of the functional configuration of a computer realized by the basic hardware configuration (FIG. 10) of a computer 90. The computer comprises at least the functional units of a control unit, a storage unit, and a communication unit.

[0143] The functional units of the computer 90 can also be realized by distributing all or part of the functional units among multiple computers 90 connected to each other via a network. The computer 90 is a concept that includes not only a single computer 90 but also a virtualized computer system.

[0144] The control unit is realized by the processor 901 reading out various programs stored in the auxiliary storage device 903, expanding the programs in the main storage device 902, and executing processes according to the programs. The control unit can realize functional units that perform various information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.

[0145] The storage unit is realized by a main storage device 902 and an auxiliary storage device 903. The storage unit stores data, various programs, and various databases. Furthermore, the processor 901 can secure a storage area corresponding to the storage unit in the main storage device 902 or the auxiliary storage device 903 in accordance with a program. Furthermore, the control unit can cause the processor 901 to execute processes of adding, updating, and deleting data stored in the storage unit in accordance with the various programs.

[0146] The term database refers to a relational database, which is used to manage sets of data called masters and tables in a tabular format structurally defined by rows and columns, by associating them with each other. In a database, a table is called a table or master, a column in a table is called a column, and a row in a table is called a record. In a relational database, relationships between tables and masters can be set and associated. Usually, a column that serves as a primary key for uniquely identifying a record is set in each table and each master, but setting a primary key in a column is not essential. The control unit can cause the processor 901 to add, delete, or update records in a specific table or master stored in the storage unit according to various programs. Furthermore, by storing data, various programs, and various databases in the storage unit, it can be considered that the information processing device and information processing system according to the present disclosure have been manufactured.

[0147] In addition, the database and master in this disclosure may include any data structure (such as a list, a dictionary, an associative array, or an object) in which information is structurally defined. The data structure also includes data that can be considered as a data structure by combining data with a function, class, method, or the like written in any programming language.

[0148] The communication unit is realized by the communication IF 991. The communication unit realizes a function of communicating with other computers 90 via a network. The communication unit can receive information transmitted from other computers 90 and input the information to the control unit. The control unit can cause the processor 901 to execute information processing on the received information in accordance with various programs. In addition, the communication unit can transmit information output from the control unit to other computers 90.

[0149] <7 Notes> In addition, the above-described embodiments are described in detail to clearly explain the present disclosure, and are not necessarily limited to those including all of the described configurations. In addition, some of the configurations of each embodiment can be added to, deleted from, or replaced with other configurations.

[0150] As an example, in the above-described embodiment of system 1, activity information of a compound for a specific receptor (capsaicin receptor TRPV1) was computer-simulated, but machine learning models 2024 for multiple receptors may be prepared in advance, and the machine learning model 2024 may be selected depending on which receptor is to be simulated.

[0151] In addition, the above-mentioned configurations, functions, processing units, processing means, etc. may be realized in part or in whole by hardware, for example, by designing them as integrated circuits. The present invention can also be realized by software program code that realizes the functions of the embodiments. In this case, a storage medium on which the program code is recorded is provided to a computer, and a processor included in the computer reads the program code stored in the storage medium. In this case, the program code itself read from the storage medium realizes the functions of the above-mentioned embodiments, and the program code itself and the storage medium storing it constitute the present invention. Examples of storage media for supplying such program code include flexible disks, CD-ROMs, DVD-ROMs, hard disks, SSDs, optical disks, magneto-optical disks, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, etc.

[0152] Furthermore, the program code for realizing the functions described in this embodiment can be implemented in a wide range of program or script languages, such as assembler, C / C++, perl, Shell, PHP, Java (registered trademark), and the like.

[0153] Furthermore, the program code of the software that realizes the functions of the embodiments may be distributed over a network and stored in a storage means such as a computer's hard disk or memory, or in a storage medium such as a CD-RW or CD-R, and the processor of the computer may read out and execute the program code stored in the storage means or storage medium.

[0154] The matters described in the above embodiments will be supplemented below. (Appendix 1) A program (2021) for operating a computer (20) having a processor (29) and memories (25, 26), the program (2021) for screening a compound based on activity information of the compound against a receptor, the memory (25, 26) storing a machine learning model (2024) having information for identifying a compound and structural information of a compound identified by the information as explanatory variables and activity information of the compound against a specific receptor as a target variable, the program (2021) causing the processor (29) to execute a first step (S701, S702) of inputting information for identifying a compound to be screened and structural information of a compound identified by the information into the machine learning model (2024) and acquiring activity information of the compound against a specific receptor as an inference result, and a second step (S704) of screening the compound based on the activity information acquired in the first step (S701, S702). (Appendix 2) The activity information of a compound is the activity concentration of a compound having a specific value, the program described in Appendix 1 (2021). (Appendix 3) The program of claim 1, wherein the receptor is a capsaicin receptor. (Appendix 4) The machine learning model (2024) was created for each receptor, using the program (2021) described in Appendix 1. (Appendix 5) A program (2021) described in Appendix 2, in which in a first step (S701, S702), multiple pieces of information for identifying each of multiple compounds and structural information of multiple compounds identified by each piece of information are input into a machine learning model (2024), and activity concentrations of each of the multiple compounds are obtained as inference results, and in a second step (S704), multiple compounds are screened based on the activity concentrations of the multiple compounds obtained in the first step (S701, S702). (Appendix 6) A program (2021) described in Appendix 5, in which in a second step (S704), at least one compound is selected from the multiple compounds based on the range of activity concentration values ​​of the multiple compounds obtained in the first step (S701, S702). (Appendix 7) In a second step (S704), at least one compound is selected from the multiple compounds based on the magnitude of the activity concentration values ​​of the multiple compounds obtained in the first step (S701, S702). A program (2021) described in Appendix 5. (Appendix 8) The program (2021) described in Appendix 1, wherein the information for identifying the compound is information for identifying structural information of the compound. (Appendix 9) The program (2021) further causes the processor (29) to execute a third step (S602) of performing machine learning for each of a plurality of candidate machine learning models (2024) using structural information of the compound and activity information of the compound against a specific receptor, each of which has a known value, as training data; a fourth step (S602) of inputting the structural information of the compound into each of the plurality of candidate machine learning models for which machine learning was performed in the third step (S602), obtaining activity information of the compound as an inference result, and calculating the predictive accuracy of the candidate machine learning model using the obtained activity information of the compound and activity information of a compound whose value corresponding to the input structural information of the compound is known; and a fifth step (S603) of determining the machine learning model (2024) to be stored in the memory (25, 26) based on the predictive accuracy calculated in the fourth step (S602). (Appendix 10) A program (2021) described in Appendix 9, in which in a fourth step (S602), a correlation coefficient is calculated between the activity information of the acquired compound and activity information of a compound having a known value corresponding to the structural information of the input compound, and this correlation coefficient is used as the prediction accuracy, and in a fifth step (S603), the candidate machine learning model having the maximum correlation coefficient is determined as the machine learning model (2024) to be stored in the memory (25, 26). (Appendix 11) The information processing device (20) includes a processor (29) and memories (25, 26), and the information processing device (20) screens compounds based on activity information of the compounds against receptors. The memories (25, 26) store a machine learning model (2024) in which information for identifying a compound and structural information of a compound identified by the information are used as explanatory variables, and activity information of the compound against a specific receptor is used as a target variable. The processor (29) inputs the information for identifying a compound to be screened and the structural information of the compound identified by the information into the machine learning model (2024) and executes a first step (S701, S702) of acquiring activity information of the compound against a specific receptor as an inference result, and a second step (S704) of screening compounds based on the activity information acquired in the first step (S701, S702). (Appendix 12) A method executed by a computer (20) having a processor (29) and memories (25, 26), the method screening a compound based on activity information of the compound against a receptor, the memory (25, 26) storing a machine learning model (2024) having information for identifying a compound and structural information of a compound identified by the information as explanatory variables and activity information of the compound against a specific receptor as a target variable, the processor (29) inputs the information for identifying a compound to be screened and the structural information of the compound identified by the information into the machine learning model (2024) and executes a first step (S701, S702) of acquiring activity information of the compound against a specific receptor as an inference result, and a second step (S704) of screening a compound based on the activity information acquired in the first step (S701, S702). (Appendix 13) A system (1) for screening compounds based on activity information of compounds against receptors, the system (1) comprising: a memory (25, 26) in which a machine learning model (2024) is stored, the machine learning model having information for identifying a compound and structural information of a compound identified by the information as explanatory variables and activity information of a compound against a specific receptor as a target variable; a means for inputting the information for identifying a compound to be screened and the structural information of a compound identified by the information into the machine learning model (2024) and acquiring activity information of the compound against a specific receptor as an inference result; and a means for screening compounds based on the activity information acquired by the means for acquiring activity information. [Explanation of symbols]

[0155] 1... system 10... terminal device 20... server 25... memory 26... storage 29... processor 40... external server 202... storage unit 203... control unit 2021... application program 2022... experiment library DB 2023... teacher data 2024... machine learning model 2031... reception control module 2032... transmission control module 2033... data acquisition module 2034... model generation module 2035... selection module 2036... screening module 2037... presentation control module 2038... presentation control module

Claims

1. A program for operating a computer having a processor and a memory, the program screening a compound based on activity information of the compound against a receptor, the memory stores a machine learning model in which information for identifying the compound and structural information of the compound identified by the information are used as explanatory variables, and activity information of the compound against a specific receptor is used as a target variable; The program causes the processor to: a first step of inputting the information for identifying the compound to be screened and structural information of the compound identified by the information into the machine learning model, and obtaining activity information of the compound against a specific receptor as an inference result; a second step of screening the compound based on the activity information obtained in the first step; A program to execute.

2. The program according to claim 1 , wherein the activity information of the compound is an activity concentration of the compound having a specific value.

3. The method of claim 1 , wherein the receptor is a capsaicin receptor.

4. The program according to claim 1 , wherein the machine learning model is created for each of the receptors.

5. In the first step, a plurality of pieces of information for identifying each of the plurality of compounds and structural information of the plurality of compounds identified by each of the pieces of information are input to the machine learning model, and the activity concentrations of each of the plurality of compounds are obtained as an inference result; In the second step, screening of the plurality of compounds is performed based on the activity concentrations of the plurality of compounds obtained in the first step. The program according to claim 2.

6. The program according to claim 5 , wherein in the second step, at least one compound is selected from the plurality of compounds based on the range of values ​​of the activity concentrations of the plurality of compounds obtained in the first step.

7. The program according to claim 5 , wherein in the second step, at least one compound is selected from the plurality of compounds based on the magnitude of the activity concentration values ​​of the plurality of compounds obtained in the first step.

8. The program according to claim 1 , wherein the information for identifying the compound is information for identifying the structural information of the compound.

9. The program further causes the processor to a third step of performing machine learning on each of a plurality of candidate machine learning models using structural information of the compound and activity information of the compound against a specific receptor, each of which has a known value, as training data; a fourth step of inputting the structural information of the compound into each of the plurality of candidate machine learning models that have been subjected to machine learning in the third step, acquiring the activity information of the compound as an inference result, and calculating the prediction accuracy of the candidate machine learning model using the acquired activity information of the compound and the activity information of the compound whose value corresponding to the input structural information of the compound is known; A fifth step of determining the machine learning model to be stored in the memory based on the prediction accuracy calculated in the fourth step; The program according to claim 1 .

10. In the fourth step, a correlation coefficient is calculated between the acquired activity information of the compound and the activity information of the compound having a known value corresponding to the input structural information of the compound, and the correlation coefficient is set as the prediction accuracy; In the fifth step, the candidate machine learning model having the maximum correlation coefficient is determined as the machine learning model to be stored in the memory. The program according to claim 9.

11. An information processing device including a processor and a memory, the information processing device screening the compound based on activity information of the compound against a receptor, the memory stores a machine learning model in which information for identifying the compound and structural information of the compound identified by the information are used as explanatory variables, and activity information of the compound against a specific receptor is used as a target variable; The processor, a first step of inputting the information for identifying the compound to be screened and structural information of the compound identified by the information into the machine learning model, and obtaining activity information of the compound against a specific receptor as an inference result; a second step of screening the compound based on the activity information obtained in the first step; An information processing device that executes the above.

12. A method implemented by a computer having a processor and a memory, the method comprising screening a compound based on activity information of the compound against a receptor, the memory stores a machine learning model in which information for identifying the compound and structural information of the compound identified by the information are used as explanatory variables, and activity information of the compound against a specific receptor is used as a target variable; The processor, a first step of inputting the information for identifying the compound to be screened and structural information of the compound identified by the information into the machine learning model, and obtaining activity information of the compound against a specific receptor as an inference result; a second step of screening the compound based on the activity information obtained in the first step; A method for performing.

13. A system for screening a compound based on activity information of the compound against a receptor, comprising: a memory storing a machine learning model in which information for identifying the compound and structural information of the compound identified by the information are used as explanatory variables, and the activity information of the compound against a specific receptor is used as a response variable; a means for inputting the information for identifying the compound to be screened and structural information of the compound identified by the information into the machine learning model, and obtaining the activity information of the compound against a specific receptor as an inference result; a means for screening the compound based on the activity information acquired by the means for acquiring activity information; The system has:

Citation Information

Patent Citations

  • Connectivity prediction method, apparatus, program, recording medium, and production method of machine learning algorithm

    JP2019028879A

  • Method for determining at least one system state using a Kalman filter - Patents.com

    JP2024512265A

  • Method for predicting presence or absence of aroma properties or olfactory receptor activation properties in substance

    WO2021200780A1