Compound library generation method, compound library generation system, computer program, and learning model generation method

JPWO2025028367A5Pending Publication Date: 2026-04-27
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2026-01-23
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

In drug discovery, searching for compounds capable of binding to a target within a large compound library is time-consuming and costly, as existing methods do not efficiently identify compounds with binding ability to targets.

Method used

A compound library generation method using a learning model trained with compound information and binding data to classify compounds into groups that can bind to a target, generating a focused library with a higher proportion of active compounds.

Benefits of technology

This approach significantly improves the efficiency of compound screening by increasing the proportion of compounds with binding ability to the target, reducing the time and cost associated with traditional methods.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Provided are a compound library generation method and the like that make it possible to construct a compound library that increases the proportion of compounds that have ability to bind to a target. A compound library generation method according to the present invention involves a computer executing processing that acquires information about a plurality of compounds stored at a first compound library, uses a learning model that has been trained to output information that indicates ability to bind to a target substance when compound information about a compound has been inputted to sort the plurality of compounds stored at the first compound library into a group of compounds that have ability to bind to a target substance and a group of compounds that do not have ability to bind to the target substance, and generates a second compound library that includes the compounds sorted into the group of compounds that have ability to bind to the target substance.
Need to check novelty before this filing date? Find Prior Art

Description

Compound library generation method, compound library generation system, computer program, and learning model generation method

[0001] The present invention relates to a compound library generation method, a compound library generation system, a computer program, and a learning model generation method.

[0002] In the field of drug discovery, it is common to search for compounds with desired properties using compound libraries containing a large number of compounds. Techniques for constructing compound libraries useful for such compound searches have been proposed (for example, Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2006-223198

[0004] A compound library contains a large number of compounds. Searching through the large number of compounds contained in a compound library for compounds that have the ability to bind to a target requires time and cost. To efficiently find useful compounds, a compound library with a high proportion of compounds that have the ability to bind to a target is desired.

[0005] A primary object of the present disclosure is to provide a compound library generation method and the like that can realize the construction of a compound library that increases the proportion of compounds that have the ability to bind to a target.

[0006] A compound library generation method according to one aspect of the present disclosure includes a computer executing a process of acquiring information on a plurality of compounds stored in a first compound library, classifying the plurality of compounds stored in the first compound library into a group of compounds that have the ability to bind to a target substance and a group of compounds that do not have the ability to bind to a target substance, using a learning model that has been trained to output information indicating the binding ability of the compounds to a target substance when the compound information of the compounds is input, and generating a second compound library including the compounds classified into the group of compounds that have the ability to bind to the target substance.

[0007] A compound library generation system according to one embodiment of the present disclosure includes a control unit that acquires information about a plurality of compounds stored in a first compound library, and classifies the plurality of compounds stored in the first compound library into a group of compounds that have the ability to bind to a target substance and a group of compounds that do not have the ability to bind to a target substance using a learning model that has been trained to output information indicating the binding ability of the compounds to a target substance when the compound information about the compounds is input, and executes a process of generating a second compound library including the compounds classified into the group of compounds that have the ability to bind to the target substance.

[0008] A computer program according to one aspect of the present disclosure causes a computer to execute a process of acquiring information on a plurality of compounds stored in a first compound library, and classifying the plurality of compounds stored in the first compound library into a group of compounds that have binding ability to a target substance and a group of compounds that do not have binding ability to a target substance, using a learning model that has been trained to output information indicating the binding ability of the compounds to a target substance when compound information on the compounds is input, and generating a second compound library including the compounds classified into the group of compounds that have binding ability to the target substance.

[0009] A method for generating a learning model according to one aspect of the present disclosure acquires training data for a plurality of compounds stored in a compound library, the training data including compound information indicating the structure or properties of the compound and information indicating the binding ability to a target substance, and generates a learning model trained to output information indicating the binding ability to a target substance when compound information is input based on the acquired training data.

[0010] According to the present disclosure, it is possible to construct a compound library that increases the proportion of compounds that have the ability to bind to a target.

[0011] FIG. 1 is a diagram showing an overview of a compound library generation system of this embodiment. FIG. 2 is a block diagram showing an example of the configuration of an information processing device and a terminal device. FIG. 3 is an explanatory diagram showing an overview of a learning model and an example of the content of information stored in a training DB. FIG. 4 is a flowchart showing an example of a processing procedure for generating a learning model. FIG. 5 is a flowchart showing an example of a processing procedure for generating a focused library. FIG. 6 is a flowchart showing an example of a processing procedure for generating a learning model executed by an information processing device of a second embodiment. FIG. 7 is a flowchart showing an example of a processing procedure for generating predicted values ​​of binding information in a third embodiment. FIG. 8 is an explanatory diagram showing an overview of a learning model of a fourth embodiment and an example of the content of information stored in a training DB.

[0012] The present disclosure will be specifically described with reference to the drawings showing embodiments thereof.

[0013] (First embodiment) Fig. 1 is a diagram showing an overview of a compound library generation system 100 of this embodiment. The compound library generation system 100 includes an information processing device 1 as a main device. The information processing device 1 is communicably connected to a terminal device 2 via a network N such as the Internet. The number of terminal devices 2 may be one or three or more.

[0014] The information processing device 1 is a device capable of various information processing and information transmission and reception, such as a server computer, a personal computer, a quantum computer, etc. The terminal device 2 is an information processing terminal used by a person in charge of a drug discovery company, which is an example of a user. The terminal device 2 is, for example, a personal computer, a smartphone, a tablet terminal, etc.

[0015] The information processing device 1 receives a basic library 31, which is a compound library held by a drug discovery company, through the terminal device 2. The information processing device 1 generates a focused library 32 according to the received basic library 31, and provides the generated focused library 32 to the drug discovery company through the terminal device 2. The basic library 31 corresponds to the first compound library, and the focused library 32 corresponds to the second compound library.

[0016] The basic library 31 is a compound library that manages compounds in a database. The basic library 31 stores a large number of compounds, including compounds obtained in past drug discovery research at a drug discovery company and compounds obtained from outside. The basic library 31 stores information about each compound, such as the compound name, structural formula, physical properties, and physiochemical characteristics. The basic library 31 may be a library independently maintained by a company, and stores information about a large number of compounds independently owned by the drug discovery company.

[0017] The focused library 32 is a compound library generated based on the basic library 31, and is a library that selectively stores compounds having desired properties from among the compounds stored in the basic library 31. More specifically, the focused library 32 selectively stores compounds that have activity against a target substance. The target substance may be any of DNA, RNA, protein, etc.

[0018] In the pharmaceutical field, compound discovery involves screening compounds in the basic library 31 to find effective candidate compounds. It takes a great deal of time and money to investigate the activity of all compounds in the basic library 31 against a target substance. This system improves the efficiency of screening work by providing a focused library 32 that increases the proportion of compounds that are active against the target substance.

[0019] In the following, as an example, a case will be described in which the target substance is RNA and a focused library 32 is generated in which the proportion of small molecule compounds that bind to the target RNA is increased. The target RNA may be untranslated RNA.

[0020] 2 is a block diagram showing an example of the configuration of the information processing device 1 and the terminal device 2. The information processing device 1 includes a control unit 11, a storage unit 12, a communication unit 13, a display unit 14, an operation unit 15, and an input / output unit 16. The information processing device 1 may be a multi-computer consisting of multiple computers, or may be a virtual machine virtually constructed by software.

[0021] The control unit 11 includes one or more arithmetic processing devices such as a central processing unit (CPU), a micro-processing unit (MPU), a graphics processing unit (GPU), etc. The control unit 11 controls each component and executes processing using built-in memories such as a read-only memory (ROM) or a random access memory (RAM), a clock, a counter, etc. Each functional unit of the information processing device 1 may be realized by software, by hardware (for example, an FPGA or an ASIC), or by a combination thereof.

[0022] The storage unit 12 includes a non-volatile memory such as a hard disk, a flash memory, or an SSD (Solid State Drive). The storage unit 12 may be an external storage device connected to the information processing device 1. The storage unit 12 stores various computer programs and data referenced by the control unit 11. The storage unit 12 of this embodiment stores a program 1P for causing a computer to execute processing related to the generation of the focused library 32, a learning model 121, and a training DB (Data Base) 122. The learning model 121 is a model generated by machine learning. The learning model 121 is expected to be used as a program module that constitutes part of artificial intelligence software.

[0023] A computer program (computer program product) including program 1P may be provided by a non-transitory recording medium 1A on which the computer program is readably recorded. The storage unit 12 stores the computer program read from the recording medium 1A by a reading device (not shown). The recording medium 1A may be, for example, a magnetic disk, an optical disk, or a semiconductor memory. Alternatively, the computer program may be downloaded from an external server connected to a communications network and stored in the storage unit 12. Program 1P may be a single computer program or may be composed of multiple computer programs, and may be executed on a single computer or on multiple computers interconnected by a communications network.

[0024] The communication unit 13 includes a communication module that communicates with external devices via the network N. The control unit 11 transmits and receives various types of information to and from the terminal device 2 via the communication unit 13.

[0025] The display unit 14 includes a display device such as a liquid crystal display, an organic EL (Electro Luminescence) display, etc. The display unit 14 displays information to be notified to the user in accordance with instructions from the control unit 11.

[0026] The operation unit 15 is an interface that accepts user operations. The operation unit 15 includes, for example, a keyboard, a mouse, a touch panel device with a built-in display, a speaker, a microphone, etc. The operation unit 15 accepts operation input from the user and sends a control signal corresponding to the operation content to the control unit 11.

[0027] The input / output unit 16 includes an input / output interface for connecting an external device by wire or wirelessly. The input / output unit 16 is connected to the detection device 4.

[0028] The detection device 4 is a device that detects the interaction between the target RNA and a compound. The detection device 4 measures the interaction between the target RNA and a compound, for example, using surface plasmon resonance (SPR). The interaction analysis method is not limited to SPR, and may be, for example, isothermal titration calorimetry, mass spectrometry, nuclear magnetic resonance spectroscopy, melting temperature measurement, absorption spectroscopy, fluorescence spectroscopy, circular dichroism spectroscopy, etc. By analyzing the interaction between the target RNA and a compound, it is possible to determine the presence or absence of binding ability to the target RNA, the degree of binding strength, the binding ratio, the presence or absence of structural changes upon binding, etc. The detection device 4 outputs detection data obtained by detection to the information processing device 1. The information processing device 1 may acquire detection data stored in a predetermined memory area via the detection device 4.

[0029] The terminal device 2 includes a control unit 21, a storage unit 22, a communication unit 23, a display unit 24, and an operation unit 25. The control unit 21 includes one or more arithmetic processing units such as a CPU, an MPU, a GPU, etc. The control unit 21 controls each component and executes processing using built-in memory such as a ROM or RAM, a clock, a counter, etc.

[0030] The storage unit 22 includes a non-volatile memory such as a hard disk, flash memory, or SSD. The storage unit 22 stores various computer programs and data referenced by the control unit 21. The storage unit 22 stores a program 2P for causing a computer to execute processing related to acquisition of a focused library 32, and a basic library 31. The storage unit 22 may also store a focused library 32 received from the information processing device 1.

[0031] The communication unit 23 includes a communication module that realizes communication via the network N. The display unit 24 includes a display device such as a liquid crystal display or an organic EL display. The operation unit 25 is an interface that accepts user operations. The operation unit 25 includes, for example, a keyboard, a mouse, a touch panel device with a built-in display, a speaker, a microphone, etc.

[0032] FIG. 3 is an explanatory diagram showing an overview of the learning model 121 and an example of the contents of information stored in the training DB 122.

[0033] The training DB 122 is a database that stores training data used for training the learning model 121. As shown in Fig. 3, the training DB 122 stores records in which compound information, binding information, etc. are linked using, for example, a compound ID that identifies a compound as a key. The group of compounds that make up the training data, i.e., the multiple compounds included in the training DB 122, are a portion of compounds extracted from the compounds stored in the basic library 31.

[0034] The compound information includes information that represents the structure or physical properties of a compound. The compound information includes, for example, molecular descriptors of the compound. Molecular descriptors are quantified to make the structural features and physicochemical properties of a substance easier to handle on a computer. Molecular descriptors can be calculated from the structural formula of a substance and can be obtained using known software such as alvaDesc, Dragon, Codessa, RDKit, and Mordred. The compound information may include values ​​of multiple types of molecular descriptors.

[0035] The compound information is not limited to molecular descriptors, but may also include, for example, structural formulas expressed as character strings according to the SMILES (Simplified Molecular Input Line Entry System) notation, molecular graphs obtained by converting chemical structural formulas into graph information, and the like.

[0036] The binding information is information indicating the binding ability to the target RNA, and includes whether or not the compound has the binding ability to the target RNA. The binding information may further include the degree of binding strength to the target RNA, the binding ratio, whether or not there is a structural change upon binding, etc. The binding information is obtained by analyzing detection data indicating the interaction between the compound and the target RNA obtained by an experiment using the detection device 4. For example, as an indicator for determining whether or not there is binding ability, it can be determined that there is binding ability when the resonance unit (RU value) of the SPR signal obtained from the compound is equal to or greater than a predetermined threshold.

[0037] The bonding information used in the training data may be a simulation value calculated by a predetermined algorithm. Examples of calculation methods for the simulation value include quantum chemical calculation and molecular dynamics calculation. The simulation value obtained by each calculation method can be calculated using known theoretical calculation software.

[0038] As shown in FIG. 3 , the learning model 121 receives compound information about a compound as input and outputs binding information indicating the binding ability of the compound to a specific target RNA. Hereinafter, the specific target RNA recognized by the learning model 121 will also be referred to as the first target RNA. The learning model 121 of this embodiment outputs a classification result indicating whether or not a compound has binding ability to the first target RNA, i.e., whether or not the compound has binding ability. The learning model 121 is used in the compound selection process when generating the focused library 32.

[0039] The learning model 121 is constructed using, for example, a random forest. The information processing device 1 creates multiple decision trees with low correlation by randomly selecting features to be used for model construction based on data sampled from training data, and generates the learning model 121 using the multiple decision trees. The learning model 121 outputs a final value that is a majority vote or an average of the estimated values ​​from each decision tree.

[0040] The compound information to be input to the learning model 121 includes a plurality of molecular descriptors. The compound information may include a SMILES string, a molecular graph, and the like.

[0041] One or more types of molecular descriptors selected from a large number of types may be input to the learning model 121. The type of molecular descriptor to be input to the learning model 121 can be set appropriately depending on the type of the first target RNA. As an example, the molecular descriptor to be input to the learning model 121 is automatically determined by the information processing device 1 based on the contribution of the input information obtained when the learning model 121 is trained.

[0042] The information processing device 1 executes an estimation process using the learning model 121 using a predetermined number of preselected molecular descriptors, thereby calculating the contribution (variable importance) of each molecular descriptor in the estimation of binding information in the learning model 121. The contribution can be calculated based on, for example, the SHAP (Shapley Additive exPlanation) value, the Gini coefficient, out-of-bag (oob) data, or Local Interpretable Model-Agnostic Explanations (LIME). Based on the calculated contribution of each molecular descriptor, the information processing device 1 selects a predetermined number of molecular descriptors in descending order of contribution, thereby determining the molecular descriptors to be used as input information to the learning model 121. The above-described process makes it possible to identify molecular descriptors suitable for input to the learning model 121 from among the many defined molecular descriptors.

[0043] The binding information that is output from the learning model 121 includes whether or not the compound has the ability to bind to the first target RNA.

[0044] The learning model 121 can be generated by preparing training data in which compound information is labeled with the presence or absence of binding ability to the first target RNA, which is the correct value, and using the training data to train an unlearned model.

[0045] The information processing device 1 inputs compound information in the training data into the learning model 121 and trains the learning model 121 so that the output from the learning model 121 approximates a correct value. The information processing device 1 generates the learning model 121 by adjusting the parameters in the learning model 121 using the presence or absence of binding ability associated with the input compound information as a correct value. Through the above-described processing, it is possible to construct a learning model 121 that has been trained to be able to appropriately output the presence or absence of binding ability to the first target RNA for the compound information.

[0046] The output information from the learning model 121 is not limited to the presence or absence of binding ability to the first target RNA. The learning model 121 may be configured to further output, for example, the degree of binding strength to the first target RNA, the binding ratio, the presence or absence of a structural change upon binding, etc.

[0047] The configuration of the learning model 121 is not limited to the example shown in Fig. 3, and may be any model that can distinguish information indicating the binding ability to the target RNA from compound information. The learning model 121 may be a model based on other learning algorithms, such as a convolution neural network (CNN), a recurrent neural network (RNN), a graph neural network (GNN), a transformer, a support vector machine, a logistic regression, or eXtreme Gradient Boosting (XGBoost).

[0048] The learning model 121 may be composed of a plurality of individual learning models constructed by different learning algorithms, for example, the learning model 121 may include a first individual learning model constructed by Random Forest and a second individual learning model constructed by XGBoost.

[0049] In the compound library generation system 100, a learning model 121 is constructed using training data generated based on compound information on a portion of the compounds stored in the basic library 31. The obtained learning model 121 is used to estimate the presence or absence of binding ability among the large number of compounds stored in the basic library 31, and compounds estimated to have binding ability are extracted to generate a new focused library 32. A processing method for generating the focused library 32 will be described in detail below.

[0050] 4 is a flowchart showing an example of a processing procedure for generating the learning model 121. The processing in each of the following flowcharts is executed by the control unit 11 in accordance with a program 1P stored in the storage unit 12 of the information processing device 1.

[0051] The control unit 11 of the information processing device 1 receives, from the user, via the terminal device 2, a selection of a target RNA to be used to generate the focused library 32 and user identification information to identify the user (step S11). The control unit 11 receives information indicating the name and structure of the target RNA from the terminal device 2, for example, based on the user's operation of the operation unit 25. The user can select any target RNA (e.g., a first target RNA) depending on the purpose of the drug discovery research.

[0052] The control unit 11 acquires a compound group containing multiple compounds to be used in the training data through the terminal device 2 (step S12). The compound group is extracted from the compounds stored in the basic library 31. Based on the user's operation of the operation unit 25, the control unit 11 receives from the terminal device 2 information on each compound included in the compound group, including information on the compound stored in the basic library 31, necessary for the processing described below (e.g., compound name, molecular structure, required physical properties, etc.). The user can arbitrarily select compounds to be used in the training data depending on the purpose of the drug discovery research. The user determines the number of compounds to be used in the training data based on the time and cost allowed for generating the training data, and extracts compounds from the basic library 31 so that the number reaches the determined number. Note that the compound group to be used in the training data may be determined by the information processing device by randomly extracting a predetermined number from the basic library 31.

[0053] The control unit 11 acquires binding information for each compound included in the acquired compound group, including the presence or absence of binding ability to the first target RNA, the degree of binding strength, the binding ratio, and the presence or absence of structural changes upon binding (step S13). The binding information for the first target RNA is obtained by analyzing the interaction between each compound and the target RNA detected using the detection device 4. The control unit 11 may acquire the binding information by, for example, receiving a manually determined result of binding ability via the operation unit 15 or the communication unit 13, or may automatically derive the binding information based on detection data received from the detection device 4 via the input / output unit 16.

[0054] The control unit 11 acquires compound information for each compound included in the compound group (step S14). For example, the control unit 11 derives a structural formula in SMILES notation from the molecular structure of the compound acquired in step S12, and calculates multiple types of molecular descriptors that are set in advance based on the derived structural formula.

[0055] The control unit 11 associates the obtained compound information with the information indicating the binding ability and stores them in the training DB 122 (step S15). By the above-described process, training data is generated.

[0056] The control unit 11 acquires training data in which the presence or absence of binding ability to the first target RNA is assigned to the compound information of the training compound based on the information stored in the training DB 122 (step S16).

[0057] The control unit 11 uses the acquired training data to generate a learning model 121 that, when compound information about a compound is input, outputs whether or not the compound has the binding ability to the first target RNA (step S17). Specifically, the control unit 11 inputs the compound information included in the training data into the learning model 121 and optimizes various parameters so that the output from the learning model 121 approximates the correct value. For example, when learning is completed because the number of learning times meets a predetermined standard, the control unit 11 stores definition information about the learned learning model 121 in the storage unit 12 as the learned learning model 121. The learning model 121 is constructed by the processes from step S16 to step S17.

[0058] The group of compounds used as training data for generating the learning model 121 is not limited to being extracted from compounds stored in the basic library 31, i.e., the library that is the target for generating the focused library 32. The learning model 121 may be constructed using information from a compound library that is different from the target for generating the focused library 32 as training data.

[0059] 5 is a flowchart showing an example of a procedure for generating the focused library 32. The control unit 11 of the information processing device 1 executes the following process after, for example, step S17 in the flowchart of FIG.

[0060] The control unit 11 of the information processing device 1 receives the basic library 31 held by the user from the terminal device 2 (step S21). In this case, the control unit 11 may receive only information necessary for the processing described below (e.g., compound name, molecular structure, etc.) from the information related to each compound included in the basic library 31. In step S11, the control unit 11 acquires information on compounds to be evaluated among the compounds included in the basic library 31. The compounds to be evaluated may be all compounds stored in the basic library 31, or a portion selected from all the compounds.

[0061] The control unit 11 acquires compound information including molecular descriptors, structural formulas, molecular graphs, etc. for each compound in the acquired basic library 31 (step S22). The control unit 11 inputs the acquired compound information for each compound into the learning model 121 (step S23), and acquires the presence or absence of binding ability output from the learning model 121 (step S24).

[0062] Based on the obtained estimation result of the presence or absence of binding ability, the control unit 11 classifies each compound in the basic library 31 into either a hit group or a non-hit group (step S25). The hit group is a group to which compounds that have the ability to bind to the first target RNA belong, and the non-hit group is a group to which compounds that do not have the ability to bind to the first target RNA belong. The processing of step S25 corresponds to a selection process for selecting compounds that have the ability to bind to the first target RNA from the basic library 31.

[0063] The control unit 11 obtains a score that quantifies the priority of each compound classified into the hit group (step S26). The priority is quantified so that the higher the probability that the compound has the ability to bind to the first target RNA, the higher the score. The control unit 11 obtains a classification accuracy value in the learning model 121 and can use the obtained classification accuracy as the score. The classification accuracy can be obtained, for example, from the proportion of decision trees that match the estimation, the confidence level for the classification class at the output node, etc.

[0064] Based on the classification results, the control unit 11 generates a focused library 32 containing the compounds classified into the hit group (step S27). In this case, the control unit 11 ranks the compounds classified into the hit group in descending order of score based on the priority score, and stores the compounds in the focused library 32 in descending order. The generated focused library 32 only needs to include information that allows at least the compounds included in the focused library 32 to be recognized, and may be in the form of a compound list that displays, for example, the names of multiple compounds in order of score. The focused library 32 may be associated with a score for each compound.

[0065] The control unit 11 transmits the generated focused library 32 to the user identified by the user identification information, that is, the terminal device 2 corresponding to the user of the basic library 31 (step S28), and ends the series of processes.

[0066] The user searches for drug discovery targets using the focused library 32 provided by the information processing device 1. The focused library 32 is composed of compounds that are estimated by the learning model 121 to have the ability to bind to the first target RNA, and therefore compounds that have activity on the first target RNA can be obtained with a high hit rate.

[0067] Although the information processing device 1 executes a series of processes in each flowchart in the above example, the information processing device 1 is not limited to a specific processor. Part of the processes in each flowchart may be executed by the terminal device 2.

[0068] Each time the information processing device 1 receives from the terminal device 2 a request to generate a focused library 32 targeting a new target RNA, it repeatedly executes a series of processes including the generation of training data, the generation of a learning model 121, and the generation of the focused library 32. That is, different learning models 121 are prepared depending on the type of target RNA, and compounds are selected using the learning models 121 corresponding to each prepared target RNA. When the information processing device 1 receives from the user a selection of a second target RNA as a target substance, it generates a learning model 121 that estimates binding information regarding the second target RNA, and generates a new focused library 32 by evaluating the binding ability of each compound to the second target RNA using the learning model 121.

[0069] In the above process, the control unit 11 may omit the score acquisition process in step S26. Alternatively, the control unit 11 may further select compounds to be stored in the focused library 32 based on the scores acquired in step S26. For example, the control unit 11 may generate the focused library 32 by extracting only compounds whose scores are equal to or greater than a threshold value.

[0070] When the learning model 121 further estimates the degree of binding strength, the binding ratio, the presence or absence of structural changes upon binding, etc., the control unit 11 may classify the compounds in the basic library 31 taking these estimation results into account. For example, the control unit 11 classifies the compounds into three or more groups based on the presence or absence of binding ability to the first target RNA and the binding strength. The control unit 11 stores compounds belonging to groups based on preset selection criteria (e.g., compounds that have binding ability and strong binding strength) in the focused library 32.

[0071] The control unit 11 may repeatedly perform classification using the learning model 121 multiple times, calculate a total score that is the sum of the priorities calculated each time, and perform ranking based on the calculated total score. The control unit 11 may finally extract a predetermined number of compounds in descending order of total score, for example, compounds whose total score is equal to or greater than a preset threshold, and store the extracted compounds in the focused library 32.

[0072] When the learning model 121 includes multiple individual learning models, the control unit 11 may classify compounds into a hit group and a non-hit group based on the output results of each individual learning model. The control unit 11, for example, acquires the presence or absence of binding ability output from each individual learning model for a certain compound. When classification results indicating that a compound has binding ability are obtained from all individual learning models or individual learning models equal to or greater than a preset threshold, the control unit 11 classifies the compound into the hit group.

[0073] The compounds to be evaluated may include compounds included in a compound group. In this case, the compounds included in the compound group may be classified into a hit group or a non-hit group based on binding information obtained by actual measurement, instead of estimation by the learning model 121.

[0074] The focused library 32 is not limited to one containing only compounds that are predicted to have the ability to bind to the first target RNA, but may also contain some compounds that are predicted not to have the ability to bind to the first target RNA.

[0075] According to this embodiment, a high-quality focused library 32 with an increased proportion of compounds that have the ability to bind to the target RNA can be generated based on the user's basic library 31. By generating the focused library 32 by extracting compounds from the basic library 31 owned by a drug discovery company, it is possible to provide a focused library 32 that can improve search efficiency while ensuring originality.

[0076] By using the learning model 121, it is possible to efficiently and accurately select compounds from the basic library 31. By generating training data by evaluating the interaction between the compound and the target substance through actual measurements, the estimation accuracy of the learning model 121 is improved. By generating the learning model 121 according to the first target RNA, it is possible to accurately estimate the binding ability to the first target RNA using the learning model 121.

[0077] By quantifying the priority of each compound in the focused library 32, it is possible to further improve the efficiency of searching using the focused library 32. By accepting the selection of a target RNA from the user, it is possible to obtain a focused library 32 suitable for any target RNA, thereby improving the convenience of the present system.

[0078] Compared to small molecular weight compounds that target proteins, research into small molecular weight compounds that target non-translated RNA has not progressed sufficiently, and in many cases, a sufficient number of compounds are not stored in the basic library 31. Therefore, the effects of the present disclosure are more pronounced when the present system and the like are applied when the target substance is target RNA.

[0079] 5 may be executed by the terminal device 2. In this case, the processes of steps S21 and S28 may be omitted.

[0080] 4 and stores the generated learning model 121 in an area accessible by the terminal device 2. The information processing device 1 may deploy the generated learning model 121 to the terminal device 2.

[0081] The terminal device 2 accesses the learning model 121 stored in an external device or refers to the memory unit 22 of its own device to read the learning model 121, and performs the processing of steps S22 to S27 using the read learning model 121, thereby performing compound selection processing, etc.

[0082] According to the above configuration, the focused library 32 can be generated on the terminal device 2 side. This eliminates the need to provide information on compounds to be evaluated from the terminal device 2 to the information processing device 1, thereby improving the confidentiality of the basic library 31.

[0083] Second Embodiment In the second embodiment, a portion of the compounds that do not have the ability to bind to the first target RNA are removed from the compounds included in the compound group. In the following embodiments, differences from the first embodiment will be mainly described, and components common to the first embodiment will be assigned the same reference numerals and detailed description thereof will be omitted.

[0084] The multiple compounds included in the training DB 122, i.e., the multiple compounds included in the compound group selected from the basic library 31, include compounds with binding ability and compounds without binding ability as binding information obtained by interaction analysis. It is assumed that the number of compounds without binding ability to the first target RNA included in the compound group is greater than the number of compounds with binding ability to the first target RNA. In other words, it is assumed that the data included in the training DB 122 is imbalanced data in which there is a bias in the ratio of compounds with binding ability to compounds without binding ability.

[0085] If the learning model 121 is generated using such imbalanced data, there is a risk that the sensitivity to compounds with binding ability will be reduced. The information processing device 1 of the second embodiment aims to improve the accuracy of the learning model 121 by eliminating some of the compounds without binding ability included in the compound group.

[0086] FIG. 6 is a flowchart showing an example of a procedure for generating the learning model 121 executed by the information processing device 1 according to the second embodiment.

[0087] The control unit 11 of the information processing device 1 executes the processes from step S11 to step S15 of the first embodiment, thereby storing the compound information and binding information of all compounds in the compound group acquired in step S15 in the training DB 122. Based on the compound information of each compound stored in the training DB 122, the control unit 11 classifies each compound included in the compound group into compounds that have binding ability to the first target RNA and compounds that do not have binding ability to the first target RNA (step S31).

[0088] Based on the classification results, the control unit 11 removes a portion of the compounds that do not have the ability to bind to the first target RNA from the compounds included in the compound group (step S32). The control unit 11 thins out the compounds that do not have the ability to bind to the first target RNA so that the ratio of compounds that have the ability to bind to the first target RNA to compounds that do not have the ability to bind to the first target RNA in the compound group (number of compounds with binding ability: number of compounds that do not have binding ability) is, for example, 1:1, 1:4, 1:8, 1:16, etc. The above ratio is an example, and the ratio in the removal process is not necessarily limited to this numerical value.

[0089] After processing step S32, the control unit 11 executes processing similar to steps S16 to S17, thereby using the data of the compound group after removing some compounds as training data and generating a learning model 121.

[0090] According to this embodiment, it is possible to suppress deterioration of the estimation accuracy of the learning model 121 due to imbalance in training data.

[0091] (Third embodiment) In the third embodiment, a learning model 121 is generated that estimates binding information for other target substances using predicted values ​​of binding information for other target substances generated based on binding information for a specific target substance and correlations between the target substances.

[0092] As described above, the learning model 121 is trained using training data in which compound information and binding information are associated with each other. In the third embodiment, predicted values ​​of binding information are used as training data instead of binding information as actual measurement data obtained using a predetermined mutual analysis method. The predicted values ​​of binding information can be calculated using binding information already obtained by actual measurement, taking into account the correlation between multiple target RNAs.

[0093] 7 is a flowchart showing an example of a processing procedure for generating predicted values ​​of binding information in the third embodiment. Hereinafter, as an example, it is assumed that the new target substance is a second target RNA in which adenine at a specific position in the RNA structure of the first target RNA is substituted with uracil. It is also assumed that the presence or absence of binding ability for the first target RNA according to actual measurement data has already been stored in the training DB 122.

[0094] The control unit 11 of the information processing device 1 calculates a correlation coefficient representing the strength of the correlation function between the first target RNA and the second target RNA (step S41). The correlation coefficient between the first target RNA and the second target RNA is calculated by, for example, comprehensively evaluating the difference in the material structure between the first target RNA and the second target RNA, specifically the difference between adenine and uracil, using various material-related indices (e.g., the number of hydrogen bond donors, the number of hydrogen bond acceptors, surface area, volume, etc.).

[0095] The control unit 11 derives a predicted value of binding information for the second target RNA based on the calculated correlation degree and the binding information for the first target RNA stored in the training DB 122 (step S42). The control unit 11 obtains a predicted value of the presence or absence of binding ability for the second target RNA, for example, by taking the calculated correlation degree into account in the presence or absence of binding ability for the first target RNA.

[0096] The control unit 11 stores the compound information and the predicted value of the binding information for the second target RNA in association with each other as training data in the training DB 122 (step S43).

[0097] Thereafter, the control unit 11 executes the same processes as steps S16 and S17 of the first embodiment to generate a learning model 121 that uses training data including predicted values ​​of binding information for the second target RNA to estimate the presence or absence of binding ability for the second target RNA for compound information. The control unit 11 uses the generated learning model 121 to select compounds that have binding ability for the second target RNA, thereby generating a focused library 32 for the second target RNA.

[0098] According to this embodiment, it is possible to reduce the time and cost required for experiments to generate training data, and to facilitate the generation of the focused library 32. It is no longer necessary to provide compound information for the generation of training data, and the burden on drug discovery companies is reduced. Even when targeting substances whose interactions are difficult to detect in practice, it is possible to generate the focused library 32, improving convenience.

[0099] Fourth Embodiment In a fourth embodiment, a learning model 121 capable of estimating binding information for various target RNAs is generated.

[0100] FIG. 8 is an explanatory diagram showing an overview of the learning model 121 according to the fourth embodiment and an example of the contents of information stored in the training DB 122.

[0101] 8 , the learning model 121 of the fourth embodiment receives compound information of a compound and target substance information of a target substance as input, and outputs binding information of the compound to the target substance. The target substance information includes information about the target substance, such as the substance name, primary structure (sequence), and secondary structure of the target substance.

[0102] The information processing device 1 accumulates, for example, binding information acquired in a generation process of a focused library 32 targeting multiple types of target substances in a training DB 122. Based on the accumulated information, the information processing device 1 generates training data in which binding information of the compound to the target substance is labeled as a correct value for compound information of the compound and target substance information of the target substance. The information processing device 1 trains a learning model 121 using the generated training data. The learning model 121 trained in this manner can estimate binding information for various target substances.

[0103] According to this embodiment, binding information for a variety of target substances can be estimated using a single learning model, eliminating the need to generate a learning model 121 for each type of target substance, making it easier to generate a focused library 32.

[0104] The embodiments disclosed herein are illustrative in all respects and should not be considered limiting. The technical features described in each embodiment can be combined with each other, and the scope of the present invention is intended to include all modifications within the scope of the claims and equivalents thereto. The sequences shown in each embodiment are not limited, and within the scope of no contradiction, each processing step may be executed in a different order, or multiple processes may be executed in parallel. The entity that performs each process is not limited, and within the scope of no contradiction, the process of each device may be executed by another device.

[0105] The matters described in each embodiment can be combined with each other. Furthermore, the independent claims and dependent claims described in the claims can be combined with each other in any and all combinations, regardless of the reference format. Furthermore, the claims use a format in which a claim references two or more other claims (multiple claim format), but this is not limited to this. A multiple claim (multi-multi claim) that references at least one other multiple claim may also be used.

[0106] 100 Compound library generation system 1 Information processing device 11 Control unit 12 Memory unit 13 Communication unit 14 Display unit 15 Operation unit 16 Input / output unit 121 Learning model 122 Training DB 1P Program 1A Recording medium 2 Terminal device 21 Control unit 22 Memory unit 23 Communication unit 24 Display unit 25 Operation unit 2P Program 2A Recording medium 31 Basic library 32 Focused library

Claims

1. Obtain information on multiple compounds stored in the first compound library, Using a learning model trained to output information indicating binding ability to a target substance when compound information of a compound is input, multiple compounds stored in the first compound library are classified into a group of compounds that have binding ability to a target substance and a group of compounds that do not have binding ability to a target substance. A second compound library is generated, containing compounds classified as having the ability to bind to the aforementioned target substance. A method for generating a compound library in which a computer performs the processing.

2. The learning model is trained based on training data that includes compound information indicating the structure or properties of the compound and information indicating the binding ability. The method for generating a compound library according to claim 1.

3. Information indicating the binding ability in the training data is obtained by analyzing the interaction between the compound and the target substance using surface plasmon resonance, isothermal titration calorimetry, mass spectrometry, nuclear magnetic resonance spectroscopy, melting temperature measurement, absorption spectroscopy, fluorescence spectroscopy, or circular dichroism spectroscopy. The method for generating a compound library according to claim 2.

4. The training data includes the compound information of the group of compounds extracted from the first compound library and the information indicating the binding ability. A method for generating a compound library according to claim 2 or claim 3.

5. Interaction analysis between the compound and the target substance revealed that some of the compounds lacked the ability to bind to the target substance. The learning model is trained based on the training data, which includes the compound information and the information indicating the binding ability of the compounds in the group of compounds after removal. The method for generating a compound library according to claim 4.

6. Based on the correlation between the first target substance and the second target substance, and information indicating the binding ability to the first target substance obtained using a predetermined interaction analysis method, information indicating the binding ability to the second target substance is derived. The learning model is trained based on training data that includes information indicating the binding ability to the derived second target substance. The method for generating a compound library according to claim 1.

7. The aforementioned compound information includes multiple molecular descriptors. The method for generating a compound library according to claim 1.

8. Based on the contribution of multiple molecular descriptors to the information indicating the binding ability, the molecular descriptors to be used as input to the learning model are determined from among the multiple molecular descriptors. The method for generating a compound library according to claim 7.

9. A learning model is prepared for each type of target substance. The method for generating a compound library according to claim 1.

10. We obtain numerical information that quantifies the priority of each compound included in the second compound library, The second compound library is generated according to the acquired numerical information. The method for generating a compound library according to claim 1.

11. The user of the first compound library selects a target substance. A second compound library is generated, containing compounds having the ability to bind to the selected target substance. The method for generating a compound library according to claim 1.

12. The target substance is the target RNA. The method for generating a compound library according to claim 1.

13. Obtain information on multiple compounds stored in the first compound library, Using a learning model trained to output information indicating binding ability to a target substance when compound information of a compound is input, multiple compounds stored in the first compound library are classified into a group of compounds that have binding ability to a target substance and a group of compounds that do not have binding ability to a target substance. A second compound library is generated, containing compounds classified as having the ability to bind to the aforementioned target substance. It includes a control unit that performs processing. Compound library generation system.

14. Obtain information on multiple compounds stored in the first compound library, Using a learning model trained to output information indicating binding ability to a target substance when compound information of a compound is input, multiple compounds stored in the first compound library are classified into a group of compounds that have binding ability to a target substance and a group of compounds that do not have binding ability to a target substance. A second compound library is generated, containing compounds classified as having the ability to bind to the aforementioned target substance. A computer program that instructs a computer to perform a process.

15. For multiple compounds stored in a compound library, training data is acquired that includes compound information indicating the structure or properties of the compound, and information indicating its binding ability to target substances. Based on the acquired training data, a learning model is generated that is trained to output information indicating the binding ability to a target substance when compound information is input. Methods for generating learning models.