Compound library generation method, compound library generation system, computer program, and learning model generation method
By using a learning model to classify and screen compounds in a compound library, the problem of low efficiency in constructing high-proportion target-binding compound libraries in existing technologies is solved, achieving efficient and accurate compound screening, which is particularly suitable for exploring compounds targeting non-coding RNA.
Patent Information
- Application Number
- CN202480049737.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-02
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-27
AI Technical Summary
In existing technologies, exploring compounds with target-binding capabilities from compound libraries requires time and costs, making it difficult to efficiently construct compound libraries with a high proportion of targets.
By acquiring compound information, a learning model is used to classify compounds into groups that have the ability to bind to the target substance and those that do not, generating a second compound library. Machine learning algorithms such as random forest and XGBoost are then used to build a learning model to screen out compounds with binding ability.
This increased the proportion of compounds with target-binding capabilities in the compound library, improving the efficiency and accuracy of compound exploration, especially for non-coding RNA targets, while reducing experimental time and costs.
Smart Images

Figure CN121586930A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for generating a compound library, a system for generating a compound library, a computer program, and a method for generating a learning model. Background Technology
[0002] In the pharmaceutical field, compound libraries containing a large number of compounds are widely used for the exploration of compounds with target properties. Techniques for constructing compound libraries useful for such compound exploration have been proposed (e.g., Patent Document 1).
[0003] Existing technical documents
[0004] Patent documents
[0005] Patent Document 1: Japanese Patent Application Publication No. 2006-223198 Summary of the Invention
[0006] The problem the invention aims to solve
[0007] Compound libraries contain a large number of compounds. Exploring compounds with target-binding capabilities from these vast libraries is time-consuming and costly. To efficiently find useful compounds, a compound library with a high proportion of target-binding compounds is desirable.
[0008] The main objective of this disclosure is to provide a method for generating a compound library, which enables the construction of a compound library that increases the proportion of compounds with target binding ability.
[0009] Solution for solving the problem
[0010] One aspect of the compound library generation method disclosed herein involves a computer performing the following processing: acquiring information on a plurality of compounds stored in a first compound library; using a learning model to classify the plurality of compounds stored in the first compound library into a group of compounds having the ability to bind to a target substance and a group of compounds not having the ability to bind to a target substance, wherein the learning model is learned to output information representing the binding ability to a target substance when compound information is input; and generating a second compound library including compounds classified into the group of compounds having the ability to bind to a target substance.
[0011] A compound library generation system according to one aspect of this disclosure includes a control unit that performs the following processing: acquiring information on a plurality of compounds stored in a first compound library; using a learning model, classifying the plurality of compounds stored in the first compound library into a group of compounds having the ability to bind to a target substance and a group of compounds not having the ability to bind to a target substance, wherein the learning model is learned to output information representing the binding ability to a target substance when compound information of the compounds is input; and generating a second compound library including compounds classified into the group of compounds having the ability to bind to a target substance.
[0012] A computer program of one aspect of this disclosure causes a computer to perform the following processing: acquiring information on a plurality of compounds stored in a first compound library; using a learning model, classifying the plurality of compounds stored in the first compound library into a group of compounds having the ability to bind to a target substance and a group of compounds not having the ability to bind to a target substance, wherein the learning model is learned to output information representing the binding ability to a target substance when compound information of the compounds is input; and generating a second compound library including compounds classified into the group of compounds having the ability to bind to a target substance.
[0013] The present disclosure discloses a method for generating a learning model, which involves acquiring training data for multiple compounds stored in a compound library. The training data includes compound information representing the structure or properties of the compounds and information representing their binding ability to target substances. Based on the acquired training data, a learning model is generated. The learning model learns to output information representing the binding ability to target substances when compound information is input.
[0014] Invention Effects
[0015] According to this disclosure, it is possible to construct a compound library that increases the proportion of compounds with target binding ability. Attached Figure Description
[0016] Figure 1 This is a diagram showing an outline of the compound library generation system of this embodiment.
[0017] Figure 2 This is a block diagram illustrating an example of the configuration of an information processing device and a terminal device.
[0018] Figure 3 This is an explanatory diagram showing an overview of the learning model and an example of the information stored in the training database.
[0019] Figure 4 This is a flowchart illustrating an example of the process of generating a learning model.
[0020] Figure 5 This is a flowchart illustrating an example of the process for generating a focus library.
[0021] Figure 6 This is a flowchart illustrating an example of a learning model generation process performed by the information processing apparatus of the second embodiment.
[0022] Figure 7 This is a flowchart illustrating an example of the process for generating predicted values by combining information in the third embodiment.
[0023] Figure 8 This is an explanatory diagram showing an outline of the learning model of the fourth embodiment and an example of the contents of the information stored in the training DB. Detailed Implementation
[0024] This disclosure will be described in detail with reference to the accompanying drawings illustrating embodiments thereof.
[0025] (First Implementation)
[0026] Figure 1 This diagram illustrates a general outline of the compound library generation system 100 according to this embodiment. The compound library generation system 100 includes an information processing device 1 as a main device. The information processing device 1 is communicatively connected to terminal devices 2 via a network N such as the Internet. The number of terminal devices 2 can be one or more.
[0027] Information processing device 1 is a device capable of various information processing, transmission, and reception, such as a server computer, personal computer, or quantum computer. Terminal device 2 is an information processing terminal used by a pharmaceutical company executive as an example user. Terminal device 2 can be, for example, a personal computer, smartphone, or tablet.
[0028] Information processing device 1 receives a base library 31, which is a compound library held by a pharmaceutical company, via terminal device 2. Information processing device 1 generates a focused library 32 corresponding to the received base library 31, and provides the generated focused library 32 to the pharmaceutical company via terminal device 2. Base library 31 corresponds to a first compound library, and focused library 32 corresponds to a second compound library.
[0029] Basic Library 31 refers to a compound database that manages compounds through database storage. Basic Library 31 stores a large number of compounds, including compounds obtained by pharmaceutical companies through previous pharmaceutical research and development, as well as compounds acquired from external sources. Basic Library 31 stores information related to these compounds, such as their names, structural formulas, physical properties, and physiological and chemical characteristics. Basic Library 31 can be an independently owned library by a pharmaceutical company, storing information on a large number of compounds owned independently by the company.
[0030] Focused Library 32 refers to a compound library generated based on Basic Library 31. It is a library that selectively stores compounds with target properties from the compounds stored in Basic Library 31. More specifically, Focused Library 32 selectively stores compounds with activity against target substances. Target substances can be any of DNA, RNA, proteins, etc.
[0031] In compound exploration in the pharmaceutical field, effective candidate compounds are screened from the compound group in the basic library 31. Investigating the activity of all compounds included in the basic library 31 against the target substance requires significant time and expense. In this system, the efficiency of the screening process is improved by providing a focusing library 32 that increases the proportion of compounds with activity against the target substance.
[0032] The following example illustrates the case of a focusing library 32 where the target substance is RNA and the generation of low-molecular-weight compounds that increase the proportion of binding to the target RNA. The target RNA can be non-coding RNA.
[0033] Figure 2 This is a block diagram illustrating an example configuration of an information processing device 1 and a terminal device 2. The information processing device 1 includes a control unit 11, a storage unit 12, a communication unit 13, a display unit 14, an operation unit 15, and an input / output unit 16. The information processing device 1 can be a multi-computer system composed of multiple computers, or it can be a virtual machine constructed by software.
[0034] The control unit 11 includes one or more computing devices such as CPUs (Central Processing Units), MPUs (Micro-Processing Units), and GPUs (Graphics Processing Units). The control unit 11 uses built-in memory such as ROM (Read Only Memory) or RAM (Random Access Memory), clocks, counters, etc., to control each component to perform processing. Furthermore, the functional units of the information processing device 1 can be implemented in software, in hardware (such as FPGAs or ASICs), or in a combination thereof.
[0035] Storage unit 12 may include, for example, non-volatile memory such as hard disk, flash memory, or SSD (Solid State Drive). Storage unit 12 may be an external storage device connected to information processing device 1. Storage unit 12 stores various computer programs and data referenced by control unit 11. In this embodiment, storage unit 12 stores a program 1P for causing the computer to perform processes related to the generation of focus library 32, a learning model 121, and a training database 122. Learning model 121 is a model generated through machine learning. It is envisioned that learning model 121 is used as a program module constituting part of artificial intelligence software.
[0036] The computer program (computer program product) including program 1P can be provided by a non-transitory storage medium 1A in a readable manner. Storage unit 12 stores the computer program read from storage medium 1A by a reading device (not shown). Storage medium 1A is, for example, a magnetic disk, optical disk, semiconductor memory, etc. Alternatively, the computer program can be downloaded from an external server connected to a communication network and stored in storage unit 12. Program 1P can be a single computer program or a program composed of multiple computer programs; furthermore, it can be executed on a single computer or on multiple computers interconnected via a communication network.
[0037] The communication unit 13 is equipped with a communication module that enables communication with external devices via network N. The control unit 11 transmits and receives various types of information with the terminal device 2 via the communication unit 13.
[0038] Display unit 14 includes, for example, a liquid crystal display (LCD) or an organic EL (Electro Luminescence) display. Display unit 14 displays information to be notified to the user according to instructions from control unit 11.
[0039] The operation unit 15 is an interface for receiving user operations. The operation unit 15 may include, for example, a keyboard, mouse, a touch panel device with a built-in display, a speaker, and a microphone. The operation unit 15 receives operation input from the user and sends control signals corresponding to the operation content to the control unit 11.
[0040] The input / output unit 16 has an input / output interface for connecting to external devices via wired or wireless means. A detection device 4 is connected to the input / output unit 16.
[0041] Detection device 4 is a device for detecting the interaction between target RNA and a compound. Detection device 4 uses, for example, surface plasmon resonance (SPR) to measure the interaction between target RNA and the compound. Furthermore, the method for analyzing the interaction is not limited to SPR; other methods include isothermal titration calorimetry, mass spectrometry, nuclear magnetic resonance spectroscopy, melting temperature measurement, absorption spectroscopy, fluorescence spectroscopy, and circular dichroism spectroscopy. By analyzing the interaction between target RNA and the compound, the presence or absence of binding ability to the target RNA, the binding strength, the binding ratio, and the presence or absence of structural changes during binding can be determined. Detection device 4 outputs the detection data obtained through detection to information processing device 1. Additionally, information processing device 1 can acquire the detection data stored in a predetermined storage area through detection device 4.
[0042] Terminal device 2 includes a control unit 21, a storage unit 22, a communication unit 23, a display unit 24, and an operation unit 25. The control unit 21 includes one or more processing units such as CPU, MPU, and GPU. The control unit 21 uses built-in memory such as ROM or RAM, clock, counter, etc., to control each component to perform processing.
[0043] Storage unit 22 may include, for example, non-volatile memory such as hard disk, flash memory, or SSD. Storage unit 22 stores various computer programs and data referenced by control unit 21. Storage unit 22 stores program 2P for causing the computer to perform processing related to acquiring focus library 32, and basic library 31. Storage unit 22 may also store focus library 32 received from information processing device 1.
[0044] The communication unit 23 includes a communication module that enables communication via network N. The display unit 24 includes, for example, a display device such as a liquid crystal display (LCD) or an organic EL display (OLED). The operation unit 25 is an interface for receiving user operations. The operation unit 25 includes, for example, a keyboard, a mouse, a touch panel device with a built-in display, a speaker, and a microphone.
[0045] Figure 3 This is an explanatory diagram showing an overview of the learning model 121 and an example of the contents of the information stored in the training DB 122.
[0046] Training DB122 is a database that stores the training data used to learn model 121. For example... Figure 3 As shown, the training DB122 stores records that associate compound information and binding information using the compound ID as a keyword. The compound groups that constitute the training data, i.e., the multiple compounds included in the training DB122, are selected from the compounds stored in the base library 31.
[0047] Compound information includes information representing the structure or physical properties of a compound. Compound information includes, for example, molecular descriptors. Molecular descriptors are numerical descriptors used to facilitate computer processing of the structural features, physicochemical properties, etc., of a substance. Molecular descriptors can be calculated from the structural formula of a substance and can be obtained using well-known software such as alvaDesc, Dragon, Codessa, RDKit, and Mordred. Compound information may include values for multiple molecular descriptors.
[0048] Compound information is not limited to molecular descriptors, but may also include, for example, structural formulas expressed as strings according to SMILES (Simplified Molecular Input Line Entry System) notation, molecular diagrams obtained by converting chemical structural formulas into graphical information, etc.
[0049] Binding information indicates the ability of a compound to bind to target RNA, including whether or not the compound binds to the target RNA. Binding information may also include the strength of binding to the target RNA, the binding ratio, and the presence or absence of structural changes during binding. Binding information is obtained by analyzing detection data, which represents the interaction between the compound and the target RNA obtained through experiments using detection device 4. For example, as a criterion for determining the presence or absence of binding ability, it can be set as follows: if the Resonance Unit (RU value) of the SPR signal obtained based on the compound is above a preset threshold, then it has binding ability.
[0050] Furthermore, the information used for training data can also be simulated values obtained through a predetermined algorithm. Examples of methods for calculating simulated values include quantum chemical calculations and molecular dynamics calculations. The simulated values obtained through these methods can be calculated using well-known theoretical calculation software.
[0051] like Figure 3 As shown, the learning model 121 takes the compound information as input and outputs binding information representing the binding ability of the compound to a specific target RNA. Hereinafter, the specific target RNA, which is the object of recognition by the learning model 121, will also be referred to as the first target. The learning model 121 of this embodiment outputs a classification result indicating whether or not the first target RNA has binding ability, i.e., whether it has binding ability or not. The learning model 121 is used for the screening of compounds when generating the focusing library 32.
[0052] The learning model 121 is constructed, for example, using a random forest. The information processing device 1 randomly selects features for model construction based on data sampled from the training data, creates multiple decision trees with low correlation, and uses these multiple decision trees to generate the learning model 121. The learning model 121 takes the majority decision or average of the estimates obtained from each decision tree as its final output.
[0053] The compound information used as input to the learning model 121 includes multiple molecular descriptors. Compound information may include SMILES strings, molecular diagrams, etc.
[0054] When inputting into the learning model 121, one or more molecular descriptors selected from a large number of molecular descriptors can be input. The types of molecular descriptors used as input to the learning model 121 can be appropriately set according to the type of the first target RNA. As an example, the molecular descriptors used as input to the learning model 121 are automatically determined by the information processing device 1 based on the contribution of the input information obtained during the learning process of the learning model 121.
[0055] The information processing device 1 performs estimation processing by the learning model 121 using a pre-selected predetermined number of molecular descriptors to calculate the contribution (variable importance) of each molecular descriptor in the combined information estimation of the learning model 121. The contribution can be calculated based on, for example, SHAP (SHapley Additive exPlanation) values, Gini coefficients, out-of-bag (OOB) data, LIME (Local Interpretable Model-Agnostic Explanations), etc. The information processing device 1 determines the molecular descriptors used for input information into the learning model 121 by preferentially selecting a predetermined number of molecular descriptors in descending order of contribution based on the calculated contribution of each molecular descriptor. Through this process, suitable molecular descriptors for input into the learning model 121 can be determined from a large number of defined molecular descriptors.
[0056] The combined information used as the output of learning model 121 includes whether or not the compound has the ability to bind to the first target RNA.
[0057] The learning model 121 can be generated by preparing training data labeled with the presence or absence of the ability to bind to the first target RNA as positive values for compound information, and using the training data to enable the unlearned model to perform machine learning.
[0058] The information processing device 1 inputs compound information from the training data into the learning model 121, and enables the learning model 121 to learn in a manner that makes the output of the learning model 121 approximate the positive value. The information processing device 1 generates the learning model 121 by adjusting the parameters in the learning model 121 by using the presence or absence of binding ability corresponding to the input compound information as the positive value. Through the above processing, a learning model 121 can be constructed that can appropriately output the presence or absence of binding ability to the first target RNA in response to compound information.
[0059] Furthermore, the output information from the learning model 121 is not limited to whether or not the binding ability to the first target RNA exists. The learning model 121 can also be configured to output, for example, the magnitude of the binding strength to the first target RNA, the binding ratio, and whether or not there are structural changes during binding.
[0060] The structure of learning model 121 is not limited to Figure 3 The example shown only needs to identify information about the binding ability of the compound to the target RNA. The learning model 121 can also be a model based on other learning algorithms, such as CNN (Convolution Neural Network), RNN (Recurrent Neural Network), GNN (Graph Neural Network), Transformer, Support Vector Machine, Logistic Regression, XGBoost (eXtreme Gradient Boosting).
[0061] The learning model 121 can also consist of multiple separate learning models constructed using different learning algorithms. For example, the learning model 121 may include a first separate learning model constructed using random forest and a second separate learning model constructed using XGBoost.
[0062] In the compound library generation system 100, a learning model 121 is constructed using training data, which is generated based on compound information of a subset of compounds stored in the base library 31. Using the obtained learning model 121, the binding ability of a large number of compounds stored in the base library 31 is estimated. By extracting compounds estimated to have binding ability, a new focusing library 32 is generated. The generation process of the focusing library 32 is described in detail below.
[0063] Figure 4 This is a flowchart illustrating an example of the generation process of the learning model 121. The processes in the following flowcharts are executed by the control unit 11 according to the program 1P stored in the storage unit 12 of the information processing device 1.
[0064] The control unit 11 of the information processing device 1 receives from the user, via the terminal device 2, the selection of the target RNA as the object of the focusing library 32, and user identification information identifying the user (step S11). The control unit 11 receives, for example, information indicating the name and structure of the target RNA from the terminal device 2 based on the operation of the user's operation unit 25. The user can select any target RNA (e.g., a first target RNA) according to the purpose of pharmaceutical research and development.
[0065] The control unit 11 acquires a compound group, including multiple compounds used for training data, via the terminal device 2 (step S12). The compound group is extracted from compounds stored in the base library 31. Based on the operation of the user's operation unit 25, the control unit 11 receives information from the terminal device 2 regarding the compounds stored in the base library 31, including information necessary for the processing described later (e.g., compound name, molecular structure, necessary physical properties, etc.). The user can arbitrarily select compounds for training data according to the purpose of pharmaceutical research and development. The user determines the number of compounds for training data based on the time and cost allowed for generating training data, and extracts compounds from the base library 31 to reach the determined number. Alternatively, the compound group for training data can also be determined by the information processing device by randomly extracting a predetermined number from the base library 31.
[0066] The control unit 11 acquires binding information for each compound in the acquired compound group. This binding information includes the presence or absence of binding ability to the first target RNA, the magnitude of the binding strength, the binding ratio, and the presence or absence of structural changes during binding (step S13). The binding information to the first target RNA is obtained by analyzing the interactions between each compound and the target RNA detected by the detection device 4. The control unit 11 can acquire the binding information, for example, by receiving the binding ability determination results obtained manually via the operation unit 15 or the communication unit 13, or it can automatically derive the binding information based on the detection data received from the detection device 4 via the input / output unit 16.
[0067] The control unit 11 acquires compound information for each compound in the compound group (step S14). For example, the control unit 11 derives a SMILES-based structural formula from the molecular structure of the compound acquired in step S12, and calculates various preset molecular descriptors based on the derived structural formula.
[0068] The control unit 11 associates the obtained compound information with the information representing binding ability and stores it in the training DB122 (step S15). Through the above processing, training data is generated.
[0069] Based on the information stored in the training DB122, the control unit 11 obtains training data (step S16) which is labeled with the presence or absence of the ability to bind to the first target RNA for the compound used for training.
[0070] The control unit 11 uses the acquired training data to generate a learning model 121, which outputs whether or not the compound has the ability to bind to the first target RNA when compound information is input (step S17). Specifically, the control unit 11 inputs the compound information included in the training data into the learning model 121 and optimizes various parameters in a way that makes the output from the learning model 121 approximate the positive solution value. For example, after learning is completed by meeting a predetermined benchmark through a certain number of learning iterations, the control unit 11 stores the definition information related to the completed learning model 121 as the completed learning model 121 in the storage unit 12. Through the processing from steps S16 to S17, the learning model 121 is constructed.
[0071] The compound set used to generate the training data for the learning model 121 is not limited to compounds extracted from the base library 31, i.e., the library that serves as the generation object of the focus library 32. The learning model 121 may also be constructed using information from a compound library that is different from the generation object of the focus library 32 as training data.
[0072] Figure 5 This is a flowchart illustrating an example of the generation process of the focus library 32. The control unit 11 of the information processing device 1, for example, is... Figure 4 After step S17 in the flowchart is completed, the following processing is performed.
[0073] The control unit 11 of the information processing device 1 receives the basic library 31 held by the user from the terminal device 2 (step S21). In this case, the control unit 11 may also receive only the information required for the processing described later (e.g., compound name, molecular structure, etc.) from the information related to each compound contained in the basic library 31. In step S11, the control unit 11 obtains information about the compounds to be evaluated from the compounds contained in the basic library 31. The compounds to be evaluated may be all the compounds stored in the basic library 31, or a selection from all the compounds.
[0074] For each compound in the acquired base library 31, the control unit 11 acquires compound information, including molecular descriptors, structural formulas, molecular diagrams, etc. (step S22). The control unit 11 inputs the acquired compound information of each compound into the learning model 121 (step S23) and acquires the presence or absence of binding ability output from the learning model 121 (step S24).
[0075] Based on the estimated results of the acquired binding ability, the control unit 11 classifies each compound in the basic library 31 into either the hit group or the miss group (step S25). The hit group consists of compounds that have the ability to bind to the first target RNA, while the miss group consists of compounds that do not have the ability to bind to the first target RNA. The process in step S25 corresponds to the screening process of selecting compounds with the ability to bind to the first target RNA from the basic library 31.
[0076] The control unit 11 obtains a score by numerically quantifying the priority of each compound classified into the hit group (step S26). The priority is quantified in such a way that the higher the probability that the compound has the ability to bind to the first target RNA, the higher the score. The control unit 11 can obtain the classification accuracy value in the learning model 121 and use the obtained classification accuracy as the score. The classification accuracy is obtained, for example, based on the proportion of decision trees consistent with the estimate, and the reliability of the classification category in the output node.
[0077] Based on the classification results, the control unit 11 generates a focus library 32 containing compounds classified into the hit group (step S27). In this case, the control unit 11 sorts the compounds classified into the hit group according to their priority scores in descending order of score, and adds them to the focus library 32 sequentially from the top-ranked compounds. The generated focus library 32 only needs to contain information that can identify the compounds included in the focus library 32; for example, it can be a list of compounds arranged and displayed in order of score. The focus library 32 may also be associated with the scores of each compound.
[0078] The control unit 11 sends the generated focus library 32 to the terminal device 2 corresponding to the user identified by the user identification information, i.e. the user of the basic library 31 (step S28), and ends a series of processes.
[0079] The user uses the focusing library 32 provided by the information processing device 1 to explore pharmaceutical targets. Since the focusing library 32 is composed of compounds estimated by the learned model 121 to have the ability to bind to the first target RNA, it is possible to obtain compounds with activity against the first target RNA with a high hit rate.
[0080] The above describes examples of how information processing device 1 executes a series of processes in each flowchart, but the processing entity of each process is not limited. Part of the process in each flowchart may also be executed by terminal device 2.
[0081] Whenever the information processing device 1 receives a request from the terminal device 2 to generate a focusing library 32 targeting a new target RNA, it repeatedly performs a series of processes including the generation of training data, the generation of learning models 121, and the generation of the focusing library 32. Specifically, different learning models 121 are prepared according to the types of target RNAs, and the prepared learning models 121 corresponding to each target RNA are used to screen compounds. When a user selects a second target RNA as the target substance, the information processing device 1 generates a learning model 121 that estimates binding information related to the second target RNA, and uses this learning model 121 to evaluate the binding ability of each compound to the second target RNA, thereby generating a new focusing library 32.
[0082] In the above process, the control unit 11 may omit the score acquisition process in step S26. Alternatively, the control unit 11 may further filter the compounds stored in the focusing library 32 based on the score obtained in step S26. For example, the control unit 11 may extract only compounds with scores above a threshold to generate the focusing library 32.
[0083] When the learning model 121 is used to further estimate the binding strength, binding ratio, and presence or absence of structural changes during binding, the control unit 11 can also consider these estimation results to classify the compounds in the basic library 31. For example, the control unit 11 classifies the compounds into three or more groups based on the presence or absence of binding ability to the first target RNA and the binding strength. The control unit 11 stores the compounds belonging to the groups that meet the preset screening criteria (e.g., having binding ability and strong binding strength) into the focusing library 32.
[0084] The control unit 11 can also repeatedly perform the classification performed by the learning model 121, calculate the total value of the priorities calculated each time, i.e., the total score, and sort them based on the calculated total score. For example, the control unit 11 will finally extract a predetermined number of compounds and compounds with a total score of more than a preset threshold in descending order of total score, and store the extracted compounds in the focusing library 32.
[0085] When the learning model 121 includes multiple individual learning models, the control unit 11 can also classify compounds into a hit group and a miss group based on the output results of each individual learning model. For example, the control unit 11 obtains the presence or absence of binding ability output by each individual learning model for a particular compound. If a classification result indicating binding ability is obtained from all individual learning models or individual learning models above a preset threshold, the control unit 11 classifies the compound into the hit group.
[0086] The compounds evaluated may also include those included in the compound group. In this case, instead of estimation by the learning model 121, the compounds included in the compound group can be classified into hit or miss groups based on binding information obtained through actual measurements.
[0087] The focus library 32 is not limited to including only compounds estimated to have binding ability to the first target RNA, but may also include some compounds estimated not to have binding ability to the first target RNA.
[0088] According to this embodiment, a high-quality focusing library 32 can be generated based on the user's base library 31, which increases the proportion of compounds with binding ability to target RNA. By generating the focusing library 32 from compounds extracted from the base library 31 owned by the pharmaceutical company, a focusing library 32 that can improve exploration efficiency while ensuring uniqueness can be provided.
[0089] By using the learning model 121, the screening of compounds from the basic library 31 can be performed efficiently and with high accuracy. Training data is generated by actually measuring and evaluating the interaction between the compounds and the target substance, thereby improving the estimation accuracy of the learning model 121. By generating the learning model 121 based on the first target RNA, the binding affinity to the first target RNA can be estimated with high accuracy using the learning model 121.
[0090] By quantifying the priority of each compound in Focus Library 32, the exploration efficiency of Focus Library 32 can be further improved. By receiving the target RNA selection from the user, Focus Library 32 suitable for any target RNA can be obtained, improving the convenience of this system.
[0091] Compared to low-molecular-weight compounds targeting proteins, research on low-molecular-weight compounds targeting non-coding RNA is still insufficient, and there are many cases where the basic library 31 does not store a sufficient number of compounds. Therefore, applying this system when the target substance is RNA will further enhance the technical effects of this disclosure.
[0092] (Modified example)
[0093] Figure 5 The generation process of the focus library 32 shown can also be performed using the terminal device 2. In this case, steps S21 and S28 can be omitted.
[0094] Information processing device 1 executes Figure 4 The learning model 121 shown is generated and stored in an area accessible to the terminal device 2. The information processing device 1 can also deploy the generated learning model 121 to the terminal device 2.
[0095] Terminal device 2 accesses the learning model 121 stored in an external device, or refers to the storage unit 22 of this device to read the learning model 121, and uses the read learning model 121 to perform the processing of steps S22 to S27, such as performing compound screening processing.
[0096] Based on the above configuration, a focusing library 32 can be generated on the terminal device 2 side. This eliminates the need to provide information about the compounds to be evaluated from the terminal device 2 to the information processing device 1, thus improving the confidentiality of the basic library 31.
[0097] (Second Implementation)
[0098] In the second embodiment, compounds that do not have the ability to bind to the first target RNA are removed from the compound group. In the following embodiments, the differences from the first embodiment will be mainly described; for components common to the first embodiment, the same reference numerals will be used and detailed descriptions will be omitted.
[0099] The compounds included in training DB122, i.e., the compounds selected from the basic library 31, represent the binding information obtained through interaction resolution. This includes both compounds with binding ability and those without. It is assumed that the number of compounds without binding ability to the first target RNA in the compound group is greater than the number of compounds with binding ability to the first target RNA. That is, it is assumed that the data included in training DB122 is unbalanced data with a skewed ratio of compounds with binding ability to those without.
[0100] When the learning model 121 is generated using such imbalanced data, it is possible to reduce the sensitivity to compounds with binding ability. The information processing apparatus 1 of the second embodiment improves the accuracy of the learning model 121 by reducing a portion of the compounds in the compound group that do not have binding ability.
[0101] Figure 6 This is a flowchart illustrating an example of the generation process of the learning model 121 performed by the information processing apparatus 1 in the second embodiment.
[0102] The control unit 11 of the information processing device 1 executes the processing steps S11 to S15 of the first embodiment, storing the compound information and binding information of all compounds in the compound group obtained in step S15 into the training DB 122. Based on the compound information of each compound stored in the training DB 122, the control unit 11 classifies each compound in the compound group into compounds that have the ability to bind to the first target RNA and compounds that do not have the ability to bind to the first target RNA (step S31).
[0103] Based on the classification results, the control unit 11 removes compounds from the compound group that do not have the ability to bind to the first target RNA (step S32). The control unit 11 removes compounds that do not have the ability to bind to the first target RNA by setting the ratio (number of compounds with binding ability : number of compounds without binding ability) in the compound group to, for example, 1:1, 1:4, 1:8, 1:16, etc. The above ratio is an example, and the ratio in the removal process is not necessarily limited to this value.
[0104] After processing in step S32, the control unit 11 performs the same processing as in steps S16 to S17, using the data of the compound group after removing some compounds as training data to generate a learning model 121.
[0105] According to this embodiment, the decrease in estimation accuracy of the learning model 121 caused by imbalance of training data can be suppressed.
[0106] (Third Implementation)
[0107] In the third embodiment, a learning model 121 for estimating the binding information of other target substances is generated using predicted values of binding information of other target substances generated based on binding information of a specific target substance and the correlation between target substances.
[0108] As described above, the learning model 121 learns using training data obtained by associating compound information with binding information. In the third embodiment, instead of the binding information being actual measurement data obtained using a predetermined mutual resolution method, the predicted value of the binding information is used as training data. The predicted value of the binding information can be obtained using binding information already obtained through actual measurement, taking into account the correlation between multiple target RNAs.
[0109] Figure 7This is a flowchart illustrating an example of the process for generating predicted values of binding information in the third embodiment. Hereinafter, as an example, it is assumed that the new target substance is a second target RNA obtained by replacing adenine at a specific position in the RNA structure of the first target RNA with uracil. Furthermore, it is assumed that the presence or absence of binding ability for the first target RNA, corresponding to actual measurement data, has been stored in the training DB122.
[0110] The control unit 11 of the information processing device 1 calculates the correlation coefficient representing the strength of the correlation function between the first target RNA and the second target RNA (step S41). The correlation coefficient between the first target RNA and the second target RNA is calculated, for example, by comprehensively evaluating the differences in the material structure of the first target RNA and the second target RNA using various material-related indicators (e.g., the number of hydrogen bond donors, the number of hydrogen bond acceptors, surface area, volume, etc.), specifically the difference between adenine and uracil.
[0111] Based on the calculated correlation coefficient and the binding information of the first target RNA stored in the training DB122, the control unit 11 derives a predicted value for the binding information of the second target RNA (step S42). For example, the control unit 11 calculates the predicted value for the presence or absence of binding ability of the second target RNA by combining the calculated correlation coefficient with the presence or absence of binding ability of the first target RNA.
[0112] The control unit 11 associates the compound information with the predicted value of the binding information to the second target RNA and stores it as training data in the training DB122 (step S43).
[0113] Subsequently, the control unit 11 performs the same processing as steps S16-S17 of the first embodiment, using training data including predicted values of binding information to the second target RNA, to generate a learning model 121 that estimates the presence or absence of binding ability to the second pair of target RNAs based on compound information. The control unit 11 uses the generated learning model 121 to screen compounds with binding ability to the second target RNA to generate a focusing library 32 related to the second target RNA.
[0114] According to this embodiment, the time and cost required for experiments to generate training data can be reduced, making it easier to generate the focus library 32. Information about the compounds is not required to generate training data, reducing the burden on pharmaceutical companies. Even when targeting substances with interactions that are difficult to detect in practice, the focus library 32 can be generated, improving convenience.
[0115] (Fourth Implementation)
[0116] In the fourth embodiment, a learning model 121 is generated that can estimate binding information for multiple target RNAs.
[0117] Figure 8 This is an explanatory diagram showing an outline of the learning model 121 of the fourth embodiment and an example of the content of the information stored in the training DB 122.
[0118] like Figure 8 As shown, the learning model 121 of the fourth embodiment takes compound information and target substance information of the target substance as input, and outputs the binding information of the compound to the target substance. The target substance information includes information related to the target substance, such as information representing the substance name, primary structure (sequence), secondary structure, etc.
[0119] Information processing device 1, for example, accumulates binding information obtained through the generation process of focus library 32 for multiple target substances into training DB 122. Based on the accumulated information, information processing device 1 generates training data in which the binding information of the aforementioned compounds to the aforementioned target substances is labeled as positive values, for both compound information and target substance information. Information processing device 1 uses the generated training data to train learning model 121. Based on the learned learning model 121, binding information for multiple target substances can be estimated.
[0120] According to this embodiment, binding information for multiple target substances can be estimated using a single learning model, thus eliminating the need to generate a learning model 121 for each target substance, making the generation of the focus library 32 easier.
[0121] It should be understood that the embodiments disclosed herein are exemplary in all respects and not restrictive. The technical features described in the various embodiments can be combined with each other, and the scope of the invention is intended to include all variations of the claims and their equivalents.
[0122] The sequences shown in each embodiment are not limited. Within the scope of non-contradiction, the order of each processing step can be changed to perform the process, and multiple processes can be performed in parallel. The processing entity of each process is not limited. Within the scope of non-contradiction, the processes of each device can also be performed by other devices.
[0123] The contents described in each embodiment can be combined with each other. Furthermore, the independent and dependent claims described in the claims can be combined with each other in all combinations, regardless of their referencing form. Moreover, the claims use a form that refers to two or more other claims (multiple reference form), but are not limited to this. A form that refers to at least one multiple reference claim (multiple reference multiple claims) can also be used.
[0124] Explanation of reference numerals in the attached figures:
[0125] 100 Compound Library Generation System
[0126] 1. Information processing device
[0127] 11 Control Department
[0128] 12 Storage Department
[0129] 13 Ministry of Communications
[0130] 14 Display Section
[0131] 15 Operations Department
[0132] 16 Input / Output Section
[0133] 121 Learning Model
[0134] 122 Training the DB
[0135] 1P program
[0136] 1A Storage Media
[0137] 2. Terminal device
[0138] 21 Control Department
[0139] 22 Storage Department
[0140] 23 Ministry of Communications
[0141] 24 Display Section
[0142] 25 Operations Department
[0143] 2P program
[0144] 2A storage medium
[0145] 31. Basic Library
[0146] 32 Focused Library
Claims
1. A method of generating a library of compounds, characterized by, The following process is executed by a computer: Obtaining information on a plurality of compounds stored in a first compound library; Classifying the plurality of compounds stored in the first compound library into a group of compounds having a binding ability to a target substance and a group of compounds not having a binding ability to the target substance using a learning model that learns to output information indicating the binding ability to the target substance when compound information on a compound is input; Generating a second compound library including compounds classified into the group of compounds having a binding ability to the target substance.
2. The compound library generation method according to claim 1, wherein the learning model is learned based on training data including the compound information indicating a structure or a property of a compound and information indicating the binding ability.
3. The compound library generation method according to claim 2, wherein the information indicating the binding ability in the training data is obtained by analyzing an interaction between a compound and a target substance using a surface plasmon resonance method, an isothermal titration calorimetry method, a mass spectrometry method, a nuclear magnetic resonance spectroscopy method, a melting temperature measurement method, an absorption spectroscopy method, a fluorescence spectroscopy method, or a circular dichroism spectroscopy method.
4. The compound library generation method according to claim 2 or 3, wherein the training data includes the compound information and information indicating the binding ability in a group of compounds extracted from the first compound library.
5. The compound library generation method according to claim 4, wherein a part of the plurality of compounds indicating no binding ability to the target substance is removed from the group of compounds by analyzing an interaction between a compound and a target substance, the learning model is learned based on the training data including the compound information and information indicating the binding ability in the group of compounds after the removal.
6. The compound library generation method according to any one of claims 1 to 5, wherein information indicating a binding ability to a second target substance is derived based on a correlation between a first target substance and the second target substance and information indicating a binding ability to the first target substance obtained using a predetermined interaction analysis method, the learning model is learned based on training data including the derived information indicating a binding ability to the second target substance.
7. The compound library generation method according to any one of claims 1 to 6, wherein the compound information includes a plurality of molecular descriptors.
8. The compound library generation method according to claim 7, wherein a molecular descriptor used for input to the learning model is determined among the plurality of molecular descriptors based on a contribution degree of the molecular descriptor to the information indicating the binding ability.
9. The compound library generation method according to any one of claims 1 to 8, wherein the learning model is prepared for each kind of the target substance.
10. The compound library generation method according to any one of claims 1 to 9, wherein acquiring numerical information obtained by numerically expressing priority of each compound included in the second compound library, generating the second compound library based on the acquired numerical information.
11. The compound library generation method according to any one of claims 1 to 10, wherein receiving selection of a target substance from a user of the first compound library, generating the second compound library including a compound having a binding ability to the selected target substance.
12. The compound library generation method according to any one of claims 1 to 11, wherein the target substance is a target RNA.
13. A compound library generation system characterized by comprising: provided with a control unit, the control unit performs the following processing: acquiring information of a plurality of compounds stored in a first compound library; using a learning model that learns to output information indicating a binding ability to a target substance when compound information of a compound is input, classifying the plurality of compounds stored in the first compound library into a compound group having a binding ability to the target substance and a compound group not having a binding ability to the target substance; generating a second compound library including a compound classified into the compound group having a binding ability to the target substance.
14. A computer program, characterized in that, causing a computer to perform the following processing: acquiring information of a plurality of compounds stored in a first compound library; using a learning model that learns to output information indicating a binding ability to a target substance when compound information of a compound is input, classifying the plurality of compounds stored in the first compound library into a compound group having a binding ability to the target substance and a compound group not having a binding ability to the target substance; generating a second compound library including a compound classified into the compound group having a binding ability to the target substance.
15. A learning model generation method, comprising: acquiring, for a plurality of compounds stored in a compound library, training data including compound information indicating a structure or a property of a compound and information indicating a binding ability to a target substance, generating a learning model that learns to output information indicating a binding ability to a target substance when compound information is input, based on the acquired training data.
Citation Information
Patent Citations
Compound library by transferase and method for producing the same
JP2006223198A