Systems and methods for molecule design

The use of a generative machine learning model automates drug discovery by iteratively generating and optimizing molecules, addressing the inefficiencies of human-driven processes and producing high-quality, diverse drug candidates efficiently.

WO2026107054A1PCT designated stage Publication Date: 2026-05-21OHIO STATE INNOVATION FOUND
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
OHIO STATE INNOVATION FOUND
Filing Date
2025-11-12
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Drug discovery is laborious and time-consuming due to the reliance on human judgment and expertise, limiting the development of effective pharmaceutical candidates.

Method used

A computer-implemented method using a generative machine learning model, such as a large language model, to automate the drug discovery process by iteratively generating, evaluating, and optimizing molecules to meet specific drug specifications, incorporating a molecule design module, reasoner module, and evaluator module for automated drug candidate generation and validation.

Benefits of technology

Accelerates drug discovery by systematically generating high-quality, diverse, and novel drug candidates with improved efficiency and reduced human bias, enhancing the drug development process through automated molecular design and evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025055128_21052026_PF_FP_ABST
    Figure US2025055128_21052026_PF_FP_ABST
Patent Text Reader

Abstract

An example method includes receiving a drug specification; selecting, by a reasoner module, a workflow based on the drug specification, wherein the reasoner module comprises a large language model configured to operate in inference mode responsive to the drug specification; executing, by a molecule design module, the workflow to produce a workflow output; evaluating, by an evaluator module, the workflow output; iteratively repeating steps of executing and evaluating until the drug specification is met by the workflow output; and outputting the workflow output, where the workflow output comprises a molecule based on the drug specification.
Need to check novelty before this filing date? Find Prior Art

Description

MCC Ref. No: 103362-063PV1T2025-038 SYSTEMS AND METHODS FOR MOLECULE DESIGN CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U. S. provisional patent application No.63 / 719,390, filed on November 12, 2024, and titled “Systems and Methods for Molecule Design,” the disclosure of which is expressly incorporated herein by reference m its entirety'.STATEMENT REGARDING FEDERALLY FUNDED RESEARCH

[0002] This invention was made with government support under R01 LM014385 awarded by the National Institutes of Health and 2133650 awarded by the National Science Foundation. The government has certain rights in the invention.BACKGROUND

[0003] Drug discovery involves identifying molecules with desirable properties that can be used as candidates for developing pharmaceuticals. Drug discovery can include many different steps, including identifying candidate molecules, simulating those molecules, and synthesizing the molecules for experimentation. Systems and methods that automate steps of the drug discovery process can benefit drug discovery'.SUMMARY

[0004] In some aspects, implementations of the present disclosure include a computer-implemented method including: receiving a drug specification; selecting, by a reasoner module, a workflow based on the drug specification, wherein the reasoner module includes a large language model configured to operate in inference mode responsive to the drug specification;MCC Ref. No: 103362-063PV1T2025-038 executing, by a molecule design module, the workflow to produce a workflow output; evaluating, by an evaluator module, the workflow output; iteratively repeating steps of executing and evaluating until the drug specification is met by the workflow output; and outputting the workflow output, wherein the workflow output includes a molecule based on the drug specification.

[0005] In some aspects, implementations of the present disclosure include a computer- implemented method, wherein the molecule design module includes a plurality of components.

[0006] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein selecting the workflow includes selecting at least one component of the molecule design module.

[0007] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein the molecule design module includes a hardware module.

[0008] In some aspects, implementations of the present disclosure include a computer- implemented method, further including storing in a memory at least one of the drug specification, workflow, and drug candidate.

[0009] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein the memory includes a key-value store of drug candidates generated by the workflow.

[0010] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein the method further includes receiving an output of the molecule design module by the reasoner, and generating a second workflow by the reasoner module.MCC Ref. No: 103362-063PV1T2025-038

[0011] In some aspects, implementations of the present disclosure include a computer- implemented method, wherein the evaluator module includes at least one of a docking simulator or a property predictor.

[0012] In some aspects, implementations of the present disclosure include a system for drug discovery’, the system including: a molecule design module; a reasoner module including a large language model, wherein the reasoner module is configured to receive a drug specification and configure the molecule design module based on the drug specification; and an evaluator module configured to evaluate an output of the molecule design module and output a molecule that satisfies the drug specification.

[0013] In some aspects, implementations of the present disclosure include a system, wherein the molecule design module includes a molecule generator configured to generate a molecule based on an instruction from the reasoner module.

[0014] In some aspects, implementations of the present disclosure include a system, wherein the molecule design module is configured to generate molecules configured to fit in a protein pocket.

[0015] In some aspects, implementations of the present disclosure include a system, wherein the evaluator module includes a docking simulator and a property predictor.

[0016] In some aspects, implementations of the present disclosure include a system, wherein the property predictor is configured to determine at least one of absorption, distribution, metabolism, excretion, or toxicity properties of a molecule generated by the molecule design module.MCC Ref. No: 103362-063PV1T2025-038

[0017] In some aspects, implementations of the present disclosure include a system, wherein the docking simulator is configured to predict a preferred orientation or interaction strength between a protein binding site and a small molecule.

[0018] In some aspects, implementations of the present disclosure include a system, wherein the reasoner module includes a first computing device configured to operate the large language model in inference mode, and wherein the molecule design module includes a second computing device in operable communication with the first computing device, and wherein the evaluator module includes a third computing device in operative communication with the first and second computing devices.

[0019] In some aspects, implementations of the present disclosure include a system, wherein the molecule design module includes a lead optimizer.

[0020] In some aspects, implementations of the present disclosure include a system, wherein the lead optimizer is configured to optimize a molecule generated by the molecule design module by modifying the molecule.

[0021] In some aspects, implementations of the present disclosure include a system, wherein the molecule design module further includes a synthesis planner.

[0022] In some aspects, implementations of the present disclosure include a system, wherein the synthesis planner is configured to generate a synthetic route of drug synthesis.

[0023] In some aspects, implementations of the present disclosure include a system, wherein the molecule design module includes a specialized hardware module configured for drug synthesis, drug discovery, or both drug synthesis and discovery.MCC Ref. No: 103362-063PV1T2025-038

[0024] It should be understood that the above-described subject matter may also be implemented as a computer-controlled apparatus, a computer process, a computing system, or an article of manufacture, such as a computer-readable storage medium.

[0025] Other systems, methods, features and / or advantages will be or may become apparent to one with skill in the art upon examination of the following drawings and detailed description. It is intended that all such additional systems, methods, features and / or advantages be included within this description and be protected by the accompanying claims,BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The components in the drawings are not necessarily to scale relative to each other, Like reference numerals designate corresponding parts throughout the several views.

[0027] FIG. 1 illustrates a flow chart of an example method for drug discovery, according to implementations of the present disclosure.

[0028] FIG. 2A illustrates an example system for drug discovery, according to an example implementation of the present disclosure.

[0029] FIG. 2B illustrates an example system for drug discovery including a self driving laboratory, according to an example implementation of the present disclosure.

[0030] FIG. 3 is an example computing device.

[0031] FIG. 4 illustrates an example of action transitions in workflows in an example implementation of the present disclosure showing the probabilities of different transitions.

[0032] FIG. 5 illustrates molecule quality and actions over iterations by an example implementation of the present disclosure.MCC Ref. No: 103362-063PV1T2025-038

[0033] FIGS. 6A-6C illustrate example action trajectories for target molecules, according to a study of an example implementation of the present disclosure.

[0034] FIG. 7 illustrates results of an ablation study on an example implementation of the present disclosure.

[0035] FIG. 8A-8C illustrate an example case study on AR / NR3C4 wherein FIG. 8A illustrates a promising molecule NL- 1 generated by an example implementation of the present disclosure, FIG. 8B illustrates a known ligand for AR / NR3C4, and FIG. 8C illustrates an example of a known approved drug for AR / NR3C4.

[0036] FIG. 9 A illustrates Docking of NL-1 within the AR / NR3C4 pocket, with hydrogen bonds shown as solid blue lines,

[0037] FIG. 9B illustrates NL-1 superimposed with the known ligand for the AR / NR3C4 pocket.

[0038] FIGS. 10A-10C illustrate example case studies of an example implementation of the present disclosure.

[0039] FIGS. 11A-11C illustrate a case study of an example implementation of the present disclosure. FIG. 11 A illustrates generated molecules. FIG. 11B illustrates a known ligand, and FIG. 11C illustrates known approved drugs.

[0040] FIG. 12A illustrates docking of NL-2 to EGFR pocket, with hydrogen bonds shown. FIG. 12B illustrates docking of NL-3 to EGFR pocket, with hydrophobic van der Waals contacts shown.MCC Ref. No: 103362-063PV1T2025-038 DETAILED DESCRIPTION

[0041] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary' skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure. As used in the specification, and in the appended claims, the singular forms “a,” “an,” “the” include plural referents unless the context clearly dictates otherwise. The term “comprising” and variations thereof as used herein is used synonymously with the term “including” and variations thereof and are open, non-limiting terms. The terms “optional” or “optionally” used herein mean that the subsequently described feature, event or circumstance may or may not occur, and that the description includes instances where said feature, event or circumstance occurs and instances where it does not. Ranges may be expressed herein as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, an aspect includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms another aspect. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. While implementations will be described for drug discovery, it will become evident to those skilled in the art that the implementations are not limited thereto, but are applicable for designing any type of molecule.

[0042] Drug discovery is highly laborious and includes many different steps. For example, identifying candidate molecules, simulating those molecules, determining if the molecules can be synthesized, synthesizing the molecules, optimizing the molecules, etc. can all involve human judgment and expertise. The use of human judgment and expertise canMCC Ref. No: 103362-063PV1T2025-038 significantly slow the development of drugs. Additional example tasks that can be required for drug discovery’ include pathway analysis, target identification, molecule generation, molecule optimization, property’ prediction, retrosynthesis prediction, yield prediction.

[0043] Implementations of the present disclosure include improvements to these and other steps of the drug discovery process that increase the rate of drug discovery. In particular, the example implementations described herein can use a generative machine learning model (e.g., a large language model) as a “reasoner” to operate a specialized molecule design module for drug discovery and iteratively output candidate molecules for experimental validation.

[0044] As used herein, a “large language model” or “LLM” is a type of machine learning model configured (e.g., trained) to generate textual outputs. As used herein, a LLM can be any language model that has been trained on sufficiently large amounts of textual data to perform the functions of the disclosed systems and / or methods. In some aspects, any available LLM can be used in the disclosed system and method, including LLMs trained as general- purpose LLMs, fine-tuned versions of general-purpose LLMs, and / or LLMs trained for drug discovery. As used herein, “generative Al” includes LLMs, and generative Al refers to machine learning models configured to produce new content based on learned patterns.

[0045] With reference to FIG. 1, an example computer-implemented method of drug discovery is shown according to implementations of the present disclosure. The method of FIG.1 can optionally' be implemented using the system 200 shown m FIG. 2A.

[0046] At step 110, the method includes receiving a drug specification. The drug specification can optionally be in the form of a user prompt with information about a target of a desired drug.MCC Ref. No: 103362-063PV1T2025-038

[0047] At step 120, the method includes selecting a workflow using a reasoner module based on the drug specification. Optionally, the reasoner module is a large language model (LLM) or is configured to access an LLM (e.g., an LLM located in memory of a remote computing device). The LLM can be configured to operate in inference mode, providing responses to user inputs. Optionally, the method can include formatting the drug specification and / or the workflow as one or more prompts to the LLM, which cause the LLM to execute the workflows described herein to generate molecules that meet the drug specification.

[0048] At step 130, the method includes executing the workflow using a molecule design module. Optionally, the molecule design module can include multiple specialized components (e.g., hardware specific to performing or accelerating molecule design), and the workflow can include a selection of which components to activate and / or what order the components can be activated in. The reasoner module 210 can be configured to determine the workflow based on the input at step 110.

[0049] The molecule design module can produce a workflow output that can represent one or more molecules or intermediate steps in designing a molecule according to the drug specification.

[0050] At step 140, an evaluator module is used to evaluate the workflow output.

[0051] At step 150, the method can include iteratively repeating step 130 of executing and step 140 of evaluating. Step 130 and step 140 can optionally be performed any number of times until the workflow output includes a molecule evaluated by the evaluator module to meet the drug specification input at step 110. Any number of workflows can be generated and any number of drug candidates can be generated based on those workflows.MCC Ref. No: 103362-063PV1T2025-038

[0052] At step 160, the method further includes outputting the workflow output, which includes the design of one or more molecules that meet the drug specification.

[0053] In some implementations, the method can further include synthesizing one or more drug candidates in the workflow. Optionally, the evaluator module can be configured to evaluate the synthesized molecules as an input to the workflow.

[0054] The method can optionally further include storing in memory any or all of the drug specification, workflow, and drug candidate. Optionally, the memory includes a key-value store of drug candidates generated by the workflow. The memory can be iteratively added to as the workflow is completed, and the results of each step stored in memory can used as inputs to future steps. For example, the memory can be used to avoid repeating the same outputs, and / or to steer any or all of the reasoner module, molecule design module, or evaluator module towards compounds stored in memory that meet the specification, or away from compounds stored in memory that fail to meet the drug specification stored in memory.

[0055] With reference to FIG. 2A, implementations of the present disclosure include systems for drug discovery. The system shown in FIG. 2A can implement an automated, agentic framework for improving the drug discovery process by combining large language models and computational tools.

[0056] The example system 200 includes a memory 204, a molecule design module 220, a reasoner module 210, and an evaluator module 230. As used herein, a “module” can be implemented as a software component executed by one or more processors that implements the methods described herein for that module (e.g., the steps, decision logic, and data transformations recited herein). Modules can be implemented as classes, functions, threads, services, containers, or microservices; they need not be physically separate and may beMCC Ref. No: 103362-063PV1T2025-038 combined, split, or co-resident in a single program binary sharing memory. In some implementations, a module is realized as circuitry or specialized hardware configured to perform some or all of the methods described herein.

[0057] The reasoner module 210 plans the system’s actions and directs the system modules to conduct drug discovery. The molecule design module 220 executes the reasoner module’s actions using state-of-the-art computational tools configured for molecule design. The evaluator module 230 assesses candidate molecules. The memory 204 keeps the information produced along the drug discovery process stored during the operation of the system according to the methods described herein. The system 200 shown in FIG. 2A enables an abstraction from conventional drug discovery, enhanced by computational tools and generative AI. Given a target of interest and property specifications on its potential drugs (e.g., at least 5 molecules, binding affinities better than -7, drug-likeness better than 0.5), implementations of the present disclosure can be used to produce a diverse set of high-quality molecules that satisfy these specifications and can be considered as potential drug candidates for the target, and to enable the discovery of novel therapeutics.

[0058] The reasoner module 210 is configured to receive an input 202, which can optionally be in the form of a user “prompt” as shown in the input 202 of FIG. 2A. The input 202 can optionally include specifications for the number of molecules to generate and any parameters that the molecules must have to meet the specification. The reasoner module 210 can be configured to receive the drug specification from the input 202 and configure the molecule design module 220 based on the drug specification.

[0059] The reasoner module 210 can be configured to perform the decision-making steps of the system. Using the information in the memory 204 (e.g., molecules under currentMCC Ref. No: 103362-063PV1T2025-038 consideration and their property profiles), The reasoner module 210 conducts reasoning and strategically plans the next actions that the system should take, leveraging the pre-trained knowledge and reasoning capabilities of LLMs. As a non-limiting example, the reasoner module 210 can be configured to perform three action types: (1) Generate to generate new molecules: (2) Optimize to optimize existing molecules; and (3) Screen to process the current molecules. These actions correspond to several key steps in hit identification (via generative Al) and lead optimization in pre-clinical drug discovery. Therefore, the reasoner module 210 can guide the system through an iterative process of molecular design, ensuring that each decision aligns with the overall requirements of drug discovery while balancing all required properties.

[0060] It should be understood that the reasoner module 210 can be configured to perform the steps described with reference to FIG. 1 and FIG. 2A in different orders. For example, the reasoner module 210 can use different tools of the molecule design module 220 in sequence or in parallel. An example flowchart of operations that can be performed is shown in FIG. 4, and described in the example herein.

[0061] The molecule design module 220 can be configured to execute the steps selected by the reasoner module 210. The molecule design module can include any number or combination of state-of-the-art computational tools tailored to different actions. In an example implementation, the molecule design module 220 includes generative models for structure-based drug design to implement the Generate action, utilizing methods such as Pocket2Mol, as a non¬ limiting example; generative models for molecular refinement to implement the Optimize action, enhancing drug-like properties and optimizing molecular structures; and at least one processor to implement the Screen action, allowing complex and logic screening, organizing and managing molecules, and identifying the most promising ones.MCC Ref. No: 103362-063PV1T2025-038

[0062] The molecule design module 220 can include any number of general-purpose computing devices (e.g., the computing device 300 of FIG. 3) as well as any number of specialized hardware modules for drug discovery. For example, the molecule design module 220 can be implemented using any or all of the components described herein for generating, screening, and optimizing molecules. For example, the molecule design module 220 can be configured to generate molecules configured to fit in a specified protein pocket. As yet another example, in some implementations, the molecule design module 220 includes a lead optimizer. For example, the lead optimizer can be configured to optimize a molecule generated by the molecule design module by modifying the molecule. Alternatively or additionally, the molecule design module 220 can optionally include a synthesis planner configured to generate a synthetic route of drug synthesis,

[0063] As shown in FI G. 2B, implementations of the present disclosure (for example, the implementation referred to herein as LIDDiA) can include automated laboratory validation of the molecules designed. For example, the evaluator module 230 described herein can implement retrosynthesis planning tools and the systems described herein can be operatively coupled to a self-driving laboratory 250 including automated laboratory equipment. Optionally, the self¬ driving laboratory 250 can be configured to enable in-vitro evaluation of molecules designed by the systems and methods described herein. The self-driving laboratory 250 can optionally be configured to generate experimental outputs based on the in-vitro evaluation of the generated molecules, which can be stored in the memory 204. This integration can optionally provide valuable real- world feedback to the reasoner module 210 by the memory’ 204 effectively enabling the systems and methods described herein to control seif-driving laboratories to further accelerate pre-clinical drug discovery.MCC Ref. No: 103362-063PV1T2025-038

[0064] The example implementation can improve over conventional systems because the molecule design module 220 leverages generative models for both hit identification (Generate) and lead optimization (Optimize). Unlike conventional drug discovery, which relies on searching and modifying existing molecular databases, Molecule design module 220 enables de novo molecular generation, expanding the chemical space beyond known molecules. This approach increases the likelihood of discovering novel, diverse, and more effective drug candidates. Moreover, by automating lead optimization through generative models, Molecule design module reduces human bias stemming from individual expertise levels and limited chemical knowledge. This ensures a more systematic, data-driven approach to improving molecular properties. By integrating generative tools, the example implementation can operate autonomously, designing superior drug candidates more efficiently than traditional human-driven search methods. This automation enhances cost- effectiveness and accelerates drug discovery, making implementations of the present disclosure powerful tools for the drug discovery process.

[0065] The system performs in silico assessments over molecules using the evaluator module 230. The evaluator module 230 assesses an array of molecule properties essential to successful drug candidates, including target binding affinity, drug-likeness, synthetic accessibility, Lipinski's rule, novelty, and diversity. For different properties, Evaluator uses the appropriate computational tools to conduct the evaluation. Evaluation results are systematically stored in the memory 204, and subsequently utilized by reasoner module 210 to refine decisionmaking and guide the next steps in the drug discovery process. Once evaluator module 230 identifies molecules satisfying all user requirements, it causes the system to terminate the search and return the most promising candidates. Thus, the evaluator module serves as a critical quality control mechanism, systematically steering the system toward identifying optimal drugMCC Ref. No: 103362-063PV1T2025-038 candidates while minimizing the exploration of suboptimal chemical spaces. In addition, the evaluator module 230 is designed to allow additional tools to be added or updated to support new property requirements.

[0066] The evaluator module 230 can be configured to evaluate the outputs of the molecule design module 220 to determine whether they meet the drug specification of the input 202. For example, the molecule design module 220 can be configured to perform docking simulation, property prediction, and to evaluate the diversity and novelty of the outputs of the molecule design module. For example, the docking simulator can be configured to predict a preferred orientation or interaction strength between a protein binding site and a small molecule. As another example, the property predictor can be configured to determine any or all of the absorption, distribution, metabolism, excretion, or toxicity properties of a molecule generated by the molecule design module.

[0067] Outputs of the molecule design module 220 that meet the drug specification of the input 202 can be output by the system as results 240. Results 240 can include any number of candidate molecules. As shown in Fig. 2A, the results 240 can include information about the molecules (e.g., information determined by the molecule design module and / or stored in the memory 204).

[0068] The memory 204 can be configured to record data from the molecule design module and the reasoner module 210. Both the molecule design module and reasoner module can optionally be configured to access the memory (e.g., to retrieve the molecules that have already been generated, the original user input, and / or any information about prior actions taken by any of the modules described herein). Optionally, any or all of the information produced throughout the operation of the system can be stored in the memory 204. This includes information providedMCC Ref. No: 103362-063PV1T2025-038 by users via prompts, such as protein target structures, property requirements, and reference molecules (e.g., known drugs). More information will be dynamically generated as the system performs the methods described herein to progress through the drug discovery' process, including the trajectories of actions that LIDDIA has taken (planned by reasoner module 210), molecules generated from prior actions, and their properties. The information in Memory is aggregated and provided to the reasoner module 210 to facilitate its planning. Memory' 204 can be dynamically changing and continuously updated to enable the reasoner module 210 to be well informed by prior knowledge and newly generated data, and thus better reflect and refine its strategies, enhancing the efficiency and effectiveness of automated drug discovery.

[0069] Implementations of the present disclosure can include additional modules as part of the molecule design module 220 and / or evaluator module 230. Additional example modules that can be used in implementations of the present disclosure include pathway retrieval, structure retrieval, sequence retrieval, description (text) retrieval, multi-step retrosynthesis (e.g. RetroStart, MEEA), single-step retrosynthesis (e.g. G2Retro, RLSynC, RSMILES), yield prediction (e.g. YieldGT), molecule generation based on protein pockets, molecule generation based on ligands (e.g. ShapeMol), molecule optimization (e.g. Modof, MolDQN), property prediction (e.g. UniMol, PubChem lookup).[00701 It should be appreciated that the logical operations described herein with respect to the various figures may be implemented (1) as a sequence of computer implemented acts or program modules (i.e., software) running on a computing device (e.g., the computing device described in FIG. 3), (2) as interconnected machine logic circuits or circuit modules (i.e., hardware) within the computing device and / or (3) a combination of software and hardware of the computing device. Thus, the logical operations discussed herein are not limited to any specificMCC Ref. No: 103362-063PV1T2025-038 combination of hardware and software. The implementation is a matter of choice dependent on the performance and other requirements of the computing device. Accordingly, the logical operations described herein are referred to variously as operations, structural devices, acts, or modules. These operations, structural devices, acts and modules may be implemented in software, in firmware, in special purpose digital logic, and any combination thereof. It should also be appreciated that more or fewer operations may be performed than shown in the figures and described herein. These operations may also be performed in a different order than those described herein.

[0071] Referring to FIG. 3, an example computing device 300 upon which the methods described herein may be implemented is illustrated. It should be understood that the example computing device 300 is only one example of a suitable computing environment upon which the methods described herein may be implemented. Optionally, the computing device 300 can be a well-known computing system including, but not limited to, personal computers, servers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, network personal computers (PCs), minicomputers, mainframe computers, embedded systems, and / or distributed computing environments including a plurality of any of the above systems or devices. Distributed computing environments enable remote computing devices, which are connected to a communication network or other data transmission medium, to perform various tasks. In the distributed computing environment, the program modules, applications, and other data may be stored on local and / or remote computer storage media.

[0072] In its most basic configuration, computing device 300 typically includes at least one processing unit 306 and system memory 304. Depending on the exact configuration and type of computing device, system memory 304 may be volatile (such as random access memoryMCC Ref. No: 103362-063PV1T2025-038 (RAM)), non-volatile (such as read-only memory (ROM), flash memory, etc.), or some combination of the two. This most basic configuration is illustrated in Fig. 3 by dashed line 302. The processing unit 306 may be a standard programmable processor that performs arithmetic and logic operations necessary' for operation of the computing device 300. The computing device 300 may also include a bus or other communication mechanism for communicating information among various components of the computing device 300.

[0073] Computing device 300 may have additional features / functionality. For example, computing device 300 may include additional storage such as removable storage 308 and non-removable storage 310 including, but not limited to, magnetic or optical disks or tapes. Computing device 300 may also contain network connection(s) 316 that allow the device to communicate with other devices. Computing device 300 may also have input device(s) 314 such as a keyboard, mouse, touch screen, etc. Output device(s) 312 such as a display, speakers, printer, etc. may also be included. The additional devices may be connected to the bus in order to facilitate communication of data among the components of the computing device 300. All these devices are well known in the art and need not be discussed at length here.

[0074] The processing unit 306 may be configured to execute program code encoded in tangible, computer-readable media. Tangible, computer-readable media refers to any media that is capable of providing data that causes the computing device 300 (i.e., a machine) to operate in a particular fashion. Various computer-readable media may be utilized to provide instructions to the processing unit 306 for execution. Example tangible, computer-readable media may include, but is not limited to, volatile media, non-volatile media, removable media and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. SystemMCC Ref. No: 103362-063PV1T2025-038 memory 304, removable storage 308, and non-removable storage 310 are all examples of tangible, computer storage media. Example tangible, computer-readable recording media include, but are not limited to, an integrated circuit (e.g., field-programmable gate array or application-specific IC), a hard disk, an optical disk, a magneto-optical disk, a floppy disk, a magnetic tape, a holographic storage medium, a solid-state device, RAM, ROM, electrically erasable program read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices,

[0075] In an example implementation, the processing unit 306 may execute program code stored in the system memory 304. For example, the bus may carry data to the system memory 304, from which the processing unit 306 receives and executes instructions. The data received by the system memory 304 may optionally be stored on the removable storage 308 or the non-removable storage 310 before or after execution by the processing unit 306.

[0076] It should be understood that the various techniques described herein may be implemented in connection with hardware or software or, where appropriate, with a combination thereof. Thus, the methods and apparatuses of the presently disclosed subject matter, or certain aspects or portions thereof, may take the form of program code (i.e., instructions) embodied in tangible media, such as floppy diskettes, CD-ROMs, hard drives, or any other machine-readable storage medium wherein, when the program code is loaded into and executed by a machine, such as a computing device, the machine becomes an apparatus for practicing the presently disclosed subject matter. In the case of program code execution on programmable computers, the computing device generally includes a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device,MCC Ref. No: 103362-063PV1T2025-038 and at least one output device. One or more programs may implement or utilize the processes described in connection with the presently disclosed subject matter, e.g., through the use of an application programming interface (API), reusable controls, or the like. Such programs may be implemented in a high level procedural or object-oriented programming language to communicate with a computer system. However, the program(s) can be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language and it may be combined with hardware implementations.

[0077] An example implementation of the present disclosure according to FIG. 2A was configured and tested. The example implementation is referred to herein as LIDDIA. As shown herein, LIDDIA consistently generates high-quality molecules across all key pharmaceutical properties. Notably, it produces the most molecules (97,2%) for each target with QED higher than average, and the most molecules ( 95.8% ) for each target with VNA higher than average, compared to other methods. Most of them ( 97.8% ) are also novel, second only to DiffSMol. With respect to known drugs, a vast majority (88.3%) of LIDDIA's molecules exhibit high synthetic accessibility (SAS) and ( 96.7% ) adhere to Lipinski's Rule of Five (LRF), comparable to general-purpose LLMs such as Claude, GPT-4o, and ol. On the most stringent metric, VNA, LIDDIA significantly outperforms other methods, with 95.8% of its generated molecules binding similarly to or better than existing drugs. In contrast, existing methods struggle to exceed 65% across these key properties. Overall, LIDDIA proves to be a robust and reliable approach, surpassing existing methods in generating high-quality drug candidates that can also have toxicity properties better than or comparable to known drugs.

[0078] Existing methods face substantial challenges to achieve multiple good properties concurrently. For instance, general-purpose LLMs - Claude, GPT4o, ol -mini, and ol-MCC Ref. No: 103362-063PV1T2025-038 exhibit a trade-off between novelty and binding. While they may sometimes achieve impressive VNA, their generations tend to resemble known drugs closely (e.g., <70% novelty). In contrast, Pocket2Mol and DiffSMol demonstrate the opposite: excel at NVT but struggle at VNA. It is possible that LLMs often anchor their generations based on prior knowledge (e.g., known ligands), thus narrowing their explorations. Meanwhile, Pocket2Mol and DiffSMol can generate new binding molecules but not better than existing drugs. Note that LIDDIA does not suffer from such compromise (e.g,, both SAS and VNA >95%),

[0079] Agent Analysis

[0080] LIDDIA action patterns

[0081] FIG. 4 illustrates an example action pattern showing example transition probabilities of actions that LIDDIA takes throughout the drug discovery process across all the targets, from start to finish. LIDDIA incorporates intelligent refinement at every stage. An example strategy of LIDDIA begins with the generation of target-binding molecules (Generate), followed by either optimization to enhance their properties (Optimize), or screening and selection over the generated molecules (Screen), each of which can be performed by the molecule design module 220 of FIG. 2A. Typically, optimization is performed, which is followed by molecule screening over the optimized molecules. Iterative optimization is possible when no viable molecules exist. Similarly, iterative molecule screening is employed when plenty of viable molecules exist but are structurally similar. For instance, LIDDIA may cluster these molecules and subsequently identify the most promising molecules within each cluster. The most common workflow covers Generate, Optimize, and then Screen toward successful outcomes. It should be understood that while these example workflows were most common in LIDDIA, otherMCC Ref. No: 103362-063PV1T2025-038 workflows may be the most common or have different probabilities in other implementations of the present disclosure.

[0082] The Screen function serves as a quality-control mechanism to enable successful outcomes. Successful molecules are only possible after Screen completes screening and selection and identifies such molecules to output. As Generate and Optimize tend to yield more molecules than necessary, allowing abundant opportunities for LIDDIA to succeed, Screen prevents LIDDIA from producing low-quality drug candidates.

[0083] Most of the generated molecules from Generate directly go through subsequent optimization by Optimize. This occurs in approximately 90% of the cases. Among all the generated molecules by Generate, it is typical that none of them satisfies all the property requirements, particularly those properties that are not integrated into the Generate tool designs. Any screening by Screen over such molecules will be futile and wasteful. Instead, LIDDIA intelligently executes Optimize, improving the likelihood of successful molecules out of Screen screening. Meanwhile, LIDDIA can still identify high-quality generated molecules and conducts screening directly over them. This clearly demonstrates the reasoning capability of LIDDIA as an effective tool for drug discovery.

[0084] LIDDIA leverages performance-driven insights to determine a most optimal action. Confronted with low-quality outputs from Generate, LIDDIA selectively pursues optimization to maximize results. Whenever additional molecules need to be considered (e.g., Screen does not identify good candidates), LIDDIA prioritizes optimizing promising molecules stored in Memory rather than generating entirely new ones. This represents a cost-effective, risk-averse strategy, balancing exploration and exploitation by refining known candidates with highMCC Ref. No: 103362-063PV1T2025-038 potential rather than investing computational resources in de novo generation with suboptimal outcomes.

[0085] LIDDIA favors refining a few highly promising candidates as it continuously progresses. FIG. 5 describes both the quality of the molecules produced and the typical actions LIDDIA takes at each step. Compared to the initial pool, the output molecules roughly achieve double the QED and VNA, thus emphasizing the importance of iterative optimization and effective screening strategies. However, diversity among these top candidates is often limited since molecules that satisfy multiple property requirements tend to converge on similar structures. This further highlights the complexity of drug discovery, FIG. 5 also highlights how LIDDIA tends to focus on refining the molecules via Optimize or Screen in later steps, mimicking a typical drug discovery workflow,

[0086] Analysis on LIDDIA action trajectories

[0087] FIGS. 6A-6C present examples of LIDDIA action trajectories to identify promising drug candidates for CDK5, a target for neurological conditions, such as Alzheimer's Dementia, along with examples of trajectories that lead to success (e g., HTR2 / X, FIRAS, DRD2) and fail (e.g., PIK3CA, MET, ADRB2) outcomes, respectively.

[0088] LIDDIA intelligently balances exploration and exploitation to identify promising candidates. In the case of CDK5, where it is highly challenging to identify a good drug candidate, LIDDIA is able to adaptively explore the chemical space via iteratively screening viable molecules and improving any property the molecules fail to meet. As shown in FIG. 6A, LIDDIA starts with Generate (step 1), proceeds with Optimize (step 2), then applies Screen (step 3) but fails to find favorable candidates. In response, LIDDIA refines the failingMCC Ref. No: 103362-063PV1T2025-038 property (step 4) and performs another screening (step 5). This process continues (steps 6, 7, 8) until the agent finally converges to a set of promising candidates.

[0089] This not an isolated case; LIDDIA consistently displays comparable intelligent decision-making behavior on other targets as observed in FIG. 6B and FIG. 6C. Notably, m cases with successful outcomes, LIDDIA methodically refines several properties before screening for candidates (DRD2), or strategically determines which molecules to prioritize and what action to take (HTR2A and HR AS), For instance, in the HRAS case, LIDDIA uses several screenings (steps 2 and 3) to identify viable candidates, optimize them (step 4), and conduct further screenings (step 5 to 8) until it identifies favorable candidates. This highlights a strength of LIDDIA- its capability to adapt to feedback (e g., molecules quality) from its Evaluator, to explore (e.g., via refinement and generation), and to exploit (e.g., screening existing molecules) the chemical space. On cases where LIDDIA yielded suboptimal results, such as PIK3CA, MET, and ADRB2, it still exhaustively performs various actions up to the action limits (e.g., 10 iterations in an example implementation).

[0090] Ablation Study

[0091] The study included additional experiments to test the effectiveness of LIDDIA. First, the study replaced Claude 3.5 Sonnet with DeepSeek-R1 in all componentsrequiring large language models to test LIDDIA's robustness to different backend LLMs.Additionally, the study compared LIDDIA to a simple deterministic loop iterating between LIDDIA's components to analyze the importance of reasoning in LIDDIA. Note that this deterministic loop is similar to LIDDIA but without any LLM. Concretely, the study ran Generate once, followed by a loop of Optimize and Screen for k number of times. The studyMCC Ref. No: 103362-063PV1T2025-038 prioritized properties that fall below requirements when optimizing molecules, and only Screen for high quality molecules. The study set k to 10. Example results are shown in FIG. 7.

[0092] Reasoning benefits successful drug discovery. LIDDIA with reasoning (e.g. Claude 3.5 and DeepSeek -Rl) achieves a much higher target success rate than without (more than 40% absolute difference), indicating the improvement that the systems and methods described herein present over LLMs alone. LIDDIA is robust to different backend LLMs.Comparing LIDDIA with Claude 3.5 and DeepSeek-R1, they both perform similarly (both with 73% TSR), emphasizing that implementations described herein are robust to different backend LLMs. Different LLMs may be used to obtain different performance with implementations of the present disclosure. For example, molecules generated by LIDDIA with DeepSeek-R1 are almost always HQ compared to Claude ( > 99% vs 84% ). However, only 77% satisfy the diversity requirements, in contrast to Claude 3.5 (90%).

[0093] Experimental Settings

[0094] Evaluation Metrics

[0095] Molecule Qualities: were used to evaluate the molecules generated by different methods. Key molecule properties were used tot evaluate the following general properties required for successful drugs: (1) drug- likeness (Bickerton et al., 2012) (QED), (2) Lipinski's Rule of Five (Lipinski et al., 2001) (LRF), (3 ) synthetic accessibility (Ertl and Schuffenhauer, 2009) (SAS), and (4) binding affinities measured by Vina scores (Trott and Olson, 2010) (VNA).

[0096] Novelty: The study measured the novelty (NVT) of a molecule m with respect to a reference set of known drugsoas follows:

[0097] NVT(m; M0) = 1 — max fsimT(m, mi)\MCC Ref. No: 103362-063PV1T2025-038 10098] whereois the reference set of known drugs, m and mi are two molecules, and simT(m, mi) is the Tanimoto similarity of m and mi's Morgan fingerprints (Morgan, 1965). High novelty indicates that new molecules are different from existing drugs, offering new therapeutic opportunities. A molecule m is considered novel if NVT(m) > 0.8.

[0099] High-quality molecules A molecule m is considered as "high quality7" ( HQ ) for a target t, if its properties satisfy QED > QEDf, LRF > LRFt, SAS < SASt, VNA < VNAt, and NVT(m) > 0.8, where the overline and the subscript t indicate the average value from all the known drugs for target t. Such multi-property requirements are typical in drug discovery.Meanwhile, this presents a significant challenge, as LIDDIA must identify molecules with key¬ properties similar to or even better than existing drugs but structurally7significantly different from them. The studied dataset includes existing drugs for targets as the gold standard for evaluation purposes.

[0100] Molecule Set Diversity. The example implementation measures the diversity (DVS) of a set of generated molecules M defined as follows,

[0101] DVS(M) = 1 -

[0102] where m, and j are two distinct molecules in AT. High diversity is preferred, as chemically diverse molecules increase the likelihood of identifying successful drug candidates. A set of molecules Af is considered diverse if DVS(Af) > 0.8. This imposes a highly stringent requirement on the diversity of the generated molecules.

[0103] Target Success Rate Target success rate: denoted as TSR, is defined as the percentage of targets for which a method can generate a diverse set of at least 5 high-quality molecules.

[0104] Protein Target DatasetMCC Ref. No: 103362-063PV1T2025-038

[0105] To evaluate LIDDIA, the study manually curated a diverse set of protein targets from OpenTargets (Ochoa et al., 2023) that are strongly associated with major human diseases: cancers, neurological conditions, cardiovascular diseases, infectious diseases, diabetes, and autoimmune diseases. For each of these protein targets, the example implementation identified an experimentally resolved structure with a small-molecule ligand from the RCSB Protein Data Bank (PDB) and extracted the binding pocket according to its ligand's position. To enable a compari on to existing drugs, the study searched ChEMBL for all known drugs targeting the selected proteins. This leads to 30 protein targets with PDB structures, ligands, and existing drugs. These targets will be used as input in the experiments herein. Table 1 presents the distribution of the targets in terms of their disease associations.

[0106] Table 1: Statistics over protein targets.Disease #Targets (%)Cancers 15(50%)Neurological Conditions 8(27%)Cardiovascular Diseases 6(20%)Infectious Diseases 4(13%)Diabetes 3(10%)Autoimmune Diseases 3(10%)

[0107] Some targets are associated with multiple categories.

[0108] Example Implementation Details

[0109] LIDDIA leveraged Claude 3,5 Sonnet (Anthropic, 2024) as the base model for its Reasoner and Evaluator since it achieves state-of-the-art performance in chemistry related tasks. The study designed and fine-tuned prompts to guide Reasoner and Evaluator modules shown in FIG. 2 A, respectively. The example implementation set the maximum number ofMCC Ref. No: 103362-063PV1T2025-038 actions taken to 10 to ensure a concise yet effective drug discovery trajectory. Any number of actions can be taken, however, in various implementations of the present disclosure.

[0110] Molecule design module executes the Generate action using Pocket2Mol. As a structure-based drug design tool, Pocket2Mol can generate molecules using only the target protein structure. This provides LIDDIA with the ability to extend to novel targets without known ligands. For efficiency, Pocket2Mol is set to generate a minimum of 100 molecules using a beam size of 300. The Optimize action is implemented via GraphGA (Jensen, 2019), a popular graph-based genetic algorithm for molecule optimization. In LIDDIA, Optimize can refine molecules on three essential properties: drug-likeness (QED), synthetic accessibility (SAS), and target binding affinity (VNA). However, the actions can be easily expanded to cover additional properties,

[0111] Baselines

[0112] The study compare LIDDIA with two types of baselines: task-specific molecule generation methods, and general-purpose LLMs. For molecule generation methods, the study used Pocket2Mol and DiffSMol. Pocket2Mol is a well-established generative method for structure-based drug design, which uses binding pocket structures as input. DiffSMol (Chen et al., 2025), on the other hand, is a state-of-the-art generative method for ligand- based drug design, requiring a binding ligand. These two methods represent distinct approaches m computational drug design, using different information to generate potential drug candidates. Notably, Pocket2Mol is used by LIDDIA in Generate actions, capitalizing on the popularity of SBDD and its ability to generate molecules without reference ligands.

[0113] For general-purpose LLMs, the example implementation used GPT4o (OpenAI et al., 2024), o1 (OpenAI, 2024a), o1-mini (OpenAI, 2024b), and Claude 3.5 SonnetMCC Ref. No: 103362-063PV1T2025-038 (Anthropic, 2024). GPT-4o and Claude 3.5 Sonnet are representative state-of-the-art language models; ol and ol-mini are specifically tailored towards scientific reasoning during their training. The study evaluated all four of these models as baselines to provide a comprehensive understanding of the performance of state-of-the-art LLMs The LLMs described herein are nonlimiting examples, and the present disclosure contemplates the use of any suitable LLM to implement the reasoner modules described herein.

[0114] Experimental Results

[0115] Table 2 presents the performance of different methods, including their success rates and the qualities of their generated molecules. LIDDIA successfully generates novel, diverse, and high-quality molecules as potential drug candidates for 73.3% of targets (TSR), significantly outperforming existing methods. Pocket2Mol, the second-best method, achieves only a 23.3% success rate, while most proprietary LLMs fail entirely. Crucially, LIDDIA excels in simultaneously optimizing all five key pharmaceutical properties - QED, LRF, SAS, VNA, and NVT- on average, 85% of the generated molecules for each target are of high quality (HQ). In contrast, GPT-4o achieves only 35% in HQ, lagging nearly 50 percentage points behind LIDDIA, while all other methods perform even worse. In terms of the qualities of the generated molecules, LIDDIA produces molecules of comparable or superior quality to the limited outputs of other methods. These results highlight LIDDIA as a highly effective and reliable framework for accelerating drug discovery, consistently outperforming existing methods in both success rate and molecule qualities. The study also compared LIDDIA with more recent state-of-the-art methods.

[0116] Table 2: Performance comparison between the baseline methods and LIDDI A. The following terms are used in Table 2. %m / t: average percentage of molecules perMCC Ref. No: 103362-063PV1T2025-038 target; #m / t: average number of molecules per target; Generated: initially generated molecules;Valid: generated molecules that are also valid; overlinet: the average value of corresponding property in the known drugs for the target t. % t: average percentage of targets among all targets; #t: average number of targets; N D E ED VS: at least 5 molecules are generated and the set is diverse; N > 5& HQ: at least 5 molecules are generated and they are of high quality; t / -L indicates higher / lower values are better.GP GPPocket! Pocket! DiffSM DiffSM T- Clau Clan Section Metric T- Mol % Mol # OL % OL# 4o4o# de % de# %Generati - 100 - 00 - 5 nitial 1 5ed98.7 4.9 initial Valid 100 100 99.9 99.9 97.3 4.9generated QED >53.4 53.4 60 60 88.2 4.4mol ecu QED t 96.7 4.8 lesgenerated LRF >99.7 99.7 72.1 72.1 95.9 4.8molecu LRF t 98.7 4.9 lesgenerated SAS <77.4 77.4 7.5 7.5 90.7 4.5molecu SASj 92.7 4.6 lesgenerated VNA <15.3 15.3 24.7 24.7 59.2 3molecu VNA_t 63.3 3.2 lesgenerated NVT >87.6 87.6 98.2 98.2 68.3 3.4molecu 0.8 46.9 2.4 lesgeneratedHQ 6.4 6.4 0.7 0.7 35 1.7molecu 30.3 1.5lesMCC Ref. No: 103362-063PV1T2025-038 amongDVS >all 100 30 100 30 90 970.8 30 9 targetsamongN>5 &all 100 30 100 30 77.7 23 27.7 8DVStargetsamongN>5 &all 27.7 8 3.3 1 10 3HQ 23.3 7 targetsamongDVS &all 23.3 7 10 3 33.3 10HQ 10 3 targetsamongall TSR 23.3 7 0 0 6.7 2 6.7 2targetsol- ol- LIDDIA LIDDIA Section Metric mini ol % ol #mini # % # %initial Generated 5 5 - 24.5 initial Valid 91.3 4.6 95.3 4.8 100 24.5 generatedQED > QED t 90.1 4.5 88.3 4.4 97.2 21.8 moleculesgeneratedLRF > LRF t 90.7 4.5 95.3 4.8 96.7 21.8 moleculesgeneratedSAS < SAS_t 81.4 4.1 92.6 4.6 88.3 17.4 moleculesgeneratedVNA < VNA_t 47.9 2.3 34.6 1.8 95.8 21.2 moleculesgeneratedNVT>0.8 64.1 3.2 55.9 2.8 97.8 22.4 moleculesgeneratedHQ 28.2 1.4 20.7 1 84 14.5 moleculesamong allDVS > 0.8 67.7 20 70 21 97.7 29 targetsamong allN>5 & DVS 43.3 13 57.7 17 90 27 targetsamong allN>5 & HQ 0 0 3.3 1 73.3 22 targetsamong allDVS & HQ 33.3 10 20 6 90 97targetsMCC Ref. No: 103362-063PV1T2025-038 among allTSR 0 0 0 0 73.3 22targetsMetric Pocket2Mol DiffSAIOL Claude GPT-4o Ol-mini O1 LIDDIA NVT t 0.87 0.89 0.77 0.82 0.79 0.8 0.86 QED ) 0.51 0.55 0.78 0.74 0.75 0.77 0.69 LRF i 4 3.43 4 3.99 3.85 4 3.93 SAS ) 2.46 6.15 2.3 2, 16 2.02 2.03 2.62 VNA i, -4.74 -4.23 -6.69 -6.56 -6.31 -5.97 -7.17DVS t 0.88 0.89 0.76 0.84 0.79 0.8 0.82

[0117] Case Study on AR / NR3C4

[0118] The study tasked LIDDIA with discovering new potential drug therapies targeting androgen receptor (AR / NR3C4), a hormone-driven transcription factor protein that plays a key role in both prostate and breast cancers (Tan et al., 2015; Giovannelli et al., 2018). LIDDIA identifies one molecule (named NL-1) with better QED, VNA, and SAS than the ligand and at least one approved drug (e.g., Flutamide) for the respective targets. They are illustrated in FIGS. 8A-8C, which show NL-1 in FIG. 8A, the ligand in FIG. 8B, and Flutamide in FIG. 8C. NL-1 has several desirable traits, such as zwitterionic (with positive and negative charged atoms on the respective ends)-a trait typical in most biological molecules and drugs. The molecule also passes several computational filters, including PAINS, BRENK, N1H, Lilly, and Lipinski, further highlighting its attractiveness as a drug. In terms of binding, the molecule has — 8.81kcal / mol for VNA, emphasizing that it can bind well to the pocket. FIG. 9A shows that the molecule is buried deep within the pocket, surrounded almost entirely by hydrophobic residues providing many van der Waals contacts. The molecule’s carboxylic acid group also engages in hydrogen bonding at one end of the pocket, further stabilizing the complex and contributing to binding affinity. Encouragingly, further evaluation reveals that NL-1 has several established syntheticMCC Ref. No: 103362-063PV1T2025-038 routes and demonstrates precedent for antagonizing stimulator of interferon genes (STING), thereby useful in the treatment of inflammatory diseases. The desirable traits, combined with its synthetic accessibility and therapeutic precedents, position the molecule as a promising candidate for the androgen receptor.

[0119] FIG. 9B shows a comparison of NL-1 to the known ligand. Both molecules feature a hydrophobic core with polar anchors on either end, but NL-1 is slightly less compact, with fewer fused rings than the known ligand. This reduces conformational rigidity, providing slightly increased flexibility to adapt to the binding pocket. Furthermore, NL-1 anchors itself with a carboxylic acid in place of the known ligand's ketone, enabling more hydrogen bonds.

[0120] Case Study on EGER

[0121] The study included another case study using LIDDIA for discovering new potential drug therapies targeting the Epidermal growth factor receptor 1 (EGFR) protein. EGFR is a transmembrane glycoprotein that plays a pivotal role in many cancers, including breast cancer, esophageal cancer, and lung cancer. Its role in cancer, as well as its accessibility on the cell membrane, has made it a prime therapeutic target. However, cancer cells mutate rapidly and can become resistant to drugs over time, leading to a need for novel drug therapies. The study compared the molecules generated for EGFR by LIDDIA with three approved drugs of EGFR -Olmutinib, Masoprocol, and Gefitinib, which exhibit the best VNA, QED and SAS among all EGFR’s approved drugs, respectively. The study also compared the outputs of LIDDIA with a known ligand for EGFR’s binding pocket. FIGS. 10A, 10B, and 10C present the overall comparison results. Note that most existing methods cannot generate any novel high-quality molecules. This emphasizes the strength of LIDDIA, which can tackle even a challenging target.MCC Ref. No: 103362-063PV1T2025-038

[0122] LIDDIA effectively generates promising novel drug candidates on EGFR. Notably, these molecules surpass the native ligand in both VNA and QED, while displaying comparable overall profiles to approved drugs. Moreover, some molecules as shown in FIG. 10A-10C, columns 1 and 2 are better than Olmutinib and Masoprocol on all metrics (i.e., VNA, QED, and SAS). The two molecules, NL-2 and NL-3 are illustrated in FIG. 11 A. These molecules possess structural features that allow them to bind the pocket well. Notably, NL-2 has hydroxyl groups on both the five- and six-membered rings from the first molecule, which form strong hydrogen bonds with the protein target on opposite sides of the pocket as shown in FIG.12A. NL-3 utilizes a different binding strategy, relying on hydrophobic packing and shape complementarity rather than polar interactions. As shown in FIG. 12B, the fluorine substituent is positioned near the pocket entrance flanked by hydrophobic residues, which serve as favorable van der Waals contacts. Meanwhile, the diarylketone moiety is buried deep within the binding pocket, anchoring the ligand through planar stacking and hydrophobic interactions, despite the absence of direct hydrogen bonding. This finding aligns with previous literature, highlighting the potency of diarylketone for antitumor drugs.

[0123] Medicinal chemists can be used for real-world deployment of LIDDIA. Despite the favorable binding and in silico properties, closer examination reveals some concerning structural features in these molecules. NL-2 contains two enol groups (the -OH near the double bond)-substructures with tautomeric instability and are highly unattractive for drugs. NL-3 contains fulvene, known to be chemically reactive, thermally unstable, sensitive to oxygen, and photosensitive. The diarylketone moiety, despite its favorable binding and potency in antitumor drugs, is known to be phototoxic. Such conflicts (e.g., favorable binding but phototoxic) are typical in drug discovery, highlight areas where the evaluator modules andMCC Ref. No: 103362-063PV1T2025-038 molecule design modules disclosed herein can be improved to further enable the identification of desirable compounds without these drawbacks by further screening and simulating the compounds generated by the molecule design modules.

[0124] Furthermore, no standalone in silico evaluation tools (e.g., computational filters sold under the trade names RDKit and Medchem) can alone detect all the issues presented in these molecules. Implementations of the present disclosure therefore contemplate the use of multiple tools, coordinated by the reasoner modules described herein, to overcome the limitations of individual tools and filters. For example, several filters (e.g., PAINS, BRENK, NIH) cannot capture the problematic features in NL-2, highlighting the limitation of existing tools. Lilly rules are able to identify the enol groups, but do not raise any alerts for NL-3.

[0125] To overcome these limitations, the present disclosure contemplates human-in-the-loop validation of the outputs of the systems and methods described herein, the use of and integration with more sophisticated in silico tools, and / or wet-lab validation of generated molecules. Human experts in the loop can perform validation for a nuanced and comprehensive evaluation of the molecules. Meanwhile, existing in silico tools, particularly in evaluation, have room for improvement. The integration of more and better tools as they are developed allows implementations of the present disclosure to generate more and better high-quality molecules. Ultimately, in vitro and in vivo in a laboratory can be used to translate LIDDIA's performance to real-world impacts.

[0126] Additional Example Results

[0127] LIDDIA was tested with two more recent task-specific molecule generation methods, TargetDiff and DecompDiff. The study observed similar results to those shown in Table 2. First, both LIDDIA with DeepSeek-Rl and Claude significantly outperformMCC Ref. No: 103362-063PV1T2025-038 both methods, highlighting the effectiveness of LIDDIA. Second, both Target Diff and DecompDiff also struggle to generate new binding molecules better than existing drugs, while LIDDIA does not.

[0128] Toxicity Predictions

[0129] ADMET-AI was used to predict the toxicity properties of LIDDIA's generated molecules. The study used toxicity properties available in ADMET-AI and compared them to drugs in the study’s dataset. The generated molecules described herein are better than or comparabl e to known drugs in terms of their toxicity properties. Note that the agent is specifically designed to generate high-quality molecules, not "safe" molecules. But the example results have promising safety profiles. This presents interesting findings that (1) the current design of LIDDIA can generate both high-quality and "safe" molecules and (2) the ability of LIDDIA to be improved through the use of automated tools to perform safety analysis.

[0130] Failure Analysis

[0131] LIDDIA yielded suboptimal results on some targets, but on all targets tested LIDDIA can generate at least one high-quality molecule, including on PIK3CA, MET, and ADRB2. However, they were considered suboptimal since: (1) the number of high-quality molecules is not sufficient (i.e., less than 5), or (2) the molecules are not diverse enough. The study hypothesized that existing tools are struggling because of the structure of the pockets. For instance, the pocket may only allow a few specific scaffolds to bind, making it extremely difficult for existing tools to generate many and diverse high-quality molecules.

[0132] FIGS. 10A-10C illustrate a case study for EGFR where each figure compares molecules generated by LIDDIA to three drugs and one binding ligand of EGFR onMCC Ref. No: 103362-063PV1T2025-038 VNA, SAS, and QED, respectively. Shaded areas indicate that the LIDDIA molecule outperforms the reference molecule on respective metrics.

[0133] FIGS. 11A-11C illustrate a case study on EGFR where FIG. 11 A illustrates LIDDIA's generated molecules (NL-2 and NL-3). FIG. 11B illustrates a known ligand for EGFR. FIG. 11 C illustrates examples of known approved drugs for EGFR. NL-2 has two enol groups and NL-3 has a fulvene, both of which are problematic as drug candidates.

[0134] Discussion

[0135] The example implementation disclosed herein is an effective agent for autonomous drug discovery. The study disclosed herein demonstrates its performance across many major therapeutical targets and revealing several key insights on its success. Furthermore, the study investigated the generated molecules on the highly critical target EGFR and show their potential as drug candidates.

[0136] The present disclosure contemplates the use of wet-lab validation to extend implementations of the present disclosure to more complicated automatic drug discovery pipelines. Moreover, the present disclosure contemplates the use of additional metrics beyond the metrics described herein. Moreover, the present disclosure contemplates that more complicated workflows can be run for any number of iterations to improve performance.

[0137] Existing LLMs generally fail to incorporate appropriate tools and combinations of tools configured for drug discovery. In particular, existing LLMs have limited intrinsic chemistry knowledge, and internet sources fail to overcome that limited chemistry knowledge when working to develop novel compounds that are not published. Accordingly, implementations of the present disclosure improve existing LLMs, as well as existing agentic loops that use LLMs to access the web, by providing domain-specific tools to LLMs to leverageMCC Ref. No: 103362-063PV1T2025-038 both the analytical capabilities of domain specific tools, as well as the general evaluation, reasoner, and coordination capabilities of LLMs, enabling an automated workflow of designing and evaluating compounds grounded by well-established computational tools for novel structurebased drug discovery. Thus, implementations of the present disclosure provide improvements to systems and methods for drug discovery, over conventional methods that used only computational tools or only LLMs with search tools.

[0138] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

MCC Ref. No: 103362-063PV1T2025-038 WHAT IS CLAIMED:

1. A computer- implemented method comprising:receiving a drug specification;selecting, by a reasoner module, a workflow based on the drug specification, wherein the reasoner module comprises a large language model configured to operate in inference mode responsive to the drug specification; executing, by a molecule design module, the workflow to produce a workflow output;evaluating, by an evaluator module, the workflow output;iteratively repeating steps of executing and evaluating until the drug specification is met by the workflow output; andoutputting the workflow output, wherein the workflow output comprises a molecule based on the drug specification.

2. The computer-implemented method of claim 1, wherein the molecule design module comprises a plurality of components.

3. The computer-implemented method of claim 1 or claim 2, wherein selecting the workflow comprises selecting at least one component of the molecule design module.

4. The computer-implemented method of any one of claims 1-3, wherein the molecule design module comprises a hardware module.MCC Ref. No: 103362-063PV1T2025-038 5. The computer-implemented method of any one of claims 1-4, further comprising storing in a memory’ at least one of the drug specification, workflow, and drug candidate.

6. The computer-implemented method of claim 5, wherein the memory comprises a key- value store of drug candidates generated by the workflow.

7. The computer-implemented method of any one of claims 1-6, wherein the method further comprises receiving an output of the molecule design module by the reasoner, and generating a second workflow by the reasoner module.

8. The computer-implemented method of any one of claims 1-7, wherein the evaluator module comprises at least one of a docking simulator or a property predictor.

9. A system for drug discovery, the system comprising:a molecule design module;a reasoner module comprising a large language model, wherein the reasoner module is configured to receive a drug specification and configure the molecule design module based on the drug specification; andan evaluator module configured to evaluate an output of the molecule design module and output a molecule that satisfies the drug specification.

10. The system of claim 9, wherein the molecule design module comprises a molecule generator configured to generate a molecule based on an instruction from the reasoner module.MCC Ref. No: 103362-063PV1T2025-038 11. The system of claim 9 or claim 10, wherein the molecule design module is configured to generate molecules configured to fit in a protein pocket.

12. The system of any one of claims 9-11, wherein the evaluator module comprises a docking simulator and a property predictor.

13. The system of claim 12, wherein the property predictor is configured to determine at least one of absorption, distribution, metabolism, excretion, or toxicity properties of a molecule generated by the molecule design module.

14. The system of claim 12, wherein the docking simulator is configured to predict a preferred orientation or interaction strength between a protein binding site and a small molecule.

15. The system of claim 9, wherein the reasoner module comprises a first computing device configured to operate the large language model in inference mode, and wherein the molecule design module comprises a second computing device in operable communication with the first computing device, and wherein the evaluator module comprises a third computing device in operative communication with the first and second computing devices.

16. The system of any one of claims 9-15, wherein the molecule design module comprises a lead optimizer.

17. The system of claim 16, wherein the lead optimizer is configured to optimize a molecule generated by the molecule design module by modifying the molecule.MCC Ref. No: 103362-063PV1T2025-03818. The system of any one of claims 9-17, wherein the molecule design module further comprises a synthesis planner.

19. The system of claim 18, wherein the synthesis planner is configured to generate a synthetic route of drug synthesis.

20. The system of any one of claims 9-19, wherein the molecule design module comprises a specialized hardware module configured for drug synthesis, drug discovery, or both drug synthesis and discovery.