Large language model for unifying text and point cloud molecular input

Through the transformer model architecture and point cloud encoder processing 3D molecular data, the limitations of existing models when generating complex 3D molecular structures are solved, and efficient and excellent quality spatial molecular generation is achieved.

CN120068941APending Publication Date: 2025-05-30INSILICO MEDICINE IP LTD +1
View PDF 16 Cites 0 Cited by

Patent Information

Application Number
CN202411647496.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-08
Filing Date
2024-11-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing models have limitations in generating and predicting complex 3D molecular structures, especially in dealing with spatial features and precise arrangement of atoms, and relying on chemical string representations to have problems with excessive length and insufficient atomic connectivity information.

Method used

The transformer model architecture is adopted, combining point cloud encoder and large language models, and 3D molecular data and text input are processed to generate molecular data in line representation format, eliminating the dependence on the reconstruction of molecular maps of external software.

Benefits of technology

It realizes the generation of physically reasonable chemical 3D structures without relying on external software, improves the efficiency and quality of spatial molecular generation tasks, and is better than the competition diffusion model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005140032250000172
    Figure BDA0005140032250000172
  • Figure BDA0005140032250000173
    Figure BDA0005140032250000173
  • Figure BDA0005140032250000231
    Figure BDA0005140032250000231
Patent Text Reader

Abstract

The invention relates to a large language model for unified text and point cloud molecular input. A converter model architecture is described. The transformer model architecture includes a point cloud input module, a text input module, a point cloud encoder module operatively coupled with the point cloud input module, a large language model module operatively coupled with the text input module and the point cloud encoder module and configured to receive data therefrom, and a text output module operatively coupled with the large language model module. The text output module is configured to output the molecular data in a row representation format.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This patent application claims priority to U.S. Provisional Application No. 63 / 604,657, filed on November 30, 2023, the entire disclosure of which is incorporated herein by specific reference. Technical Field

[0003] This disclosure relates to computer-implemented protocols for generating data related to molecular properties. Background Art

[0004] Previously, computer-based models have been developed to generate a wide variety of physically reasonable models of three-dimensional (3D) structures. Although these computer-based models have proven sufficient for some tasks, there is still a need to improve computer-based models for describing and predicting more complex 3D structures.

[0005] An artificial neural network (ANN) is a computational system inspired by the biological neural networks that make up animal brains. ANNs are based on a series of interconnected units or nodes called artificial neurons, which loosely mimic the neurons in a biological brain. Each connection, like a synapse in a biological brain, can transmit a signal to other neurons. Artificial neurons receive signals, process them, and can signal neurons to which they are connected. The "signal" at a connection is a real number, and the output of each neuron is computed by some non-linear function of the sum of its inputs. The connections are called edges. Neurons and edges typically have weights that are adjusted as learning proceeds. The weights increase or decrease the strength of the signal at the connection. Neurons can have a threshold such that a signal is sent only if the aggregated signal crosses that threshold. Typically, neurons are grouped into layers. Different layers may perform different transformations on their inputs. Signals travel from the first layer (input layer) to the last layer (output layer), possibly after passing through these layers multiple times.

[0006] A Deep Neural Network (DNN) is an ANN with one or more hidden layers. Due to their complex structure and large number of trainable parameters, these networks enable more efficient problem-solving. Autoencoders are a subset of DNNs that learn hidden representations of objects. The objects can be objects in different mathematical representations, such as strings, graphs, or pictures. An autoencoder consists of two parts - an encoder and a decoder. The encoder is an encoding function that maps an object to a point (e.g., a latent point) in a numerical space with a specified dimension. This numerical space is called the latent space. The decoder is a decoding function that maps a point in the latent space to an object in the object space. For training, these networks use a reconstruction loss, which is a function that penalizes the model for the difference between the input (encoder input) and output (decoder output) representations of the object.

[0007] Generative models (GMs) are a subclass of DNNs that enable the generation of objects. Different from standard DNNs that predict the attributes of objects, these networks are trained in such a way that in the future, new objects can be generated without input data. These models learn the distribution of objects (e.g., distribution learning) and then try to generate samples from this distribution.

[0008] Distribution learning generative models default to generating random molecules. However, sometimes people hope to generate objects that satisfy given properties. This formulation of the problem is called conditional generation.

[0009] Transformer models are a subset of DNNs that transform an input sequence into an output sequence. These models can determine the context of the input sequence data to generate new data in the output. The transformer model architecture generally includes an input sequence encoder and a decoder. The input sequence encoder outputs a tokenized (e.g., numerical encoding matrix) representation of the input sequence. The decoder receives the tokenized representation of the input sequence and iteratively generates the output. In the encoder, the multi-head self-attention mechanism enables the transformer model to associate each word in the tokenized input sequence with other tokens / words and focus on different tokens / parts of the input sequence when processing each token by amplifying the signals of key tokens / parts. The output of the decoder can be transformed into a sequence representing the prediction. Summary of the Invention

[0010] This document describes a transformer model architecture. In one embodiment, the transformer model architecture includes a point cloud input module and a text input module. The transformer model architecture further includes a point cloud encoder module operatively coupled to the point cloud input module, and a large language model module that is operatively coupled to the text input module and the point cloud encoder module and is configured to receive data therefrom. A text output module is operatively coupled to the large language model module and is configured to output molecular data in a row notation format. The large language model module and the point cloud encoder module can be combined into a single model.

[0011] In some embodiments, the point cloud encoder can aggregate spatial position information from the point cloud and includes a graph neural network configured to process the relative distances between points of the point cloud.

[0012] In some embodiments, the transformer model architecture can include multiple graph neural network layers, where each graph neural network layer aggregates information from connected nodes and edges and processes global information from the entire graph.

[0013] In some embodiments, the graph neural network can include an attention mechanism to compute attention biases related to the relative positions and edge features between points of the point cloud and update the point embeddings based on the computed attention biases.

[0014] In some embodiments, the point cloud encoder can be trained in an unsupervised manner using 3D molecular data including the point cloud received via the point cloud input module.

[0015] In some embodiments, for each point of the point cloud, the 3D molecular data can include spatial position data and data representing one or more molecular features. The one or more molecular features can include one or more of the following: atomic symbol, atomic charge, atomic name, or the corresponding amino acid name.

[0016] In some embodiments, the 3D molecular data can include data in at least one of the following chemical language formats: Simplified Molecular-Input Line Entry System (SMILES) format, Self-referencing Embedded Strings (SELFIES) format, or XYZ format.

[0017] In some embodiments, 3D molecular data can represent large ligand or protein pocket structures and can be downsampled based on one or more priority points of a point cloud. One or more priority points of the point cloud can include at least one of the following: a ligand, an alpha carbon (C-α) atom, or a terminal atom of a protein amino acid.

[0018] In some embodiments, a point cloud encoder can be configured to infer a point description based on one or more points in the neighborhood of a point; and determine the three-dimensional (3D) position of the point relative to one or more points in the neighborhood of the point. In some examples, the point and one or more points in the neighborhood of the point may not be directly edge-connected.

[0019] In one embodiment, a transformer model is provided that has a point cloud input module, a text input module, a point cloud encoder module operatively coupled to the point cloud input module, a large language model module operatively coupled to the text input module and the point cloud encoder module and configured to receive data therefrom, and a text output module configured to output molecular data in a row notation format. 3D molecular data including a point cloud is fed into the point cloud encoder module in an unsupervised manner. The point cloud encoder module infers a point description based on one or more points in the neighborhood of a point. The 3D position of the point is determined relative to one or more points in the neighborhood of the point. At least some of the one or more points are masked or blurred. At least some of the masked or blurred point features are predicted by a pre-trained point cloud encoder module. Random points are sampled, and the distances between the random points are predicted by the pre-trained point cloud encoder module. In some embodiments, a masked or blurred loss value can be determined based on the prediction of the masked or blurred point features. A distance loss value can be determined based on the prediction of the distances between the random points, and a weighted sum of the masked or blurred loss and the distance loss can be minimized based on the masked loss and the distance loss values, respectively.

[0020] In some embodiments, the point cloud can be encoded into one or more point embeddings, and embeddings of an input text sequence can be prepared. The one or more point embeddings and the input text sequence embeddings can be combined to obtain a fused input, and the fused input can be fed into the point cloud encoder module.

[0021] In one embodiment, shape-conditioned generation includes training a transformer model to recover a molecule from a blurred region of the 3D space in which the molecule of a molecular point cloud is located, where some parts of the molecular point cloud are blurred while selected parts of the molecule are unchanged and not blurred. 3D molecular data including a point cloud is input via the point cloud input module. Text representing the desired molecule is input via the text input module. The 3D molecular data and the text are processed using the trained point cloud encoder module, and text representing the molecule is output in a row notation.

[0022] In one embodiment, the linker design generation includes training a transformer model to recover a removed portion of a molecule when the transformer model architecture does not receive a spatial description of the removed portion of the molecule. Data representing one or more molecules is input into the trained transformer model. The trained transformer model is used to generate one or more options for a linker portion of the molecule to replace the removed portion of the molecule, and for the linker portion of the molecule, one or more molecules with the generated one or more options are obtained.

[0023] In one embodiment, a computer system includes a trained transformer model and one or more non-transitory computer-readable media that store instructions that, when executed by one or more processors, cause the computer system to perform one or more of the above operations, including, for example, operating the transformer model architecture or its pre-training or its implementation to perform shape-conditioned generation or linker design generation for at least one molecule.

[0024] The foregoing summary is illustrative only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, additional aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above and following information, as well as other features of the present disclosure, will be more fully apparent from the following description and the appended claims, in conjunction with the drawings. It is to be understood that these drawings depict only several embodiments in accordance with the disclosure and are not to be considered as limiting its scope, and the present disclosure will be described more specifically and in detail by using the drawings.

[0026] Figure 1 An overview of a spatial molecule generation task is illustrated according to an embodiment.

[0027] Figure 2A and Figure 2B A block diagram of a point cloud transformer model architecture based on functional features and modular components is illustrated according to an embodiment.

[0028] Figure 3 A block diagram of a point cloud encoder module of a point cloud transformer model is illustrated according to an embodiment.

[0029] Figure 4 A process of point embedding calculation is illustrated according to an embodiment.

[0030] Figure 5 A schematic diagram of an edge embedding is illustrated according to an embodiment.

[0031] Figure 6Block diagram illustrating the structure of a graph neural network layer according to an embodiment.

[0032] Figure 7 Point cloud pre-training objective illustrated according to an embodiment.

[0033] Figure 8 Block diagram illustrating the pre-training process of a point cloud transformer model according to an embodiment.

[0034] Figure 9 Flowchart illustrating pre-training of a transformer model according to an embodiment.

[0035] Figure 10 Training and validation results of model performance on a shape-conditioned generation task illustrated according to an embodiment.

[0036] Figure 11 Training and validation results of model performance on a pocket-conditioned generation task in terms of non-blurry segments illustrated according to an embodiment.

[0037] Figure 12 Example of a molecular structure for shape-conditioned generation illustrated according to an embodiment.

[0038] Figure 13 Flowchart illustrating shape-conditioned generation using a trained transformer model according to an embodiment.

[0039] Figure 14 Flowchart illustrating linker design generation using a trained transformer model according to an embodiment.

[0040] Figure 15A and Figure 15B Example of molecular structure generation results compared with other models and ground truth illustrated according to an embodiment.

[0041] Figure 16A Example of molecular structure generation results in a linker design task illustrated according to an embodiment.

[0042] Figure 16B Example of molecular structure generation results in a pocket-conditioned generation task illustrated according to an embodiment.

[0043] Figure 17A and Figure 17B Example of molecular structure generation results of a PC transformer model compared with a SQUID model illustrated according to an embodiment.

[0044] Figure 18 Example of a computing system supporting techniques for implementing a point cloud transformer model illustrated according to an embodiment.

[0045] Figure 19 Visualization of learned relative biases illustrated according to an embodiment.

[0046] Figure 20 t-SNE visualizations of text atom token embeddings and text amino acid token embeddings are illustrated according to an embodiment.

[0047] Figure 21 A graphical comparison of model performance on shape-conditioned generation is illustrated according to an embodiment.

[0048] Figure 22A - 22D Input point cloud formation is shown.

[0049] The elements and components in the drawings may be arranged according to at least one embodiment described herein, and the arrangement may be modified by those of ordinary skill in the art based on the disclosure provided herein. Detailed Description

[0050] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, like symbols generally identify like components, unless the context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It is readily understood that the aspects of the present disclosure, which are generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in many different configurations, all of which are explicitly contemplated herein.

[0051] Language models (LMs) have demonstrated excellent natural language understanding and performance in a variety of natural language tasks, including language translation, question answering, and code generation. LMs also exhibit efficiency when acting as conversation agents and engaging in meaningful conversations. For example, large language models (LLMs) have shown overwhelming performance and advantages in natural language tasks. For instance, they can learn meaningful word representations based on text context and perform tasks such as translation from one language to another, and can maintain the conversation even when the conversation switches between different topics. LLMs can also accept input text queries, as well as optional images, to generate text answers. Additionally, recent research has shown that, in addition to text, LMs (e.g., the models released by Gill and IDEFICS) are also capable of integrating various data modalities, including images, videos, audio, and point clouds. Thus, LMs can be generalized and used as a foundation model for a variety of applications. The integration of data modalities can be achieved by including additional domain-specific encoders and decoders that convert non-text data into hidden embeddings, which can be added to the model input together with text token embeddings.

[0052] Recent research has shown that LMs can handle specialized chemical languages such as SMILES and self-referencing embedding strings (SELFIES) representations. These models have demonstrated a high proficiency in understanding and manipulating the textual representations of chemical data, enabling their application to a variety of tasks. For example, LMs have been utilized for molecular property prediction, molecular generation, and chemical reaction prediction. LLMs have shown the ability to understand not only natural language but also chemical languages such as SMILES (Galactica, MolFormer, ChemFormer, MolT5, T5Chem) and amino acid sequences (ProGen, RGN2). Thus, the models can work with text chemical data as input (e.g., molecular property prediction), as output (e.g., molecular generation), or both (e.g., reaction prediction).

[0053] While SMILES and SELFIES representations capture the structural properties of compounds, they have limitations in representing spatial features. Accurately determining the precise arrangement of atoms within a compound and their interactions with the surrounding environment is a key element in various drug discovery methods. Recent research has shown that LMs can effectively generate meaningful chemical 3D structures in text format by leveraging specialized formats such as Protein Data Bank (PDB), Crystallographic Information File (CIF), and XYZ, in which each line represents atomic coordinates, elements, and additional features. While this approach shows some promise, it is often limited by its length, for example, requiring dozens of tokens to 3D describe a single atom. While appropriate for molecules with a moderate number of atoms, these representations become redundant and impractical for larger structures with hundreds of atoms such as proteins. Additionally, these formats lack information about atomic connectivity and require the use of external software tools (e.g., Open Babel chemical toolkit software) to determine chemical bonds. In practice, this external software is highly sensitive to the quality of atomic positions, and even minor noise in the atomic coordinates can significantly alter the reconstructed chemical graph or cause the molecule to break into unconnected fragments.

[0054] In addition, there is a significant gap between structural two-dimensional (2D) and spatial three-dimensional (3D) chemical data. For example, when developing the recent AlphaFold 3D protein structure prediction architecture, the protein 3D structure prediction problem is one of the most fundamental and difficult computational tasks. In most drug discovery tasks, the 3D representation of protein pockets (i.e., protein structures with appropriate properties to bind ligands) is an important component in discovering ligand compounds that can bind and alter protein activity (e.g., see the Molecular Diffusion (MolDiff) and DiffLinker models).

[0055] Although recent progress suggests promise in integrating LMs into the drug discovery pipeline, existing models have some limitations, such as relying on chemical string representations (e.g., SMILES and SELFIES), which lack spatial features crucial for drug discovery. Additionally, attempts to convert chemical 3D structures into text formats have encountered problems such as excessive length and insufficient atomic connectivity information.

[0056] Thus, there is a need for a model, such as the novel point cloud transformer model described herein, that can work with chemical text data (as input, output, or both) and spatial (e.g., three-dimensional) data.

[0057] The point cloud transformer model described herein (also referred to herein as the "PC transformer" model or "transformer" model) combines a domain-specific encoder and a text representation (tokens) of the spatial atomic arrangement for generation tasks involving 3D molecular structures, including tasks that require efficient processing of spatial atomic configurations as input and output. In one embodiment, the transformer model architecture includes a point cloud input module and a text input module. The point cloud encoder module is operatively coupled to the point cloud input module, and the large language model module is operatively coupled to the text input module and the point cloud encoder module and is configured to receive data therefrom. The text output module is operatively coupled to the large language model module and is configured to output molecular data in line notation format. In some embodiments, the large language model module and the point cloud encoder module may be combined into a single model.

[0058] Specifically, the transformer model utilizes a point cloud encoder to achieve a concise and order-invariant representation of molecular and protein 3D structures. The text format is used to generate spatial molecular structures by first generating a chemical language (e.g., SMILES) representation and then generating atomic coordinates based on the chemical language sequence. This format eliminates the dependence on external software for reconstructing chemical bonds and allows the model to autonomously determine the molecular graph. Additionally, the encoder may include unique features, such as point embedding calculations and adjustments to position bias calculations, to adapt to the characteristics of domain-specific point cloud data.

[0059] A novel pre-training method for molecular point clouds extracts data from a spatial molecular structure dataset, i.e., 3D molecular data. For example, the pre-training scheme may include training a model to predict missing parts of an incomplete molecular point cloud by using a dropout technique that operates on entire sub-fragments of the 3D structure. Strategies for creating the input point cloud may include masking specific fragments and / or blurring selected fragments. In one embodiment, 3D molecular data including point clouds is fed into a point cloud encoder module in an unsupervised manner. Point descriptors are inferred by the pre-trained point cloud encoder module based on one or more points in the neighborhood of a point. The 3D position of a point is determined relative to one or more points in the neighborhood of the point. At least some of the one or more points are masked or blurred. At least some of the masked or blurred point features are predicted by the pre-trained point cloud encoder module. Random points are sampled, and the distances between the random points are predicted by the pre-trained point cloud encoder module. In some embodiments, a masking or blurring loss value may be determined based on the prediction of the masked or blurred point features. A distance loss value may be determined based on the prediction of the distances between the random points, and the weighted sum of the masking or blurring loss and the distance loss may be minimized based on the masking loss and the distance loss values respectively.

[0060] After fine-tuning the model within single-task and multi-task frameworks, the model exhibits excellent performance in several established spatial molecular generation tasks, outperforming competing diffusion models in terms of the quality of the generated samples and the efficiency of the training time. For example, in experiments conducted on six spatial molecular generation tasks, the Transformer model exhibits excellent performance, or at least shows results comparable to the LM baseline and the current diffusion scheme.

[0061] Figure 1 An overview diagram of a spatial molecular generation task is illustrated according to an embodiment. Each spatial molecular generation task, including conformation generation, pocket condition generation, linker design, can be broadcast into the Transformer model 100 as feed text and / or point cloud input data, and the model 100 is pre-trained and trained to generate a target point cloud represented as text data. Generally speaking, the types of tasks considered herein include text-to-text tasks (distribution learning 102 and conformation generation 104), molecular point cloud + text-to-text (shape condition generation 106 and linker design 108), and molecular / protein point cloud + text-to-text (e.g., pocket condition generation 110 and scaffold decoration 112). Additionally, those skilled in the art will understand that various other spatial molecular generation tasks may also fall within the scope of various embodiments.

[0062] Language models in chemistry

[0063] The sequential properties of molecules enable the use of transformer models and pre-training methods, in combination with models such as ChemBERTa, T5Chem (T5), ChemFormer, and BARTSmiles, which utilize masked language modeling for molecular SMILES representation. Recent advancements have introduced domain-specific LMs based on T5. MolT5 uses initial pre-training on a collection of molecular SMILES and text, followed by single-task fine-tuning on molecular annotation (molecule-to-text) and text-based molecular generation (text-to-molecule) tasks. On the other hand, Text+Chem T5 is a cross-domain, multi-task T5 model that is fine-tuned on five tasks including forward reaction prediction and retrosynthesis. Another recent model is the multi-domain nach0 large language model, which has undergone fine-tuning on a diverse collection of 28 task-dataset pairs, employing instruction fine-tuning in a multi-task manner. Different from MolT5 and Text+Chem T5, nach0 uses separate tokenization for chemical atoms and natural language tokens. BioT5 utilizes custom tokenization for SELFIES and natural language sequences and is fine-tuned on 15 tasks related to molecular and protein property prediction, drug-target interaction, and protein-protein interaction. DrugChat encodes molecular graphs using a graph neural network (GNN), employs a large LM (LLM), and uses adapters to transform the graph representation for LLM compatibility. Following Transformer-M, MolLM provides a unified pre-training framework that has a text transformer encoder and a molecular transformer encoder, pre-trained on molecular graphs, and processes both 2D and 3D structures through an attention mechanism that includes edge features and 3D spatial relationships. Those skilled in the art will understand that one or more of these language models can be employed to implement the various embodiments described herein. Additionally, various other language models can be adapted to implement the various embodiments and should be considered to be within the scope of the various embodiments.

[0064] Spatial molecular structure generation model

[0065] Most of the recently published spatial molecular structure generation models are based on the Denoising Diffusion Probabilistic Model (DDPM) paradigm. These models adopt a sequential denoising process, where the initial step involves the assignment of positions sampled from a Gaussian distribution model, followed by iterative noise removal to construct the molecular structure. The Entity Data Model (EDM) follows this approach to address the learning of spatial molecular distributions. However, a drawback of the diffusion scheme in generating 3D molecules is that it relies on external software (such as the open-source OpenBabel software) to reconstruct molecular bonds based solely on atomic coordinates. Even a slight error in the generated atomic positions can severely affect the reconstructed molecular graph. This limitation has been addressed in several subsequent works. The MolDiff model integrates an additional bond predictor to guide the diffusion and ensure bond consistency along with the 3D atomic coordinates. The MotionDiffusion (MDM) model integrates the diffusion model with a SchNet encoder and a scoring network to ensure edge consistency and enhance sample diversity. The diffusion model can be adopted in the conditional setting by predefining atomic types, positions, bonds, and / or freezing some point positions based on 3D conditions. The Geometric Diffusion (GeoDiff) and Torsion Diffusion models use the molecular graph to perform the conformational generation task. GeoMol is an end-to-end, non-autoregressive, and SE(3)-invariant machine learning scheme for generating the distribution of low-energy molecular 3D conformers. The DiffLinker and LinkerNet models are 3D equivariant diffusion models that can learn to generate linker fragments between given disconnected molecular subfragments in the linker design task. These models can find stable linkers and connect the fragments to produce low-energy conformations for the entire molecule. The DiffDec model is proposed for the scaffold (molecular core) decoration task and generates R-groups for a given molecular scaffold. In a task called shape-conditioned generation, the reference ligand molecule is represented as a shape, which is a fuzzy spatial region approximating the volume inside the molecular surface. The ShapeMol model uses the diffusion scheme to propose new molecules with shapes very similar to the reference molecule. The pocket-conditioned generation task involves creating molecules that can fit seamlessly into a specified pocket space, avoiding clashes while establishing interactions to enhance binding affinity. The D3FG and TargetDiff diffusion models can solve this task by de novo designing novel 3D molecules that will effectively bind to the specified protein pocket. The DecompDiff model proposes to enhance the diffusion process by leveraging data-dependent decomposition priors, reflecting the natural segmentation of ligand molecules into functional regions.However, as described above, there is a need for an improved model to process domain-specific data representing molecular features and spatial data.

[0066] The following sections refer to Figure 2A and Figure 2B which illustrate, according to embodiments, a functional-feature-based and modular-component-based block diagram of a point cloud transformer model architecture 200.

[0067] Referring to Figure 2A the point cloud transformer model architecture 200 is a multimodal architecture that processes both text data and 3D molecular point clouds. It receives these two types of input data and processes them to produce a desired output. The text data undergoes a standard tokenization process that breaks the text into smaller segments, or tokens. This process may be similar to the process used in general language models such as T5, Bidirectional Encoder Representations from Transformers (BERT), and BART. On the other hand, the 3D molecular point cloud is processed using a domain-specific point cloud encoder.

[0068] Large language models 202, also known as pre-trained language models or large-scale language models, are trained on large amounts of unlabeled text and spatial chemical structure data. These models can understand and interpret complex language patterns and structures, making them well-suited for processing the text data received by the PCTransformer (e.g., sometimes referred to as nach0.pc). These models are pre-trained, meaning they have been trained on large datasets before being used in the PCTransformer. This pre-training enables the models to process text data more efficiently and accurately.

[0069] Spatial graph neural networks 204, also known as graph neural networks for spatial data or spatial GNNs, work with 3D molecular point clouds. These networks are designed to process 3D data, making them well-suited for processing point clouds. These networks work by processing the relative distances between points within the point cloud. This enables the networks to accurately capture the spatial information contained within the point cloud.

[0070] The domain-specific point cloud encoder 206 processes the 3D molecular point cloud. This encoder can be specifically designed to process 3D data, ensuring accurate capture and processing of the information contained within the point cloud. The encoder maintains rotational, translational, and reflection invariance in the point embeddings, ensuring that the spatial information within the point cloud is preserved regardless of the orientation or position of the point cloud.

[0071] The graph neural network 208 processes the relative distances between points within the point cloud. This network is designed to work with relative distances, making it well-suited for processing the spatial information contained within the point cloud. The network represents points as nodes in a graph, with each node having associated features. The connectivity of the graph is described using an edge list and edge lengths, which describe the connections between the nodes.

[0072] The GaussianSmearing layer and the Pytorch Geometric library 210 (also referred to as the general layer and general library) are used to encode the edge lengths of the graph to obtain embeddings. These embeddings are then processed by a multi-layer perceptron (also referred to as an MLP or multi-layer neural network 212) to produce the final output of the point cloud encoder: a set of hidden representations, also referred to as hidden layer values or hidden layer outputs.

[0073] The TransformerConv and AttentionalAggregation operations 214 are used to update the hidden information of the nodes within the graph. These operations can be performed using the PyTorch Geometric library, which provides a range of tools and functions for working with graph data. These operations aggregate information from connected nodes and edges, and read global information from the entire graph, updating the hidden information of the nodes in the process. The output of these operations is a set of hidden representations, which represent the final output of the point cloud encoder.

[0074] Reference Figure 2B As shown in FIG. 1, the process-based block diagram of the point cloud transformer model 200 includes a point cloud encoder module 220, and a large language model module 230 including a text encoder 232 and a text decoder 234. In one embodiment, the architecture of the transformer model 200 extends the large language model module 230 with a domain-specific molecular point cloud encoder module 220 for generating token embeddings based on 3D molecular data including point clouds. For example, the base text encoder-decoder for the large language model module 230 may include the above-described T5 architecture. For example, using such a standardized architecture may allow users to train the large language model module 230 from scratch, or fine-tune a pre-trained large language model module 230, while using the point cloud encoder module 220 for both text and point cloud tasks. In one embodiment, the large language model module 230 may be initialized with natural language and chemical language models for multi-domain tasks, such as the nach0 LM from Insilico Medicine.

[0075] Input / output data format

[0076] In one embodiment, the transformer model 200 is configured to receive text data via the text input module 240 and optionally receive 3D molecular data via the molecular point cloud input module 250 to generate a text output via the text output module 260, which can be converted into a 3D molecular structure 270. In some aspects, the model 200 can consider dehydrogenated molecular structures.

[0077] As Figure 3 shown in more detail, the 3D molecular data including the point cloud received via the point cloud input module 350 may include an unordered set of points, denoted as {p i}, where each point p i = (c i , f i ) is described by a spatial position 352, e.g., Cartesian coordinates c i = (x i , y i , z i ). Each feature f i j can be regarded as a word token, or in other words, an unordered set of tokens 354, f i = {f i j}, which can be used to represent one or more point features. For example, the feature 356 of a point corresponding to an atom of a ligand may include its atomic symbol, atomic charge, whether it is in an aromatic ring (e.g., f i = {'atom_N', 'charge_0', 'aromatic',...}), and so on. For example, a ligand having a positively charged nitrogen atom can be represented as f i = {'ligand', 'N', '+'}. In another example, the features of the atoms of a protein can be extended to include its atom name and its corresponding amino acid name, e.g., f i = {'pocket', 'C', 'GLY', 'CA'}. Each point feature, such as an atomic symbol, atomic charge, atom name, or corresponding amino acid name, can be represented as a word token 354 and can be utilized in text input and / or output.

[0078] 3D molecular data can be represented in one or more of various chemical language formats. For example, at least one of the following chemical language formats can be used or combined: Simplified Molecular-Input Line-Entry System (SMILES) format, Self-Referencing Embedded Strings (SELFIES) format, or XYZ format text. In one example, 3D molecular data can be represented by combining SMILES and XYZ formats, where the SMILES format can be used to describe the molecular graph, and the lines of text can be concatenated in XYZ format to describe the position of each atom in the same order as it appears in the SMILES. This direct fusion of formats can eliminate the need for external software to reconstruct the molecular graph. In some embodiments, to optimize token counting, the number of digits after the decimal point can be limited (e.g., limited to two digits), and each coordinate can be tokenized by, for example, splitting the digits after the decimal point (‘-1.23’ to [‘-1’,‘.23’]), such that each coordinate can be described by two tokens. Those skilled in the art will understand that additional formats using additional tokens (e.g., three or more tokens) are possible.

[0079] In some embodiments, the 3D molecular data can be downsampled based on one or more priority points of the point cloud. For example, the points representing the ligand, alpha-carbon (C-α) atoms, or the terminal atoms of protein amino acids can be retained, while other points are ignored. This process can reduce the number of points the model has to process and can improve the memory efficiency of the model for larger point clouds.

[0080] Point cloud encoder

[0081] Continuing to refer Figure 3 , the point cloud encoder module 300 can be adapted to the text input data received via the text input module 340, and optionally, the point cloud input data (e.g., 3D molecular data) received via the point cloud input module 350. In one embodiment, the token 354 represents the feature of a specific spatial location 352. The token 354 is embedded via the token embedding layer 360, followed by sum pooling 362A and 362B to optimize memory and processing efficiency. In some embodiments, Scalar Sinusoidal Embeddings (SSE) 364 and 366 respectively integrate the continuous spatial coordinates (determined from the distance matrix 365 of the spatial location) and the relative pairwise distances, ensuring translational and rotational invariance.

[0082] In one embodiment, the point cloud encoder module 300 can include a language model text encoder with several modifications inspired by the nature of the point cloud data. Different from traditional text, the basic data unit is a point rather than a word token. Each point has a feature and a three-dimensional position.

[0083] One difference between the point cloud encoder 300 and the standard language model text encoder architecture is the point embedding calculation. For natural language, sequentially positioned text 342 is converted into word tokens 344, which are transformed using a token embedding layer 360. In one embodiment, the token embedding layer 360 can also be applied to 3D molecular data as point token embeddings. For example, point cloud features can be converted into hidden vectors by the token embedding layer 360 and then summed to form node-level hidden vectors The formula is as follows In Figure 4 shows the process for point embedding calculation.

[0084] Return reference Figure 3 , after the token embedding layer 360, point tokens can be aggregated, for example, by sum pooling operations 362a and 362b, to combine multiple point token embeddings. As a result, compared to the traditional text representation of the same data, fewer hidden vectors can be used to represent the point cloud including 3D molecular data. For example, dozens of embeddings can be used to represent a molecular point cloud, while hundreds of embeddings can be used to represent text tokens. This modification significantly reduces memory usage while enhancing the speed of attention mechanisms that may be sensitive to the input size, as shown in graph neural networks 368a and 368b. For example, the attention mechanism can calculate attention biases related to the relative positions and edge features between points of the point cloud and update the point embeddings based on the calculated attention biases.

[0085] In one embodiment, the point space coordinates 352 are another difference between the point cloud encoder module 300 and the standard language model text encoder architecture. Different from text tokens 344 whose token indices are discrete and form an increasing sequence, the spatial position of each point in the point cloud 350 (e.g., spatial positions 356 and 358) can be depicted using continuous Cartesian coordinates. For example, the point cloud relative position bias calculation 367 can be modified to embed pairwise distances instead of relative sequence positions, as is done for the relative text sequence position matrix 346 in the text relative position bias calculation 348. Additionally, before passing the embeddings to the decoder 380, the spatial coordinates of the points (e.g., points 356 or 358) can be embedded and summed with the point token embeddings 369a and 369b to ensure that the self-attention layer of the graph neural network 368b is invariant to any point cloud translation or rotation.

[0086] As described above, scalar sine embeddings (SSE) 364 and 366 can be used to embed scalar continuous values of coordinates and distances respectively. The inclusion of SSE draws on positional sine embeddings and extends the concept to map continuous scalar values into high-dimensional vectors. As represented in the following equation, the input scalar s is divided by the wavelength w initialized uniformly on a logarithmic grid i, and the generated sine and cosine function values are used to form the resulting embedding vector.

[0087] SSE 2i (s) = sin(s / w i ), SSE 2i+1 (s) = cos(s / w i )

[0088] The following shows an example point cloud encoding algorithm according to various embodiments:

[0089] Point cloud encoder algorithm

[0090] Input: point feature f i = {f i j} and coordinate c i = (x i , y i , z i ) c i = (x i , y i , z i )

[0091] Output: point embedding

[0092] 1: Embed and aggregate point features

[0093] 2: for l = 1, 2, …, L do

[0094] 3: For each head h, calculate the relative attention bias

[0095] 4: n l = SelfAttentionBlock l (n l-1 , b l ) Update the point embedding

[0096] 5: end for

[0097] 6: Embed coordinates

[0098] Graph neural network

[0099] Graph neural networks are an alternative way to construct the architecture of a point cloud encoder. However, as described herein, any network, technique, or method can be used to construct the architecture of a point cloud encoder. One constraint of the point cloud encoding process is that the point embeddings should be invariant to rotation, translation, and reflection of the point cloud in 3D. To overcome this problem, a graph neural network 368a and 368b is used, which directly utilizes the relative distances between points rather than the original 3D coordinates.

[0100] The inputs to graph neural networks 368a and 368b are graphs with nodes as points (including their corresponding features), and the connectivity of the graph is described by an edge list and edge lengths. Connecting all points to each other with edges may cause redundancy and consume resources. Therefore, in practice, the connections can be limited to groups of points, as well as the connections between each point and its m nearest neighbors within the point group. If a point belongs to several point groups, the point is connected to the nearest neighbors in each group. Thus, if a point belongs to d point groups, it may be connected to m*d other nodes. However, in practice, the nearest neighbors in different groups may overlap, and the actual number of connected nodes may be much less. To obtain the embeddings {e k} of the edges, the edge lengths can be encoded, for example, using the GaussianSmearing layer from the Pytorch geometric library, followed by a multi-layer perceptron ( Figure 5 ).

[0101] To update the hidden information of the nodes, a sequence of l graph neural network layers 368a and 368b can be applied. The structure of one layer can be found in Figure 6 . In one embodiment, each layer (1) aggregates information from the connected nodes and edges, and (2) processes (reads out) the global information from the entire graph. For example, TransformerConv and AttentionalAggregation from the PyTorch geometric library can be used for these two operations.

[0102] The output of the point cloud encoder module 300 is a set of hidden representations whose size is equal to the number of points in the point cloud, with one hidden representation vector for each point.

[0103] Molecular point cloud pre - training

[0104] In one embodiment, the point cloud encoder module 300 can be pre-trained in an unsupervised manner on a large set of 3D chemical data to enhance the generality of the entire PC Transformer model. The pre-training objectives of the point cloud encoder are (1) to perform reasoning based on point-to-point descriptions in the neighborhood of points, and (2) to understand the relative positions of points in 3D, even if these points are not connected by direct edges.

[0105] Figure 7 The pre-training process is illustrated according to the embodiments. For pre-training, a part of the points M can be masked, and further prediction can be made on the masked point features (as in the pre-training techniques used in LLM):

[0106]

[0107] In addition, random point pairs P = {(i, j)} are sampled to predict the distance d between them ij = ‖c i - c j ‖ 2 . In practice, the widely used MSE and MAE losses are greatly affected by point pairs that are far from each other. To overcome this bias and the neglect of close point pairs, the following loss function is utilized:

[0108]

[0109] where

[0110] In the pre-training stage, the weighted sum of these two losses is minimized:

[0111] L pretraining = λ mask L mask + λ dist L dist

[0112] In one embodiment, the pre-trained language model module and the point cloud encoder module are fused together into one model. To form the fused text-point cloud input of the model, the point cloud {p i} is encoded by the point cloud input module 350 to obtain point embeddings Meanwhile, the embeddings {w k} of the input text sequence {t k} can be prepared by the text input module 340. It should be noted that and {w k} can both be unordered sets of embeddings, because the point cloud does not have an ordered set of points, while the text embeddings include the sequence position by adding position encoding to each token embedding. To form the text-point cloud fused input, these two sets can form a union: Then the fused input can be passed to the large language model module (e.g., the LLM module 230). For example, the language model can include an encoder-decoder transformer model.

[0113] Depending on the generation task, the molecular point cloud may require additional preprocessing steps such as downsampling, augmentation, or adding noise. In the linker design task, some parts of the molecule are removed, and the model is trained to recover the removed parts. For example, the BRICS algorithm can be used to split the molecule into several fragments and remove a random number of fragments (at least one should be retained). For shape-conditioned generation, the random fragments can be blurred instead of completely removed. In this task, the selected atoms are replicated several times and high-variance noise is added.

[0114] Figure 8 A block diagram of a pre-training process 800 of a point cloud transformer model is illustrated according to an embodiment. For example, the point cloud encoder module 810 can be pre-trained using the spatial molecular structure 820 and its SMILES representation 830.

[0115] In one embodiment, the input spatial molecular structure dataset can be formed by: (i) deleting selected fragments and (ii) confusing the selected fragments by omitting point features (such as atom symbols, charges, and valences), replicating the points multiple times, and introducing Gaussian noise into the coordinates of these copies. For example, to reconstruct the randomly missing molecular fragment 840 of a molecular structure with known parts 850, a blurring operation (e.g., by Gaussian noise) 860 can be used to mask parts of the corresponding 3D point cloud, and a masking operation 870 can be used to omit the point features of the corresponding 3D point cloud. The point cloud encoder module 810 can be configured to extract knowledge from an unlabeled spatial molecular structure dataset including the point cloud. In one embodiment, the point cloud encoder module 810 can be fed an incomplete spatial molecular structure dataset and trained to reconstruct the missing parts. In particular, a dropout training method can operate on entire sub-fragments of the molecule rather than individual atoms. For example, the BRICS algorithm can be used to split the molecular representation into several fragments, and a subset of these fragments is randomly selected with a predefined probability (ensuring that at least one is selected) and excluded from the molecular representation. Additionally, the pre-training process can be flexible, handling both datasets including spatial molecular structures and datasets with protein pocket and ligand pairs. In the latter case, for example, the ligand can be masked while the protein pocket remains unchanged in the input point cloud.

[0116] As shown in 880, the model 810 can be configured to generate missing molecular fragments 840 during pre-training, including their SMILES representations, attachment points, and atomic coordinates. During pre-training, the task of the model is to accurately recover missing and / or ambiguous components, specifying their SMILES representations as well as their atomic coordinates. For example, each recovered fragment may include a connection point indicated by the symbol "*". In cases where the model is supposed to reconstruct multiple missing fragments, these fragments are separated by the "." token. In the experiments described below, the pre-training process 800 enhances the performance of the model on downstream tasks. Analysis of the impact of the pre-training process 800 is included in Example H below.

[0117] Figure 9 A flowchart 900 for pre-training a transformer model is illustrated according to an embodiment. At block 910, a transformer model is provided. For example, a transformer model, such as the transformer model 200 described above, may include a point cloud input module, a text input module, a point cloud encoder module operatively coupled to the point cloud input module, a large language model module operatively coupled to the text input module and the point cloud encoder module and configured to receive data therefrom, and a text output module configured to output molecular data in line notation format. In some embodiments, the large language model module and the point cloud encoder module may be combined into a single module, for example, as shown by the point cloud encoder module 300.

[0118] At block 920, the point cloud encoder module is pre-trained in an unsupervised manner using 3D molecular data including a point cloud. For example, the point cloud encoder module 810 can be pre-trained using the obtained representation of the spatial molecular structure 820 and its SMILES representation 830. At block 930, the point cloud encoder module infers a point description based on one or more points in the neighborhood of the point. At block 940, the 3D position of the point is determined relative to one or more points in the neighborhood of the point. At block 950, at least some of the one or more points are masked or blurred, and at block 960, at least some of the masked or blurred point features are predicted by the point cloud encoder module. In some embodiments, various points can be masked and blurred, as Figure 8 shown. For example, the pre-training data point cloud can be encoded into one or more point embeddings, and embeddings of the input text sequence can be prepared (see Figure 3 ). The one or more point embeddings and the input text sequence embeddings can be combined to obtain a fused input, and the fused input can be fed into the point cloud encoder module.

[0119] At block 970, random points are sampled, and at block 980, the distances between the random points are predicted by the point cloud encoder module. In some embodiments, a masked or blurred loss value may be determined based on the prediction of masked or blurred point features. For example, a distance loss value may be determined based on the prediction of the distances between the random points, and a weighted sum of the masked or blurred loss and the distance loss may be minimized based on the masked loss and the distance loss values, respectively.

[0120] Experiments

[0121] The proposed PC Transformer model was evaluated on several molecular generation tasks. The MOSES, GEOM, and CrossDocked2020 molecular datasets were used to train and validate the model on distribution learning, linker design, and shape-conditioned generation tasks.

[0122] MOSES is a benchmark platform that provides a molecular dataset and a set of metrics for comparing molecular generation models trained on this dataset. The dataset includes approximately 2 million molecules, which are divided into training, test, and scaffold test sections. Each molecule is represented in the text SMILES format. This unlabeled dataset was used to pre-train the model in an unsupervised manner so that it can understand the distribution of synthetically accessible and drug-like molecules, as well as the SMILES format structure and rules.

[0123] The GEOM dataset provides an ensemble of molecular conformations calculated using a combination of the CREST and GFN2-xTB software. The GEOM dataset contains two subsets: (1) conformations calculated for 133,258 molecules from the QM9 dataset, and (2) conformations calculated for 304,466 molecules from the AICures challenge. All conformations were calculated for a vacuum environment.

[0124] The CrossDocked2020 dataset is a large repository of molecular poses docked in various protein pockets. It contains 22.5 million ligand poses that are docked into multiple similar binding pockets on the Protein Data Bank.

[0125] Experiments were performed to evaluate the ability of the PC Transformer model to solve various molecular generation tasks. Substructures from the MOSES dataset were adopted to pre-train the text part of the proposed model. During the pre-training phase, the model was trained according to the rules of the SMILES format and the distribution of drug-like molecules.

[0126] Shape - conditioned generation

[0127] For shape-conditioned generation, the model is trained to recover a molecule from a blurred region (referred to as "shape") of the 3D space in which the molecule lies. In practice, some parts of the molecular point cloud are blurred, e.g., there is strong noise, and thus it is difficult for even human experts to identify the blurred segments. The model trained for this task can be used to generate known ligand surrogates that are spatially adaptable but have different molecular structures. If desired, some parts of the molecular point cloud can be kept unchanged and unblurred.

[0128] The training and validation of the model were conducted on the GEOM and CrossDocked2020 datasets. The results ( Figure 10 ) showed that the surrogates generated by the model had an average Tanimoto similarity of 0.62 (GEOM) / 0.61 (CrossDocked2020) and an average shape similarity of 0.76 (GEOM) / 0.70 (CrossDocked2020). The model retained the non-blurred segments ( Figure 11 ) in 97.32% (GEOM) / 92.91% (CrossDocked2020) of the cases.

[0129] Figure 13 A flowchart of shape-conditioned generation using the trained transformer model is illustrated according to an embodiment. For example, a computer system as described below Figure 18 may include a trained transformer model and one or more non-transitory computer-readable media that store instructions which, when executed by one or more processors, may cause the computer system to perform Figure 13 one or more operations described in

[0130] , including, for example, operating the transformer model or its pre-training or its implementation to perform shape-conditioned generation of at least one molecule. Figure 13 Referring to

[0131] Figure 12

[0132] , shape-conditioned generation 1300 includes, at block 1310, training a transformer model to recover a molecule from a blurred region of the 3D space in which a certain molecule of the molecular point cloud lies. For example, some parts of the molecular point cloud may be blurred, and selected parts of the molecule may be kept unchanged and unblurred. At block 1320, 3D molecular data including the point cloud is input via a point cloud input module. At block 1330, text representing the desired molecule is input via a text input module. At block 1340, the 3D molecular data and the text are processed using a trained point cloud encoder module, and at block 1350, text representing the molecule is output in line notation.

[0131] Examples of shape-conditioned molecules generated as described herein can be found in Figure 12

[0132] Linker design

[0133] In the linker design task, the model is trained to recover the removed part of the molecule. This task is more complex than shape-conditioned generation because the model does not receive a spatial description (size, curvature, etc.) of the removed fragment. Similar to the previous tasks, the model trained to solve the linker design problem is able to generate molecular alternatives for known ligands. It can also be applied to cases where there is no known ligand, but it is known where the active fragment should be located in 3D space.

[0134] The proposed model was trained and validated on the GEOM dataset to solve the linker design problem. In 96.39% of the cases, the generated molecules contain the desired fragment conditions. Examples of aligned ground truth / generated pairs and fragment conditions / generated molecules can be found in Figure 16A .

[0135] Figure 14 A flowchart of linker design generation using the trained transformer model is illustrated according to an embodiment. For example, a computer system as described below Figure 18 may include a trained transformer model and one or more non-transitory computer-readable media that store instructions that, when executed by one or more processors, may cause the computer system to perform Figure 14 one or more operations described in

[0136] , including, for example, operating the transformer model architecture or its pre-training or its implementation to perform linker design generation for at least one molecule. Figure 14 Referring to

[0137] the linker design 1400 generates, including at block 1410, training a transformer model to recover the removed part of the molecule when the transformer model architecture does not receive a spatial description of the removed part of the molecule. At block 1420, data representing one or more molecules is input into the trained transformer model. At block 1430, the trained transformer model is used to generate one or more options for the linker part of the molecule to replace the removed part of the molecule. At block 1440, for the linker part of the molecule, one or more molecules with the generated one or more options are obtained.

[0138] Experimental results

[0139] The effectiveness of the point cloud transformer model has been evaluated in the following spatial molecular generation tasks: (i) 3D molecular structure generation (spatial molecular distribution learning, conformation generation); (ii) molecular point cloud completion (linker design, scaffold decoration); (iii) shape-conditioned generation; and (iv) pocket-conditioned generation. Details of the datasets, models, and metrics used to obtain the experimental results can be found in Example Subsections B, C, D, and F below.

[0140] Spatial molecular distribution learning and conformation generation tasks

[0141] The 3D molecular structure generation evaluation models the ability of the model to generate structurally diverse and physically reasonable spatial molecular objects.

[0142] The spatial molecular distribution learning task evaluates whether the model can generate novel 3D molecular structures with a distribution close to the ground truth. High-quality datasets such as the GEOM-Drugs (also referred to as "GEOM" in this paper) dataset were used to provide conformational ensembles generated using metadynamics in CREST. The following valid molecules were evaluated from five perspectives: drug likeness, 3D substructure, bonds, ring distribution, and root mean square distribution (RMSD) with respect to the template conformation.

[0143] In conformation generation, the focus is on generating reasonable conformations for a given molecular graph. The conformation of a molecule is its energetically favorable 3D structure, and each conformation represents a local minimum on the potential energy surface. As in the previous task, the generated conformational ensembles were evaluated on the GEOM dataset. The average minimum RMSD (AMR) and coverage metrics were followed and adopted. These metrics evaluate the recall rate (R), which indicates the degree to which the generated ensemble covers the ground truth ensemble, and the precision (P), which measures the quality within the generated conformations.

[0144] Consistent with the baselines mentioned, 2K conformations were generated for molecules with K ground truth conformations, and the training / validation / test split was utilized. The baselines for the spatial molecular distribution learning task were retrained on this data split, and the metrics on 10K sampled objects were reported for fair comparison. This includes the retrained MolDiff model with all necessary elements, which originally supported only molecules with the major elements (C, N, O, F, P, S, and Cl).

[0145] The generative capabilities of models trained in single-task and multi-task settings were evaluated by (i) comparing with the MolDiff model for the first task and (ii) comparing with the GeoMol, GeoDiff, and Torsion Diffusion (Tor.Diff.) models for the second task. Based on the open-source LLaMa model architecture from HuggingFace, an LM baseline model was adopted for both tasks. The Open LLaMa implementation was utilized and trained on the same GEOM-Drugs dataset (see the configuration details in Example E.2).

[0146] The results of the two tasks are presented in Tables 1 and 2.

[0147]

[0148]

[0149] Table 1: Spatial molecular distribution learning performance metrics on GEOM-DRUGS.

[0150]

[0151] Table 2: Quality of the conformational ensembles generated on GEOM-DRUGS

[0152] As shown in Table 1, the performance of the LMs (PC Transformer model and OpenLLaMa) is superior to that of the MolDiff diffusion model, producing a molecular distribution closer to the ground truth in terms of 3D substructures, bonds, and ring distributions. In particular, the PC Transformer model generates more accurate molecular structures, performing better than OpenLLaMa on 3D substructure groups, while being better or comparable on other metrics ( Figure 15A and Figure 15B ). In the conformational generation task, the only model that outperforms the PC Transformer model on all metrics is the Torsion Diffusion model, which utilizes samples from an external software (RDKit) as the initial generation points during the inference stage. Nevertheless, the PC Transformer model still shows stronger results than all other pure neural network baselines that generate molecular conformations from scratch.

[0153] Linker generation and scaffold decoration tasks

[0154] The molecular point cloud completion (linker design, scaffold decoration) task evaluates the ability of the model to complete disjoint or partially defined molecular structures.

[0155] In the linker design task, the model operates on several disconnected fragments and should generate small molecule structures that are spatially and chemically connected to the given fragments (as Figure 16Aas shown), and completed as a connected chemical structure. A subset of the ZINC dataset was adopted, including 250K random molecules with conformations generated using RDKit. This dataset also provided the input fragments and linkers for each molecule.

[0156] In the second task, the model takes the core part of the molecule, i.e., the scaffold, and adds specific side-chain motifs called R-groups. This molecular completion task is called scaffold decoration. Usually, scaffold decoration is employed to enhance some molecular properties, such as the binding affinity to a specific protein. A multi-R-group decoration task on the CrossDocked dataset containing 100K ligand-protein pairs was used, where each ligand was split into a scaffold and an R-group.

[0157] The PC Transformer model takes molecular fragments / scaffolds and protein pockets (if any) as input point clouds and generates linkers / R-groups without repeating input atoms. The generated molecular substructures contain attachment points described by the symbol "*" and coordinates, as referred to above Figure 8 (block 880). These attachment points are used to combine the input fragments with the generated fragments into a coherent molecule. The model can generate several R-groups in the scaffold decoration task. Each R-group also contains attachment points and is separated by the symbol ".". For a comparison of pocket-conditioned generation, see Figure 16B .

[0158] The comparison of the PC Transformer model (nach0-pc) with the DeLinker, 3DLinker, and DiffLinker models in the linker generation task, and with the LibINVENT, FLAG, and DiffDec models in scaffold decoration is discussed below.

[0159] As shown in Tables 3 and 4 below, for both tasks, the PC Transformer model can complete the input molecular point clouds with a high success rate, and the generated molecules can pass 2D filters, such as PAINS. Although the structural diversity is moderate, the PC Transformer model generates spatially diverse molecules. In addition, it enhances the scaffold binding affinity, comparable to other current state-of-the-art scaffold decoration models.

[0160]

[0161] Table 3: Model performance evaluation for the linker design task (baseline metrics from

[41] ).

[0162] Shape - conditioned generation

[0163] Shape-conditioned generation focuses on generating molecules that are spatially similar but structurally dissimilar to a reference structure. This is achieved by representing the reference molecule as a shape - i.e., the region where the nuclei and electron clouds of the molecule are located. In practice, in drug discovery tasks, molecules with similar shapes are more likely to have similar properties, even if the structures are chemically distinct. The molecular shape is represented as a point cloud. Each atom is replicated a number of times, and Gaussian noise with a predefined standard deviation σ is added to the atomic positions while removing all point features completely. By alternating the parameter σ, a balance between spatial similarity and chemical similarity can be achieved.

[0164] Training and test phase sampling is performed using conformations calculated from the MOSES dataset using RDKit. Figure 17A and Figure 17B A comparison of the PC Transformer model with the SQUID model is shown. The noise injection parameters are alternated, being the standard deviation σ for the PC Transformer model (nach0-pc) and the prior interpolation coefficient λ for SQUID, to show the available trade-off between structural and shape similarity. It is shown that the PC Transformer model provides a wider range of available trade-offs, covering the region from high structural similarity to high shape similarity. Additionally, compared to SQUID, the PC Transformer model is shown to produce objects with higher spatial similarity for lower values of structural similarity.

[0165] Pocket - conditioned generation

[0166] For pocket-conditioned generation, the PC Transformer model is trained to generate novel high-affinity structures for a given protein pocket. The CrossDocked dataset is used, and only high-affinity ligand-protein pairs are retained. During testing, for each protein pocket in the test set, one hundred (100) molecules are randomly sampled. The generation quality of the PC Transformer model (nach0-pc) is evaluated. A comparison with the AR, Pocket2Mol, and TargetDiff models is provided in Table 4.

[0167]

[0168] Table 4: (Left) Pocket-conditioned generation performance metrics (metrics for baselines) (Right) Scaffold decoration task metrics (metrics for baselines).

[0169] Although there is a performance gap between the PC Transformer model and the TargetDiff model, based on docking scores and binding affinities, the PC Transformer model (nach0-pc) shows results comparable to AR and Pocket2Mol.

[0170] Therefore, the PC Transformer model (nach0-pc) is adept at generating diverse and physically plausible chemical 3D structures. By combining a domain-specific point cloud encoder with an encoder-decoder language model, the PC Transformer model (nach0-pc) effectively addresses the challenges associated with handling chemical 3D structures and SMILES sequences. Through extensive fine-tuning within single-task and multi-task frameworks, the PC Transformer model (nach0-pc) exhibits superior or comparable performance relative to various state-of-the-art diffusion models.

[0171] In some embodiments, the substance is a molecule, such as a small molecule, a macromolecule, a polypeptide, a protein, an antibody, an oligonucleotide, a nucleic acid (e.g., RNA, DNA, etc.), a polypeptide, a carbohydrate, a lipid, or a combination thereof, whether natural or synthetic.

[0172] In some embodiments, the generated molecules are analyzed, and one or more specific molecules that meet the specified criteria are selected. Then the selected one or more molecules are selected and synthesized, and then tested using one or more cells to determine whether the synthesized molecules actually meet the criteria.

[0173] Once one or more molecules are generated, the model can classify the molecules according to any desired profile. Specific physical properties, such as specific chemical groups or 3D structures, can be prioritized, and then molecules with profiles that match the desired profile are selected and synthesized. In this way, an object selector (e.g., a molecule selector), which can be a software module, selects at least one molecule for synthesis, which can be done through filtering as described herein. The selected molecule is then provided to an object synthesizer, and the selected object (e.g., the selected molecule) is synthesized in the synthesizer. The synthesized object (e.g., the molecule) is then provided to an object validator (e.g., a molecule validator), which tests whether the object meets the criteria or properties, or tests whether it has biological activity for a specific use. For example, a synthesized object that is a molecule can be tested with live cell cultures or other validation techniques to verify that the synthesized molecule meets the desired properties. In this way, an object such as a molecule or other substance has a real-world version obtained, for example, by purchasing, preparing, synthesizing, or otherwise obtaining the object in its physical form.

[0174] Once the generated object is selected, the method then includes validating the selected object. The validation can be performed in the manner described herein. When the object is a molecule, the validation can include synthesis and then testing with live cells.

[0175] In some embodiments, a method may include selecting a selected substance corresponding to the selected generated data or to a desired property; and validating the selected substance. In some embodiments, the method may include: obtaining a physical version of the selected substance; and testing the physical version to obtain a desired property or biological activity. Additionally, in any method, obtaining a physical version of a substance may include synthesizing, purchasing, extracting, refining, derivatizing, or otherwise obtaining at least one of physical substances. The physical substance may be a molecule or other substance. These methods may include tests involving assaying the physical substance in cell culture. These methods may also include assaying the physical substance by genotyping, transcriptomic typing, 3D mapping, ligand-receptor docking, pre- and post-perturbation, initial state analysis, final state analysis, or a combination thereof. When the physical substance is a new molecular entity, preparing a physical version of the selected generated substance often may include synthesis. Thus, these methods may include selecting what is not part of an original data set or previously known generation objects.

[0176] In some embodiments, the method may include: obtaining a physical form of the selected compound including synthesizing, purchasing, extracting, refining, derivatizing, or otherwise obtaining at least one of physical objects; and / or testing including assaying the physical form of the selected object in cell culture; and / or assaying the physical form of the selected compound by genotyping, transcriptomic typing, 3D mapping, ligand-receptor docking, pre- and post-perturbation, initial state analysis, final state analysis, or a combination thereof.

[0177] Those skilled in the art will appreciate that for the processes and methods disclosed herein, the functions performed in the processes and methods may be implemented in a different order. Additionally, the steps and operations outlined are provided as examples only, and some steps and operations may be optional, combined into fewer steps and operations, or expanded into additional steps and operations without detracting from the essence of the disclosed embodiments.

[0178] In one embodiment, aspects of the method may be performed on a computing system. As such, the computing system may include a memory device having computer-executable instructions for performing these methods. The computer-executable instructions may be part of a computer program product that includes one or more algorithms for performing any method of any claim.

[0179] In one embodiment, any operation, process, or method described herein may be performed or caused to be performed in response to the execution of computer-readable instructions stored on a computer-readable medium and executable by one or more processors. The computer-readable instructions may be executed by processors of a variety of computing systems, including desktop computing systems, portable computing systems, tablet computing systems, handheld computing systems, and network elements and / or any other computing device. The computer-readable medium is non-transitory. The computer-readable medium is a physical medium on which computer-readable instructions are stored for physical reading by a computer / processor from the physical medium.

[0180] The processes and / or systems and / or other technologies described herein may be implemented via various carriers (e.g., hardware, software, and / or firmware), and the preferred carrier may vary depending on the context in which these processes and / or systems and / or other technologies are deployed. For example, if the implementer determines that speed and accuracy are critical, the implementer may choose a carrier dominated by hardware and / or firmware; if flexibility is critical, the implementer may choose a software-dominated implementation; or, as an alternative, the implementer may choose some combination of hardware, software, and / or firmware.

[0181] The various operations described herein can be implemented singly and / or in combination by a variety of hardware, software, firmware, or nearly any combination thereof. In one embodiment, several parts of the subject matter described herein can be implemented via an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), or other integrated format. However, some aspects (all or part) of the embodiments disclosed herein can be equivalently implemented with integrated circuits, e.g., as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or as nearly any combination thereof, and it is possible to design the circuitry and / or write the code for the software and / or firmware in accordance with this disclosure. Additionally, the mechanisms of the subject matter described herein can be distributed in a variety of forms as a program product, and illustrative embodiments of the subject matter described herein are applicable regardless of the particular type of signal bearing medium used to actually effect the distribution. Examples of physical signal bearing media include, but are not limited to, the following: recordable type media such as floppy disks, hard disk drives (HDD), compact discs (CD), digital versatile discs (DVD), digital tapes, computer memories, or any other non-transitory or non-transmissive physical medium. Examples of physical media with computer readable instructions do not include transitory or transmissive type media such as digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.).

[0182] Devices and / or processes are typically described in the manner set forth herein, and then such described devices and / or processes are integrated into a data processing system using engineering practices. That is, at least a portion of the devices and / or processes described herein can be integrated into a data processing system via a reasonable amount of experimentation. A typical data processing system generally includes one or more system unit enclosures, a video display device, memory, such as volatile and non-volatile memory, processors, such as microprocessors and digital signal processors, computing entities, such as operating systems, drivers, graphical user interfaces, and applications, one or more interaction devices, such as touchpads or screens, and / or control systems, including feedback loops and control motors (e.g., feedback for sensing position and / or speed; control motors for moving and / or adjusting components and / or amounts). A typical data processing system can be implemented using any suitable commercially available components, such as those generally included in data computing / communication and / or network computing / communication systems.

[0183] The subject matter described herein sometimes illustrates that different components are included within or connected to different other components. The architectures so depicted are merely exemplary, and in fact, many other architectures can be implemented to achieve the same functionality. Conceptually, any arrangement of components for achieving the same functionality is in fact "associated" so as to achieve the desired functionality. Thus, any two components combined herein for achieving a particular functionality can be regarded as being "associated" with each other so as to achieve the desired functionality, regardless of the architecture or intermediate components. Similarly, any two components so associated can also be regarded as being "operably connected" or "operably coupled" to each other to achieve the desired functionality, and any two components capable of being so associated can also be regarded as being "operably coupled" to each other to achieve the desired functionality. Specific examples of operably coupled include, but are not limited to: physically mating and / or physically interacting components and / or wirelessly interacting and / or wirelessly interactive components and / or logically interacting and / or logically interactive components.

[0184] Figure 18 An example computing device 600 (e.g., a computer) that can be arranged to perform the methods (or portions thereof) described herein is shown. In a very basic configuration 602, the computing device 600 generally includes one or more processors 604 and a system memory 606. A memory bus 608 can be used for communication between the processor 604 and the system memory 606.

[0185] Depending on the desired configuration, processor 604 can be of any type, including but not limited to: a microprocessor (μP), a microcontroller (μC), a digital signal processor (DSP), or any combination thereof. Processor 604 can include one or more levels of cache, such as first-level cache 610 and second-level cache 612, processor core 614, and registers 616. Example processor core 614 can include an arithmetic logic unit (ALU), a floating-point unit (FPU), a digital signal processing core (DSP core), or any combination thereof. Example memory controller 618 can also be used with processor 604, or in some implementations, memory controller 618 can be an internal part of processor 604.

[0186] Depending on the desired configuration, system memory 606 can be of any type, including but not limited to: volatile memory (such as RAM), non-volatile memory (such as ROM, flash memory, etc.), or any combination thereof. System memory 606 can include an operating system 620, one or more applications 622, and program data 624. Application 622 can include a determination application 626, which is arranged to perform operations as described herein, including those described for the methods herein. Determination application 626 can obtain data, such as pressure, flow rate, and / or temperature, and then determine changes to the system to change pressure, flow rate, and / or temperature.

[0187] Computing device 600 can have additional features or functionality, as well as additional interfaces, to facilitate communication between the basic configuration 602 and any desired devices and interfaces. For example, bus / interface controller 630 can be used to facilitate communication between the basic configuration 602 and one or more data storage devices 632 via a storage interface bus 634. Data storage device 632 can be a removable storage device 636, a non-removable storage device 638, or a combination thereof. Examples of removable and non-removable storage devices include: disk devices such as floppy disk drives and hard-disk drives (HDDs), optical disk drives such as compact disk (CD) drives or digital versatile disk (DVD) drives, solid state drives (SSDs), and tape drives, to name a few. Example computer storage media can include: volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data.

[0188] System memory 606, removable storage 636, and non-removable storage 638 are examples of computer storage media. Computer storage media includes, but is not limited to: RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs), or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices, or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by computing device 600. Any such computer storage media can be part of computing device 600.

[0189] Computing device 600 may also include interface bus 640 for facilitating communication from various interface devices (e.g., output device 642, peripheral interface 644, and communication device 646) via bus / interface controller 630 to basic configuration 602. Example output device 642 includes graphics processing unit 648 and audio processing unit 650, which may be configured to communicate with various external devices such as a display or speakers via one or more A / V ports 652. Example peripheral interface 644 includes serial interface controller 654 or parallel interface controller 656, which may be configured to communicate with external devices such as input devices (e.g., keyboard, mouse, pen, voice input device, touch input device, etc.) or other peripheral devices (e.g., printer, scanner, etc.) via one or more I / O ports 658. Example communication device 646 includes network controller 660, which may be arranged to facilitate communication with one or more other computing devices 662 via one or more communication ports 664 over a network communication link.

[0190] A network communication link can be an example of a communication medium. Communication media typically can be implemented by computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and can include any information delivery media. A "modulated data signal" can be a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media can include wired media such as a wired network or direct-wired connection, and wireless media such as sound, radio frequency (RF), microwave, infrared (IR), and other wireless media. The term "computer-readable medium" as used herein can include both storage media and communication media.

[0191] The computing device 600 can be implemented as part of a small form factor portable (or mobile) electronic device such as the following: cellular phone, personal data assistant (PDA), personal media player device, wireless network watch device, personal earpiece device, dedicated device, or a hybrid device that includes any of the above functions. The computing device 600 can also be implemented as a personal computer that includes both laptop and non-laptop computer configurations. The computing device 600 can also be any type of network computing device. The computing device 600 can also be an automation system as described herein.

[0192] The embodiments described herein may include the use of a special purpose or general purpose computer that includes various computer hardware or software modules.

[0193] Embodiments within the scope of the present invention also include computer-readable media for carrying or storing thereon computer-executable instructions or data structures. Such computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general purpose or special purpose computer. When information is transferred or provided to a computer via a network or another communication connection (either hardwired, wireless, or a combination of hardwired or wireless), the computer properly views the connection as a computer-readable medium. Thus, any such connection is properly termed a computer-readable medium. Combinations of the above should also be included within the scope of computer-readable media.

[0194] Computer-executable instructions, for example, include instructions and data that cause a general purpose computer, special purpose computer, or special purpose processing device to perform a particular function or a set of functions. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. For example, one or more neural network configurations can be used to perform the specific features and behaviors described above.

[0195] In some embodiments, a computer program product may include a non-transitory tangible memory device having computer-executable instructions that, when executed by a processor, cause the methods described herein to be performed. The non-transitory tangible memory device may also have other executable instructions for any of the methods or method steps described herein. Additionally, the instructions may be instructions for performing non-computational tasks, such as the synthesis of molecules and / or experimental protocols for verifying molecules. Other executable instructions may also be provided.

[0196] The present disclosure is not limited to the specific embodiments described in this application, which are intended to be illustrative of various aspects. Many modifications and variations can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. In addition to those enumerated herein, functional equivalent methods and apparatuses within the scope of the present disclosure will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to be included within the scope of the appended claims. The present disclosure is limited only by the terms of the appended claims and the full scope of equivalents to which such claims are entitled. It is to be understood that the present disclosure is not limited to the particular methods, reagents, compound components, or biological systems, of course, which may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.

[0197] For substantially any plural and / or singular terms used herein, those skilled in the art may convert from plural to singular and / or from singular to plural as appropriate, depending on the context and / or application. For clarity, various singular / plural forms may be explicitly recited herein.

[0198] Those skilled in the art will understand that, in general, the terms used herein, and especially the terms used in the appended claims (e.g., the subject matter of the appended claims), are generally intended to be "open" terms (e.g., the term "comprising" should be interpreted as "comprising but not limited to", the term "having" should be interpreted as "having at least", the term "including" should be interpreted as "including but not limited to", etc.). Those skilled in the art will understand that if the intention is to specify the number of introduced claim recitations, such intention will be explicitly recited in the claim, and in the absence of such recitation, there is no such intention. For example, as an aid to understanding, the following appended claims may contain the use of introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed to mean that the introduction of a claim recitation by the indefinite article "a" limits any particular claim containing such introduced claim recitation to only embodiments containing one such recitation, even when the same claim includes an introductory phrase "one or more" or "at least one" and an indefinite article such as "a" (e.g., "a" when interpreted to mean "at least one" or "one or more"); this also holds for the use of the definite article to introduce claim recitations. In addition, even if a specific number of introduced claim recitations is explicitly recited, those skilled in the art will also recognize that such recitation should be interpreted as meaning at least the recited number (e.g., merely reciting "two recitations" without further modifiers means at least two recitations, or two or more recitations). In addition, in cases where idiomatic expressions similar to "at least one of A, B, and C, etc." are used, generally such constructions are intended to have the meaning of such idiomatic expressions that those skilled in the art will understand (e.g., "a system having at least one of A, B, and C" will include, but not be limited to, a system having only A, a system having only B, a system having only C, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having all of A, B, and C, etc.). In cases where idiomatic expressions similar to "at least one of A, B, or C, etc." are used, generally such constructions are intended to have the meaning of such idiomatic expressions that those skilled in the art will understand (e.g., "a system having at least one of A, B, or C" will include, but not be limited to, a system having only A, a system having only B, a system having only C, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having all of A, B, and C, etc.). Those skilled in the art will also understand that almost any disjunctive word and / or phrase presenting two or more alternative terms - whether in the specification, claims, or drawings - should be understood to contemplate the possibility of including one of such terms, any one of such terms, or both of such terms.For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B".

[0199] In addition, in the case where the features or aspects of the present disclosure are described in terms of Markush groups, those skilled in the art will recognize that the present disclosure is thereby also described in terms of any individual member of the Markush group or a subgroup of the members.

[0200] Those skilled in the art will understand that, for any and all purposes, such as for providing a written description, all ranges disclosed herein also cover any and all possible sub-ranges and combinations of sub-ranges thereof. Any recited range can be readily recognized as fully describing and enabling the same range to be at least divided into equal halves, thirds, quarters, fifths, tenths, and so on. By way of non-limiting example, each range discussed herein can be readily divided into a lower third, a middle third, and an upper third, and so on. Those skilled in the art will also understand that all language such as "up to", "at least", etc. includes the recited number and refers to a range that can then be divided into sub-ranges as described above. Finally, those skilled in the art will understand that a range includes each individual member. Thus, for example, a group having 1 - 3 units refers to a group having 1, 2, or 3 units. Similarly, a group having 1 - 5 units refers to a group having 1, 2, 3, 4, or 5 units, and so on.

[0201] From the foregoing, it will be appreciated that various embodiments of the present disclosure are described herein for purposes of illustration, and various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various embodiments disclosed herein are not intended to be limiting, and the true scope and spirit are indicated by the appended claims.

[0202] All references cited herein are hereby incorporated by specific reference in their entirety.

[0203] Examples

[0204] A. Experimental constraints

[0205] In this work, nach0 pre-trained on SMILES data was utilized. A significant drawback of the SMILES format is the lack of a one-to-one correspondence between molecules and SMILES strings. Due to variations in the starting atom, molecular graph traversal, and kekulization, a single molecule can have multiple SMILES representations. Secondly, scaling the PC Transformer model for long chemical sequences can be a challenge due to the quadratic nature of the attention mechanism. Finally, it is important to recognize that the academic datasets used in this study mainly consist of existing drugs and known chemical probes, covering only a small fraction of the vast predictive chemical space. Additionally, they do not account for the testing of novel chemical diversities that are different from the molecules reported in the literature.

[0206] B. Dataset statistics and examples of inputs and outputs

[0207] Several datasets were employed to evaluate and benchmark our generative models for different molecular generation tasks. These datasets provide comprehensive and high-quality molecular data, facilitating tasks such as property prediction, conformation generation, and drug design. The following is a detailed description of the datasets used. Table 5 shows examples of the input and output for each task.

[0208]

[0209]

[0210] Table 5

[0211] "No input PC" means there is no input point cloud. The listed datasets were utilized during the pre-training phase. To address the imbalanced dataset sizes, a balanced batch construction process was employed, which includes samples from different tasks with equal probability of forming the final training batch.

[0212] B.1 GEOM

[0213] For the tasks of learning spatial molecular distributions and conformational generation, the GEOM-Drugs dataset was utilized, which provides conformational ensembles generated by metadynamics in CREST. The Geometric Ensemble of Molecules (GEOM) dataset is a comprehensive collection of molecular conformations generated using advanced sampling and semi-empirical density functional theory (DFT). This dataset contains over 37 million molecular conformations of more than 450,000 molecules. The GEOM dataset includes conformations of 133,000 species from QM9 and 317,000 species with experimental data related to biophysics, physiology, and physical chemistry. The dataset also contains ensembles of 1,511 species, where BACE-1 inhibition data are labeled with high-quality DFT free energies in implicit water solvent, and 534 ensembles are further optimized by DFT. This dataset can be used for model training and benchmarking in tasks such as molecular property prediction, conformational generation, and drug design.

[0214] B.2 ZINC dataset

[0215] For the linker design task, a subset of the ZINC dataset was utilized, which includes 250,000 randomly selected molecules with low-energy conformations generated using RDKit. The ZINC dataset is a comprehensive collection of commercially available compounds, mainly selected for virtual screening applications. It contains molecular graphs, each with up to 38 heavy atoms. This dataset is widely used in computational chemistry and drug discovery research for tasks such as molecular property prediction. It provides a diverse and well-characterized set of molecules that helps in developing models that produce chemically valid, unique, and novel compounds with desired properties.

[0216] Each molecule undergoes fragmentation by disconnecting all double bonds of acyclic single bonds, and the resulting partitions are filtered according to a set of predefined rules. Due to this fragmentation process, a molecule can produce various combinations of two fragments with a linker in between. In the experiment, the focus was on generating linkers without specifying the anchor points (connection points in the input fragments).

[0217] B.3 MOSES

[0218] For the shape-conditioned generation task, the MOSES dataset was utilized. Initially, this dataset was filtered to retain only the molecules containing the 100 most prevalent fragments, and then RDKit was used to generate 3D conformers for the remaining molecules. The MOSES dataset is a benchmarking platform that provides a huge dataset and a set of metrics for evaluating generative models in the context of unconditional molecular generation tasks. The dataset includes nearly 2 million molecular samples, which were filtered using MCF, PAINS, and other rules to ensure quality. The provided metrics evaluate the performance of generative models from multiple perspectives: the validity of the generated structures, the quality of the molecular distribution match, and the ability of the model to generate novel and diverse molecules.

[0219] B.4 CrossDocked

[0220] For the scaffold decoration and pocket-conditioned generation tasks, the CrossDocked dataset was adopted. It includes 22.5 million protein-ligand complexes. The dataset was preprocessed by retaining only 100,000 ligand-protein pairs with the highest binding affinities and clipping the protein pocket regions within the ligands. Additionally, the LibINVENT slicing method with 37 custom reaction-based rules was applied to the ligands to obtain scaffolds and R groups.

[0221] C Evaluation metrics

[0222] C.1 Molecular distribution learning

[0223] Valid and complete molecules were evaluated from four perspectives: drug-likeness, 3D structure, bonds, and rings.

[0224] The drug-likeness of the generated molecules was evaluated using the following metrics: (1) QED, representing the quantitative estimate of drug-likeness; (2) SA, indicating the synthetic accessibility score (higher values indicate easier synthesis); and (3) Lipinski, measuring the number of Lipinski's five-rule criteria satisfied by the molecule.

[0225] The key difference between 3D generation and 2D molecular graph generation lies in the determination of atomic positions, which makes it crucial to measure its accuracy. The minimum RMSD between the generated 3D molecules and 100 potential conformations predicted by the RDKit toolkit was calculated. Next, the most common bonds, bond pairs, and bond triples were identified in the validation set. Bond lengths, bond angles, and dihedral angles were measured for both the generated molecules and the molecules in the validation set. The Jensen-Shannon (JS) divergence was used to quantify the difference in distribution between the generated molecules and the validation set.

[0226] The bond-related properties of the generated molecules were examined. Initially, the distribution of per-atom bond counts between the generated molecules and the validation set was compared to determine if the model produced an excessive or insufficient number of bonds. Next, the distribution of various bond types was analyzed, including the basic bond types (single, double, triple, and aromatic) as well as common bond types, bond pairs, and bond triplets used in 3D structure evaluation.

[0227] The bond analysis was extended to include rings. First, the distribution of ring counts for each molecule was compared between the generated molecules and the validation set. The distribution of counts of rings of various sizes (n-sized rings) was compared between the generated molecules and the validation set using the JS divergence, averaged for n∈{3,4,...,9}.

[0228] C.2 Conformation generation

[0229] The evaluation was performed on the GEOM dataset, which provides an ensemble of conformers generated using metadynamics in CREST. As evaluation metrics for conformer generation, a scheme was adopted that uses the average minimum RMSD (AMR) and coverage (COV) for precision (P) and recall (R) measured when generating twice as many conformers as provided by CREST. For K = 2L, let and {C k} k∈[1,K] denote the collections of ground truth and generated conformers respectively:

[0230]

[0231] Where δ is the coverage threshold. Following TorDiff, using The accuracy metric is obtained by exchanging the ground truth and the generated conformers.

[0232] C.3 Linker generation

[0233] The validity, uniqueness, recovery, RMSD and SCRDKit of the samples were calculated. Subsequently, it was determined whether the generated linkers obeyed the 2D filters utilized when generating the ZINC training set.

[0234] Validity - performs cleanup and additionally verifies that the molecule includes all atoms from the fragment. For all other metrics, only a subset of valid samples is considered.

[0235] Uniqueness - SMILES of the entire molecule are compared and the number of unique molecules sampled for each input fragment pair is counted.

[0236] Recovery - The SMILES of each molecule sampled for a given fragment pair is compared with the SMILES of the corresponding ground truth molecule. Before comparison, hydrogen atoms and stereochemistry are removed from the molecules.

[0237] RMSD - To evaluate the 3D conformation of the sampled molecules with respect to the ground truth molecules, the root mean square deviation (RMSD) between the coordinates of the generated linkers and the coordinates of the true molecules is determined in the case of accurate recovery of the original molecules. For calculating the RMSD, only the recovered molecules are considered and they are aligned with the corresponding ground truth molecules using the RDKit function rdkit.Chem.rdMolAlign, which returns the optimal RMSD used to align the two molecules.

[0238] The SCRDKit metric was utilized to measure the geometric and chemical similarity between the generated molecules and their ground truth counterparts.

[0239] 2D filters, including synthetic accessibility, ring aromaticity (RA), and pan-assay interference compound (PAINS), were applied to create the ZINC and CASF datasets. The RA filter ensures the correct covalent bond order in ring structures, while the PAINS filter detects compounds that are prone to producing false positive results in high-throughput screening.

[0240] C.4 Scaffold decoration

[0241] The same metrics as those for linker generation were adopted for scaffold decoration. Additionally, the Vina score and high affinity metrics were included.

[0242] The Vina score is calculated by the QVina software and is used to measure the binding affinity. QVina is designed to perform the same functions as AutoDock Vina - predicting the binding affinity and orientation of ligands to protein receptors - but is optimized to perform these calculations more rapidly. The Vina score serves as a metric to evaluate the binding affinity between ligands and receptors in docking simulations. It represents the predicted binding energy quantitatively, and a lower score indicates a stronger binding affinity. This score is used to assess the likelihood of successful ligand-receptor interactions and is particularly valuable in computational drug discovery and virtual screening efforts.

[0243] C.5 Shape - conditioned generation

[0244] To evaluate the shape-conditioned generation task, graph similarity (SimG) and shape similarity (SimS) metrics were utilized.

[0245] SimG is utilized to define the graph (chemical) similarity between two molecules, simG ∈ [0,1], as the Tanimoto similarity, which is calculated by RDKit using the default settings with 2048-bit fingerprints.

[0246] SimS is utilized to obtain the shape similarity metric from the ESP-Sim software package.

[0247] C.6 Pocket - conditioned generation

[0248] Vina docking is utilized for the pocket condition generation task. To calculate the Vina docking values, AutoDockVina is used, which is a widely adopted software for molecular docking research.

[0249] The high-affinity metric is used to measure the strength of ligand-receptor binding in molecular docking or biochemical assays compared to a reference ligand. It helps predict successful interactions between molecules. It takes into account factors such as the stability of the ligand-receptor complex, the strength of their interactions (such as hydrogen bonds), and the overall energy of the interaction. The Vina docking values are compared for the generated molecules and the reference molecules, and these docking values provide the percentage of cases where the generated molecules bind better than the reference molecules.

[0250] D Model parameters and training details

[0251] A model based on the nach0 architecture is utilized. The experiment involves a basic model variant with 370 million parameters, characterized by 12 layers, a 768-dimensional hidden state, a 2048-dimensional feed-forward hidden state, and 12 attention heads.

[0252] For the selected model, pre-training is performed using the language modeling (LM) objective and subsequent fine-tuning is carried out. The base model is trained using two NVIDIA A6000 GPUs. The probability of molecular fragment dropout in the pre-training phase is 0.2.

[0253] The pre-training and fine-tuning phases are executed using the following hyperparameters: the batch size for both pre-training and fine-tuning is 64, the learning rate is set to 1e-4, the weight decay is 0.01, and cosine scheduling is used. The pre-training phase lasts for 100,000 steps, while the fine-tuning phase also lasts for 100,000 steps.

[0254] E Additional related work

[0255] E.1 Linker generation

[0256] The following three models emerged as significant contributing factors to linker generation: DeLinker, 3DLinker, and DiffLinker.

[0257] DeLinker is a model suitable for linker design. It specifically preserves 3D structure information and generates linkers by providing two input fragments. This is one of the first attempts to apply graph neural networks (GNNs) to linker design.

[0258] 3DLinker is an E(3)-equivariant variational autoencoder for molecular linker design. It generates small "linkers" to physically attach two independent molecules with their different functions. The generation of linkers is conditioned on two given molecules, and the linker largely depends on the anchor atoms of the two molecules to be connected. It predicts the anchor atoms and jointly generates the linker graph and its 3D structure.

[0259] DiffLinker is an equivariant 3D conditional diffusion model for molecular linker design. Given a set of 3D disconnected fragments, DiffLinker places the missing atoms in between and designs a molecule containing all the initial fragments. Different from previous schemes that could only connect pairs of molecular fragments, DiffLinker can link any number of fragments.

[0260] E.2 LLaMa baseline

[0261] To establish a comparative single-domain baseline, the model and tokenizer were trained from scratch using a LLaMa-like architecture. For both cases, the implementation of Open LLaMa from the HuggingFace framework was adopted, and training was performed on the GEOM dataset.

[0262] In this study, the tokenizer was kept consistent with the Open LLaMa version, although it was only retrained on the most recent dataset. No improvement was made to the tokenization process for numbers or molecules, and the vocabulary size was limited to 512 tokens.

[0263] Regarding the model, modifications were applied to its configuration to align with the dimensions of nach0-pc. That is, the number of hidden layers and attention heads was set to 12, the hidden size was set to 768, and the intermediate size was set to 2048. The overall parameter count of the model reached nearly 90 million. During the training process, the global batch size was set to 32, the learning rate was set to 1e-4, a linear scheduler with warm-up steps was used, and the weight decay was set to 1e-2. The model was trained for 10 epochs using the causal language modeling objective.

[0264] E.3 Attention - based models for point clouds

[0265] Recently, several studies have been conducted on applying transformer architectures to point cloud analysis. Point Cloud Transformer (PCT)

[13] uses permutation invariant transformers to replace self-attention mechanisms for managing unstructured and disordered point data in irregular domains. Similarly, the transformer-based network (TR-Net) PointConT

[71] exploits the locality of points in the feature space by clustering sampled points with similar features into the same class and computing self-attention within each class. This design aims to capture long-range dependencies within the point cloud while maintaining computational efficiency.

[72] adopts a neighborhood embedding strategy and a residual backbone featuring skip connections to enhance context-aware and spatial-aware features. The network uses an offset attention operator on the point cloud spatial information to refine the attention weights, thus improving the extraction of global features.

[0266] E.4 Non - diffusion schemes

[0267] Several neural generative models have been proposed to generate spatial molecular structures, including generators that directly work with atomic density grids and voxels. A conditional variational autoencoder was trained on the atomic density grid representation of cross-docked protein-ligand structures

[73] . To construct valid molecular conformations from the generated atomic densities, atomic fitting and bond inference processes were utilized. Pocket2Mol

[64] is an E(3)-equivariant generative network consisting of two modules: 1) a graph neural network (GNN) that captures the spatial and bonding relationships between atoms in the binding pocket; and 2) an algorithm for sampling drug candidates based on the pocket representation from a tractable distribution. VoxMol

[74] samples noisy density grids from a smooth distribution using underdamped Langevin Markov chain Monte Carlo and denoises the noisy grids in a single step to refine the exact atomic positions. Different from point cloud diffusion models, VoxMol is simpler to train, does not require prior knowledge of the number of atoms, and does not treat features as distinct distributions.

[0268] F Computational resources

[0269] F.1 Hardware computational resources

[0270] The PC transformer (nach0-pc) model was utilized for various experiments, using state-of-the-art computational hardware to ensure efficient and effective training and evaluation. The hardware configuration of the CoreWeave cloud service provider includes:

[0271] GPU: Two NVIDIA RTX A6000 GPUs with 48GB of memory for PC Transformer model training; and two RTX A4000 GPUs with 16GB of memory for comparative model training. These GPUs are designed specifically for deep learning tasks, providing high throughput and large memory capacity, which are crucial for handling the large computational requirements of our models. CPU: AMD EPYC 7413 processors with 24 cores and a base clock speed of 2.65GHz, supporting multi-threaded operations and high parallelism. RAM: 128GB of DDR4 memory to support large batch processing and storage of large model parameters. Storage: High-speed NVMe SSD with a total capacity of 1TB to ensure fast data access and model checkpoints.

[0272] F.2 Training and evaluation time

[0273] Training and evaluating the PC Transformer model requires the following computational times:

[0274] Training time: The initial pre-training phase takes 40 hours. This phase includes preprocessing the dataset, training the model across multiple rounds, and hyperparameter tuning. Fine-tuning time: The fine-tuning phase takes approximately 60 hours. Evaluation time: The evaluation phase involves running inference, calculating performance metrics, and validating the results, taking an additional 6 hours (excluding ablation study sampling). Multiple runs were conducted to verify the consistency of the results and ensure robustness and accuracy.

[0275] In summary, deploying the PC Transformer model on advanced computing resources, while costly and time-consuming, is crucial for achieving high-performance results in our experiments. Investing in state-of-the-art hardware facilitates efficient training and rigorous evaluation, highlighting the importance of adequate resources in modern machine learning research.

[0276] G Point cloud encoder details and visualization

[0277] Figure 19 Visualization of the learned relative biases is illustrated according to an embodiment. The green line corresponds to a typical C-C single covalent bond. The yellow line corresponds to a typical hydrogen bond length. The red line corresponds to the typical distance between the resulting CA-CA atoms in a protein. As shown, the activation values of the relative biases are plotted as a function of the distance between atoms. On the X-axis, the distance is represented, ranging from 0 to The Y-axis depicts the corresponding activation values, ranging from -25 offset to +25 offset for the top row of plots and from -10 offset to +10 offset for the remaining rows of plots. This visualization is carried out using synthetic data, where the distances between atomic coordinates are systematically varied. The dashed lines represent typical atomic distances: the green dashed line represents the typical length of a C-C single covalent bond, the yellow dashed line corresponds to the typical hydrogen bond length, and the red dashed line represents the typical distance between consecutive CA-CA atoms in a protein. Notably, the activations in some of the attention heads within a particular layer peak exactly in regions corresponding to these typical bond lengths, highlighting the sensitivity of the model to biologically relevant atomic distances.

[0278] Figure 20 Illustrated is the t-SNE visualization 2100 of the text atom token embeddings 2110 and the text amino acid token embeddings 2120. These embeddings are projected into a 2D space where similar tokens are located based on their common chemical and structural properties. For the atom tokens 2110, the visualization highlights how atoms with similar properties or roles within a molecule are represented. Similarly, the amino acid token embeddings 2120 reflect their structural similarities, demonstrating the model's ability to capture and encode the underlying chemical and structural information within the text token embeddings.

[0279] H Analysis: Design choices of the PC Transformer model (nach0 - pc)

[0280] This section explores the role of the language model (LM) within the PC Transformer model. First, the LM component is selected, where the PC Transformer model is initialized with nach0, a top-tier chemical LM. Comparative analysis reveals that, across various tasks including distribution learning, conformation generation, linker design, and pocket-conditioned generation, the performance of the models leveraging nach0 is superior to those using other LMs. Notably, the nach0-based models consistently exhibit excellent performance, highlighting the effectiveness of domain-specific pre-training. This section provides more details on the design choices of the PC Transformer model for tasks of spatial molecular distribution learning (Table 6), conformation generation (Table 7), linker design (Table 8), shape-conditioned generation ( Figure 21 ) and pocket-conditioned generation (Table 9).

[0281] H.1 LM components

[0282] The first design choice for the PC transducer model is the LM component. For all experiments presented, the LM component of the PC transducer model was initialized with nach0, a state-of-the-art chemical LM for multi-domain tasks. In this section, several LMs are compared: (i) the state-of-the-art cross-domain nach0, (ii) the general-domain FLAN, and (iii) random initialization.

[0283] Regarding the distribution learning task, several observations can be made from Table 6.

[0284]

[0285]

[0286] Table 6

[0287] First, in terms of 3D structure, bond, and ring metrics, the single-task and multi-task PC transducer models with nach0 outperform other models. Second, the single-task model has the lowest JS divergence (0.146), followed by the multi-task model (0.205). The single-task model has the best performance (JS bond angle metric of 0.100), while the multi-task model is close behind (0.107). The FLAN-based model shows the highest divergence (0.247). Overall, the single-task model consistently achieves the lowest JS divergence across various metrics, indicating high accuracy in bond and ring structures. The PC transducer model with nach0 and the performance of the LM component based on randomness are both better than the general-domain LM component FLAN. This shows that in downstream tasks in single-task and multi-task settings, the PC transducer model with domain-specific nach0 provides a significant improvement over the model with general-domain FLAN.

[0288] Similar observations were made for the conformation generation task. Table 7 is a comparison of the model performance for conformation generation, indicating that including multi-task and domain-specific pre-training significantly enhances the quality metrics of the task of generating conformations.

[0289]

[0290] Table 7

[0291] This table evaluates the performance of the PC transducer model in multi-task and single-task settings with a nach0 backbone, as well as with FLAN and random LM components and without a pre-trained model.

[0292] In addition to the improvements observed in the task of generating conformations, Table 8 also highlights the great benefits of including multi-task and domain pre-training for the linker design task.

[0293]

[0294]

[0295] Table 8

[0296] Table 8 shows a comparison of the model performance on the linker design. The table evaluates the PC Transformer model in multi-task and single-task settings with a nach0 backbone, as well as the performance with FLAN and random LM components and without a pre-trained model. Similar to the enhancements seen in conformation generation, including multi-task learning and pre-training techniques significantly improved the quality metrics of the linker design.

[0297] Table 9 shows a comparison of the model performance on the pocket condition generation task.

[0298]

[0299] Table 9

[0300] The table evaluates the PC Transformer model in multi-task and single-task settings with a nach0 backbone, as well as the performance with FLAN and random LM components and without a pre-trained model. The better scores in all models are highlighted in bold.

[0301] When examining the results of the pocket condition generation task, corresponding results emerge. The integration of multi-task and pre-training methods brought about a significant improvement in the quality metrics.

[0302] Another way to compare the impact of the implemented enhancements is to analyze the dependence curve, which describes the relationship between the average Tanimoto similarity and the average shape similarity for various α values. For the same level of shape similarity, a robust model is expected to produce a lower Tanimoto similarity. Figure 21 Figure 2200 graphically compares the model performance on shape condition generation according to an embodiment. The table evaluates the PC Transformer model in multi-task and single-task settings with a nach0 backbone, as well as the performance with FLAN and random LM components and without a pre-trained model. Figure 2200 illustrates these curves for models with different ablation levels. This comparison shows that the performance of the multi-task model is inferior to that of the single-task model, as evidenced by its higher Tanimoto similarity for comparable shape similarities.

[0303] H.2 Impact of 3D pre - training

[0304] Another significant contribution of the PC Transformer model is its 3D pre-training scheme. Considering the three fundamental aspects of machine learning - data and task complexity - pre-training has an advantage when the amount of data available for downstream tasks is relatively small compared to the complexity of the task.

[0305] As shown in Table 6, in terms of 3D structure, bond, and ring metrics, the pre-trained and fine-tuned PC Transformer model provides a significant improvement compared to the fine-tuned PC Transformer model without pre-training. Surprisingly, in the conformational generation task, model pre-training achieves better results than the single-task PC Transformer model. Regarding the linker design task, similar observations can be made as in the distribution learning task: 3D pre-training helps the model perform downstream tasks.

[0306] H.3 Classical and non - isomorphic SMILES

[0307] The PC Transformer model enhances isomeric SMILES including stereotags. This enables the model to utilize some 3D information present in the stereotags of isomeric SMILES and better generalize through enhancement. An ablation study was conducted to compare this option with the non-enhanced version of isomeric SMILES (referred to as classical SMILES in Table 7) and non-isomeric SMILES without enhancement. Table 7 shows that the main multi-task PC Transformer model performs better in conformational generation than the classical isomeric SMILES scheme without enhancement but is comparable to the non-isomeric SMILES scheme without enhancement. The results indicate that enhancement improves performance. However, for unconditional conformational generation with a relatively large and diverse dataset, the non-isomeric SMILES option may also be an alternative.

[0308] I Model training time and CO2 impact

[0309] Table 10 shows the GPU computation time and CO2 emissions of the PC Transformer model and the state-of-the-art diffusion model.

[0310]

[0311]

[0312] Table 10

[0313] In Table 10, the timings (marked with *) were extracted from the reference documents. Single-task fine-tuning was excluded from the total time and total CO2 emissions of the PC Transformer model and is only present here for the purpose of direct comparison with the single-task model.

[0314] As shown in Table 10, the computational resources required for the pre-training phase of the PC Transformer model are quantified in GPU hours (total training duration). This covers pre-training on the dataset, fine-tuning of the model on all datasets, and generating molecules using the model.

[0315] In particular, the fine-tuned models originating from the pre-trained checkpoints exhibit enhanced performance while consuming only about half of the computational resources required by models trained from scratch. This efficient resource allocation approach validates the feasibility in implementing the PC Transformer model technology. Additionally, it is fully necessary to perform the initial resource allocation for model pre-training because these models can be reused in many applications once trained, thus enhancing their practicality and cost-effectiveness. All experiments were executed using the CoreWeave infrastructure. For the above PC Transformer model training and evaluation, a cumulative computational time of 164.5 hours was executed on Nvidia RTX A6000 48GB hardware (with a TDP of 300W). The total emissions estimate for this model is 20.73 kgCO2eq. When comparing the PC Transformer model with the other SOTA diffusion models discussed, it also took 1020 hours on Nvidia RTX A4000 16GB (with a TDP of 140W) and had emissions of 9436 kgCO2eq. These estimates were obtained using the Machine Learning Impact Calculator.

[0316] The PC Transformer model is multi-task and was trained sequentially for all six tasks. As shown in Table 10, when directly compared with other single-task models, the PC Transformer model shows effective results in terms of computational time and CO2 emissions. Based on the total time and CO2 emissions divided by the number of tasks, the PC Transformer model is 3.5 times more efficient than the second most efficient single-task model while maintaining acceptable quality across all six tasks.

[0317] The following datasets were used: 1) GEOM dataset, 2) MOSES dataset, 3) ZINC dataset, and CrossDocked dataset. The following models were used: 1) nach0, and 2) FLAN model. To perform the ablation study, the following model architectures were used: 1) OpenLLaMa source, and 2) MolDiff model source code.

[0318] The following incorporated references are cross-referenced: US11,568,961; US11,403,521; US 2015 / 0178442; US2020 / 0090049; US2020 / 0082916; US2020 / 0258594; US2022 / 0310196; US2021 / 0233621; US2021 / 0271980; US2021 / 0287067; US2021 / 0383898; US2022 / 0172802; US2022 / 0406404; EP 3289501; WO 2021 / 165887; and WO 2021 / 229454.

[0319] This document incorporates the following references by specific reference.

[0320] References

[0321] [1] Jacob Devlin, et al.;BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics.

[0322] [2] Colin Raffel, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140): 1–67, 2020.

[0323] [3] Mike Lewis, et al.; BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online, July 2020. Association for Computational Linguistics.

[0324] [4] Zhangyin Feng, et al.; CodeBERT: A pre-trained model for programming and natural languages. In Trevor Cohn, Yulan He, and Yang Liu, editors, Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1536–1547, Online, November 2020. Association for Computational Linguistics.

[0325] [5] Tom Brown, et al.; Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc., 2020.

[0326] [6]Jean-Baptiste Alayrac,et al.;Flamingo:a visual language modelforfew-shot learning.In S.Koyejo,S.Mohamed,A.Agarwal,D.Belgrave,K.Cho,andA.Oh,editors,Advances in Neural Information Processing Systems,volume35,pages23716–23736.Curran Associates,Inc.,2022.

[0327] [7]Jing Yu Koh,Ruslan Salakhutdinov,and DanielFried.Groundinglanguage models to images for multimodal inputs and outputs.InAndreasKrause,Emma Brunskill,Kyunghyun Cho,Barbara Engelhardt,Sivan Sabato,and Jonathan Scarlett,editors,Proceedings of the 40th InternationalConferenceon Machine Learning,volume 202 of Proceedings of MachineLearningResearch,pages 17283–17300.PMLR,23–29 Jul 2023.

[0328] [8]Jiasen Lu,Dhruv Batra,Devi Parikh,and Stefan Lee.ViLBERT:Pretraining task-agnostic visiolinguistic representations for vision-and-languagetasks.In H.Wallach,H.Larochelle,Beygelzimer,F.d'Alché-Buc,E.Fox,andR.Garnett,editors,Advances in Neural Information Processing Systems,volume32.Curran Associates,Inc.,2019.

[0329] [9]Chen Sun,Austin Myers,Carl Vondrick,Kevin Murphy,andCordeliaSchmid.VideoBERT:A joint model for video and languagerepresentationlearning.In 2019 IEEE / CVF International Conference onComputer Vision(ICCV),pages 7463–7472,Los Alamitos,CA,USA,nov 2019.IEEE Computer Society.

[0330]

[10] Yi Ren,et al.;FastSpeech:Fast,robust and controllable texttospeech.In H.Wallach,H.Larochelle,A.Beygelzimer,F.d'Alché-Buc,E.Fox,andR.Garnett,editors,Advances in Neural Information Processing Systems,volume32.Curran Associates,Inc.,2019.

[0331]

[11] Alec Radford,et al.;Robust speech recognition via large-scaleweak supervision.In Andreas Krause,Emma Brunskill,Kyunghyun Cho,BarbaraEngelhardt,Sivan Sabato,and Jonathan Scarlett,editors,Proceedings ofthe 40thInternational Conference on Machine Learning,volume 202 ofProceedings ofMachine Learning Research,pages 28492–28518.PMLR,23–29 Jul 2023.

[0332]

[12] Hengshuang Zhao,et al.;.Point Transformer.In 2021 IEEE / CVFInternational Conference on Computer Vision(ICCV),pages 16239–16248,2021.

[0333]

[13] Meng-Hao Guo,et al.;.PCT:Point cloud transformer.ComputationalVisual Media,7(2):187–199,Jun 2021.

[0334]

[14] Xumin Yu,et al.;.Point-BERT:Pre-training 3D pointcloudtransformers with masked point modeling.In 2022 IEEE / CVF ConferenceonComputer Vision and Pattern Recognition(CVPR),pages 19291–19300,2022.

[0335]

[15] Daniel Flam-Shepherd, Kevin Zhu, and Alán Aspuru-Guzik. Language models can learn complex molecular distributions. Nature Communications, 13(1):3293, Jun 2022.

[0336]

[16] Carl Edwards, et al.;Translation between molecules and natural language. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 375–413, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics.

[0337]

[17] Qizhi Pei, et al.;BioT5: Enriching cross-modal integration in biology with chemical knowledge and natural language associations. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 1102–1123, Singapore, December 2023. Association for Computational Linguistics.

[0338]

[18] Dimitrios Christofidellis, et al.;. Unifying molecular and textual representations via multi-task language modelling. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 6140–6157. PMLR, 23–29 Jul 2023.

[0339]

[19] Micha Livne, et al.;. nach0: multimodal natural and chemical languages foundation model. Chem. Sci., pages–, 2024.

[0340]

[20] Mario Krenn, et al.; Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1(4): 045024, oct 2020.

[0341]

[21] Ross Irwin, et al.; Chemformer: a pretrained transformer for computational chemistry. Machine Learning: Science and Technology, 3(1): 015022, jan 2022.

[0342]

[22] Daniel Flam-Shepherd and Alán Aspuru-Guzik. Language models can generate molecules, materials, and protein binding sites directly in three dimensions as XYZ, CIF, and PDB files, 2023.

[0343]

[23] Noel M. O’Boyle, et al.;Open Babel: An open chemical toolbox. Journal of Cheminformatics, 3(1):33, Oct 2011.

[0344]

[24] Seyone Chithrananda, et al.;ChemBERTa: large-scale self-supervised pre-training for molecular property prediction. arXiv preprint arXiv:2010.09885, 2020.

[0345]

[25] Youwei Liang, et al.;Drugchat: Towards enabling chatgpt-like capabilities on drug molecule graphs. TechRxiv, 2023.

[0346]

[26] Shengjie Luo, et al.;One transformer can understand both 2d&3d molecular data. In The Eleventh International Conference on Learning Representations, 2022.

[0347]

[27] Xiangru Tang, et al.;Mollm: A unified language model to integrate biomedical text with 2d and 3d molecular representations. Bioinformatics, 2024.

[0348]

[28] Jascha Sohl-Dickstein,et al.;Deep unsupervised learningusingnonequilibrium thermodynamics.In Francis Bach and David Blei,editors,Proceedings of the 32nd International Conference on Machine Learning,volume37 of Proceedings of Machine Learning Research,pages 2256–2265,Lille,France,07–09 Jul 2015.PMLR.

[0349]

[29] Jonathan Ho,Ajay Jain,and Pieter Abbeel.Denoisingdiffusionprobabilistic models.In H.Larochelle,M.Ranzato,R.Hadsell,M.F.Balcan,and H.Lin,editors,Advances in Neural Information Processing Systems,volume33,pages 6840–6851.Curran Associates,Inc.,2020.

[0350]

[30] Emiel Hoogeboom,et al.;Equivariant diffusion formoleculegeneration in 3D.In Kamalika Chaudhuri,Stefanie Jegelka,Le Song,CsabaSzepesvari,Gang Niu,and Sivan Sabato,editors,Proceedings of the39thInternational Conference on Machine Learning,volume 162 of ProceedingsofMachine Learning Research,pages 8867–8887.PMLR,17–23 Jul 2022.

[0351]

[31] Xingang Peng,et al.;MolDiff:Addressing the atom-bondinconsistencyproblem in 3D molecule diffusion generation.In Andreas Krause,Emma Brunskill,Kyunghyun Cho,Barbara Engelhardt,Sivan Sabato,andJonathan Scarlett,editors,Proceedings of the 40th International Conference onMachine Learning,volume202 of Proceedings of Machine Learning Research,pages 27611–27629.PMLR,23–29Jul 2023.

[0352]

[32] Lei Huang,et al.;MDM:Molecular diffusion model for 3Dmoleculegeneration.Proceedings of the AAAI Conference on ArtificialIntelligence,37(4):5105–5112,Jun.2023.

[0353]

[33] Kristof Schütt,et al.;SchNet:A continuous-filterconvolutionalneural network for modeling quantum interactions.In I.Guyon,U.VonLuxburg,S.Bengio,H.Wallach,R.Fergus,S.Vishwanathan,and R.Garnett,editors,Advances in Neural Information Processing Systems,volume30.CurranAssociates,Inc.,2017.Minkai Xu,Lantao Yu,Yang Song,Chence Shi,StefanoErmon,and Jian Tang.GeoDiff:A geometric diffusion model formolecularconformation generation.In International Conference onLearningRepresentations,2022.

[0354]

[34] Bowen Jing,et al.;Torsional diffusion for molecularconformergeneration.In S.Koyejo,S.Mohamed,A.Agarwal,D.Belgrave,K.Cho,andA.Oh,editors,Advances in Neural Information Processing Systems,volume 35,pages 24240–24253.Curran Associates,Inc.,2022.

[0355]

[35] Octavian Ganea,et al.;GeoMol:Torsional geometric generationofmolecular 3d conformer ensembles.In M.Ranzato,A.Beygelzimer,Y.Dauphin,P.S.Liang,and J.Wortman Vaughan,editors,Advances in NeuralInformationProcessing Systems,volume 34,pages 13757–13769.Curran Associates,Inc.,2021.

[0356]

[36] Ilia Igashov,et al.;Equivariant 3d-conditional diffusion modelformolecular linker design.Nature Machine Intelligence,Apr 2024.

[0357]

[37] Jiaqi Guan,et al.;LinkerNet:Fragment poses and linker co-designwith 3D equivariant diffusion.In A.Oh,T.Neumann,Globerson,K.Saenko,M.Hardt,and S.Levine,editors,Advances in Neural InformationProcessingSystems,volume 36,pages 77503–77519.Curran Associates,Inc.,2023.

[0358]

[38] Ziqi Chen,et al.;Shape-conditioned 3D molecule generationviaequivariant diffusion models.In NeurIPS 2023 Generative AI and Biology(GenBio)Workshop,2023.

[0359]

[39] Haitao Lin,et al.;Functional-group-based diffusion for pocket-specific molecule generation and elaboration.In A.Oh,T.Neumann,A.Globerson,K.Saenko,M.Hardt,and S.Levine,editors,Advances in NeuralInformationProcessing Systems,volume 36,pages 34603–34626.CurranAssociates,Inc.,2023.

[0360]

[40] Jiaqi Guan,et al.;3D equivariant diffusion for target-awaremolecule generation and affinity prediction.In The EleventhInternationalConference on Learning Representations,2023.

[0361]

[41] Jiaqi Guan,et al.;DecompDiff:Diffusion models withdecomposedpriors for structurebased drug design.In Andreas Krause,EmmaBrunskill,Kyunghyun Cho,Barbara Engelhardt,Sivan Sabato,and JonathanScarlett,editors,Proceedings of the 40th International Conference on MachineLearning,volume 202 of Proceedings of Machine Learning Research,pages 11827–11846.PMLR,23–29 Jul 2023.

[0362]

[42] Ashish Vaswani,et al.;Attention is all you need.In I.Guyon,U.VonLuxburg,S.Bengio,H.Wallach,R.Fergus,S.Vishwanathan,and R.Garnett,editors,Advances in Neural Information Processing Systems,volume30.Curran Associates,Inc.,2017.

[0363]

[43] Degen,et al.;On the art of compiling and using’drug-like’chemical fragment spaces.ChemMedChem,3(10):1503–1507,2008.

[0364]

[44] Simon Axelrod and Rafael Gómez-Bombarelli.GEOM,energy-annotatedmolecular conformations for property prediction andmoleculargeneration.Scientific Data,9(1):185,Apr 2022.

[0365]

[45] Hugo Touvron,et al.;LLaMa:Open and efficient foundationlanguagemodels,2023.

[0366]

[46] Fergus Imrie,et al.;Deep generative models for 3D linkerdesign.Journal of Chemical Information and Modeling,60(4):1983–1995,2020.PMID:32195587.

[0367]

[47] Yinan Huang,et al.;3DLinker:An e(3)equivariant variational autoencoder for molecular linker design.In Kamalika Chaudhuri,Stefanie Jegelka,Le Song,Csaba Szepesvari,Gang Niu,and Sivan Sabato,editors,Proceedings of the 39th International Conference on Machine Learning,volume 162 of Proceedings of Machine Learning Research,pages 9280–9294.PMLR,17–23 Jul 2022.

[0368]

[48] Zaixi Zhang,et al.;Molecule generation for target protein binding with structural motifs.In The Eleventh International Conference on Learning Representations,2023.

[0369]

[49] Keir Adams and Connor W.Coley.Equivariant shape-conditioned generation of 3D molecules for ligand-based drug design.In The Eleventh International Conference on Learning Representations,2023.

[0370]

[50] Daniil Polykovskiy,et al.;Molecular Sets(MOSES):A benchmarking platform for molecular generation models.Frontiers in Pharmacology,11,2020.

[0371]

[51] Shitong Luo,et al.;A 3D generative model for structure-based drug design.In M.Ranzato,et al.editors;Processing Systems,volume 34,pages 6229–6239.Curran Associates,Inc.,2021.

[0372]

[52] Xingang Peng,et al.;Pocket2Mol:Efficient molecular sampling based on 3D protein pockets.In Kamalika Chaudhuri,Stefanie Jegelka,Le Song,Csaba Szepesvari,Gang Niu,and Sivan Sabato,editors,Proceedings of the 39th International Conference on Machine Learning,volume 162 of Proceedings of Machine Learning Research,pages 17644–17655.PMLR,17–23 Jul 2022.

[0373]

[53] Peter Ertl and Ansgar Schuffenhauer.Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions.Journal of Cheminformatics,1(1):8,Jun 2009.

[0374]

[54] Amr Alhossary,et al.;Fast,accurate,and reliable molecular docking with QuickVina 2.Bioinformatics,31(13):2214–2216,02 2015.

[0375]

[55] Giovanni Bolcato,et al.;On the value of using 3D shapeandelectrostatic similarities in deep generative methods.Journal ofChemicalInformation and Modeling,62(6):1388–1398,2022.PMID:35271260.

[0376]

[56] Luyao Liu,et al.;TR-Net:a transformer-based neural networkforpoint cloud processing.Machines,10(7):517,2022.

[0377]

[57] Matthew Ragoza,et al.;Generating 3D molecules conditionalonreceptor binding sites with deep generative models.Chem.Sci.,13:2701–2713,2022.

[0378]

[58] Pedro O.Pinheiro,et al.;3D molecule generation by denoisingvoxelgrids.CoRR,abs / 2306.07473,2023.

[0379]

[59] Hyung Won Chung,et al.;Scaling instruction-finetunedlanguagemodels,2022.

[0380]

[60] Alexandre Lacoste,et al.;Quantifying the carbon emissionsofmachine learning.arXiv preprint arXiv:1910.09700,2019.

Claims

1. A transformer model architecture, comprising: Point cloud input module; Text input module; a point cloud encoder module operatively coupled to the point cloud input module; a large language model module operatively coupled to the text input module and the point cloud encoder module and configured to receive data therefrom; as well as A text output module is operatively coupled to the large language model module and configured to output the molecule data in a line representation format.

2. The transformer model architecture of claim 1, wherein: The point cloud encoder module aggregates spatial position information from a point cloud.

3. The transformer model architecture of claim 1, wherein: The point cloud encoder module includes a graph neural network configured to process relative distances between points of a point cloud.

4. The transformer model architecture of claim 3, further comprising a plurality of graph neural network layers.

5. The transformer model architecture of claim 4, wherein: Each graph neural network layer aggregates information from connected nodes and edges and processes global information from the entire graph.

6. The transformer model architecture of claim 3, wherein: The graph neural network includes an attention mechanism to: calculating an attention bias related to relative positions and edge features between points of the point cloud; and Update the point embedding based on the computed attention bias.

7. The transformer model architecture of claim 1, wherein: The point cloud encoder module is trained in an unsupervised manner using 3D molecular data.

8. The transformer model architecture of claim 1, wherein: The point cloud input module is configured to receive 3D molecular data comprising a point cloud.

9. The transformer model architecture of claim 8, wherein: The 3D molecular data includes, for each point of the point cloud, spatial position data and data representing one or more molecular features.

10. The transformer model architecture of claim 9, wherein: The one or more molecular features include one or more of the following: atomic symbols, atomic charges, atom names, or corresponding amino acid names.

11. The transformer model architecture of claim 8, wherein: The 3D molecular data includes data in at least one of the following chemical language formats: Simplified Molecular Input Line Entry System (SMILES) format, Self-Referencing Embedded String (SELFIES) format, or XYZ format.

12. The transformer model architecture of claim 8, wherein: The 3D molecular data represent large ligand or protein pocket structures.

13. The transformer model architecture of claim 8, wherein: The 3D molecular data is downsampled based on one or more prioritized points of the point cloud.

14. The transformer model architecture of claim 13, wherein: The one or more priority points of the point cloud include at least one of: a ligand, an alpha carbon (C-α) atom, or a terminal atom of an amino acid of a protein.

15. The transformer model architecture of claim 1, wherein: The point cloud encoder module is configured as: inferring a description of the point based on one or more points in a neighborhood of the point; and A three-dimensional (3D) position of the point relative to one or more points in a neighborhood of the point is determined.

16. The transformer model architecture of claim 15, wherein: The point and one or more points in the neighborhood of the point are not connected by a direct edge.

17. The transformer model architecture of claim 1, wherein: The large language model module and the point cloud encoder module are combined into a single model.

18. A method for pre-training a transformer model, the method comprising: providing a transformer model having a point cloud input module, a text input module, a point cloud encoder module operatively coupled to the point cloud input module, a large language model module operatively coupled to the text input module and the point cloud encoder module and configured to receive data therefrom, and a text output module configured to output molecule data in a row representation format; Feeding 3D molecular data including point clouds into the point cloud encoder module in an unsupervised manner; Inferring, by the point cloud encoder module, a point description for the point based on one or more points in a neighborhood of the point; determining a 3D position of the point relative to one or more points in a neighborhood of the point; masking or blurring at least some of the one or more points; predicting, by the point cloud encoder module, at least some of the masked or blurred point features; Sampling random points; and The distances between the random points are predicted by the point cloud encoder module.

19. The method of claim 18, further comprising: determining a masking or blurring loss value based on a prediction of the masked or blurred point features; determining a distance loss value based on the prediction of the distance between the random points; and A weighted sum of the masking or blurring loss and the distance loss is minimized based on the masking loss value and the distance loss value, respectively.

20. The method of claim 18, further comprising: encoding the point cloud into one or more point embeddings; Prepare embeddings for input text sequences; combining one or more points of the point embedding and the input text sequence embedding to obtain a fused input; and The fused input is fed into the point cloud encoder module.

21. A method for generating shape conditions, comprising: A transformer model is trained to recover a molecule from a blurred region of a 3D space in which the molecule is located in a molecular point cloud, wherein some portions of the molecular point cloud are blurred and selected portions of the molecule are not altered and are not blurred, wherein the transformer model architecture comprises: Point cloud input module; Text input module; a point cloud encoder module operatively coupled to the point cloud input module; a large language model module operatively coupled to the text input module and the point cloud encoder module and configured to receive data therefrom; and a text output module configured to output molecular data in a line representation format; Inputting 3D molecular data including point clouds via the point cloud input module; inputting text representing a desired molecule via the text input module; processing the 3D molecular data and text by the point cloud encoder and large language model modules using the trained point cloud encoder module; and The text output module outputs text representing the molecule in line representation.

22. A method for designing and generating a linker, comprising: Training a transformer model to recover the removed portion of a molecule when the transformer model does not receive a spatial description of the removed portion of the molecule, wherein the transformer model architecture comprises: Point cloud input module; Text input module; a point cloud encoder module operatively coupled to the point cloud input module; a large language model module operatively coupled to the text input module and the point cloud encoder module and configured to receive data therefrom; and a text output module configured to output molecular data in a line representation format; inputting data representing one or more molecules into the trained transformer model; generating one or more options for the linker portion of the molecule to replace the removed portion of the molecule using the trained transformer model; and One or more molecules are obtained having one or more options generated for the linker portion of the molecule.

23. One or more non-transitory computer-readable media storing instructions that, in response to being executed by one or more processors, cause a computer system to perform operations comprising: A transformer model is provided, wherein the transformer model has a point cloud input module; a text input module; a point cloud encoder module operatively coupled to the point cloud input module; a large language model module operatively coupled to the text input module and the point cloud encoder module and configured to receive data therefrom; and a text output module configured to output the molecule data in a line representation format; pre-training the point cloud encoder module in an unsupervised manner using 3D molecular data including point clouds; Inferring a point description based on one or more points in the point's neighborhood by a pre-trained point cloud encoder module; determining a 3D position of the point relative to one or more points in a neighborhood of the point; masking or blurring at least some of the one or more points; predicting at least some of the masked or blurred point features by a pre-trained point cloud encoder module; Sampling random points; and The distances between the random points are predicted by a pre-trained point cloud encoder module.

24. The computer readable medium of claim 23, wherein: The operations also include: For shape condition generation or linker design generation of at least one molecule, the transformer model or its pre-training or its implementation is operated.

25. A computer system comprising: one or more processors; as well as One or more non-transitory computer-readable media storing instructions that, in response to being executed by the one or more processors, cause the computer system to perform operations, the operations comprising: Pre-training the point cloud encoder module in an unsupervised manner using 3D molecular data including point clouds; Inferring a point description based on one or more points in the point's neighborhood by a pre-trained point cloud encoder module; determining a 3D position of the point relative to one or more points in a neighborhood of the point; masking or blurring at least some of the one or more points; predicting at least some of the masked or blurred point features by a pre-trained point cloud encoder module; Sampling random points; and The distances between the random points are predicted by a pre-trained point cloud encoder module.

26. The computer system of claim 18, further comprising an operation to: For shape condition generation or linker design generation of at least one molecule, the transformer model or its pre-training or its implementation is operated.

Citation Information

Patent Citations

  • Physics-based computational methods for predicting compound solubility

    EP3289501A2

  • Mutual information adversarial autoencoder

    US11403521B2

  • System and method for accelerating FEP methods using a 3D-restricted variational autoencoder

    US11568961B2

  • Methods and systems for calculating free energy differences using a modified bond stretch potential

    US20150178442A1

  • Entangled conditional adversarial autoencoder for drug discovery

    US20200082916A1