Generation of molecular structures
Patent Information
- Application Number
- PCT/EP2026/057792
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-19
- Publication Date
- 2026-10-01
Smart Images

Figure IMGF000023_0001 
Figure IMGF000024_0001 
Figure IMGF000024_0002
Abstract
Description
[0001] R. 417936
[0002] - 1 -
[0003] Description
[0004] Title
[0005] GENERATION OF MOLECULAR STRUCTURES
[0006] Technical Field
[0007] The presently disclosed subject matter relates to a method for generating molecular structures, a method for training a neural network to transform noisy spatial functions into structured spatial density functions, a computer-readable medium, and a system.
[0008] Background
[0009] Geometric graphs are typically processed using Graph Neural Networks (GN Ns) for inference or generative tasks. In a generative setting, this implies that we need to specify the number of nodes in the graph beforehand, thus conditioning the network on this information.
[0010] Graphs are common structures that have found extensive applications in various domains, including computer science, social networks, and biological systems. Traditionally, graphs are represented as discrete objects, comprising of a finite set of nodes and edges connecting them. In this formalism, the existence of nodes and edges is binary - they are either present or not. In this work, we propose a novel perspective on graphs, treating them as continuous objects embedded in a high-dimensional space, typically Euclidean space. By extending the notion of node existence to a continuous domain, we can model the uncertainty and probabilistic nature of many real-world systems. This continuous graph representation, for example, is particularly relevant in the context of molecular modeling, where particles, such as electrons, inherently exhibitR. 417936
[0011] - 2 -
[0012] probabilistic behavior and do not have precisely defined positions due to quantum effects.
[0013] There is a need for a more versatile method to generate molecular structures.
[0014] Summary
[0015] It would be advantageous to have an improved method for generating molecular structures. For example, a generated molecular structure may comprise a plurality of molecular substructure positions. Typically, a molecular structure is a molecule, and molecular substructures are atoms; however other examples are provided herein.
[0016] The method may comprise progressively denoising a spatial function by iteratively applying a trained neural network. The neural network is trained to transform noisy density functions into structured spatial density functions based on training data comprising molecular structures represented as spatial density functions. Initially, the spatial function may be generated as a noisy spatial function; however, during denoising, the spatial function is transformed into a spatial density function. The integral over the spatial density function after denoising may be used to obtain an integer value representing the number of molecular substructures.
[0017] Note that the number of molecular substructures need not be defined before or during the denoising process. Indeed, the number of molecular substructures may increase or decrease as denoising progresses. After denoising the resulting spatial density function may be discretized by fitting a sum of predefined density distributions. Typically, the predefined density distributions are Gaussians, but other examples are given herein. The centers of the fitted density distributions, e.g., their means, may be taken as positions of the molecular substructures of the molecular structure.R. 417936
[0018] - 3 -
[0019] Interestingly, the process allows representing an arbitrary number of molecular substructure positions, which may increase and / or decrease during the denoising.
[0020] An effective way to represent a spatial function as a data structure is to represent it as a series of positions of a plurality of points and values of the spatial function. For example, the trained neural network may be configured to modify the positions and / or values of the plurality of points.
[0021] A molecular structure generating system and / or a system for training the neural network therefor is an electronic device or system of electronic devices, e.g., one or more computers.
[0022] The molecular generation method described herein may be applied in a wide range of practical applications. Such practical applications include: generating molecular structures for exploration, e.g., finding novel and useful molecular structures, e.g., for pharmacological or other chemical applications; inferring the configuration of a molecule from chemical measurements; and debugging, testing, and / or benchmarking chemical software.
[0023] An aspect is a method for generation and / or training. An embodiment of the method may be implemented on a computer as a computer-implemented method, in dedicated hardware, or in a combination of both. Executable code for an embodiment of the method may be stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Preferably, the computer program product comprises non-transitory program code stored on a computer-readable medium for performing an embodiment of the method when said program product is executed on a computer. For example, a computer-readable medium may be a computer-readable storage medium.
[0024] In an embodiment, the computer program comprises computer program code adapted to perform all or part of the steps of an embodiment of the method whenR. 417936
[0025] - 4 -
[0026] the computer program is run on a computer. Preferably, the computer program is embodied on a computer-readable medium.
[0027] Brief Description
[0028] Further details, aspects, and embodiments will be described, by way of example only, with reference to the drawings. Elements in the figures are illustrated for simplicity and clarity and have not necessarily been drawn to scale. In the figures, elements which correspond to elements already described may have the same reference numerals. In the drawings,
[0029] Figure 1 schematically shows an example of an embodiment of molecular structure system,
[0030] Figure 2a schematically shows an example of an embodiment of a molecular structure generator system,
[0031] Figure 2b schematically shows an example of an embodiment of a denoising training system,
[0032] Figure 3a schematically shows an example of an embodiment of generation of a molecular structure,
[0033] Figure 3b schematically shows an example of an embodiment of an initial noisy spatial density function,
[0034] Figure 3c schematically shows an example of an embodiment of intermediate molecular structure,
[0035] Figure 3d schematically shows an example of an embodiment of a molecular structure generated before discretization,R. 417936
[0036] - 5 -
[0037] Figure 4 schematically shows an example of an embodiment of a continuous graph representing a molecular structure,
[0038] Figure 5a schematically shows an example of an embodiment of a continuous graph representing a molecular structure,
[0039] Figures 5b and 5c schematically show an example of an embodiment of a disturbed continuous graph,
[0040] Figure 6 schematically shows an example of an embodiment of a method for generating molecular structures,
[0041] Figure 7a schematically shows a computer-readable medium having a writable part comprising a computer program according to an embodiment,
[0042] Figure 7b schematically shows a representation of a processor system according to an embodiment.
[0043] Reference signs list
[0044] The following list of references and abbreviations corresponds to Figures 1 -5c, and is provided for facilitating the interpretation of the drawings and shall not be construed as limiting the claims.
[0045] 100 a molecular structure system
[0046] 110 molecular structure generator computer
[0047] 120 denoising training computer
[0048] 130 continuous graph inference computer
[0049] 111.121.131 a processor system
[0050] 112.122.132 storage
[0051] 113. 123. 133 communication interface
[0052] 200 a molecular structure generator system
[0053] 210 initial spatial function generatorR. 417936
[0054] 211 initial spatial function
[0055] 220 denoising network
[0056] 230 integrator
[0057] 231 molecular substructure number
[0058] 240 fitting unit
[0059] 241 first molecular structure
[0060] 250 connectivity inferring unit
[0061] 251 second molecular structure
[0062] 270 molecular structure training data
[0063] 280 density function conversion unit
[0064] 281 density function representation
[0065] 290 diffusor
[0066] 291 -293 diffusion training data
[0067] 1000, 1001 a computer-readable medium
[0068] 1010 a writable part
[0069] 1020 a computer program
[0070] 1110 integrated circuit(s)
[0071] 1120 a processing unit
[0072] 1122 a memory
[0073] 1124 a dedicated integrated circuit
[0074] 1126 a communication element
[0075] 1130 an interconnect
[0076] 1140 a processor system
[0077] Detailed Description
[0078] While the presently disclosed subject matter is susceptible to embodiment in many different forms, there are shown in the drawings and will herein be described in detail one or more specific embodiments, with the understanding that the present disclosure is to be considered as exemplary of the principles of the presently disclosed subject matter and not intended to limit it to the specific embodiments shown and described.R. 417936
[0079] - 7 -
[0080] In the following, for the sake of understanding, elements of embodiments are described in operation. However, it will be apparent that the respective elements are arranged to perform the functions being described as performed by them.
[0081] Further, the subject matter that is presently disclosed is not limited to the embodiments only but also includes every other combination of features described herein or recited in mutually different dependent claims.
[0082] Some embodiments are directed to generating molecular structures, which may include progressively denoising a spatial density function by iteratively applying a trained neural network. The neural network is trained to transform noisy spatial density functions into structured spatial density functions based on training data comprising molecular structures represented as spatial density functions. After denoising, a sum of predefined density distributions is fitted to the denoising result to obtain the molecular structure.
[0083] Figure 1 schematically shows an example of an embodiment of a molecular structure generator computer 110, an embodiment of a denoising training computer 120, and an embodiment of a continuous graph inference computer 130.
[0084] Shown is a molecular structure generator computer 110, denoising training computer 120, and continuous graph inference computer 130. Molecular structure generator computer 110 and denoising training computer 120 may be comprised in a molecular structure system 100. Optionally, molecular structure system 100 may also comprise continuous graph inference computer 130.
[0085] Molecular structure generator computer 110 is configured to generate molecular structures. A generated molecular structure comprises at least a plurality of molecular substructure positions, e.g., atom positions, amino acid positions, chemical group positions, etc. A generated molecular structure may comprise additional information, e.g., bonding information between the molecular substructure positions, atom types, chemical and other information, as furtherR. 417936
[0086] - 8 -
[0087] explained herein. Generating molecular structures is useful for a variety of applications, ranging from technical debugging of other software applications to compound discovery.
[0088] Molecular structure generator computer 110 uses a neural network trained to progressively denoise a spatial density function. Denoising training computer 120 may be used to train that neural network.
[0089] Intermediate data structures of molecular structure generator computer 110 use a continuous representation of a graph structure. In the output of molecular structure generator computer 110, this data structure may be discretized and / or transformed into traditional representations of molecular structures, e.g., a graph or a sequence of substructure positions, e.g., atom positions. Computations, e.g., estimating chemical properties of the generated molecular structure, may then be performed by using traditional graph computation algorithms, in particular graph neural networks. Interestingly, the inventors found that neural network computations may be performed directly on continuous graph representations.
[0090] Molecular structure generator computer 110 may comprise a processor system 111 , a storage 112, and a communication interface 113. Denoising training computer 120 may comprise a processor system 121 , a storage 122, and a communication interface 123. Continuous graph inference computer 130 may comprise a processor system 131, a storage 132, and a communication interface 133.
[0091] In the various embodiments of communication interfaces 113, 123, and / or 133, the communication interfaces may be selected from various alternatives. For example, the interface may be a network interface to a local or wide area network, e.g., the Internet, a storage interface to an internal or external data storage, an application interface (API), etc.
[0092] Molecular structure generator computer 110, denoising training computer 120, and continuous graph inference computer 130 are represented here as single devices, though any of them could just as well be implemented as systems, e.g.,R. 417936
[0093] - 9 -
[0094] a geographically distributed system, e.g., a cloud computing system, e.g., a system comprising multiple computers. Further, any one of molecular structure generator computer 110, denoising training computer 120, and continuous graph inference computer 130 could be implemented as a process running on a computer, e.g., a cloud computing system.
[0095] Storage 112, 122, and 132 may be, e.g., electronic storage, magnetic storage, etc. The storage may comprise local storage, e.g., a local hard drive or electronic memory. Storage 112, 122, and 132 may comprise non-local storage, e.g., cloud storage. In the latter case, storage 112, 122, and 132 may comprise a storage interface to the non-local storage. Storage may comprise multiple discrete substorages together making up storage 112, 122, and 132.
[0096] Storage 112, 122, and / or 132 may be non-transitory storage. For example, storage 112, 122, and / or 132 may store data in the presence of power, such as a volatile memory device, e.g., a Random Access Memory (RAM). For example, storage 112, 122, and / or 132 may store data in the presence of power as well as outside the presence of power, such as a non-volatile memory device, e.g., Flash memory. Storage may comprise a volatile writable part, say a RAM, and / or a non-volatile writable part, e.g., Flash. Storage may comprise a non-volatile non-writable part, e.g., ROM, e.g., storing part of the software.
[0097] Devices 110, 120, and / or 130 may communicate internally, with each other, with other devices, external storage, input devices, output devices, and / or one or more sensors over a computer network. The computer network may be an internet, an intranet, a LAN, a WLAN, a WAN, etc. The computer network may be the Internet. Devices 110, 120, and / or 130 may comprise a connection interface which is arranged to communicate within molecular structure system 100 or outside of molecular structure system 100 as needed. For example, the connection interface may comprise a connector, e.g., a wired connector, e.g., an Ethernet connector, an optical connector, etc., or a wireless connector, e.g., an antenna, e.g., a Wi-Fi, 4G, or 5G antenna.R. 417936
[0098] - 10 -
[0099] Communication interface 113 may be used to send or receive digital data, e.g., data representing the initial noisy spatial function, intermediate denoised spatial density functions, and computed integrals for further processing. Communication interface 123 may be used to send or receive digital data, e.g., training data sets and neural network parameter updates for refining the denoising process.
[0100] Communication interface 133 may be used to send or receive digital data, e.g., continuous graph representations, neural networks trained to compute on continuous graph representations.
[0101] Molecular structure generator computer 110, denoising training computer 120, and continuous graph inference computer 130 may have a user interface, which may include well-known elements such as one or more buttons, a keyboard, a display, a touch screen, etc. The user interface may be arranged for accommodating user interaction for initiating molecular structure generation operations, training, and / or continuous graph computations.
[0102] The execution of devices 110, 120, and / or 130 may be implemented in a processor system. Devices 110, 120, and / or 130 may comprise functional units to implement aspects of embodiments. The functional units may be part of the processor system. For example, functional units shown herein may be wholly or partially implemented in computer instructions stored in a storage of the device and executable by the processor system.
[0103] The processor system may comprise one or more processor circuits, e.g., microprocessors, CPUs, GPUs, etc. Devices 110, 120, and / or 130 may comprise multiple processors. A processor circuit may be implemented in a distributed fashion, e.g., as multiple sub-processor circuits. For example, devices 110, 120, and / or 130 may use cloud computing.
[0104] Typically, molecular structure generator computer 110, denoising training computer 120, and continuous graph inference computer 130 each comprise one or more microprocessors which execute appropriate software stored at the device; for example, that software may have been downloaded and / or stored in aR. 417936
[0105] - 11 -
[0106] corresponding memory, e.g., a volatile memory such as RAM or a non-volatile memory such as Flash.
[0107] Instead of using software to implement a function, devices 110, 120, and / or 130 may, in whole or in part, be implemented in programmable logic, e.g., as a field-programmable gate array (FPGA). The devices may be implemented, in whole or in part, as a so-called application-specific integrated circuit (ASIC), e.g., an integrated circuit (IC) customized for their particular use. For example, the circuits may be implemented in CMOS, e.g., using a hardware description language such as Verilog, VHDL, etc. In particular, molecular structure generator computer 110, denoising training computer 120, and continuous graph inference computer 130 may comprise circuits, e.g., for cryptographic processing and / or arithmetic processing.
[0108] In hybrid embodiments, functional units are implemented partially in hardware, e.g., as coprocessors, e.g., arithmetic coprocessors, and partially in software stored and executed on the device.
[0109] Embodiments use continuous graph representations, a novel approach that extends the concept of traditional geometric graphs. The mathematical formalism allows for both inference and generative tasks to be performed on these continuous graph representations.
[0110] By treating graphs as continuous objects, this method eliminates the requirement to predefine the number of nodes in a graph. Furthermore, this approach is resolution-adaptable, enabling the use of different resolutions for the continuous graph representations during both training and inference phases.
[0111] The continuous graph representations can be used for inference tasks on geometric graphs. For example, they can predict the energy of a molecule or classify meshes. The method also allows for generative modeling of both graphs and meshes.R. 417936
[0112] - 12 -
[0113] Continuous graph representations can be used for generating graphs and meshes. Given a dataset, it can be used to generate new data either conditionally or unconditionally. In an embodiment, this may be done in a resolution-free way.
[0114] The motivating example and current primary use for continuous graph representations are in molecule / drug design. A further use case is in mesh generation for simulation models. For example, molecule configurations may be generated for drug development or other chemical engineering processes.
[0115] In a possible molecule design process, existing 3D configurations of molecules are used as training data, and candidates for plausible 3D configurations of molecules are generated. For example, molecules may be tested experimentally, e.g., in a wet lab.
[0116] A new 3D configuration of a molecule can be used as input geometry for a three-dimensional partial differential equation (PDE) simulation, where its spatial coordinates define the simulation domain and boundary conditions. This simulation may comprise numerically solving a three-dimensional partial differential equation to compute properties of the molecule, such as electron density or electrostatic potential.
[0117] For example, an embodiment may be configured to generate 3D molecular structures, e.g., molecules. These may be novel molecular structures, which do not need manual domain expert generation. Generated molecular structures may be evaluated according to various criteria, as further set out herein. Selected molecular structures may be synthesized.
[0118] In an embodiment, discrete graph structures are generalized to continuous field representations, which can represent both discrete graphs and intermediate constructs, e.g., between two traditional discrete graphs. Properties such as node density, edge density, and other features may be represented as continuous fields. Interestingly, these continuous field representations, also referred to as excitation fields, can be used to generate new graphs with an arbitrary number ofR. 417936
[0119] - 13 -
[0120] nodes; in particular without predetermining a set number of nodes. In a motivating example, the graphs represent molecular structures. For example, a graph may represent a molecule and the graph nodes may represent atoms.
[0121] Graphs are represented as continuous objects in space. Consider a geometric graph where each node is embedded in a n-dimensional space, and has a position vector r e R” assigned to it. The key element is a function pn(r) which represents node density at a location r. For a graph with N nodes, this density should integrate to N.
[0122] Optionally, additional functions may be used.
[0123] h(r): A feature field evaluated at the point r. The feature field can be considered as a generalization of a feature-vector, as each location in space now has a vector value assigned to it.
[0124] pe(ri,r2): Represents the edge density existing between the locations r4and r2.
[0125] For a graph with E edges, this density should integrate to E. This function is optional. In fact in many embodiments, the function is omitted. In such a case the location of edges is implicit from the location of the graph nodes, e.g., atoms.
[0126] We shall interchangeably refer to this continuous representation of graphs as a field. Thus, when we refer to a graph field, we are referring to the set of node densities, and optionally feature field(s), and / or edge densities.
[0127] Conventional discrete graphs can be expressed within the framework of continuous graphs. Consider a graph comprising of a set of nodes
[0128]
[0129] and edges e, ,, where e, , = 1 indicates that nodes i and / are connected. Each node i is positioned at and possesses a feature vector h;. In this context, the field functions can be represented as follows:
[0130] pn(r) = 2^ <5 (r - rj, e.g., it may be a sum of Dirac delta functions. This is because the nodes have deterministic positions.R. 417936
[0131] - 14 -
[0132] j°e(
[0133]
[0134] ri<r2) = 1X1 £7=i <5 (ri_ ri)<^(r2_ r;)eij:This can be understood as follows: an edge can only exist between two points if both of them are nodes (indicated by the two delta functions), and if those two points share an edge in the original graph, denoted by e^.
[0135] h(r) = Z ' i 3 (r - r h;: This function may be a weighted sum of Dirac delta functions, where each node’s feature vector h; is associated with its respective position fj.
[0136] It is important to highlight that the densities mentioned, e.g., the node density, are not normalized. This fact is used to advantage when generating graphs, e.g., a molecular structure, that do not have a predefined number of nodes — it is not necessary to define before generating how many nodes should be generated. Rather the number of nodes can be changed organically during generation. This leads to significantly more natural structures that better bit the empirical distribution of the training data.
[0137] In an embodiment, the node density has a total density mass of N representing the number of nodes. If an edge density is used, it has a mass of E indicating the total number of edges. To constrain the vast space of potential functions, some reasonable assumptions regarding the edge density function pe(r1,r2) rnay be introduced. Intuitively, for two random points in space, the likelihood of an edge existing between them is low unless nodes are present at those specific locations. Therefore, the first important property of this function is that it should assign higher values to pairs of locations where the node existence densities are higher. This relationship naturally emerges in the limiting case of a regular fully-connected graph (eij = 1 for every pair of nodes), as we can observe that in such cases, the edge and node densities are related:
[0138] N N
[0139] Pefa, r2) = 6 (ri - r;)<5(r2- r7) = jon(ri)jon(r2)
[0140]
[0141] i=l j=l
[0142] Motivated by this observation, the edge function may be modelled as:R. 417936
[0143] - 15 -
[0144] e(ri,r2) = PnCxOpn^ftr^,
[0145] where f (r1;r2) is a function.
[0146] In many scenarios, explicit edge indices may not be present. For example, in molecular structures, edges may be implicit. For example, e.g., as observed in molecular structures edges, e.g., edges representing an atom bonds, may be inferred, e.g., based on proximity.
[0147] For example, a quick approximation, which may be used for graph computations such as further disclosed herein a function may be used that scales edge density interaction based on the distance between nodes. For example, such function could be: /
[0148]
[0149] (Ji, r2) =e“fri-r2)2. However, in general, we will not necessarily assume this specific form of the function f (r1;r2). In particular, in many generation embodiments, no explicit modelling of edges is done or needed, as edges may be inferred using traditional computational chemistry software.
[0150] Figure 2a schematically shows an example of an embodiment of a molecular structure generator system 200. System 200 uses continuous graph representation, e.g., such as described above, to generate realistic molecular structures; for example, the molecular structures may be obtained as drawn from an empirical distribution as defined by a training set.
[0151] Embodiments may be applied to various types of molecular structures. A molecular structure comprises at least a plurality of molecular substructure positions. A molecular structure may further comprise bonding information between the molecular substructures and / or other information regarding the molecular structure.
[0152] In our motivating example, which was found to work effectively in a prototype, the molecular structures may be molecules. In this case, the molecular substructures may comprise atoms. Accordingly, in an embodiment, molecular structures areR. 417936
[0153] - 16 -
[0154] generated comprising at least a plurality of atom positions of the atoms in the molecule.
[0155] Other molecular structures are possible. In an embodiment, the molecular structure comprises a protein, wherein the molecular substructures are amino acids. In an embodiment, the molecular structures comprise chemicals, wherein individual chemical groups (such as hydroxyl, methyl, or carboxyl groups) act as molecular substructures. Embodiment described for molecules and atoms may be adapted to operate on these other molecular structures and substructures.
[0156] In an embodiment, a molecular structure, e.g., a molecule, is represented as a density function. In particular, the positions of the molecular substructures are represented in the density function. For example, a molecular structure of a physical molecular structure may be represented as a density function as follows. First, the molecular substructures of the physical molecular structure are represented in a spatial domain, e.g., a part of 2D or 3D space, preferably with positions corresponding to the physical positions of the physical molecular substructures. Typically, 3D space is used for the spatial domain. Using 3D space has the advantage that modelled positions naturally correspond to physical positions. However, the generation process may be used for any positive dimension.
[0157] The molecular substructures can be represented as a sum of density functions, wherein the density functions correspond to the molecular substructures, and the probability density function is centered on the position of the corresponding molecular substructure, e.g., an atom.
[0158] Figure 4 schematically shows an example of an embodiment of a continuous graph representing a molecular structure. In this case, a water molecule has been represented in 2D space, which in turn is represented as a sum of density functions. The figure shows a 3D representation of the resulting density function for the water molecule. The atom positions and bonds are indicated in the figure as a traditional graph, superimposed above the density function. Note that the integral over the density function is 3.R. 417936
[0159] - 17 -
[0160] In an embodiment, the molecular structures are molecules, and the molecular substructures are atoms. A spatial density function represents the position of the atoms. In an embodiment, the molecular structures are proteins. Distinct peaks of the spatial density function correspond to individual amino acid centers.
[0161] Structures that are in the process of generation are likewise spatial functions, e.g., functions from the spatial domain, e.g., a part of 2D or 3D space, to the real numbers. However, during generation, the function does not necessarily have to be a density function; that is, the function may be allowed to map to negative numbers. Note that a density function is a function from the spatial domain to non-negative numbers.
[0162] In an embodiment, the spatial functions are spatial density functions throughout generation, though, as said, this is not necessary.
[0163] There are various ways to represent a spatial function and / or spatial density function in a computer. In a first approach, a regular grid of voxels is applied to the spatial domain. In this case, the spatial function is sampled at the positions of the regular grid, e.g., a Cartesian grid, a honeycomb grid, etc. During generation, the values at the regular grid are refined to converge on a molecular structure. In this case, the number of values needed to represent a spatial function grows with the third power of the resolution. When a larger resolution is used, e.g., in order to reliably represent larger molecules, the amount of sampled values grows quickly. As a result, denoising operations operate on more data, and thus become slower. Furthermore, the neural network used for denoising will require more parameters, making it slower and requiring more training data.
[0164] The inventors found an alternative way to represent a spatial function. Here, a spatial function is represented as a series of tuples
[0165]
[0166] wherein the
[0167]
[0168] are vectors in the spatial domain, and ptis a number, e.g., a real number. For a spatial density function the ptare non-negative or preferably, positive numbers. It was found that this representation presents many advantages. First of all, the number of samples does not grow with the size of the spatial domain. Although aR. 417936
[0169] - 18 -
[0170] higher resolution of the spatial function will use more samples, the number of samples needed turns out to grow much smaller. It was found that 0(n) samples, wherein n represents the maximal number of molecular substructures, appear to work well. For example, in an embodiment, for an upper bound of molecular substructures of n, 10On or fewer samples are used. In an embodiment, 10n or fewer samples are used. For example, between lOn and lOOn samples.
[0171] Even if some tolerance is allowed for increased complexity that may arise with an increasing number of molecular substructures, e.g., O(nlogn) samples may be used, e.g., 100 nlogn or fewer samples, this still this grows significantly slower than the third power. For example, between nlogn and lOOnlogn samples, e.g., between lOnlogn and lOOnlogn samples. Note, that in voxel representation increasing resolution, that is an increasing number of voxels, is needed for increasingly large molecules. In sample representation, increasing accuracy may be obtained by enlarging the spatial domain, or increasing the accuracy of the representation of the positions, e.g., an increased number of digits.
[0172] Generation: Initial Noisy Spatial Density Function
[0173] Initial spatial function generator 210 is configured for generating an initial noisy spatial function 211 and representing the spatial function in a data structure. The data structure may be according to a grid or, preferably, according to a list of sampled values.
[0174] Returning to Figure 2a, system 200 comprises an initial spatial function generator 210. For example, system 200 may generate a sequence of
[0175]
[0176] wherein the Xi are randomly selected from the spatial domain, and the ptare randomly selected from a suitable interval. The values pt
[0177]
[0178] and are treated as a tuple, because the value of ptis determined by the value of the density of the field at the location x so the locations and values are entangled.
[0179] A suitable interval may be derived from the training process, as the interval for the values ptshould correspond to the maximally disturbed values in the disturbed training data (see below). The range for ptmay be constraint in that theR. 417936
[0180] - 19 -
[0181] field integrates to a particular number. Such a constraint will influence though typically not determine the number of nodes in a graph. This may be used to coax the search into a particular direction. For example, a user may set the particular number, after which the tuples are selected so that an integral over the field represented by the tuples equals or at least approximates the particular number.
[0182] As an example, the values ptmay be selected from -1 to 4, though larger or smaller intervals are possible. Typically, as more denoising iterations are used, a larger interval for the values ptis used. The vector values of
[0183]
[0184] are less critical. For example, they may be chosen to correspond to values observed in training. For example, the spatial domain may be the cubed unit interval. In an embodiment, the cubed interval [-4, 6] is used for the points xt. As resolution depends on the number of samples and the accuracy of the real number representation, but not on the size of the spatial domain, this is less critical. Various distributions for xtand ptare possible. For example, the distribution may be uniform. The distribution may be Gaussian, preferably with a large variation.
[0185] Generation: Progressive Denoising
[0186] The noisy spatial function 211 is progressively denoised by iteratively applying a trained neural network 220. Various numbers of iterations can be used.
[0187] Interestingly, even with comparatively low numbers of iterations, good results can be obtained; probably due to the fitting process further improving the result. For example, in an embodiment, at least 10, at least 100, or at least 1000 iterations are used. In the worked example shown in Figures 3a-3d, 150 iterations were used. In an embodiment, a different number of iterations may be used, either a higher or lower number. In a more efficient embodiment, fewer than 150 iterations are used. In an embodiment, at least 10 iterations are used, e.g., between 10 and 150 iterations.
[0188] The neural network is trained to transform noisy density functions eventually into structured spatial density functions based on training data comprising molecularR. 417936
[0189] - 20 -
[0190] structures represented as spatial density functions. A possible training scheme for neural network 220 is presented below.
[0191] Interestingly, the number of molecular substructures, e.g., atoms, in the final spatial density function corresponds to an integral over the density function. Interestingly, the integral over the spatial function can vary during the denoising process. This means that the number of molecular substructures, e.g., atoms, that are represented is not fixed but can vary. As a practical limit, a higher number of atoms needs more resolution to be adequately represented, but this can be accounted for by a linear increase in the number of samples.
[0192] Interestingly, by repeatedly applying a denoising network, the initial random configuration converges on a configuration drawn from the empirical distribution of the training data.
[0193] The neural network may work on point sets. For example, in an embodiment, the neural network receives as input a point cloud rather than a fixed grid. In an embodiment, the neural network comprises one or more convolution operators defined to operate on a point set. For each input point (x^.p^, a local neighborhood is defined — for example, the k-nearest neighbors or all points within a fixed radius. Then, a weighted sum of the densities pj of the neighboring points Xj is computed. However, the kernel weights are not fixed numbers in a grid; but it given by a function k(x X ) that depends on the relative spatial positions xtand xj and is defined by the network.
[0194] Mathematically, for each point x the convolution-like operation could be defined as: ZjEN x.)k(xi,Xj') • pj. Here, / V(xj) denotes the neighborhood of x and k is defined by a neural network. This operation aggregates the local information in a way that is invariant to the ordering of the points. The results of the point cloud convolutions are input to further layers, e.g., further convolutions, and / or pooling operations.R. 417936
[0195] - 21 -
[0196] In an embodiment, the point cloud technique disclosed in “On the Utility of Equivariance and Symmetry Breaking in Deep Learning Architectures on Point Clouds” by Vadgama et al. is used.
[0197] In an embodiment, the spatial functions are presented as samples taken along a regular grid. Although this is slower, it does allow the use of conventional convolutions.
[0198] Feature Fields
[0199] Denoising spatial functions, e.g., to obtain spatial density functions that represent molecular structures is surprisingly effective. The inventors found two complementary ways to include more information during the generation: conditional generation and feature fields. Feature fields are particularly effective, as they provide more information after generation is complete, e.g., information beyond the positions of molecular substructures. Moreover, feature fields allow the incorporation of more chemical data during training. This leads to improved performance of the denoising neural network and thus more accurate results.
[0200] A feature field is a spatial function from the same spatial domain. The feature field may have values of categorical type, e.g., atom type, and numerical type, e.g., charge. For categorical data, the feature field may map to vectors, in particular, one-hot vectors. For example, a one-hot vector may encode an atom type, using a different vector element for each atom type. This allows the feature field to vary continuously during generation while converging to a one-hot vector.
[0201] Like the spatial function, a feature field is also progressively denoised by iteratively applying a trained neural network. Preferably, this uses the same neural network as the one denoising node density, to take advantage of the additional structural information in the feature fields.
[0202] Feature fields may be used to represent a range of information, e.g., one or more of the following: atom type, charge, electronegativity, partial atomic charge, hybridization state, van der Waals radius, bonding preferences, and polarizability.R. 417936
[0203] - 22 -
[0204] For example, in the case of the water molecule of Figure 4, a feature field may encode the atom types. This may be regarded as a mapping from
[0205] R2(coordinates) to R2(two atom types). However, there is also the node density, which is a scalar-valued field. Note that typically, feature fields map from R3if a 3D spatial domain is used. In the example of Figure 4, the feature field uses a one-hot vector of dimension two, however in other molecular tasks, there may be more than two dimensions, e.g., because there are more atom types.
[0206] Interestingly, it was found that no regularization is needed to keep the density function and feature fields consistent. The model learns that there are correlations between the two. It turned out not to be necessary to impose conditions on the density function and feature field(s). In an embodiment, the neural network takes as input both a spatial function, possibly specified as a list of pairs of a vector and a value, e.g., a density, and a feature field. The feature field may also be represented as a list of pairs and a value or vector of values. The value or values may be a density. The neural network is configured to produce as output both a denoised spatial field corresponding to spatial density and a denoised feature field.
[0207] Below a number of embodiments are described using more mathematical language. In an embodiment, spatial functions, e.g., excitation fields, are generated that can later be converted to discrete graphs. During training, graphs are converted to fields, as described herein. This creates the ability to calculate losses during training directly in a continuous space, rather than having to convert fields back to graphs, which would typically require some non-differentiable operation. On a high-level, the sampling process can be summarized as:
[0208] Initiate the process with a noisy node density p(r), and optionally noisy feature fields h(r).
[0209] Denoise both the node densities and optionally the feature fields.R. 417936
[0210] - 23 -
[0211] As a result, a node density p(r) is obtained, and optionally also a feature field h(r).
[0212] To avoid the representation of these fields as images and / or voxels, we generative models, e.g., denoising neural networks, may be used that operate directly on a sampled field representations, e.g., wherein the final representation is of the form of a sequence of pairs (%;,
[0213]
[0214] for the node density, and a sequence of pairs (%;, h xi')') for the feature field. Preferably, / ?(%;) is a nonnegative or even positive function. Preferably, the integral p corresponds to the number of molecular substructures, e.g., atoms. This condition may also be imposed on h(xi'), e.g, non-negative or positive values or vectors with nonnegative or positive values. However for the feature field this is not necessary, especially if the modeled values may be negative. It is noted that for the intermediate spatial functions during the denoising process non-negativity constraints may be relaxed. In particular, negative values may be included in the initial spatial function to start the denoising process.
[0215] Formally, Denoising Diffusion Probabilistic Models (DDPMs) are a type of latent variable model. They aim to learn a parametric data distribution p0(xo) from an empirical distribution of finite samples q(x0). This is achieved by reversing the forward diffusion process that generates latent variables x1;Tby gradually adding noise to the data, typically Gaussian noise. We develop x0~ q(x0) over T timesteps as follows:
[0216]
[0217] q(xt|xt_i) : = N(xt; (1 - at)Z)
[0218] Here, atis the cumulative product of fixed variances that are obtained from a noise scheduler. For our purposes, we may use a simple linear scheduler.
[0219] By reparametrizing the previous equation as xt=
[0220]
[0221] + V1 > canbe shown that the simplified loss can be used for training:
[0222]
[0223] ^9 ^t~|O,7’|,xo~q(xo),6~J\T(O, / ) [ll^—+ ^ / 1—t)||R. 417936
[0224] - 24 -
[0225] where eeis the denoising model parametrized by 9. The process is well-defined when the data x0is a vector from an ^-dimensional vector space ]RW. However, in our setup, we are interested in generating densities p(r), which are continuous functions. The generalization to this case is not straightforward. Nonetheless, a rigorous mathematical formulation of such a problem exists.
[0226] We propose an improved version of these methods, serving as a natural generalization, wherein a Gaussian process (GP) is used instead of independent Gaussian noise. A Gaussian process may be regarded as a random function with spatial correlations defined by a kernel. Formally, we can express the noisy density p(r) at time step t as:
[0227] P
[0228]
[0229] t(r) = V^PoCr) + 71 - atg(r),
[0230] where g(v) is a Gaussian process evaluated at the coordinate r. Consequently, the optimization target may be modified to:
[0231] L
[0232]
[0233] e = JEt~[o,T],po(r)~Q(Po(r)),5~GP0,ri~x [kCn) - + ^1 - t)||2],
[0234] where additionally the expectation is taken over the field evaluations at a finite number of potentially non-uniformly sampled locations 17 on a domain X, and we sample a random function g r) from a GP parameterized by 0. Note that for generating both densities p(r) and feature fields h(r), an embodiment follows the same procedure as described above. Interestingly, this can be implemented by concatenating them and treating the result as a single vector field.
[0235] Another approach, which may be used together or instead of feature fields is conditional generation. In an embodiment the generation of the molecular structure is conditioned on one or more conditioning parameters. For example, the one or more conditioning parameters may specify desired molecular properties or structural constraints. For example, the one or more conditioning parameters may be obtained from a user, read from a file, or the like.R. 417936
[0236] - 25 -
[0237] The conditioning parameters can include both quantitative molecular descriptors and qualitative structural constraints. For instance, these parameters may represent continuous values such as molecular weight, logP, polar surface area, or the number of hydrogen bond donors / acceptors, as well as categorical constraints like the presence of specific functional groups, ring sizes, or connectivity patterns.
[0238] Conditioning parameters may comprise, e.g., a class label, or a text description, etc. In an embodiment, conditioning parameters are preprocessed, e.g., continuous values may be normalized, while categorical attributes may be encoded, e.g., via one-hot encoding or learned embeddings.
[0239] Optionally, the conditioning parameters may be encoded into a vector, possibly using an embedding layer or an encoder network. This vector serves as a compact representation of the auxiliary data.
[0240] To introduce conditioning into this framework, the network architecture and / or the training procedure may be adapted in various ways so that the denoising process is informed by the conditioning parameters. The neural network is trained to transform noisy spatial density functions into structured spatial density functions conditioned by the one or more conditioning parameters.
[0241] Note that conditioning parameters are used both during generation and during training. During training, the network learns to generate molecular structures that satisfy the specified properties and constraints by minimizing the prediction error between the denoised outputs and the target structures.
[0242] Integrating the conditioning parameters into the denoising network can be achieved through various methods. For example, the input of the denoising network (e.g., the noisy spatial function) may be extended with one or more conditioning parameters. For example, the conditioning vector may be appended to the noise input or intermediate feature maps. The one or more conditioningR. 417936
[0243] - 26 -
[0244] parameters may be provided to the trained neural network when applied to the spatial function.
[0245] This is preferably done at each denoising step. Other approaches include incorporating conditional normalization layers into the network (e.g., FiLM layers) that use the conditioning vector to modulate activations, and / or incorporating cross-attention layers that allow the network to selectively focus on different parts of the conditioning vector during the denoising process.
[0246] Discretizing
[0247] Returning to Figure 2a ,at some point, the iterative denoising by repeatedly applying a denoising neural network is terminated. Termination may be indicated in various ways. For example, the system may be configured for a fixed number of iterations. The fixed number may be the same as or proportional to the number of disturbing steps used during training. Another approach is to use a convergence criterion. The denoising process may be terminated when the difference between successive denoised outputs becomes minimal, indicating convergence. For example, a norm, say L2 or L1 norm, of the differences may be below a threshold to decide that the changes are minimal.
[0248] Once denoising has terminated, the continuous graph representation may be converted to a discrete graph. In this case, the discrete graph may not have explicit edges. The edges may be left implicit, as they are suggested by molecular substructure positions and possibly feature fields. For example, molecular substructure positions may be atom positions. For example, a feature field may be atom type.
[0249] In an embodiment, an integral of the spatial function produced when the last denoising step is computed. As the targets during training are spatial density functions, it is very likely that the final spatial function is a spatial density function.R. 417936
[0250] - 27 -
[0251] An embodiment may explicitly check if the spatial function is not a spatial density function, and if so, an error may be reported. However, assuming the integral over the spatial function is positive, the discretization may still proceed.
[0252] Based on the integration of the spatial function, an integer number is derived. Typically, this number is a rounding of the integration result, e.g., rounding to the nearest integer, rounding down to an integer, etc. In a variant, a discretization may be performed for slightly different values, e.g., a rounding of the integration plus a small number, e.g., plus one. This will produce slightly different discretizations.
[0253] In an embodiment, system 200 comprises an integrator 230 configured to compute an integral over the last spatial function computed by the denoising neural network. Integrator 230 produces a molecular substructure number 231. If the spatial function is represented as values along a regular grid, then the integral may be approximated by adding the values and multiplying with the size of the voxels. If the spatial function is represented as a series of samples, the integral can also be computed, e.g., using then importance sampling or other Monte Carlo techniques may be used.
[0254] Next, a sum of a specified number of predefined density distributions is fitted to the final spatial function. In an embodiment, system 200 comprises a fitting unit 240 configured to produce a first molecular structure 241. First molecular structure 241 may comprise a list of positions of the molecular substructures, e.g., atoms.
[0255] For example, assume f (x) is the final spatial function, and m is specified number, that is the integer number based on the integration of the spatial function. The objective may be defined as finding
[0256]
[0257] so that 27=1 / V( i) fits the function f (x). Here N(jii) is a distribution with mean i- Typically, a Gaussian with fixed standard deviation, e.g., unit standard deviation is used. The distribution used is typically the same as the distribution used during training when generating the training data. Other distributions than Gaussian can be used, e.g., a Gamma distribution.R. 417936
[0258] - 28 -
[0259] For example, an objective function may quantify the difference between the output of the denoising network and the function to be fitted. When the neural network output is given as a list of pairs (x^.p^, then the objective function may, e.g., be defined as: L(p , ...,Pn =
[0260]
[0261] / <)>Pt2- Note that here the values ...,Pn are points in the spatial domain, typically a 3D domain.
[0262] An optimizer may be used to optimize the objective function and derive the means. For example, a gradient descent optimizer may be used.
[0263] In an embodiment, first an initial estimate of the means is obtained using a clustering algorithm, e.g., using k-means, to get an initial estimate of the means, followed by fitting a weighted mixture of gaussians to further approximate the means better.
[0264] Once the distributions have been fitted to the data, the centers of the fitted density distributions are identified as positions of the molecular substructures of the molecular structure. In the example, above the means, e.g., the fitted values of ...,Pn may be regarded as the centers of the fitted density distributions.
[0265] Once the means are known a molecular structure has been generated. It is possible to enhance the molecular structure in various ways. For example, connectivity between identified molecular substructure positions may be inferred by applying a computational bond determination method to construct a molecular graph. In an embodiment, system 200 comprises a connectivity inferring unit 250 configured to produce a second molecular structure 251. Second molecular structure 251 may comprise the same information as first molecular structure 241 and in addition bonding information. Alternatively, second molecular structure 251 may be a pure graph representation of the molecule with atoms as nodes, bonds as edges. Second molecular structure 251 may also be in traditional formats, e.g., a connection table (Ctab). Second molecular structure 251 may comprise information describing the structural relationships and properties of a collection of atoms. The atoms may be wholly or partially connected by bonds. Such collections may, for example, describe molecules, molecular fragments,R. 417936
[0266] - 29 -
[0267] substructures, substituent groups, polymers, alloys, formulations, mixtures, and unconnected atoms. Second molecular structure 251 may be in MDL file format.
[0268] Various options are known in the art to do this, ranging from traditional computational methods to machine learning approaches. One approach is detailed in “Prediction of organic homolytic bond dissociation enthalpies at near chemical accuracy with sub-second computational cost” , by Peter C. St. John, et al., included herein by reference; and in “Investigating Molecular Geometry: A Comprehensive Analysis Using VSEPR Theory”, by Aejaz Ahmad Khan, included herein by reference.
[0269] It is noted that bonding information may be applied in discretized form, e.g., as edges in a graph, wherein the molecular substructures are nodes, or as numerical information. Possibly for each edge between molecular substructure nodes, numerical bonding information may be provided. An example of such numerical information is bond order, which quantifies the strength and multiplicity of a bond between two molecular substructure nodes, or bond dissociation enthalpy, which represents the energy required to break the bond.
[0270] Instead of or in addition to providing additional information to the discretized generated molecular structure, bonding information could also be computed for a continuous graph representation, e.g., just before discretizing. In this case, continuous values can be used for the presence or absence of a bond. This has the advantage that continuous graph neural network computations can be used, as described herein.
[0271] Another way to provide additional information to the discretized molecular structure, with or without explicit bonding, e.g., with or without explicit edges, is to extract information from feature fields that were computed, e.g., denoised, during the generation of the molecular structure.
[0272] For example, after discretizing the spatial function indicating the density of molecular substructures, e.g., atoms, one or more denoised feature fields may beR. 417936
[0273] - 30 -
[0274] sampled at the identified positions of the molecular substructures to obtain feature values associated with each molecular substructure.
[0275] The discretized molecular structure may be saved to a file and / or reported to a user. Additional information, e.g., inferred edges such as bonding information, and / or sampled feature fields may likewise be saved to a file and / or reported to a user. Various applications may be provided with the molecular structure information, whether in discretized or continuous form, and / or with or without additional information.
[0276] In an embodiment, the method of the invention is used to reconstruct molecular structures, e.g., molecules, based on chemical information. For example, such an embodiment may be used in a forensic setting to determine which molecules were likely to be present based on chemical information. For example, such an embodiment may be used in a chemical testing and / or research setting, where it is needed to know which molecules were present based on chemical information, e.g., to investigate the source of a leak, to determine the identity of a useful molecule or molecular structure, etc.
[0277] For example, chemical data is obtained regarding the molecular structure, typically experimentally. For example, experimental data regarding a present molecule may be collected; however, the molecule in question may be unknown. For example, the experimental data may comprise any one of the following: full or partial NMR spectra, full or partial mass spectra, or full or partial MS fragmentation. The experimental data may be provided as conditioning data during the generation process. Generated molecules are likely to fit the experimental data and correspond with the physical — though possibly unknown — molecule. In an embodiment, multiple molecules are generated, and the top-k most occurring molecules are presented.
[0278] A further application relates to the testing of chemical software. Chemical software requires rigorous testing to ensure accurate and reliable performance. Such software includes, for example, algorithms for inferring molecular bonds from atomic positions and computational methods for determining chemicalR. 417936
[0279] - 31 -
[0280] properties such as dipole moments or reaction energies. However, testing such software is challenging due to the limited availability of molecular data in public and proprietary databases. Furthermore, the molecular structures and properties present in these databases are often already used during software development, making them unsuitable for independent validation. As a result, testing efforts may be biased or incomplete.
[0281] Using random data to test, debug, or benchmark chemical software is generally ineffective. The vast majority of randomly generated structures do not correspond to chemically meaningful molecules, often leading to early termination due to structural errors or unrealistic bonding configurations. Such an approach does not effectively exercise the full functionality of the software, as most test cases fail prematurely rather than probing complex or edge-case behaviors.
[0282] Consequently, to test chemical software in a way that enables meaningful stresstesting and validation, there is a need for a large amount of realistic molecular structures.
[0283] Generating a plurality of molecular structures according to an embodiment provides a solution by creating realistic and chemically valid test cases. Applying a further chemical software, configured to receive a molecular structure, to the plurality of molecular structures allows systematic evaluation of its behavior. During testing, errors may arise due to unexpected molecular configurations or computational failures. The software may be monitored for outright errors, e.g., crashes of the software, memory leaks, etc. Furthermore, the software's output can be monitored for implausible results, e.g., inconsistencies such as physically implausible bond lengths or unrealistic property calculations. When errors or other problems are detected, the specific molecular structure causing the issue can be identified and flagged, allowing developers to analyze the failure.
[0284] Additionally, a plurality of molecular structures can be used to benchmark different chemical software implementations by applying various types of further software and comparing their performance under controlled conditions. By timing computations on the same or similar hardware, discrepancies in efficiency can be detected. Differences in execution time, memory usage, or computational stabilityR. 417936
[0285] - 32 -
[0286] across similar molecules may indicate inefficiencies or implementation flaws in a given software. Comparative benchmarking can also reveal variations in computed chemical properties, e.g., cases where one implementation diverges from others.
[0287] In an embodiment, conditioning information may be used to ensure the generated molecular structures span a diverse chemical space, e.g., covering a wide range of bonding patterns, functional groups, and / or steric configurations.
[0288] A further application relates to the synthesis of molecular structures. Generating a molecular structure, in particular a molecule, enables the exploration of new chemical compounds with potential applications in various fields. Applying a retrosynthesis tool to the molecular structure provides a synthetic route by identifying feasible precursor molecules and reaction steps. By following this synthetic route, the molecule can be synthesized and experimentally validated. Synthesis may involve a human chemist carrying out the required reactions in a laboratory, following established procedures and adjusting conditions as needed, or it can be fully automated using robotic synthesis platforms that execute reactions under software control, e.g., leveraging automated reagent dispensing, reaction monitoring, and purification techniques.
[0289] In an embodiment, the synthesized molecule is tested for chemical and / or pharmaceutical properties. In an embodiment, the synthesis outcome is monitored for verification of the retrosynthesis tool’s accuracy, identifying cases where predicted pathways are impractical, inefficient, or ineffective.
[0290] Figure 3a schematically shows an example of an embodiment of the generation of a molecular structure.
[0291] Figure 3a shows various intermediate points during generation of a molecule. A percentage below each figure indicates at which point during generation the snapshot was taken.R. 417936
[0292] - 33 -
[0293] In each of the images, the spatial function is represented as a series of positions of a plurality of points and values of the spatial function. The plurality of points are chosen in a 3D spatial domain. The function is visualized in Figure 3 by plotting the 3D positions of each of the points and indicating the value of the spatial function with a grayscale.
[0294] The image at 0% denotes the initial sampling of points before the denoising starts. Note that at this point it is not required that the spatial function, e.g., a function from the 3D spatial domain to the real numbers, is a density function, that is, negative values are allowed. The reason for allowing this is that in the course of training, although at one end of the spectrum density functions are sampled corresponding to physical molecules, during training these points are disturbed, so that at the other end of the spectrum the spatial function represented by the points may have devolved to a spatial function rather than a spatial density function. Indeed, the scale indicated in the lower right corner of Figure 3 allows for values up to -1.
[0295] In an embodiment, the 0% image is obtained by sampling positions and their features from a standard normal distribution with zero mean.
[0296] After generating the initial 0% spatial functions, in the form of a representation of position-value pairs, the denoising network is repeatedly applied. The neural network is trained to modify the positions and / or values of the plurality of points to become closer to the empirical distribution of physical molecules in the training data.
[0297] As can be seen in Figure 3, as more iterations are processed, the points start to converge more and more on an actual molecule. Three images of Figure 3a have been produced on a larger scale.
[0298] Figure 3b schematically shows an example of an embodiment of an initial noisy spatial density function at 0%R. 417936
[0299] - 34 -
[0300] Figure 3c schematically shows an example of an embodiment of intermediate molecular structure at 54%
[0301] Figure 3d schematically shows an example of an embodiment of a molecular structure generated before discretization at 100%,
[0302] Training
[0303] Figure 2b schematically shows an example of an embodiment of a denoising training system 200.
[0304] To train the denoising neural network, molecular structure training data 270 is needed, that is a plurality of molecular structures comprising a plurality of molecular substructure positions. Various public or proprietary datasets may be used. For example, the worldwide protein data bank archive may be used for plurality 270. Furthermore, the QM9 dataset, or GEOM-drug dataset may be used.
[0305] The molecular structures in the training set are first converted to spatial density functions. In an embodiment, system 200 comprises a density function conversion unit 280 configured to produce a density function representation 281.
[0306] For example, a suitable spatial domain may be chosen, e.g., the unit interval cubed, or the like.
[0307] To convert a molecular structure, its molecular substructures, e.g., atoms, first positional information is mapped to the spatial domain, e.g., by applying a scaling factor; the molecular structure is embedded in the spatial domain. Thus obtaining a set of means: A spatial density function is then defined as (%) =
[0308]
[0309] Hi). Here N(x; ^i) is a distribution with mean z;. Typically, a Gaussian with fixed standard deviation, e.g., unit standard deviation is used. Other distributions than Gaussian can be used, e.g., a Gamma distribution. For example, the spatial domain may be a cubed interval of a size proportional to theR. 417936
[0310] - 35 -
[0311] standard deviation, e.g., using at least 5 times, at least 10 times, etc., the standard deviation, e.g., use 10 if the standard deviation is 1.
[0312] The spatial function f is then represented in a data structure, e.g., density function representation 281. For example, a series of values of f along a regular grid, e.g., voxels in the spatial domain may be used for representation 281.
[0313] Preferably however, a f is regarded as a non-normalized probability distribution and sampled according to its distribution. Various approaches exist, e.g., rejection sampling, or mixture sampling, e.g., Gaussian mixture sampling. This produces a list of points xt, with preference given to regions with higher density. This is advantageous since regions with higher density are exactly the regions where molecular substructures are located and thus where it is desired to have the most information. For each point the density is computed producing a list of pairs (X / (%i)).
[0314] Finally, representation 281 is processed by diffusor 290. Diffusor repeatedly applies a disturbance to the elements of representation 281, producing a list of increasingly disturbed spatial function representations 291-293. The number of produced disturbed spatial functions may be equal to the number of training steps performed for a given training element. To disturb a spatial function represented by a regular grid, e.g., a voxel based approach, it is sufficient to disturb the density value, e.g., by adding a value sampled from a Gaussian distribution. To disturb a spatial function represented by a sequence of values both position and density value may be disturbed, e.g., by adding a value sampled from a Gaussian to the density, and adding a value sampled from a Gaussian or Gaussian process to the position.
[0315] Once the disturbed spatial function representations are obtained, they are used to train the denoising neural network. The network is trained as a generative model or conditional generative model, learning to predict the original spatial function representation 281 given progressively noisier inputs. A loss function, such as mean squared error (MSE) or a variation of Kullback-Leibler divergence, may be used to quantify the difference between the predicted and ground truthR. 417936
[0316] - 36 -
[0317] representations. During training, a sequence of corrupted spatial function representations is provided to the network, which learns to iteratively refine them by minimizing the loss. Gradient-based optimization, typically using Adam or another variant of stochastic gradient descent (SGD), is employed to update the model parameters. The network may be trained over multiple epochs with batch normalization, dropout, or other regularization techniques to prevent overfitting. Once training is complete, the model can be used to denoise new spatial function representations by reversing the applied disturbances and reconstructing a molecular structure.
[0318] Through the use of spatial functions, e.g., excitation fields, new graphs can be generated that have an arbitrary number of nodes — the number of nodes need not be fixed before generation, though it could be controlled through conditioning data. It is important to note that this generative approach differs from graph neural networks, even graph neural networks configured to work on continuous graph representations, which primarily focus on solving traditional graph-related inference tasks using this generalized graph concept. There are two distinct ways to approach the generative process. The first approach is through field diffusion as shown above, while the second one uses inverse heat dissipation, as further described herein. Note that these two different approaches are used for generating fields, while the graph conversion that we will finally introduce is agnostic to the model used for field generation.
[0319] Instead of using a stochastic equation to progressively disperse the spatial density function over time, it is also possible to apply a heat dissipation process to the spatial density function for perturbation. Figure 5a schematically shows an example of an embodiment of a continuous graph representing a molecular structure. A molecule with three atoms has been transformed into a density function, in this case with a two dimensional spatial domain. The graph of Figure 5a can be sampled to obtain a data structure representing it. Figures 5b and 5c schematically show an example of an embodiment of a disturbed continuous graph. The density function of Figure 5a has been disturbed according to a heat dispersion process. In Figure 5b, an intermediate stage of the disturbanceR. 417936
[0320] - 37 -
[0321] process is shown. In Figure 5c, the density function has been fully flattened. Each of the disturbed density functions can be sampled to obtain training data.
[0322] It is noted that both the embedding of the original molecular structure in the training data, and the disturbing process can be performed in many ways, e.g., under the control of a random number generator. For example, different positions may be chosen to represent the same molecular structure. For example, different samples may be drawn from a Gaussian distribution to disperse a spatial function in different ways. . In an embodiment, the embedding and / or dispersing is repeated multiple times to enlarge the training set.
[0323] Continuous Graph Neural Networks Computations
[0324] Once a discretized molecular structure has been determined, e.g., first molecular structure 241 or second molecular structure 251, various properties can be computed for it. Such computing could be done using a graph neural network.
[0325] A Graph Neural Network (GNN) can compute molecular properties by treating the molecular structure as a graph, where atoms are nodes and bonds are edges. The GNN iteratively updates node (atom) and edge (bond) representations through message passing, where each node aggregates information from its neighbors using learnable transformation functions. This process captures local and global structural patterns, enabling the model to encode chemical interactions and spatial dependencies. After multiple message-passing iterations, a graph-level representation is obtained via pooling or readout functions, which can be used for property prediction tasks such as energy levels, reactivity, or solubility.
[0326] It would be desirable to apply graph neural network directly on the undiscretized molecular structure. Through the discretizing soft information is lost which may be useful during further computations.
[0327] Examples of properties that can be computed using either a discrete or continuous graph neural network computation include: solubility, toxicity, bindingR. 417936
[0328] - 38 -
[0329] affinity, e.g., drug-target interactions — all of these are important drug development, electronic properties, thermodynamic stability, amongst others.
[0330] In an embodiment, connectivity between the identified molecular substructure positions is inferred by applying a computational bond determination method to construct a spatial density function indicating a likelihood of a connection existing between two points. For example, for a given spatial domain D, the function may be a function from peD x D -> ]RSO. The function gives a likelihood of the presence of a bond between the two positions. In an embodiment, the function Pe(x>y) is constructed as pe(x, y) = c(x, y)p (x)p(x), wherein p is the generated spatial density function. Next one or more noisy feature fields are generated, each noisy feature field being a function from a spatial domain to real numbers. The noisy feature fields could be generated before the denoising and updated during the denoising. The noisy feature fields could instead be generated after the denoising. In both cased the one or more feature fields are iteratively updated in a process after the denoising. The process comprising integrating over the spatial density function indicating the likelihood of connection existing, the one or more noisy feature fields, and a trained function. The function pemay instead by approximated by a function that exponentially decays as a function of the distances between the two input points.
[0331] We define operations on these fields and prove their convergence. We now demonstrate how our approach can be applied to tasks commonly tackled by graph neural networks (GN Ns). These tasks include graph-level predictions, such as estimating the energy of a molecule, as well as node-level tasks, for example classifying individual nodes. Embodiments are capable of handling these tasks effectively.
[0332] Instead of using the process on molecular structures produced through denoising, it can also be applied by converting a discrete graph into a continuous representation.
[0333] Assuming we have a field feature at layer k, denoted as h(k)(r), we propose the following operation to calculate the field feature at the next layer:R. 417936
[0334] - 39 -
[0335] h
[0336]
[0337] (k+1)(r;) = Wj h(fe)(r7)pe(ri,r7) dr7,
[0338] Here, W represents a linear pointwise transformation. We can understand this convolutional update as a two-step process. First, we perform a convolutional step where we aggregate neighboring spatial information through the integral operation, and then we pointwise transform features using the linear transformation W. This can be thought of as an operation using which each point r; aggregates the features from all other spatial domains, and the contribution is weighed by the density of that point in space having an edge with the point 1 . Note that for simplicity, we have omitted the nonlinearity applied after the linear transformation. In 4 we prove that this converges to regular message passing in the limit of
[0339] Continuous graph predictions
[0340] Suppose we have performed K layers of continuous graph convolutions, and we now want to obtain the feature value associated with each node. For example, this may be done for tasks like node classification or obtaining each node’s feature representation. We propose the following method:
[0341] The final node representation for node i after K layers, denoted as h®, can be obtained from the feature field hw(r) using the following integral:
[0342] h® = JhW(r (r)dr
[0343] Here, pn.(r) corresponds to the density distribution for the j-th node. If we view the total node density function pnas a mixture model, this operation can be interpreted as the expected value of the feature field for the j-th component of the mixture. With these node-level vectors, we can proceed with predictions as one typically would in a (GNN).R. 417936
[0344] - 40 -
[0345] On the other hand, if we wish to predict a graph-level feature H (equivalent to global pooling in GNNs), we have two options. We can either sum-pool the feature vectors obtained using the previous equations. However, since the model is a mixture model, a linear property holds:
[0346] H = f hw(r)pn.(r) dr = f hw(r) pn. (r) dr = f hw(r)pn(r) dr,
[0347]
[0348] i i i
[0349] In other words, we can view the graph-level feature as the expected value of the feature field.
[0350] The inventors realized that computing on continuous representations of graph has utility independent of how the continuous representation is obtained. The following clauses represent useful embodiments.
[0351] Clause 1. A computer-implemented method (600) for performing computations on molecular structures, a molecular structure comprising multiple molecular substructures,
[0352] the method comprising:
[0353] - obtaining a first spatial density function and a second spatial density function representing a molecular structure, wherein the first spatial density function is from a spatial domain to non-negative numbers indicating a likelihood a molecular substructure is present, and the second spatial density function is from the spatial domain squared to non-negative numbers indicating a likelihood a connection, e.g., a bond, exists between two points in the spatial domain,
[0354] generating one or more noisy feature fields, each noisy feature field being a function from a spatial domain to real numbers;
[0355] - iteratively updating the one or more feature fields by integrating over the spatial density function indicating the likelihood of connection existing, the one or more noisy feature fields, and a trained function.R. 417936
[0356] - 41 -
[0357] Clause 2. The method according to Clause 1 , wherein
[0358] - the first and / or second spatial density function is represented by sampling at a regular grid, or
[0359] - the first and / or second spatial function is represented as a series of positions of a plurality of points, and / or pairs of points and values of the spatial function.
[0360] It is noted that any one of the clauses 1 and 2 may be combined with features and embodiments as described herein, e.g., in the description and / or claims.
[0361] In our method, feature fields are treated as continuous objects. However, representing and performing operations on these continuous objects within a neural network setting is not simple. Moreover, executing operations like integral transforms, as proposed in the previous subsection, is non-trivial. One approach is to discretize the grid and work with discrete representations of these objects, but this comes at the cost of losing some of the advantageous properties of continuous representations.
[0362] Fortunately, we’ve observed that for translation-invariant convolutions, continuous convolutions effectively become a standard convolution operations. This enables us to perform these operations in the Fourier domain, which offers several advantages. Firstly, these operations are resolution-independent. In other words, if we learn transformations, as we do in methods such as Fourier Neural Operators (FNO), the learned kernels in the Fourier domain remain invariant to the initial discretization of our representations. This strikes a balance between continuous and discretized approaches, allowing us to use the benefits of both approaches. It is possible to train a model that is invariant to discretization.
[0363] For translation-invariant kernels, we show that these operations can be efficiently implemented in the Fourier domain, enabling resolution-independent learning. Our method allows for bidirectional conversion between discrete graphs andR. 417936
[0364] - 42 -
[0365] continuous representations, facilitating both node-level and graph-level predictions.
[0366] Convergence to Message Passing
[0367] To verify whether the proposed graph neural network convolution operation converges to the standard message passing equation when the edge functions correspond to those of a regular graph, we can examine the expression step by step:
[0368] h(k+1)(r;) = Wf h(k)(r7)joe(r;, r7) dr7
[0369] N N
[0370] =5~rj')h(k)(rj)ei'J'dr7 ir=l p=i N N
[0371] = W 5 - ri^h(k)(rJ')efj' ir=l p=i N / N \ = 6 (rf- 17, ) ( Wh(k)(r7,) eifp )
[0372]
[0373] ir = l \J' = 1 /
[0374] To understand this expression, note that the first sum over i' ensures that we consider node features at all valid locations. The term within the brackets represents the typical aggregation step of the transformed features of neighboring nodes. Therefore, this is equivalent to the graph convolutions observed in regular graphs.
[0375] Decomposed Convolutions
[0376] After decomposing the edge function, the convolution operation becomes:
[0377] h(k+i) (r.)= wjh(fc) ^pe dr,
[0378] = Wj dry
[0379]
[0380] = Wj h(k)(r7)pn(ri) / (ri,r7)p„(r7) dr7R. 417936
[0381] - 43 -
[0382] If we first perform the integral transformation and then apply the pointwise transformation (along with the nonlinearity which is omitted here), the integral from the previous equation would have the form:
[0383] h
[0384]
[0385] (k+1)(r;) = f h(k\r7)pn(ri)f(ri,r7)pn(r7) dr7= f h(k)(r7)jf(ri,r7) dry
[0386] In this order, we can recognize that this is equivalent to an integral transform with the kernel K
[0387]
[0388] (ri,rj') = pn(.rdf(ri,rj)pn(^.
[0389] Translation-invariant Convolutions
[0390] We can make further assumptions on the function
[0391]
[0392] ry). Specifically, one may assume that it can be written in a translation-invariant form
[0393] pr„r;) = X0,r, - r;) = pr, - r;).
[0394] For example, the physically inspired function, would naturally fall into this category. With this assumption, the integral becomes:
[0395] h(k+1)(r;) = f h(k)(r7)pn(ri)f(ri- ry)p„(ry) dry
[0396] = PnCri) / h(k)(r7)f(r; - ry)p„(ry) dry
[0397] =
[0398]
[0399] Pn fri) / u(k)(r7)f(r; - ry) dr7,
[0400] with u(k)(
[0401]
[0402] r7) = h(k)(r7)p„(r7). We can recognize that this is a convolution between functions u and f . The latter makes the approach more practical, since it allows computation using the Fourier transform, e.g., FFT, which is faster.
[0403] Figure 6 schematically shows an example of an embodiment of a method (600) for generating molecular structures. Method 600 may be computer-implemented and comprises
[0404] - generating (610) an initial noisy spatial function and representing the spatial function in a data structure;R. 417936
[0405] - 44 -
[0406] - progressively (620) denoising the spatial function by iteratively applying a trained neural network, wherein the neural network is trained to transform noisy density functions into structured spatial density functions based on training data comprising molecular structures represented as spatial density functions,
[0407] - computing (630) an integral of the spatial density function obtained from the denoising and determining an integer value based on the result;
[0408] - fitting (640) a sum of a specified number of predefined density distributions, wherein the specified number corresponds to the determined integer value;
[0409] - identifying (650) centers of the fitted density distributions as positions of the molecular substructures of the molecular structure.
[0410] For example, the data structure representing the spatial function, parameters, e.g., weights, representing the neural network, and identified centers of the fitted density distributions may be stored in an electronic storage, e.g., a memory, a hard drive, etc. For example, applying a neural network to the spatial function, during generation or training, and / or adjusting the stored parameters to train the network may be done using an electronic computing device, e.g., a computer.
[0411] The neural networks, either during training and / or during applying may have multiple layers, which may include, e.g., convolutional layers and the like. For example, the neural network may have at least 2, 5, 10, 15, 20 or 40 hidden layers, or more, etc. The number of neurons in the neural network may, e.g., be at least 10, 100, 1000, 10000, 100000, 1000000, or more, etc.
[0412] Many different ways of executing the method are possible, as will be apparent to a person skilled in the art. For example, the steps can be performed in the shown order, but the order of the steps can be varied, or some steps may be executed in parallel. Moreover, in between steps other method steps may be inserted. The inserted steps may represent refinements of the method such as described herein, or may be unrelated to the method. For example, some steps may beR. 417936
[0413] - 45 -
[0414] executed, at least partially, in parallel. Moreover, a given step may not have finished completely before a next step is started.
[0415] Embodiments of the method may be executed using software, which comprises instructions for causing one or more computers, e.g., a processor system, to perform an embodiment of method 600. The software may include only those steps taken by a particular sub-entity of the system. The software and / or other data according to an embodiment may be stored in a non-transitory storage medium, such as a hard disk, a floppy disk, a memory, an optical disc, read-only memory, random access memory, CD-ROMs, magnetic tape, optical data storage devices, etc. Transitory signals and carrier waves are excluded from non-transitory media.
[0416] The software may be sent as a transitory signal along a wire or wirelessly, e.g., sent as a transitory signal over a data network, e.g., the Internet. For example, signals and / or carrier waves may serve as a transitory medium for carrying information. For example, a modulated electromagnetic wave may carry a signal bearing the software and / or other data according to an embodiment.
[0417] The software may be made available for download and / or for remote usage on a server. Embodiments of the method may be executed using a bitstream arranged to configure programmable logic, e.g., a field-programmable gate array (FPGA), to perform an embodiment of the method.
[0418] It will be appreciated that the presently disclosed subject matter also extends to computer programs, particularly computer programs on or in a carrier, adapted for putting the presently disclosed subject matter into practice. The program may be in the form of source code, object code, code intermediate between source and object code, such as partially compiled code, or in any other form suitable for use in the implementation of an embodiment of the method. An embodiment relating to a computer program product comprises computer-executable instructions corresponding to each of the processing steps of at least one of the methods set forth. These instructions may be subdivided into subroutines and / or be stored in one or more files that may be linked statically or dynamically.R. 417936
[0419] - 46 -
[0420] Another embodiment relating to a computer program product comprises computer-executable instructions corresponding to each of the devices, units, and / or parts of at least one of the systems and / or products set forth.
[0421] Figure 7a shows a computer-readable medium 1000 having a writable part 1010, and a computer-readable medium 1001 also having a writable part Computer-readable medium 1000 is shown in the form of an optically readable medium. Computer-readable medium 1001 is shown in the form of an electronic memory, in this case a memory card. Computer-readable mediums 1000 and 1001 may store data 1020 wherein the data may indicate instructions which, when executed by a processor system, cause a processor system to perform an embodiment of a method for generating molecular structures and / or a method for training a neural network to transform noisy spatial functions into structured spatial density functions. The computer program 1020 may be embodied on the computer-readable medium 1000 as physical marks or by magnetization of the computer-readable medium 1000. However, any other suitable embodiment is conceivable as well. Furthermore, it will be appreciated that, although the computer-readable medium 1000 is shown here as an optical disc, the computer-readable medium 1000 may be any suitable computer-readable medium, such as a hard disk, solid-state memory, flash memory, etc., and may be non-recordable or recordable. The computer program 1020 comprises instructions for causing a processor system to perform an embodiment of said method for generating molecular structures and / or said method for training a neural network to transform noisy spatial functions into structured spatial density functions.
[0422] Figure 7b shows a schematic representation of a processor system 1140 according to an embodiment. The processor system comprises one or more integrated circuits 1110. The architecture of the one or more integrated circuits 1110 is schematically shown in Figure 7b. Integrated circuits 1110 comprises a processing unit 1120, e.g., a processor, a CPU, for running computer program components to execute a method according to an embodiment and / or implement its modules or units. Integrated circuits 1110 comprises a memory 1122 for storing programming code, data, etc. Part of memory 1122 may be read-only. Integrated circuits 1110 may comprise a communication element 1126, e.g., anR. 417936
[0423] - 47 -
[0424] antenna, connectors, or both, and the like. Integrated circuits 1110 may comprise a dedicated integrated circuit 1124 for performing part or all of the processing defined in the method. Processing unit 1120, memory 1122, dedicated IC 1124 and communication element 1126 may be connected to each other via an interconnect 1130, such as a bus. The processor system 1140 may be arranged for contact and / or contactless communication, using an antenna and / or connectors, respectively.
[0425] For example, in an embodiment, processor system 1140, e.g., the molecular structure generating system or device, and / or training system or device for training a neural network to transform noisy spatial functions into structured spatial density functions, may comprise a processor circuit and a memory circuit, the processor being arranged to execute software stored in the memory circuit. The memory circuit may be a ROM circuit, or a non-volatile memory, e.g., a flash memory. The memory circuit may be a volatile memory, e.g., an SRAM memory. In the latter case, the device may comprise a non-volatile software interface, e.g., a hard drive, a network interface, etc., arranged for providing the software.
[0426] While system 1140 is shown as including one of each described component, the various components may be duplicated in various embodiments. For example, the processing unit 1120 may include multiple microprocessors that are configured to independently execute the methods described herein or are configured to perform elements or subroutines of the methods described herein such that the multiple processors cooperate to achieve the functionality described herein. Further, where the system 1140 is implemented in a cloud computing system, the various hardware components may belong to separate physical systems. For example, the processing unit 1120 may include a first processor in a first server and a second processor in a second server.
[0427] It should be noted that the above-mentioned embodiments illustrate rather than limit the presently disclosed subject matter, and that those skilled in the art will be able to design many alternative embodiments.R. 417936
[0428] - 48 -
[0429] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. Use of the verb ‘comprise’ and its conjugations does not exclude the presence of elements or steps other than those stated in a claim. The article ‘a’ or ‘an’ preceding an element does not exclude the presence of a plurality of such elements. Expressions such as “at least one of’ when preceding a list of elements represent a selection of all or of any subset of elements from the list. For example, the expression, “at least one of A, B, and C” should be understood as including only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C. The presently disclosed subject matter may be implemented by hardware comprising several distinct elements, and by a suitably programmed computer. In the device claim enumerating several parts, several of these parts may be embodied by one and the same item of hardware. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0430] In the claims references in parentheses refer to reference signs in drawings of exemplifying embodiments or to formulas of embodiments, thus increasing the intelligibility of the claim. These references shall not be construed as limiting the claim.
Claims
R. 417936- 49 -Claims1. A computer-implemented method (600) for generating molecular structures, a generated molecular structure comprising a plurality of molecular substructure positions, the method comprising:generating (610) an initial noisy spatial density function and representing the spatial function in a data structure;progressively (620) denoising the spatial function by iteratively applying a trained neural network, wherein the neural network is trained to transform noisy density functions into structured spatial density functions based on training data comprising molecular structures represented as spatial density functions,computing (630) an integral of the spatial density function obtained from the denoising and determining an integer value based on the result;fitting (640) a sum of a specified number of predefined density distributions, wherein the specified number corresponds to the determined integer value; identifying (650) centers of the fitted density distributions as positions of the molecular substructures of the molecular structure.
2. A method as in Claim 1 , whereina noisy spatial density function is configured to represent an arbitrary number of molecular substructures positions, and / orthe trained neural network is configured to increase and / or decrease an integral of the spatial function in a denoising iteration.
3. A method as in any one of the preceding claims, comprisinginferring connectivity between identified molecular substructure positions by applying a computational bond determination method to construct a molecular graph.R. 417936- 50 -4. The method according to any one of the preceding claims, further comprising:generating one or more noisy feature fields, each feature field being a function from a spatial domain to real numbers or vector of real numbers; progressively denoising the one or more noisy feature fields by iteratively applying the trained neural network;sampling the one or more denoised feature fields at the identified positions of the molecular substructures to obtain feature values associated with each molecular substructure.
5. The method of Claim 4, wherein the feature values represent one or more of:atom type, charge, electronegativity, partial atomic charge, hybridization state, van der Waals radius, bonding preferences, and polarizability.
6. The method of any one of the preceding claims, wherein the spatial density function is defined over a 2- or 3-dimensional space.
7. The method of Claim 3, comprising applying a graph neural network to the molecular graph to predict molecular properties.
8. The method according to any one of the preceding claims, comprising inferring connectivity between the identified molecular substructure positions by applying a computational bond determination method to construct a spatial density function indicating a likelihood of a connection existing between two points,generating one or more noisy feature fields, each noisy feature field being a function from a spatial domain to real numbers;iteratively updating the one or more feature fields by integrating over the spatial density function indicating the likelihood of connection existing, the one or more noisy feature fields, and a trained function.
9. The method according to any one of the preceding claims, whereinthe spatial density function is represented by sampling at a regular grid, orR. 417936- 51 -the spatial function is represented as a series of positions of a plurality of points, and values of the spatial function, and wherein the trained neural network is configured to modify the positions and / or values of the plurality of points.
10. The method of any one of the preceding claims, wherein the generation of the molecular structure is conditioned on one or more conditioning parameters, the method further comprising:receiving one or more conditioning parameters specifying desired molecular properties or structural constraints;modifying the initial noisy spatial function using the one or more conditioning parameters and / or providing the one or more conditioning parameters to the trained neural network when applied to the spatial function, wherein the neural network is trained to transform noisy spatial functions into structured spatial density functions conditioned by the one or more conditioning parameters.
11. The method of Claim 10, wherein the one or more conditioning parameters comprise experimental data (e.g., comprising full or partial NMR spectra, mass spectra, MS fragmentation), wherein the method is configured to generate a molecular structure corresponding to the experimental data.
12. The method according to any one of the preceding claims, comprising generating a plurality of molecular structures according to the method according to any one of the preceding claims,applying a further chemical software, configured to receive a molecular structure, to the plurality of molecular structures to test, debug, stress-test, or benchmark the further chemical software13. The method according to any one of the preceding claims, comprising applying a retrosynthesis tool to the molecular structure, obtaining a synthetic route, and synthesizing a molecule according to the molecular structure.R. 417936- 52 -14. A computer-implemented method for training a neural network to transform noisy spatial functions into structured spatial density functions, the method comprisingobtaining a plurality of molecular structures comprising a plurality of molecular substructure positions, the molecular structures corresponding to structures of physical molecules,representing the molecular structures as spatial density functions, by adding a number of predefined density distributions fitted to the plurality of molecular substructure positions, wherein the number corresponds to the number of molecular substructures in the plurality of molecular substructures, and wherein means of the fitted density distributions correspond to positions of the molecular substructures,generating diffusion training data by progressively applying a noise perturbation process to the spatial density function, wherein the noise perturbation process iteratively transforms the spatial density function into a noisy spatial function;training the neural network using the diffusion training data, wherein the neural network is trained to reconstruct the original spatial density function by minimizing a loss function that penalizes deviations between the denoised spatial function and the original structured spatial density function.
15. The method of Claim 14, whereinthe noise perturbation process comprises applying a diffusion process to the spatial density function, wherein the diffusion process is governed by a stochastic equation that progressively disperses the spatial density function over time, and / orthe noise perturbation process comprises applying a heat dissipation process to the spatial density function.
16. One or more non-transitory computer-readable media and / or one or more transitory computer-readable media storing computer-executable instructions that, when executed by a computing system, cause the computing system to perform the method according to any one of the preceding claims.R. 417936- 53 -17. A system comprising: one or more processors; and one or more storage devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations according to the method according to any one of the preceding claims.