Methods, systems and computer programs for training and applying a machine-learning model for prediction of molecular properties of a molecule

The physics-informed semi-supervised learning method addresses the computational and generalization challenges in molecular property prediction by generating pseudo-labels from spatially perturbed atomic coordinates, enhancing model robustness and reducing computational overhead for accurate simulations in drug development and material synthesis.

WO2025157429A1PCT designated stage Publication Date: 2025-07-31NEC LAB EURO GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/070085
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-24
Filing Date
2024-07-16
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Conventional methods for predicting molecular properties using machine-learning models require significant computational resources and struggle with generalization and robustness issues due to sensitivity to atomic structure perturbations, limiting their application in simulations like molecular dynamics and material synthesis.

Method used

A physics-informed semi-supervised learning (PISSL) approach that generates pseudo-labels using spatially perturbed atomic coordinates, incorporating physics-based relationships to train machine-learning models with reduced computational overhead and improved robustness, utilizing methods like Taylor expansion and active learning to enhance model accuracy.

Benefits of technology

This approach reduces the computational burden and improves the accuracy and robustness of molecular property predictions, enabling longer and more stable simulations for applications such as drug development, material synthesis, and healthcare, by leveraging physics-informed pseudo-labels to augment training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024070085_31072025_PF_FP_ABST
    Figure EP2024070085_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method, system and computer program for training a machine-learning model for prediction of one or more molecular properties of a molecule, and a method, system and computer program for applying such a machine-learning model. The present invention optimizes molecular simulation. It can be used in a variety of applications including, but not limited to, several anticipated use cases in drug development, material synthesis, and in healthcare. The computer-implemented method for training the machine-learning model comprises obtaining (110) an initial training sample for training the machine-learning model, the initial training sample representing a first molecule and comprising atomic coordinate information, chemical element information, and information on one or more molecular properties of the first molecule, generating (120) spatially perturbed atomic coordinate information for a spatially perturbed variation of the first molecule, deterministically (140, 145) calculating or estimating at least one molecular property of the spatially perturbed variation of the first molecule based on the initial training sample and based on a spatial deviation between the atomic coordinate information of the first molecule and the spatially perturbed atomic coordinate information, calculating (160) a first partial loss function between the information on the one or more molecular properties of the first molecule included in the initial training sample and predicted information on the one or more molecular properties of the first molecule predicted by the machine-learning model, calculating (170) a second partial loss function between the deterministically calculated or estimated at least one molecular property of the spatially perturbed variation of the first molecule and a corresponding prediction of the at least one molecular property predicted by the machine-learning model, and adjusting (190) the machine-learning model based on a result of the first and second partial loss function.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS, SYSTEMS AND COMPUTER PROGRAMS FOR TRAINING AND APPLYING A MACHINE-LEARNING MODEL FOR PREDICTION OF MOLECULAR PROPERTIES OF A MOLECULE

[0002] The present invention relates to a method, system, and computer program for training a machine-learning model for prediction of one or more molecular properties of a molecule, and a method, system and computer program for applying such a machinelearning model.

[0003] Explicitly solving electronic Schrodinger equation (SE) is one of the most popular methods to obtain chemical properties of interacting many-body systems, such as molecules and metals, in both industrial and academic communities. However, quantum mechanical solvers, such as the density functional theory (DFT), in general, demands a huge numerical overhead, scaling often cubically OQV3) with the number of atoms N in the atomic system. This prohibits practitioners to perform the simulation for realistic system sizes and demands them to utilize high-performance computing system. To solve this challenge, the interest on applying ML (Machine Learning) methods to estimate chemical properties has recently been growing as an alternative method (V. Zaverkin and J. Kastner, Gaussian moments as physically inspired molecular descriptors for accurate and scalable machine learning potentials. J. Chem. Theory Comput. 16, 5410-5421 (2020)). Thanks to the ability to emulate physics of an atomic system in a data-driven way, ML models can make iterative calculation of atomic forces in large systems, required to propagate the systems during an atomistic simulation and computing molecular properties, feasible.

[0004] However, in present ML approaches, a large number of numerical simulations, involving computationally expensive solutions of the Schrodinger equitation, are still necessary to get reasonably sized training data. Moreover, the generalization (extrapolation) challenge appears to be of immense importance, as most ML models are sensitive to outlier and local perturbations of atomic structures, given the fact that no data generation technique can guarantee complete coverage of physical space of interest. It is an object of the present invention to improve and further develop methods, systems, and computer programs for determining molecular properties of a molecule, which require a lower number and / or complexity of computations and / or which address the generalization challenge.

[0005] In accordance with the invention, the aforementioned object is accomplished by a method comprising the features of claim 1. According to this claim, a computer- implemented method for training a machine-learning model for prediction of one or more molecular properties of a molecule comprises obtaining an initial training sample for training the machine-learning model. The initial training sample represents a first molecule and comprises atomic coordinate information, chemical element information, and information on one or more molecular properties of the first molecule. The method comprises generating spatially perturbed atomic coordinate information for a spatially perturbed variation of the first molecule. The method comprises deterministically calculating or estimating at least one molecular property of the spatially perturbed variation of the first molecule based on the initial training sample and based on a spatial deviation between the atomic coordinate information of the first molecule and the spatially perturbed atomic coordinate information. The method comprises calculating a first partial loss function between (i.e. , based on, e.g., based on a difference between) the information on the one or more molecular properties of the first molecule included in the initial training sample and predicted information on the one or more molecular properties of the first molecule predicted by the machinelearning model. The method comprises calculating a second partial loss function between the deterministically calculated or estimated at least one molecular property of the spatially perturbed variation of the first molecule and a corresponding prediction of the at least one molecular property predicted by the machine-learning model. The method comprises adjusting the machine-learning model based on a result of the first and second partial loss function.

[0006] This way, the impact of each training sample is duplicated - both the molecule included in the initial training sample and the spatially perturbed variation that is derived, using a physics-informed method, from the molecule included in the initial training sample are used to train the machine-learning model, with only a low computational overhead for deriving the spatially perturbed variation of the first molecule from the first molecule. This lowers the computational effort required for training the machine-learning model, and thus the overall computational effort. Examples of the present disclosure thus provide methods and systems address the above objective by utilizing a physics-informed approach to obtain pseudo-labels (i.e. , the at least one molecular property of the spatially perturbed variation of the first molecule). These labels increase the effective size of the training dataset and improve the robustness of ML-based models to extrapolative samples. The present invention improves the accuracy of molecular simulation. It can be used in a variety of applications including, but not limited to, several anticipated use cases in drug development, material synthesis, and in healthcare.

[0007] In general, the machine-learning model may be trained to output, based on the atomic coordinate information, and based on the chemical element information, the one or more molecular properties, with the one or more molecular properties comprising the at least one molecular property corresponding to the deterministically calculated or estimated at least one molecular property. This way, the machine-learning model can be used to predict the one or more molecular properties, with a computational effort that is smaller than the effort required for determining the one or more properties using the Schrodinger equation.

[0008] In some examples, the one or more molecular properties comprise at least one of a potential energy of the molecule and a force-field of the molecule. Both the potential energy and the force-field of the molecule can be deterministically calculated or deterministically estimated with a low computational effort, for a spatially perturbed variation of the first molecule being derived from a first molecule, from the molecular properties of the first molecule.

[0009] For example, the at least one molecular property being deterministically calculated for the spatially perturbed atomic coordinate information may comprise the potential energy of the spatially perturbed variation of the first molecule. For example, the potential energy may be deterministically calculated using a Taylor expansion, based on the potential energy of the first molecule, based on a force-field of the first molecule, and based on a spatial deviation between the atomic coordinates of the first and spatially perturbed variation of the first molecules. This way, the potential energy can be calculated at a low computational complexity and effort and be used to further train the machine-learning model. For example, the at least one molecular property being deterministically estimated (e.g., by deterministically calculating the potential energy and estimating / approximating the force field using deterministic calculations) for the spatially perturbed atomic coordinate information may comprise the force-field of the spatially perturbed variation of the first molecule. In particular, the force-field of the spatially perturbed variation of the first molecule may be deterministically estimated for the spatially perturbed atomic coordinate information based on a gradient of the potential energy of the spatially perturbed variation of the first molecule, with the potential energy being deterministically calculated for the spatially perturbed atomic coordinate information. Again, both the determination of the potential energy for the spatially perturbed variation of the first molecule and the derivation of the force-field can be performed at a low computational complexity and effort and be used to further train the machine-learning model.

[0010] In some examples, the initial training sample comprises the molecular property or properties that are relevant for deterministically calculating or estimating the at least one molecular property of the spatially perturbed variation of the first molecule. In this case, the at least one molecular property of the spatially perturbed variation of the first molecule may be deterministically calculated based on the information on the one or more molecular properties included in the initial training sample. This way, no additional calculations are necessary for determining the one or more molecular properties being used for the deterministic calculation or estimation of the at least one molecular property of the spatially perturbed variation of the first molecule. In some examples, however, the necessary information might not be part of the initial training sample. In this case, the one or more molecular properties may be predicted by the machine-learning model itself. Accordingly, the at least one molecular property of the spatially perturbed variation of the first molecule may be deterministically calculated based on predicted information on at least one of the one more molecular properties predicted by the machine-learning model. This way, the proposed concept can be implemented even if the one or more molecular properties are not contained in the initial training sample.

[0011] The spatial deviation between the first molecule and spatially perturbed variation of the first molecule are used to deterministically derive the at least one molecular property of the spatially perturbed variation of the first molecule from the one or more molecular properties of the first molecule. This spatial deviation may be quantified using a spatial deviation vector. For example, generating the spatially perturbed atomic coordinate information may comprise generating a spatial deviation vector representing the spatial deviation between the atomic coordinate information of the first molecule and the spatially perturbed atomic coordinate information. The spatial deviation vector may then be used for determining the at least one molecular property of the spatially perturbed variation of the first molecule.

[0012] In some examples, not only the deterministic calculation of the at least one molecular property may be used to augment the training data, but a second approach may be used. For example, the method may comprise generating second spatially perturbed atomic coordinate information for a second spatially perturbed variation of the first molecule and predicting information on the one or more molecular properties of the second spatially perturbed variation of the first molecule using the machine-learning model based on the second spatially perturbed atomic coordinate information. The method may comprise deterministically calculating or estimating a first version of at least one molecular property of a target spatially perturbed variation of the first molecule (such as the spatially perturbed variation of the first molecule) based on the predicted information on the one or more molecular properties of the second spatially perturbed variation of the first molecule. The method may comprise predicting a second version of the at least one molecular property of the target spatially perturbed variation of the first molecule based on the initial training sample or based on predicted information on the one or more molecular properties of a third spatially perturbed variation of the first molecule. The method may comprise calculating a third partial loss function based on a difference between the first and second version of the at least one molecular property of the target spatially perturbed variation of the first molecule. The method may comprise adjusting the machine-learning model based on a result of the first, second and third partial loss function. In this approach, in the following also denoted “consistency pseudo label generation”, one or more additional data points are generated as basis for the estimation of the at least one molecular property of the target spatially perturbed variation of the first molecule, which may be the spatially perturbed variation of the first molecule or another molecule. To increase the training impact, the atomic coordinate information might not be varied randomly, but into a direction that is expected to have a larger impact. For example, to generate the spatially perturbed atomic coordinate information and / or second spatially perturbed atomic coordinate information, the atomic coordinate information of the first molecule may be varied in a direction that leads to larger changes in the loss function compared to other directions, or in a direction in which an impact of the variation on the loss function is more uncertain compared to other directions. This way, a fewer number of training samples may be required, and / or generalization of the resulting model may be improved.

[0013] In some examples, the method may comprise storing an additional training sample representing the spatially perturbed variation of the first molecule. This additional training sample may be used, for example, in additional training sessions and / or for training a different machine-learning model.

[0014] For example, the spatially perturbed atomic coordinate information may be generated by an active learning algorithm. Active learning is a subfield of machine learning that focuses on how machines can learn more efficiently. In particular, by letting an active learning algorithm choose which training samples are used to train the machinelearning model, a greater accuracy with fewer training samples can be achieved. Active learning can be achieved by training the machine-learning model on a small subset of labels, selecting a set of additional (unlabeled) data points that promise to have a large impact on the training (e.g., using uncertainty sampling, query by committee etc.), labeling the selected additional data points, and retraining or finetuning the machine-learning model. This process is repeated, with the active learning algorithm selecting the data points to be labeled. In the present context, this can be done by letting the active learning algorithm generate the spatial deviation vector for the (second) spatially perturbed atomic coordinate information.

[0015] In general, the training of machine-learning models is based on a large number of training samples. Accordingly, the computer-implemented method may be repeated for a plurality of initial training samples of a set of training data for training the machine-learning model. Some aspects of the present disclosure relate to a method for applying the machinelearning model trained for prediction of one or more molecular properties of a molecule. The method comprises obtaining the machine-learning model trained according to the above method and predicting at least one molecular property of at least one molecule using the machine-learning model. This way, the trained machinelearning model may be applied on new molecules, to determine their properties with a computational effort that is lower than conventional calculations using Schrodinger’s equations.

[0016] The determined property or properties may be used for various applications. For example, the method may comprise using the predicted at least one molecular property for at least one of a molecular dynamic simulation, a Monte-Carlo-based molecular simulation, identifying a material with a desired property, and a simulation of a short-term dynamic evolution of a protein.

[0017] Another aspect of the present disclosure relates to a system comprising one or more processors and one or more storage devices. The system is configured to perform at least one of the above methods.

[0018] Another aspect of the present disclosure relates to a computer program comprising instructions which, when the program may be executed by a computer, cause the computer to carry out at least one of the above methods.

[0019] Another aspect of the present disclosure relates to a non-transitory, computer- readable medium comprising a program code that, when the program code may be executed on a processor, a computer, or a programmable hardware component, causes the processor, computer, or programmable hardware component to perform at least one of the above methods.

[0020] There are several ways how to design and further develop the teaching of the present invention in an advantageous way. To this end, it is to be referred to the patent claims subordinate to patent claim 1 on the one hand and to the following explanation of preferred examples of embodiments of the invention, illustrated by the drawing on the other hand. In connection with the explanation of the preferred embodiments of the invention by the aid of the drawing, generally preferred embodiments and further developments of the teaching will be explained. In the drawing

[0021] Figs. 1 a and 1 b show flow charts of examples of a computer-implemented method for training a machine-learning model for prediction of one or more molecular properties of a molecule;

[0022] Fig. 1c shows a schematic diagram of a system for training a machinelearning model for prediction of one or more molecular properties of a molecule;

[0023] Figs. 2a and 2b show schematic illustrations of Physics-Informed Semi-Supervised Learning (PISSL)-based training of ML-based models;

[0024] Fig. 3 shows a schematic illustration of employing the model during inference;

[0025] Fig. 4a shows a flow chart of an example of a computer-implemented method for applying a machine-learning model trained for prediction of one or more molecular properties of a molecule;

[0026] Fig. 4b shows a schematic diagram of a system for applying a machinelearning model trained for prediction of one or more molecular properties of a molecule;

[0027] Fig. 5 shows a schematic drawing of an atomic coordinate perturbation;

[0028] Fig. 6a to 6c show schematic diagrams of three examples of a pseudo-label generation block;

[0029] Figs. 7a to 7c show schematic illustrations of examples data points generated via the consistency loss;

[0030] Fig. 8 shows a schematic diagram of an example of a Consistency Pseudo¬

[0031] Label Generation block; Fig. 9 shows a table of root mean square errors on the Ani-1x dataset;

[0032] Fig. 10 shows plots of the performance gain by PICPS approach with respect to different numbers of training samples;

[0033] Fig. 11 shows a table of root mean square errors on the TiO2 dataset;

[0034] Fig. 12 shows a table of root mean square errors on Aspirin and Benzene in the rMD17 dataset;

[0035] Fig. 13 shows a result of an ablation study;

[0036] Fig. 14 shows a table of a comparison of a performance of a physics- informed pseudo-label and a relational-consistent pseudo label; and

[0037] Fig. 15 shows an experimental result of the spatial-deviation vector selection dependence.

[0038] Embodiments of the present disclosure relate to a computer-implemented physics- informed semi-supervised learning (PISSL) method and system to train machine learning (ML) models for molecular property prediction. PISSL helps make ML-based models robust to the local perturbations of the atomic structure. The improved robustness allows performing stable atomistic simulations for a longer time, making the property prediction task simpler and data set generation easier. Embodiments of the present disclosure allow smaller training data sets than conventional supervised learning approaches, reducing the computational overhead for applications where obtaining new labels (e.g., reference energies, atomic forces, etc., at more accurate levels of theory, such as coupled cluster method) is challenging.

[0039] Let us consider the case of developing an ML model for a new material, enabling the prediction of its potential energy and atomic forces. Such a model is often referred to as force-field model. Conventionally, the training procedure starts from performing single point calculations, i.e., solving SE for candidate samples, to obtain required training data. Because of the high numerical cost of the quantum mechanical solver, it is usually difficult to collect large amount of data in the reasonable amount of time and the obtained ML models often suffer from generalization problem. Although utilizing active learning method can alleviate this problem, the computational demand for solving SE cannot be negligible, in particular, when considering a system with the large number of atoms, such as proteins.

[0040] In this document, a physics-informed pseudo-labelling method is proposed. According to an embodiment, the method includes estimating a pseudo-label in the neighborhood of the training data sample utilizing physical and mathematical relations. Because of considering physical and mathematical principles, the accuracy of the obtained pseudo-label can be rigorously validated, which allows to train ML models even with small training data sets. Moreover, embodiments of the method disclosed herein can be combined with any existing state-of-the-art models, i.e. , it is model agnostic. Finally, embodiments of the method disclosed herein improve the “robustness” of the ML potential model prediction, that is, reducing the cases of unphysical predictions by the ML potential model thanks to the pseudo-label generated with respect to a perturbation on the atomic position of the molecules.

[0041] For example, the proposed method may be implemented by the method and system of Figs. 1 a to 1c.

[0042] Figs. 1 a and 1 b show flow charts of examples of a computer-implemented method for training a machine-learning model for prediction of one or more molecular properties of a molecule. Such a machine-learning model the machine-learning model may be trained to output, based on atomic coordinate information, and based on chemical element information, one or more molecular properties, including a property being described, in the following, as deterministically calculated or estimated at least one molecular property. For example, the machine-learning model may be machinelearning model for predicting one or more properties, such as potential energy and / or force-field, of molecules being input into the machine-learning model.

[0043] The method comprises obtaining 110 an initial training sample for training the machine-learning model, with the initial training sample representing a first molecule. The initial training sample comprises atomic coordinate information (i.e., how the atoms of the molecule are arranged vis-a-vis each other), chemical element information (e.g., the chemical elements of the atoms), and information on one or more molecular properties of the first molecule (such as potential energy and / or force field). The method comprises generating 120 spatially perturbed atomic coordinate information for a spatially perturbed variation of the first molecule. The method comprises deterministically calculating 140 or estimating 145 (e.g., using a physics- informed calculation) at least one molecular property of the spatially perturbed variation of the first molecule based on the initial training sample and based on a spatial deviation between the atomic coordinate information of the first molecule and the spatially perturbed atomic coordinate information. The method comprises calculating 160 a first partial loss function between the information on the one or more molecular properties of the first molecule included in the initial training sample and predicted information on the one or more molecular properties of the first molecule predicted by the machine-learning model. The method comprises calculating 170 a second partial loss function between the deterministically calculated or estimated at least one molecular property of the spatially perturbed variation of the first molecule and a corresponding prediction of the at least one molecular property predicted by the machine-learning model. The method comprises adjusting 190 the machine-learning model based on a result of the first and second partial loss function. For example, the computer-implemented method may be repeated for a plurality of initial training samples of a set of training data for training the machine-learning model.

[0044] Fig. 1c shows a schematic diagram of a corresponding system 10, e.g., a computer system, for training a machine-learning model for prediction of one or more molecular properties of a molecule, with the system being configured to perform the method of Fig. 1 a and / or 1 b. The system 10 comprises an optional interface 12 and one or more processors 14 to perform the method of Figs. 1 a and / or 1 b. In some examples, the system 10 may further perform the method of Fig. 4a. Features introduced in connection with the methods of Fig. 1 a, 1 b and / or 4a, and with respect to Figs. 2a-3, 5 to 15 may likewise be performed by the system 10 of Fig. 1c. For example, the system 10 may comprise machine-readable instructions for providing the functionality of the method, with the one or more processors 14 executing the machine-readable instructions to perform the method of Figs. 1 a and / or 1 b. The one or more processors 14 are coupled with the interface 12 and with memory and / or storage, such as one or more storage devices 16. For example, the one or more processors 14 may be configured to provide the functionality of the system 10, e.g., in conjunction with the interface 12 (for exchanging information) and / or with the memory / storage 16 (for storing information).

[0045] For example, the interface 12 may include or correspond to a network interface circuitry and / or a device interface circuitry configured to be communicatively coupled to one or more other devices, such as the one or more processors 14. For example, the interface 12 may include a transmitter, a receiver, or a combination thereof (e.g., a transceiver), and may enable wired communication, wireless communication, or a combination thereof. For example, the one or more processors 14 may include or correspond to a digital signal processor circuitry (DSP), a graphical processing unit (GPU), and / or a central processing unit (CPU). The one or more processors 14 may be coupled to the memory / storage circuitry(s) 16. The memory / storage circuitry 16 may include instructions (e.g., executable instructions), such as computer-readable instructions or processor circuitry-readable instructions. The instructions may include one or more instructions that are executable by a computer, such as by the one or more processors 14. For example, the memory / storage circuitry 16 may include or correspond to volatile or nonvolatile storage circuitry, such as Random Access Memory (RAM), magnetic disks, optical disks, or flash memory devices. The memory / storage circuitry 16 may include both removable and non-removable memory devices.

[0046] Embodiments of the present disclosure relate to a method and to a system for performing the method, with the system comprising an ML potential model (the machine-learning model), a Pseudo-Label Generation block (for deterministically calculating and / or estimating the at least one molecular property, and an optional Consistent-Pseudo-Label Generation block, that may work together as follows:

[0047] 1 ) ML potential model: This model may accept the atomic coordinate ( / ?) and chemical element value (Z) information of the input molecule and estimates molecular properties, such as the potential energy (E), force-field (F), etc. (in the following, denotes other possible molecular or material properties).

[0048] 2) Pseudo-Label generation block: This block may accept the atomic positions and atomic numbers of an input molecule and generates new atomic positions and the corresponding potential energy. 3) Consistent-Pseudo-Label Generation block: This optional block may accept atomic positions of the input structure, new atomic positions generated by the pseudo-label generation block, and chemical element value of the input molecule and the potential energy of the molecule with the newly generated atomic positions with a different path (see Figs. 7a to 7c).

[0049] 4) Datasets: A storage of molecular data samples comprising (R, Z; E, F, Mi) for each sample where R is the atomic coordinate, Z is chemical element information, E is the potential energy of the molecule, F is the force-field, and Mi denotes other possible molecular or material properties.

[0050] Figs. 2a and 2b show schematic illustrations of a Physics-Informed Semi-Supervised Learning (PISSL)-based training of ML-based models, such as the machine-learning model. Figs. 2a and 2b show the aforementioned components, i.e. , the datasets, the ML model (F) and the Pseudo-Label generation block (PL). Fig. 2b additionally shows the Consistent-Pseudo-Label Generation block (CPL).

[0051] The ML potential model parameters may be updated (i.e., adjusted) using a gradient descent method, such as the stochastic gradient descent:

[0052] Li, L2 , L3 are the loss functions, such as the mean squared error and the mean absolute error, as proposed by Cooper et al.: “Efficient training of ANN potentials by including atomic forces via Taylor expansion and application to water and a transitionmetal oxide.” npj Computational Materials 6 (1 ), 54 (2020). For example, L1 may correspond to the first partial loss function of the method of Fig.1 a / 1 b, L2 may correspond to the second partial loss function of the method of Fig. 1 a / ab, and L3 may correspond to an optional third partial loss function of the method of Fig. 1 b. The detailed expression of L2 and L3 are described below.

[0053] During inference, the proposed system becomes the usual ML potential model, which accepts the atomic positions and numbers of an input molecule and output the molecule properties, such as the potential energy. Fig. 3 illustrates an example of using the model during inference. In the example shown in Fig. 3, the atomic coordinates and chemical element information is sued as input for the ML potential model F, which is used to predict the molecular properties of the molecule, such as the potential energy Epred or the force-field Fpred.

[0054] Application (i.e., inference) of the machine-learning model is also shown in Figs. 4a and 4b. Fig. 4a shows a flow chart of an example of a computer-implemented method for applying a machine-learning model trained for prediction of one or more molecular properties of a molecule. The method comprises obtaining 410 a machine-learning model trained according to the method of Figs. 1a and / or 1 b. The method comprises predicting 420 at least one molecular property of at least one molecule using the machine-learning model. Optionally, the method may further comprise using 430 the predicted at least one molecular property for at least one of a molecular dynamic simulation, a Monte-Carlo-based molecular simulation, identifying a material with a desired property, and a simulation of a short-term dynamic evolution of a protein.

[0055] Fig. 4b shows a schematic diagram of a corresponding system for applying a machine-learning model trained for prediction of one or more molecular properties of a molecule, with the system being configured to perform the method of Fig. 4a. The system 40 comprises an optional interface 42 and one or more processors 44 to perform the method of Fig. 4a. In some examples, the system 40 may further perform the method of Figs. 1a and / or 1 b. Features introduced in connection with the methods of Fig. 1 a, 1 b and / or 4a, and with respect to Figs. 2a-3, 5 to 15 may likewise be performed by the system 40 of Fig. 4b. For example, the system 40 may comprise machine-readable instructions for providing the functionality of the method, with the one or more processors 44 executing the machine-readable instructions to perform the method of Fig. 4a. The one or more processors 44 are coupled with the interface 42 and with memory and / or storage, such as one or more storage devices 46. For example, the one or more processors 44 may be configured to provide the functionality of the system 40, e.g., in conjunction with the interface 42 (for exchanging information) and / or with the memory / storage 46 (for storing information).

[0056] For example, the interface 42 may include or correspond to a network interface circuitry and / or a device interface circuitry configured to be communicatively coupled to one or more other devices, such as the one or more processors 44. For example, the interface 42 may include a transmitter, a receiver, or a combination thereof (e.g., a transceiver), and may enable wired communication, wireless communication, or a combination thereof. For example, the one or more processors 44 may include or correspond to a digital signal processor circuitry (DSP), a graphical processing unit (GPU), and / or a central processing unit (CPU). The one or more processors 44 may be coupled to the memory / storage circuitry(s) 46. The memory / storage circuitry 46 may include instructions (e.g., executable instructions), such as computer-readable instructions or processor circuitry-readable instructions. The instructions may include one or more instructions that are executable by a computer, such as by the one or more processors 44. For example, the memory / storage circuitry 46 may include or correspond to volatile or nonvolatile storage circuitry, such as Random Access Memory (RAM), magnetic disks, optical disks, or flash memory devices. The memory / storage circuitry 46 may include both removable and non-removable memory devices.

[0057] In the following, the deterministic calculation or estimation of the at least one molecular property, and in particular of the potential energy and the force-field, are introduced in more detail with respect to the Pseudo-Label Generation block.

[0058] The Pseudo-Label Generation block estimates a pseudo label, that is, the potential energy and / or force field of the input molecule with a perturbed atomic coordinate (see Fig. 5, which shows a schematic drawing of an atomic coordinate perturbation). According to an example, the pseudo label (and in particular the potential energy) may be estimated utilizing the Taylor expansion as:

[0059] F(F + <5F,Z) = F(F) + VF(F) ■ 6R + O(<5 ?2) = F(F) - F(F) ■ 6R + O(<5 ?2) (3)

[0060] In other words, the potential energy may be deterministically calculated 140 using a Taylor expansion, based on the potential energy of the first molecule F(F), based on a force-field of the first molecule F(F), and based on a spatial deviation 6R between the atomic coordinates of the first molecule and the spatially perturbed variation of the first molecule. For example, generating 120 the spatially perturbed atomic coordinate information may comprise generating the spatial deviation vector 6R representing the spatial deviation between the atomic coordinate information of the first molecule and the spatially perturbed atomic coordinate information. Then, the loss function L2 can be set as: where 12 is the loss function, such as the mean squared error and the mean absolute error.

[0061] In addition to the potential energy, the force-field of the molecule with the perturbed atomic coordinate (e.g., the spatially perturbed variation of the first molecule) may be determined. The force-field of the spatially perturbed variation of the first molecule may be deterministically estimated 145 for the spatially perturbed atomic coordinate information based on a gradient, such as the spatial derivative, of the potential energy of the spatially perturbed variation of the first molecule, with the potential energy being deterministically calculated for the spatially perturbed atomic coordinate information (e.g., using the above Taylor expansion). For example, to derive the force-field at each atomic position, one approach is to compute the gradient of the potential energy with respect to the atomic position rz: F, = -VjEj. Alternatively, the force-field may be estimated directly from the potential energy of the spatially perturbed variation of the first molecule or from the atomic coordinate information and element information of the spatially perturbed variation of the first molecule, e.g., using another model. This way, the at least one molecular property being deterministically estimated 145 for the spatially perturbed atomic coordinate information may comprise the force-field of the spatially perturbed variation of the first molecule.

[0062] One way of implementing Pseudo-Label Generation block is shown in Fig. 6a. In this case, the potential energy and force field at the original atomic coordinate R are the true value saved in the Dataset. In other words, the at least one molecular property of the spatially perturbed variation of the first molecule may be deterministically calculated 140 or estimated 145 based on the information on the one or more molecular properties included in the initial training sample (i.e., the training sample included in the dataset). First, the spatial deviation vector generation block generates the spatial deviation vector 6R which can be generated as random vector, the adversarial direction (the most sensitive direction of the loss function in terms of an input), and the most uncertain direction with utilizing uncertain score, such as the standard deviation of the multiple model prediction. In other words, to generate the spatially perturbed atomic coordinate information and / or second spatially perturbed atomic coordinate information, the atomic coordinate information of the first molecule may be varied in a direction that leads to larger changes in the loss function compared to other directions, or in a direction in which an impact of the variation on the loss function may be more uncertain compared to other directions. Next, the Taylor Expansion block calculates the new potential energy E R + 6R) following Equation (3).

[0063] Another way of implementing Pseudo-Label Generation block is given in Fig. 6b. In this case, the potential energy and force field at the original atomic coordinate R are estimated by the ML potential model directly, without utilizing auto-differentiation of the potential field to obtain the force field (see also Fig. 6c). In other words, the at least one molecular property of the spatially perturbed variation of the first molecule may be deterministically calculated 140 or estimated 145 based on predicted information on at least one of the one more molecular properties predicted by the machine-learning model. The other part is the same as the Example 1 .

[0064] Another way of implementing Pseudo-Label Generation block is given in Fig. 6c. In this case, the potential energy at the original atomic coordinate R is estimated by the ML potential model but force field F is estimated by performing the gradient of the potential energy by the Auto-Differentiation block. Again, the at least one molecular property of the spatially perturbed variation of the first molecule may be deterministically calculated 140 or estimated 145 based on predicted information on at least one of the one more molecular properties predicted by the machine-learning model. The other part is the same as the Example 1 .

[0065] In the context of the present disclosure, the term ..deterministically calculated" is used for determining the potential energy, e.g., using the above equations, and ..deterministically estimated" is used for determining the force field, which is determined, via an approximation, using the gradient of the potential energy, which is a deterministic estimation or approximation. In the context of the present disclosure, the terms ..deterministically calculated" and “deterministically estimated” indicate that the respective properties are not determined using probabilistic or predictive (machine-learning) models. Optionally, the pseudo-labels may be saved into a storage. Accordingly, as further shown in Fig. 1 b, the method may comprise storing 150 an additional training sample representing the spatially perturbed variation of the first molecule.

[0066] In the following, an example of the Consistency Pseudo-Label Generation block is given. In the context of the method of Fig. 1 b, the Consistency Pseudo-Label Generation block is used to calculate a third loss function, which is based on the insight that the prediction of molecular properties should yield the same result, regardless of which molecule is used as a starting point. For this purpose, the properties of a molecule are determined starting from two different starting points (i.e. , two spatially different versions of the molecule, having the same atomic elements, but different atomic coordinates). The loss is based on the difference between the results yielded by the two different starting points.

[0067] For example, with respect to the method of Fig. 1 b, the method may comprise generating 130 second spatially perturbed atomic coordinate information for a second spatially perturbed variation of the first molecule. The method may comprise predicting 132 information on the one or more molecular properties of the second spatially perturbed variation of the first molecule using the machine-learning model based on the second spatially perturbed atomic coordinate information. This second spatially perturbed variation of the first molecule is denoted estimation data point B. EB, FB). The method may comprise deterministically calculating 140 or estimating 145 a first version of at least one molecular property of a target spatially perturbed variation of the first molecule based on the predicted information on the one or more molecular properties of the second spatially perturbed variation of the first molecule (i.e., the pseudo-label is determined starting from the second spatially perturbed variation of the first molecule, whose properties are predicted by the machine-learning model). The method may further comprise predicting 134 (e.g., using the pseudolabel generation block) a second version of the at least one molecular property of the target spatially perturbed variation of the first molecule.

[0068] Predicting 134 the second version of the at least one molecular property may be done based on the initial training sample (see Fig. 7a, in this case, the spatially perturbed variation of the first molecule may be the target spatially perturbed variation of the first molecule) or based on predicted information on the one or more molecular properties of a third spatially perturbed variation of the first molecule (see Fig. 7b). This third spatially perturbed variationof the first molecule, denoted estimation point D. in Fig. 7b, can be used instead of the (non-spatially perturbed) first molecule as a starting point for calculating the second version of the at least one molecular property. The method may comprise calculating 180 a third partial loss function (L3) based on a difference between the first and second version of the at least one molecular property of the target spatially perturbed variation of the first molecule. In Fig. 7a, data point A. (F^F^) corresponds to the first molecule, extrapolation point C. (EC, FC) corresponds to the spatially perturbed variation of the first molecule, and estimation point B (Fs, Fs) corresponds to the second spatially perturbed variation of the first molecule. For example, in Fig. 7a, extrapolation point C. may be the target spatially perturbed variation. For example, one of upcoming equations 5 to 7 can be used to determine the third partial loss function. In Fig. 7b, estimation point D. ED, FD) corresponds to the third spatially perturbed variation of the first molecule, while data point A. (F^F^) corresponds to the first molecule, extrapolation point C. (EC, FC) corresponds to the spatially perturbed variation of the first molecule, and estimation point B. EB, FB) corresponds to the second spatially perturbed variation of the first molecule. In Fig. 7b, data point C. EC, FC) is the target spatially perturbed variation, with estimation points B. (Fs, FB) and D. (FD, FD) being used as starting points for the first and second version of the at least one molecular property. For example, equation 8 may be used to determine the third partial loss function. In Fig. 7c, data point A. (F^F^) corresponds to the first molecule, data point A’ EA', FAE) corresponds to a second spatially perturbed variation of the third molecule, for which ground-truth data exists, and extrapolation point C. EC, FC) corresponds to the spatially perturbed variation of the first molecule. In this case, equation 9 may be used to determine the third partial loss function. The method may comprise adjusting 190 the machine-learning model based on a result of the first, second and third partial loss function.

[0069] For the consistency pseudo label generation, a first example is provided in the Fig. 7a where A, B, C are the original atomic coordinate data point, the estimation point, and the target extrapolation point, respectively. At the estimation point B, the potential energy and the force field are estimated by the ML potential model (i.e., the second spatially perturbed atomic coordinate information of the estimation point (i.e., the second spatially perturbed variation of the first molecule) is determined, and used as input for the ML potential model to predict the potential energy and force field of the second spatially perturbed variation of the first molecule). The extrapolation point C corresponds to the point where the Pseudo label generation block estimates the potential and force field (i.e. , the spatially perturbed variation of the first molecule), so that the 6R in the Pseudo Label generation block corresponds with 6RACin Fig. 7a. According to embodiments, the consistency pseudo label generation block may estimate the potential field at the extrapolation point C with two different ways: the one is using Equation (3) from the data point A and the other is also using Equation (3) from the data point B. The consistency pseudo-label loss function L3 in Equation

[0070] (2) can be written as:

[0071] Here the potential field value obtained by Equation (3) is defined as VC\Awhere C and A are the extrapolation and data points, respectively.

[0072] Another way of the consistency pseudo potential loss function Equation (5) can be given as: where the potential energy at the extrapolation point C is evaluated by the Equation

[0073] (3) from the point B and directly by ML potential model (i.e., without using equation (3)).

[0074] Another way of the consistency pseudo potential loss function Equation (5) can be given as:

[0075] ^3=^3 (^8 < ^B|C) (7) where the potential energy at the estimation point B is evaluated by the Equation (3) from the point C and directly by ML potential model. Note that in the present disclosure, it is assumed that the pseudo label loss function Equation (4) is utilized at the point C but not at the point B.

[0076] Another way of the consistency pseudo potential loss function Equation (5) can be given as: where the potential energy at the extrapolation point C is evaluated by the Equation (3) from the point B and D. Note that the point D is also an estimation point depicted in Fig. 8. Another way of the consistency pseudo potential loss function Equation (5) can be given as:

[0077] L3= l3vC\A, VA,',) (9) where the potential energy at the extrapolation point C is evaluated by the Equation (3) from the point A and A’ both of which are data point. Note that this is only applicable if a training dataset includes several samples with the same molecule.

[0078] One way of implementing the consistency pseudo-label generation block is depicted in Fig. 8. In this example, the block accepts the atomic coordinate R, charge Z, and the spatial deviation vector 6R generated by the Pseudo Label generation block in Fig. 2a, 2b, 6a to 6c. Those inputs are accepted by the spatial deviation vector generation block that generates a new spatial deviation vector 6R to determine the coordinate of the estimation point B (the spatially perturbed variation of the first molecule) in the Figs. 7a and 7b. Note that in the case of Fig. 7b the spatial deviation vector generation block generates two vectors to determine the coordinate B and D (the spatially perturbed variation of the first molecule). Utilizing the generated spatial deviation vector, or the estimation point B and D, the pseudo-label generation block generates another estimation of the potential energy at the extrapolation point C, described as E’ R + 6R) in Fig. 7c. Again, the spatial deviation vectors can be determined by utilizing a random vector, the adversarial direction, and the most uncertain direction.

[0079] One way of implementing ML potential model is to use neural network (NN)-based models, including graph NNs ones, such as GMNN [1], GemNet, SchNet, Graphomer, and Equiformer.

[0080] For example, the PI-SSL method may be combined with active learning methods. For example, the spatially perturbed atomic coordinate information and / or second spatially perturbed atomic coordinate information may be generated by an active learning algorithm. Active learning is one of the popular methods to obtain an ML potential model with better accuracy using minimal amount of data generated by performing numerical calculation of quantum mechanics equation (V. Zaverkin, D. Holmuller, H. Christiansen, F. Errica, F. Alesiani, M. Takamoto, M. Niepert, J. Kastner. Uncertainty-biased molecular dynamics for learning uniformly accurate interatomic potentials. Submitted to npj Computational Materials, preprint: https: / / arxiv.org / abs / 2312.01416. (2023)). Embodiments of the PI-SSL method disclosed herein allow to reduce the necessary data generation process, reducing the total time necessary for obtaining an ML potential model with a sufficient accuracy.

[0081] Machine learning is a branch of artificial intelligence that involves the development of algorithms and models that allow computers to learn and make predictions or decisions without being explicitly programmed. It focuses on creating systems that can improve their performance over time by learning from data.

[0082] Training a machine-learning model refers to the process of teaching the model to make accurate predictions or decisions. During training, the model is exposed to a large amount of data, which is used to adjust the model's internal parameters or weights. The model learns patterns, relationships, or rules from the training data, allowing it to generalize and make predictions on new, unseen data.

[0083] Training data is the set of examples or instances that is used to teach a machinelearning model. It is often labeled data, meaning that each example is associated with a known outcome or target value. The training data consists of both input features and the corresponding output or target variable. The model learns from this data by analyzing the patterns and relationships between the input features and the target variable. Training algorithms, such as supervised learning, semi-supervised learning, unsupervised learning, or reinforcement learning may be used for training the machine-learning model.

[0084] In the present context, semi-supervised learning is used. The training is semisupervised as only the dataset with the initial training samples is originally known. The other elements being used for the training, such as the pseudo-labels (the at least one molecular property being deterministically determined), are generated to augment the original dataset. These additional elements are initially unlabeled, with the labels being determined using the physics-informed equations. Once the pseudolabels are generated, training can be performed using supervised learning.

[0085] Machine-learning models, such as the machine-learning model being trained in the present disclosure, are often implemented as Artificial Neural Networks (ANNs), and in particular Deep Neural Networks, Support Vector Machines, Decision Tree models, or Random Forest models.

[0086] The proposed concept may be applied as follows:

[0087] For example, the proposed concept may be used for molecular simulations for new drug discovery. Molecular dynamic (MD) simulations are used to simulate a process of an interaction between source and target molecules, such as drug and protein. For the moment, the MD simulation is accurate but computationally very heavy, which prohibits an efficient search of new drug. On the other hand, machine learning technique has started to be used in this area. However, people usually use classical force field that is a crude approximation of the quantum mechanical potential field because of the heavy numerical cost for performing the numerical calculation on the protein. In addition, the large size of the protein molecule prohibits to perform the numerical calculation on the sufficient amount of protein molecule to train ML potential field. Embodiments of the PI-SSL method disclosed herein can alleviate those problem because our method increases the label data without performing heavy numerical calculations. In addition, the ML potential field has a problem when applying to MD simulation, that is, unphysical prediction because of coming across unseen data during the training Embodiments of the PI-SSL method disclosed herein alleviate this problem thanks to the improvement of the model robustness and allows to perform longer and stable MD simulation that is essential for identifying an efficient drug.

[0088] For example, the proposed concept may be used for Molecular simulation for new material design. Similar to the above, our method can also be applied to find a new material with a desirable property. MD simulations are also used for finding the properties of a new material, such as elasticity. An ML potential model allows to obtain those properties much better accuracy than the classical force field and our method allows to obtain them with smaller numerical cost than the usual training method of the model.

[0089] For example, the proposed concept may be used for Monte-Carlo based Molecular simulation for new material discovery (H2-catalyst). The proposed concept can be applied not only for MD simulations but also Monte-Carlo based molecular simulation (MC simulation). In this case, the force-field for MD updater is substituted by a ML potential model trained with the PI-SSL method. The remaining part can be the same as the previous cases.

[0090] For example, the proposed concept may be used for Molecular simulation for Protein and mRNA Evolution. The method can also be used to simulate the short-term dynamic evolution of proteins and mRNA, which is important to understand the dynamic properties of the proteins, such as a target protein for a newly developed drug and the structural stability of mRNA vaccine. Although it is popular to use the classical force field for the dynamic simulation for proteins, it has been recently found that the classical force-field sometimes fails to simulate real proteins due to its insufficient accuracy, which prompts us to utilize the ML potential model for this application. In this case, the proposed PI-SSL method is very beneficial because of the very high numerical cost to perform quantum mechanical simulations on a large molecule, such as the protein, which prohibits us to obtain a large number of data point to train an ML potential model.

[0091] The proposed physics-informed semi-supervised learning (PISSL) method and system may comprise one or more of the following steps / components: A collection of experimental and simulation data for target molecules (i.e., the training data set). It may include building a new network for the ML potential model. It may comprise performing the Pseudo-Label generation process to collect the pseudo-label label into a storage, e.g., to reduce the memory usage during the model training. It may comprise training of the ML potential model over the dataset utilizing the PI-SSL method. The perturbed direction on the atomic position may be evaluated, utilizing random / adversarial / most uncertain direction. The Pseudo-Label generation process may be performed to collect the pseudo-label and store into memory and calculate the new loss functions for end-to-end training. For example, the trained ML potential model may be used for the MD / MC simulation or an inference task.

[0092] The proposed concept may use a new physics-informed semi-supervised training method to develop the ML potential model that is used to simulate a real-world system. The physics-informed semi-supervised training method allows generation of pseudo-labels with better accuracy than a pure ML driven method because of the physics informed method. This results in a smaller data size, as the pseudo-label generation and consistent pseudo-label generation allow to either reduce the necessary training data size or improve the accuracy with a fixed training data size. In contrast to other pseudo-labeling methods, the pseudo label accuracy itself improves during the training with the accuracy improvement of the ML potential model (providing end-to-end training. It can be used with any existing ML potential models. Increasing the amount of training data using the pseudo-labels reduces unphysical prediction of the model because of less unseen data during training, improving the robustness of the training. Robustness of the model prediction also allows a longer and more stable MD / MC simulations, allowing to find new long-time behavior.

[0093] The proposed concept allows to reduce a necessary data size or quantum mechanical simulations to obtain an ML potential model with a sufficient accuracy. It also improves the accuracy of an ML potential model using a data set with a fixed size. It may accelerate application of an ML potential model to a real-world problem, while improving the robustness of the ML potential model prediction, which may result in longer and more stable MD / MC simulations.

[0094] Many modifications and other embodiments of the invention set forth herein will come to mind the one skilled in the art to which the invention pertains having the benefit of the teachings presented in the foregoing description and the associated drawings. Therefore, it is to be understood that the invention is not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the invention. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0095] Further details and embodiments of the present invention are described in the following scientific discussion titled “A Consistency-Driven Physics-Informed Method for Improving Generalizability and Robustness to Perturbations of Machine-Learned Interatomic Potentials”:

[0096] Machine Learning (ML) models play an important role in the computational chemistry community as they help replace computationally demanding numerical simulations that involve explicit solving of the electronic Schrodinger equation. Despite their utility, these ML models confront challenges related to generalizability and robustness, leading to unphysical predictions that impede their real-world applicability, However, such ML models suffer from generalizability and robustness issues that result in unphysical prediction and hinder a real-world application, in particular such as the molecular dynamic simulation with sufficiently long-time evolution. A physics-informed consistent-pseudo-labelling approach is proposed for the training of ML models for molecular property prediction. The proposed method employs physics-based relationships to generate pseudo-labels for interatomic potential energy and formulate them as two distinct loss functions, physics-informed relation-consistency loss and physics-informed spatial-consistency loss. In particular, the proposed method enables the training of ML models with reduced training data requirements without the need for pretraining expensive models because of the physics-informed nature. Through a series of comprehensive experiments, we demonstrate the efficacy of our proposed method, revealing consistent enhancements over baseline models across diverse datasets.

[0097] Explicitly solving electronic Schrodinger equation is one of the most widely employed approaches for deriving chemical properties in interacting many-body systems, such as molecules and metals, within both industrial and academic communities. Nevertheless, quantum mechanical solvers, such as the density functional theory (DFT) typically demands substantial numerical overhead, exhibiting a cubic scaling of O( / V3) with the number of atoms / V in the atomic system. This computational demand prohibits practitioners to simulate realistic system sizes, forcing them to rely on high- performance computing system. To address this problem, there has been a growing interest in employing ML methods as an alternative method to estimate chemical properties. Thanks to the ability to emulate the physics of an atomic system in a data- driven manner, ML models enable iterative calculation of atomic forces in large systems. This capability is crucial for propagating systems during atomistic simulations and computing molecular properties.

[0098] Nevertheless, contemporary ML approaches encounter challenges, including the need for a substantial number of numerical simulations, which involve computationally expensive solutions to the Schrodinger equitation to obtain reasonably sized training datasets. Consequently, the development of ML potential models becomes an inherent small-data problem, particularly when investigating a new material lacking a comparable molecule database for transfer learning and semi-supervised learning. To mitigate this challenge, one prevalent strategy is the application of active learning (AL) methods. Nonetheless, AL still demands a non-negligible computation of quantum mechanical solvers, prompting active exploration of methods for further reducing the reliance on quantum calculations. Additionally, the issue of generalization (extrapolation) appears to be of immense importance, as most ML models are sensitive to outliers and local perturbations in atomic structures. This sensitivity arises due to the inherent limitation that no data generation technique can guarantee complete coverage of the physical space of interest.

[0099] In the present disclosure, an approach is proposed to address these issues through the utilization of a physics-informed consistent pseudo-labelling (PICPL) method. This method aims to derive pseudo-labels for interatomic potential energy, ensuring their quality in accordance with the principle of physics. This method employs physicsbased relationships to generate pseudo-labels for interatomic potential energy and formulate them as two distinct loss functions, physics-informed relation-consistency (PIRC) and physics-informed spatial-consistency (PISC) losses. The proposed method enables the training of ML models with reduced training data requirements without the need for pretraining expensive models and access to large training datasets because of the physics-informed nature. Specifically, the incorporation of pseudo-labels alleviates the sparsity issue associated with limited data availability and enhances the robustness of machine learning models to extrapolative samples.

[0100] The following contributions are made A novel method to estimate pseudo-label in physics-informed manner without accessing large training datasets is explored. Based on the formulation, two distinct loss functions are defined, the physics-informed relation-consistency loss and physics-informed spatial-consistency loss. Extensive experiments were performed to evaluate the effectiveness of the proposed method.

[0101] Recently, an interest on utilizing ML models to solve the Schrodinger equation for atomistic simulations is growing in order to simulate molecular systems more accurately while avoiding huge numerical cost. The initial work for this line of research has already appeared more than two decades ago in (Blank et al., 1995). A considerable number of investigations are conducted to develop ML models for estimating interatomic potential energy and force-field. To simulate the quantum system accurately, the community developed two breakthroughs. The first one is to develop the higher order local descriptor (Drautz, 2019; Zaverkin & Kastner, 2020) which enables to construct high body order complete polynomial basis function. The other one is to utilize the equivariant internal features in message-passing neural networks (MPNNs) which allows ML models to avoid learning Physics-Informed Machine Learning (Cooper et al., 2020): ANN PES with Taylor expansion.

[0102] In the following, a physics-informed Consistent-Pseudo-Labelling Method is proposed. In the following, a molecular system consisting of / V atoms is considered, where each atomic position in the molecule is represented as r = [ri, ..., rw] e R3XN, and each chemical element is denoted as z = e RN. Following

[0103] (Behler & Pamnello, 2007), it is assumed that the total energy E of a molecular system can be decomposed into the sum of atomic contribution as: where NN is a neural network, such as a graph neural network and Q is the trainable weight parameter.

[0104] To derive the force-field at each atomic position, one approach is to compute the gradient of the potential energy with respect to the atomic position rz:

[0105] Fi = - VJEJ (A2)

[0106] However, it is worth noting that some recent models do not utilize Equation A2 but directly estimate the force-field by the same neural network for the potential energy but with an additional output for the force-field as demonstrated by (Hu et al., 2021 ).

[0107] In the following, the proposed simple but very effective physics-informed relationconsistency (PIRC) loss methodology is explained, the proposed objective is to estimate a new potential energy for a molecule in the training dataset but with different atomic coordinate: r + <5r where <5r is a spatial-deviation vector for each atom. By performing the Taylor expansion with respect to <5r, the relationship between new potential energy E(r + <5r) and the original potential energy E(r) can be expressed as: where a physics relation, Equation A2, is utilized to derive the resulting equation. By neglecting the second-order small value 0(<5r2), the following loss function is proposed: (-44) where L denotes a distance metric function, such as mean- absolute error and root- mean-square error loss functions and Epred, Fpred are the predicted potential energy and force-field by a ML potential model. Note that, distinct from conventional SSL with pseudo-label approaches (Cooper et al., 2020), Equation A4 is purely composed by the predicted potential energy and force-field. This allows the ML model to learn the consistency between the predicted potential energy and force-field to satisfy Equation A3. And it also enables a gradient flow in terms of the force-field prediction F, enhancing accuracy not only in potential energy but also in force-field. The comparison to the conventional method is presented in connection with Figs. K to Q. The appropriate choice of the spatial deviation vector 6r and the validity of the neglecting the second-order small value 0(<5r2) are also discussed in connection with Figs. K to Q.

[0108] Next, another physics-informed consistent-pseudo-labelling loss methodology is explained, namely, the physics-informed spatial-consistency (RISC) loss function. For simplicity, initially, a scenario involving a single atom is considered. In contrast to the physics-informed pseudo-label loss, which consistently involves a two-point relation between the original point at r and the extrapolated point at r + <5r, physics principles advocate that the potential value is inherently local and should remain unaffected by the specific path of extrapolation, as expressed in Equation A3. To satisfy this fundamental physics principle, a 3-point configuration is introduced, illustrated in Fig. 7a. This figure includes the following 3-points: (1 ) the data point (A), representing the original atomic position from the training data; (2) the estimation point (B) that is located at a neighboring position to the data point, where the potential energy and force-field are estimated by a ML potential model without utilizing Equation A3; (3) the extrapolation point (C), where the potential energy is estimated from A, and the physics-informed pseudo-label loss Equation A4 is imposed. Fig. 7a shows a schematic illustration describing the relation among the data point (A), the estimation point (B), and extrapolation point (C) in the context of physics-informed spatial- consistency loss. According to the locality of potential energy, the potential energy at the extrapolation point (C) should be consistent regardless of whether it is extrapolated from A or B, leading to the formulation of the novel physics-informed consistency loss:

[0109] Here, L is also a distance metric loss, and EA,\Ais a potential energy at A’ derived from the potential energy at point A using the relation Equation A3 as: where SPAA' is a spatial vector between the points A and A’. Note that the physics- informed consistency loss Equation A5 implicitly enables the provision of pseudolabels at B and C, in particular, without explicitly obtaining a potential value at point B because of its consistency nature. While the proposed explanation of the physics- informed consistency loss function assumes a single-atom scenario, extending it to a molecular case is straightforward by applying the configuration of the 3-point system in Fig. 7a to all atoms in the molecule. Note that the flexibility of the spatial-consistency loss allows us the opportunity to explore configurations beyond the one depicted in Fig. 7a, e.g., as shown in Fig. 7b and 7c.

[0110] In the following, experimental results are provided to show the effectiveness of the physics-informed consistent-pseudo-labelling (PICPL) method. In particular, the small-data regime is presented, that is, the training sample number is either 100 or 1000 in order to simulate the real-world application when the huge computational cost prohibits generating enough training samples Note that the choice of the training sample number is also common for active learning method for the interatomic potential.

[0111] The following representative models that are provided in the OpenCatalyst project code base (Chanussot* et al., 2021 ) were trained: SchNet (Schutt et al., 2017), PaiNN (Schu tt et al., 2021 ), SpinConv (Shuaibi et al., 2021 ), eSCN (Passaro & Zitnick, 2023), and Equiformer v2 (Liao et al., 2023). To evaluate the effect and dependency of the physics-informed pseudo-label approach in detail, the training was performed on various datasets (The results on further datasets are provided in Appendix D): Ani- 1x (Smith et al., 2020) as a comprehensive molecular dataset, TiO2 (Artrith & Urban, 2016) as a dataset for metals, and the revised MD17 (rMD17) dataset (Chmiela et al., 2017; 2018) as a dataset with smaller and single- molecule dataset. In terms of benchmarks, as a baseline, the proposed physics-informed pseudo-label approach was compared with the vanilla case and the augmentation via de-noising the perturbation in the atomic-coordinate (NoisyNode) case (Godwin et al., 2021 ).

[0112] The Ani-1x: Large Molecule Dataset (Smith et al., 2020) includes 63865 molecules whose size ranges from 4 to 64 and the ML model requires to learn quantum mechanical feature (potential energy and force-field) on various molecules from small number of samples for each molecule. The result trained is provided in the table of Fig. 9.

[0113] Fig. 9 shows a table of root mean square errors on the Ani-1x dataset (Smith et al., 2020). Energy (E, kcal / mol) and force (F, kcal / A7mol) errors of different models on trained either 100 or 1000 configurations with three different initial weight parameters to suppress statistical fluctuations.

[0114] The result shows that the proposed approach clearly improves the performance of the models in almost all the cases. In particular, the performance gain in the potential energy reaches 10% to even more than 50%. Interestingly, the performance gain can be observed not only in the potential energy but also in the force-field, though the proposed approach does not explicitly provide the pseudo-label information in the force-field. For the most cases, the augmentation method (NoisyNode) reduces the performance. This is a natural consequence because the augmentation prohibits the ML models to learn the correct reaction of the potential energy and force- field in terms of the deviation of the atomic-coordinate. The exceptional cases are observed in the case of SchNet. This can be attributed to the invariant nature of SchNet which neglects the angular information of the atomic coordinate but only takes into account the scalar distance between atoms, resulting in losing the ability to distinguish the direction of the spatial perturbation in the atomic coordinate.

[0115] In the following, the performance gain dependence on the training sample number is discussed. The training of ML potential models was performed with training sample number ranging as: [50, 102, 103, 104]. The results are plotted in Fig. 10. Fig. 10 shows plots of the performance gain by PICPS approach with respect to the training sample number: Ntrain= [50, 102, 103, 104]. In the top row, the ratio of the root mean square error (RMSE) of PICPS to vanilla cases in terms of the force-field is shown. In the bottom row, the ratio of RMSE of PICPS to vanilla cases in terms of the potential energy is shown.

[0116] Although the model dependence is strong, the performance gain shows a weak decrease with the training sample number in most cases. Note that this is an expected result because the extrapolated data regime by the pseudo-label is gradually covered by the real training sample as the training sample number increases. It is also found that the performance gain in the potential energy is larger than that in the force-field, which is only indirectly trained through the consistency constraint in PIRC (Equation A4). Finally, it is shown that the performance gain is larger when considering ML models with large capacity which gains more benefit from the increase of training data.

[0117] In the following, benchmark results are discussed for titanium dioxide (TiC ). TiO2 is an industrially relevant and well-studied material. A TiO2 dataset (Artrith & Urban, 2016) includes 7815 structures of several titanium dioxide phases whose reference energy and forces were obtained from DFT calculations. The number of atoms in the metal is typically 95 with the periodic boundary condition. ML models are required to learn quantum mechanical features (potential energy and force-field) on various phases, in addition to the periodic nature. The result of the training is provided in Fig. 11 . Fig. 11 shows a table of root mean square errors on the TiO2 dataset (Artrith & Urban, 2016). Fig. 11 shows energy (E, kcal / mol) and force (F, kcal / A7mol) errors of different models on trained either 100 or 1000 configurations with three different initial weight parameters to suppress statistical fluctuations. In the case of PaiNN, the prediction becomes a NaN value when added perturbation in atomic potential, which is represented by N / A in the list. As observed in the case of Ani-1x dataset, the proposed approach enables an improvement of the performance for both the potential energy and the force-field.

[0118] The rMD17 Small-Molecular Dynamics Trajectory dataset (Chmiela et al., 2018) includes ten relatively small-size molecules each of which has 105samples collected by performing MD simulations. ML models are required to learn quantum mechanical feature (potential energy and force-field) on one-molecule in a steady-state. In contrast to the Ani-1x dataset, it is aimed at analyzing the effect of the physics- informed consistency approach on smaller molecules with smaller variation than Ani- 1x dataset. For this purpose, Aspirin (natom = 21 ) and Benzene (natom = 12) were picked up as middle and small size molecules. The result of the training is provided in Fig. 12. Fig. 12 shows a table of root mean square errors on the Aspirin and Benzene in rMD17 dataset (Chmiela et al., 2018). Energy (E, kcal / mol) and force (F, kcal / A° / mol) errors of different models were determined using ML models trained either on 100 or 1000 configurations with three different initial weight parameters to suppress statistical fluctuations. The result indicates that the proposed approach becomes more effective with increasing molecule size and decreasing training sample number. In particular, the PIC performance in the case of Benzene has nearly no gain relative to the vanilla case. It is believed that this can be attributed by the too small variation in the Benzene dataset, which can be too easy for ML models to learn. Note that this can be supported by the small performance difference between the cases of 100 and 1000 training data samples.

[0119] In the following, several detailed analyses on the proposed approach are provided. For this purpose, the Equiformer v2 and PaiNN as the recent SOTA and light- weighted equivalent networks, which are trained on Ani-1x dataset with / V = 1000 training samples, were considered.

[0120] In the following, the effect of PIRC and PISC losses is considered via an ablation study. The result of the ablation study is provided in Fig. 13. The listed numerical values are the root mean square errors on the Ani-1x dataset (Artrith & Urban, 2016). Energy (E, kcal / mol) and force (F, kcal / A° / mol) errors of different models on trained 1000 configurations with three different initial weight parameters to suppress statistical fluctuations. Fig. 13 shows that the performance gain is mainly obtained by PIPS loss. In terms of the PISC loss, the result indicates that the training with only PISC loss does not provide performance gain while making the training difficult. However, it also shows that the PISC loss stabilized the training with PIRC loss also contributes a small amount of the performance gain.

[0121] In the following, the performance between a physics-informed pseudo-label and the proposed relational-consistent pseudo label is discussed. In the conventional case, the conventional pseudo-label loss is defined as: where Eiabei, Fiabei are the potential energy and force-field provided in the training dataset as labels, respectively. For simplicity, only the physics-informed pseudo- label type losses, Equation A4 and Equation A7 are considered. For a fair comparison, the following two cases are considered: the first case is the training with force-field label in which the ML potential model should learn not only the potential energy but also the force-field. And the second case is the training without force-field in which the ML potential model can concentrate on fitting the potential energy only For the case without force-field, the coefficient of the force-field loss was set as 0. For the case with the force-field, the numeric coefficient of the PIRC loss was set to 1 .0 and for the case without force-field, the numeric coefficient is set as 0.1 as a tuned value. The results are provided in Fig. 14. Fig. 14 shows a table of a comparison with the pseudolabel case. The listed numerical values are the root mean square errors on the Ani- 1x dataset (Artrith & Urban, 2016). Energy (E, kcal / mol) and force (F, kcal / A° / mol) errors of different models on trained 1000 configurations with three different initial weight parameters to suppress statistical fluctuations.

[0122] First, it was noticed, that the proposed PIRC loss shows the best performance in all the cases. Interestingly, PaiNN failed to learn the potential energy in the case of the conventional loss with force-field. It is believed that this is because of the in-balance of the training between the potential energy and the force-field because the conventional loss only trains the potential energy and also partly due to the not enough capacity in PaiNN to develop independent weight parameters for the forcefield prediction branch. This is supported by the observation of the case without forcefield training where the potential energy error is reduced in comparison to the vanilla case, though the PIRC case still shows the better performance. In conclusion, the proposed PIRC loss enables ML models to learn both the potential energy and forcefield with the help of consistently satisfying the physics law Equation A3.

[0123] In the following, the performance dependence on the choice of the spatial-deviation vector <5r in Equation A4 is discussed. In the present context, the spatial-deviation vector was defined as:

[0124] 8r = die (A9) where dl is the length and e is the unit vector of the spatial-deviation vector. In the previous experiments, the vector was determined using uniform random number with dl < dlo where dlo is the maximum deviation length. (Goodfellow et al., 2014; Miyato et al., 2018). An adversarial direction for the spatial-deviation vector can be defined as the most sensitive direction of the distance-measure between the ML prediction and label values in terms of the input x. If the distance-measure D by L2 norm function is considered, the adversarial direction can be approximated by: where yPred and yiabei are the ML model prediction and the label values. Fig. 15 shows the experimental result of the spatial-deviation vector selection dependence. The listed numerical values are the root mean square errors on the Ani-1x dataset (Artrith & Urban, 2016). Energy (E, kcal / mol) and force (F, kcal / A° / mol) errors of different models on trained 1000 configurations with three different initial weight parameters to suppress statistical fluctuations. ’’Vanilla” means that ML model without PIC, ’’random ” means that trained with PIC utilizing randomly determined direction and length, and ”adv” means along with the adversarial direction. The result shows that both choices improve the performance in comparison with the vanilla case, though the improvement depends on the model selection.

[0125] In the proposed methodology, two distinct physics-informed consistent-pseudo- labelling loss functions are introduced for molecular interatomic potential energy and force field. These losses operate explicitly (PIRC loss) and implicitly (PISC loss), which enables any ML potential models to im- prove their accuracy, particularly in scenarios characterized by limited data, the proposed extensive experiments demonstrated no- table efficacy and efficiency of the proposed method from various aspects: dependence on the training dataset size, dependence on the variability in molecular structures within the dataset, and selection of the deviation vector. Future development should explore an application of the developed model to molecular dynamic simulations for real-world problems and evaluate the robustness of the predicted force-field.

[0126] The proposed PICPS method scope is restricted to ML models for predicting molecular properties. It might not be applicable to other general graph tasks that neither physics nor chemistry is related. Many modifications and other embodiments of the invention set forth herein will come to mind to the one skilled in the art to which the invention pertains having the benefit of the teachings presented in the foregoing description and the associated drawings. Therefore, it is to be understood that the invention is not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0127] L i s t o f r e f e r e n c e s i g n s:

[0128] System Interface Processor Storage device System Interface Processor Storage device Obtaining an initial training sample Generating spatially perturbed atomic coordinate information

[0129] Generating second spatially perturbed atomic coordinate information Predicting one or more molecular properties of a second spatially perturbed variation of the first molecule

[0130] Predicting a second version of at least one molecular property Deterministically calculating a potential energy Deterministically estimating a force-field Storing an additional training sample Calculating a first loss function Calculating a second loss function Calculating a third loss function Adjusting a machine-learning model Obtaining a machine-learning model Predicting at least one molecular property Using the predicted at least one molecular property

Claims

C l a i m s1. A computer-implemented method for training a machine-learning model for prediction of one or more molecular properties of a molecule, the method comprising: obtaining (110) an initial training sample for training the machine-learning model, the initial training sample representing a first molecule and comprising atomic coordinate information, chemical element information, and information on one or more molecular properties of the first molecule; generating (120) spatially perturbed atomic coordinate information for a spatially perturbed variation of the first molecule; deterministically (140, 145) calculating or estimating at least one molecular property of the spatially perturbed variation of the first molecule based on the initial training sample and based on a spatial deviation between the atomic coordinate information of the first molecule and the spatially perturbed atomic coordinate information; calculating (160) a first partial loss function between the information on the one or more molecular properties of the first molecule included in the initial training sample and predicted information on the one or more molecular properties of the first molecule predicted by the machine-learning model; calculating (170) a second partial loss function between the deterministically calculated or estimated at least one molecular property of the spatially perturbed variation of the first molecule and a corresponding prediction of the at least one molecular property predicted by the machine-learning model; and adjusting (190) the machine-learning model based on a result of the first and second partial loss function.

2. The computer-implemented method according to claim 1 , wherein the machine-learning model is trained to output, based on the atomic coordinate information and based on the chemical element information, the one or more molecular properties, with the one or more molecular properties comprising the at least one molecular property corresponding to the deterministically calculated or estimated at least one molecular property.

3. The computer-implemented method according to one of the claims 1 or 2, wherein the one or more molecular properties comprise at least one of a potential energy of the molecule and a force-field of the molecule.

4. The computer-implemented method according to one of the claims 1 to 3, wherein the at least one molecular property being deterministically calculated (140, 135) or estimated for the spatially perturbed atomic coordinate information comprises the potential energy of the spatially perturbed variation of the first molecule.

5. The computer-implemented method according to claim 4, wherein the potential energy is deterministically calculated (140) using a Taylor expansion, based on the potential energy of the first molecule, based on a force-field of the first molecule, and based on a spatial deviation between the atomic coordinates of the first and spatially perturbed variation of the first molecules.

6. The computer-implemented method according to one of the claims 1 to 5, wherein the at least one molecular property being deterministically estimated (145) for the spatially perturbed atomic coordinate information comprises the force-field of the spatially perturbed variation of the first molecule.

7. The computer-implemented method according to one of the claims 1 to 6, wherein the force-field of the spatially perturbed variation of the first molecule is deterministically estimated (145) for the spatially perturbed atomic coordinate information based on a gradient of the potential energy of the spatially perturbed variation of the first molecule, with the potential energy being deterministically calculated for the spatially perturbed atomic coordinate information.

8. The computer-implemented method according to one of the claims 1 to 7, wherein the at least one molecular property of the spatially perturbed variation of the first molecule is deterministically calculated (140) or estimated (145) based on the information on the one or more molecular properties included inthe initial training sample, and / or based on predicted information on at least one of the one more molecular properties predicted by the machine-learning model.

9. The computer-implemented method according to one of the claims 1 to 8, wherein generating (120) the spatially perturbed atomic coordinate information comprises generating a spatial deviation vector representing the spatial deviation between the atomic coordinate information of the first molecule and the spatially perturbed atomic coordinate information.

10. The computer-implemented method according to one of the claims 1 to 9, wherein the method comprises generating (130) second spatially perturbed atomic coordinate information for a second spatially perturbed variation of the first molecule, predicting (132) information on the one or more molecular properties of the second spatially perturbed variation of the first molecule using the machine-learning model based on the second spatially perturbed atomic coordinate information, deterministically calculating (140) or estimating (145) a first version of at least one molecular property of a target spatially perturbed variation of the first molecule, such as the spatially perturbed variation of the first molecule, based on the predicted information on the one or more molecular properties of the second spatially perturbed variation of the first molecule, predicting (134) a second version of the at least one molecular property of the target spatially perturbed variation of the first molecule based on the initial training sample or based on predicted information on the one or more molecular properties of a third spatially perturbed variation of the first molecule, calculating (180) a third partial loss function based on a difference between the first and second version of the at least one molecular property of the target spatially perturbed variation of the first molecule, and adjusting (190) the machine-learning model based on a result of the first, second and third partial loss function.

11. The computer-implemented method according to one of the claims 1 to 10, wherein to generate the spatially perturbed atomic coordinate information and / or second spatially perturbed atomic coordinate information, the atomic coordinate information of the first molecule is varied in a direction that leads tolarger changes in the loss function compared to other directions, or in a direction in which an impact of the variation on the loss function is more uncertain compared to other directions.

12. The computer-implemented method according to one of the claims 1 to 12, wherein the method comprising storing (150) an additional training sample representing the spatially perturbed variation of the first molecule, and / or wherein the spatially perturbed atomic coordinate information is generated by an active learning algorithm, and / or wherein the computer-implemented method is repeated for a plurality of initial training samples of a set of training data for training the machine-learning model.

13. A computer-implemented method for applying a machine-learning model trained for prediction of one or more molecular properties of a molecule, the method comprising: obtaining (410) a machine-learning model trained according to the method of one of the claims 1 to 12; and predicting (420) at least one molecular property of at least one molecule using the machine-learning model.

14. A system (10, 40) comprising one or more processors (13, 44) and one or more storage devices (16, 46), wherein the system (10, 40) is configured to perform at least one of the method of one of the claims 1 to 12 and the method of claim 13.

15. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out at least one of the method of one of the claims 1 to 12 and the method of claim 13.