Machine learning based low-scale coarsening modeling method and related apparatus
By employing a machine learning-based low-scaling coarse-grained modeling method, the problems of high cost in full-atom simulation and difficulty in coarse-grained simulation modeling are solved, enabling efficient and accurate simulation of complex multi-component systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGCHUN INSTITUTE OF APPLIED CHEMISTRY CHINESE ACADEMY OF SCIENCES
- Filing Date
- 2026-03-13
- Publication Date
- 2026-05-26
AI Technical Summary
Existing all-atom simulation methods are computationally expensive and limited in scale. Coarse-grained simulation methods are difficult to model in multi-component systems and cannot achieve large-scale, long-term simulation of complex multi-component systems.
We employ a machine learning-based low-scale coarse-grained modeling approach, constructing a coarse-grained model using all-atomic data. We then utilize machine learning to automatically optimize the potential function parameters, reducing the degrees of freedom while preserving key interaction mechanisms.
It improves the simulation efficiency and accuracy of complex multi-component systems, enhances the stability and cross-system transferability of the model, and enables efficient simulations on a large scale and over long periods of time.
Smart Images

Figure CN122090976A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of molecular dynamics simulation technology, and in particular to a method, apparatus, device, and computer-readable storage medium for low-scale coarse-grained modeling based on machine learning. Background Technology
[0002] Currently, molecular dynamics simulations are mainly used to simulate and study molecular systems such as ions, small molecules, and polymers. Commonly used molecular dynamics simulation methods include all-atom simulations and coarse-grained simulations.
[0003] All-atom simulations, by explicitly describing each atom in a system and its interactions, can accurately characterize intermolecular mechanisms, making them particularly suitable for studying local structures and long-range (e.g., electrostatics) and short-range (e.g., van der Waals) interactions in ionic, small molecule, and polymeric systems. However, the number of degrees of freedom required for all-atom simulations increases rapidly with system size, and their computational cost grows non-linearly with the number of particles and simulation time, significantly limiting their application in large-scale or long-term simulations.
[0004] To reduce computational complexity and expand the simulable time and spatial scales, researchers have proposed a coarse-grained simulation method. This method reduces the system's degrees of freedom by merging multiple atoms into one or more coarse-grained particles and describes the interactions between these particles using an effective interaction potential function. Through coarse-graining, the statistical properties of the system can be preserved to some extent while significantly improving computational efficiency.
[0005] Existing coarse-grained simulation methods typically approximate these complex interactions by fitting empirical parameters or specific all-atom data. In simple or single-component systems, these models can achieve reasonable results under certain conditions. However, when the system contains multiple components such as ions, small molecules, and polymers, the coupling between interactions of different scales and types makes the construction of coarse-grained models and the selection of parameters much more complex.
[0006] Therefore, with the continuous development of molecular dynamics simulation technology towards larger scales and more complex systems, how to efficiently construct coarse-grained models suitable for multi-component systems while ensuring physical rationality has become one of the urgent technical problems to be solved in this field. Summary of the Invention
[0007] The purpose of this application is to provide a method, apparatus, device, and computer-readable storage medium for low-scale coarse-grained modeling based on machine learning, thereby solving the technical problems of high computational cost, scale limitation, and difficulty in modeling coarse-grained simulation methods in existing all-atom simulation methods.
[0008] To achieve the above objectives, this application provides a low-scale coarse-grained modeling method based on machine learning, comprising:
[0009] Step 1: Obtain all atomic data of the target system;
[0010] Step 2: Construct the coarse-grained model to be learned; the coarse-grained model to be learned adopts a potential function form including a potential energy expression for describing polar correlation interaction, and the coarse-grained model to be learned is a model that has been verified by the full-atom data;
[0011] Step 3: Using the all-atom data as the optimization target, the potential function parameters to be determined in the coarse-grained model to be learned are automatically iteratively optimized using machine learning methods until the preset conditions are met, and the optimized coarse-grained model is output.
[0012] Optionally, step 1 includes:
[0013] Step 11: Construct a molecular system model of the target system;
[0014] Step 12: Specify an all-atom force field model for the molecular system model;
[0015] Step 13: Based on the molecular system model and the all-atom force field model, perform all-atom molecular dynamics simulation calculations;
[0016] Step 14: Collect the dynamic data related to interactions generated during the all-atom molecular dynamics simulation to obtain the simulated trajectory data;
[0017] Step 15: Analyze the simulation results of the simulated trajectory data and extract reference physical quantities as the all-atom data.
[0018] Optionally, step 2 includes:
[0019] Step 21: Perform coarse-grained modeling on the target system, map multiple sets of atoms in the target system to one or more coarse-grained particles, and determine the spatial coordinates, type identifiers, and topological connections of the coarse-grained particles;
[0020] Step 22: Determine the feature quantities based on the spatial coordinates, the type identifier, and the topological connection relationship;
[0021] Step 23: Define the potential function form for the interaction between the coarse-grained particles to obtain the initial coarse-grained model;
[0022] Step 24: Validate the initial coarse-grained model using the full-atom data, and determine the validated initial coarse-grained model as the coarse-grained model to be learned.
[0023] Optionally, step 23 includes:
[0024] Based on Stockmayer fluid theory, dipole moments are set for the coarse-grained particles corresponding to the polar components in the target system, and the potential energy expressions for the first polar correlation interaction and the second polar correlation interaction are defined according to the dipole moments to obtain the initial coarse-grained model.
[0025] The first polar correlation interaction represents the interaction between the coarse-grained particle corresponding to the ion and the coarse-grained particle having the dipole moment, and the second polar correlation interaction represents the interaction between two coarse-grained particles having the dipole moment.
[0026] Optionally, step 24 includes:
[0027] Based on the initial coarse-grained model, perform coarse-grained molecular dynamics simulation calculations and verify whether the difference between the coarse-grained molecular dynamics simulation results and the all-atom data meets the preset range;
[0028] When the preset range is met, the initial coarse-grained model is determined as the coarse-grained model to be learned.
[0029] Optionally, step 3 includes:
[0030] Using the all-atomic data as the optimization target, machine learning is performed based on the Bayesian optimization method. The potential function parameters to be determined in the coarse-grained model to be learned are automatically iteratively optimized until the preset conditions are met, and the optimized coarse-grained model is output.
[0031] Optionally, step 3 includes:
[0032] Step 31: Construct the objective function; the objective function is a function used to measure the performance of the coarse-grained model;
[0033] Step 32: Based on the current potential function parameters and their corresponding parameter evaluation results, establish a probabilistic proxy model for the objective function;
[0034] Step 33: Define the acquisition function based on the probabilistic proxy model;
[0035] Step 34: Automatically select the next set of candidate potential function parameters from the preset parameter space using the acquisition function;
[0036] Step 35: Substitute the candidate potential function parameters into the coarse-grained model to be learned, perform coarse-grained molecular dynamics simulation calculations, and determine the objective function value based on the coarse-grained molecular dynamics simulation results and the objective function to obtain the parameter evaluation results;
[0037] Step 36: Feed back the candidate potential function parameters and their corresponding parameter evaluation results to step 32, and repeat steps 32 to 35 until the objective function value converges or the preset termination condition is met, and output the optimized coarse-grained model.
[0038] To achieve the above objectives, this application also provides a machine learning-based low-scale coarse-grained modeling apparatus, comprising:
[0039] The all-atom data acquisition module is used to perform step 1: acquiring all-atom data of the target system;
[0040] The initial machine learning modeling module is used to perform step 2: constructing a coarse-grained model to be learned; the coarse-grained model to be learned adopts a potential function form including a potential energy expression for describing polar correlation interaction, and the coarse-grained model to be learned is a model that has been validated by the full-atom data;
[0041] The machine learning module is used to perform step 3: taking the all-atomic data as the optimization target, automatically iteratively optimizing the potential function parameters to be determined in the coarse-grained model to be learned using machine learning methods until the preset conditions are met, and outputting the optimized coarse-grained model.
[0042] To achieve the above objectives, this application also provides a machine learning-based low-scale coarse-grained modeling device, comprising:
[0043] Memory, used to store computer programs;
[0044] A processor is used to implement the steps of the machine learning-based low-scale coarse-grained model modeling method described above when executing the computer program.
[0045] To achieve the above objectives, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the machine learning-based low-scale coarse-grained model modeling method described above.
[0046] Existing all-atom simulation methods require explicit descriptions of interactions between all atoms in a system, resulting in computational complexity that increases rapidly with system size and simulation time. Therefore, they struggle to achieve large-scale and long-term simulations of complex multi-component systems containing ions, small molecules, and polymers. While traditional coarse-grained simulation methods reduce computational freedom, they typically rely on empirical parameters or simplified electrostatic force treatments, leading to insufficient descriptions of charge correlations, dipole effects, and multi-component interactions, thus limiting model accuracy and transferability. Therefore, relying solely on existing all-atom or coarse-grained simulation methods is insufficient to simultaneously meet the technical requirements of computational efficiency and physical accuracy.
[0047] To address the aforementioned issues, this application provides a machine learning-based low-scale coarse-grained modeling method. Based on all-atom simulation data, it introduces machine learning methods to automatically learn and construct a coarse-grained model incorporating polar correlation interactions, achieving unified modeling of ionic, small molecule, and polymeric systems. This method significantly reduces the degrees of freedom of the coarse-grained model while preserving key electrostatic and dipole physics mechanisms, and reduces human intervention through a highly automated process. This not only improves the computational efficiency of simulating complex systems but also enhances the model's accuracy, stability, and cross-system transferability, providing a solution with significant technical advantages for the efficient simulation of complex multi-component systems. This application also provides a machine learning-based low-scale coarse-grained modeling apparatus, device, and computer-readable storage medium, which possess the aforementioned beneficial effects. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0049] Figure 1 A flowchart illustrating a low-scale coarse-grained modeling method based on machine learning, provided for embodiments of this application;
[0050] Figure 2 A flowchart illustrating a low-scale coarse-grained modeling method based on machine learning, provided in an embodiment of this application;
[0051] Figure 3 This is a structural block diagram of a low-scale coarse-grained modeling device based on machine learning, provided in an embodiment of this application. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] Molecular dynamics (MD) simulation is a method of studying the evolution of atomic or molecular systems over time by solving Newton's equations of motion using numerical computer methods. Researchers can successfully obtain microscopic information about the structure, energy, transport, thermodynamics, and dynamics of matter or materials using this method. With the development of computing resources, molecular dynamics simulation has become an important research tool in physics, chemistry, materials science, life sciences, energy science, bionics, and other fields, and is widely used to study the microscopic properties of solution systems, melt systems, solid systems, interfaces and confined processes, as well as various related complex systems.
[0054] For simulation studies of molecular systems such as ions, small molecules, and polymers, two main types of molecular dynamics simulation methods are used: all-atom simulation and coarse-grained simulation.
[0055] All-atom simulation methods, based on classical force fields or first-principles calculations, explicitly model each atom in the system and solve the particle motion equations through numerical integration to obtain structural and dynamic information of the system during its time evolution. In this type of method, the interactions between atoms are typically described by bonding, non-bonding, and electrostatic terms, with charge distribution given by pre-defined fixed local charges of atoms. All-atom simulation methods can accurately characterize ionic solvation structures, conformational changes in small molecules, and local conformational features of polymers.
[0056] Coarse-grained simulation methods reduce the degrees of freedom of a system by mapping several atoms to a coarse-grained particle and use an effective potential function to describe the interactions between these particles. The construction of a coarse-grained model in this method typically involves determining the particle mapping scheme, selecting target properties, and fitting the potential function parameters. In ionic systems, coarse-grained models usually introduce a simplified electrostatic term into the effective potential, treat the medium outside the ions as a dielectric continuum, and only consider the Coulomb interactions between ionic components. For polar molecules, the local charge polarity can be approximated by setting an equivalent dipole. These models can reproduce experimental or all-atom simulation results well under specific system and parameter conditions and significantly reduce computational costs.
[0057] Existing simulation methods have the following drawbacks:
[0058] First, because constructing a full-atom model requires explicitly handling all atomic degrees of freedom, its computational cost increases significantly with the system size and simulation time. Therefore, in practical applications, it is usually only suitable for small-scale or short-term simulations. This technical characteristic directly makes it difficult for full-atom simulation methods to undertake large-scale, long-term studies of complex multi-component systems, and their main role is often limited to providing reference data.
[0059] Furthermore, since coarse-grained models rely on full-atom simulation results or empirical rules for parameter construction, their effective particle size and potential function are essentially approximations of interactions under specific systems and conditions. When the system composition, temperature, concentration, or charge environment changes, the original parameters often fail to maintain consistent applicability, leading to a decrease in model prediction accuracy. This parameter dependence limits the transferability of existing coarse-grained models across different systems. Simultaneously, to reduce computational complexity, existing techniques often employ simplified charge models or fixed-form effective potential functions to approximate charge correlations and dipole effects. This approach makes it difficult for such models to effectively match real-world systems in multi-component systems.
[0060] Furthermore, the construction process of existing coarse-grained models typically involves manually setting mapping schemes, manually selecting target properties, and repeatedly adjusting parameters, making it highly dependent on researchers' experience. When the research object expands from a single-component system to a multi-component system containing ions, small molecules, and macromolecules, the difficulty of model construction and parameter optimization increases significantly, leading to reduced modeling efficiency and difficulty in ensuring repeatability and consistency.
[0061] Therefore, this application provides a coarse-grained modeling method, which is an automated coarse-grained modeling method based on machine learning. While reducing the simulation degrees of freedom and computational costs, it can systematically retain the key interaction mechanisms between components. By performing machine learning on key information such as local charge, atomic configuration and atomic forces in the full-atom model, the coarse-grained model is systematically optimized, realizing accurate, cross-scale and automated mapping from the full-atom model to the coarse-grained model, providing an improved technical solution for the efficient simulation of complex systems.
[0062] Please refer to Figure 1 , Figure 1 A flowchart illustrating a machine learning-based low-scale coarse-grained model modeling method provided in this application embodiment, the method may include:
[0063] Step 1: Obtain the full atomic data of the target system.
[0064] This embodiment allows selection of the specific type of target system based on the object under study. The target system can be a single-component system or a multi-component mixture system. In one possible implementation, the target system can be a molecular system comprising ions, small molecules, and / or polymeric components.
[0065] It should be noted that the modeling method proposed in this embodiment is applicable to ionic, small molecule and polymer systems under the same theoretical and computational framework, and can handle complex multi-component systems, with good versatility and scalability.
[0066] This embodiment does not limit the specific method of obtaining full-atom data, as long as it ensures that full-atom data can be obtained from the target system. In one possible implementation, step 1 may include:
[0067] Step 11: Construct a molecular system model of the target system.
[0068] This embodiment does not limit the specific method of constructing the molecular system model, as long as it can describe each atom of the target system. In one possible implementation, full-atom modeling can be performed on each component of the target system, and the initial configuration, composition ratio, and simulation conditions of each component can be set.
[0069] This embodiment does not limit the specific way of setting the initial configuration, composition ratio and simulation conditions, and can be set according to research needs.
[0070] In one possible implementation, the initial configuration of the simulation system can be generated by an automated program; this program can be written in Perl; alternatively, Python, Fortran, C / C++, Java, or any other computer language can be used to generate the simulation input data. This initial configuration generation method allows for flexible adjustment of the system component concentrations and initial distribution algorithms according to different research needs.
[0071] In one possible implementation, simulation conditions can include settings for parameters such as system temperature, pressure, simulation step size, boundary conditions, and ensemble type. Specifically, by controlling temperature and pressure, the system can be maintained under isothermal, isobaric, or other thermodynamic constraints. Depending on research needs and system characteristics, various common temperature control methods can be selected, including but not limited to Langevin baths, Nose-Hoover baths, Hoover chain baths, DPD baths, Berendsen baths, and Andersen baths—any commonly used molecular dynamics simulation bath. The simulation step size determines the temporal accuracy of particle position and velocity updates, ensuring the stability and accuracy of numerical integration. Boundary conditions can include periodic boundaries, fixed boundaries, free boundaries, or semi-open boundaries to adapt to different research systems. The ensemble type specifies the statistical averaging method for the system under specific thermodynamic constraints, ensuring that the simulation results are consistent with the expected experimental conditions. The simulation dimension can be adjusted to two or three dimensions as needed. In this way, this scheme supports arbitrary molecular dynamics simulation condition settings.
[0072] Step 12: Specify an all-atom force field model for the molecular system model.
[0073] This embodiment does not limit the specific parameters of the all-atomic force field model, as long as the all-atomic force field can be defined. In one possible implementation, the all-atomic force field model may include bonding interaction parameters, non-bonding interaction parameters, and parameters such as charge distribution and atomic forces.
[0074] This embodiment does not limit the specific type of all-atomic force field used in the all-atomic force field model; the all-atomic force field can be selected arbitrarily according to the system type. In one possible implementation, the force field can be an existing known force field used to describe the interaction behavior between atoms, or a verified combination of parameters. In one possible implementation, the all-atomic force field in this embodiment can be entirely based on the OPLS-AA force field, i.e., the atomic force field model adopts the OPLS-AA force field model; in addition, any all-atomic molecular dynamics model such as the AMBER force field model or the CHARMM force field model can also be used.
[0075] Step 13: Perform all-atom molecular dynamics simulation calculations based on the molecular system model and the all-atom force field model.
[0076] It should be noted that all-atom molecular dynamics simulation refers to a numerical simulation method that uses classical Newtonian mechanics equations to calculate the evolution of particles in a molecular system over time, and is used to study the microstructure and dynamic properties of the system.
[0077] This embodiment does not limit the specific software used for all-atom molecular dynamics simulations, as long as it can perform molecular dynamics simulation operations. In one possible implementation, all-atom molecular dynamics simulations can be based on the LAMMPS (Large-scale Atomic / Molecular Massively Parallel Simulator) software platform. LAMMPS is an open-source molecular dynamics simulation software used for simulation calculations of atoms, molecules, and coarse-grained systems. In addition, the automated simulation and analysis workflow proposed in this embodiment can be adapted to or extended to other molecular dynamics simulation platforms, including but not limited to GROMACS, AMBER, NAMD, HOOMD-blue, OpenMM, Desmond, ESPResSo, Galamost, and molecular dynamics programs compiled or developed by researchers themselves.
[0078] It should be noted that, through compatibility and expansion of the aforementioned software, this embodiment can cover a wide range of simulation needs, from the atomic scale to the coarse-grained scale, and from small systems to large systems with millions of systems, demonstrating broad applicability and flexible scalability.
[0079] This embodiment does not limit the specific method of performing all-atom molecular dynamics simulations, as long as it ensures that the time-series simulation of the target system can be achieved. In one possible implementation, step 13 may include: based on the molecular system model and the all-atom force field model, using numerical integration methods to obtain the simulated trajectory of the target system evolving on a continuous time scale under given temperature, pressure or other thermodynamic conditions.
[0080] It should be noted that in molecular dynamics, numerical integration methods are used to numerically solve Newton's equations of motion. Calculating the evolution of a particle's velocity and position over time based on the forces acting on it is a core step in simulating system dynamics. In one possible implementation, the numerical integration method can employ any molecular dynamics integration algorithm, such as the Verlet algorithm, the Velocity-Verlet algorithm (which balances accuracy and stability and is the most commonly used), or the Leap-frog algorithm (suitable for long-term stable simulations). These algorithms can efficiently advance the evolution of the system over time while ensuring energy conservation and numerical stability.
[0081] In one possible implementation, the time step can be adjusted appropriately based on the smoothness of the force field.
[0082] In one possible implementation, this embodiment can support multi-core CPUs, GPU acceleration, and distributed computing clusters to automatically adjust the number of parallel tasks and load balancing strategies.
[0083] In one possible implementation, this embodiment may also integrate an external field (electric field, magnetic field, shear flow) simulation module to achieve multi-physics field coupling simulation.
[0084] Step 14: Collect the dynamic data related to interactions generated during the all-atom molecular dynamics simulation to obtain the simulated trajectory data.
[0085] This implementation does not limit the specific method of obtaining simulated trajectory data, as long as it ensures that the simulated trajectory data can be obtained from the simulation results. In one possible implementation, the simulated trajectory can be sampled during or after the all-atom molecular dynamics simulation to collect dynamic data related to interactions.
[0086] It should be noted that the dynamic data related to interactions collected in this embodiment can be used as the raw dataset for subsequent modeling. The embodiment does not limit the types of dynamic data related to interactions, which can be determined according to actual circumstances. In one possible implementation, the dynamic data related to interactions may include atomic coordinates, velocities, charge distribution, and other dynamic data related to interactions.
[0087] Step 15: Analyze the simulation results of the simulated trajectory data and extract the reference physical quantities as all-atom data.
[0088] It should be noted that the simulation results analysis in this embodiment is used for statistical and structural characterization of the particle trajectories and related physical quantities obtained from molecular dynamics simulations. Possible analytical methods include, but are not limited to:
[0089] Radial distribution function (RDF): used to describe the probability distribution of particles relative to a reference particle at different distances, characterizing the local structure of the system;
[0090] Structure factor (S(q)): Used to describe the statistical properties of particle distribution in reciprocal space of the system, and can simulate scattering experimental data;
[0091] Mean square displacement (MSD): used to calculate particle diffusion coefficient and analyze kinetic properties;
[0092] Bond length, bond angle, and dihedral angle distribution: used to analyze changes in the internal structure of molecules;
[0093] Cluster analysis is used to identify aggregates of particles or molecules.
[0094] Energy analysis includes changes in potential energy, kinetic energy, and total energy over time, used to determine the equilibrium state and stability of the system; other analyses include system property parameters such as diffusion coefficient and viscosity calculated using MSD or velocity autocorrelation function.
[0095] Intermediate scattering function, diffusion coefficient, and other arbitrary conventional dynamic and structural information.
[0096] In one possible implementation, the simulation results analysis can be performed using a separate analysis program. This embodiment does not limit the specific analysis program used, as long as it can perform the analysis. In one possible implementation, the analysis program can be written in Fortran; alternatively, it can be written in Python, Perl, C / C++, Java, or any other computer language for statistical analysis of the structure, thermodynamics, and kinetic properties of the system, and can be extended according to the research objectives.
[0097] It should be noted that the analysis process in this embodiment can be flexibly combined and customized, and the output results can be directly used for statistical analysis and machine learning.
[0098] It should be noted that the reference physical quantities in this embodiment can be used as training targets or constraints for machine learning modeling. This embodiment does not limit the specific types of reference physical quantities. In one possible implementation, the physical quantities may include, but are not limited to, radial distribution function, structure factor, mean square displacement, mean square radius of gyration, and dipole orientation; in addition, they may include, but are not limited to, various structural, mechanical, thermodynamic, and dynamic properties such as system energy, interparticle force information, structural statistics, coordination number, and dielectric correlation quantities.
[0099] Step 2: Construct the coarse-grained model to be learned; the potential function form used in the coarse-grained model to be learned includes the potential energy expression used to describe the polar correlation interaction, and the coarse-grained model to be learned is a model that has been validated by full-atom data.
[0100] It should be noted that the coarse-grained model to be learned constructed in this embodiment includes potential function parameters to be determined, which need to be determined by machine learning methods in step 3.
[0101] This embodiment does not limit the specific method of constructing the coarse-grained model to be learned, as long as it includes a potential energy expression to describe the polar correlation interaction and is validated by all-atom data. In one possible implementation, step 2 may include:
[0102] Step 21: Perform coarse-grained modeling of the target system, mapping multiple sets of atoms in the target system to one or more coarse-grained particles, and determining the spatial coordinates, type identifiers, and topological connections of the coarse-grained particles.
[0103] It should be noted that, in this embodiment, step 21 can reduce the number of degrees of freedom while maintaining the main structural features of the system.
[0104] In one possible implementation, multiple sets of atoms in the target system can be mapped to one or more coarse-grained particles based on a preset mapping rule.
[0105] Step 22: Determine the feature quantities based on spatial coordinates, type identifiers, and topological connectivity.
[0106] It should be noted that feature quantities refer to numerical quantities extracted from all-atom or coarse-grained simulation data to characterize the system architecture or interaction state, serving as input variables for machine learning models. In one possible implementation, feature quantities may include particle size, charge, dipole moment, bond length, and LJ potential energy parameters; in addition, they may include, but are not limited to, any parameters reflecting local or global information and interaction of the system, such as inter-particle distance, angular relationship, local density, charge or dipole distribution characteristics, particle forces, and potential energy surface, to reflect key physical information of the system and support machine learning models in learning and predicting potential function parameters.
[0107] Step 23: Define the potential function form for the interaction between coarse-grained particles to obtain the initial coarse-grained model.
[0108] This embodiment does not limit the specific form of the coarsening potential function. In one possible implementation, step 23 may include:
[0109] Based on Stockmayer fluid theory, dipole moments are set for coarse-grained particles corresponding to polar components in the target system, and the potential energy expressions for the first polar correlation interaction and the second polar correlation interaction are defined according to the dipole moments to obtain the initial coarse-grained model.
[0110] The first polar correlation interaction represents the interaction between the coarse-grained particle corresponding to the ion and the coarse-grained particle with a dipole moment (this interaction is an electrostatic interaction, which can be simply referred to as the ion-dipole interaction). The second polar correlation interaction represents the interaction between two coarse-grained particles with a dipole moment (this interaction is also an electrostatic interaction, which can be simply referred to as the dipole-dipole interaction).
[0111] It should be noted that Stockmayer fluid theory is a molecular model that introduces point dipole-dipole interactions on top of the classical Lennard-Jones interaction, used to describe anisotropic interactions between particles with permanent dipole moments. This theory can simultaneously characterize short-range repulsion and dispersion as well as long-range dipole coupling effects, and is often used to simulate polar molecular systems. This embodiment introduces a Stockmayer-type dipole interaction term to describe permanent dipole-dipole interactions, which can retain an effective description of charge correlation, dipole effects, and multi-component systems.
[0112] It should be noted that, in addition to ion-dipole electrostatic interactions and dipole-dipole electrostatic interactions, the interactions between the components in this embodiment may also include particle repulsion and short-range interactions, electrostatic interactions between charged components, and bond connections between polymer chains. Among these, ion-dipole electrostatic interactions, dipole-dipole electrostatic interactions, particle repulsion and short-range interactions, and electrostatic interactions between charged components are all non-bonded interactions; the bond connections between polymer chains are bonded interactions.
[0113] It should be noted that molecular dynamics simulations describe the interactions between particles using potential functions. In one possible implementation, the potential function can also include the Soft potential, Coulomb potential, Lennard-Jones potential, and FENE potential. The Soft potential is a flexible repulsive potential, often used for initial model optimization to gradually eliminate particle overlap and prevent excessive system energy. The Coulomb potential characterizes long-range electrostatic interactions between charged particles. The Lennard-Jones potential comprehensively describes short-range repulsion and medium-range van der Waals attraction between particles and is the most widely used non-bonded potential. The FENE (Finitely Extensible Nonlinear Elastic) potential is commonly used in polymer systems to simulate nonlinear elastic bond connections between chain segments. Furthermore, bond stretching potentials, bond angle potentials, and dihedral potentials can be introduced in molecular dynamics as needed.
[0114] Step 24: Validate the initial coarse-grained model using all-atom data, and determine the validated initial coarse-grained model as the coarse-grained model to be learned.
[0115] This embodiment does not limit the specific method of verification, and can be determined according to the actual situation. In one possible implementation, step 24 may include:
[0116] Coarse-grained molecular dynamics simulations were performed based on the initial coarse-grained model, and the differences between the coarse-grained molecular dynamics simulation results and the all-atom data were verified to meet the preset range.
[0117] When a preset range is detected, the initial coarse-grained model is determined as the coarse-grained model to be learned.
[0118] It should be noted that step 24 in this embodiment can be understood as the process of initially training the coarse-grained potential function parameters based on the feature quantity and potential function form, using the reference physical quantity extracted from the full-atom simulation, before performing machine learning. The purpose of performing this step is to ensure that the constructed coarse-grained model to be learned can statistically reproduce the structural or mechanical distribution characteristics of the full-atom system.
[0119] It should be noted that this invention introduces multiple types of reference physical quantities, including structural, thermodynamic, and dynamical, as constraints during the construction and parameter optimization of the coarse-grained model. By using multiple physical quantities as joint constraints and multiple feature quantities as machine learning inputs, the coarse-grained model can effectively approximate the all-atom model in terms of structure preservation, charge correlation, and dynamic behavior, thereby improving the overall accuracy and physical consistency of the model and the force field.
[0120] Step 3: Using all-atom data as the optimization target, the potential function parameters to be determined in the coarse-grained model to be learned are automatically iteratively optimized using machine learning methods until the preset conditions are met, and the optimized coarse-grained model is output.
[0121] It should be noted that machine learning methods are a class of methods that automatically build models based on data. Their core idea is to learn the mapping relationship between input and output from existing sample data through algorithms, rather than relying on manually set empirical formulas. This embodiment combines all-atom molecular simulation with machine learning methods to construct a technical route that automatically generates coarse-grained models starting from all-atom data. This can significantly reduce the involvement of human experience and achieve a high degree of automation in the model building process.
[0122] In one possible implementation, machine learning methods can employ supervised learning for modeling, establishing a mapping relationship between potential function parameters and target physical quantities by using known full-atom simulation reference data and the corresponding coarse-grained model output. Alternatively, semi-supervised or unsupervised learning methods can be used, or a combination of these methods can be employed to adapt to different data scales and modeling needs.
[0123] In one possible implementation, the machine learning method may employ a regression-based model algorithm to fit the relationship between coarse-grained potential function parameters and a reference physical quantity; in other possible implementations, the machine learning algorithm may also include, but is not limited to, neural network models, Gaussian process models, kernel regression models, decision tree models, or combinations thereof, to predict or update the potential function parameters.
[0124] This embodiment does not limit the specific process of the machine learning method, as long as it can automatically generate the potential function parameters of the coarse-grained model. In one possible implementation, step 3 may include:
[0125] Using all-atomic data as the optimization target, machine learning is performed based on the Bayesian optimization method. The potential function parameters to be determined in the coarse-grained model to be learned are automatically iteratively optimized until the preset conditions are met, and the optimized coarse-grained model is output.
[0126] It should be noted that Bayesian optimization is a probabilistic statistical method used for optimizing high-cost, black-box functions. It constructs a probabilistic surrogate model of the objective function and combines this with a data collection function to balance exploration and exploitation, thereby efficiently finding the optimal parameters with a limited number of evaluations.
[0127] In one possible implementation, the Bayesian optimization method can be based on an automated operation framework; an automated operation framework refers to the automated operation of a process through full-process programming, which can avoid tedious human operations or erroneous manual interventions during the process.
[0128] In one possible implementation, the program for the automated execution framework can be written in Python; alternatively, it can be written in Fortran, C / C++, Java, or any other computer language to achieve Bayesian optimization.
[0129] This embodiment does not limit the specific method of Bayesian optimization, as long as the optimal function parameters can be found. In one possible implementation, step 3 may further include:
[0130] Step 31: Construct the objective function; the objective function is a function used to measure the performance of the coarse-grained model.
[0131] It should be noted that the objective function in this embodiment is used to quantitatively describe the difference between the coarse-grained molecular dynamics simulation results and the all-atom reference data, and this difference is used as the subsequent loss function. This embodiment does not limit the specific form of the objective function. In one possible implementation, the objective function can be in the form of structural error, force matching error, energy deviation, or a weighted combination of the above errors.
[0132] Step 32: Based on the current potential function parameters and their corresponding parameter evaluation results, establish a probabilistic surrogate model for the objective function.
[0133] It should be noted that a probabilistic surrogate model is a statistical model used to approximate complex or computationally expensive objective functions. Its output not only provides predicted values but also estimates of uncertainty. During optimization, this model acts as a substitute for the true objective function, guiding the parameter search direction and reducing the number of direct calls to costly computations or experiments. By combining the probabilistic surrogate model with Bayesian optimization to achieve automatic parameter optimization, coarse-grained models can approximate fully atomic reference results at a lower computational cost.
[0134] It should be noted that in this embodiment, the probabilistic surrogate model is used to approximate the mapping relationship between the coarse-grained potential function parameters and the objective function value. This embodiment does not limit the specific type of probabilistic surrogate model. In one possible implementation, the probabilistic surrogate model can adopt a Gaussian process model to simultaneously obtain the predicted value of the objective function and the corresponding uncertainty estimate. The Gaussian process model is a nonparametric regression model based on probability statistics, used to model unknown functions. It defines that any finite set of points follows a multivariate Gaussian distribution, achieving a unified description of the overall shape and uncertainty of the function. Gaussian processes are often used as a specific implementation of probabilistic surrogate models, suitable for optimization problems with a limited number of samples but high cost per evaluation.
[0135] It should be noted that in this embodiment, when establishing the probabilistic surrogate model for the first time, it can be established based on the preset potential function parameters and their corresponding parameter evaluation results; otherwise, all subsequent probabilistic surrogate models are established based on the potential function parameters obtained in step 35 and their corresponding parameter evaluation results.
[0136] Step 33: Define the acquisition function based on the probabilistic proxy model.
[0137] It should be noted that the acquisition function is a decision function used to guide the next step of parameter selection during the optimization process. It comprehensively utilizes the predicted values and uncertainty information provided by the surrogate model, balancing between "exploring unknown regions" and "utilizing existing optimal solutions." By maximizing the acquisition function, parameter optimization can be efficiently advanced within a limited computational budget.
[0138] This embodiment does not limit the specific type of acquisition function. In one possible implementation, the acquisition function uses expected improvement, which comprehensively considers prediction performance and model uncertainty. It can be used to balance exploring new parameter regions with utilizing currently known superior parameters. Expected improvement is a commonly used form of acquisition function used to measure the expected performance improvement relative to the current best result under given parameter conditions. It simultaneously considers the mean and uncertainty of the prediction error, enabling the optimization process to both refine the known superior parameter region and actively explore potential better solutions.
[0139] Step 34: Automatically select the next set of candidate potential function parameters from the preset parameter space using the acquisition function. It should be noted that, in this embodiment, the preset parameter space refers to the set of all possible potential function parameter values.
[0140] It should be noted that in this embodiment, the potential function parameters used for subsequent evaluation are selected by acquiring the function, which can avoid exhaustive search of the entire parameter space.
[0141] Step 35: Substitute the candidate potential function parameters into the coarse-grained model to be learned, perform coarse-grained molecular dynamics simulation calculations, and determine the objective function value based on the coarse-grained molecular dynamics simulation results and the objective function to obtain the parameter evaluation results.
[0142] It should be noted that this embodiment compares and verifies the coarse-grained molecular dynamics simulation results with all-atom reference data, automatically searches and updates the coarse-grained potential function parameters, and feeds the verification results back to the machine learning and automatic parameter optimization stages, forming a repeatable and convergent modeling closed loop, which can improve the reliability of the model.
[0143] It should be noted that the coarse-grained molecular dynamics simulation calculation process in this embodiment can refer to the above-mentioned all-atom molecular dynamics simulation calculation, and will not be repeated here.
[0144] Step 36: Feed back the candidate potential function parameters and their corresponding parameter evaluation results to step 32, and repeat steps 32 to 35 until the objective function value converges or the preset termination condition is met, and output the optimized coarse-grained model.
[0145] It should be noted that the purpose of feeding back the new potential function parameters and their corresponding parameter evaluation results to step 32 in this embodiment is to update the predictive capability of the probabilistic surrogate model.
[0146] This embodiment does not limit the specific types of preset termination conditions. In one possible implementation, the preset termination condition may include an error threshold, the magnitude of parameter change, or the maximum number of iterations.
[0147] It should be noted that this embodiment, through the aforementioned Bayesian optimization process, achieves automatic search and optimization of coarse-grained potential function parameters. This reduces the number of model evaluations while improving parameter optimization efficiency and model stability. Furthermore, this Bayesian optimization process can utilize an automation framework program written in Python. This program coordinates the above steps and the scheduling and invocation of the program, enabling fully automated management of the entire process from selecting candidate parameters, submitting simulation tasks, calculating the objective function, updating the surrogate model, and iterative convergence judgment.
[0148] It should be noted that the coarse-grained force field (determined by the potential function) of the coarse-grained model in this embodiment can be automatically learned from all-atom data using machine learning methods. In one possible implementation, by explicitly introducing Stockmayer-type ion-dipole interaction terms and dipole-dipole interaction terms into the coarse-grained model to describe particle polarity and charge correlation effects, a potential function including Stockmayer-type ion-dipole interaction terms and dipole-dipole interaction terms can be obtained. In addition to the coarse-grained force field determined by this potential function, through the machine learning process, the coarse-grained force field can also automatically select or combine different interaction forms according to reference data to adapt to different systems and simulation requirements. In one possible implementation, any coarse-grained molecular dynamics model or its derivatives, such as conventional empirical force fields, Martini force fields, dissipative forceon dynamics (DPD) models, MD-LB (molecular dynamics-lattice Boltzmann), and dissipative forceon dynamics (MPCD), can be established.
[0149] Based on the above embodiments, this application uses all-atom simulation data as a foundation and introduces machine learning methods to automatically learn and construct coarse-grained models that include polar correlation interactions, achieving unified modeling of ionic, small molecule, and polymeric systems. This method significantly reduces the degrees of freedom of the coarse-grained model while preserving key electrostatic and dipole physics mechanisms, and reduces human intervention through a highly automated process. This not only improves the computational efficiency of simulating complex systems but also enhances the accuracy, stability, and cross-system transferability of the model, providing a solution with significant technical advantages for the efficient simulation of complex multi-component systems.
[0150] The following examples illustrate the coarse-grained modeling process described above. Please refer to them. Figure 2 , Figure 2 This is a flowchart illustrating a low-scale coarse-grained modeling method based on machine learning, provided in an embodiment of this application. The overall process of this embodiment includes three main stages: all-atomic data acquisition, initial machine learning modeling, and Bayesian optimization to determine parameters. These stages can be executed sequentially or iteratively updated when preset conditions are met. The specific steps of the above scheme are as follows:
[0151] I. Acquisition of All-Atom Data
[0152] 1.1: System Construction
[0153] Select a target system based on the object to be studied, and construct a molecular system model containing ions, small molecules, and / or polymeric components. The target system can be a single-component system or a multi-component mixture system. The initial configuration, composition ratio, and simulation conditions of each component in the target system can be set according to the research requirements. The initial structure generation program can be written using Perl.
[0154] 1.2: Force Field and Parameter Setting
[0155] A full-atomic force field model is specified for the molecular system model. The full-atomic force field model may include bonding interaction parameters, non-bonding interaction parameters, and parameters such as charge distribution and atomic forces. The full-atomic force field used can be an existing known force field or a verified combination of parameters to describe the interaction behavior between atoms. Preferably, the full-atomic force field adopts the OPLS-AA force field.
[0156] 1.3: Full Atom Simulation Run
[0157] Based on molecular system models and all-atom force field models, all-atom molecular dynamics simulations are performed, and numerical integration methods are used to obtain the simulated trajectory of the target system evolving on a continuous time scale under given temperature, pressure or other thermodynamic conditions; among them, all-atom molecular dynamics simulations can be performed using LAMMPS software;
[0158] 1.4: Trajectory Data Acquisition
[0159] During or after the all-atom molecular dynamics simulation, the simulation trajectory is sampled to collect atomic coordinates, velocities, charge distributions, and other dynamic data related to interactions, forming the original dataset for subsequent modeling.
[0160] 1.5: Extraction of Reference Physical Quantities
[0161] The simulation results of the simulated trajectory data are analyzed, and reference physical quantities are extracted as all-atom data. The reference physical quantities may include, but are not limited to, radial distribution function, structure factor, mean square displacement, mean square radius of gyration, dipole orientation or other structural and thermodynamic properties, which are used as training targets or constraints for machine learning modeling. The analysis program can be written in Fortran.
[0162] II. Initial Machine Learning Modeling
[0163] 2.1: Atom-to-coarse-grained mapping
[0164] Based on the preset mapping rules, multiple sets of atoms in the target system are mapped to one or more coarse-grained particles, and the spatial coordinates, type identifiers and topological connections of the coarse-grained particles are determined, thereby reducing the number of degrees of freedom while maintaining the main structural features of the system.
[0165] 2.2: Feature Generation
[0166] Feature quantities are determined based on spatial coordinates, type identifiers, and topological connections. These feature quantities may include inter-particle distances, particle sizes, charge distribution characteristics, and dipole correlation descriptors, which are used to characterize the structure and interaction environment between particles and serve as input variables for machine learning models.
[0167] 2.3: Formal Definition of Potential Function
[0168] The potential function form is set for the interaction between coarse-grained particles to obtain the initial coarse-grained model. The potential function form can adopt the Soft potential, Coulomb potential, Lennard-Jones potential and FENE potential, and the Stockmayer-type dipole interaction term describing the permanent dipole-dipole interaction is introduced, which can retain the effective description of charge correlation, dipole effect and multi-component system.
[0169] 2.4: Initial Machine Learning Training
[0170] The initial coarse-grained model is validated using all-atom data, and the validated initial coarse-grained model is determined as the coarse-grained model to be learned. By initially training the coarse-grained potential function parameters based on the feature quantities and potential function form and using the reference physical quantities extracted from the all-atom simulation before machine learning, it can be ensured that the constructed coarse-grained model to be learned can statistically reproduce the structural or mechanical distribution characteristics of the all-atom system.
[0171] III. Bayesian Optimization for Parameter Determination
[0172] 3.1: Construction of the objective function
[0173] An objective function is constructed to measure the performance of the coarse-grained model. The objective function is used to quantitatively describe the difference between the coarse-grained molecular dynamics simulation results and the all-atom reference data, and this difference is used as the subsequent loss function. The objective function can take the form of structural error, force matching error, energy deviation, or a weighted combination of the above errors.
[0174] 3.2: Establishing the Proxy Model
[0175] Based on the current potential function parameters and their corresponding parameter evaluation results, a probabilistic surrogate model for the objective function is established. The probabilistic surrogate model is used to approximate the mapping relationship between the coarse-grained potential function parameters and the objective function value. A Gaussian process model can be used to simultaneously obtain the predicted value of the objective function and the corresponding uncertainty estimate.
[0176] 3.3: Definition of Acquisition Function
[0177] Based on the probabilistic surrogate model, an acquisition function is defined to guide the parameter search. The acquisition function can be improved with expectation, which takes into account both prediction performance and model uncertainty. It can be used to weigh the trade-off between exploring new parameter regions and utilizing currently known good parameters.
[0178] 3.4: Candidate Parameter Selection
[0179] By using the acquisition function, the next set of candidate potential function parameters can be automatically selected from the preset parameter space for subsequent evaluation, thus avoiding an exhaustive search of the entire parameter space.
[0180] 3.5: Parameter Evaluation and Update
[0181] The candidate potential function parameters are substituted into the coarse-grained model to be learned, and coarse-grained molecular dynamics simulation calculations are performed. The objective function value is determined based on the coarse-grained molecular dynamics simulation results and the objective function, and the parameter evaluation results are obtained.
[0182] 3.6: Iterative Convergence Judgment
[0183] After feeding back the candidate potential function parameters and their corresponding parameter evaluation results to 3.2, repeat steps 3.2 to 3.5 until the objective function value converges or the preset termination condition is met, and output the optimized coarse-grained model; wherein, the preset termination condition may include an error threshold, parameter change magnitude or maximum number of iterations.
[0184] The following describes a coarse-grained modeling apparatus, device, and computer-readable storage medium provided in the embodiments of this application. The coarse-grained modeling apparatus, device, and computer-readable storage medium described below can be referred to in correspondence with the coarse-grained modeling method described above.
[0185] Please refer to Figure 3 , Figure 3 A structural block diagram of a machine learning-based low-scale coarse-grained model modeling device provided in this application embodiment, the device may include:
[0186] The all-atom data acquisition module 100 is used to perform step 1: acquiring all-atom data of the target system;
[0187] The initial machine learning modeling module 200 is used to perform step 2: constructing a coarse-grained model to be learned; the coarse-grained model to be learned adopts a potential function form including a potential energy expression for describing polar correlation interaction, and the coarse-grained model to be learned is a model that has been validated by full-atom data.
[0188] The machine learning module 300 is used to perform step 3: taking all atomic data as the optimization target, automatically iteratively optimizes the potential function parameters to be determined in the coarse-grained model to be learned using machine learning methods until the preset conditions are met, and outputs the optimized coarse-grained model.
[0189] Based on the above embodiments, this application uses all-atom simulation data as a foundation and introduces machine learning methods to automatically learn and construct coarse-grained models that include polar correlation interactions, achieving unified modeling of ionic, small molecule, and polymeric systems. This method significantly reduces the degrees of freedom of the coarse-grained model while preserving key electrostatic and dipole physics mechanisms, and reduces human intervention through a highly automated process. This not only improves the computational efficiency of simulating complex systems but also enhances the accuracy, stability, and cross-system transferability of the model, providing a solution with significant technical advantages for the efficient simulation of complex multi-component systems.
[0190] Based on the above embodiments, step 1 may include:
[0191] Step 11: Construct a molecular system model of the target system;
[0192] Step 12: Specify the all-atom force field model for the molecular system model;
[0193] Step 13: Perform all-atom molecular dynamics simulation calculations based on the molecular system model and the all-atom force field model;
[0194] Step 14: Collect the dynamic data related to interactions generated during the all-atom molecular dynamics simulation to obtain the simulated trajectory data;
[0195] Step 15: Analyze the simulation results of the simulated trajectory data and extract the reference physical quantities as all-atom data.
[0196] Based on the above embodiments, step 2 may include:
[0197] Step 21: Perform coarse-grained modeling of the target system, mapping multiple sets of atoms in the target system to one or more coarse-grained particles, and determining the spatial coordinates, type identifiers, and topological connections of the coarse-grained particles;
[0198] Step 22: Determine the feature quantities based on spatial coordinates, type identifiers, and topological connectivity;
[0199] Step 23: Define the potential function form for the interaction between coarse-grained particles to obtain the initial coarse-grained model;
[0200] Step 24: Validate the initial coarse-grained model using all-atom data, and determine the validated initial coarse-grained model as the coarse-grained model to be learned.
[0201] Based on the above embodiments, step 23 may include:
[0202] Based on Stockmayer fluid theory, dipole moments are set for coarse-grained particles corresponding to polar components in the target system, and the potential energy expressions for the first polar correlation interaction and the second polar correlation interaction are defined according to the dipole moments to obtain the initial coarse-grained model.
[0203] The first polar correlation interaction represents the interaction between the coarse-grained particle corresponding to the ion and the coarse-grained particle with a dipole moment, while the second polar correlation interaction represents the interaction between two coarse-grained particles with a dipole moment.
[0204] Based on the above embodiments, step 24 may include:
[0205] Coarse-grained molecular dynamics simulations were performed based on the initial coarse-grained model, and the differences between the coarse-grained molecular dynamics simulation results and the all-atom data were verified to meet the preset range.
[0206] When a preset range is detected, the initial coarse-grained model is determined as the coarse-grained model to be learned.
[0207] Based on the above embodiments, step 3 may include:
[0208] Using all-atomic data as the optimization target, machine learning is performed based on the Bayesian optimization method. The potential function parameters to be determined in the coarse-grained model to be learned are automatically iteratively optimized until the preset conditions are met, and the optimized coarse-grained model is output.
[0209] Based on the above embodiments, step 3 may include:
[0210] Step 31: Construct the objective function; the objective function is a function used to measure the performance of the coarse-grained model;
[0211] Step 32: Based on the current potential function parameters and their corresponding parameter evaluation results, establish a probabilistic surrogate model for the objective function;
[0212] Step 33: Define the acquisition function based on the probabilistic surrogate model;
[0213] Step 34: Automatically select the next set of candidate potential function parameters from the preset parameter space using the acquisition function;
[0214] Step 35: Substitute the candidate potential function parameters into the coarse-grained model to be learned, perform coarse-grained molecular dynamics simulation calculations, and determine the objective function value based on the coarse-grained molecular dynamics simulation results and the objective function to obtain the parameter evaluation results;
[0215] Step 36: Feed back the candidate potential function parameters and their corresponding parameter evaluation results to step 32, and repeat steps 32 to 35 until the objective function value converges or the preset termination condition is met, and output the optimized coarse-grained model.
[0216] Based on the above embodiments, this application also provides a coarse-grained modeling device, including: a memory and a processor, wherein the memory is used to store a computer program; the processor is used to execute the computer program to implement the steps of the coarse-grained modeling method of the above embodiments. Of course, the coarse-grained modeling device may also include various necessary network interfaces, power supplies, and other components.
[0217] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the coarse-grained modeling method described in the above embodiments. The storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0218] This document uses specific examples to illustrate the principles and implementation methods of this application, and the various embodiments are progressively related. Each embodiment focuses on the differences from other embodiments, and similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, please refer to the corresponding method section description. The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of this application. For those skilled in the art, several improvements and modifications can be made to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0219] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
Claims
1. A low-scale coarse-grained modeling method based on machine learning, characterized in that, include: Step 1: Obtain all atomic data of the target system; Step 2: Construct the coarse-grained model to be learned; The coarse-grained model to be learned adopts a potential function form including a potential energy expression for describing polar correlation interactions, and the coarse-grained model to be learned is a model that has been verified by the full-atom data. Step 3: Using the all-atom data as the optimization target, the potential function parameters to be determined in the coarse-grained model to be learned are automatically iteratively optimized using machine learning methods until the preset conditions are met, and the optimized coarse-grained model is output.
2. The low-scale coarse-grained modeling method based on machine learning according to claim 1, characterized in that, Step 1 includes: Step 11: Construct a molecular system model of the target system; Step 12: Specify an all-atom force field model for the molecular system model; Step 13: Based on the molecular system model and the all-atom force field model, perform all-atom molecular dynamics simulation calculations; Step 14: Collect the dynamic data related to interactions generated during the all-atom molecular dynamics simulation to obtain the simulated trajectory data; Step 15: Analyze the simulation results of the simulated trajectory data and extract reference physical quantities as the all-atom data.
3. The low-scale coarse-grained modeling method based on machine learning according to claim 1, characterized in that, Step 2 includes: Step 21: Perform coarse-grained modeling on the target system, map multiple sets of atoms in the target system to one or more coarse-grained particles, and determine the spatial coordinates, type identifiers, and topological connections of the coarse-grained particles; Step 22: Determine the feature quantities based on the spatial coordinates, the type identifier, and the topological connection relationship; Step 23: Define the potential function form for the interaction between the coarse-grained particles to obtain the initial coarse-grained model; Step 24: Validate the initial coarse-grained model using the full-atom data, and determine the validated initial coarse-grained model as the coarse-grained model to be learned.
4. The low-scale coarse-grained modeling method based on machine learning according to claim 3, characterized in that, Step 23 includes: Based on Stockmayer fluid theory, dipole moments are set for the coarse-grained particles corresponding to the polar components in the target system, and the potential energy expressions for the first polar correlation interaction and the second polar correlation interaction are defined according to the dipole moments to obtain the initial coarse-grained model. The first polar correlation interaction represents the interaction between the coarse-grained particle corresponding to the ion and the coarse-grained particle having the dipole moment, and the second polar correlation interaction represents the interaction between two coarse-grained particles having the dipole moment.
5. The low-scale coarse-grained modeling method based on machine learning according to claim 3, characterized in that, Step 24 includes: Based on the initial coarse-grained model, perform coarse-grained molecular dynamics simulation calculations and verify whether the difference between the coarse-grained molecular dynamics simulation results and the all-atom data meets the preset range; When the preset range is met, the initial coarse-grained model is determined as the coarse-grained model to be learned.
6. The low-scale coarse-grained modeling method based on machine learning according to claim 1, characterized in that, Step 3 includes: Using the all-atomic data as the optimization target, machine learning is performed based on the Bayesian optimization method. The potential function parameters to be determined in the coarse-grained model to be learned are automatically iteratively optimized until the preset conditions are met, and the optimized coarse-grained model is output.
7. The low-scale coarse-grained modeling method based on machine learning according to claim 6, characterized in that, Step 3 includes: Step 31: Construct the objective function; the objective function is a function used to measure the performance of the coarse-grained model; Step 32: Based on the current potential function parameters and their corresponding parameter evaluation results, establish a probabilistic proxy model for the objective function; Step 33: Define the acquisition function based on the probabilistic proxy model; Step 34: Automatically select the next set of candidate potential function parameters from the preset parameter space using the acquisition function; Step 35: Substitute the candidate potential function parameters into the coarse-grained model to be learned, perform coarse-grained molecular dynamics simulation calculations, and determine the objective function value based on the coarse-grained molecular dynamics simulation results and the objective function to obtain the parameter evaluation results; Step 36: Feed back the candidate potential function parameters and their corresponding parameter evaluation results to step 32, and repeat steps 32 to 35 until the objective function value converges or the preset termination condition is met, and output the optimized coarse-grained model.
8. A low-scale coarse-grained modeling device based on machine learning, characterized in that, include: The all-atom data acquisition module is used to perform step 1: acquiring all-atom data of the target system; The initial machine learning modeling module is used to perform step 2: constructing a coarse-grained model to be learned; the coarse-grained model to be learned adopts a potential function form including a potential energy expression for describing polar correlation interaction, and the coarse-grained model to be learned is a model that has been validated by the full-atom data; The machine learning module is used to perform step 3: taking the all-atomic data as the optimization target, automatically iteratively optimizing the potential function parameters to be determined in the coarse-grained model to be learned using machine learning methods until the preset conditions are met, and outputting the optimized coarse-grained model.
9. A low-scale coarse-grained modeling device based on machine learning, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the machine learning-based low-scale coarse-grained model modeling method as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the machine learning-based low-scale coarse-grained model modeling method as described in any one of claims 1 to 7.