Information processing apparatus, information processing system, and method
The information processing apparatus addresses the challenge of simulating atomic structures by generating a second model optimized for a large number of atoms, achieving high precision and scalability in molecular dynamics simulations.
Patent Information
- Application Number
- JP2023211218
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-26
AI Technical Summary
Existing methods for simulating atomic structures face a trade-off between accuracy and the number of atoms that can be simulated, with conventional classical potentials offering high atom counts but low accuracy, and quantum chemical calculations providing high accuracy but limited atom counts.
An information processing apparatus that generates a second model based on parameters from a first model trained using specific data, where the second model is optimized for analyzing atomic structures and can handle a larger number of atoms with improved precision.
Enables high-precision simulations of atomic structures for a large number of atoms, improving both accuracy and scalability in molecular dynamics simulations.
Smart Images

Figure 2025095299000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to an information processing apparatus, an information processing system, and a method.
Background Art
[0002] Conventionally, as methods for simulating atomic structures, there are Coupled-cluster singles-and-doubles (CCSD), simulations using classical molecular dynamics potentials (hereinafter referred to as classical potentials), quantum chemical calculations such as Density Functional Theory (DFT), and machine learning potentials learned with DFT as correct data.
[0003] However, in simulations using conventional classical potentials, although the number of atoms that can be simulated is large, the accuracy may be low. Also, in CCSD, DFT, and machine learning potentials learned with DFT as correct data, although the accuracy is high, the number of atoms that can be simulated is limited. For this reason, a technique that enables high-precision simulations targeting a large number of atoms has been demanded.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] The problem to be solved by the present disclosure is to enable high-precision simulations targeting a large number of atoms.
Means for Solving the Problems
[0006] The information processing apparatus according to the embodiment includes at least one processor and at least one memory. The at least one processor generates a second model different from the first model using at least some of the parameters of the first model trained using the first data. The second model is a model that outputs an analysis result of the atomic structure when the atomic structure is input. At least some of the parameters are determined based on the type of atoms to be analyzed by the second model.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Modes for Carrying Out the Invention
[0008] Hereinafter, embodiments will be described in detail with reference to the drawings.
[0009] (Embodiment) FIG. 1 is a block diagram showing an example of the configuration of an information processing system S according to the present embodiment. As shown in FIG. 1, the information processing system S of the present embodiment includes an information processing apparatus 1 and an external apparatus 9A.
[0010] The information processing apparatus 1 is a computer that generates an artificial intelligence model that executes MD (Molecular Dynamics) simulation and provides a simulation using the artificial intelligence model. Details of the functions of the information processing apparatus 1 will be described later with reference to FIG. 2 and subsequent figures. The artificial intelligence model is a non-limiting example of the "model" in the present embodiment. Also, the MD simulation is a non-limiting example of the "simulation" in the present embodiment. The information processing apparatus 1 is an example of the first information processing apparatus in the present embodiment.
[0011] The information processing apparatus 1 includes, for example, a computer 30 and an external apparatus 9B connected to the computer 30 via a device interface 39. The computer 30 includes, as an example, a processor 31, a main storage device (memory) 33, an auxiliary storage device (memory) 35, a network interface 37, and a device interface 39. The information processing apparatus 1 may be realized as a computer 30 in which the processor 31, the main storage device 33, the auxiliary storage device 35, the network interface 37, and the device interface 39 are connected via a bus 41.
[0012] The computer 30 shown in FIG. 1 includes one of each component, but may include a plurality of the same components. Also, in FIG. 1, one computer 30 is shown, but software may be installed on a plurality of computers, and each of the plurality of computers may execute the same or different parts of the software. In this case, it may be in the form of distributed computing in which each computer communicates via a network interface 37 or the like to execute processing. That is, the information processing apparatus 1 in the present embodiment may be configured as a system in which one or a plurality of computers execute instructions stored in one or a plurality of storage devices to realize various functions described later. Also, the information transmitted from the terminal may be processed by one or a plurality of computers provided on the cloud, and the processing result may be transmitted to a terminal such as a display device (display unit) corresponding to the external device 9B.
[0013] The various operations of the information processing apparatus 1 in the present embodiment may be executed in parallel using one or a plurality of processors or using a plurality of computers via a network. Also, the various operations may be allocated to a plurality of arithmetic cores in the processor and executed in parallel. Also, part or all of the processing, means, etc. of the present disclosure may be executed by at least one of a processor and a storage device provided on the cloud that can communicate with the computer 30 via a network. Thus, the various operations described later in the present embodiment may be in the form of parallel computing by one or a plurality of computers.
[0014] The processor 31 may be an electronic circuit (processing circuit, Processing circuit, Processing circuitry, CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit), etc.) including a control device and an arithmetic device of the computer 30. Further, the processor 31 may be a semiconductor device including a dedicated processing circuit, etc. The processor 31 is not limited to an electronic circuit using electronic logic elements, and may be realized by an optical circuit using optical logic elements. Further, the processor 31 may include an arithmetic function based on quantum computing.
[0015] The processor 31 can perform arithmetic processing based on data and software (program) input from each device, etc. of the internal configuration of the computer 30, and output the arithmetic result and control signal to each device, etc. The processor 31 may control each component constituting the computer 30 by executing the OS (Operating System) of the computer 30 and applications, etc.
[0016] The information processing apparatus 1 in the present embodiment may be realized by one or a plurality of processors 31. Here, the processor 31 may refer to one or a plurality of electronic circuits arranged on one chip, or may refer to one or a plurality of electronic circuits arranged on two or more chips or two or more devices. When using a plurality of electronic circuits, each electronic circuit may communicate by wire or wirelessly.
[0017] The main memory device 33 is a storage device that stores instructions executed by the processor 31 and various data, etc., and the information stored in the main memory device 33 is read by the processor 31. The auxiliary storage device 35 is a storage device other than the main memory device 33. Note that these storage devices mean any electronic components capable of storing electronic information, and may be semiconductor memories. The semiconductor memory may be either a volatile memory or a non-volatile memory. The storage device for storing various data used in the information processing apparatus 1 according to the present embodiment may be realized by the main memory device 33 or the auxiliary storage device 35, or may be realized by a built-in memory built into the processor 31. For example, the storage unit in the present embodiment may be realized by the main memory device 33 or the auxiliary storage device 35.
[0018] A plurality of processors may be connected (coupled) to one storage device (memory), or a single processor 31 may be connected. A plurality of storage devices (memories) may be connected (coupled) to one processor. When the information processing apparatus 1 in the present embodiment is configured by at least one storage device (memory) and a plurality of processors connected (coupled) to this at least one storage device (memory), the configuration may include that at least one of the plurality of processors is connected (coupled) to at least one storage device (memory). Further, this configuration may be realized by storage devices (memories) and the processor 31 included in a plurality of computers. Furthermore, the configuration in which the storage device (memory) is integrated with the processor 31 (for example, a cache memory including L1 cache and L2 cache) may be included.
[0019] The network interface 37 is an interface for connecting to the communication network 5, either wirelessly or by wire. The network interface 37 may use an appropriate interface, such as one that conforms to an existing communication standard. Information exchange may be performed between the external device 9A connected via the communication network 5 and the network interface 37. Note that the communication network 5 may be any one of a WAN (Wide Area Network), LAN (Local Area Network), PAN (Personal Area Network), or a combination thereof, as long as information exchange can be performed between the computer 30 and the external device 9A. Examples of a WAN include the Internet, etc., examples of a LAN include IEEE802.11, Ethernet (registered trademark), etc., and examples of a PAN include Bluetooth (registered trademark), NFC (Near Field Communication), etc.
[0020] The device interface 39 is an interface such as a USB (Universal Serial Bus) for directly connecting to an output device such as a display device, an input device, and an external device 9B. Note that the output device may have a speaker that outputs sound, etc.
[0021] The external device 9A is a device connected to the computer 30 via the communication network 5. The external device 9B is a device directly connected to the computer 30 via the device interface 39. Note that the information processing system S may include a plurality of external devices 9A and / or external devices 9B. In this case, the information processing apparatus 1 may be communicably connected to each of the plurality of external devices 9A and / or external devices 9B.
[0022] As an example, the external device 9A may be a computer in which at least one processor, a main storage device, an auxiliary storage device, a network interface, and a device interface are connected via a bus. The external device 9A is an example of another information processing apparatus and a second information processing apparatus in the present embodiment.
[0023] The information processing apparatus 1 has a trained artificial intelligence model according to this embodiment, and may be owned by a business operator that provides the use of the trained artificial intelligence model to users as a service. In this case, the external device 9A may be used by a user who uses the trained artificial intelligence model. The user can operate the external device 9A to access the information processing apparatus 1 and execute a simulation using the trained artificial intelligence model stored in the information processing apparatus 1. Further, the processor of the external device 9A may execute a process of transmitting information on the type of atom to be analyzed desired by the user and other information related to learning or analysis to the information processing apparatus 1.
[0024] Also, as another example, the external device 9A or the external device 9B may be an input device (input unit). The input device is, for example, a device such as a camera, a microphone, a motion capture, various sensors, a keyboard, a mouse, or a touch panel, and provides the acquired information to the computer 30. Further, the external device 9A or the external device 9B may be a device including an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0025] Also, as another example, the external device 9A or the external device 9B may be an output device (output unit). The output device may be, for example, a display device (display unit) such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), or an organic EL (Electro Luminescence) panel, or may be a speaker that outputs sound or the like. Further, the external device 9A or the external device 9B may be a device including an output device, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0026] As another example, the external device 9A or the external device 9B may be a storage device (memory). For example, the external device 9A may be a network storage or the like, and the external device 9B may be a storage such as an HDD.
[0027] As another example, the external device 9A or the external device 9B may be a device having some functions of the components of the information processing apparatus 1 in the present embodiment. That is, the computer 30 may transmit or receive part or all of the processing results of the external device 9A or the external device 9B.
[0028] FIG. 2 is a diagram showing an example of a functional block of the processor 31 according to the present embodiment. The functions realized by the processor 31 include, for example, a first training data generation unit 311, a first training unit 312, an acquisition unit 313, an editing unit 314, a second training data generation unit 315, a second training unit 316, an evaluation unit 317, and an inference unit 318. The functions realized by the first training data generation unit 311, the first training unit 312, the acquisition unit 313, the editing unit 314, the second training data generation unit 315, the second training unit 316, the evaluation unit 317, and the inference unit 318 are stored as programs in, for example, the main storage device 33 or the auxiliary storage device 35. The processor 31 reads and executes the program stored in the main storage device 33 or the auxiliary storage device 35 or the like, whereby the functions related to the first training data generation unit 311, the first training unit 312, the acquisition unit 313, the editing unit 314, the second training data generation unit 315, the second training unit 316, the evaluation unit 317, and the inference unit 318 can be realized.
[0029] The first training data generation unit 311 generates training data (learning data) for pre-training the first model. In the present embodiment, as an example, the first model to be pre-trained is an MTP (Moment Tensor Potential). In the present embodiment, the first model before pre-training is also referred to as an initial model. The training data for pre-training is an example of the first data in the present embodiment.
[0030] In this embodiment, the first training data generation unit 311 generates training data for pre-training using a trained model. The model used for generating the training data for pre-training is a model different from the first model and the second model described later. As the model for generating the training data, for example, an NNP (Neural Network Potential) trained with DFT (Density Functional Theory) as correct answer data can be adopted. The trained NNP used for generating the training data for pre-training is, for example, a general-purpose neural network potential capable of corresponding to any atomic structure. The trained NNP includes, for example, an input layer, one or more graph convolution (graph convolution) layers, and an output layer. The trained NNP is an example of the third model in this embodiment. The training data for pre-training is an example of the first data in this embodiment.
[0031] Since the processing by the trained NNP is faster than the processing by DFT, by using the trained NNP trained with DFT as correct answer data in advance for creating the training data, the time required for generating the training data can be shortened compared to directly using DFT.
[0032] The training data for pre-training includes an atomic structure and an analysis result of the atomic structure. The atomic structure includes the types (elements) of a plurality of atoms and position information (atomic coordinates). The analysis result of the atomic structure includes at least information regarding the energy or force of the atomic structure. Further, the analysis result may further include an index for evaluating the stress, density, or other physical property values of the atomic structure. As an index for evaluating the physical property values of the atomic structure, for example, there is an elastic modulus.
[0033] In addition, the training data may further include information about the Cell, periodic boundary conditions, or the type of simulation. The Cell is a box in which the environment where the simulation is executed is defined, and is set for each type of simulation. The type of simulation, also called the use case, is defined by, for example, the preconditions when analyzing the atomic structure. Specifically, it is the temperature that is a precondition for the simulation, or the surface structure of the substance that constitutes the atomic structure, etc. For example, as an example of the type of simulation, there is "a simulation that gives the initial atomic coordinates and shows how the structure changes when the temperature is raised to xx degrees from there".
[0034] The first training unit 312 generates a trained first model by training using the first data. More specifically, the first training unit 312 generates a trained first model by training the initial model with the training data for pre-training generated by the first training data generation unit 311.
[0035] The acquisition unit 313 acquires information used for generating a second model, which will be described later, from an external device 9A. More specifically, the acquisition unit 313 acquires, for example, information including the type of atoms to be analyzed from the external device 9A. In the present embodiment, the type of atoms may include the type of elements to be analyzed and differences in the environments in which the atoms to be analyzed are placed. For example, when the environments in which the atoms are placed are different, they may be regarded as different types of atoms. The atoms or elements to be analyzed are, for example, the atoms or elements included in the atomic structure that the user desires to analyze. The atoms or elements to be analyzed become the analysis targets of the second model generated based on the first model. The second model will be described later. Note that the acquisition source of the type of atoms to be analyzed is not limited to the external device 9A. For example, the acquisition unit 313 may acquire the type of atoms to be analyzed from an external device 9B. Alternatively, the type of atoms to be analyzed may be stored in advance in an auxiliary storage device 35 or the like. Further, the acquisition unit 313 may acquire the atomic structure to be analyzed, such as a molecular structure, a crystal structure, a surface structure, or a combination thereof, from the external device 9A or other acquisition sources. Furthermore, the acquisition unit 313 may acquire the phenomena or physical properties to be analyzed, such as elastic modulus, chemical reaction, density, vibration characteristics, diffusion, etc., or the environment to be analyzed, such as temperature and pressure, from the external device 9A or other acquisition sources.
[0036] The editing unit 314 generates a second model different from the first model using at least some parameters of the trained first model trained using the training data for pre-training. Since the second model is generated based on the first model, the first model and the second model are artificial intelligence models of the same type. Specifically, both the first model and the second model in the present embodiment are MTP models.
[0037] Since the MTP model does not have processing corresponding to multiple layers like other neural networks, such as GNN (Graph Neural Network), the processing load is low. Therefore, in the present embodiment, by configuring the first model and the second model with MTP models, the processing can be speeded up.
[0038] FIG. 3 is a diagram for explaining an example of the configuration of the first model 20 according to the present embodiment. The first model 20 is a model that outputs an analysis result of an atomic structure when receiving an input of the atomic structure.
[0039] As described above, since the type of the first model 20 is an MTP model, it includes a plurality of structural descriptors 201 and an ML (Machine Learning) model 202. The plurality of structural descriptors 201 correspond to the elements input to the first model 20. The plurality of structural descriptors 201 are an example of parameters in the present embodiment. Also, the ML model 202 of the trained first model 20 has parameters based on the first training data and hyperparameters that can be set by the user. The parameters based on the first training data that the ML model 202 has are, for example, regression coefficients when the ML model is a linear regression model. The various parameters that the ML model 202 has may also be an example of parameters in the present embodiment.
[0040] The structural descriptor 201 is a parameter determined based on rules for each type of atom. More specifically, the structural descriptor 201 is a parameter determined based on rules for each type of element, and the parameter is determined for each pair of two elements. For example, the input data of the first model 20 in the inference process is an atomic structure, and more specifically, it is information in which a plurality of atomic types (elements) and their position information (coordinate information, etc.) are defined. One structural descriptor 201 receives the input of the two elements and the position information of the two elements for each pair of the types of two elements included in the input atomic structure.
[0041] The ML model 202 is located after the plurality of structure descriptors 201, and executes processing based on the data obtained from the plurality of structure descriptors 201 to output the analysis result of the atomic structure. Specifically, the ML model 202 of the present embodiment is a regression model based on a polynomial basis function. The ML model 202 is updated by training.
[0042] The editing unit 314 of the present embodiment generates a second model including a part of the plurality of structure descriptors 201 included in the trained first model 20 and the ML model 202. The plurality of structure descriptors 201 included in the second model are the structure descriptors 201 corresponding to the elements to be analyzed acquired by the acquisition unit 313. In other words, the editing unit 314 determines the structure descriptors 201 included in the second model based on the types of atoms to be analyzed acquired by the acquisition unit 313.
[0043] FIG. 4 is a diagram for explaining an example of the configuration of the second model 21 according to the present embodiment. The second model 21 is a model that outputs the analysis result of the atomic structure when receiving the input of the atomic structure.
[0044] In the present embodiment, as an example, it is assumed that the element to be analyzed acquired by the acquisition unit 313 does not contain element B. In this case, the editing unit 314 excludes the structure descriptor 201 corresponding to element B from the plurality of structure descriptors 201 included in the trained first model 20. Then, the editing unit 314 generates a second model 21 including the plurality of structure descriptors 201 corresponding to the elements to be analyzed and the ML model 202. In the present embodiment, it is not necessary to change the ML model 202 when generating the second model 21. Since the second model 21 is untrained at the time of generation, the various parameters of the ML model 202 are the same as those of the trained first model 20.
[0045] In the second model 21, only some of the plurality of structure descriptors 201 that the trained first model 20 has and that are extracted by the editing unit 314 are included. Therefore, the number of structure descriptors 201 included in the second model 21 is smaller than the number of structure descriptors 201 included in the first model 20. For this reason, in the present embodiment, the first model 20 can analyze more types of atoms than the second model 21 can analyze.
[0046] In the inference process in the MTP model, processing also occurs for the structure descriptors 201 corresponding to elements not included in the input atomic structure. Therefore, in the second model 21, the processing speed can be improved by excluding unnecessary structure descriptors 201 in advance.
[0047] Returning to FIG. 2, the second training data generation unit 315 generates training data for fine-tuning the second model 21. In the present embodiment, the second training data generation unit 315 generates training data for fine-tuning, for example, using an NNP trained with DFT as correct data. The trained NNP used for generating the training data for fine-tuning may be the same model as the trained NNP used for generating the above-described training data for pre-training. The trained NNP is also an example of the third model in the present embodiment. Note that the model used for generating the training data for pre-training and the model used for generating the training data for fine-tuning may be different models. The training data for fine-tuning is an example of the second data in the present embodiment.
[0048] The second training data generation unit 315 in the present embodiment generates second data including an atomic structure including the element to be analyzed acquired by the acquisition unit 313 and the analysis result of the atomic structure by a trained NNP trained with DFT as correct data. Further, the atomic structure included in the second data does not include elements other than the element to be analyzed. Further, the second data may further include information about Cell, periodic boundary conditions, or the type of simulation.
[0049] The second training unit 316 generates a trained second model 21 by training the second model 21 before training using second data. The trained second model 21 is a model that outputs an analysis result of an atomic structure when the atomic structure is input. The second training unit 316 stores the generated trained second model 21 in, for example, the auxiliary storage device 35.
[0050] More specifically, the second training unit 316 fine-tunes the second model 21 before training according to the element to be analyzed. By fine-tuning, the ML model 202 included in the second model 21 before training is trained, and the parameters of the ML model 202 are updated.
[0051] In the above pre-training, in order to train the first model 20 in a general-purpose manner to be applicable to any atomic structure, it is desirable to include various atomic structures in the training data without restricting the atomic structure including the element to be analyzed assumed in actual use. Therefore, the types of atomic structures included in the first data, which is the training data for pre-training, are more than the types of atomic structures included in the second data, which is the training data for fine-tuning.
[0052] On the other hand, in fine-tuning, by training the second model 21 specialized for the element to be analyzed assumed in actual use, the generality decreases, but the processing speed improves with respect to the target element. Therefore, in the inference process regarding the target element, the number of atoms that can be simulated in one inference process by the second model 21 increases compared to the first model 20.
[0053] The second training unit 316 may perform training up to a predetermined number of epochs. Alternatively, the second training unit 316 may divide the training data (second data) into a learning set and a validation set, observe the changes in the respective losses (magnitudes of errors) during training, and continue training until it is determined that the loss has stopped decreasing. Note that the condition for completion of training may be specified by the user via the external device 9A, for example. Alternatively, the condition for completion of training may be stored in the auxiliary storage device 35 or the like of the information processing apparatus 1 in advance based on the accuracy and processing speed required for the second model 21 after training. For example, the user may perform operations such as changing the learning rate during training or changing the number of epochs according to the model capacity. The model capacity is the size of the structure of the MTP model. For example, the larger the model capacity, the larger the number of atoms for which the second model 21 after training analyzes relationships. The higher the accuracy of the analysis processing as the model capacity increases, but the slower the processing speed.
[0054] Note that in the pre-training by the first training data generation unit 311 described above, the completion of training may also be determined based on the number of epochs or the losses of the training data and the validation data during learning, or the user may be accepted to set or change the end condition of training. The conditions for completion of training may be different between pre-training and fine-tuning.
[0055] The evaluation unit 317 compares the analysis result output from the second model 21 with the verification data and outputs the comparison result as a numerical value.
[0056] The verification data is generated by a generation means different from the first model 20 and the second model 21. More specifically, the verification data is an analysis result by a third model that generated at least one of the first data and the second data. For example, when the first data, which is training data for pre-training, and the second data, which is training data for fine-tuning, are generated by a trained NNP trained with DFT as correct answer data, the evaluation unit 317 may generate verification data by the trained NNP.
[0057] The verification data includes, for example, an atomic structure and an analysis result of the atomic structure, similar to the training data. As described above, since the trained NNP is faster in processing speed than DFT, by using a trained NNP that has been trained with DFT as the correct data for creating the verification data, the time required for generating the verification data can be shortened compared to directly using DFT. Note that the evaluation unit 317 may use a part of the generated first data or second data as the verification data. Further, the verification data may further include a Cell, periodic boundary conditions, or the type of simulation.
[0058] The evaluation unit 317 inputs the atomic structure of the verification data into the trained second model 21, and compares the analysis result of the atomic structure output from the trained second model 21 with the analysis result of the atomic structure of the verification data. Further, information about a Cell, periodic boundary conditions, or the type of simulation may be input into the trained second model 21 as input data during verification.
[0059] The evaluation unit 317 evaluates the analysis result of the trained second model 21 with the analysis result of the atomic structure of the verification data being regarded as correct. Specifically, the evaluation unit 317 compares the analysis result of the atomic structure output from the trained second model 21 with the analysis result of the atomic structure of the verification data, and evaluates whether or not they match.
[0060] Further, the evaluation unit 317 evaluates whether or not the simulation process of the trained second model 21 fails. The simulation process failing means, for example, that the simulation process of the trained second model 21 input with the input data for verification does not end normally and no analysis result can be obtained.
[0061] The numerical values representing the comparison results are, for example, the success rate of the simulation process, the MAE (Mean Absolute Error) of the analysis results, or the time required for the simulation process. The success rate of the simulation process is the ratio of the number of simulation processes that ended normally to the total number of simulation processes. The MAE of the analysis results is the average of the absolute values of the differences between the analysis results of the atomic structure of the verification data, which is the true value, and the analysis results of the trained second model 21. For example, when the analysis result is the density of the atomic structure, the evaluation unit 317 calculates the average of the absolute values of the differences between the density values of the training data generated using the trained NNP and the density output from the trained second model 21. Note that the indices used for evaluating the trained second model 21 by the evaluation unit 317 are not limited to these, and any index can be adopted.
[0062] The method for outputting the analysis results by the evaluation unit 317 is not particularly limited. It may be displayed on a display, stored as data in various storage devices, or transmitted to the external device 9A or the external device 9B. The evaluation unit 317 may, for example, output the analysis results output from the trained second model 21 and the analysis results of the atomic structure of the verification data side by side for each evaluation index to a display or the like.
[0063] In addition, the evaluation unit 317 may evaluate the trained second model 21 separately for each type of simulation. In this case, the type of simulation executed by the trained second model 21 is the same as the type of simulation of the trained NNP at the time of generating the verification data.
[0064] If the evaluation of the trained second model 21 by the evaluation unit 317 does not meet the predetermined level, fine-tuning may be additionally performed on the atomic structure or simulation type with a low evaluation. For example, the second training data generation unit 315 may additionally generate training data (second data) for fine-tuning based on the evaluation by the evaluation unit 317. Further, the second training unit 316 may further fine-tune the trained second model 21 based on the additionally generated second data. That is, the evaluation unit 317, the second training data generation unit 315, and the second training unit 316 may perform active learning of the second model 21.
[0065] The inference unit 318 executes simulation processing by the trained second model 21. More specifically, the inference unit 318 obtains an analysis result of the atomic structure output from the fine-tuned second model 21 by inputting the atomic structure to the fine-tuned second model 21. The inference unit 318 executes simulation processing, for example, when receiving a user operation input by the external device 9A. The atomic structure to be analyzed in the simulation processing may be acquired from, for example, the external device 9A or the external device 9B. Alternatively, the atomic structure to be analyzed in the simulation processing may be stored in the auxiliary storage device 35 of the information processing apparatus 1.
[0066] When the user operates the external device 9A to access the information processing apparatus 1 and executes a simulation using the trained second model 21 stored in the information processing apparatus 1, the dynamics calculation using energy or force output as an analysis result from the trained second model 21 is executed using the computing resources on the external device 9A side.
[0067] Further, as input data for the simulation processing, information about Cell, periodic boundary conditions, or the type of simulation may be input to the second model 21 after training.
[0068] In addition, when the external device 9A is a computer used by a user, the structure and parameters of the second model 21 may not be viewable by the user who uses the external device 9A. For example, the inference unit 318 may display, via a browser or other application, a screen including an input field where the user can input the atomic structure to be analyzed and other input data, and an execution button for simulation processing, on the display of the external device 9A. In this case, even if the user cannot directly refer to the second model 21, the user can execute a simulation using the trained second model 21. For the transfer of data between the application viewable by the user and the trained second model 21, technologies such as socket communication or shared memory can be adopted, for example. By transferring data through socket communication or shared memory, high-speed communication becomes possible without the user directly accessing the trained second model 21. Also, regarding the first model 20, the structure and parameters may not be viewable by the user who uses the external device 9A. Further, the first model 20 and the second model 21 before fine-tuning may not be accessible or used by the user.
[0069] The inference unit 318 outputs the analysis result of the simulation process to, for example, the external device 9A or the external device 9B. The method of outputting the analysis result by the inference unit 318 is not particularly limited, and it may be displayed on a display or stored in various storage devices as data.
[0070] Next, the processing flow from training to evaluation executed by the information processing apparatus 1 of the present embodiment configured as described above will be described. FIG. 5 is a flowchart showing an example of the processing flow from training to evaluation executed by the information processing apparatus 1 according to the present embodiment.
[0071] First, the first training data generation unit 311 generates pre-training training data (first data) (S1). Specifically, the first training data generation unit 311 inputs an atomic structure into an NNP trained with DFT as correct answer data, and obtains an analysis result of the atomic structure output from the trained NNP. When generating the pre-training training data, not only the elements to be analyzed but also atomic structures containing various elements are input into the NNP. The first training data generation unit 311 generates a plurality of sets of training data with the input atomic structure and the analysis result of the atomic structure output as one set. Further, the first training data generation unit 311 may input information about Cell, periodic boundary conditions, or the type of simulation into the trained NNP and include this information in the training data.
[0072] Then, the first training unit 312 pre-trains the first model 20 with the first data (S2).
[0073] Then, the acquisition unit 313 acquires the type of element to be analyzed, for example, from an external device 9A (S3).
[0074] Then, the editing unit 314 extracts a structure descriptor 201 corresponding to the atomic structure to be analyzed acquired by the acquisition unit 313 from the pre-trained first model 20, and generates a second model 21 including the extracted structure descriptor 201 and the ML model 202 (S4).
[0075] The second training data generation unit 315 generates training data for fine-tuning (second data) based on the type of element to be analyzed acquired by the acquisition unit 313 (S5). Specifically, the second training data generation unit 315 inputs an atomic structure containing only the element to be analyzed into the NNP trained with DFT as the correct answer data, and obtains the analysis result of the atomic structure output from the trained NNP. The second training data generation unit 315 generates a plurality of sets of training data with the input atomic structure and the analysis result of the atomic structure output as one set. Further, the second training data generation unit 315 may input information about Cell, periodic boundary conditions, or the type of simulation into the trained NNP and include this information in the training data. The trained NNP used for generating the training data for fine-tuning may be the same as the model used in the process of generating the pre-training data in S1.
[0076] The second training unit 316 fine-tunes the second model 21 using the second data (S6).
[0077] Then, the evaluation unit 317 generates verification data using the trained NNP trained with DFT as the correct answer data (S7), and evaluates the fine-tuned second model 21 using the verification data (S8). Specifically, the evaluation unit 317 inputs the atomic structure input to the trained NNP during the generation of the verification data into the fine-tuned second model 21, and compares the analysis result output from the fine-tuned second model 21 with the verification data. Further, the evaluation unit 317 evaluates whether the simulation process of the trained second model 21 fails. The evaluation unit 317 outputs the evaluation result, and the processing of this flowchart ends.
[0078] If the evaluation of the trained second model 21 does not meet the predetermined level, the processing from S5 to S7 may be repeatedly executed until the evaluation reaches the predetermined level.
[0079] Also, in FIG. 5, the processes from the generation to the evaluation of the pre-training training data are described continuously, but there may be a time interval between each process. For example, after the processes up to pre-training are executed in advance, fine-tuning may be executed at an arbitrary timing according to the requests of individual users.
[0080] As described above, the information processing apparatus 1 according to the present embodiment generates a second model 21 different from the first model 20 using at least some parameters of the first model 20 trained using the first data. Specifically, the second model 21 according to the present embodiment has at least a part of a plurality of structure descriptors 201 included in the trained first model 20. Therefore, according to the information processing apparatus 1 of the present embodiment, by further generating the second model 21 from the trained first model 20, the processing speed of the simulation can be improved, and a high-precision simulation targeting a large number of atoms can be enabled.
[0081] Also, the information processing apparatus 1 according to the present embodiment determines, based on the type of element to be analyzed in the second model 21, the parameters of the trained first model 20, for example, among the plurality of structure descriptors 201, those included in the second model 21. Therefore, according to the information processing apparatus 1 of the present embodiment, parameters suitable for the type of element to be analyzed in the second model 21 can be adopted. Specifically, the information processing apparatus 1 according to the present embodiment extracts some of the plurality of structure descriptors 201 included in the first model 20 to generate the second model 21. Some of the structure descriptors 201 to be extracted are determined based on the type of element to be analyzed in the second model 21. Therefore, the information processing apparatus 1 according to the present embodiment can improve the speed of the simulation process by reducing the structure descriptors 201 related to elements outside the analysis target.
[0082] In addition, the information processing apparatus 1 of the present embodiment trains the second model 21 using the second data. That is, the information processing apparatus 1 of the present embodiment further fine-tunes the second model 21 generated from the pre-trained first model 20. Therefore, according to the information processing apparatus 1 of the present embodiment, the accuracy of the simulation can be improved with respect to the processing related to the second data.
[0083] In addition, the first data and the second data of the present embodiment are data generated using a trained third model different from the first model 20, specifically, an NNP trained with DFT as correct answer data. The processing by the trained NNP is faster than the processing by DFT. Therefore, by using the trained NNP trained with DFT as correct answer data in advance for creating the first data and the second data, the time required for generating the first data and the second data can be shortened compared to directly using DFT.
[0084] In addition, the first data and the second data of the present embodiment include an atomic structure and an analysis result of the atomic structure, and the types of atomic structures included in the first data are more than the types of atomic structures included in the second data. In pre-training, in order to train the first model 20 to be generally applicable to any atomic structure, it is desirable to include various atomic structures in the training data without restricting to the atomic structures including the elements to be analyzed assumed in actual use. Therefore, as in the present embodiment, since the types of atomic structures included in the first data, which is the training data for pre-training, are more than the types of atomic structures included in the second data, which is the training data for fine-tuning, the first model 20 can be pre-trained to be applicable to various atomic structures.
[0085] Also, the information processing apparatus 1 of the present embodiment determines, based on the information acquired from the external device 9A, which of the plurality of structure descriptors 201 included in the first model 20 to include in the second model 21. The information acquired from the external device 9A includes the types of elements to be analyzed in the second model 21. For example, when the user inputs from the external device 9A the elements included in the atomic structure for which analysis is desired, the information processing apparatus 1 of the present embodiment can configure the structure descriptor 201 of the second model 21 according to the user's needs.
[0086] Also, the first model 20 and the second model 21 of the present embodiment are models of the same type. Therefore, according to the information processing apparatus 1 of the present embodiment, the parameters of the ML model 202 of the first model 20 generated by pre-training can be inherited by the second model 21.
[0087] Also, the number of structure descriptors 201 included in the second model 21 of the present embodiment is smaller than the number of structure descriptors 201 included in the first model 20. Since the processing amount is reduced as the number of structure descriptors 201 is smaller, the information processing apparatus 1 of the present embodiment can improve the speed of the simulation processing of the second model 21.
[0088] Also, the first model 20 and the second model 21 of the present embodiment are MTP models. Since the MTP model does not have processing corresponding to a plurality of layers like other neural networks, for example, GNN (Graph Neural Network), the processing load is low. Therefore, the information processing apparatus 1 of the present embodiment can speed up the processing by configuring the first model 20 and the second model 21 with MTP models.
[0089] In addition, the analysis result of the atomic structure output from the second model 21 of the present embodiment includes at least information regarding the energy or force of the atomic structure. Analyzing the atomic structure has a high processing load, and there is a need to speed up the processing in order to analyze a larger number of atoms. The information processing apparatus 1 of the present embodiment can provide an MD simulation of the atomic structure using the fine-tuned second model 21 in response to such a need.
[0090] In addition, the information processing apparatus 1 of the present embodiment compares the analysis result output from the second model 21 with verification data generated by generation means different from the first model 20 and the second model 21, and outputs the comparison result as a numerical value. Therefore, according to the information processing apparatus 1 of the present embodiment, it is possible to objectively determine whether or not the second model 21 has been sufficiently fine-tuned. For this reason, it is possible to take measures such as additional training for the second model 21 according to the evaluation result.
[0091] In addition, the verification data of the present embodiment is the analysis result by a third model that generated at least one of the first data and the second data, that is, a trained NNP trained with DFT as correct data. Therefore, according to the information processing apparatus 1 of the present embodiment, the time required for generating the verification data can be shortened by using the trained NNP.
[0092] (Modification example) In the above-described embodiments, the processes described as functions of the information processing apparatus 1 may be executed by the same apparatus or by different apparatuses. For example, the training of the first model 20, the training of the second model 21, the generation of training data for pre-training, the generation of training data for fine-tuning, the generation of verification data, and the inference using the trained second model 21 may be executed by the same apparatus as in the above-described embodiments, or may be executed by different apparatuses respectively. Specifically, in the above-described embodiments, the information processing apparatus 1 has performed the generation and pre-training of the training data for pre-training the first model 20, but these processes do not necessarily have to be performed within the information processing apparatus 1. For example, the information processing apparatus 1 may acquire and use the pre-trained first model 20 from another information processing apparatus.
[0093] Also, in the above-described embodiments, as an example, it has been described that the NNP trained with the DFT as the correct answer data generates the training data for pre-training and the verification data, but the training data for pre-training and the verification data may be generated by the DFT.
[0094] Also, in the above-described embodiments, the verification data used for the evaluation of the second model 21 is assumed to be generated by the trained NNP trained with the DFT as the correct answer data, but the generation method of the verification data is not limited to this. For example, the information processing apparatus 1 may acquire data such as the energy, force, stress, density, or elastic value of the atomic structure obtained by experiments and use it as verification data. By using the data obtained by experiments as verification data, it is possible to eliminate the influence of the accuracy of the model for generating the verification data on the evaluation results.
[0095] In the above-described embodiment, the various parameters of the ML model 202 of the second model 21 before training are the same as those of the ML model 202 of the trained first model 20. However, the parameters of the ML model 202 of the second model 21 before training and the ML model 202 of the trained first model 20 may be different. In other words, the editing unit 314 of the information processing apparatus 1 may change not only the structure descriptor 201 of the pre-trained first model 20 but also the parameters of the ML model 202.
[0096] In the above-described embodiment, the ML model 202 included in the first model 20 and the second model 21 is a regression model based on a polynomial basis function. However, the ML model 202 is not limited to this, and any model can be adopted. For example, the ML model 202 may be an NN (Neural Network).
[0097] In the above-described embodiment, the first model 20 and the second model 21 are configured as MTP models. However, other models may be adopted. For example, a GNN can be adopted as the first model 20 and the second model 21. In this case, the information processing apparatus 1 may execute atomic type embedding before the GNN. Also, as the first model 20 and the second model 21, a neural network potential that represents the environment of each atom by a symmetry function, for example, a Behler-Parrinello type neural network, may be adopted, or a Gaussian approximation potential may be adopted.
[0098] In the above-described embodiment, the trained second model 21 is stored in the auxiliary storage device 35 of the information processing apparatus 1. However, the trained second model 21 may be stored in an information processing apparatus owned by the user. For example, when the external device 9A is the user's computer that uses the trained second model 21, the information processing apparatus 1 may transmit the trained second model 21 to the external device 9A. Note that it is possible to execute the inference (analysis process) using the trained second model 21 on the information processing apparatus 1 side at a higher speed than when the external device 9A executes the inference using the trained second model 21.
[0099] In the above-described embodiment, the case where one second model 21 is generated from one trained first model 20 has been described. However, a plurality of different second models 21 may be generated from one trained first model 20. The plurality of different second models 21 may be, for example, those having different atomic structures or simulation types to be analyzed.
[0100] The differences between the plurality of different second models may be differences in the fine-tuning data, differences in the structure descriptors 201, or differences in the ML models 202.
[0101] For example, the editing unit 314 of the information processing apparatus 1 may generate a plurality of second models 21 having different structure descriptors 201 for each atomic structure to be analyzed.
[0102] Also, a plurality of different trained second models 21 may be generated from the same second model before training. For example, the second training data generation unit 315 may generate a plurality of pieces of fine-tuning data (second data) including different atomic structures. In this case, the second training unit 316 may generate a plurality of different trained second models 21 using the plurality of pieces of fine-tuning data including different atomic structures.
[0103] In addition, when a plurality of second models 21 after training are generated for each atomic structure to be analyzed, when the inference unit 318 executes the inference process, the model name for specifying the trained second model 21 to be executed is also included in the input data of the simulation process. The model name for specifying the trained second model 21 to be executed may be input by the user from, for example, the external device 9A or the external device 9B.
[0104] A part or all of each device (information processing device 1, external devices 9A and 9B) in the above-described embodiments may be configured by hardware, or may be configured by information processing of software (program) executed by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like. When configured by information processing of software, software that realizes at least some functions of each device in the above-described embodiments is stored in a non-temporary storage medium (non-temporary computer-readable medium) such as a CD-ROM (Compact Disc-Read Only Memory) or a USB (Universal Serial Bus) memory, and the computer may read it to execute the information processing of the software. Further, the software may be downloaded via a communication network. Furthermore, by implementing all or part of the software processing in a circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), the information processing by the software may be executed by hardware.
[0105] The storage medium for storing the software may be a removable one such as an optical disk, or a fixed storage medium such as a hard disk or a memory. Further, the storage medium may be provided inside the computer (main storage device, auxiliary storage device, etc.) or outside the computer.
[0106] In this specification (including the claims), when expressions such as "at least one (one) of a, b, and c" or "at least one (one) of a, b, or c" (including similar expressions) are used, it includes any of a, b, c, a - b, a - c, b - c, or a - b - c. Also, for any element, multiple instances may be included, such as a - a, a - b - b, a - a - b - b - c - c, etc. Furthermore, it also includes adding other elements outside the enumerated elements (a, b, and c), such as having d like a - b - c - d.
[0107] In this specification (including the claims), when expressions such as "using data as input / based on data / in accordance with data / in response to data" (including similar expressions) are used, unless otherwise specified, it includes cases where various data themselves are used as input, and cases where something obtained by performing some processing on various data (for example, data with noise added, normalized data, intermediate representations of various data, etc.) is used as input. Also, when it is described that some result is obtained "based on data / in accordance with data / in response to data", it includes cases where the result is obtained based only on the said data, and may also include cases where the result is obtained under the influence of other data, factors, conditions, and / or states other than the said data. Further, when it is described that "data is output", unless otherwise specified, it includes cases where various data themselves are used as output, and cases where something obtained by performing some processing on various data (for example, data with noise added, normalized data, intermediate representations of various data, etc.) is output.
[0108] In this specification (including the claims), when the terms "connected" and "coupled" are used, they are intended as non-limiting terms that include any of direct connection / coupling, indirect connection / coupling, electrical connection / coupling, communicative connection / coupling, operative connection / coupling, physical connection / coupling, etc. The terms should be appropriately interpreted according to the context in which they are used, but connection / coupling forms that are not intentionally or naturally excluded should be interpreted non-limitingly as being included in the terms.
[0109] In this specification (including the claims), when the expression "A configured to B" is used, it may include that the physical structure of element A has a configuration capable of performing operation B, and that the permanent or temporary setting / configuration of element A is set to actually perform operation B. For example, when element A is a general-purpose processor, it suffices that the processor has a hardware configuration capable of performing operation B and is set to actually perform operation B by a permanent or temporary program (instruction) setting. Also, when element A is a dedicated processor, a dedicated arithmetic circuit, etc., it suffices that the circuit structure, etc. of the processor is implemented to actually perform operation B regardless of whether control instructions and data are actually attached.
[0110] In this specification (including the claims), when terms meaning containment or possession (such as "comprising / including", "having", etc.) are used, they are intended as open-ended terms, including cases where the object contains or possesses things other than the object indicated by the object of such terms. When the object of these terms meaning containment or possession does not specify a quantity or is an expression suggesting a singular number (an expression with "a" or "an" as the article), such expression should be construed as not being limited to a specific number.
[0111] In this specification (including the claims), even if expressions such as "one or more", "at least one", etc. are used in one place and expressions that do not specify a quantity or suggest a singular number (expressions with "a" or "an" as the article) are used in other places, the latter expression is not intended to mean "one". Generally, expressions that do not specify a quantity or suggest a singular number (expressions with "a" or "an" as the article) should be construed as not necessarily being limited to a specific number.
[0112] In this specification, when it is described that a specific effect (advantage / result) is obtained for a specific configuration of a certain embodiment, unless there are other reasons, it should be understood that the same effect can also be obtained for one or more other embodiments having the same configuration. However, the presence or absence of such effect generally depends on various factors, conditions, and / or states, and it should be understood that the effect is not necessarily obtained by such configuration. The effect is only obtained by the configuration described in the embodiment when various factors, conditions, and / or states are satisfied, and in the invention related to the claim defining such configuration or a similar configuration, the effect is not necessarily obtained.
[0113] In this specification (including the claims), when terms such as "maximize / maximization" are used, it includes obtaining a global maximum value, obtaining an approximation of the global maximum value, obtaining a local maximum value, and obtaining an approximation of the local maximum value, and should be appropriately interpreted according to the context in which the term is used. Also included is obtaining an approximation of these maximum values probabilistically or heuristically. Similarly, when terms such as "minimize / minimization" are used, it includes obtaining a global minimum value, obtaining an approximation of the global minimum value, obtaining a local minimum value, and obtaining an approximation of the local minimum value, and should be appropriately interpreted according to the context in which the term is used. Also included is obtaining an approximation of these minimum values probabilistically or heuristically. Similarly, when terms such as "optimize / optimization" are used, it includes obtaining a global optimum value, obtaining an approximation of the global optimum value, obtaining a local optimum value, and obtaining an approximation of the local optimum value, and should be appropriately interpreted according to the context in which the term is used. Also included is obtaining an approximation of these optimum values probabilistically or heuristically.
[0114] In this specification (including the claims), when a plurality of hardware components perform a predetermined process, each hardware component may cooperate to perform the predetermined process, or some of the hardware components may perform all of the predetermined process. Also, some of the hardware components may perform a part of the predetermined process and other hardware components may perform the remainder of the predetermined process. In this specification (including the claims), when expressions such as "one or more hardware components perform a first process and the one or more hardware components perform a second process" (including similar expressions) are used, the hardware components performing the first process and the hardware components performing the second process may be the same or different. That is, it is sufficient that the hardware components performing the first process and the hardware components performing the second process are included in the one or more hardware components. Note that the hardware components may include electronic circuits, devices including electronic circuits, and the like.
[0115] In this specification (including the claims), when a plurality of storage devices (memories) store data, each of the plurality of storage devices may store only a part of the data or may store all of the data. Also, a configuration in which some of the plurality of storage devices store data may be included.
[0116] As described above in detail for the embodiments of the present disclosure, the present disclosure is not limited to the individual embodiments described above. Various additions, changes, replacements, partial deletions, etc. are possible without departing from the conceptual ideas and spirit of the present invention derived from the content defined in the claims and their equivalents. For example, in the embodiments described above, when numerical values or mathematical formulas are used for the description, these are shown for illustrative purposes and do not limit the scope of the present disclosure. Also, the order of each operation shown in the embodiments is also illustrative and does not limit the scope of the present disclosure.
[0117] Regarding the above embodiments, the following supplementary notes are disclosed as one aspect and optional features of the invention. (Supplementary Note 1) at least one processor; at least one memory, and The at least one processor uses at least some parameters of a first model trained using first data to generate a second model different from the first model, wherein the second model is a model that outputs an analysis result of an atomic structure when the atomic structure is input, and the at least some parameters are determined based on the type of atom to be analyzed by the second model, an information processing apparatus. (Appendix 2) The at least one processor trains the second model using second data. (Appendix 3) The second data is data generated using a third model different from the first model. (Appendix 4) The first data is data generated using the third model. (Appendix 5) The third model is a neural network potential. (Appendix 6) The first data and the second data include an atomic structure and an analysis result of the atomic structure, and the type of atomic structure included in the first data is more than the type of atomic structure included in the second data. (Appendix 7) The at least one processor acquires information on the type of atom to be analyzed by the second model from another information processing apparatus. (Appendix 8) The first model and the second model are models of the same type. (Appendix 9) The first model and the second model each include a plurality of parameters, and the number of parameters included in the second model is less than the number of parameters included in the first model. (Appendix 10) The type of the model is the MTP (Moment Tensor Potential) model. (Appendix 11) The types of atoms that can be analyzed by the first model are more than those of the atoms that can be analyzed by the second model. (Appendix 12) The at least one processor compares the analysis result output from the second model with verification data generated by generation means different from the first model and the second model, and outputs the comparison result as a numerical value. (Appendix 13) The verification data is the analysis result by a third model that generates at least one of the first data and the second data. (Appendix 14) The verification data is data obtained by experiments. (Appendix 15) The analysis result of the atomic structure includes at least information regarding the energy or force of the atomic structure. (Appendix 16) A model generation method for generating the second model using the above information processing apparatus. Model generation method. (Appendix 17) An information processing system including a first information processing apparatus and a second information processing apparatus, wherein the second information processing apparatus sends information including types of atoms to be analyzed by the second model to the first information processing apparatus, and the first information processing apparatus generates a second model using at least some parameters of a first model trained using first data based on the information received from the second information processing apparatus, wherein the second model is a model that outputs an analysis result of the atomic structure when the atomic structure is input. Information processing system. (Appendix 18) The first information processing apparatus trains the second model using second data. (Appendix 19) The second data is data generated using a third model different from the first model. (Appendix 20) The first data and the second data include an atomic structure and an analysis result of the atomic structure. The types of atomic structures included in the first data are more than the types of atomic structures included in the second data. (Appendix 21) The first model and the second model each include a plurality of parameters. The number of parameters included in the second model is less than the number of parameters included in the first model. (Appendix 22) The analysis result of the atomic structure includes at least information regarding the energy or force of the atomic structure. (Appendix 23) A method executed by at least one processor, using at least some parameters of a first model trained using first data to generate a second model different from the first model, wherein the second model is a model that outputs an analysis result of an atomic structure when the atomic structure is input, and the some parameters are determined based on the types of atoms to be analyzed by the second model. Method.
Explanation of Reference Numerals
[0118] 1 Information processing apparatus 5 Communication network 9A, 9B External device 20 First model 21 Second model 30 Computer 31 Processor 33 Main storage device 35 Auxiliary storage device 37 Network interface 39 Device interface 41 buses 201 Structure descriptor 202 ML model 311 First training data generation unit 312 First training unit 313 Acquisition unit 314 Editing unit 315 Second training data generation unit 316 Second training unit 317 Evaluation unit 318 Inference unit S Information processing system
Claims
1. At least one processor and, At least one memory, and is provided with, The at least one processor, Using at least some of the parameters of the first model trained using the first data, generates a second model different from the first model, The second model is a model that outputs an analysis result of the atomic structure when the atomic structure is input, The at least some of the parameters are determined based on the type of atom to be analyzed by the second model, An information processing apparatus.
2. The at least one processor, Trains the second model using second data, The information processing apparatus according to claim 1.
3. The second data is data generated using a third model different from the first model, The information processing apparatus according to claim 2.
4. The first data is data generated using the third model, The information processing apparatus according to claim 3.
5. The third model is a neural network potential, The information processing apparatus according to claim 3.
6. The first data and the second data include an atomic structure and an analysis result of the atomic structure, The type of atomic structure included in the first data is more than the type of atomic structure included in the second data, The information processing apparatus according to claim 2.
7. The at least one processor, Obtains information on the type of atom to be analyzed by the second model from another information processing apparatus, The information processing apparatus according to claim 1.
8. The first model and the second model are models of the same type, The information processing apparatus according to claim 1.
9. The first model and the second model each include a plurality of parameters, The number of parameters included in the second model is less than the number of parameters included in the first model, The information processing apparatus according to claim 8.
10. The type of the model is an MTP (Moment Tensor Potential) model, The information processing apparatus according to claim 8.
11. The types of atoms that the first model can analyze are more than the types of atoms that the second model can analyze, The information processing apparatus according to claim 1.
12. The at least one processor, Compare the analysis result output from the second model with verification data generated by generation means different from the first model and the second model, and output the comparison result as a numerical value. The information processing apparatus according to claim 2.
13. The verification data is an analysis result by a third model that generated at least one of the first data and the second data. The information processing apparatus according to claim 12.
14. The verification data is data obtained by experiments. The information processing apparatus according to claim 12.
15. The analysis result of the atomic structure includes information regarding at least the energy or force of the atomic structure. The information processing apparatus according to any one of claims 1 to 14.
16. Generate the second model using the information processing apparatus according to any one of claims 1 to 14. Model generation method.
17. An information processing system including a first information processing apparatus and a second information processing apparatus, wherein the second information processing apparatus transmits information including the types of atoms to be analyzed by the second model to the first information processing apparatus, and the first information processing apparatus generates a second model using at least some parameters of a first model trained using first data based on the information received from the second information processing apparatus, wherein the second model is a model that outputs an analysis result of the atomic structure when the atomic structure is input. Information processing system.
18. The first information processing apparatus trains the second model using second data. The information processing system according to claim 17.
19. The second data is data generated using a third model different from the first model. The information processing system according to claim 18.
20. The first data and the second data include an atomic structure and an analysis result of the atomic structure, and the types of atomic structures included in the first data are more than the types of atomic structures included in the second data. The information processing system according to claim 18.
21. The first model and the second model each include a plurality of parameters, and the number of parameters included in the second model is less than the number of parameters included in the first model. The information processing system according to claim 17.
22. The analysis result of the atomic structure includes at least information regarding the energy or force of the atomic structure. The information processing system according to any one of claims 17 to 21. **Claim 23** A method executed by at least one processor, generating a second model different from the first model using at least some parameters of the first model trained using first data, wherein the second model is a model that outputs an analysis result of the atomic structure when the atomic structure is input, and the some parameters are determined based on the type of atoms to be analyzed by the second model. Method.