Information processing device, trained model generation device, information processing method, trained model generation method, information processing program, and trained model generation program

The Neural Structure Field (NeSF) approach addresses the challenge of representing and decoding crystal structures by using continuous vector fields, achieving improved accuracy and efficiency in reconstructing complex crystal structures compared to existing methods.

JP2025090756APending Publication Date: 2025-06-17OMRON CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025039567
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-21
Filing Date
2025-03-12
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing methods struggle to accurately represent and decode crystal structures using neural networks, particularly due to the challenge of handling variable numbers of atoms and maintaining spatial resolution without excessive computational complexity.

Method used

The proposed Neural Structure Field (NeSF) approach represents crystal structures as continuous vector fields, using position and type fields to implicitly represent atomic positions and types, thereby overcoming the limitations of grid-based discrete space representations.

Benefits of technology

NeSF achieves superior performance in reconstructing crystal structures compared to existing grid-based methods, with significant improvements in position and type error rates, and can effectively represent complex crystal structures without the trade-off between spatial resolution and computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025090756000001_ABST
    Figure 2025090756000001_ABST
Patent Text Reader

Abstract

To enable representation of a field showing a structure of a substance represented by an atomic point cloud.SOLUTION: A field showing a structure of a substance represented by an atomic point cloud is represented using a neural network model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, a learned model generation apparatus, an information processing method, a learned model generation method, an information processing program, and a learned model generation program.

Background Art

[0002] Conventionally, techniques related to an autoencoder-based generative deep representation learning pipeline for 3D crystal structures have been known (see, for example, Non-Patent Document 1).

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to represent a field representing the structure of a substance composed of atomic point groups.

Means for Solving the Problems

[0005] To achieve the above object, an information processing apparatus according to the present disclosure is an information processing apparatus including a processing unit that uses a neural network model to represent a field representing the structure of a substance composed of an atomic point group.

[0006] Further, an information processing method according to the present disclosure is an information processing method in which a computer executes a process of using a neural network model to represent a field representing the structure of a substance composed of an atomic point group.

[0007] Further, an information processing program according to the present disclosure is an information processing program for causing a computer to execute a process of using a neural network model to represent a field representing the structure of a substance composed of an atomic point group.

Advantages of the Invention

[0008] According to the information processing apparatus, information processing method, and information processing program of the present disclosure, it is possible to represent a field representing the structure of a substance composed of an atomic point group.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Mode for Carrying Out the Invention

[0010] Hereinafter, an example of an embodiment of the present disclosure will be described with reference to the drawings. In this embodiment, an information processing apparatus according to the present disclosure will be described as an example. In each of the drawings, the same or equivalent components and parts are given the same reference numerals. Also, the dimensions and ratios in the drawings are exaggerated for the convenience of explanation and may differ from the actual ratios.

[0011] <Overview> Crystals, and among them, the inverse design method of materials, can contribute to a next-generation approach for exploring materials with desired properties without relying on "luck" or "serendipity". In this embodiment, as an accurate and practical approach for expressing crystal structures using neural networks, a Neural Structure Field (NeSF) is proposed. Inspired by the concept of vector fields in physics and implicit neural representations in computer vision, the proposed NeSF regards the crystal structure not as a discrete set of atoms but as a continuous field. Different from existing grid-based discrete space representations, NeSF overcomes the trade-off between spatial resolution and computational complexity and can represent any crystal structure. In this embodiment, to evaluate NeSF, an autoencoder for crystal structures that can recover various crystal structures such as perovskite-structured materials and cuprate superconductors is proposed. Extensive quantitative results show the excellent performance of NeSF compared with existing grid-based approaches. Note that the numbers in [] below represent the reference numbers shown in FIGS. 21 to 23.

[0012] <1. Introduction> In the basic paradigm of materials science, assuming that material properties are closely related to the crystal structure, the relationship between structure and properties is considered. Therefore, in the conventional approach in materials science, theoretical and experimental analyses of the relationship between the structure and properties of materials are carried out to develop new materials with excellent properties. New materials are being explored [1,2]. However, these traditional approaches rely on labor-intensive human analysis and even "serendipity". To automate or assist in the analysis and development of materials, data-driven approaches have been actively studied in materials science, and the field of materials informatics (MI) has been established [3-6]. Different from traditional approaches based on the estimation of physical laws, MI aims to reveal material knowledge (such as laws governing the relationship between structure and properties) from the collected material datasets through statistical and machine learning (ML) methods. In recent years, due to the technological progress of ML and the emergence of large-scale material databases, MI has been developing rapidly [6,7]. Therefore, powerful neural network-based ML methods are becoming important elements in MI research [6]. Applications of MI include the prediction of material properties from material property data such as crystal structure [6] and composition [8-11], the automatic analysis of experimental data [12-14], and natural language processing for knowledge retrieval from scientific literature

[15] .

[0013] Many MI studies have focused on the prediction of properties of specific materials, such as band gap, Seebeck coefficient, elastic modulus, etc. [16-18]. Considering this type of task as the discovery of the relationship between structure and properties among materials, there is another important type of task, that is, the discovery of the relationship between properties and structure that constitutes the inverse problem [19-23].

[0014] Despite the potential usefulness of this inverse approach in materials development, there have been few studies addressing this [19-23] or the underlying problems [24-26], that is, the estimation of crystal structure under specific conditions. For MI and ML, a decisive difference occurs depending on whether the crystal structure is input or output (that is, whether the crystal structure is encoded or decoded in MI and ML terms). The encoding of crystal structure is properly established using graph neural networks [10,16,17,27], but there remain technical bottlenecks in the decoding of crystal structure. In this investigation, the bottleneck was addressed.

[0015] The crystal structure of an inorganic material is a regular and periodic arrangement of atoms in three-dimensional (3D) space. This arrangement is usually described by the 3D positions and types of atoms within the unit cell, and the lattice constants that define the translation of the unit cell in 3D space. The atoms within the unit cell do not have a definite order, and the number can vary from one to hundreds. Since ML models, including neural networks, generally accept tensors that are fixed-dimensional and consistently ordered for processing, it is not straightforward to handle crystal structures with ML models

[16] , and it is even more difficult to determine crystal structures through the models.

[0016] In this embodiment, a general representation of the crystal structure is proposed that enables a neural network to decode or determine such a structure. The key concept underlying our approach is shown in Fig. 1. Here, the crystal structure is represented not as a discrete set of atoms, but as a continuous vector field associated with 3D space. We call our approach the Neural Structure Field (NeSF). NeSF uses two types of vector fields, a position field and a type field, to implicitly represent the positions and types of atoms within the unit cell of the crystal structure, respectively.

[0017] To explain the concept of NeSF, assume that the target material information is given as a fixed-dimensional vector z, and consider the problem of restoring the crystal structure from z. The input z can specify, for example, information about the crystal structure of the material or some desired criteria for the generated material. In NeSF, instead of directly outputting the crystal structure as f(z) using a neural network f, the neural network f is used as an implicit function to indirectly represent the crystal structure embedded in z. Specifically, f is treated as a vector field over 3D Cartesian coordinates p, conditioned on the target material information z.

[0018]

Number

[0019] In the crystal structure of interest, the position field is trained to output a 3D vector that points to the atomic position a closest to the query point p. Thus, the output s is expected to be a - p. If the position field is ideally trained, the position a of the closest atom at any query point p can be obtained as p + f(p,z).

[0020] Mathematically, the position field can be interpreted as the gradient vector field -φ(p) of a scalar potential

[0021]

Number

[0022] . The scalar potential is expressed by the square of the distance between the query point p and the position {a i} of the closest atom. Similarly, the type field is trained to output a categorical probability distribution indicating the type of the closest atom. Thus, the output dimension of the type field is the number of candidate atomic species.

[0023] The proposed NeSF is inspired by the concept of vector fields in classical physics and implicit neural representations in computer vision [30-34]. To address some representation problems in 3D computer vision applications, such as 3D object shape estimation [31-33] and free-viewpoint image synthesis [30,34], implicit neural representations have been recently proposed. In 3D shape estimation, directly outputting a 3D mesh or point cloud to a neural network causes representation problems similar to those occurring in crystal structures. To overcome these problems, the signed distance function (SDF) is utilized in DeepSDF

[31] , where the neural network f(p) models the 3D shape by indicating whether the query point p is outside or inside the object volume with positive or negative signs and outputs each scalar. NeSF follows the basic idea of implicit neural representations and further extends it to the estimation of crystal structures described by atomic positions and species. The accurate description of atomic positions in crystal structures is very important in materials science. Therefore, NeSF outputs a vector pointing to the nearest atom, representing the atomic positions more directly than existing implicit neural representations of 3D geometry.

[0024] Our idea of representing 3D geometric crystal structures as continuous vector fields has been considered implicitly and partially in recent MI research [19, 20, 22, 24-26] using grid-based discretization (i.e., voxelization), but there has been no explicit consideration as a discretized vector field. In these studies, the 3D space within a unit cell is discretized into voxels, and an electron density is assigned to each voxel. The electron density basically represents the presence or absence of atoms around the voxel. However, compared to 1D (such as audio signals) or 2D (such as images) data, the discretization of 3D data is quite troubled by the trade-off between spatial resolution and computational complexity in terms of both computation time and memory space. For example, the ICSG3D method

[24] represents crystal structures using 32×32×32 voxels and estimates them using 3D convolutional neural networks (CNNs). Since voxel-based 3D CNNs consume a large amount of computation and memory, a resolution of 32×32×32 voxels is approximately the limit for training voxel-based models on standard computing systems. On the other hand, existing crystal structures contain dozens or more atoms in the unit cell, or the unit cell is elongated or distorted. Therefore, a sufficiently high resolution is required to accurately represent diverse crystal structures in voxels. Furthermore, voxel-based models can only indirectly provide atomic positions in representations such as the peaks of scalar fields of electron density discretized in voxel space. The proposed NeSF overcomes the limitations of voxelization. In NeSF, there is basically no trade-off between spatial resolution and required memory. Theoretically, NeSF can achieve infinitely high spatial resolution using a compact (high memory and parameter efficiency) neural network instead of costly 3D CNNs. Furthermore, NeSF can effectively represent any crystal structure including elongated or distorted unit cells. Additionally, NeSF can directly provide the Cartesian coordinates of atomic positions rather than the peaks of scalar fields. The proposed NeSF is believed to break through the technical bottleneck of the MI approach for crystal structure estimation and contribute to the progress of MI research in this direction.

[0025] In the following section, the proposed NeSF is described in detail, and its representational power for various crystal structures is shown through numerical experiments. In particular, NeSF has been successful in recovering various crystal structures, from relatively basic structures of perovskite materials to complex structures of cuprate superconductors. The results of extensive quantitative evaluations indicate that NeSF is superior to the voxelization approach of ICSG3D

[24] .

[0026] <2 Results and Discussion> First, the procedure for estimating the crystal structure using NeSF and training NeSF is described. Next, as an application of NeSF, an autoencoder for crystal structures is presented. In this autoencoder, the crystal structure is embedded into a vector z (referred to as the latent vector) via the encoder, and NeSF functions as the decoder to reconstruct the input crystal structure from z. The performance of the NeSF-based autoencoder is quantitatively analyzed by evaluating the reconstruction accuracy that exceeds the voxelization-based ICSG3D baseline. The space of the learned vector z is qualitatively analyzed, and this analysis by the proposed autoencoder shows that the learned space reflects the similarity between crystal structures rather than simply embedding crystal structures randomly.

[0027] <2.1 Crystal Structure by NeSF> When the material information of the estimation target is given as the vector z, estimating the crystal structure from z amounts to estimating the positions and species of atoms within the unit cell along with the lattice constants. The lattice constants are modeled as lengths a, b, and c and angles α, β, and γ, and are estimated by a simple multi-layer perceptron (MLP) using the input z. On the other hand, the atomic positions and atomic species are, as described in the previous section, the position field f of NeSF p and the species field f sThey are estimated respectively by. These fields can also be implemented as simple MLPs, each taking the query position p and the vector z as inputs and predicting the field values (i.e., 3D pointing vectors or categorical probability distributions). An overview of the NeSF network architecture is shown on the right side of Figure 1(b).

[0028] When the vector z is given from the encoder, the estimation of atomic positions and species using NeSF is shown in Figure 2 and summarized in five steps.

[0029] 1. Initialize the particles. First, estimate the lattice constant via an MLP. Next, at the 3D grid points within the bounding box, periodically expand the initial query points {p 0 i} called particles. The bounding box is common to each dataset and is specified to roughly enclose the atoms of all training samples.

[0030] 2. Move the particles. Update the position of each particle p t+1 i = p t i + f p (p t i , z) according to the position field, and obtain {p t i} as candidates for atomic positions. Since the position field is expected to point to the nearest atomic position, the particles move towards the nearest atom through this process. i}

[0031] 3. Score the particles. Score each particle p i and filter out outliers. The norm ||f p (p i , z)|| of the output of the position field indicates the estimated distance from p i to the nearest atom, so score each particle p i by ||f p (pi Score with ||(x, y, z), and if the score exceeds a specific threshold (set to 0.9 Å in this case), discard the particle.

[0032] 4. Detect atoms. By this point, the particles are expected to form clusters around the atoms. Therefore, apply a simple clustering algorithm to detect each cluster position as the atomic position and determine the number of atoms in the crystal structure. Here, use the clustering algorithm well-known in object detection, Non-Maximum Suppression. Specifically, initialize the list of candidate particles as B c = {p i}, and initialize another list of accepted particles as B a = {}. 1) Select the particle with the lowest score (estimated to be closest to the atom) from B c and move it to B a . 2) Delete the particles within the spherical region around the selected particle from B c (the sphere radius was set to 0.5 Å in this study). By repeating these steps until B c becomes empty, obtain the atomic positions {a a} stored in B i .

[0033] 5. Estimate the type. Finally, use the type field f s (p, z) to estimate the atomic type of each atomic position a i . Instead of directly using a i as a query for robust estimation against the error of a i , intensively spread new particles around each a i as a query to the type field. Therefore, multiple probability distributions are obtained, each predicting the type of atom a i . The atomic type with the highest frequency among them is selected as the final estimated value.

[0034] <2.2 Training of NeSF> The training of NeSF is much simpler than the above-mentioned estimation algorithm. The 3D query points {p s i} within the unit cell are randomly sampled, and the loss values of the field outputs at these points are calculated. Therefore, f p (p s i ,z) and f s (p s i ,z) are monitored to indicate the positions and types of the nearest atoms respectively. However, due to the practical limitation of memory usage, it is not possible to sample the query points densely. Therefore, a sampling strategy for query points is required for training. Existing implicit neural representations for 3D shape estimation such as DeepSDF

[31] sample training query points near the surface. Curriculum DeepSDF

[35] further introduces curriculum learning, and the sampling density is enhanced near the surface as the training progresses.

[0035] To explore a desirable sampling strategy for training the position field and the type field, their dynamics are examined with the proposed algorithm. 1) The particles repeatedly move within the position field towards the nearest atom. Therefore, the position field must be accurate enough everywhere for the particles to flow to the destination and very accurate near the atom. 2) The type field is queried only around the atoms. Therefore, it does not need to be accurate everywhere but must be robust to errors in the estimated atom positions.

[0036] To meet the above requirements, two sampling methods are introduced to train the position field and the type field. 1) Global grid sampling: In this method, 3D grid points that uniformly cover the entire unit cell are considered, and the points are sampled with perturbations following a Gaussian distribution. 2) Local grid sampling: In this method, local 3D grid points centered at each atomic position are considered, and the points are sampled with perturbations following a Gaussian distribution.

[0037] To train the position field, both sampling methods are combined. Thus, query points are sampled uniformly over the entire unit cell and densely around the atoms. To train the type field, local grid sampling is used to concentrate the training query points in the vicinity of the atoms.

[0038] <2.3 Crystal Structure Autoencoder> To demonstrate and evaluate the representational power of NeSF, a crystal structure autoencoder is proposed. Similar to other general autoencoders, the proposed NeSF-based autoencoder consists of an encoder and a decoder. The encoder is a neural network that transforms the input crystal structure (i.e., the positions and types of atoms within the unit cell, and the lattice constants) into an abstract latent vector z. The decoder using NeSF reconstructs the input crystal structure from the latent vector z. Autoencoders are usually used to learn the latent vector representation of data through self-supervised learning. In this learning, the input data can monitor the learning through the reconstruction loss.

[0039] We focus on the decoding of crystal structures, which are studied under MI. Since crystal structures are essentially a collection of atoms, their encoding needs to handle variable numbers of atoms that are invariant to permutations. In ML, such encoders are generally called set functions [28,29]. Among them, the family of graph neural networks [10,16,17,27] functions as a general crystal structure encoder. However, these networks implicitly represent the positions of atoms as edges and encode the distances between atoms while discarding the exact coordinates. This distance-based graph representation is key to ensuring invariance to coordinate systems, but the loss of information in the input may inadvertently hinder the performance of reconstruction. Therefore, we adopt the basic encoder architectures of PointNet

[28] and DeepSets

[29] . This architecture not only represents the simplest type of set function-based network but also preserves the information of the input crystal structure, making it suitable for evaluating the performance of the NeSF decoder. The proposed autoencoder architecture is detailed in Section 4.2 of Figure 28.

[0040] <Training and Evaluation Procedures> We trained and evaluated an autoencoder and ICSG3D24 (baseline) on three material datasets: ICSG3D, limited cell size 6 Å (LCS6 Å), and YBCO-like datasets. These datasets collect the crystal structures of materials from the Materials Project and are designed to have various levels of difficulty. The ICSG3D dataset

[24] is a material collection consisting of three datasets containing 7897 materials with limited crystal systems (cubic) and prototypes (i.e., AB, ABX2, and ABX3). The LCS6 Å dataset is composed of 6005 materials with unit cell sizes of 6 Å or less in the x, y, and z directions, with no restrictions on crystal systems and prototypes. The YBCO-like dataset is composed of 100 materials with narrow unit cells along the c-axis. These structures usually include the structures of yttrium barium copper oxide (YBCO) superconductors. Due to the complex structures and relatively few samples, the YBCO-like dataset is the most difficult among the three evaluated datasets. For details of these datasets, please refer to Section 4.1 of Figure 27.

[0041] For training and evaluation, each dataset was randomly split into training (90.25%), validation (4.75%), and test (5%) sets. The training set was only used for training the ML model. The validation set was used to pre-validate the trained ML model, and the test set was used to calculate the final evaluation score after training, validation, and hyperparameter tuning. The hyperparameters were tuned based on the validation score of the LCS6Å dataset. To reduce performance variations due to validation score randomness (such as randomness in the initialization of network weights), training and evaluation were repeated 10 times with different random seeds, and the performance was evaluated using the average and standard deviation of the scores. Since datasets like YBCO only had 100 samples, they were processed in a slightly different way from the other two evaluated datasets. To reduce performance variations due to data splitting, 20-fold cross-validation was employed for datasets like YBCO, but it was evaluated only once instead of 10 times. The iterative training of the neural network was performed using stochastic gradient descent with Adam

[36] as the optimizer. The detailed training procedure, including the definition of the loss function, is provided in Section 4.3 of Figure 28.

[0042] The reconstruction performance was measured with respect to errors in the number, position, and type of atoms. The error in the number of atoms is the proportion of materials for which the number of atoms in the unit cell is not correctly estimated. The position error is the average error in the reconstructed atom positions. Depending on the denominator of the metric, the position error was evaluated in two ways. Using the actual metric, the average position error at the actual atom sites of the crystal structure was evaluated by calculating the shortest distance to the estimated atom sites. In contrast, the detected metric was used to evaluate the error at the estimated sites by calculating the shortest distance to the actual atom sites. The actual metric is more sensitive to errors related to underestimation of the number of atoms, while the detected metric is more sensitive to errors related to overestimation. The type error is the average proportion of atoms for which the type is not correctly estimated. Similar to the position error, the type error was evaluated using the actual and the detected measurement criteria. Lower values of these metrics indicate better performance.

[0043] <Quantitative Performance Comparison with ICSG3D> Table 1 shows the reconstruction errors of the proposed NeSF-based autoencoder and the ICSG3D baseline in the test sets of three datasets. Overall, the proposed method is consistently better than ICSG3D in all evaluation metrics, with significant improvements in the type error for all datasets and all metrics for datasets such as YBCO. Figure 3 shows the crystal structures from three evaluated datasets, comparing the test samples with the reconstruction results by the proposed autoencoder and ICSG3D.

[0044]

Table 1

[0045] In the case of the simplest ICSG3D dataset among the three datasets, ICSG3D achieved excellent performance against the number-of-atoms error and position error, but the type error increased significantly (by about 65% in both the actual metric and the detected metric). In contrast, with the proposed method, the position error and the number-of-atoms error were slightly improved, and the type error decreased significantly (by about 4%). This is considered to be because ICSG3D estimates atomic species through electron density maps, while the proposed method represents atomic species more directly as a categorical distribution. Extending ICSG3D to estimate the categorical distribution of each voxel would require about 100×323 times more memory in the output, which is not practical (i.e., in addition to one electron density map, 100 categories are required for every 323 voxels).

[0046] For datasets such as LCS6Å and YBCO, which are more difficult than the ICSG3D dataset, the performance advantages of the proposed method are even more evident, especially regarding the position error. The LCS6Å dataset contains various crystal structures (non-cubic structures, distorted crystal structures, etc.), while the YBCO-like dataset contains a very narrow crystal structure. Furthermore, since datasets such as YBCO contain few samples, there is a possibility of leading to overfitting of the model (i.e., the performance of the test set may decrease significantly). Despite these difficulties, the proposed method can accurately estimate the positions and species of atoms.

[0047] To analyze in detail the relationship between the performance of the method and the structural complexity, Figure 4 shows the distribution of reconstruction errors in 10 test runs for materials from the ICSG3D and LCS6Å datasets, corresponding to the number of atoms given as the median (point) and the 68% range (colored region) around them. Datasets such as YBCO are excluded from this analysis because they contain only materials with 13 atoms in the unit cell. Figures 4a and 4b show the signed error between the detected number of atoms and the actual number of atoms. These results show that both methods correctly estimate the number of atoms for most (i.e., more than 68%) of the samples within the ICSG3D dataset (Figure 4a). However, the ICSG3D method underestimates the number of atoms in the LCS6Å dataset (Figure 4b). Similarly, Figures 4c and 4d show the distribution of position errors, and Figures 4e and 4f show the distribution of type errors. In most cases, the number of atoms is either correctly estimated or underestimated by both methods, so errors are reported based on actual measurement criteria. Examining the distribution at x = 2 in Figure 4e allows us to judge the tendency of type errors in the two-atom structures of the ICSG3D dataset.

[0048] The proposed method provides the correct types of both atoms for more than 68% of the two-atom substances, while ICSG3D often incorrectly estimates the type of one of the two atoms. Overall, the three types of errors by both methods tend to increase with the number of atoms, but for materials with varying numbers of atoms, our method consistently gives better results than ICSG3D. Compared to our method, the performance of ICSG3D tends to decline more significantly for materials with a large number of atoms. ICSG3D tends to underestimate the number of atoms (Figure 4b), which is thought to be because the spatial resolution of ICSG3D is limited to 32×32×32 voxels. This suggests that ICSG3D cannot capture multi-atom structures.

[0049] The remarkable performance of the proposed method is thought to be due to two reasons. First, our method does not use discretization, so it is more advantageous than the grid-based ICSG3D for estimating complex crystal structures. In grid-based methods, the spatial resolution is limited by the computationally and memory-intensive cubic increase. The proposed NeSF is freed from such a trade-off between resolution and computational complexity and can thus effectively represent complex structures. Second, the model size of the proposed NeSF using MLP is much smaller than that of the 3DCNN architecture of ICSG3D. In general, the number of training samples required for an ML model correlates with the number of trainable parameters. Grid-based methods use layers of 3D convolutional filters that contain many trainable parameters. In contrast, NeSF uses implicit neural representations to indirectly describe 3D space as a field rather than as voxels. Therefore, it is efficiently implemented by MLP with fewer parameters than 3DCNN. Specifically, the NeSF-based autoencoder has 760,000 parameters, which is only 2.24% of the number of parameters of the 3DCNN-based ICSG3D (34 million parameters). This difference makes NeSF more advantageous than grid-based methods, especially for small datasets such as the YBCO dataset.

[0050] <2.4 Latent Space Interpolation> By examining the learned latent space of crystal structures, the characteristics of the proposed NeSF-based autoencoder were qualitatively analyzed. In general, a good latent space should closely map similar items (in terms of properties, characteristics, categories, etc.) within the space. This provides a latent data representation that facilitates human and machine analysis. To evaluate the construction of the latent space of crystal structures, the transitions in the latent space were visualized as a series of crystal structures. If the latent space is trained to capture the relationships between materials from the perspective of structural similarity, interpolating between two points in the latent space should generate a series of materials with similar crystal structures.

[0051] The interpolation analysis is carried out as follows.

[0052] 1. Set the ICSG3D test set as the source and destination materials, and obtain their latent vectors as z src and z dst through the trained encoder.

[0053] 2. Linearly interpolate between z src and z dst in the latent space to obtain a sequence of latent vectors.

[0054] 3. Decode each latent vector through the trained decoder (NeSF) to obtain its crystal structure.

[0055] Since the intermediate crystal structures between the source and destination are reconstructed from the latent vectors through the trained decoder, these structures may not be displayed in the dataset.

[0056] To facilitate the interpretation of the analysis, the well-known zinc-blende and rock-salt structure families were selected as benchmark materials. Both families have compositions given by AX, where A is the cation and X is the anion, and the crystal structures are based on the cubic system. Therefore, if the source and destination materials belong to any of these families, the characteristic configurations and structure prototypes need to be maintained throughout the interpolation path.

[0057] As a first example, the interpolation results from ZnS (mp-10695) to CdS (mp-2469) are shown in Figure 5. The obtained compositional transitions are ZnS → MgZn3S4 (Mg 0.25 Zn 0.75 S) → MgZnS2 (Mg 0.5 Zn 0.5 S) → Mg3ZnS4 (Mg 0.75 Zn 0.25 S) → MgS → MgCd3S4 (Mg 0.25 Cd0.75 It is S) → CdS.

[0058] As a second example, the results of interpolation from MgO (mp - 1265) to NaCl (mp - 22862) are shown in Fig. 6. The transition of the obtained compositional formula is MgO → NaMgO2 (Na 0.5 Mg 0.5 O) → NaO → Na2ClO (NaCl 0.5 O 0.5 ) → Na4C l3 O( NaCl 0.75 O 0.25 ) → NaCl.

[0059] Furthermore, in SI, the results of interpolation from NaCl (mp - 22862) to PbS, from MgO (mp - 1265) to CaO (mp - 2605), and from PbS (mp - 21276) to CaO (mp - 2605) are provided.

[0060] In the interpolation examples, the composition AX and the cubic structure are mostly retained. Furthermore, the configuration changes continuously without collapsing along the interpolation path. These results suggest that our encoder learns a meaningful continuous representation by capturing the characteristics of the crystal structure in an abstract space, and that the proposed NeSF model can successfully decode these representations into crystal structures.

[0061] <2.5 Limitations and Future Directions> This study mainly focused on the development of a basic approach for estimating crystal structures using implicit neural representations. To suggest further room for improvement and important directions in future work, three main limitations of NeSF were identified.

[0062] First, NeSF adopts a relatively simple network architecture, and the options for architecture design have not been fully considered. For example, NeSF treats elements as mutually independent categorical (one-hot) vectors. Therefore, the physical properties and similarities of elements are not explicitly used to train the model. On the other hand, in other developments, attempts have been made to explicitly inject the physical properties and functions of elements into the model. For example, existing crystal structure encoders

[16] use fingerprints such as group and period, electron density, atomic radius, and electronegativity to represent input elements instead of using one-hot vectors. In ICSG3D24, the output of atomic species is trained using the mean squared error of atomic numbers instead of the categorical loss considered in our method. By incorporating these schemes into the model, it may be possible to reflect the properties and similarities between elements in the latent space and the reconstructed crystal structure. Another architectural choice regarding NeSF is whether to use the conventional (current choice) or primitive cell to represent the crystal structure. Both forms have advantages and disadvantages depending on the application. The conventional unit cell provides a more intuitive visualization that facilitates manual analysis, while the primitive unit cell provides a more compact structural representation that may facilitate computer processing. In future work, it is necessary to thoroughly analyze the design of the network architecture.

[0063] Another limitation is that the proposed NeSF does not explicitly consider the symmetry of the space group. Therefore, the local spatial arrangement of atoms in the conventional unit cell estimated by NeSF does not necessarily conform to the space group symmetry. ICSG3D

[24] also has the same limitation, but symmetry is an important concept in crystallography. Therefore, incorporating the constraints of space group symmetry into NeSF is an important direction for future work.

[0064] Finally, for evaluation purposes, instead of generative models such as variational autoencoders [16, 25, 37] and adversarial generative networks [26, 37 - 39], an autoencoder architecture was adopted. These generative models intentionally disrupt the potential structural representation to generate diverse structures not shown in the dataset. This aspect of generative models is more suitable for discovering new structures, but the lack of ground truth structures hinders quantitative and reliable performance analysis.

[0065] NeSF needs to be applied to generative models in future work using appropriate performance analysis.

[0066] <3 Summary> In this embodiment, NeSF is proposed to estimate crystal structures using neural networks. It is difficult to directly determine crystal structures using neural networks because these structures are basically represented as unordered sets containing various numbers of atoms. NeSF overcomes this problem by treating the crystal structure not as a discrete set of atoms but as a continuous vector field. The idea of NeSF borrows from vector fields in physics and recent implicit neural representations in computer vision. Implicit neural representations are ML techniques that use neural networks to represent 3D geometry. NeSF extends this approach by introducing a position field and a type field to estimate the atomic positions and types of the crystal structure respectively. Different from existing grid - based approaches for representing crystal structures, NeSF can represent any crystal structure without a trade - off between spatial resolution and computational complexity.

[0067] NeSF was applied as an autoencoder for crystal structures and demonstrated its performance and expressiveness on a dataset with diverse crystal structures. Quantitative performance analysis showed clear advantages of the NeSF-based autoencoder over existing grid-based methods, especially in estimating complex crystal structures. Furthermore, qualitative analysis of the learned latent space revealed that the autoencoder captures similarities between crystal structures rather than randomly mapping them.

[0068] In materials science, the design and construction of crystal structures are fundamental processes in searching for materials with desired properties. ML has advanced rapidly with the development of neural networks, and representing any crystal structure using these networks is essential for next-generation materials development. For example, NeSF can be easily incorporated into powerful deep generative models such as variational autoencoders and adversarial generative networks to discover new crystal structures. Such crystal structure generation models are important for materials inverse design, which is a major challenge in MI. NeSF can overcome the technical bottleneck of ML in crystal structure estimation and open the way for next-generation materials development.

[0069] FIG. 7 is a block diagram showing the hardware configuration of the information processing apparatus 10 according to the present embodiment. As shown in FIG. 7, the information processing apparatus 10 includes a CPU (Central Processing Unit) 12, a memory 14, a storage device 16, an input / output I / F (Interface) 18, a storage medium reader 20, and a communication I / F 22. Each component is connected to be communicable with each other via a bus 24.

[0070] The storage device 16 stores an information processing program for executing each process described later. The CPU 12 is a central processing unit that executes various programs and controls each component. That is, the CPU 12 reads a program from the storage device 16 and executes the program using the memory 14 as a work area. The CPU 12 performs the above-described various arithmetic processes according to the program stored in the storage device 16.

[0071] The memory 14 is composed of a RAM (Random Access Memory) and temporarily stores programs and data as a working area. The storage device 16 is composed of a ROM (Read Only Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc., and stores various programs including an operating system and various data.

[0072] The input / output I / F 18 is an interface that performs input of data from an external device and output of data to an external device. Also, for example, input devices for performing various inputs such as a keyboard and a mouse, and output devices for outputting various information such as a display and a printer may be connected. By adopting a touch panel display as the output device, it may function as an input device.

[0073] The storage medium reader 20 reads data stored in various storage media such as a CD (Compact Disc)-ROM, a DVD (Digital Versatile Disc)-ROM, a Blu-ray disc, a USB (Universal Serial Bus) memory, etc., and writes data to the storage medium.

[0074] The communication I / F 22 is an interface for communicating with other devices, and for example, standards such as Ethernet (registered trademark), FDDI, Wi-Fi (registered trademark) are used.

[0075] The information processing apparatus 10 of the present embodiment learns an autoencoder including an encoder and a decoder using the method as described above. Thereby, a field representing the structure of a substance composed of an atomic point group can be expressed using a neural network model.

[0076] Next, the functional configuration of the information processing apparatus 10 will be described. As shown in FIG. 8, the information processing apparatus 10 functionally includes a learning acquisition unit 102, a learning unit 104, an acquisition unit 108, and a processing unit 110. Further, a data storage unit 100 and a learned model storage unit 106 are provided in a predetermined storage area of the information processing apparatus 10. Each functional configuration is realized by the CPU 12 reading out each program stored in the storage device 16, expanding it in the memory 14, and executing it.

[0077] First, the information processing apparatus 10 causes a neural network model that represents a field representing the structure of a substance composed of an atomic point group to be learned.

[0078] The data storage unit 100 stores learning crystal data representing the crystal structure of a substance. The learning crystal data is configured to include position data representing the positions of atoms constituting the crystal of the substance, type data representing the types of atoms constituting the crystal of the substance, and lattice constant data of the crystal of the substance.

[0079] FIG. 9 is a diagram showing an example of a plurality of learning crystal data stored in the data storage unit 100. As shown in FIG. 9, one piece of learning crystal data is data representing the crystal structure of a certain substance. As shown in FIG. 9, one piece of learning crystal data is stored in association with position data representing the positions of a plurality of atoms constituting the crystal of a substance, type data representing the types of a plurality of atoms constituting the crystal of the substance, and lattice constant data of the crystal of the substance. The position data is three-dimensional position coordinate data of each of the plurality of atoms. Also, the type data is label data representing the type of each of the plurality of atoms. The lattice constant data includes the length of the crystal axis and the angle between the axes.

[0080] FIG. 10 is a diagram for explaining the structure of the autoencoder of the present embodiment and the outline of the processing executed by the information processing apparatus 10 of the present embodiment. FIG. 10 is a more simplified diagram of FIG. 1(b).

[0081] As shown in FIG. 10, the autoencoder AE of the present embodiment includes an encoder E, a first decoder D1, a second decoder D2, and a third decoder D3.

[0082] As shown in FIG. 10, when a combination of position data P representing the positions of atoms constituting the crystal of a substance, type data S representing the types of atoms constituting the crystal of the substance, and lattice constant data L of the crystal of the substance is input, the encoder E outputs a latent vector z. Note that for one crystal structure of a substance, one crystal data representing a combination of the position data P, the type data S, and the lattice constant data L is set.

[0083] Also, as shown in FIG. 10, when a combination of a query point p which is a point of interest in the substance and the latent vector z is input, the first decoder D1 outputs a position field f representing a field of positions of atoms constituting the crystal of the substance. p Based on this position field f p estimated position data P of atoms constituting the crystal of the substance e is calculated. The calculation method will be described later.

[0084] Also, as shown in FIG. 10, when a combination of a query point p which is a point of interest in the substance and the latent vector z is input, the second decoder D2 outputs a type field f representing a field of types of atoms constituting the crystal of the substance. s Based on this type field f s estimated type data S of atoms constituting the crystal of the substance e is calculated. The calculation method will be described later.

[0085] Also, as shown in FIG. 10, when a query point p which is a point of interest in the substance and the latent vector z are input, the third decoder D3 outputs estimated lattice constant data L of the crystal of the substance. eOutput. In this embodiment, the case where there is one third decoder D3 will be described as an example, but there may be two third decoders D3. In this case, for example, the third decoder D3 is composed of a decoder that outputs the length of the crystal axis and a decoder that outputs the interaxial angle.

[0086] The information processing apparatus 10 of this embodiment makes each parameter of the autoencoder AE learned by unsupervised machine learning (or self-supervised learning) so that the combination of the position data P, the type data S, and the lattice constant data L input to the autoencoder AE matches the combination of the estimated position data P e and the estimated type data S e and the estimated lattice constant data L e output from the autoencoder AE. Thereby, a learned autoencoder AE is obtained. Also, a learned first decoder D1, a learned second decoder D2, and a learned third decoder D3, which are components of the learned autoencoder AE, are obtained.

[0087] When the learning acquisition unit 102 receives an instruction signal to learn the autoencoder AE, it reads out the learning crystal data stored in the data storage unit 100. Also, the learning acquisition unit 102 sets a learning query point that is a point of interest in the learning substance. Note that the learning acquisition unit 102 may set the learning query point by randomly sampling positions in the space within the learning substance.

[0088] When the learning unit 104 learns the autoencoder AE using unsupervised machine learning, it inputs the learning crystal data to the encoder E in the autoencoder AE to obtain a latent vector z representing the crystal structure of the learning substance. Note that the encoder E can be realized, for example, by using the ideas of PointNet

[28] or DeepSets

[29] described above.

[0089] Next, the learning unit 104 inputs a combination of the latent vector z and the learning query point p to the first decoder D1 of the autoencoder AE, thereby obtaining a position field f representing the positions of the atoms constituting the crystal of the learning material. p Then, the learning unit 104 estimates the positions of the atoms constituting the crystal of the learning material based on the position field f output from the first decoder D1. p

[0090] FIG. 11 is a diagram for explaining the setting of the learning query points and the position field. Note that FIG. 11 is the same as FIG. 2 but is reproduced for explanation. As shown in FIG. 11, a plurality of learning query points are set in the space M within the material. The white circles shown in FIG. 11 represent the learning query points. Also, A1, A2, and A3 shown in FIG. 11(a) represent the actual positions of the atoms in the space M within the material.

[0091] The learning unit 104 inputs a combination of the learning query point p and the latent vector z to the first decoder D1, thereby obtaining a position field f as shown in FIG. 11(b). Each of the arrows shown in FIG. 11 corresponds to the position field f. p p

[0092] Then, as shown in FIGS. 11(b) and (c), the learning unit 104 repeatedly updates the position of each of the plurality of learning query points according to the position field f output from the first decoder D1. Specifically, the learning unit 104 updates the position of each of the plurality of learning query points with respect to the position field f. p pBy adding the vectors represented thereby, the positions of new learning query points are generated. As a result, as shown in FIG. 11(d), the positions of each of the plurality of learning query points converge to the positions of the actual atoms. Then, based on the positions of each of the plurality of learning query points, the learning unit 104 estimates the positions of the atoms as shown in FIG. 11(e), for example, by using Non-max Suppression. P e 1,P e 2,P e 3 is obtained.

[0093] Next, the learning unit 104 inputs a combination of the latent vector z and the learning query point p to the second decoder D2 of the autoencoder AE, thereby obtaining a type field fs representing the field of the types of atoms constituting the crystal of the learning substance. Then, the learning unit 104 estimates the types of atoms constituting the crystal of the learning substance based on the type field fs output from the second decoder D2.

[0094] FIG. 12 is a diagram for explaining the setting of the learning query points and the type field. Note that FIG. 12 is the same as FIG. 2, but is reproduced for explanation. As shown in FIG. 12, in the space M within the substance, a plurality of learning query points are set. Similar to FIG. 12, the white circles shown in FIG. 12 represent the learning query points.

[0095] Specifically, as shown in FIGS. 12(f) and (g), the learning unit 104 determines the estimated positions P e 1,P e 2,P e 3 of the atoms and the positions around the estimated positions P e 1,P e 2,P e 3, and sets a plurality of learning query points.

[0096] Then, the learning unit 104 inputs a combination of the learning query point p and the latent vector z to the second decoder D2, thereby obtaining the type field f sis acquired. As shown in Fig. 12(g), the type field f s corresponds to the probability representing the type of atom obtained for each learning query point. In the example shown in Figs. 12(g) and (h), it is shown that the probability that the type of atom located at a certain learning query point is iron (Fe) is the highest, and the probability that the type of atom located at another learning query point is copper (Cu) is the highest.

[0097] Then, as shown in Figs. 12(g) and (h), the learning unit 104 estimates the type of atom corresponding to the position of each of the plurality of learning query points according to the type field f s output from the second decoder D2. For example, the learning unit 104 determines, for each of the estimated positions P e 1, P e 2, P e 3, the element with the highest probability at the learning query point corresponding to the estimated position and the element with the highest probability at the learning query points around the estimated position. Then, the learning unit 104 estimates the most frequent element among the plurality of learning query points set for each of the estimated positions P e 1, P e 2, P e 3 as the type of atom at the estimated positions P e 1, P e 2, P e 3. For this reason, the estimated type S e 1 of the atom existing at the estimated position P e 1, the estimated type S e 2 of the atom existing at the estimated position P e 2, and the estimated type S e 3 of the atom existing at the estimated position P e 3 are obtained.

[0098] For example, as shown in Fig. 12(i), it is estimated that the type of the atom located at the estimated position P e 1 is silicon (Si), it is estimated that the type of the atom located at the estimated position P e 2 is iron (Fe), and it is estimated that the type of the atom located at the estimated position P e 3 is copper (Cu).

[0099] Next, the learning unit 104 inputs the latent vector z to the third decoder D3 of the autoencoder AE to obtain the lattice constant data L of the crystal of the learning material. e The lattice constant data L output from the third decoder D3 e includes the lengths of the crystal axes and the angles between the axes.

[0100] Then, the learning unit 104 uses unsupervised machine learning to train the autoencoder AE so that the combination of the estimated atomic positions, the estimated atomic types, and the lattice constant data output from the third decoder D3 corresponds to the combination of the atomic positions, the atomic types, and the lattice constant data in the learning crystal data, thereby generating the learned first decoder D1 and the learned second decoder D2.

[0101] For example, in the above example, the learning unit 104 makes the estimated positions P e 1, P e 2, P e 3 of the estimated atoms match the atomic position data P1, P2, P3 in the learning crystal data, and the estimated types S e 1, S e 2, S e 3 of the estimated atoms match the atomic type data S1, S2, S3 in the learning crystal data, and the lattice constant data L e output from the third decoder D3 matches the lattice constant data L in the learning crystal data. By training the autoencoder AE using unsupervised machine learning, the learned first decoder D1 and the learned second decoder D2 are generated.

[0102] The learning unit 104 stores the learned autoencoder AE in the learned model storage unit 106. Since the learned autoencoder AE also includes the learned first decoder D1, the learned second decoder D2, and the learned third decoder D3, those learned models are also stored in the learned model storage unit 106.

[0103] The learned encoder E outputs a latent vector z representing the crystal of the substance when crystal data representing the crystal structure of the substance, which includes position data representing the positions of the atoms constituting the crystal of the substance, type data representing the types of the atoms constituting the crystal of the substance, and lattice constant data of the crystal of the substance, is input.

[0104] As described above, the field data representing the structure field of the substance of the present embodiment is represented by a position field representing the position field of the atoms constituting the crystal of the substance and a type field representing the type field of the atoms constituting the crystal of the substance. Therefore, the learned first decoder D1, which is an example of the learned first neural network model, outputs a position field f p corresponding to the query point when an arbitrary vector and a query point are input. Also, the second decoder D2, which is an example of the learned second neural network model, outputs a type field f s corresponding to the query point when an arbitrary vector and a query point are input. Note that the learned first decoder D1 and the learned second decoder D2 are examples of the learned neural network models that output field data representing the structure field of the substance at the query point when an arbitrary vector and a query point are input.

[0105] The learned third decoder outputs lattice constant data L e of the crystal of the substance when a combination of an arbitrary vector replacing the latent vector z and a query point is input.

[0106] When the acquisition unit 108 receives an instruction signal to acquire the field data of the crystal structure of the target substance, the acquisition unit 108 reads out the learned first decoder D1 and the learned second decoder D2 stored in the learned model storage unit 106. Also, the acquisition unit 108 acquires, for example, a target vector input from the user and a query point that is the point of interest within the substance.

[0107] Note that the target vector can be any vector that replaces the latent vector z, and can be any vector as long as it describes the properties of the material in the same way as the latent vector. For example, the target vector may be a vector representing the physical properties desired by the user. For example, the target vector may be a vector representing the physical properties of a substance in which [bandgap = xxx] is stored in the first component and [formation energy = xxx] is stored in the second component. In this case the learned first decoder D1 and the learned second decoder D2 are treated as pre-trained models, and by re-training with a vector representing the physical properties of the material, useful field data will be output from the learned first decoder D1 and the learned second decoder D2.

[0108] The processing unit 110 inputs the target vector and the query point acquired by the acquisition unit 108 to the learned first decoder D1 and the learned second decoder D2, thereby acquiring field data corresponding to the query point.

[0109] Specifically, the processing unit 110 inputs the target vector and the query point acquired by the acquisition unit 108 to the learned first decoder D1, thereby acquiring the position field f p corresponding to the query point. Also, the processing unit 110 inputs the target vector and the query point acquired by the acquisition unit 108 to the learned second decoder D2, thereby acquiring the type field f s corresponding to the query point.

[0110] The field data output from the learned first decoder D1 and the learned second decoder D2 is data representing the field of the crystal structure of the material according to the target vector and the query point. By inputting a target vector, which is an arbitrary vector instead of the latent vector, to the learned first decoder D1 and the learned second decoder D2, it may be possible to predict the crystal structure represented by the target vector.

[0111] For example, when the target vector is a vector representing desired physical properties as described above, the field data corresponding to the vector is output. By inputting a plurality of query points into the learned second decoder D2 and the learned second decoder D2, field data at those plurality of query points is acquired, and based on that field data, it becomes possible to generate the crystal structure of a substance having the desired physical properties.

[0112] As the target vector, for example, it is also possible to use measurement data of XRD (X-ray diffraction). In this case, by inputting the measurement data of XRD of the target substance into the learned first decoder D1 and the learned second decoder D2, it becomes possible to predict the crystal structure of the target substance. Also, it is possible to use in combination the learned first decoder D1 and the learned second decoder D2 of the present embodiment and an encoder that processes other modalities such as text data.

[0113] <Operation of Information Processing Apparatus 10> Next, with reference to the drawings, the operation of the information processing apparatus 10 of the embodiment will be described. When the information processing apparatus 10 receives learning crystal data, it stores it in the data storage unit 100. Then, when the information processing apparatus 10 receives an instruction signal to start the learning process, it executes the learned model generation processing routine shown in FIG. 13.

[0114] <Learned Model Generation Processing Routine> In step S100, the learning acquisition unit 102 acquires a plurality of learning crystal data stored in the data storage unit 100.

[0115] In step S102, the learning acquisition unit 102 sets one piece of learning crystal data from the plurality of learning crystal data acquired in step S100 above. Then, the learning unit 104 sets a plurality of learning query points p in the space within the substance represented by the set learning crystal data. i Here, i is an index for identifying the learning query point.

[0116] In step S104, the learning unit 104 obtains the latent vector z by inputting the position data P, type data S, and lattice constant data L among the learning crystal data set in step S102 to the encoder E of the autoencoder AE.

[0117] In step S106, the learning unit 104 i inputs the combination of the learning query point p set in step S102 and the latent vector z set in step S104 to the first decoder D1 to obtain the position field f. p Note that the learning unit 104 obtains the position field f for each of a plurality of learning query points. p

[0118] In step S108, the learning unit 104 i inputs the combination of the learning query point p set in step S102 and the latent vector z set in step S104 to the second decoder D2 to obtain the type field f. s Note that the learning unit 104 obtains the type field f for each of a plurality of learning query points. s

[0119] In step S110, the learning unit 104 i inputs the combination of the learning query point p set in step S102 and the latent vector z set in step S104 to the third decoder D3 to obtain the lattice constant data L. e Note that the learning unit 104 obtains the lattice constant data L for each of a plurality of learning query points. e

[0120] In step S112, the learning unit 104 uses the position field f obtained in step S106. pEstimate the positions of the atoms based on this. Note that the learning unit 104 updates the positions of a plurality of learning query points by the method described above, and the final estimated position P e is obtained.

[0121] In step S114, the learning unit 104 estimates the type of the atom based on the estimated position P of the atom obtained in step S112 e and the type field f obtained in step S108 s . Thus, the learning unit 104 obtains the estimated type S of the atom existing at the estimated position P e . e is obtained.

[0122] In step S116, the learning unit 104 causes the autoencoder AE to be learned using unsupervised machine learning so that the estimated position P of the atom obtained in step S112 e matches the position data P of the atom in the learning crystal data, and the estimated type S of the atom obtained in step S114 e matches the type data S of the atom in the learning crystal data, and the lattice constant data L obtained in step S110 e matches the lattice constant data L in the learning crystal data, thereby generating a learned first decoder D1 and a learned second decoder D2.

[0123] The processes of steps S102 to S116 are repeated until the end condition of the machine learning is satisfied. For example, as the end condition of the machine learning, conditions such as whether the machine learning process has been executed a predetermined number of times or whether the error between the data output from the autoencoder AE and the learning crystal data is equal to or less than a predetermined threshold value can be adopted. Note that in the above description, the case where one learning crystal data is set and learning is executed has been described as an example, but the present invention is not limited to this, and machine learning may be executed at once using a plurality of learning crystal data.

[0124] In step S118, the learning unit 104 determines whether or not the above-described end condition is satisfied. If the end condition is satisfied, the process proceeds to step S120. On the other hand, if the end condition is not satisfied, the process returns to step S102.

[0125] In step S120, the learning unit 104 stores the learned autoencoder AE obtained by the machine learning process of steps S102 to S116 in the learned model storage unit 106.

[0126] Next, when the information processing apparatus 10 receives a target vector and a query point, it executes the estimation processing routine shown in FIG. 14.

[0127] <Estimation processing routine> In step S200, the acquisition unit 108 acquires a target vector and a query point.

[0128] In step S202, the processing unit 110 reads out the learned first decoder D1 and the learned second decoder D2 stored in the learned model storage unit 106.

[0129] In step S204, the processing unit 110 inputs the target vector and the query point acquired in step S200 to the learned first decoder D1 read out in step S202, thereby obtaining a position field f p corresponding to the target vector and the query point.

[0130] In step S206, the processing unit 110 inputs the target vector and the query point acquired in step S200 to the learned second decoder D2 read out in step S202, thereby obtaining a type field f s corresponding to the target vector and the query point.

[0131] In step S208, the processing unit 110 uses the position field f acquired in step S204 pand the type field f acquired in step S206 s are output as a result.

[0132] As described above, according to the information processing apparatus 10, it is possible to train a neural network model that represents a field representing the structure of a substance composed of an atomic point group. Further, according to the information processing apparatus 10, it is possible to represent a field representing the structure of a substance composed of an atomic point group by using the neural network model.

[0133] Specifically, the information processing apparatus 10 acquires learning crystal data representing the crystal structure of a learning substance, the learning crystal data including position data representing the positions of atoms constituting the crystal of the learning substance, type data representing the types of atoms constituting the crystal of the learning substance, and lattice constant data of the crystal of the learning substance, and a learning query point which is a point of interest in the learning substance. When the information processing apparatus 10 trains an autoencoder using unsupervised machine learning, it inputs the learning crystal data to the encoder of the autoencoder to acquire a latent vector representing the crystal structure of the learning substance. The information processing apparatus 10 inputs a combination of the latent vector and the learning query point to the first decoder of the autoencoder to acquire a position field representing the field of positions of atoms constituting the crystal of the learning substance. The information processing apparatus 10 estimates the positions of atoms constituting the crystal of the learning substance based on the position field output from the first decoder. The information processing apparatus 10 inputs a combination of the latent vector and the learning query point to the second decoder of the autoencoder to acquire type field data representing the field of types of atoms constituting the crystal of the learning substance. The information processing apparatus 10 estimates the types of atoms constituting the crystal of the learning substance based on the type field output from the second decoder. The information processing apparatus 10 inputs a combination of the latent vector and the learning query point to the third decoder of the autoencoder to acquire the lattice constant data of the crystal of the learning substance. The information processing apparatus 10 trains the autoencoder using unsupervised machine learning so that the combination of the estimated atom positions, the estimated atom types, and the lattice constant data output from the third decoder corresponds to the combination of the position data, the type data, and the lattice constant data in the learning crystal data, thereby generating a trained first decoder and a trained second decoder.

[0134] In this embodiment, by adopting the method as described above, it has become possible to represent the field representing the structure of a substance by a neural field. Also, it has become possible to generate crystal data from a fixed-length vector.

[0135] Further, the information processing apparatus 10 acquires a target vector and a query point which is a point of interest within a substance, and inputs the target vector and the query point to a first decoder and a second decoder which are examples of a learned neural network model, thereby acquiring field data corresponding to the query point. This field data is represented by a position field representing the field of positions of atoms constituting the crystal of the substance and a type field representing the field of types of atoms constituting the crystal of the substance. When an arbitrary vector replacing the latent vector and the query point are input to the first decoder, the first decoder outputs a position field corresponding to the query point. When an arbitrary vector replacing the latent vector and the query point are input to the second decoder, the second decoder outputs a type field corresponding to the query point. The information processing apparatus 10 inputs the target vector and the query point to the first decoder, thereby acquiring a position field corresponding to the query point. Also, the information processing apparatus 10 inputs the target vector and the query point to the second decoder, thereby acquiring a type field corresponding to the query point. For example, when the target vector is a vector representing a desired physical property, field data corresponding to that vector is output. By inputting a plurality of query points to the learned second decoder D2 and the learned second decoder D2, field data at those plurality of query points is acquired, and thus it is also possible to generate a crystal structure of a substance having a desired physical property based on that field data.

[0136] Note that when changing the target vector, it is preferable to re-learn the learned first decoder D1 and learned second decoder D2. In this case, the learned first decoder D1 and learned second decoder D2 are used as pre-learned models.

[0137] In the above embodiment, the case where the field representing the crystal structure is modeled by a neural field has been described as an example, but the present invention is not limited thereto. Not only crystal structures including a repeating structure, but also structures of substances composed of atomic point groups that do not include a repeating structure may be modeled by a neural field. In this case, since it may not be necessary to estimate the lattice constant, the Lattice Decorder shown in FIG. 1 may not be necessary.

[0138] Also, in the above-described embodiment, the case where the vector z input to the decoder is the latent vector output from the encoder has been described as an example, but the present invention is not limited thereto. For example, a vector representing physical properties desired by a user may be set as z, and the vector z may be configured to be given to the decoder (for example, conditions such as bandgap = xxx, formation energy = xxx are expressed as vectors). Further, a configuration may be adopted in which random noise is used as the vector z.

[0139] In the above embodiment, the case where an autoencoder including an encoder and a decoder is used as the neural network model has been described as an example, but the present invention is not limited thereto. Other models may be used as the neural network model, and for example, a generative adversarial network (GAN) may be used.

[0140] Also, FIGS. 15 to 30 show diagrams for explaining the details of the present embodiment.

[0141] In addition, in the above-described embodiment, each process executed by the CPU by reading software (program) may be executed by various processors other than the CPU. Examples of the processor in this case include a PLD (Programmable Logic Device) whose circuit configuration can be changed after manufacture, such as an FPGA (Field-Programmable Gate Array), and a dedicated electric circuit which is a processor having a circuit configuration dedicated to executing specific processes, such as an ASIC (Application Specific Integrated Circuit). Further, each process may be executed by one of these various processors, or may be executed by a combination of two or more processors of the same type or different types (for example, a plurality of FPGAs, a combination of a CPU and an FPGA, etc.). More specifically, the hardware structure of these various processors is an electric circuit combining circuit elements such as semiconductor elements.

[0142] In the above-described embodiment, an aspect in which each program is pre-stored (installed) in the storage device has been described, but the present invention is not limited to this. The program may be provided in a form stored in a storage medium such as a CD-ROM, a DVD-ROM, a Blu-ray disk, a USB memory, or the like. Further, the program may be in a form downloaded from an external device via a network.

[0143] (Supplementary Note) Hereinafter, aspects of the present disclosure will be appended.

[0144] (Supplementary Note 1) An information processing apparatus including a processing unit that uses a neural network model to represent a field representing the structure of a substance composed of an atomic point group.

[0145] (Supplementary Note 2) Further including an acquisition unit that acquires a target vector and a query point that is a point of interest in the substance, The neural network model is a trained neural network model that outputs field data representing the structure field of the substance at the query point when an arbitrary vector and the query point are input. The processing unit inputs the target vector and the query point acquired by the acquisition unit to the trained neural network model, thereby acquiring the field data corresponding to the query point. The information processing apparatus according to Supplementary Note 1.

[0146] (Supplementary Note 3) The field data is represented by a position field representing the position field of the atoms constituting the crystal of the substance and a type field representing the type field of the atoms constituting the crystal of the substance. The trained neural network model includes a trained first neural network model and a trained second neural network model. When the arbitrary vector and the query point are input, the trained first neural network model outputs the position field corresponding to the query point. When the arbitrary vector and the query point are input, the trained second neural network model outputs the type field corresponding to the query point. The processing unit Inputs the target vector and the query point acquired by the acquisition unit to the trained first neural network model, thereby acquiring the position field corresponding to the query point. Inputs the target vector and the query point acquired by the acquisition unit to the trained second neural network model, thereby acquiring the type field corresponding to the query point. The information processing apparatus according to Supplementary Note 2.

[0147] (Supplementary Note 4) The pre-trained neural network model is a pre-trained model generated in advance by machine learning based on learning crystal data that represents the crystal structure of the substance and includes position data representing the positions of atoms constituting the crystal of the substance, type data representing the types of atoms constituting the crystal of the substance, and lattice constant data of the crystal of the substance. The information processing apparatus according to appendix 2 or appendix 3.

[0148] (Appendix 5) The pre-trained neural network model is a pre-trained first decoder and a pre-trained second decoder among pre-trained autoencoders. When crystal data that represents the crystal structure of the substance and includes position data representing the positions of atoms constituting the crystal of the substance, type data representing the types of atoms constituting the crystal of the substance, and lattice constant data of the crystal of the substance is input to the pre-trained encoder among the pre-trained autoencoders, a latent vector representing the crystal of the substance is output. When the arbitrary vector replacing the latent vector and the query point are input to the pre-trained first decoder, the position field corresponding to the query point is output. When the arbitrary vector replacing the latent vector and the query point are input to the pre-trained second decoder, the type field corresponding to the query point is output. The pre-trained first decoder and the pre-trained second decoder are pre-trained models obtained by performing unsupervised machine learning on an autoencoder including an encoder, a first decoder, and a second decoder based on the crystal data. The information processing apparatus according to appendix 3.

[0149] (Appendix 5) The autoencoder further includes a third decoder. The learned third decoder is a learned model that outputs lattice constant data of the crystal of the substance when any vector replacing the latent vector is input. The information processing apparatus according to Supplementary Note 4.

[0150] (Supplementary Note 7) A learned model generation apparatus including a learning unit that learns a neural network model representing a field representing the structure of a substance composed of an atomic point group.

[0151] (Supplementary Note 8) Learning crystal data representing the crystal structure of the substance for learning, including position data representing the positions of atoms constituting the crystal of the substance for learning, type data representing the types of atoms constituting the crystal of the substance for learning, and lattice constant data of the crystal of the substance for learning, and a learning acquisition unit that acquires a learning query point that is a point of interest in the substance for learning. The learning unit When learning an autoencoder using unsupervised machine learning, By inputting the learning crystal data to the encoder of the autoencoder, a latent vector representing the crystal structure of the substance for learning is obtained. By inputting a combination of the latent vector and the learning query point to the first decoder of the autoencoder, a position field representing the field of positions of atoms constituting the crystal of the substance for learning is obtained. Based on the position field output from the first decoder, the positions of atoms constituting the crystal of the substance for learning are estimated. By inputting a combination of the latent vector and the learning query point to the second decoder of the autoencoder, type field data representing the field of types of atoms constituting the crystal of the substance for learning is obtained. Based on the type field output from the second decoder, the types of atoms constituting the crystal of the substance for learning are estimated. For the third decoder among the autoencoders, by inputting a combination of the latent vector and the learning query points, lattice constant data of the crystal of the learning substance is obtained. Using unsupervised machine learning to train the autoencoder so that the combination of the estimated atomic positions, the estimated atomic types, and the lattice constant data output from the third decoder corresponds to the combination of the position data, the type data, and the lattice constant data in the learning crystal data, a trained first decoder and a trained second decoder are generated. The learned model generation device according to Supplementary Note 7.

[0152] (Supplementary Note 9) An information processing method executed by a computer, using a neural network model to represent a field representing the structure of a substance composed of atomic point groups.

[0153] (Supplementary Note 10) A learned model generation method executed by a computer for training a neural network model representing a field representing the structure of a substance composed of atomic point groups.

[0154] (Supplementary Note 11) An information processing program for causing a computer to execute processing using a neural network model to represent a field representing the structure of a substance composed of atomic point groups.

[0155] (Supplementary Note 12) A learned model generation program for causing a computer to execute processing for training a neural network model representing a field representing the structure of a substance composed of atomic point groups.

Explanation of Signs

[0156] 10 Information processing device 10 Information processing device 100 Data storage unit 102 Learning acquisition unit 104 Learning Department 106 Model Memory Unit 108 Acquisition Unit 110 Processing Unit

Claims

1. An information processing device including a processing unit that uses a neural network model to represent a field that represents a structure of a substance constituted by a group of atomic points.

2. The method further includes an acquisition unit that acquires a target vector and a query point that is a point of interest within the material, the neural network model is a trained neural network model that, when an arbitrary vector and the query point are input, outputs field data representing a field of a structure of the material at the query point; the processing unit inputs the target vector and the query point acquired by the acquisition unit to the trained neural network model, thereby acquiring the field data corresponding to the query point; The information processing device according to claim 1 .

3. the field data is represented by a position field representing a field of a position of an atom constituting the crystal of the substance, and a type field representing a field of a type of an atom constituting the crystal of the substance; the trained neural network model includes a trained first neural network model and a trained second neural network model; the trained first neural network model, when the arbitrary vector and the query point are input, outputs the position field corresponding to the query point; the second trained neural network model, when the arbitrary vector and the query point are input, outputs the type field corresponding to the query point; The processing unit includes: acquiring the position field corresponding to the query point by inputting the target vector and the query point acquired by the acquisition unit into the trained first neural network model; acquiring the type field corresponding to the query point by inputting the target vector and the query point acquired by the acquisition unit into the trained second neural network model; The information processing device according to claim 2 .

4. The trained neural network model is A trained model is generated in advance by machine learning based on training crystal data, the training crystal data including: crystal data representing a crystal structure of the substance, the crystal data including position data representing positions of atoms constituting the crystal of the substance, type data representing types of atoms constituting the crystal of the substance, and lattice constant data of the crystal of the substance.

4. The information processing device according to claim 2.

5. The trained neural network model is a trained first decoder and a trained second decoder of a trained autoencoder, When a trained encoder of the trained autoencoder receives crystal data representing a crystal structure of the substance, the crystal data including position data representing positions of atoms constituting the crystal of the substance, type data representing types of atoms constituting the crystal of the substance, and lattice constant data of the crystal of the substance, the trained encoder outputs a latent vector representing a crystal of the substance; the trained first decoder, when the arbitrary vector replacing the latent vector and the query point are input, outputs the position field corresponding to the query point; the trained second decoder, when the arbitrary vector replacing the latent vector and the query point are input, outputs the type field corresponding to the query point; The trained first decoder and the trained second decoder are trained models obtained by performing unsupervised machine learning on an autoencoder including an encoder, a first decoder, and a second decoder based on the crystal data. The information processing device according to claim 3 .

6. the autoencoder further comprising a third decoder; The trained third decoder is a trained model that outputs lattice constant data of a crystal of the substance when the arbitrary vector instead of the latent vector is input. The information processing device according to claim 5 .

7. A trained model generating device including a learning unit that trains a neural network model that expresses a field that represents the structure of a substance composed of atomic point groups.

8. The learning acquisition unit further includes: learning crystal data representing a crystal structure of the learning substance, the learning crystal data including position data representing the positions of atoms constituting the crystal of the learning substance, type data representing the types of atoms constituting the crystal of the learning substance, and lattice constant data of the crystal of the learning substance; and a learning query point which is a point of interest within the learning substance; The learning unit is When training an autoencoder using unsupervised machine learning, The training crystal data is input to an encoder of the autoencoder to obtain a latent vector representing a crystal structure of the training substance; A combination of the latent vector and the training query point is input to a first decoder of the autoencoder to obtain a position field representing a position field of atoms constituting a crystal of the training substance; estimating positions of atoms constituting the crystal of the learning material based on the position field output from the first decoder; A combination of the latent vector and the training query point is input to a second decoder of the autoencoder to obtain a type field representing a field of types of atoms constituting a crystal of the training substance; estimating the type of atoms constituting the crystal of the learning material based on the type field output from the second decoder; A combination of the latent vector and the training query point is input to a third decoder of the autoencoder to obtain lattice constant data of a crystal of the training material; generating a trained first decoder and a trained second decoder by training an autoencoder using unsupervised machine learning so that a combination of the estimated atomic positions, the estimated atomic types, and the lattice constant data output from the third decoder corresponds to a combination of the position data, the type data, and the lattice constant data in the training crystal data; The trained model generating device according to claim 7.

9. An information processing method in which a computer executes a process to represent a field that represents the structure of a substance composed of a group of atomic points using a neural network model.

10. A trained model generation method in which a computer executes processing to train a neural network model that expresses a field that represents the structure of a substance composed of atomic point groups.

11. An information processing program that uses a neural network model to cause a computer to execute processing that represents a field that represents the structure of a substance composed of atomic point groups.

12. A trained model generation program that trains a neural network model that expresses a field that represents the structure of a substance composed of atomic point groups, and allows a computer to execute processing.