Model generation method, data presentation method, data generation method, estimation method, model generation device, data presentation device, data generation device, and estimation device

The model generation method in Materials Informatics uses machine learning to process crystal structure data, overcoming the challenges of conventional methods by predicting material properties accurately and cost-effectively without extensive correct answer information.

JP7674636B2Active Publication Date: 2025-05-12OMRON CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021157205
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-27
Publication Date
2025-05-12
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

Conventional Methods for Materials Informatics face challenges in achieving accurate and cost-effective predictions of material properties due to the complexity of first-principles calculations and the need for extensive correct answer information.

Method used

A model generation method using machine learning that acquires and processes first and second data related to crystal structures, employing encoders to convert this data into feature vectors, allowing for the generation of trained models that can predict material properties without requiring extensive correct answer information.

Benefits of technology

This approach enables the accurate and cost-effective prediction of material properties by mapping similar materials into nearby areas in a feature space, reducing the need for expensive correct answer information and improving the efficiency of material development processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007674636000006
    Figure 0007674636000006
  • Figure 0007674636000007
    Figure 0007674636000007
  • Figure 0007674636000008
    Figure 0007674636000008
Patent Text Reader

Abstract

To attain a new perception about a material at low cost.SOLUTION: A method for generating a model according to one aspect of the present invention acquires first data and second data related to the crystal structure of a material and conducts a mechanical learning of a first encoder and a second encoder by using the first data and the second data. The second data shows the nature of the material with an index different from that of the first data. The first encoder is formed to convert the first data into a first feature vector and the second encoder is formed to convert the second data into a second feature vector. The dimension of the first feature vector is the same as that of the second feature vector. In the mechanical learning, the values of the feature vectors of a positive sample of the first encoder and the second encoder are positioned close to each other, and the values of the feature vector of a negative sample is positioned far away in comparison with the value of the feature vector of the positive sample.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a model generation method, a data presentation method, a data generation method, an estimation method, a model generation device, a data presentation device, a data generation device, and an estimation device. [Background technology]

[0002] In recent years, information processing technologies including machine learning have been utilized in materials development. This field is called materials informatics (MI) and has made a great contribution to the efficiency of new materials development. A typical method for predicting material properties by information processing is a method using first-principles calculations disclosed in Non-Patent Document 1 and the like. First-principles calculations are a method for calculating the state of electrons in a material according to the Schrödinger equation of quantum mechanics. According to first-principles calculations, the properties of a material can be predicted based on the state of electrons calculated under various conditions. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Masanori Kayama, "Current status and prospects of computational materials science: focusing on application to material interfaces", Hyomen Gijutsu, 2013, Vol. 64, No. 10, pp. 524-530. Summary of the Invention [Problem to be solved by the invention]

[0004] The present inventors have found that the conventional MI method has the following problems. That is, since the calculation of the Schrodinger equation in a real material (multi-body electron system) is extremely complicated, an approximation calculation using density functional theory or the like is used. The accuracy depends on the approximation calculation employed. With the current capabilities of general computers, it is difficult to perform a highly accurate first-principles calculation in a realistic time, so the more complicated the target material is, the more difficult it is to predict its properties. Therefore, a method is being developed in which knowledge of the properties of known materials, characteristic parts of crystal structures, etc. is given as correct answer information, a trained inference model is generated by performing machine learning, and new knowledge such as the composition and properties of a new material is obtained using the generated trained inference model. However, with such a method, it is difficult to obtain new knowledge with high accuracy within the range where correct answer information is not given. In addition, it is extremely costly to provide correct answer information for all known materials. Therefore, with a machine learning method that provides correct answer information for known materials, it is difficult to obtain new knowledge with high accuracy at low cost.

[0005] In one aspect, the present invention has been made in consideration of the above circumstances, and an object of the present invention is to provide a technique for obtaining new knowledge about materials at low cost, and a method for utilizing the same. [Means for solving the problem]

[0006] In order to solve the above-mentioned problems, the present invention employs the following configuration.

[0007] That is, a model generation method according to one aspect of the present invention is an information processing method including a step of a computer acquiring first data and second data related to a crystal structure of a material, and a step of the computer using the acquired first data and the acquired second data to perform machine learning of a first encoder and a second encoder. The second data is configured to indicate the properties of the material with an index different from that of the first data. The acquired first data and second data include a positive sample and a negative sample. The positive sample is composed of a combination of the first data and the second data for the same material. The negative sample is composed of at least one of the first data and the second data for a material different from the material of the positive sample. The first encoder is configured to convert the first data into a first feature vector, and the second encoder is configured to convert the second data into a second feature vector. The dimension of the first feature vector is the same as the dimension of the second feature vector. The machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder so that the values ​​of the first feature vector and the second feature vector calculated from the first data and the second data of the positive samples are positioned close to each other, and the value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative samples is positioned far away from the value of at least one of the first feature vector and the second feature vector calculated from the positive samples.

[0008] In the experimental example described below, trained encoders were generated by machine learning to map each of a plurality of different types of data related to crystal structures into a feature space of the same dimension. In this machine learning, each encoder was trained so that the feature vectors of various data (positive samples) of the same material were positioned close to each other in the feature space, and the feature vectors of data (negative samples) of different materials were positioned far from the feature vectors of the positive samples. Then, when various data were mapped into the feature space using each of the generated trained encoders, various data of each material having similar characteristics were mapped into a nearby range in the feature space. From the results of this experimental example, it was found that, according to each of the trained encoders generated by such machine learning, it is possible to evaluate the similarity of materials based on the positional relationship in the feature space, even without providing knowledge of the composition, characteristics, etc. of known materials, and to accurately obtain new knowledge of materials from the evaluation results.

[0009] As described above, when generating a highly accurate trained model that directly derives the properties of a material from data on the crystal structure, it takes a lot of effort to provide correct answer information for all known materials. In contrast, in the model generation method according to the present configuration, it is possible to prepare positive samples and negative samples to be used in machine learning depending on whether the materials are the same or not, and the effort required to provide correct answer information for all known materials can be omitted. Therefore, according to the model generation method according to the present configuration, trained encoders (first encoder and second encoder) that map the first data and second data, respectively, to the above-mentioned feature space can be generated at low cost. As a result, new knowledge about materials can be obtained at low cost by each of the generated trained encoders. In addition, since it is not necessary to provide correct answer information, it is possible to prepare a large amount of positive samples and negative samples to be used in machine learning at low cost. Therefore, it is possible to generate trained encoders for accurately obtaining new knowledge about materials at low cost.

[0010] The model generation method according to the above aspect may further include a step of the computer performing machine learning of a first decoder. The machine learning of the first decoder may be configured by training the first decoder such that a result of the first data being restored by the first decoder from a first feature vector calculated from the first data by using the first encoder matches the first data. With this configuration, a trained first decoder that has acquired the ability to restore the first data can be generated. By using the generated trained first decoder and trained second encoder, first data can be generated from the second data for material that is known in the second data but unknown in the first data.

[0011] The model generation method according to the above aspect may further include a step of the computer performing machine learning of a second decoder. The machine learning of the second decoder may be configured by training the second decoder such that a result of the second decoder restoring the second data from a second feature vector calculated from the second data by using the second encoder matches the second data. With this configuration, a trained second decoder that has acquired the ability to restore the second data can be generated. By using the generated trained second decoder and trained first encoder, second data can be generated from the first data for material that is known in the first data but unknown in the second data.

[0012] The model generation method according to the above aspect may further include a step in which the computer performs machine learning of an estimator. In the step of acquiring the first data and the second data, the computer may further acquire correct answer information indicating properties of the material. The machine learning of the estimator may be configured by training the estimator using the first encoder and the second encoder so that a result of estimating the properties of the material from at least one of the first feature vector and the second feature vector calculated from the acquired first data and the second data matches the correct answer information.

[0013] According to this configuration, a trained estimator for estimating the properties of a material can be generated. In this configuration, although correct answer information may be provided for all learning materials, the feature space mapped by each trained encoder contains information on the similarity of the materials. The estimator is configured to estimate the properties of the material from the feature vector in the feature space, and the information can be used when estimating the properties of the material. Therefore, even if correct answer information is not prepared for all materials, a trained estimator capable of estimating the properties of the material with high accuracy can be generated. Therefore, according to this configuration, a trained estimator capable of estimating the properties of the material with high accuracy can be generated at low cost.

[0014] In the model generation method according to the above aspect, the first data may indicate information on the local structure of the crystal of the material, and the second data may indicate information on the periodicity of the crystal structure of the material. In this configuration, data indicating the properties of the material based on a local perspective of the crystal structure is adopted as the first data. Also, data indicating the properties of the material based on an overall bird's-eye view is adopted as the second data. As a result, in the feature space mapped by the trained encoder generated, the similarity of the materials can be evaluated from both the local and bird's-eye views, and new knowledge of the material can be obtained with high accuracy from the evaluation results.

[0015] In the model generation method according to the above aspect, the first data may be composed of at least one of three-dimensional atomic position data, Raman spectroscopy data, nuclear magnetic resonance spectroscopy data, infrared spectroscopy data, mass spectroscopy data, and X-ray absorption spectroscopy data, as data indicating the properties of the material based on a local viewpoint of the crystal structure. Alternatively, the first data may be composed of three-dimensional atomic position data, and the three-dimensional atomic position data may be configured to express the state of atoms in the material by at least one of a probability density function, a probability distribution function, and a probability mass function. With these configurations, it is possible to appropriately prepare the first data indicating the properties of the material based on a local viewpoint of the crystal structure.

[0016] In the model generating method according to the above aspect, the second data may be composed of at least one of X-ray diffraction data, neutron diffraction data, electron beam diffraction data, and total scattering data as data showing the properties of the material from an overall bird's-eye view. With this configuration, the second data showing the properties of the material from an overall bird's-eye view can be appropriately prepared.

[0017] The present invention is not limited to a model generation method configured to execute the above series of information processing by a computer. One aspect of the present invention may be a data processing method using a trained machine learning model generated by the model generation method according to any of the above aspects.

[0018] For example, a data presentation method according to one aspect of the present invention is an information processing method including the steps of: acquiring at least one of first data and second data related to the crystal structure of each of a plurality of target materials by a computer; converting at least one of the acquired first data and second data of each of the target materials into at least one of a first feature vector and a second feature vector by the computer using at least one of a trained first encoder and a trained second encoder; mapping each value of at least one of the obtained first feature vector and the second feature vector of each of the target materials onto a space; and outputting each value of at least one of the first feature vector and the second feature vector of each of the target materials mapped onto the space by the computer. The trained first encoder and the trained second encoder may be generated by machine learning using the first data and the second data for learning in any of the above model generation methods.

[0019] In the data presentation method according to the above aspect, in the mapping step, the computer may convert each of the values ​​of at least one of the first feature vector and the second feature vector of each of the target materials obtained to a lower dimension so as to maintain a positional relationship between the values, and then map each of the converted values ​​in space. In the outputting step, the computer may output each of the converted values ​​of at least one of the first feature vector and the second feature vector of each of the target materials. According to this configuration, when outputting each value of a feature vector to obtain new knowledge about a material, by converting to a lower dimension so as to maintain a positional relationship between the values, it is possible to reduce the impact on information regarding the similarity of the materials and to improve the efficiency of output resources (for example, space saving in the information output range, improved visibility, etc.).

[0020] Also, for example, a data generation method according to one aspect of the present invention is an information processing method for generating second data from first data. The first data and the second data are related to a crystal structure of a target material. The second data is configured to indicate the properties of the material with an index different from that of the first data. The data generation method includes a step of acquiring first data of the target material by a computer, a step of converting the acquired first data of the target material into a first feature vector by the computer using a trained first encoder, and a step of generating second data of the target material by restoring second data from at least one of the value of the first feature vector obtained by the conversion and its neighboring values ​​by the computer using a trained decoder. The trained first encoder may be generated by machine learning using the first data and the second data for learning together with the second encoder in any of the above model generation methods. The trained decoder (second decoder) may be generated by machine learning using the second data for learning in any of the above model generation methods. The first data may indicate information on the local structure of a crystal of the target material, and the second data may indicate information on the periodicity of the crystal structure of the target material.

[0021] Also, for example, a data generation method according to one aspect of the present invention is an information processing method for generating first data from second data. The first data and the second data are related to a crystal structure of a target material. The second data is configured to indicate the properties of the material with an index different from that of the first data. The data generation method includes a step of acquiring second data of the target material by a computer, a step of converting the acquired second data of the target material into a second feature vector by the computer using a trained second encoder, and a step of generating first data of the target material by restoring the first data from at least one of the value of the second feature vector obtained by the conversion and its neighboring values ​​by the computer using a trained decoder. The trained second encoder may be generated by machine learning using the first data and second data for learning together with the first encoder in any of the above model generation methods. The trained decoder (first decoder) may be generated by machine learning using the first data for learning in any of the above model generation methods. The first data may indicate information on the local structure of a crystal of the target material, and the second data may indicate information on the periodicity of the crystal structure of the target material.

[0022] Also, for example, an estimation method according to one aspect of the present invention is an information processing method including the steps of: a computer acquiring at least one of first data and second data related to a crystal structure of a target material; a computer converting at least one of the acquired first data and second data into at least one of a first feature vector and a second feature vector using at least one of a trained first encoder and a trained second encoder; and a computer estimating a property of the target material from at least one of the obtained values ​​of the first feature vector and the second feature vector using a trained estimator. The trained first encoder and the trained second encoder may be generated by machine learning using the first data and the second data for learning in any of the above model generation methods. The trained estimator may be generated by machine learning using further ground truth information indicating the property of the learning material in any of the above model generation methods.

[0023] As another embodiment of each of the information processing methods according to the above embodiments, one aspect of the present invention may be an information processing device that realizes all or part of each of the above configurations, an information processing system, a program, or a storage medium that stores such a program and is readable by a computer or other device, machine, etc. Here, a storage medium that is readable by a computer, etc. is a medium that stores information such as a program by electrical, magnetic, optical, mechanical, or chemical action.

[0024] For example, a model generation device according to one aspect of the present invention is an information processing device including a learning data acquisition unit configured to acquire first data and second data related to a crystal structure of a material, and a machine learning unit configured to perform machine learning of a first encoder and a second encoder using the acquired first data and second data.

[0025] Also, for example, a data presentation device according to one aspect of the present invention is an information processing device including: a target data acquisition unit configured to acquire at least one of first data and second data relating to a crystal structure of each of a plurality of target materials; a conversion unit configured to acquire at least one of a first feature vector and a second feature vector by performing at least one of a process of converting the first data into a first feature vector using a trained first encoder and a process of converting the second data into a second feature vector using a trained second encoder; and an output processing unit configured to map each value of at least one of the first feature vector and the second feature vector obtained for each of the target materials onto a space, and to output each value of at least one of the first feature vector and the second feature vector for each of the target materials mapped onto the space.

[0026] Also, for example, a data generation device according to one aspect of the present invention is an information processing device configured to generate second data from first data, the data generation device including: a target data acquisition unit configured to acquire first data of a target material, a conversion unit configured to convert the acquired first data of the target material into a first feature vector using a trained first encoder, and a restoration unit configured to generate the second data of the target material by restoring second data from at least one of the value of the first feature vector obtained by conversion and values ​​near the first feature vector using a trained decoder.

[0027] Also, for example, a data generation device according to one aspect of the present invention is an information processing device configured to generate first data from second data, the data generation device includes a target data acquisition unit configured to acquire second data of a target material, a conversion unit configured to convert the acquired second data of the target material into a second feature vector using a trained second encoder, and a restoration unit configured to generate the first data of the target material by restoring the first data from at least one of the value of the second feature vector obtained by conversion and its neighboring values ​​using a trained decoder.

[0028] Also, for example, an estimation device according to one aspect of the present invention is an information processing device including: a target data acquisition unit configured to acquire at least one of first data and second data related to a crystal structure of a target material; a conversion unit configured to convert the acquired at least one of the first data and second data into at least one of a first feature vector and a second feature vector using at least one of a trained first encoder and a trained second encoder; and an estimation unit configured to estimate properties of the target material from the obtained values ​​of at least one of the first feature vector and the second feature vector using a trained estimator. Effect of the Invention

[0029] According to the present invention, it is possible to provide a technique for obtaining new knowledge about materials at low cost and a method for utilizing the technique. [Brief description of the drawings]

[0030] [Figure 1] FIG. 1 shows a schematic diagram of an example of a situation in which the present invention is applied. [Diagram 2] FIG. 2 is a schematic diagram illustrating an example of a hardware configuration of the model generating device according to the embodiment. [Diagram 3] FIG. 3 is a schematic diagram illustrating an example of a hardware configuration of a data processing device according to an embodiment. [Figure 4] FIG. 4 illustrates an example of a software configuration of the model generating device according to the embodiment. [Figure 5A] FIG. 5A illustrates an example of a machine learning process for a first decoder by the model generation device according to the embodiment. [Figure 5B] FIG. 5B illustrates an example of a machine learning process for a second decoder by the model generation device according to the embodiment. [Figure 5C] FIG. 5C illustrates an example of a process of machine learning of an estimator by the model generation device according to the embodiment. [Figure 6]FIG. 6 illustrates an example of a software configuration of the data processing device according to the embodiment. [Figure 7A] FIG. 7A illustrates an example of a process of data presentation processing by the data processing device according to the embodiment. [Figure 7B] FIG. 7B illustrates an example of a process of data generation processing by the data processing device according to the embodiment. [Figure 7C] FIG. 7C illustrates an example of a process of data generation processing by the data processing device according to the embodiment. [Figure 7D] FIG. 7D illustrates an example of the process of the estimation process performed by the data processing device according to the embodiment. [Figure 8] FIG. 8 is a flowchart illustrating an example of a processing procedure of the model generating device according to the embodiment. [Figure 9] FIG. 9 is a flowchart illustrating an example of a processing procedure related to the data presentation method of the data processing device according to the embodiment. [Figure 10A] FIG. 10A is a flowchart illustrating an example of a processing procedure related to a data generating method of the data processing device according to the embodiment. [Figure 10B] FIG. 10B is a flowchart illustrating an example of a processing procedure related to the data generating method of the data processing device according to the embodiment. [Figure 11] FIG. 11 is a flowchart illustrating an example of a processing procedure relating to the estimation method of the data processing device according to the embodiment. [Figure 12] FIG. 12 is a schematic diagram showing an example of the configuration of an encoder according to another embodiment. [Figure 13] FIG. 13 shows the results of checking the range in which elements corresponding to materials containing each element of the periodic table exist in the data distribution on the feature space created by the experimental example. [Figure 14A] FIG. 14A shows the results of coloring each element according to the value (eV) of the physical property (energy above the hull) in the data distribution on the feature space created by the experimental example. [Figure 14B]FIG. 14B shows the results of coloring each element according to the value (eV) of the physical property (band gap) in the data distribution on the feature space created by the experimental example. [Figure 14C] FIG. 14C shows the result of coloring each element according to the value (T) of the physical property (magnetization) in the data distribution on the feature space created by the experimental example. [Figure 15A] FIG. 15A shows the composition of the materials used for the queries in the experimental examples. [Figure 15B] FIG. 15B shows the compositions of materials extracted in the nearest neighborhood in the feature space by the query shown in FIG. 15A. [Figure 15C] FIG. 15C shows the composition of materials extracted as the second nearest neighbors in feature space by the query shown in FIG. 15A. [Figure 16A] FIG. 16A shows the composition of the materials used for the queries in the experimental examples. [Figure 16B] FIG. 16B shows the compositions of materials extracted in the nearest neighborhood in the feature space by the query shown in FIG. 16A. [Figure 16C] FIG. 16C shows the composition of materials extracted as the second nearest neighbors in feature space by the query shown in FIG. 16A. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0031] An embodiment according to one aspect of the present invention (hereinafter, also referred to as "the present embodiment") will be described below with reference to the drawings. However, the present embodiment described below is merely an example of the present invention in every respect. It goes without saying that various improvements and modifications can be made without departing from the scope of the present invention. In other words, in implementing the present invention, a specific configuration according to the embodiment may be appropriately adopted. Note that while the data appearing in this embodiment is described in natural language, more specifically, it is specified in pseudo-language, commands, parameters, machine language, etc. that can be recognized by a computer.

[0032] §1 Examples of application 1 is a schematic diagram showing an example of a situation in which the present invention is applied. As shown in FIG. 1, an information processing system 100 according to the present embodiment includes a model generating device 1 and a data processing device 2.

[0033] The model generating device 1 according to the present embodiment is at least one computer configured to generate a trained machine learning model. Specifically, the model generating device 1 acquires first data 31 and second data 32 related to a crystal structure of a material. The second data 32 indicates the properties of the material using an index different from that of the first data 31. As an example, the first data 31 may indicate information related to the local structure of the crystal of the material. The second data 32 may indicate information related to the periodicity of the crystal structure of the material.

[0034] The acquired first data 31 and second data 32 include positive samples and negative samples. The positive samples are composed of a combination of first data 31p and second data 32p for the same material. The negative samples are composed of at least one of first data 31n and second data 32n for a material different from that of the positive samples.

[0035] The model generating device 1 uses the acquired first data 31 and second data 32 to perform machine learning of a first encoder 51 and a second encoder 52. The first encoder 51 is a machine learning model configured to convert the first data into a first feature vector. The second encoder 52 is a machine learning model configured to convert the second data into a second feature vector. The dimension of the first feature vector is the same as the dimension of the second feature vector.

[0036] The machine learning of the first encoder 51 and the second encoder 52 is configured by training the first encoder 51 and the second encoder 52 so that the values ​​of the first feature vector 41p and the second feature vector 42p calculated from the first data 31p and the second data 32p of the positive sample are positioned close to each other, and at least one value of the first feature vector 41n and the second feature vector 42n calculated from at least one of the first data 31n and the second data 32n of the negative sample is positioned far from at least one value of the first feature vector 41p and the second feature vector 42p calculated from the positive sample. As a result of this machine learning, a trained first encoder 51 and a trained second encoder 52 are generated.

[0037] On the other hand, the data processing device 2 according to the present embodiment is at least one computer configured to execute data processing using a trained machine learning model generated by the model generation device 1. The data processing device 2 may be called, for example, a data presentation device, a data generation device, an estimation device, etc., depending on the content of the information processing to be executed. FIG. 1 shows an example of a scene in which the data processing device 2 operates as a data presentation device.

[0038] Specifically, the data processing device 2 acquires at least one of the first data 61 and the second data 62 related to the crystal structure of each of the plurality of target materials. The data processing device 2 converts at least one of the acquired first data 61 and the second data 62 of each target material into at least one of the first feature vector 71 and the second feature vector 72 using at least one of the trained first encoder 51 and the trained second encoder 52. The data processing device 2 maps each value of at least one of the acquired first feature vector 71 and the second feature vector 72 of each target material onto a space. Then, the data processing device 2 outputs each value of at least one of the first feature vector 71 and the second feature vector 72 of each target material mapped onto a space.

[0039] As described above, in this embodiment, it is possible to prepare positive samples and negative samples to be used for machine learning depending on whether the materials are the same or not. Therefore, in the model generating device 1, the trained first encoder 51 and the trained second encoder 52 can be generated at low cost. In addition, by the above machine learning, the trained first encoder 51 and the trained second encoder 52 can acquire the ability to map the first data and the second data of materials having similar characteristics to a nearby range in the feature space. As a result, in the data processing device 2, new knowledge about materials can be obtained by using at least one of the generated trained first encoder 51 and the trained second encoder 52.

[0040] In one example, as shown in Fig. 1, the model generation device 1 and the data processing device 2 may be connected to each other via a network. The type of network may be appropriately selected from, for example, the Internet, a wireless communication network, a mobile communication network, a telephone network, a dedicated network, and the like. However, the method of exchanging data between the model generation device 1 and the data processing device 2 is not limited to this example, and may be appropriately selected depending on the embodiment. In another example, data may be exchanged between the model generation device 1 and the data processing device 2 using a storage medium.

[0041] 1, the model generating device 1 and the data processing device 2 are separate computers. However, the configuration of the information processing system 100 according to this embodiment is not limited to this example and may be appropriately determined according to the embodiment. In another example, the model generating device 1 and the data processing device 2 may be an integrated computer. In yet another example, at least one of the model generating device 1 and the data processing device 2 may be configured by multiple computers.

[0042] §2 Configuration Example [Hardware configuration] <Model generation device> Fig. 2 shows a schematic diagram of an example of a hardware configuration of the model generating device 1 according to the present embodiment. As shown in Fig. 2, the model generating device 1 according to the present embodiment is a computer to which a control unit 11, a storage unit 12, a communication interface 13, an external interface 14, an input device 15, an output device 16, and a drive 17 are electrically connected. Note that in Fig. 2, the communication interface and the external interface are written as "communication I / F" and "external I / F". Similar notations are used in Fig. 3 described later.

[0043] The control unit 11 includes a hardware processor such as a CPU (Central Processing Unit), a RAM (Random Access Memory), and a ROM (Read Only Memory), and is configured to execute information processing based on programs and various data. The control unit 11 (CPU) is an example of a processor resource. The storage unit 12 is an example of a memory resource, and is configured, for example, by a hard disk drive, a solid state drive, or the like. In this embodiment, the storage unit 12 stores various information such as a model generation program 81, first data 31, second data 32, learning result data 125, and the like.

[0044] The model generation program 81 is a program for causing the model generation device 1 to execute information processing (FIG. 8 described later) for generating a trained machine learning model. The model generation program 81 includes a series of instructions for the information processing. The first data 31 and the second data 32 are used for machine learning. The learning result data 125 indicates information on the trained machine learning model generated by machine learning. In this embodiment, the learning result data 125 is generated as a result of executing the model generation program 81.

[0045] The communication interface 13 is, for example, a wired LAN (Local Area Network) module, a wireless LAN module, etc., and is an interface for performing wired or wireless communication via a network. The model generation device 1 may perform data communication with other computers via the communication interface 13.

[0046] The external interface 14 is, for example, a USB (Universal Serial Bus) port, a dedicated port, etc., and is an interface for connecting to an external device. The type and number of the external interfaces 14 may be selected arbitrarily. The model generation device 1 may be connected to a device for obtaining each data (31, 32) via the communication interface 13 or the external interface 14.

[0047] The input device 15 is a device for performing input, such as a mouse, a keyboard, etc. The output device 16 is a device for performing output, such as a display, a speaker, etc. An operator can operate the model generation device 1 by using the input device 15 and the output device 16. The input device 15 and the output device 16 may be integrally configured, for example, by a touch panel display, etc.

[0048] The drive 17 is, for example, a CD drive, a DVD drive, or the like, and is a drive device for reading various information such as a program stored in a storage medium 91. At least one of the model generation program 81, the first data 31, and the second data 32 may be stored in the storage medium 91.

[0049] The storage medium 91 is a medium that accumulates various information such as a program by electrical, magnetic, optical, mechanical, or chemical action so that a computer or other device, machine, etc. can read the stored information such as the program. The model generation device 1 may obtain at least one of the model generation program 81, the first data 31, and the second data 32 from the storage medium 91.

[0050] 2 illustrates a disk-type storage medium such as a CD or a DVD as an example of the storage medium 91. However, the type of the storage medium 91 is not limited to the disk type, and may be other than the disk type. An example of a storage medium other than the disk type is a semiconductor memory such as a flash memory. The type of the drive 17 may be appropriately selected depending on the type of the storage medium 91.

[0051] Regarding the specific hardware configuration of the model generating device 1, components can be omitted, replaced, and added as appropriate depending on the embodiment. For example, the control unit 11 may include multiple hardware processors. The hardware processor may be configured with a microprocessor, a field-programmable gate array (FPGA), a digital signal processor (DSP), or the like. The storage unit 12 may be configured with a RAM and a ROM included in the control unit 11. At least one of the communication interface 13, the external interface 14, the input device 15, the output device 16, and the drive 17 may be omitted. The model generating device 1 may be configured with multiple computers. In this case, the hardware configurations of the computers may or may not be the same. In addition, the model generating device 1 may be an information processing device designed specifically for the service provided, a general-purpose server device, a general-purpose PC (Personal Computer), or the like.

[0052] <Data processing device> Fig. 3 is a schematic diagram showing an example of a hardware configuration of the data processing device 2 according to the present embodiment. As shown in Fig. 3, the data processing device 2 according to the present embodiment is a computer in which a control unit 21, a storage unit 22, a communication interface 23, an external interface 24, an input device 25, an output device 26, and a drive 27 are electrically connected.

[0053] The control unit 21 to the drive 27 and the storage medium 92 of the data processing device 2 may be configured similarly to the control unit 11 to the drive 17 and the storage medium 91 of the model generating device 1. The control unit 21 includes a CPU, RAM, ROM, etc., which are hardware processors, and is configured to execute various information processing based on programs and data. The control unit 21 (CPU) is an example of a processor resource. The storage unit 22 is an example of a memory resource, and is configured, for example, by a hard disk drive, a solid state drive, etc. In this embodiment, the storage unit 22 stores various information such as a data processing program 82, learning result data 125, etc.

[0054] The data processing program 82 is a program for causing the data processing device 2 to execute information processing (FIGS. 9 to 11 described below) on data related to the crystal structure of a target material using a trained machine learning model. The data processing program 82 includes a series of instructions for the information processing. At least one of the data processing program 82 and the learning result data 125 may be stored in the storage medium 92. The data processing device 2 may acquire at least one of the data processing program 82 and the learning result data 125 from the storage medium 92.

[0055] The data processing device 2 may perform data communication with other computers via the communication interface 23. The data processing device 2 may be connected to a device for obtaining the first data or the second data via the communication interface 23 or the external interface 24. The data processing device 2 may receive operations and inputs from an operator by using the input device 25 and the output device 26.

[0056] Regarding the specific hardware configuration of the data processing device 2, components can be omitted, replaced, and added as appropriate depending on the embodiment. For example, the control unit 21 may include multiple hardware processors. The hardware processor may be configured with a microprocessor, FPGA, DSP, etc. The storage unit 22 may be configured with a RAM and a ROM included in the control unit 21. At least one of the communication interface 23, the external interface 24, the input device 25, the output device 26, and the drive 27 may be omitted. The data processing device 2 may be configured with multiple computers. In this case, the hardware configurations of the computers may or may not be the same. Furthermore, the data processing device 2 may be an information processing device designed exclusively for the service provided, as well as a general-purpose server device, a general-purpose PC, etc.

[0057] [Software configuration] <Model generation device> 4 is a schematic diagram showing an example of a software configuration of the model generating device 1 according to the present embodiment. The control unit 11 of the model generating device 1 loads the model generating program 81 stored in the storage unit 12 in the RAM. Then, the control unit 11 executes instructions included in the model generating program 81 loaded in the RAM by the CPU. As a result, as shown in FIG. 4, the model generating device 1 according to the present embodiment operates as a computer including a learning data acquiring unit 111, a machine learning unit 112, and a storage processing unit 113 as software modules. That is, in the present embodiment, each software module of the model generating device 1 is realized by the control unit 11 (CPU).

[0058] The learning data acquisition unit 111 is configured to acquire first data 31 and second data 32 for learning. The first data 31 and second data 32 relate to the crystal structure of a material and indicate the properties of the material with different indices. The acquired first data 31 and second data 32 include a plurality of positive samples and a plurality of negative samples. Each position sample is composed of a combination of first data 31p and second data 32p for the same material. Each negative sample is composed of at least one of first data 31n and second data 32n for a material different from the material of the corresponding positive sample (any of the plurality of positive samples).

[0059] The machine learning unit 112 is configured to perform machine learning of the first encoder 51 and the second encoder 52 using the acquired first data 31 and second data 32. The first encoder 51 is configured to convert the first data into a first feature vector. The second encoder 52 is configured to convert the second data into a second feature vector of the same dimension as the dimension of the first feature vector. That is, each encoder (51, 52) is configured to map the first data and the second data, respectively, into a feature space of the same dimension.

[0060] The machine learning of the first encoder 51 and the second encoder 52 is configured by training the first encoder 51 and the second encoder 52 so that the values ​​of the first feature vector 41p and the second feature vector 42p calculated from the first data 31p and the second data 32p of each positive sample are positioned close to each other, and the value of at least one of the first feature vector 41n and the second feature vector 42n calculated from at least one of the first data 31n and the second data 32n of each negative sample is positioned far away from the value of at least one of the first feature vector 41p and the second feature vector 42p calculated from the corresponding positive sample.

[0061] That is, in the machine learning, the first encoder 51 and the second encoder 52 are trained so that the first distance between the feature vectors (41p, 42p) of the positive samples is relatively shorter than the second distance between the feature vectors of the corresponding negative samples. This training may be configured by at least one of adjusting the first distance to be smaller and adjusting the second distance to be larger. Note that the second distance may be configured by at least one of the distance between the first feature vectors (41p, 41n) of the corresponding positive samples and negative samples, the distance between the first feature vector 41p and the second feature vector 42n, the distance between the second feature vector 42p and the first feature vector 41n, and the distance between the second feature vectors (42p, 42n). The first feature vectors (41p, 41n) are calculated from the first data (31p, 31n) using the first encoder 51. A second feature vector (42p, 42n) is calculated from the second data (32p, 32n) using the second encoder 52. As a result of the machine learning, a trained first encoder 51 and a trained second encoder 52 are generated.

[0062] 5A to 5C, the model generating device 1 according to the present embodiment may be configured to further generate at least one of a trained first decoder 55, a trained second decoder 56, and a trained estimator 58. The first decoder 55 corresponds to the first encoder 51 and is configured to restore first data from a first feature vector. The second decoder 56 corresponds to the second encoder 52 and is configured to restore second data from a second feature vector. The estimator 58 is configured to estimate material properties from at least one of the first feature vector and the second feature vector.

[0063] 5A is a schematic diagram showing an example of a machine learning process of the first decoder 55 by the model generating device 1 according to the present embodiment. When the model generating device 1 is configured to generate a trained first decoder 55, the machine learning unit 112 may be configured to further perform machine learning of the first decoder 55 using the first data 31. The machine learning of the first decoder 55 is configured by training the first decoder 55 so that a result of restoring the first data 31 by the first decoder 55 from a first feature vector calculated from the first data 31 using the first encoder 51 matches the first data 31. As a result of this machine learning, a trained first decoder 55 can be generated.

[0064] 5B is a schematic diagram showing an example of a machine learning process of the second decoder 56 by the model generating device 1 according to the present embodiment. When the model generating device 1 is configured to generate a trained second decoder 56, the machine learning unit 112 may be configured to further perform machine learning of the second decoder 56 using the second data 32. The machine learning of the second decoder 56 is configured by training the second decoder 56 so that a result of restoring the second data 32 by the second decoder 56 from a second feature vector calculated from the second data 32 using the second encoder 52 matches the second data 32. As a result of this machine learning, a trained second decoder 56 can be generated.

[0065] FIG. 5C is a schematic diagram showing an example of a process of machine learning of the estimator 58 by the model generating device 1 according to the present embodiment. When the model generating device 1 is configured to generate a trained estimator 58, the learning data acquiring unit 111 may be configured to further acquire correct answer information (correct answer label) 35 indicating the property (true value) of the material. The machine learning unit 112 may be configured to further perform machine learning of the estimator 58 using the correct answer information 35 and at least one of the first data 31 and the second data 32. The machine learning of the estimator 58 is configured by training the estimator 58 so that a result of estimating the property of the material by the estimator 58 from at least one of the first feature vector calculated from the first data 31 by using the first encoder 51 and the second feature vector calculated from the second data 32 by using the second encoder 52 matches the corresponding correct answer information 35. As a result of this machine learning, a trained estimator 58 can be generated.

[0066] 4 and 5A to 5C, the storage processing unit 113 is configured to generate information about a trained machine learning model generated by machine learning (in this embodiment, the first encoder 51, the second encoder 52, the first decoder 55, the second decoder 56, and the estimator 58) as learning result data 125, and to store the generated learning result data 125 in an arbitrary storage area. The learning result data 125 may be appropriately configured to include information for reproducing the trained machine learning model.

[0067] (An example of a machine learning model) In this embodiment, the first encoder 51, the second encoder 52, the first decoder 55, the second decoder 56, and the estimator 58 are configured by a machine learning model having one or more calculation parameters used for each calculation. As long as each of the above calculations can be executed, the type and structure of the machine learning model adopted for each of them may not be particularly limited and may be appropriately selected according to the embodiment. As an example, each of the first encoder 51, the second encoder 52, the first decoder 55, and the second decoder 56 may be configured by a neural network or the like. The estimator 58 may be configured by a neural network, a support vector machine, a regression model, a decision tree model, or the like.

[0068] Training consists of adjusting (optimizing) values ​​of calculation parameters so as to derive an output that matches the training data (first data 31 / second data 32) from the training data. The machine learning method may be appropriately selected depending on the type of machine learning model to be adopted. As an example, the machine learning method may be a method such as backpropagation, solving an optimization problem, or performing regression analysis.

[0069] When a neural network is employed, typically, the first encoder 51, the second encoder 52, the first decoder 55, the second decoder 56, and the estimator 58 are configured to have an input layer, one or more intermediate layers (hidden layers), and an output layer. For each layer, any type of layer, such as a fully connected layer, may be employed. The number of layers included in each, the type of each layer, the number of nodes (neurons) in each layer, and the connection relationship of the nodes may be appropriately determined according to the embodiment. The weight of the connection between each node, the threshold value of each node, etc. are examples of the above-mentioned calculation parameters. In the following, an example of a training process in the case where a neural network is employed for each of the first encoder 51, the second encoder 52, the first decoder 55, the second decoder 56, and the estimator 58 will be described.

[0070] (A) Encoder training As shown in FIG. 4, as an example of a training process in the case where each encoder (51, 52) is configured by a neural network, the machine learning unit 112 inputs the first data 31p of each positive sample to the first encoder 51 and executes a forward propagation arithmetic process of the first encoder 51. As a result of this arithmetic process, the machine learning unit 112 acquires a first feature vector 41p corresponding to the first data 31p of each positive sample from the first encoder 51. Similarly, the machine learning unit 112 inputs the second data 32p of each positive sample to the second encoder 52 and executes a forward propagation arithmetic process of the second encoder 52. As a result of this arithmetic process, the machine learning unit 112 acquires a second feature vector 42p corresponding to the second data 32p of each positive sample from the second encoder 52.

[0071] Furthermore, when the first data 31n is included in the negative sample corresponding to each positive sample, the machine learning unit 112 inputs the first data 31n of the corresponding negative sample to the first encoder 51, and executes the forward propagation arithmetic process of the first encoder 51. As a result of this arithmetic process, the machine learning unit 112 acquires a first feature vector 41n corresponding to the first data 31n from the first encoder 51. Similarly, when the second data 32n is included in the negative sample corresponding to each positive sample, the machine learning unit 112 inputs the second data 32n of the corresponding negative sample to the second encoder 52, and executes the forward propagation arithmetic process of the second encoder 52. As a result of this arithmetic process, the machine learning unit 112 acquires a second feature vector 42n corresponding to the second data 32n from the second encoder 52.

[0072] The machine learning unit 112 calculates an error from the calculated values ​​of each feature vector so as to achieve at least one of the following operations: reducing the first distance (bringing the vector values ​​of the positive samples closer to each other) and increasing the second distance (moving the vector values ​​between the positive samples and the negative samples farther apart). Any loss function may be used to calculate the error as long as it is possible to achieve at least one of the operations: reducing the first distance and increasing the second distance. Examples of loss functions that can achieve the operations include Triplet Loss, Contrastive Loss, Lifted Structure Loss, N-Pair Loss, Angular Loss, and Divergence Loss.

[0073] The machine learning unit 112 calculates the gradient of the calculated error. Next, the machine learning unit 112 backpropagates the gradient of the calculated error by the error backpropagation method, thereby calculating the error in the values ​​of the calculation parameters of the first encoder 51 and the second encoder 52. Then, the machine learning unit 112 updates the values ​​of the calculation parameters based on the calculated error.

[0074] Through this series of update processes, the machine learning unit 112 adjusts the values ​​of the calculation parameters of the first encoder 51 and the second encoder 52 so that the first distance between the feature vectors (41p, 42p) of each positive sample is shorter than the second distance between the feature vector of each positive sample and the feature vector of the corresponding negative sample. The adjustment of the values ​​of the calculation parameters may be repeated until a predetermined condition is satisfied, such as, for example, a specified number of times or the sum of the calculated errors satisfies a predetermined index. Furthermore, the machine learning conditions, such as the learning rate, may be appropriately set according to the embodiment. Through this machine learning process, it is possible to generate a trained first encoder 51 and a trained second encoder 52 that have acquired the ability to map the first data and the second data of the same material to a nearby position in the feature space and the first data and the second data of different materials to a distant position.

[0075] (B) Training the first decoder 5A, as an example of a training process in the case where the first decoder 55 is configured by a neural network, the machine learning unit 112 inputs each of the first data 31 to the first encoder 51 and executes a forward propagation arithmetic process of the first encoder 51. As a result of this arithmetic process, the machine learning unit 112 acquires a first feature vector corresponding to each of the first data 31 from the first encoder 51. The machine learning unit 112 inputs each of the acquired first feature vectors to the first decoder 55 and executes a forward propagation arithmetic process of the first decoder 55. As a result of this arithmetic process, the machine learning unit 112 acquires an output value corresponding to a result of restoring the first data 31 from each of the first feature vectors from the first decoder 55.

[0076] The machine learning unit 112 calculates an error between the acquired output value and the corresponding first data 31, and further calculates a gradient of the calculated error. The machine learning unit 112 backpropagates the gradient of the calculated error by the error backpropagation method, thereby calculating an error in the value of the calculation parameter of the first decoder 55. Then, the machine learning unit 112 updates the value of the calculation parameter of the first decoder 55 based on the calculated error.

[0077] Through this series of update processes, the machine learning unit 112 adjusts the values ​​of the calculation parameters of the first decoder 55 so that the sum of errors between the restoration result (output value) and the true value (corresponding first data 31) for each first data 31 is reduced. This adjustment of the values ​​of the calculation parameters may be repeated until a predetermined condition is satisfied, such as, for example, performing the adjustment a specified number of times or the sum of the calculated errors being equal to or less than a threshold value. Furthermore, machine learning conditions such as a loss function and a learning rate may be appropriately set according to the embodiment. Through this machine learning process, a trained first decoder 55 that has acquired the ability to restore the corresponding first data from the first feature vector obtained by the first encoder 51 can be generated.

[0078] In addition, as long as it is possible to generate a trained first decoder 55 that has acquired the ability to restore the first data from the first feature vector, the timing of executing the machine learning of the first decoder 55 is not particularly limited and may be appropriately selected according to the embodiment. In one example, the machine learning of the first decoder 55 may be executed after the machine learning of the first encoder 51 and the second encoder 52. In this case, the trained first encoder 51 may be used for the machine learning of the first decoder 55. In another example, the machine learning of the first decoder 55 may be executed simultaneously with the machine learning of the first encoder 51 and the second encoder 52. In this case, the machine learning unit 112 may back-propagate the gradient of the error in the machine learning of the first decoder 55 to the first encoder 51 and calculate the error of the value of the calculation parameter of the first encoder 51. Then, the machine learning unit 112 may update the value of the calculation parameter of the first encoder 51 together with the first decoder 55 based on the calculated error.

[0079] (C) Training the second decoder 5B, as an example of a training process in the case where the second decoder 56 is configured by a neural network, the machine learning unit 112 inputs each second data 32 to the second encoder 52 and executes a forward propagation arithmetic process of the second encoder 52. As a result of this arithmetic process, the machine learning unit 112 acquires a second feature vector corresponding to each second data 32 from the second encoder 52. The machine learning unit 112 inputs each of the acquired second feature vectors to the second decoder 56 and executes a forward propagation arithmetic process of the second decoder 56. As a result of this arithmetic process, the machine learning unit 112 acquires an output value corresponding to a result of restoring the second data 32 from each second feature vector from the second decoder 56.

[0080] The machine learning unit 112 calculates an error between the acquired output value and the corresponding second data 32, and further calculates a gradient of the calculated error. The machine learning unit 112 backpropagates the gradient of the calculated error by the error backpropagation method, thereby calculating an error in the value of the calculation parameter of the second decoder 56. Then, the machine learning unit 112 updates the value of the calculation parameter of the second decoder 56 based on the calculated error.

[0081] Through this series of update processes, the machine learning unit 112 adjusts the values ​​of the calculation parameters of the second decoder 56 so that the sum of errors between the restoration result (output value) and the true value (corresponding second data 32) for each second data 32 is reduced. This adjustment of the values ​​of the calculation parameters may be repeated until a predetermined condition is satisfied, such as, for example, performing the adjustment a specified number of times or the calculated sum of errors being equal to or less than a threshold value. In addition, machine learning conditions such as a loss function and a learning rate may be appropriately set according to the embodiment. Through this machine learning process, a trained second decoder 56 that has acquired the ability to restore the corresponding second data from the second feature vector obtained by the second encoder 52 can be generated.

[0082] In addition, as long as it is possible to generate a trained second decoder 56 that has acquired the ability to restore the second data from the second feature vector, the timing of executing the machine learning of the second decoder 56 is not particularly limited and may be appropriately selected according to the embodiment. In one example, the machine learning of the second decoder 56 may be executed after the machine learning of the first encoder 51 and the second encoder 52. In this case, the trained second encoder 52 may be used for the machine learning of the second decoder 56. In another example, the machine learning of the second decoder 56 may be executed simultaneously with the machine learning of the first encoder 51 and the second encoder 52. In this case, the machine learning unit 112 may back-propagate the gradient of the error in the machine learning of the second decoder 56 to the second encoder 52 as well, and may also calculate the error of the value of the calculation parameter of the second encoder 52. Then, the machine learning unit 112 may update the value of the calculation parameter of the second encoder 52 together with the second decoder 56 based on the calculated error.

[0083] Also, in one example, the machine learning of the second decoder 56 may be executed in parallel with the machine learning of the first decoder 55. In another example, the machine learning of the second decoder 56 may be executed separately from the machine learning of the first decoder 55. In this case, the machine learning process executed first may be either the first decoder 55 or the second decoder 56.

[0084] (D) Training the estimator 5C, a plurality of data sets each composed of a combination of at least one of the first data 31 and the second data 32 and the corresponding material correct answer information 35 are used for the machine learning of the estimator 58. An example of a training process in the case where the estimator 58 is composed of a neural network is shown below.

[0085] When training the estimator 58 to estimate material properties from the first feature vector, the machine learning unit 112 inputs the first data 31 of each data set to the first encoder 51 and executes a forward propagation arithmetic process of the first encoder 51. As a result of this arithmetic process, the machine learning unit 112 obtains a first feature vector corresponding to each first data 31 from the first encoder 51. The machine learning unit 112 inputs each obtained first feature vector to the estimator 58 and executes a forward propagation arithmetic process of the estimator 58. As a result of this arithmetic process, the machine learning unit 112 obtains an output value corresponding to a result of estimating the property of each material from the estimator 58.

[0086] When training the estimator 58 to estimate material properties from the second feature vectors, the machine learning unit 112 inputs the second data 32 of each data set to the second encoder 52 and executes a forward propagation arithmetic process of the second encoder 52. As a result of this arithmetic process, the machine learning unit 112 obtains a second feature vector corresponding to each second data 32 from the second encoder 52. The machine learning unit 112 inputs each obtained second feature vector to the estimator 58 and executes a forward propagation arithmetic process of the estimator 58. As a result of this arithmetic process, the machine learning unit 112 obtains an output value corresponding to a result of estimating the property of each material from the estimator 58.

[0087] The estimator 58 may be configured to accept both the first feature vector and the second feature vector as input, or may be configured to accept only one of the first feature vector and the second feature vector as input. When configured to accept both the first feature vector and the second feature vector as input, the machine learning unit 112 inputs the first feature vector and the second feature vector derived from the first data 31 and the second data 32 of the same material to the estimator 58, and obtains from the estimator 58 an output value corresponding to the result of estimating the properties of the material.

[0088] Next, the machine learning unit 112 calculates an error between the acquired output value and a true value indicated by the corresponding correct answer information 35, and further calculates a gradient of the calculated error. The machine learning unit 112 backpropagates the gradient of the calculated error by the error backpropagation method, thereby calculating an error in the value of the calculation parameter of the estimator 58. Then, the machine learning unit 112 updates the value of the calculation parameter of the estimator 58 based on the calculated error.

[0089] Through this series of update processes, the machine learning unit 112 adjusts the values ​​of the calculation parameters of the estimator 58 so that the sum of errors between the output value of the estimation result derived from at least one of the first data 31 and the second data 32 and the true value indicated by the corresponding correct answer information 35 for each data set is reduced. This adjustment of the values ​​of the calculation parameters may be repeated until a predetermined condition is satisfied, such as, for example, performing the process a specified number of times or the sum of the calculated errors being equal to or less than a threshold value. In addition, machine learning conditions such as a loss function and a learning rate may be appropriately set according to the embodiment. Through this machine learning process, a trained estimator 58 that has acquired the ability to estimate the properties of a material from at least one of the first feature vector and the second feature vector can be generated.

[0090] In addition, as long as it is possible to generate a trained estimator 58 that has acquired the ability to estimate the properties of a material, the timing of executing the machine learning of the estimator 58 is not particularly limited and may be appropriately selected according to the embodiment. In one example, the machine learning of the estimator 58 may be executed after the machine learning of the first encoder 51 and the second encoder 52. In this case, when training to estimate the properties of a material from the first feature vector, the trained first encoder 51 may be used for the machine learning of the estimator 58. When training to estimate the properties of a material from the second feature vector, the trained second encoder 52 may be used for the machine learning of the estimator 58. In another example, the machine learning of the estimator 58 may be executed simultaneously with the machine learning of the first encoder 51 and the second encoder 52. In this case, when training to estimate the properties of a material from the first feature vector, the machine learning unit 112 may back-propagate the gradient of the error in the machine learning of the estimator 58 to the first encoder 51 as well, and may also calculate the error of the value of the calculation parameter of the first encoder 51. Then, the machine learning unit 112 may update the values ​​of the calculation parameters of the first encoder 51 together with the estimator 58 based on the calculated error. Furthermore, when training to estimate material properties from the second feature vector, the machine learning unit 112 may back-propagate the gradient of the error in the machine learning of the estimator 58 to the second encoder 52 as well, and calculate the error of the values ​​of the calculation parameters of the second encoder 52. Then, the machine learning unit 112 may update the values ​​of the calculation parameters of the second encoder 52 together with the estimator 58 based on the calculated error.

[0091] Also, in one example, the machine learning of the estimator 58 may be executed simultaneously with at least one of the machine learning of the first decoder 55 and the second decoder 56. In another example, the machine learning of the estimator 58 may be executed separately from the machine learning of the first decoder 55 and the second decoder 56. In this case, the machine learning executed first may be either the estimator 58 or each of the decoders (55, 56).

[0092] In another example, the estimator 58 may be configured by a machine learning model other than a neural network, such as a support vector machine or a regression model. In this case, the machine learning of the estimator 58 is configured by adjusting the values ​​of the calculation parameters of the estimator 58 so that the output value of the estimation result derived from at least one of the first data 31 and the second data 32 for each data set approaches (for example, matches) the true value indicated by the corresponding correct answer information 35. A method for adjusting the values ​​of the calculation parameters of the estimator 58 may be appropriately selected depending on the machine learning model to be adopted. As an example, a method such as solving an optimization problem or performing a regression analysis may be adopted as a method for adjusting the values ​​of the calculation parameters of the estimator 58.

[0093] (Preservation Treatment) The storage processing unit 113 stores the trained machine learning model (the first encoder 51, the second encoder 52, the first decoder 55, the second decoder 56, and the estimator 58) generated by each of the above machine learning processes as the learning result data 125. As long as the information for executing the above calculation of the trained machine learning model can be held, the configuration of the learning result data 125 is not particularly limited and may be appropriately determined according to the embodiment. As an example, the learning result data 125 may be configured to include information indicating the configuration of the machine learning model (for example, the structure of a neural network, etc.) and the values ​​of the calculation parameters adjusted by the above machine learning. The learning result data 125 may be stored in any storage area. The learning result data 125 may be appropriately referred to in order to set the trained machine learning model in a usable state on a computer.

[0094] In the example of FIG. 4 and FIG. 5A to FIG. 5C, for convenience of explanation, information on all of the first encoder 51, the second encoder 52, the first decoder 55, the second decoder 56, and the estimator 58 is included in the learning result data 125. However, the format for holding the learning result is not limited to such an example. Information on at least any of the first encoder 51, the second encoder 52, the first decoder 55, the second decoder 56, and the estimator 58 may be held as separate learning result data. In another example, independent learning result data may be generated for each of the first encoder 51, the second encoder 52, the first decoder 55, the second decoder 56, and the estimator 58.

[0095] <Data processing device> 6 is a schematic diagram showing an example of a software configuration of the data processing device 2 according to the present embodiment. The control unit 21 of the data processing device 2 loads the data processing program 82 stored in the storage unit 22 in the RAM. Then, the control unit 21 executes instructions included in the data processing program 82 loaded in the RAM by the CPU. As a result, as shown in FIG. 6, the data processing device 2 according to the present embodiment operates as a computer including a target data acquisition unit 211, a conversion unit 212, a restoration unit 213, an estimation unit 214, and an output processing unit 215 as software modules. That is, in the present embodiment, each software module of the data processing device 2 is realized by the control unit 21 (CPU) in the same manner as the model generating device 1.

[0096] By providing at least one of the trained first encoder 51 and the trained second encoder 52 generated by the model generating device 1, a data presentation device that presents the value of a feature vector calculated from at least one of the first data and the second data can be configured. By providing the trained first encoder 51 and the trained second decoder 56, a data generation device that generates second data from the first data can be configured. By providing the trained second encoder 52 and the trained first decoder 55, a data generation device that generates first data from the second data can be configured. By providing at least one of the trained first encoder 51 and the trained second encoder 52 and the trained estimator 58, an estimation device that estimates material properties from at least one of the first data and the second data can be configured. FIG. 6 shows an example in which the data processing device 2 is configured to be able to execute the operations of all the devices.

[0097] (A) Data presentation device FIG. 7A shows an example of the process of the data presentation process (that is, a scene in which the data presentation device 2 operates as a data presentation device).

[0098] In this case, the target data acquisition unit 211 is configured to acquire at least one of the first data 61 and the second data 62 related to the crystal structure of each of the multiple target materials. The conversion unit 212 includes at least one of the trained first encoder 51 and the trained second encoder 52 by holding the learning result data 125. The conversion unit 212 is configured to acquire at least one of the first feature vector 71 and the second feature vector 72 by performing at least one of a process of converting the first data 61 of each target material acquired using the trained first encoder 51 into a first feature vector 71 and a process of converting the second data 62 of each target material acquired using the trained second encoder 52 into a second feature vector 72.

[0099] The output processing unit 215 is configured to map at least one of the first and second feature vectors 71 and 72 of each of the target materials obtained onto the space VS, and output at least one of the first and second feature vectors 71 and 72 of each of the target materials mapped onto the space VS. In one example, the output processing unit 215 may be configured to map at least one of the first and second feature vectors 71 and 72 of each of the target materials obtained onto the space VS as is. In another example, the output processing unit 215 may be configured to convert at least one of the first and second feature vectors 71 and 72 of each of the target materials obtained into a lower dimension than the original dimension in the mapping process so as to maintain the positional relationship of the values, and then map each of the converted values ​​onto the space VS. In this case, the output processing unit 215 may be configured to output at least one of the converted values ​​of the first and second feature vectors 71 and 72 of each of the target materials in the process of outputting each value. This makes it possible to reduce the impact on information regarding the similarity of each target material while improving the efficiency of output resources (for example, reducing the space required for information output, improving visibility, etc.).

[0100] The data processing device 2 may be configured to present, in the space VS, both the first feature vector 71 and the second feature vector 72. Alternatively, the data processing device 2 may be configured to present, in the space VS, only one of the first feature vector 71 and the second feature vector 72.

[0101] (B) A data generating device for generating second data from first data FIG. 7B illustrates an example of a process for generating second data 64 from first data 63 (that is, a scene in which the data processing device 2 operates as a data generating device that generates second data from the first data).

[0102] In this case, the target data acquisition unit 211 is configured to acquire the first data 63 of the target material. The conversion unit 212 includes a trained first encoder 51 by holding the learning result data 125. The conversion unit 212 is configured to convert the acquired first data 63 of the target material into a first feature vector 73 by using the trained first encoder 51. The restoration unit 213 includes a trained second decoder 56 by holding the learning result data 125. The restoration unit 213 is configured to generate the second data 64 by restoring the second data 64 from at least one of the value of the first feature vector 73 obtained by the conversion and its neighboring values ​​by using the trained second decoder 56. The output processing unit 215 is configured to output the generated second data 64.

[0103] (C) A data generating device for generating the first data from the second data FIG. 7C illustrates an example of a process for generating first data 66 from second data 65 (that is, a scene in which the data processing device 2 operates as a data generating device that generates the first data from the second data).

[0104] In this case, the target data acquisition unit 211 is configured to acquire the second data 65 of the target material. The conversion unit 212 includes a trained second encoder 52 by holding the learning result data 125. The conversion unit 212 is configured to convert the acquired second data 65 of the target material into a second feature vector 75 using the trained second encoder 52. The restoration unit 213 includes a trained first decoder 55 by holding the learning result data 125. The restoration unit 213 is configured to generate the first data 66 by restoring the first data 66 from at least one of the value of the second feature vector 75 obtained by the conversion and its neighboring values ​​using the trained first decoder 55. The output processing unit 215 is configured to output the generated first data 66.

[0105] (D) Estimation device FIG. 7D illustrates an example of a process for estimating a property of a target material from at least one of the first feature vector and the second feature vector (that is, a scene in which the data processing device 2 operates as an estimation device).

[0106] In this case, the target data acquisition unit 211 is configured to acquire at least one of the first data 67 and the second data 68 related to the crystal structure of the target material. The conversion unit 212 includes at least one of the trained first encoder 51 and the trained second encoder 52 by holding the learning result data 125. When the data processing device 2 is configured to estimate the characteristics of the target material from the first feature vector, the conversion unit 212 is configured to include the trained first encoder 51. When the data processing device 2 is configured to estimate the characteristics of the target material from the second feature vector, the conversion unit 212 is configured to include the trained second encoder 52. The conversion unit 212 is configured to convert at least one of the acquired first data 67 and the second data 68 into at least one of the first feature vector 77 and the second feature vector 78 using at least one of the trained first encoder 51 and the trained second encoder 52. The estimation unit 214 includes a trained estimator 58 by holding the learning result data 125. The estimation unit 214 is configured to estimate the property of the target material from at least one of the values ​​of the obtained first feature vector 77 and second feature vector 78, using the trained estimator 58. The output processing unit 215 is configured to output the result of estimating the property of the target material.

[0107] <Each data> The first data (31, 61, 63, 66, 67) and the second data (32, 62, 64, 65, 68) are configured to indicate information on the crystal structure of the material. The first data 31 and the second data 32 are used for machine learning and relate to the learning material. The first data (61, 63, 67) and the second data (62, 65, 68) are used for each inference process such as the data presentation and relate to the material (target material) that is the target of each inference process. The material is a substance that has a structure in which atoms or molecules are arranged (and thereby exhibits a function). As long as the first data and the second data can be obtained, it does not matter whether the material actually exists or is a virtual substance on a computer. The first data (31, 61, 63, 67) and the second data (32, 62, 65, 68) may be obtained by actual measurement or may be obtained by simulation.

[0108] The first data (31, 61, 63, 66, 67) and the second data (32, 62, 64, 65, 68) indicate the properties of the material using different indices. The types of the data may be appropriately selected according to the embodiment. As an example, the first data (31, 61, 63, 66, 67) may indicate the properties of the material based on a local viewpoint of the crystal structure. As a specific example, the first data (31, 61, 63, 66, 67) may indicate information on the local structure of the crystal of the material. The second data (32, 62, 64, 65, 68) may indicate the properties of the material based on an overall bird's-eye view. As a specific example, the second data (32, 62, 64, 65, 68) may indicate information on the periodicity of the crystal structure of the material. The periodicity of the crystal structure may be expressed by the presence or absence of periodicity, the state of periodicity (the state of the periodic characteristics exhibited by the crystal structure), etc. The material may be periodic or non-periodic.

[0109] As an example of data showing information on the local structure, the first data (31, 61, 63, 66, 67) may be composed of at least one of three-dimensional atomic position data, Raman spectroscopy data, nuclear magnetic resonance spectroscopy data, infrared spectroscopy data, mass spectrometry data, and X-ray absorption spectroscopy data. When the first data (31, 61, 63, 66, 67) is configured to include three-dimensional atomic position data, the three-dimensional atomic position data may be configured to express the state of atoms in the material (e.g., position, type, etc.) by at least one of a probability density function, a probability distribution function, and a probability mass function. That is, in the three-dimensional atomic position data, the probability regarding the state of atoms, such as the probability that a target atom exists at a target position, the probability that an atom of a target type is included, etc., may be represented by at least one of a probability density function, a probability distribution function, and a probability mass function. According to these configurations, it is possible to appropriately prepare first data showing the characteristics of a material based on a local viewpoint of a crystal structure.

[0110] As an example of data indicating information regarding periodicity, the second data (32, 62, 64, 65, 68) may be composed of at least one of X-ray diffraction data, neutron diffraction data, electron beam diffraction data, and total scattering data. This makes it possible to appropriately prepare second data indicating the properties of the material based on an overall bird's-eye view.

[0111] Each feature vector is a sequence of numbers of a fixed length (for example, a length of about 10 to 1000) that is easily handled by a computer and is generated by each encoder (51, 52). Each feature vector is often configured so that it is difficult for a human to directly understand its meaning. Basically, one feature vector is generated for each of the first data and second data of each material.

[0112] The range of material properties estimated from the feature vector when operating as an estimation device depends on the correct answer information 35 used in the machine learning. The content and number of material properties estimated from the feature vector are not particularly limited and may be appropriately determined depending on the embodiment. The material properties may be, for example, catalytic properties, electron mobility, band gap, thermal conductivity, thermoelectric properties, mechanical properties (e.g., Young's modulus, sound speed, etc.), etc.

[0113] <Other> Each software module of the model generating device 1 and the data processing device 2 will be described in detail in an operation example described later. In this embodiment, an example is described in which each software module of the model generating device 1 and the data processing device 2 is realized by a general-purpose CPU. However, a part or all of the above software modules may be realized by one or more dedicated processors. That is, each of the above modules may be realized as a hardware module. Furthermore, with regard to the software configurations of the model generating device 1 and the data processing device 2, software modules may be omitted, replaced, or added as appropriate depending on the embodiment.

[0114] §3 Example of operation [Model generation device] 8 is a flowchart showing an example of a processing procedure of the model generation device 1 according to this embodiment. The processing procedure of the model generation device 1 described below is an example of a model generation method. However, the processing procedure of the model generation device 1 described below is merely an example, and each step may be changed as much as possible. Furthermore, steps may be omitted, replaced, or added to the processing procedure of the model generation device 1 described below as appropriate depending on the embodiment.

[0115] (Step S101) In step S101, the control unit 11 operates as a learning data acquisition unit 111 and acquires first data 31 and second data 32 for learning, which include a plurality of positive samples and a plurality of negative samples. Each position sample is composed of a combination of first data 31p and second data 32p for the same material. Each negative sample is composed of at least one of first data 31n and second data 32n for a material different from the material of the corresponding positive sample.

[0116] The first data 31 and the second data 32 may be obtained by actual measurement or may be obtained by simulation. A measuring device corresponding to each data (31, 32) may be used to measure each data (31, 32). The type of measuring device and the simulation method may be appropriately selected according to the type of each data (31, 32). For example, first-principles calculation, molecular dynamics calculation, etc. may be used as the simulation method.

[0117] In one example, the control unit 11 may directly acquire the first data 31 and the second data 32 from the corresponding measuring devices. Alternatively, the control unit 11 may acquire the first data 31 and the second data 32 by executing a simulation. In another example, the control unit 11 may acquire the first data 31 and the second data 32 from a storage area of ​​another computer or an external storage device, for example, via a network, a storage medium 91, or the like. In this case, the first data 31 and the second data 32 may be stored in the same storage area (storage device, storage medium), or may be stored in different storage areas. The number of samples of the first data 31 and the second data 32 to be acquired may be appropriately selected depending on the embodiment.

[0118] In this embodiment, the control unit 11 further acquires correct answer information 35 indicating the properties of the material corresponding to at least one of the first data 31 and the second data 32. The correct answer information 35 may be generated manually or by any mechanical method. In one example, the correct answer information 35 may be generated in the model generation device 1. In another example, the control unit 11 may acquire the correct answer information 35 from a storage area of ​​another computer or an external storage device via, for example, a network, a storage medium 91, or the like. Note that the timing of acquiring the correct answer information 35 is not limited to such an example. The process of acquiring the correct answer information 35 may be executed at any timing before the machine learning of the estimator 58 in step S104 described later is performed.

[0119] When the first data 31, the second data 32, and the correct answer information 35 are acquired, the control unit 11 advances the process to the next step S102.

[0120] (Step S102) In step S102, the control unit 11 operates as the machine learning unit 112, and performs machine learning on the first encoder 51 and the second encoder 52 using the acquired first data 31 and second data 32. As described above, the control unit 11 optimizes the values ​​of the calculation parameters of the first encoder 51 and the second encoder 52 by machine learning so that the first distance between the feature vectors of each positive sample is shorter than the second distance between the feature vector of each positive sample and the feature vector of the corresponding negative sample.

[0121] The optimization in this machine learning may be configured by at least one of adjusting to decrease the first distance and adjusting to increase the second distance. In addition, in this machine learning, the control unit 11 may optimize values ​​of the calculation parameters of the first encoder 51 and the second encoder 52 so that the first feature vector 41p and the second feature vector 42p of each positive sample match each other (i.e., the first distance approaches 0).

[0122] As a result of the machine learning, it is possible to generate a trained first encoder 51 and a trained second encoder 52 that have acquired the ability to map the first data and the second data of the same material to nearby positions in the feature space and to map the first data and the second data of different materials to distant positions. When the machine learning of the first encoder 51 and the second encoder 52 is completed, the control unit 11 proceeds to the next step S103.

[0123] (Step S103) In step S103, the control unit 11 operates as the machine learning unit 112 and performs machine learning on the first decoder 55 using the first data 31. As described above, the control unit 11 optimizes the values ​​of the calculation parameters of the first decoder 55 by machine learning so as to reduce the sum of errors between the output value indicating the restoration result and the corresponding first data 31 for each piece of first data 31. As a result of this machine learning, it is possible to generate a trained first decoder 55 that has acquired the ability to restore the corresponding first data from the first feature vector obtained by the first encoder 51.

[0124] In addition, the control unit 11 operates as a machine learning unit 112 and performs machine learning of the second decoder 56 using the second data 32. As described above, the control unit 11 optimizes the values ​​of the calculation parameters of the second decoder 56 by machine learning so that the sum of errors between the output value indicating the restoration result and the corresponding second data 32 is reduced for each second data 32. As a result of this machine learning, it is possible to generate a trained second decoder 56 that has acquired the ability to restore the corresponding second data from the second feature vector obtained by the second encoder 52. When the machine learning of the first decoder 55 and the second decoder 56 is completed, the control unit 11 proceeds to the next step S104.

[0125] The timing of executing the machine learning of each of the first decoder 55 and the second decoder 56 may not be limited to such an example. In another example, the machine learning of at least one of the first decoder 55 and the second decoder 56 may be executed simultaneously with the machine learning of the above step S102. When the machine learning of the first decoder 55 is executed simultaneously with the machine learning of the above step S102, the control unit 11 may also optimize the value of the calculation parameter of the first encoder 51 based on the error of the above restoration. When the machine learning of the second decoder 56 is executed simultaneously with the machine learning of the above step S102, the control unit 11 may also optimize the value of the calculation parameter of the second encoder 52 based on the error of the above restoration.

[0126] Furthermore, the first data 31 used in the machine learning of the first decoder 55 may not completely match the first data (31p, 31n) that may be used in the machine learning of each of the encoders (51, 52). Similarly, the second data 32 used in the machine learning of the second decoder 56 may not completely match the second data (32p, 32n) that may be used in the machine learning of each of the encoders (51, 52).

[0127] (Step S104) In step S104, the control unit 11 operates as the machine learning unit 112 and performs machine learning of the estimator 58 using a plurality of data sets. As described above, the control unit 11 optimizes the values ​​of the calculation parameters of the estimator 58 by machine learning so that the sum of errors between the output value of the estimation result derived from at least one of the first data 31 and the second data 32 and the true value indicated by the corresponding correct answer information 35 is reduced for each data set. As a result of this machine learning, a trained estimator 58 that has acquired the ability to estimate the characteristics of a material from at least one of the first feature vector and the second feature vector can be generated. When the machine learning of the estimator 58 is completed, the control unit 11 advances the process to the next step S105.

[0128] The timing of executing the machine learning of the estimator 58 may not be limited to such an example. In another example, the machine learning of the estimator 58 may be executed before the machine learning of at least one of the first decoder 55 and the second decoder 56. In another example, the machine learning of the estimator 58 may be executed simultaneously with the machine learning of the above step S102. In this case, when the estimator 58 is configured to estimate the characteristics of the material from the first feature vector, the control unit 11 may also optimize the values ​​of the calculation parameters of the first encoder 51 based on the error of the above estimation. Similarly, when the estimator 58 is configured to estimate the characteristics of the material from the second feature vector, the control unit 11 may also optimize the values ​​of the calculation parameters of the second encoder 52 based on the error of the above estimation.

[0129] Furthermore, the first data 31 and the second data 32 that can be used for the machine learning of the estimator 58 do not have to completely match the first data (31p, 31n) and the second data (32p, 32n) that can be used for the machine learning of each encoder (51, 52).

[0130] (Step S105) In step S105, the control unit 11 operates as the storage processing unit 113, and generates information about the trained machine learning models (the first encoder 51, the second encoder 52, the first decoder 55, the second decoder 56, and the estimator 58) generated by each machine learning as learning result data 125. Then, the control unit 11 stores the generated learning result data 125 in an arbitrary storage area.

[0131] The learning result data 125 may be stored in, for example, a RAM in the control unit 11, the storage unit 12, an external storage device, a storage medium, or a combination of these. The storage medium may be, for example, a CD, a DVD, or the like, and the control unit 11 may store the learning result data 125 in the storage medium via the drive 17. The external storage device may be, for example, a data server such as a NAS (Network Attached Storage). In this case, the control unit 11 may use the communication interface 13 to store the learning result data 125 in the data server via a network. The external storage device may also be, for example, an external storage device connected to the model generation device 1 via the external interface 14.

[0132] When the storage of the learning result data 125 is completed, the control unit 11 ends the processing procedure of the model generating device 1 according to this operation example.

[0133] The generated learning result data 125 may be provided to the data processing device 2 at any timing. In one example, the control unit 11 may transfer the learning result data 125 to the data processing device 2 as part of the process of step S105 or separately from the process of step S105. The data processing device 2 may acquire the learning result data 125 by receiving this transfer. In another example, the data processing device 2 may acquire the learning result data 125 by accessing the model generating device 1 or a data server via a network using the communication interface 23. In another example, the data processing device 2 may acquire the learning result data 125 via the storage medium 92. In another example, the learning result data 125 may be incorporated in the data processing device 2 in advance.

[0134] Furthermore, the control unit 11 may update or newly create a trained machine learning model by repeating the processes of steps S101 to S105 on a regular or irregular basis. In this case, the control unit 11 may update or newly create all of the machine learning models. Alternatively, the control unit 11 may update or newly create only a part of the machine learning models. Furthermore, during the repetition, at least a part of the first data 31 and the second data 32 that can be used for machine learning may be appropriately changed, modified, added, deleted, or the like. Then, the control unit 11 may provide the updated or newly created learning result data 125 to the data processing device 2 in any method and at any timing. As a result, the learning result data 125 (trained machine learning model) held by the data processing device 2 may be updated.

[0135] [Data processing device] (A) Data presentation process FIG. 9 is a flowchart showing an example of a processing procedure for presenting a feature vector by the data processing device 2 according to this embodiment. The processing procedure for presenting the feature vector below is an example of a data presenting method. The command portion of the processing procedure for presenting the feature vector below in the data processing program 82 is an example of a data presenting program. However, the processing procedure for presenting the feature vector below is merely an example, and each step may be modified as much as possible. Furthermore, steps may be omitted, replaced, or added to the processing procedure for presenting the feature vector below as appropriate depending on the embodiment.

[0136] (Step S201) In step S201, the control unit 21 operates as the target data acquisition unit 211, and acquires at least one of the first data 61 and the second data 62 relating to the crystal structure of each of a plurality of target materials.

[0137] The first data 61 and the second data 62 are the same type as the first data 31 and the second data 32 for learning. Like the first data 31 and the second data 32, the first data 61 and the second data 62 may be obtained by actual measurement or may be obtained by simulation. When acquiring the first data 61, at least a part of the acquired first data 61 may overlap with the first data 31 for learning. Similarly, when acquiring the second data 62, at least a part of the acquired second data 62 may overlap with the second data 32 for learning. In one example, at least one of the first data 61 and the second data 62 to be processed may be specified by an operator in any manner.

[0138] In one example, the control unit 21 may acquire at least one of the first data 61 and the second data 62 directly from the corresponding measuring device, or may acquire them as a result of executing a simulation. In another example, the control unit 21 may acquire at least one of the first data 61 and the second data 62 from a storage area of ​​another computer or an external storage device, for example, via a network, a storage medium 92, or the like. In this case, when acquiring both, the first data 61 and the second data 62 may be stored in the same storage area (storage device, storage medium), or may be stored in different storage areas. The number of samples of at least one of the first data 61 and the second data 62 to be acquired may be appropriately selected depending on the embodiment.

[0139] When at least one of the first data 61 and the second data 62 for each target material is acquired, the control unit 21 advances the process to the next step S202.

[0140] (Step S202) In step S202, the control unit 21 operates as a conversion unit 212 and performs at least one of a process of converting the acquired first data 61 into a first feature vector 71 and a process of converting the acquired second data 62 into a second feature vector 72 using at least one of the trained first encoder 51 and the trained second encoder 52.

[0141] Specifically, when acquiring the first data 61 and converting the acquired first data 61 into the first feature vector 71, the control unit 21 sets the trained first encoder 51 with reference to the learning result data 125. Then, the control unit 21 inputs the first data 61 of each target material to the trained first encoder 51 and executes arithmetic processing of the trained first encoder 51. As a result of this arithmetic processing, the control unit 21 acquires the first feature vector 71 of each target material.

[0142] Similarly, when acquiring second data 62 and converting the acquired second data 62 into a second feature vector 72, the control unit 21 sets the trained second encoder 52 with reference to the learning result data 125. Then, the control unit 21 inputs the second data 62 of each target material to the trained second encoder 52 and executes arithmetic processing of the trained second encoder 52. As a result of this arithmetic processing, the control unit 21 acquires the second feature vector 72 of each target material.

[0143] When at least one of the first feature vector 71 and the second feature vector 72 for each target material is acquired through the above processing, the control unit 21 advances the processing to the next step S203.

[0144] (Step S203) In step S203, the control unit 21 operates as the output processing unit 215, and maps at least one of the first feature vector 71 and the second feature vector 72 of each target material obtained onto the space VS. The space VS is for displaying the positional relationship of the feature vectors.

[0145] In one example, the control unit 21 may map each value of at least one of the first feature vector 71 and the second feature vector 72 of each target material obtained directly onto the space VS. In another example, the control unit 21 may convert each value of at least one of the first feature vector 71 and the second feature vector 72 of each target material obtained into a lower dimension so as to maintain the positional relationship of each value, and then map each converted value onto the space VS. As an example of the conversion, the original dimension of each feature vector (71, 72) may be several tens to 1000. In contrast, the dimension after conversion may be two or three dimensions. As long as the positional relationship of the feature vectors can be maintained as much as possible, the conversion method is not particularly limited and may be appropriately selected depending on the embodiment. The conversion method may employ, for example, t-SNE (t-distributed stochastic neighbor embedding), NMF (non-negative matrix factorization), PCA (principal component analysis), ICA (independent component analysis), Fast ICA (a fast algorithm for ICA), MDS (multidimensional scaling), Spectral Embedding, random projection, UMAP (uniform manifold approximation and projection), etc. The space VS into which each converted value is mapped may be referred to as, for example, a visualization space, a reduced-dimensional feature space, etc.

[0146] When the mapping of each feature vector to the space VS is completed, the control unit 21 advances the process to the next step S204.

[0147] (Step S204) In step S204, the control unit 21 operates as the output processing unit 215, and outputs each value of at least one of the first feature vector 71 and the second feature vector 72 of each target material mapped onto the space VS. In the process of step S203, when each value of at least one of the first feature vector 71 and the second feature vector 72 is converted to a lower dimension, the control unit 21 outputs each value of the feature vector converted to a lower dimension.

[0148] The output destination and the output format may be appropriately selected according to the embodiment. The output destination may be, for example, the output device 26, an output device of another computer, etc. The output format may be, for example, a screen output, a printout, etc. The control unit 21 may also execute any information processing when outputting the feature vector. As an example of information processing, the control unit 21 may accept the selection of one or more attention materials from among a plurality of target materials. The attention material may be selected, for example, by designating it from a list of target materials, designating a feature vector displayed on the space VS, or the like. Then, the control unit 21 may output the selected attention material in distinction from other target materials. The control unit 21 may also output a list of other target materials whose feature vectors exist in a neighborhood range of the feature vector of the selected attention material. The neighborhood range may be appropriately designated. The other target materials existing in the neighborhood range may be sorted in order of proximity on the space VS and then output.

[0149] When the output of each value of the feature vector is completed, the control unit 21 ends the processing procedure related to the data presentation according to this operation example. The control unit 21 may repeatedly execute the processes of steps S201 to S204 described above at any timing, such as when receiving an instruction from an operator. During this repetition, at least a part of the data acquired in step S201 (at least one of the first data 61 and the second data 62) may be appropriately changed, modified, added, deleted, etc. As a result, the data output in step S204 may be changed.

[0150] (B) Processing for generating second data from first data FIG. 10A is a flowchart showing an example of a processing procedure for generating second data 64 from first data 63 by the data processing device 2 according to this embodiment. The following data generation processing procedure is an example of a data generation method. The command portion of the following data generation processing procedure in the data processing program 82 is an example of a data generation program. However, the following data generation processing procedure is merely an example, and each step may be changed as much as possible. Furthermore, steps may be omitted, replaced, or added to the following data generation processing procedure as appropriate depending on the embodiment.

[0151] (Step S301) In step S301, the control unit 21 operates as the target data acquisition unit 211 and acquires first data 63 for at least one target material. The first data 63 is the same type as the first data 31 for learning. Like the first data 31, the first data 63 may be acquired by actual measurement or may be acquired by simulation. The number of first data 63 to be acquired may be appropriately determined depending on the embodiment.

[0152] In one example, the control unit 21 may obtain the first data 63 directly from the measurement device, or may obtain the first data 63 as a result of executing a simulation. In another example, the control unit 21 may obtain the first data 63 from a storage area of ​​another computer or an external storage device, for example, via a network, a storage medium 92, etc. When the control unit 21 obtains the first data 63, the control unit 21 proceeds to the next step S302.

[0153] (Step S302) In step S302, the control unit 21 operates as the conversion unit 212 and converts the acquired first data 63 into a first feature vector 73 using the trained first encoder 51. Specifically, the control unit 21 sets the trained first encoder 51 with reference to the learning result data 125. The control unit 21 inputs the acquired first data 63 into the trained first encoder 51 and executes arithmetic processing of the trained first encoder 51. As a result of this arithmetic processing, the control unit 21 acquires a first feature vector 73 of the target material. When the first feature vector 73 is acquired, the control unit 21 proceeds to the next step S303.

[0154] (Step S303) In step S303, the control unit 21 operates as the restoration unit 213, and restores the second data 64 from at least one of the value of the first feature vector 73 obtained by the conversion and its neighboring values, using the trained second decoder 56. That is, the control unit 21 handles at least one of the value of the first feature vector 73 obtained by the processing of step S302 and its neighboring values ​​as the value of the second feature vector, thereby performing the restoration of the second data 64.

[0155] Specifically, the control unit 21 sets the trained second decoder 56 with reference to the learning result data 125. The control unit 21 also determines one or more input values ​​for the trained second decoder 56 from the value of the first feature vector 73 obtained by the process of step S302 and its neighborhood range. The neighborhood range may be set appropriately. As an example, the maximum value of the first distance of the position samples may be calculated using the trained first encoder 51 and the trained second encoder 52. The neighborhood range may be set based on the maximum value of the first distance. The control unit 21 may use the obtained value of the first feature vector 73 as an input value as it is, or may use the neighborhood value of the obtained first feature vector 73 as an input value. The neighborhood value may be determined appropriately from the neighborhood range of the first feature vector 73.

[0156] Then, the control unit 21 inputs the determined input value to the trained second decoder 56, and executes the calculation process of the trained second decoder 56. As a result of this calculation process, the control unit 21 can generate second data 64 of the target material (i.e., obtain the restored second data 64 from the trained second decoder 56). In the process of this step S303, one or more input values ​​may be selected, thereby generating one or more pieces of second data 64 for one piece of first data 63. After generating the second data 64, the control unit 21 proceeds to the next step S304.

[0157] (Step S304) In step S304, the control unit 21 operates as the output processing unit 215 and outputs the generated second data 64. The output destination and the output format may be appropriately selected depending on the embodiment. The output destination may be, for example, the RAM, the storage unit 22, the output device 26, an output device of another computer, a storage area of ​​another computer, etc. The output format may be, for example, data output, screen output, printing, etc.

[0158] When the output of the generated second data 64 is completed, the control unit 21 ends the processing procedure related to data generation according to this operation example. The control unit 21 may repeatedly execute the processing of steps S301 to S304 described above at any timing, such as when receiving an instruction from an operator. During this repetition, the first data 63 to be processed may be appropriately selected in the processing of step S301.

[0159] (C) Processing for generating first data from second data FIG. 10B is a flowchart showing an example of a processing procedure for generating first data 66 from second data 65 by the data processing device 2 according to this embodiment. The following data generation processing procedure is an example of a data generation method. The command portion of the following data generation processing procedure in the data processing program 82 is an example of a data generation program. However, the following data generation processing procedure is merely an example, and each step may be changed as much as possible. Furthermore, steps may be omitted, replaced, or added to the following data generation processing procedure as appropriate depending on the embodiment.

[0160] (Step S401) In step S401, the control unit 21 operates as the target data acquisition unit 211 and acquires second data 65 of at least one target material. The second data 65 is the same type as the second data 32 for learning. Like the second data 32, the second data 65 may be acquired by actual measurement or may be acquired by simulation. The number of second data 65 to be acquired may be appropriately determined depending on the embodiment.

[0161] In one example, the control unit 21 may obtain the second data 65 directly from the measurement device, or may obtain the second data 65 as a result of executing a simulation. In another example, the control unit 21 may obtain the second data 65 from a storage area of ​​another computer or an external storage device, for example, via a network, a storage medium 92, etc. When the control unit 21 obtains the second data 65, the control unit 21 proceeds to the next step S402.

[0162] (Step S402) In step S402, the control unit 21 operates as the conversion unit 212 and converts the acquired second data 65 into a second feature vector 75 using the trained second encoder 52. Specifically, the control unit 21 sets the trained second encoder 52 with reference to the learning result data 125. The control unit 21 inputs the acquired second data 65 into the trained second encoder 52 and executes arithmetic processing of the trained second encoder 52. As a result of this arithmetic processing, the control unit 21 acquires the second feature vector 75 of the target material. When the second feature vector 75 is acquired, the control unit 21 proceeds to the next step S403.

[0163] (Step S403) In step S403, the control unit 21 operates as the restoration unit 213, and restores the first data 66 from at least one of the value of the second feature vector 75 obtained by the conversion and its neighboring values, using the trained first decoder 55. That is, the control unit 21 performs restoration of the first data 66 by treating at least one of the value of the second feature vector 75 obtained by the processing of step S402 and its neighboring values ​​as the value of the first feature vector.

[0164] Specifically, the control unit 21 sets the trained first decoder 55 with reference to the learning result data 125. The control unit 21 also determines one or more input values ​​for the trained first decoder 55 from the value of the second feature vector 75 obtained by the process of step S402 and its neighborhood range. As in the above step S303, the neighborhood range may be set appropriately. As an example, the neighborhood range may be set based on the maximum value of the first distance calculated by the trained first encoder 51 and the trained second encoder 52. The control unit 21 may use the obtained value of the second feature vector 75 as an input value as it is, or may use the neighborhood value of the obtained second feature vector 75 as an input value. The neighborhood value may be determined appropriately from the neighborhood range of the second feature vector 75.

[0165] Then, the control unit 21 inputs the determined input value to the trained first decoder 55, and executes the calculation process of the trained first decoder 55. As a result of this calculation process, the control unit 21 can generate first data 66 of the target material (i.e., obtain the restored first data 66 from the trained first decoder 55). In the process of this step S403, one or more input values ​​may be selected, thereby generating one or more first data 66 for one second data 65. After generating the first data 66, the control unit 21 advances the process to the next step S404.

[0166] (Step S404) In step S404, the control unit 21 operates as the output processing unit 215 and outputs the generated first data 66. The output destination and the output format may be appropriately selected according to the embodiment. The output destination may be, for example, the RAM, the storage unit 22, the output device 26, an output device of another computer, a storage area of ​​another computer, etc. The output format may be, for example, data output, screen output, printing, etc.

[0167] When the output of the generated first data 66 is completed, the control unit 21 ends the processing procedure related to data generation according to this operation example. The control unit 21 may repeatedly execute the processing of steps S401 to S404 described above at any timing, such as when receiving an instruction from an operator. During this repetition, the second data 65 to be processed may be appropriately selected in the processing of step S401.

[0168] (D) Characteristics estimation process FIG. 11 is a flowchart showing an example of a processing procedure for estimating the properties of a target material by the data processing device 2 according to this embodiment. The processing procedure for the following property estimation is an example of an estimation method. The command portion of the processing procedure for the following property estimation in the data processing program 82 is an example of an estimation program. However, the processing procedure for the following property estimation is merely an example, and each step may be changed as much as possible. Furthermore, steps may be omitted, replaced, or added to the processing procedure for the following property estimation as appropriate depending on the embodiment.

[0169] (Step S501) In step S501, the control unit 21 operates as the target data acquisition unit 211 and acquires at least one of the first data 67 and the second data 68 related to the crystal structure of the target material. The first data 67 and the second data 68 are the same type as the first data 31 and the second data 32 for learning. Like the first data 31 and the second data 32, the first data 67 and the second data 68 may be acquired by actual measurement or may be acquired by simulation.

[0170] In one example, the control unit 21 may obtain at least one of the first data 67 and the second data 68 directly from a corresponding measuring device, or may obtain the data as a result of executing a simulation. In another example, the control unit 21 may obtain at least one of the first data 67 and the second data 68 from a storage area of ​​another computer or an external storage device, for example, via a network, a storage medium 92, etc. When the control unit 21 obtains at least one of the first data 67 and the second data 68 of the target material, the control unit 21 advances the process to the next step S502.

[0171] (Step S502) In step S502, the control unit 21 operates as a conversion unit 212 and performs at least one of a process of converting first data 67 obtained using a trained first encoder 51 into a first feature vector 77, and a process of converting second data 68 obtained using a trained second encoder 52 into a second feature vector 78.

[0172] Specifically, when the trained estimator 58 is configured to estimate the properties of the target material from the first feature vector, the control unit 21 sets the trained first encoder 51 with reference to the learning result data 125. The control unit 21 inputs the acquired first data 67 to the trained first encoder 51 and executes the calculation process of the trained first encoder 51. As a result of this calculation process, the control unit 21 acquires the first feature vector 77 of the target material.

[0173] Similarly, when the trained estimator 58 is configured to estimate the properties of the target material from the second feature vector, the control unit 21 sets the trained second encoder 52 with reference to the learning result data 125. The control unit 21 inputs the acquired second data 68 to the trained second encoder 52 and executes the calculation process of the trained second encoder 52. As a result of this calculation process, the control unit 21 acquires the second feature vector 78 of the target material.

[0174] When at least one of the first feature vector 77 and the second feature vector 78 of the target material is acquired through the above processing, the control unit 21 advances the processing to the next step S503.

[0175] (Step S503) In step S503, the control unit 21 operates as the estimation unit 214, and estimates the characteristics of the target material from at least one of the obtained values ​​of the first feature vector 77 and the second feature vector 78 using the trained estimator 58. Specifically, the control unit 21 sets the trained estimator 58 with reference to the learning result data 125. The control unit 21 inputs at least one of the obtained values ​​of the first feature vector 77 and the second feature vector 78 to the trained estimator 58, and executes the calculation process of the trained estimator 58. As a result of this calculation process, the control unit 21 acquires an output value corresponding to the result of estimating the characteristics of the target material from the trained estimator 58. After acquiring the estimation result, the control unit 21 proceeds to the next step S504.

[0176] (Step S504) In step S504, the control unit 21 operates as the output processing unit 215 and outputs information related to the result of estimating the properties of the target material. The output destination and the output format may be appropriately selected depending on the embodiment. The output destination may be, for example, the RAM, the storage unit 22, the output device 26, an output device of another computer, a storage area of ​​another computer, etc. The output format may be, for example, data output, screen output, audio output, printing, etc.

[0177] When the output of the result of estimating the property of the target material is completed, the control unit 21 ends the processing procedure related to property estimation according to this operation example. The control unit 21 may repeatedly execute the processing of steps S501 to S504 described above at any timing, such as when receiving an instruction from an operator. During this repetition, at least one of the first data 67 and the second data 68 to be processed may be appropriately selected in the processing of step S501.

[0178] [Features] As described above, in this embodiment, it is possible to prepare positive samples and negative samples to be used for machine learning depending on whether the materials are the same or not. Therefore, in the model generating device 1, the trained first encoder 51 and the trained second encoder 52 can be generated at low cost by the above steps S101 and S102. In the data processing device 2, by the processing of the above steps S201 to S204, at least one of the trained first encoder 51 and the trained second encoder 52 can be used to map at least one of the first data 61 and the second data 62 of each of the multiple target materials into a feature space. In this feature space, the similarity of the materials can be evaluated based on the positional relationship of the feature vectors. Based on this evaluation result, new knowledge of the material can be obtained.

[0179] In this embodiment, data showing the properties of the material based on a local viewpoint of the crystal structure may be adopted as the first data 31, and data showing the properties of the material based on an overall bird's-eye view may be adopted as the second data 32. This makes it possible to generate trained encoders (51, 52) that have acquired the ability to map each piece of data into a feature space that allows the similarity of materials to be evaluated from both the local and bird's-eye views. In the data processing device 2, at least one of such trained first encoder 51 and trained second encoder 52 is used in the processes of steps S201 to S204, thereby making it possible to obtain new knowledge of the material with higher accuracy.

[0180] Furthermore, in this embodiment, the model generating device 1 can generate a trained first decoder 55 that has acquired the ability to restore the first data by the process of step S103 above. As a result, in the data processing device 2, using the trained second encoder 52 and trained first decoder 55 generated by the processes of steps S401 to S403 above, it is possible to generate valid first data from the second data of a target material for a material that is known in the second data but unknown in the first data.

[0181] Furthermore, in this embodiment, the model generating device 1 can generate a trained second decoder 56 that has acquired the ability to restore the second data by the process of step S103 above. As a result, in the data processing device 2, using the trained first encoder 51 and trained second decoder 56 generated by the processes of steps S301 to S303 above, it is possible to generate valid second data from the first data of a target material for a material that is known in the first data but unknown in the second data.

[0182] Furthermore, in this embodiment, in the model generating device 1, the process of step S104 described above can generate a trained estimator 58 that has acquired the ability to estimate material properties from at least one of the first feature vector and the second feature vector. As a result, in the data processing device 2, the process of steps S501 to S503 described above can estimate target material properties from at least one of the first data and the second data using at least one of the trained first encoder 51 and the trained second encoder 52 and the trained estimator 58.

[0183] In this embodiment, although the correct answer information 35 may be provided for all learning materials, the feature space mapped by each trained encoder (51, 52) contains information regarding the similarity of the materials. Since the estimator 58 is configured to estimate the material properties from the feature vectors in the feature space, the information can be taken into consideration when estimating the material properties. Therefore, even if the correct answer information 35 is not provided for all materials, a trained estimator 58 capable of estimating the material properties with high accuracy can be generated. Therefore, according to this embodiment, a trained estimator 58 capable of estimating the material properties with high accuracy can be generated at low cost.

[0184] § 4 Variations Although the embodiment of the present invention has been described in detail above, the above description is merely an example of the present invention in every respect. It goes without saying that various improvements or modifications can be made without departing from the scope of the present invention. For example, the following modifications are possible. In the following, the same reference numerals are used for the same components as in the above embodiment, and the description of the same points as in the above embodiment is omitted as appropriate. The following modifications can be combined as appropriate.

[0185] <4.1> In the above embodiment, the data, the encoder, and the decoder are referred to as "first" and "second". However, these references do not indicate that the number of these components is limited to two. In other words, "third" and subsequent data, encoders, and decoders may appear.

[0186] Fig. 12 shows an example of the configuration of an encoder according to another embodiment as an example of a scene in which the "third" and subsequent components appear. In the example of Fig. 12, in addition to the first encoder 51 and the second encoder 52, there is a third encoder 53 configured to convert third data into a third feature vector of the same dimension as the first feature vector and the second feature vector. The third data, like the first data and the second data, indicates information about the crystal structure of a material.

[0187] In this modification, the model generating device 1 may acquire multiple types of data related to the crystal structure of the material. Each type of data may indicate the properties of the material with an index different from that of the other types of data. The acquired multiple types of learning data may include multiple positive samples and negative samples. Each positive sample may be composed of a combination of multiple types of data for the same material. Each negative sample may be composed of at least one of multiple types of data for a material different from the material of the corresponding position sample.

[0188] The model generating device 1 may perform machine learning of a plurality of encoders using the acquired plurality of types of data. At least one encoder may correspond to each type of data. Each encoder may be configured to correspond to any one of the plurality of types of data and convert the corresponding type of data into a feature vector of the same dimension as the other encoders. The machine learning of the plurality of encoders may be configured by training the plurality of encoders so that, by using each encoder, the values ​​of the plurality of feature vectors calculated from the plurality of types of data of each positive sample are positioned close to each other, and the value of the feature vector calculated from at least one of the plurality of types of data of each negative sample is positioned far from at least one of the values ​​of the plurality of feature vectors calculated from the corresponding positive sample. The first data 31 and the second data 32 in the above embodiment may each be any one of the plurality of types of data. The first encoder 51 and the second encoder 52 may each be any one of the plurality of encoders.

[0189] The data processing device 2 may acquire at least one of a plurality of types of data related to the crystal structure of each of a plurality of target materials. The data processing device 2 may convert at least one of the acquired plurality of types of data of each target material into a feature vector using at least one of a plurality of trained encoders. The data processing device 2 may map the value of the acquired feature vector of each target material onto the space VS, and output each value of the feature vector mapped onto the space VS.

[0190] Furthermore, the model generating device 1 may perform machine learning of at least one decoder corresponding to each encoder. The machine learning of the at least one decoder may be configured by training the at least one decoder so that a result of restoring the corresponding type of data by the at least one decoder from a feature vector calculated from the corresponding type of data by using the corresponding encoder matches the corresponding type of data. In response to this, the data processing device 2 may generate other data (second data 64 / first data 66) from target data (the above-mentioned first data 63 / second data 65) among the multiple types of data.

[0191] Furthermore, the model generating device 1 may generate a trained estimator that has acquired the ability to estimate the material properties from at least one of a plurality of feature vectors by machine learning. The machine learning of the estimator may be configured by training the estimator so that the result of estimating the material properties from at least one of a plurality of feature vectors calculated from at least one of a plurality of types of data using at least one of a plurality of encoders matches the true value indicated by the corresponding correct answer information. Correspondingly, the data processing device 2 may estimate the target material properties from at least one of a plurality of types of data.

[0192] <4.2> In the data processing device 2 according to the above embodiment, at least any of the data presentation process, the process of generating second data from first data, the process of generating first data from second data, and the estimation process may be omitted.

[0193] When the process of generating the second data from the first data is omitted, the process of generating the trained second decoder 56 in step S103 may be omitted in the model generation device 1. Information on the trained second decoder 56 may be omitted from the learning result data 125.

[0194] When the process of generating the first data from the second data is omitted, the process of generating the trained first decoder 55 in step S103 may be omitted in the model generation device 1. Information on the trained first decoder 55 may be omitted from the learning result data 125.

[0195] When the estimation process is omitted, the process of generating the trained estimator 58 (step S104) may be omitted in the model generation device 1. Information on the trained estimator 58 may be omitted from the learning result data 125.

[0196] In the data processing device 2, when the trained first encoder 51 is not used, information on the trained first encoder 51 may be omitted from the learning result data 125. In the data processing device 2, when the trained second encoder 52 is not used, information on the trained second encoder 52 may be omitted from the learning result data 125.

[0197] In response to the omission of each process, components for executing the corresponding process may be omitted in each software module of the model generating device 1 and the data processing device 2. As an example, when the data presentation process is omitted, in the software configuration of the data processing device 2, the parts related to the data presentation process of the target data acquisition unit 211, the conversion unit 212, and the output processing unit 215 may be omitted. As another example, when both data generation processes are omitted, in the software configuration of the model generating device 1, the parts generating the trained first decoder 55 and the trained second decoder 56 may be omitted. In the software configuration of the data processing device 2, the parts related to the data generation process of the target data acquisition unit 211, the conversion unit 212, and the output processing unit 215, and the restoration unit 213 may be omitted. As another example, when the estimation process is omitted, in the software configuration of the model generating device 1, the parts generating the trained estimator 58 may be omitted. In the software configuration of the data processing device 2, the parts related to the estimation process of the target data acquisition unit 211, the conversion unit 212, and the output processing unit 215, and the estimation unit 214 may be omitted.

[0198] In addition, at least one of the data presentation process, the process of generating second data from first data, the process of generating first data from second data, and the estimation process may be executed by a different computer. As an example, the data presentation process, the process of generating second data from first data, the process of generating first data from second data, and the estimation process may be executed by different computers. In this case, the computer that executes each process may be configured similarly to the data processing device 2 described above.

[0199] <4.3> In the above embodiment, a trained estimator 58 is generated. A trained converter that estimates at least one of the first feature vector and the second feature vector from information indicating the characteristics of the target material may be generated corresponding to the trained estimator 58. The trained converter can be generated by machine learning in which the input and output of the estimator 58 are reversed. That is, the machine learning of the converter may be configured by training the converter so that at least one of the first feature vector and the second feature vector estimated by the converter from the characteristics indicated by the correct answer information 35 matches at least one of the first feature vector and the second feature vector calculated from at least one of the corresponding first data 31 and second data 32 using at least one of the first encoder 51 and the second encoder 52. The trained converter may be generated by the model generating device 1, or may be generated by another computer.

[0200] Thus, at least one of a first feature vector and a second feature vector of a material having the property of the target may be estimated from information indicating the property of the target using the trained converter. Then, at least one of a first trained decoder 55 and a second trained decoder 56 may be used to restore at least one of the first data and the second data from the estimated at least one of the first feature vector and the second feature vector. This process of restoring data from the property of the material may be executed by the data processing device 2 or may be executed by another computer.

[0201] §5 Experimental Examples In order to verify the effectiveness of the present invention, a trained first encoder and a trained second encoder according to the following experimental examples were generated, although the present invention is not limited to the following experimental examples.

[0202] (1) First Experimental Example First, 122,543 pieces of inorganic material data consisting of five or less elements were collected (downloaded) from inorganic material data registered in the Materials Project database (https: / / materialsproject.org / ). The three-dimensional atomic position data included in the collected inorganic material data was adopted as the first data. In addition, X-ray diffraction data obtained from this three-dimensional atomic position data by a simulation based on Bragg's law (using the Python library "pymatgen") was adopted as the second data. Then, a trained first encoder and a trained second encoder according to the first experimental example were generated by the same method as in the above embodiment. The first encoder uses a convolutional neural network with a convolutional layer (References: Charles R. Qi, Li Yi Hao Su, Leonidas J. Guibas, "PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space", 31st Conference on Neural Information Processing Systems (NIPS 2017) / Tian Xie, Jeffrey C. Grossman, "Crystal Graph Convolutional Neural Networks for an Accurate and Interpretable Prediction of Material Properties", Phys. Rev. Lett. 120, 145301, 6 April 2018). The second encoder uses a one-dimensional convolutional neural network. Each encoder is configured to convert each data into a 1024-dimensional feature vector. The loss function in the machine learning of each encoder is Triplet Loss. Specifically, the error L was calculated by the following equations 1 to 3, and the parameters of each encoder were optimized by the error backpropagation method.

[0203]

number

number

number

[0204] Using the generated trained first encoder, the first data (three-dimensional atomic position data) of each material used for machine learning was converted into a first feature vector. Next, the dimension of each first feature vector was converted from 1024 dimensions to two dimensions using t-SNE, and the values ​​of each feature vector were mapped into a two-dimensional visualization space and output to the screen. The resulting map (data distribution) was then analyzed using two methods: (A) global distribution analysis and (B) local neighborhood analysis.

[0205] (A) Global distribution analysis To confirm how the elements corresponding to each material are distributed on the map, the extent to which elements corresponding to materials containing each element in the periodic table exist in the obtained map was analyzed. In addition, the correspondence between the distribution of each element and the physical properties (energy above the hull, band gap, magnetization) was analyzed by color-coding each element in the obtained map according to the value of the physical properties.

[0206] Figure 13 shows the results of checking the range in the obtained map where elements corresponding to materials containing each element of the periodic table exist. Figures 14A to 14C show the results of coloring each element in the obtained map according to the value of the physical property (Figure 14A: energy above the hull, Figure 14B: band gap, Figure 14C: magnetization). Note that "na" in Figure 13 indicates that the corresponding element does not exist.

[0207] As shown in FIG. 13, the range of elements corresponding to materials containing each element was similar in both the vertical and horizontal directions of the periodic table. From this result, it was found that the obtained map adequately captured the similarity of the behavior of elements in each material. Also, as shown in FIG. 14A to FIG. 14C, elements having similar physical properties formed clusters on the obtained map. For example, as shown in FIG. 14A, a cluster of unstable compounds with large energy values ​​was confirmed in the upper left part of the map. In addition, the results of FIG. 14B and FIG. 14C confirmed that substances with similar band gap or magnetization values ​​formed multiple clusters, and each cluster was a group of substances with similar structure or composition. For example, the result of FIG. 14B confirmed that metals with low band gaps and nonmetals with high band gaps formed large clusters throughout the map. Also, the result of FIG. 14C confirmed that a cluster of rare earth permanent magnet materials with strong magnetization was confirmed in the upper right part of the map. These results indicate that the obtained maps adequately capture the similarities in the physical properties of each material.

[0208] (B) Local neighborhood analysis Next, to confirm what elements are located near each element on the obtained map (i.e., whether the map captures the similarity of materials), the two selected materials, "Hg-1223 (HgBa2Ca2Cu3O8)" and "LiCoO2," were used as queries to search for materials that exist near the queries.

[0209] In addition, feature vectors (feature representations of each material) according to the first and second comparative examples were generated using two types of descriptors, "Ewald Sum Matrix" and "Sine Coulomb Matrix," proposed in the reference "Faber, F., Lindmaa, A., von Lilienfeld, OA & Armiento, R. "Crystal structure representations for machine learning models of formation energies". Int. J. Quantum Chem. 115, 1094-1101 (2015)". These feature vectors were generated by calculating eigenvalue vectors in which the eigenvalues ​​of the matrix are arranged in descending order of absolute value from two types of descriptors expressed by matrices. Then, the feature vectors according to each comparative example were used to search for materials existing in the vicinity of each query.

[0210] [Table 1]

[0211] Table 1 shows the 1st to 50th nearest materials extracted for the query "Hg-1223" by the first experimental example and each comparative example. FIG. 15A shows the composition of the query "Hg-1223". FIG. 15B shows the nearest (1st) material extracted by the first experimental example, "Hg-1234 (HgBa2Ca3Cu4O 10 FIG. 15C shows the composition of the second most nearby extracted material, Hg-1212 (HgBa2CaCu2O6), from the first experimental example.

[0212] The query "Hg-1223" is a known superconductor with the highest critical temperature Tc. In the first experimental example, superconductors "Hg-1234" and "Hg-1212" with high critical temperatures Tc were extracted in the first and second neighborhoods of the query. As shown in Figs. 15A to 15C, "Hg-1234" and "Hg-1212" extracted as the first and second neighborhoods have a structure similar to that of the query "Hg-1223". In addition, in the first example, Tl-based superconductors with high critical temperatures Tc "Tl-2234" (fourth), "Tl-2212" (sixth), "Tl-1234" (seventh), and "Tl-1212" (19th) were extracted. Furthermore, in the first example, most of the nearby materials extracted up to the top 50 were superconductors. In contrast, the methods of the comparative examples extracted relatively large amounts of unrelated materials rather than superconductors.

[0213] [Table 2]

[0214] Table 2 shows the 1st to 50th nearest materials extracted for the query "LiCoO2" by the first experimental example and each comparative example. FIG. 16A shows the composition of the query "LiCoO2". FIG. 16B shows the composition of the material "LiCuO2" extracted as the nearest (1st) by the first experimental example. FIG. 16C shows the composition of the material "LiNiO2" extracted as the second nearest by the first experimental example.

[0215] The query "LiCoO2" is one of the most important cathode materials for lithium-ion batteries. In the first experimental example, materials "LiCuO2" and "LiNiO2" that have the same layer structure as the query but different transition metal elements were extracted in the first and second vicinity of the query (see Figures 16A to 16C). In the first example, the top seven neighboring materials have the same layer structure as the query but contain different transition metal elements. These neighboring materials included the actually important lithium-ion battery materials "LiNiO2" and "LiFeO2". In other words, in the first example, other important lithium-ion battery materials could be extracted from "LiCoO2". Also, in the first example, most of the neighboring materials extracted up to the top 50 were lithium oxides. In contrast, in the methods of each comparative example, inconsistent materials were extracted.

[0216] (C) Summary From the analysis results of the above two methods, it was found that the similarity of material properties can be evaluated based on the positional relationship in the feature space mapped by the trained encoder, even if information indicating the properties of the material, such as its structure, is not given. In other words, it was found that the above machine learning can generate a trained encoder that has acquired the ability to map data on crystal structures into a feature space capable of discovering similarities in material properties, even without giving information indicating the properties of the material. As a result, it was found that the trained encoder may be able to obtain new knowledge, such as the properties of new materials and the search for promising alternative materials.

[0217] (2) Second Experimental Example A trained first encoder and a trained second encoder for the second experimental example were generated under the same conditions as those of the first experimental example, except that the number of materials used for training was changed to 98,035 (80% of the total data). Using the generated trained first encoder, the first data of each of the 24,508 materials (20% of the total data) that were not used for training was converted into a first feature vector, and a map similar to that of the first experimental example was generated. In addition, using the generated trained second encoder, the second data of each of the same 24,508 materials was converted into a second feature vector. Then, the obtained second feature vector of each material was used as a query to extract the neighboring elements (materials) of the query in the generated map of the first feature vector. As a result, it was evaluated whether or not it was possible to search for the same material as the query by the second feature vector on the map of the first feature vector.

[0218] As a result of the evaluation, the probability that the same material would be extracted in the top 1 was 56.628%. The probability that the same material would be extracted in the top 5 was 95.203%. The probability that the same material would be extracted in the top 10 was 99.078%. In addition, if elements were randomly extracted from the obtained map, the probability that the same material would be extracted was 0.0041% (1 / 24,508). Therefore, it was found that the first data and the second data of the same material can be mapped to a nearby range with a high probability by the trained first encoder and the trained second encoder generated by the above machine learning. In other words, it was found that the first feature vector and the second feature vector of the same material have similar values ​​and can be replaced. From this result, it was found that if a trained decoder corresponding to each encoder is generated, it is possible to generate one of the first data and the second data from the other without significantly losing information.

[0219] (3) Supplementary Information In each experimental example, three-dimensional atomic position data was adopted as the first data, and X-ray diffraction data was adopted as the second data. Three-dimensional atomic position data is a type of data showing information on the local structure of the crystal of a material. X-ray diffraction data is a type of data showing information on the periodicity of the crystal structure of a material. Therefore, it was presumed that the same results as above would be obtained even if data showing information on the local structure of the crystal of a material other than three-dimensional atomic position data was adopted as the first data, and data showing information on the periodicity of the crystal structure of a material other than X-ray diffraction data was adopted as the second data. Examples of other data showing information on the local structure of the crystal of a material include Raman spectroscopy data, nuclear magnetic resonance spectroscopy data, infrared spectroscopy data, mass spectroscopy data, and X-ray absorption spectroscopy data. Examples of other data showing information on the periodicity of the crystal structure of a material include neutron diffraction data, electron beam diffraction data, and total scattering data.

[0220] In addition, it is possible to evaluate the properties of a material without necessarily being based on both a local viewpoint and a bird's-eye viewpoint of the crystal structure. Therefore, it was inferred that, as long as the first data and the second data indicate the properties of the material using different indices, it is possible to obtain results similar to those described above without adopting data showing information on the local structure of the crystal of the material as the first data, or without adopting data showing information on the periodicity of the crystal structure of the material as the second data. [Explanation of symbols]

[0221] 1...Model generation device, 11: control unit, 12: storage unit, 13: communication interface, 14...External interface, 15...input device, 16...output device, 17...drive, 81... generation program, 91... storage medium, 111...learning data acquisition unit, 112...machine learning unit, 113...storage processing section, 125...Learning outcome data, 2...Data processing device, 21: control unit, 22: storage unit, 23: communication interface, 24…External interface, 25...input device, 26...output device, 27...drive, 82...data processing program, 92...storage medium, 211: target data acquisition unit; 212: conversion unit; 213: restoration unit, 214: estimation unit, 215: output processing unit, 31: first data, 32: second data, 35: correct answer information, 41: first feature vector; 42: second feature vector; 51: first encoder; 52: second encoder; 55...first decoder, 56...second decoder, 58... Estimator, 61...first data, 62...second data, 71...first feature vector, 72...second feature vector, 63...first data, 64...second data, 73…first feature vector, 65...second data, 66...first data, 75…second feature vector, 67...first data, 68...second data, 77…First feature vector, 78…Second feature vector

Claims

1. A computer obtains first data and second data related to a crystal structure of the material, the second data indicates a property of the material in an index different from that of the first data; the acquired first data and second data include positive samples and negative samples; the positive sample is comprised of a combination of first data and second data for the same material; and the negative sample is composed of at least one of first data and second data about a material different from a material of the positive sample; Steps and The computer performs machine learning of a first encoder and a second encoder using the acquired first data and the acquired second data, the first encoder configured to convert the first data into a first feature vector; the second encoder configured to convert the second data into a second feature vector; the dimensionality of the first feature vector is the same as the dimensionality of the second feature vector; and The machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder such that values ​​of a first feature vector and a second feature vector calculated from the first data and the second data of the positive sample are positioned close to each other, and a value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative sample is positioned far from a value of at least one of the first feature vector and the second feature vector calculated from the positive sample. Steps and Equipped with Model generation method.

2. The method further comprises the step of: The machine learning of the first decoder is configured by training the first decoder so that a result of restoring the first data by the first decoder from a first feature vector calculated from the first data by using the first encoder matches the first data. The method of claim 1 .

3. The computer further comprises the step of performing machine learning of a second decoder; The machine learning of the second decoder is configured by training the second decoder so that a result of restoring the second data by the second decoder from a second feature vector calculated from the second data by using the second encoder matches the second data. The model generating method according to claim 1 or 2.

4. The computer further comprises a step of performing machine learning of an estimator; In the step of acquiring the first data and the second data, the computer further acquires correct answer information indicating properties of the material, The machine learning of the estimator is configured by training the estimator using the first encoder and the second encoder so that a result of estimating the material characteristics from at least one of the first feature vector and the second feature vector calculated from the acquired first data and the second data matches the correct answer information. A method for generating a model according to any one of claims 1 to 3.

5. the first data indicates information about a local structure of a crystal of the material; The second data indicates information regarding the periodicity of the crystal structure of the material. A method for generating a model according to any one of claims 1 to 4.

6. The first data is composed of at least one of three-dimensional atomic position data, Raman spectroscopy data, nuclear magnetic resonance spectroscopy data, infrared spectroscopy data, mass spectrometry data, and X-ray absorption spectroscopy data. The method of claim 5 .

7. the first data is composed of three-dimensional atomic position data; The three-dimensional atomic position data is configured to represent the state of atoms in the material by at least one of a probability density function, a probability distribution function, and a probability mass function. The method of claim 5 .

8. The second data is composed of at least one of X-ray diffraction data, neutron diffraction data, electron diffraction data, and total scattering data. A model generating method according to any one of claims 5 to 7.

9. A computer acquires at least one of first data and second data related to a crystal structure of each of a plurality of target materials; converting, by the computer, at least one of the acquired first data and second data for each of the target materials into at least one of a first feature vector and a second feature vector using at least one of a trained first encoder and a trained second encoder; The computer maps each value of at least one of the first feature vector and the second feature vector of each of the target materials obtained onto a space; outputting the respective values ​​of at least one of the first feature vector and the second feature vector of each of the target materials mapped onto the space; A method of presenting data comprising: The second data indicates a property of the material in an index different from that of the first data, the dimension of the first feature vector is the same as the dimension of the second feature vector; The first trained encoder and the second trained encoder are generated by machine learning using first data and second data for training, The first and second learning data include positive and negative samples, the positive sample is composed of a combination of first data and second data for the same material; the negative sample is comprised of at least one of first and second data about a material different from a material of the positive sample; and The machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder such that values ​​of a first feature vector and a second feature vector calculated from the first data and the second data of the positive sample are positioned close to each other, and a value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative sample is positioned far from a value of at least one of the first feature vector and the second feature vector calculated from the positive sample. How data is presented.

10. In the mapping step, the computer converts each of the values ​​of at least one of the first feature vector and the second feature vector of each of the target materials obtained into a low-dimensional form so as to maintain a positional relationship between the values, and then maps the converted values ​​onto a space; In the step of outputting each of the values, the computer outputs each of the transformed values ​​of at least one of the first feature vector and the second feature vector of each of the target materials. The method of claim 9 .

11. A data generation method for generating second data from first data, comprising the steps of: the first data and the second data relate to a crystal structure of a target material; The second data indicates a property of the material in an index different from that of the first data, The data generation method includes: A computer acquires first data of the target material; converting the acquired first data of the target material into a first feature vector using a trained first encoder; generating the second data by recovering the second data from at least one of the values ​​of the first feature vector obtained by the transformation and values ​​in the vicinity of the values ​​of the first feature vector using a trained decoder; Equipped with the trained first encoder, together with the second encoder, is generated by machine learning using first and second training data; the second encoder configured to convert the second data into a second feature vector; the dimension of the first feature vector is the same as the dimension of the second feature vector; The first and second learning data include positive and negative samples, the positive sample is composed of a combination of first data and second data for the same material; the negative sample is comprised of at least one of first data and second data about a material different from a material of the positive sample; the machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder such that values ​​of a first feature vector and a second feature vector calculated from the first data and the second data of the positive sample are positioned close to each other, and a value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative sample is positioned far from a value of at least one of the first feature vector and the second feature vector calculated from the positive sample; the trained decoder was generated by machine learning using the second data for training; and The machine learning of the decoder is configured by training the decoder so that a result of restoring the second data by the decoder from a second feature vector calculated from the second data for learning by using the second encoder matches the second data for learning. Data generation method.

12. A data generation method for generating second data from first data, comprising the steps of: The first data indicates information about a local structure of a crystal of a target material, The second data indicates information regarding a periodicity of a crystal structure of the target material, The data generation method includes: A computer acquires first data of the target material; converting the acquired first data of the target material into a first feature vector using a trained first encoder; generating the second data by recovering the second data from at least one of the values ​​of the first feature vector obtained by the transformation and values ​​in the vicinity of the values ​​of the first feature vector using a trained decoder; Equipped with the trained first encoder, together with the second encoder, is generated by machine learning using first and second training data; the second encoder configured to convert the second data into a second feature vector; the dimension of the first feature vector is the same as the dimension of the second feature vector; The first and second learning data include positive and negative samples, the positive sample is composed of a combination of first data and second data for the same material; the negative sample is comprised of at least one of first data and second data about a material different from a material of the positive sample; the machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder such that values ​​of a first feature vector and a second feature vector calculated from the first data and the second data of the positive sample are positioned close to each other, and a value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative sample is positioned far from a value of at least one of the first feature vector and the second feature vector calculated from the positive sample; the trained decoder was generated by machine learning using the second data for training; and The machine learning of the decoder is configured by training the decoder so that a result of restoring the second data by the decoder from a second feature vector calculated from the second data for learning by using the second encoder matches the second data for learning. Data generation method.

13. A data generation method for generating first data from second data, comprising the steps of: The first data indicates information about a local structure of a crystal of a target material, The second data indicates information regarding a periodicity of a crystal structure of the target material, The data generation method includes: acquiring second data of the target material by a computer; converting the acquired second data of the target material into a second feature vector using a second trained encoder; generating the first data by restoring the first data from at least one of the values ​​of the second feature vector obtained by the transformation and values ​​in the vicinity of the values ​​of the second feature vector using a trained decoder; Equipped with the trained second encoder is generated by machine learning using the first data and the second data for training together with the first encoder; the first encoder configured to convert the first data into a first feature vector; the dimension of the first feature vector is the same as the dimension of the second feature vector; The first and second learning data include positive and negative samples, the positive sample is composed of a combination of first data and second data for the same material; the negative sample is comprised of at least one of first data and second data about a material different from a material of the positive sample; the machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder such that values ​​of a first feature vector and a second feature vector calculated from the first data and the second data of the positive sample are positioned close to each other, and a value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative sample is positioned far from a value of at least one of the first feature vector and the second feature vector calculated from the positive sample; the trained decoder was generated by machine learning using the first data for training; and The machine learning of the decoder is configured by training the decoder so that a result of restoring the first data by the decoder from a first feature vector calculated from the first data for learning by using the first encoder matches the first data for learning. Data generation method.

14. A computer acquires at least one of first data and second data related to a crystal structure of a target material; converting the acquired first and / or second data into at least one of a first feature vector and a second feature vector using at least one of a trained first encoder and a trained second encoder; the computer estimating a property of the target material from values ​​of at least one of the first and second feature vectors obtained using a trained estimator; An estimation method comprising: The second data indicates a property of the material in an index different from that of the first data, the dimension of the first feature vector is the same as the dimension of the second feature vector; The first trained encoder and the second trained encoder are generated by machine learning using first data and second data for training, The first and second learning data include positive and negative samples, the positive sample is composed of a combination of first data and second data for the same material; the negative sample is comprised of at least one of first data and second data about a material different from a material of the positive sample; the machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder such that values ​​of a first feature vector and a second feature vector calculated from the first data and the second data of the positive sample are positioned close to each other, and a value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative sample is positioned far from a value of at least one of the first feature vector and the second feature vector calculated from the positive sample; The trained estimator is generated by machine learning using further ground truth information that indicates characteristics of the training material; and The machine learning of the estimator is configured by training the estimator using at least one of the first encoder and the second encoder so that a result of estimating the characteristics of the material for learning from at least one of the first feature vector and the second feature vector calculated from at least one of the first data and second data for learning matches the correct answer information. Estimation method.

15. A learning data acquisition unit configured to acquire first data and second data related to a crystal structure of a material, the second data indicates a property of the material in an index different from that of the first data; the acquired first data and second data include positive samples and negative samples; the positive sample is comprised of a combination of first data and second data for the same material; and the negative sample is composed of at least one of first data and second data about a material different from a material of the positive sample; A learning data acquisition unit; A machine learning unit configured to perform machine learning of a first encoder and a second encoder using the acquired first data and the acquired second data, the first encoder configured to convert the first data into a first feature vector; the second encoder configured to convert the second data into a second feature vector; the dimensionality of the first feature vector is the same as the dimensionality of the second feature vector; and The machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder such that values ​​of a first feature vector and a second feature vector calculated from the first data and the second data of the positive sample are positioned close to each other, and a value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative sample is positioned far from a value of at least one of the first feature vector and the second feature vector calculated from the positive sample. Machine Learning Department, Equipped with Model generation device.

16. a target data acquisition unit configured to acquire at least one of first data and second data related to a crystal structure of each of a plurality of target materials; a conversion unit configured to perform at least one of a process of converting the first data into a first feature vector using a trained first encoder and a process of converting the second data into a second feature vector using a trained second encoder to obtain at least one of a first feature vector and a second feature vector; an output processing unit configured to map each value of at least one of the first feature vector and the second feature vector of each of the target materials obtained onto a space, and to output each value of at least one of the first feature vector and the second feature vector of each of the target materials mapped onto the space; A data presentation device comprising: The second data indicates a property of the material in an index different from that of the first data, the dimension of the first feature vector is the same as the dimension of the second feature vector; The first trained encoder and the second trained encoder are generated by machine learning using first data and second data for training, The first and second learning data include positive and negative samples, the positive sample is composed of a combination of first data and second data for the same material; the negative sample is comprised of at least one of first and second data about a material different from a material of the positive sample; and The machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder such that values ​​of a first feature vector and a second feature vector calculated from the first data and the second data of the positive sample are positioned close to each other, and a value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative sample is positioned far from a value of at least one of the first feature vector and the second feature vector calculated from the positive sample. Data presentation device.

17. 1. A data generating device configured to generate second data from first data, comprising: the first data and the second data relate to a crystal structure of a target material; The second data indicates a property of the material in an index different from that of the first data, The data generating device includes: a target data acquisition unit configured to acquire first data of the target material; a transformer configured to transform the acquired first data of the target material into a first feature vector using a trained first encoder; a reconstruction unit configured to generate the second data by reconstructing the second data from at least one of the values ​​of the first feature vector obtained by the transformation and values ​​in the vicinity of the values ​​of the first feature vector using a trained decoder; Equipped with the trained first encoder, together with the second encoder, is generated by machine learning using first and second training data; the second encoder configured to convert the second data into a second feature vector; the dimension of the first feature vector is the same as the dimension of the second feature vector; The first and second learning data include positive and negative samples, the positive sample is composed of a combination of first data and second data for the same material; the negative sample is comprised of at least one of first data and second data about a material different from a material of the positive sample; the machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder such that values ​​of a first feature vector and a second feature vector calculated from the first data and the second data of the positive sample are positioned close to each other, and a value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative sample is positioned far from a value of at least one of the first feature vector and the second feature vector calculated from the positive sample; the trained decoder was generated by machine learning using the second data for training; and The machine learning of the decoder is configured by training the decoder so that a result of restoring the second data by the decoder from a second feature vector calculated from the second data for learning by using the second encoder matches the second data for learning. Data generation device.

18. a target data acquisition unit configured to acquire at least one of first data and second data related to a crystal structure of the target material; a conversion unit configured to convert the acquired first and / or second data into a first and / or second feature vector using a trained first encoder and / or a trained second encoder; an estimator configured to estimate a property of the target material from values ​​of at least one of the first and second feature vectors obtained using a trained estimator; An estimation device comprising: The second data indicates a property of the material in an index different from that of the first data, the dimension of the first feature vector is the same as the dimension of the second feature vector; The first trained encoder and the second trained encoder are generated by machine learning using first data and second data for training, The first and second learning data include positive and negative samples, the positive sample is composed of a combination of first data and second data for the same material; the negative sample is comprised of at least one of first data and second data about a material different from a material of the positive sample; the machine learning of the first encoder and the second encoder is configured by training the first encoder and the second encoder such that values ​​of a first feature vector and a second feature vector calculated from the first data and the second data of the positive sample are positioned close to each other, and a value of at least one of the first feature vector and the second feature vector calculated from at least one of the first data and the second data of the negative sample is positioned far from a value of at least one of the first feature vector and the second feature vector calculated from the positive sample; The trained estimator is generated by machine learning using further ground truth information that indicates characteristics of the training material; and The machine learning of the estimator is configured by training the estimator using at least one of the first encoder and the second encoder so that a result of estimating the characteristics of the material for learning from at least one of the first feature vector and the second feature vector calculated from at least one of the first data and second data for learning matches the correct answer information. Estimation device.

Citation Information

Patent Citations

  • Converting apparatus and program

    JP2021099713A

  • Methods and apparatus for multi-modal prediction using a trained statistical model

    WO2019231624A2