Material Element Species Estimation System Using Spectral Data of Material

The MoE model system, incorporating one-dimensional CNN expert models and a sparse gate circuit, effectively addresses the challenge of identifying material types from spectral data, achieving high accuracy in estimating elemental species even with a large number of classes.

JP7683457B2Active Publication Date: 2025-05-27TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021176900
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-28
Publication Date
2025-05-27
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

Existing methods for identifying the type of a material from spectral data face challenges due to the enormous number of known material types, leading to decreased accuracy in classification.

Method used

A system utilizing a Mixture of Experts (MoE) model, comprising one-dimensional CNN expert models for each elemental species and a sparse gate circuit for combining feature vectors, along with a multi-layer perceptron unit for determining elemental presence, improves classification accuracy.

Benefits of technology

The system achieves high accuracy in estimating elemental species in materials, capable of identifying hundreds of thousands of material types with improved precision compared to single CNN models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683457000002
    Figure 0007683457000002
  • Figure 0007683457000003
    Figure 0007683457000003
  • Figure 0007683457000004
    Figure 0007683457000004
Patent Text Reader

Abstract

To provide a system capable of estimating element species contained in an arbitrary material from spectral data of the arbitrary material.SOLUTION: A system for estimating element species contained in a material includes: an expert model that is provided for each element species and outputs a feature quantity vector obtained using an algorithm including a one-dimensional convolutional neural network algorithm, when spectral data of a material is input; and a multi-layer perceptron unit that determines, on the basis of weighted linear combination of feature quantity vectors from the expert models, presence or absence of each element in the material for each element species using the neural network algorithm. An expert model for each element species is learned to be able to identify a type of a material containing an element thereof, and the multi-layer perceptron unit is learned, for each element type, so that it can be determined whether or not an element thereof is contained in the material for each element species.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for analyzing spectral data obtained from various materials. More specifically, spectral data of any material obtained by powder X-ray diffraction, infrared spectroscopy, etc. is collated with the spectral data of known materials stored in a spectral database using machine learning techniques to estimate the elemental species contained in any material. The spectral data of the material is any spectral data in which spectral intensity values are measured for any one-dimensional variable (it may be full-spectrum or line-spectrum). Specifically, in addition to powder X-ray diffraction spectra and infrared spectra, it may be a mass spectrum, NMR spectrum, absorption spectrum, emission spectrum, etc., and such cases also belong to the scope of the present invention.

Background Art

[0002] In identifying any material, spectral data of the material measured by any measurement method is measured, and the spectral data is compared with the spectral data of known materials. When comparing the spectral data of the material to be identified with known materials, the number of spectral data of known materials is now enormous, and manual comparison of spectral data requires an enormous amount of time, labor, and cost. Therefore, various attempts have been made to automatically execute the comparison of the spectral data of materials by a computer using machine learning technology or AI technology. For example, in Patent Document 1, a feature amount is extracted from X-ray diffraction spectral data obtained from any material, and a machine learning model trained to be able to select a diffraction pattern with a high degree of similarity from the diffraction patterns of known materials based on the feature amount is used to identify such an arbitrary material. A system has been proposed. Further, in Non-Patent Document 1, the effectiveness of several machine learning algorithms for the problem of identifying the dimensionality and space group of the crystal structure of a material from the X-ray diffraction spectral data of the material has been examined, and it has been reported that a discriminator using a one-dimensional a-CNN (all convolutional neural network) can accurately identify the dimensionality and space group of the crystal structure. In such a document, in the learning stage of a-CNN, as learning data, existing data of the X-ray diffraction spectrum of each material and data generated by artificially and randomly changing the existing data of each material by a method called physics-informed data augmentation are used. It has been reported that even when there is only one or two X-ray diffraction spectral data for each material of the existing data, a discriminator capable of achieving accurate identification of the dimensionality and space group of the crystal structure from the spectral data could be constructed. In Non-Patent Document 2, it has been shown that using a convolutional neural network (CNN), the constituent phases of quaternary compounds composed of Sr (strontium), Li (lithium), Al (aluminum), and O (oxygen) can be identified from the powder X-ray diffraction data.

[0003] Note that, as a machine learning technique for accurately achieving classification with a huge number of classes, a technique called deep metric learning has been proposed. Regarding such deep metric learning, various algorithms named CosFace, ArcFace, AdaFace, etc. have been devised for recognizing the faces of a large number of people (see Non-Patent Documents 3, 4, etc.). In those algorithms, the feature vector of an image obtained from a CNN of a two-dimensional image is input into the above-mentioned deep metric learning algorithm, and it has been reported that individual face images can be identified and classified with higher accuracy than in the case of using only a CNN.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Non-Patent Documents

[0005]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Non-Patent Document 4

Non-Patent Document 5

Non-Patent Document 6

Summary of the Invention

Problems to be Solved by the Invention

[0006] For the purpose of identifying the type of a material whose type is to be identified (unidentified material) by comparing the spectral data of the material with the existing data of known materials, when constructing an identifier that identifies and classifies the type of material from spectral data using machine learning techniques as in the above-mentioned literature, as already mentioned, since the number of types of known materials is enormous, the number of classes to be classified by the identifier also becomes enormous. For example, the materials registered in the Inorganic Crystal Structure Database (ICSD) provided by the National Institute of Standards and Technology in the United States, etc., as of 2021, number 136,899. Therefore, if all of them can be identified by X-ray diffraction data, the number of classes to be classified in the identifier into which the X-ray diffraction data is input is also the number of registered types in the database. In this regard, generally, in an identifier that identifies and classifies the type of material from spectral data using machine learning techniques as described above, the accuracy of identifying different types decreases as the number of classes to be classified, that is, the number of types of materials, increases.

[0007] By the way, in the technique of classifying classes by a neural network (NN), as a method for improving classification accuracy and efficiency when the number of classes becomes enormous, a model called a Mixture of Experts model has been proposed (see, for example, Non-Patent Document 5. Hereinafter, referred to as the "MoE model"). Briefly stated, in such a MoE model, a plurality of NN models are configured, each of which is a model (expert model) learned so as to be able to classify a part of the classes to be classified. A vector obtained by combining the feature vectors obtained from each of these multiple expert models is further input into an NN model (MLP (multi-layer perceptron) model: multi-layer perceptron model) learned so as to be able to classify the classes to be classified, and the MLP model is configured to output a classification result. According to the configuration of this MoE model, it may be possible to achieve high-precision classification even when the number of classes is so large that it cannot be fully classified by a single NN or CNN model alone.

[0008] Therefore, the inventor of the present invention examined the applicability of a CNN having a configuration of a MoE model in the identification of any material from spectral data. As a result, the inventor of the present invention configured, for each of various elements, an identifier using a one-dimensional CNN learned to be able to identify the type of material from the spectral data of the material containing the element as an expert model, and combined the feature vectors obtained from the expert models for each of those various elements according to the algorithm of a sparse gate circuit (see Non-Patent Document 5), and input the resulting vector into an MLP model. By configuring an identifier in the form of a MoE model, the inventor succeeded in configuring an identifier capable of accurately estimating the types (element types) of elements contained in an arbitrary material from the spectral data of the arbitrary material. This finding is utilized in the present invention.

[0009] Thus, one object of the present invention is to provide a system capable of estimating the element types contained in an arbitrary material from the spectral data of the arbitrary material.

Means for Solving the Problems

[0010] According to the present invention, the above object is achieved by a system for estimating the element types contained in a material from the spectral data of the material, comprising a data input layer to which the spectral data of the material is input, a vector output layer that outputs a feature vector of a predetermined dimension, and an expert model provided for each element type that calculates the feature vector based on the spectral data, and a sparse gate circuit unit that outputs a combined feature vector, which is a vector generated by a weighted linear combination of the feature vectors output from the vector output layer of the expert model for each element type, and a vector input layer to which the combined feature vector is input, a determination output layer that outputs a determination of whether or not each of the element types is contained in the material, and a contained element determination multi-layer perceptron unit that outputs a determination of whether or not each of the element types is contained in the material based on the combined feature vector. A system including the above It is achieved by

[0011] In the above configuration, The "spectral data of the material" may be data in which spectral intensity values are measured for any one-dimensional variable, such as a powder X-ray diffraction spectrum, an infrared absorption spectrum, a mass spectrum, an NMR spectrum, an absorption spectrum, or an emission spectrum, obtained from an arbitrary material. The material may be any inorganic or organic substance, and the elemental species may be any element, and may be all elemental species that may be contained in the material inspected by the system of the present invention, and thus may be all elemental species contained in ordinary materials. In an embodiment, the elemental species in the present invention may be 75 elemental species contained in ordinary inorganic or organic substances. The "expert model" is means provided for each elemental species, which, when spectral data is input to the data input layer, outputs, in the vector output layer, a feature vector corresponding to the spectral data obtained by calculation using an algorithm including an algorithm of a one-dimensional convolutional neural network based on the spectral data. Each of the expert models uses, for each elemental species, the spectral data of known materials containing that element as elemental species-specific learning data, and when each of the elemental species-specific learning data is input to the data input layer, is a model learned to output the probability for each material type of the elemental species-specific learning data such that the probability for the material type of the input elemental species-specific learning data is maximized, and the feature vector output to the vector output layer is a logit giving the probability for each material type of the elemental species-specific learning data. The "element content determination multi-layer perceptron unit" is configured such that when a combined feature vector is input to the vector input layer, the probability of each element type contained in the material is calculated by the neural network algorithm. When the probability of that element type exceeds a predetermined value, it is determined that the element type is contained in the material, and when the probability of that element is below the predetermined value, it is determined that the element type is not contained in the material, and this determination is output to the determination output layer. The "sparse gate circuit unit" and the "element content determination multi-layer perceptron unit" use spectral data of various known materials as overall learning data. When each of the overall learning data is input to each of the expert models for each element type, the probability of each element type contained in the element type calculated in the element content determination multi-layer perceptron unit exceeds the predetermined value for the element types contained in the material of the input overall learning data, and is learned to be below the predetermined value for the element types not contained in the material of the input overall learning data.

[0012] In the system of the present invention described above, the "expert model" provided for each element type is, as described above, simply put, a model that is learned to be able to identify the type of material from the spectral data of the material containing each element type for each element type. That is, for example, the expert model provided for lithium (Li prediction model) is configured to be able to identify the type of the material when spectral data of a material containing lithium is input, and the expert model provided for oxygen (O prediction model) is configured to be able to identify the type of the material when spectral data of a material containing oxygen is input. And in the system of the present invention, such expert models are to be prepared for all element types for which it is desired to detect whether they are contained in the material to be inspected.

[0013] Each of the above-mentioned expert models is basically trained to identify the type of material from the spectral data of the material according to a machine learning algorithm including the algorithm of a one-dimensional convolutional neural network. Here, the "algorithm of a one-dimensional convolutional neural network" may be an algorithm that performs operations by the algorithm of a convolutional neural network (CNN) using, as input data, data with values such as intensity values or luminance values assigned to one-dimensional variables, such as one-dimensional spectral data. Specifically, in the data input layer of each expert model, spectral values for each arbitrarily set predetermined variable interval in the spectral data are assigned to one neuron (perceptron). For example, in the case of X-ray diffraction data, the intensity value for each predetermined angle is input to one neuron in the input layer (therefore, in the range of 0 to 120 degrees, in the case of a configuration where an intensity value is given to one neuron every 0.02 degrees, 6000 neurons are prepared in the input layer). Usually, in a CNN, the identification result or the result numerical value of a regression operation is output in the final output layer. However, in the present invention, as described above, since a feature vector of an arbitrarily set predetermined dimension, for example, 1024 dimensions, is used for subsequent operations, the output layer is composed of neurons of such a predetermined number of dimensions, and a feature vector is composed of the output values of each of these neurons as elements (that is, the operation of the neural network of the present invention reaches a stage before the output stage of the identification or regression result in the normal case). According to the research of the inventors of the present invention, for the purpose of the present invention, it has been found that the algorithm of 1D-RegNet (Non-Patent Document 6) is advantageously used as the algorithm of a one-dimensional CNN.

[0014] Also, in the "expert model", the algorithm for calculating the feature vector preferably concatenates with the 1D CNN algorithm and executes the "deep metric learning" algorithm (AdaCos, CosFace, ArcFace, SphereFace, etc.) proposed for human face recognition described in Non-Patent Documents 3, 4, etc. already mentioned. Deep metric learning technology is a technology advantageous for solving class classification problems. In this technology, to put it simply, when the cosine similarity between the feature vector obtained from the CNN operation and the weight vector is calculated and the feature vector is mapped onto a multi-dimensional spherical surface, the distance on the spherical surface between the mapped points of the feature vectors is such that the weights are learned to be separated as the features of the data corresponding to the feature vectors are different. Then, using the cosine similarity between the feature vector and the weight vector, the probability that the data corresponding to the feature vector is classified into each class is calculated. In the case of the "expert model" of the present invention, after the feature vector is calculated from the spectral data of the material by the 1D CNN algorithm, the feature vector may be further processed and converted according to the deep metric learning algorithm so as to obtain more accurate results. As will be described in more detail in the section of the later embodiments, when executing the deep metric learning algorithm in the expert model, the "hierarchical metric learning" algorithm that executes the deep metric learning algorithm twice in succession may be executed so as to obtain more accurate results.

[0015] In the learning process of the expert model for each element type, as described above, for each element type, the spectral data of known materials containing the element is used as learning data (learning data for each element type). When each of the learning data is input to the data input layer, learning is executed so as to output the probability for each material type of the learning data such that the probability for the material type of the input learning data is maximized. The method of the learning process here may be executed according to an arbitrary method, for example, the backpropagation method. And, as already mentioned, in the subsequent arithmetic processing of the expert model for each element type, a feature vector of a predetermined dimension output to the vector output layer is used. In such a feature vector, logits (of the aforementioned predetermined number of dimensions) that give the probability for each material type of the learning data are used.

[0016] Thus, the feature vectors obtained by the expert model for each element type are then integrated as "combined feature vectors" by weighted linear combination in the sparse gate circuit section (therefore, here, the dimension of the vector does not change), and the combined feature vectors are input to the "element-containing determination multi-layer perceptron section" having the configuration of a neural network. And in the element-containing determination multi-layer perceptron section, based on the combined feature vectors, for each element type, the probability that each element type is contained in the material of the spectral data is calculated. For each element type, when the probability exceeds an arbitrarily set predetermined value (for example, 0.5), it is determined that the element type is contained in the material, and when the probability is less than the predetermined value, it is determined that the element type is not contained in the material.

[0017] The above sparse gate circuit unit and the contained element determination multi-layer perceptron unit are configured by machine learning using existing spectral data of various known materials as learning data (overall learning data). In a specific learning process, as described above, when each of the learning data of various known materials is input to each of the expert models for each element type, the probability of each element type contained in the element type calculated in the contained element determination multi-layer perceptron unit exceeds a predetermined value for the element types contained in the material of the input learning data, and is below the predetermined value for the element types not contained in the material of the input learning data. The weight parameters when calculating the weighted linear combination of the feature vectors in the sparse gate circuit unit and the weight parameters when executing the neural network algorithm in the contained element determination multi-layer perceptron unit are determined. The method of the learning process here may also be executed according to an arbitrary method, for example, the backpropagation method.

[0018] As the spectral data of the known materials used in the learning of each expert model and the learning of the sparse gate circuit unit and the contained element determination multi-layer perceptron unit, data registered or stored in any database may be used. For example, in the case of estimating the element types contained in a material from the X-ray diffraction data of an inorganic substance, the spectral data of the known materials may be data registered in the ICSD or the like.

[0019] In the configuration of the material contained element type estimation system of the present invention described above, to put it simply, an expert model configured to be able to identify the types of materials containing each element using a one-dimensional CNN as a basic algorithm for each element type, and a contained element determination multi-layer perceptron unit that performs weighted linear combination of the feature vectors obtained from these expert models and further determines the presence or absence of each element type in various materials using the neural network algorithm, a so-called MoE model is constructed. According to such a configuration, the performance of class classification is improved, and it becomes possible to accurately estimate the element types contained in each of a large number of materials that cannot be fully identified by a single CNN model.

[0020] Incidentally, usually, for the data of known materials registered in a database, only one or two standard examples are registered for each material. However, in the spectrum data obtained by actual measurement, various variations such as the shift of variables (shift of peak positions) such as the angle and frequency at which peaks occur in the data, the fluctuation of the intensity ratio of multiple peaks, the disappearance of peaks, and the splitting of peaks can occur. Therefore, if learning is performed only with the standard data registered in the database in the system of the present invention, it becomes difficult to accurately identify the material type and estimate the presence or absence of element types for the data with the above-described changes. Thus, in the configuration of the system of the present invention described above, as learning data, not only the spectrum data of known materials but also another spectrum data (extended data) generated by performing a process of artificially and virtually adding various changes as described above to the spectrum data of known materials ("physical-based data augmentation") may be used.

[0021] In such physical-based data augmentation processing, specifically, the following processing may be executed. (i) Peak shift processing - Randomly shift the variables at which peaks are observed in the spectrum data of known materials. (ii) Peak intensity ratio change processing - Randomly change the intensity ratios of multiple peaks observed in the spectrum data of known materials. (iii) Peak disappearance processing - Randomly erase the peaks observed in the spectrum data of known materials. (iv) Peak splitting processing - Randomly split the peaks observed in the spectrum data of known materials into two or more. (v) Combinations of at least two of (i) to (iv) In addition, changes to various data in each process may be executed to various degrees by adaptation, taking into account the degree of variation that can occur in actual measurement data. Preferably, in the algorithm for physical-based data augmentation, all of peak shift processing, peak intensity ratio change processing, peak disappearance processing, and peak splitting processing may be executed to generate extended data.

[0022] In general, the more extended data used as learning data, the more the effect of improving the discrimination accuracy of the discriminator is expected. However, it is not necessary to execute the generation of extended data and the learning process a useless number of times. Therefore, in the embodiment, the generation of extended data and the learning process using the same may be executed until the identification accuracy of the material type or the estimation accuracy of the elemental species contained in the material reaches a predetermined value. That is, the learning process of the expert model, the learning process of the sparse gate circuit unit, and the learning process of the multi-layer perceptron unit may generate different extended data by the algorithm of physical-based data augmentation and be repeatedly executed using the different extended data as learning data until the result accuracy for any material reaches a predetermined value.

[0023] Also, as described above, in the configuration of the above-mentioned expert model, when the deep distance learning algorithm is executed, the probabilities for various material types may be calculated by the "hierarchical distance learning" algorithm that continuously executes deep distance learning twice. In the "hierarchical distance learning" algorithm, the feature vector output by the 1D CNN is transformed at least twice by the deep distance learning algorithm, and then the probability that the data corresponding to the feature vector is classified into each class (in the case of the present invention, the type of material) is calculated. More specifically, in the case of the configuration where the deep distance learning algorithm is continued twice in the hierarchical distance learning algorithm, when the feature vector is input to the processing unit that executes the first deep distance learning, the feature vector is transformed by the operation according to the first deep distance learning algorithm, and a second feature vector of a predetermined dimension that may be arbitrarily set is calculated. Such a second feature vector may be a logit used to calculate the probability in the case of a configuration that calculates the probability for each class by the first deep distance learning algorithm. Then, the second feature vector (the logit of the first deep distance learning algorithm) is further input to the processing unit that executes the second deep distance learning algorithm, and the operation according to the deep distance learning algorithm is performed to further transform the second feature vector, and the probability for each class is calculated from the obtained feature vector, that is, the logit in the second deep distance learning algorithm. For example, a softmax function may be used to calculate the probability for each class from the logit. It has been found that the identification accuracy of the material type is improved by using such a hierarchical distance learning algorithm. When the hierarchical distance learning algorithm is adopted, the feature vector passed from each expert model to the sparse gate circuit unit may be the above-mentioned second feature vector, that is, the logit calculated in the first deep distance learning process.

[0024] The algorithm used in the above hierarchical distance learning algorithm may be any deep distance learning algorithm, specifically, it may be selected from AdaCos, CosFace, ArcFace, SphereFace, etc. As exemplified in the column of the following embodiments, according to the research of the inventor of the present invention, as a hierarchical distance learning algorithm, AdaCos (Non-Patent Document 4) is used as the first deep distance learning algorithm, and CosFace (Non-Patent Document 3) is used as the second deep distance learning algorithm. It has been found that the identification of material types can be achieved with higher accuracy. In that case, in the distance learning processing unit, according to the AdaCos algorithm, the feature vector output by the 1D CNN processing unit is converted into the logit of AdaCos, and the logit of AdaCos is further converted into the logit of CosFace according to the CosFace algorithm, and various probabilities for various material types can be calculated from the logit of CosFace. And in a system for estimating the elemental species in a material from the spectral data of the material, the feature vector passed from each expert model to the sparse gate circuit unit may be the logit calculated by AdaCos.

Advantages of the Invention

[0025] In the configuration of the system of the present invention described above, in estimating the elemental species contained in a material from the spectral data of the material, the configuration of the MoE model is adopted, whereby it is possible to estimate the presence or absence of elemental species for a huge number of materials. In fact, in the research of the inventor of the present invention, it has been found that according to the configuration of the present invention, the identification of hundreds of thousands of materials (the number of types is six digits) can be achieved with high accuracy. Also, according to the system of the present invention, for a very large number of materials, the elemental species contained in the materials can be estimated by one system. Therefore, when it is desired to identify any material type, first, the elemental species contained in the material can be estimated by the system of the present invention, and then, using the expert model of the elemental species estimated to be contained, it is possible to detect the material type in detail. Thus, it is expected that the identification of the material types corresponding to the spectral data of many arbitrary materials can be achieved with higher accuracy.

[0026] Other objects and advantages of the present invention will become apparent from the following description of the preferred embodiments of the present invention.

Brief Description of the Drawings

[0027]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Explanation of Reference Numerals

[0028] 1... Computer main body 2... Computer terminal 3... Monitor 4... Keyboard, mouse (input device) 10... Database

Best Mode for Carrying Out the Invention

[0029] Configuration of the computer device The system for estimating the elemental species contained in a material from the spectral data of an arbitrary material according to this embodiment may be realized by the operation of a computer program in the computer device 1 in a form commonly used in this field, such as that illustrated in FIG. 1. The computer device 1 is equipped with a CPU, a storage device, and an input / output device (I / O) interconnected by a bidirectional common bus in a normal manner. The storage device includes a memory storing each program for executing the arithmetic processing used in the arithmetic operations of this embodiment, a work memory, and a data memory used during the arithmetic operations. Further, the instructions by the implementer to the computer device 1 and the display and output of calculation results and other information are performed through the computer terminal device 2 connected to the computer device 1. The computer terminal device 2 is provided with a monitor 3 and input devices 4 such as a keyboard and a mouse in a normal manner. When the program is started, the implementer can perform various instructions and inputs to the computer device 1 using the input device 4 according to the display on the monitor 3 according to the procedure of the program, and can visually confirm the calculation state and calculation results from the computer device 1 on the monitor 3. Furthermore, in this embodiment, as will be described later, existing spectral data of known materials are used, and those existing spectral data may be acquired from an arbitrary database. Therefore, the computer device 1 may be able to access an arbitrary database 10 by an arbitrary method. It should be understood that other peripheral devices (such as a printer for outputting results, a storage device for inputting and outputting calculation conditions and calculation result information, etc.) not shown may be equipped in the computer device 1 and the computer terminal device 2. When executing various processes or calculations described below using the computer device 1, in a normal manner, programs necessary for various processes or calculations are started, and the implementer inputs data necessary for the calculation, calculation conditions during the calculation, and other various settings according to the input procedure prepared in the program in the computer terminal device 2, and the calculation is started. Then, during or after the execution of the calculation, the calculation result can be output through the computer terminal device 2 as appropriate.

[0030] In the above computer device 1, when estimating the elemental species contained in a material from spectral data, to put it simply, similar to the arithmetic processing using ordinary machine learning techniques, first, a learning process (determination of parameters such as weights used in the calculation) using existing data as learning data is executed. As a result, a model (discriminator) for estimating the elemental species in the material is constructed. Subsequently, spectral data of an arbitrary material for which the contained elemental species is to be estimated is input to the learned discriminator, and the discriminator outputs the elemental species contained in the material as the discrimination result. Thus, when executing the learning process, the setting of hyperparameters used in the calculation, the reading of existing data of known materials from the database, etc. are executed by the input operation of the operator from the input device 4, and the computer device 1 executes the learning process using the read existing data as learning data according to the program stored in the program memory. Then, when the learning process is completed, spectral data of the material for which the elemental species is to be estimated is input by the input operation of the operator from the input device 4, and the computer device 1 estimates the elemental species contained in the input spectral data material by the learned discriminator according to the program stored in the program memory, and the result is output to the computer terminal device 2 and displayed on a monitor or the like.

[0031] In this embodiment, as the spectral data of the material, as already mentioned, it may be data in which spectral values or intensity values are measured for any one-dimensional variable, such as a powder X-ray diffraction spectrum, infrared absorption spectrum, mass spectrum, NMR spectrum, absorption spectrum, emission spectrum, etc., and the material may be an inorganic or organic substance. The existing data of known materials used as learning data may be data obtained by any method (measurement or simulation) or data obtained from any database. For example, in the case of an inorganic powder X-ray diffraction spectrum, it may be data available from the ICSD provided by the National Institute of Standards and Technology, etc.

[0032] Configuration of the identifier Referring to FIG. 2, the identifier in the system for estimating the elemental species contained in a material from the spectral data of the material according to the present embodiment basically includes, for all elemental species whose presence in the material is to be inspected, a group of expert models provided for each elemental species, a sparse gate circuit section, and a multi-layer perceptron (MLP) model section (element-containing determination multi-layer perceptron section). In such a configuration, to put it simply, when the spectral data of the material to be inspected is input to each of the expert models for each elemental species, each expert model calculates a feature vector of a predetermined dimension that may be arbitrarily set based on the spectral data of the material, and these are passed to the sparse gate circuit section, where they are integrated into one feature vector (combined feature vector) by weighted linear combination. Then, when the MLP model section receives the combined feature vector, for all elemental species whose presence in the material is to be inspected, it calculates the probability that the element is contained in the material of the spectral data in which the element was input for each elemental species, and an elemental species whose probability exceeds a predetermined value is determined to be contained in the material.

[0033] (1) Basic configuration of the expert model In the configuration of the above system, first, each of the expert models provided for each elemental species is, to put it simply, an identifier configured to be able to identify the type of the material when receiving the spectral data of the material containing each elemental species using machine learning techniques. However, when incorporated into the system for estimating the elemental species of the present embodiment, the part from the data input layer to which the spectral data is input to the layer (vector output layer) where a feature vector of a predetermined number of dimensions is calculated is used.

[0034] The discriminator used in the expert model may basically be a discriminator that identifies the type of material from the spectral data of the material using a one-dimensional CNN algorithm, and it may use only one-dimensional CNN, but preferably, it may have a configuration in which the algorithm of deep metric learning is connected to the one-dimensional CNN algorithm. Specifically, as schematically depicted in FIG. 3, in the basic configuration of the discriminator used in the expert model, a one-dimensional CNN processing unit that inputs spectral data and calculates a feature vector, and a metric learning processing unit that receives the feature vector from the one-dimensional CNN processing unit and determines the probability for each material may be included.

[0035] More specifically, first, the one-dimensional CNN processing unit may be an arithmetic processing unit that implements the one-dimensional CNN algorithm in a normal manner. Specifically, for each of the plurality of neurons in the input layer of the CNN, the spectral values (intensity values, luminance values, etc.) for each variable of the spectral data of the material are input, and through operations in the convolutional layer, pooling layer, and fully connected layer, a feature vector with a predetermined number of dimensions is calculated in the output layer. For example, when the spectral data is a powder X-ray diffraction spectrum, 6000 neurons may be prepared in the input layer so that the intensity values every 0.02 degrees in the range of 0 to 120 degrees of 2θ, which is a variable, are given to each neuron. Also, the dimension of the feature vector in the output layer may be set as appropriate. In the case of the example where 6000 values are given to the above input layer, for example, the number of dimensions of the feature vector in the output layer may be 1024 or the like. According to the research of the inventor of the present invention, in this embodiment, it has been found that the 1D-RegNet algorithm (Non-Patent Document 6) is advantageously used as the one-dimensional CNN algorithm.

[0036] As already mentioned, the distance learning processing unit may be an arithmetic processing unit that implements algorithms of deep distance learning techniques such as AdaCos, CosFace, ArcFace, and SphereFace. Specifically, it receives the feature vector output by the 1D CNN processing unit, calculates the cosine similarity between the feature vector and the weight vector, calculates the logit using such cosine similarity and hyperparameters, and calculates, for each class, the probability that the data corresponding to the feature vector is classified into each class from such logit. Here, each class classified by the distance learning processing unit is, in this embodiment, the type of material including the corresponding element type of each expert model and is the type of material to which each data can correspond. For the calculation from the logit to the probability of each class, typically, the softmax function may be used, but it is not limited thereto. Then, among the probabilities of each class, the class giving the maximum probability is selected, and the material type of the selected class is identified as the material type of the input spectral data. Note that, as will be described later, in the distance learning processing unit in this embodiment, the algorithm of "hierarchical distance learning" that continuously executes the algorithm of the deep distance learning technique at least twice may be advantageously adopted.

[0037] The learning of the discriminator including the one-dimensional CNN processing unit and the distance learning processing unit may be performed in the so-called "End-to-end training" manner using the existing data of known materials as learning data, as described above. Here, the known materials are known materials containing the corresponding elemental species of each expert model. In the learning process, specifically, for example, when performing the learning process with a certain learning data, the spectral data of the learning data is input to the input layer of the one-dimensional CNN processing unit, and in the output layer of the distance learning processing unit that gives the probability of each class, the correct label (the probability of the class of the material type corresponding to the input learning data is set to 1, and the probabilities of other classes are set to 0.) is given. In the output layer of the distance learning processing unit, a loss function is calculated, and learning for updating the parameters in the one-dimensional CNN processing unit and the distance learning processing unit may be performed according to the algorithm of the error backpropagation method. Here, the loss function may typically be the cross-entropy error, but is not limited thereto. Since the learning of the discriminator is supervised learning using the existing data of known materials as described above, the classes that can be classified by the discriminator, that is, the types of materials, are the types of known materials.

[0038] (2) Configuration of hierarchical distance learning in the distance learning unit of the expert model As described above, in the distance learning processing unit of the expert model of the present embodiment, the algorithm of "hierarchical distance learning" is adopted, and further improvement in the ability to identify the type of material from data may be achieved. In such an algorithm of hierarchical distance learning, as already mentioned, the arithmetic processing in the deep distance learning algorithm is executed continuously at least twice. Specifically, usually, in the deep distance learning algorithm, the logit of each class (k dimensions) is calculated from the cosine similarity between the feature vector (i dimensions) output by the one-dimensional CNN processing unit and the weight vector (for each class) and the hyperparameter, and the probability of each class is calculated as it is. On the other hand, in the case of hierarchical distance learning, as shown in FIG. 4(A), using the feature vector (i dimensions) output by the one-dimensional CNN processing unit, the logit (j dimensions) is calculated in the first distance learning processing unit, and then the logit is passed to the second distance learning processing unit, where the logit of each class (k dimensions) is calculated, and the probability of each class is calculated from such logits using a softmax function or the like (when the deep distance learning algorithm is further continued, the logit is passed to the next distance learning processing unit). Note that the j dimension of the logit calculated in the first distance learning processing unit may be arbitrarily set and may be the same as or different from the i dimension of the feature vector output by the one-dimensional CNN processing unit.

[0039] According to the above hierarchical distance learning algorithm, it is considered that data with similar features can be classified in more detail. Referring to the conceptual diagram of FIG. 4(B), at the stage where the feature vector output by the one-dimensional CNN processing unit of various data is mapped onto a virtual spherical surface, like each point (●, ▲, ■, etc.) depicted on the left of FIG. 4(B), the distance between the feature vectors of each data is not very far apart and the discrimination ability is not very high. However, the feature vector is transformed by the first distance learning process so that the distance between vectors with different features is separated, like each point depicted in FIG. 4(B). Further, by the second distance learning process, like each point depicted on the right of FIG. 4(B), the distance between vectors with similar features is also separated. That is, since the distance opens not only between vectors with different features but also between vectors with similar features, it is considered that the discrimination ability for vectors with similar features is improved thereby.

[0040] Note that the algorithm adopted in the hierarchical distance learning algorithm may be selected from any deep distance learning algorithm, for example, AdaCos, CosFace, ArcFace, SphereFace. According to the research of the inventor of the present invention, it has been found that very good discrimination accuracy can be achieved by using AdaCos for the first distance learning process and CosFace for the second distance learning process. In that case, the feature vector passed from each expert model to the sparse gate circuit unit, which will be described in detail later, may be the logit in AdaCos.

[0041] As described above, by adopting the one-dimensional CNN and the deep distance learning algorithm in the configuration of each expert model, the discrimination ability to identify the type of material from the spectral data of the material is improved.

[0042] (3) Configuration of the sparse gate circuit unit Referring to FIG. 2 again, in the sparse gate circuit unit, a weighted linear combination of the feature vectors calculated by each expert model is calculated. Specifically, the weighted linear combination may be given by the following formula. X i = Σw 元素 i · x 元素 i …(1) Here, x 元素 i is the i-th component of the feature vector calculated by the expert model for each element type, w 元素 i is the i-th component of the weight vector multiplied by the feature vector of each element type, and X i is the i-th component of the combined feature vector, and Σ is the sum over all element types to be inspected. The weight vector multiplied by the feature vector of each element type is determined by the learning process described later.

[0043] (4) Configuration of the MLP model part The MLP model part is configured to calculate, for each element type of all element types to be inspected, the probability that the element type is contained in the material, from the combined feature vector received from the sparse gate circuit part at the vector input layer, using the algorithm of a neural network (fully connected layer). Then, an element type whose such probability exceeds a predetermined value that may be arbitrarily set is determined to be contained in the material, and an element type whose probability is below the predetermined value is determined not to be contained in the material. More specifically, in the operation of the algorithm of the neural network, each component of the combined feature vector is assigned to each neuron of the vector input layer, at the output side of the network, a logit for each element type is calculated, and from each logit, for example, using the softmax function, the probability for each element type is calculated (probability output layer), and at the determination output layer, a determination for each element type is output according to the magnitude relationship between the probability and the predetermined value. The predetermined value for determining the presence or absence of an element type from the probability may be, for example, 0.5, but is not limited thereto.

[0044] (5) Learning process of the sparse gate circuit part and the MLP model part The learning processes of the sparse gate circuit section and the MLP model section may be executed by means of a so-called "End-to-end training" method using existing data of known materials as learning data. Here, the known materials are all types of materials that can be inspection targets. Also, since each known material can naturally contain multiple types of elements, in the probability output layer of the MLP model section, probabilities can exceed a predetermined value at multiple neurons simultaneously. That is, the MLP model section is configured as a discriminator that solves a multi-class multi-label problem. In the learning process, specifically, for example, when executing the learning process with certain learning data, the spectral data of the learning data is input to the vector input layer of each expert model, and in the probability output layer that gives the probabilities of each class (element type) of the MLP model section, the correct label (assuming the probability of the class of the element type contained in the material corresponding to the input learning data = 1 and the probability of the class of the element type not contained = 0.) is given, and in the probability output layer of the MLP model section, a loss function is calculated, and learning to update the parameters in the MLP model section and the sparse gate circuit section may be executed according to the algorithm of the error backpropagation method. Here, the loss function may typically be cross-entropy error, but is not limited to this. Note that in the learning stage of the MLP model section and the sparse gate circuit section, the parameters in each expert model are not updated.

[0045] Generation of extended data by physical-based data augmentation The number of available existing data of existing materials used for learning of the identifier is, in most cases, small for each material. For example, as for the data of inorganic substances registered in ICSD, usually one to two standard data are registered for each material. On the other hand, as already described, in actual measurement, various fluctuations occur in the measured values, and in the measurement data, compared with the existing data, variables such as the angle and frequency at which peaks occur shift (shift of peak position), fluctuations in the intensity ratio of multiple peaks, disappearance of peaks, splitting of peaks, etc. may occur. Therefore, if only one to two standard data registered in the database for each material are used as the data for learning of the identifier, it becomes difficult to accurately classify and identify the data with the above-mentioned fluctuations.

[0046] Therefore, in the present embodiment, preferably, data obtained by artificially giving various fluctuations that may occur in actual measurement to the existing data of existing materials is generated on a computer (physical-based data augmentation process), and these data may be used as learning data together with the existing data (original data) of existing materials. According to such a configuration, it is expected that actual measurement data having various fluctuations different from the standard data for a certain material can be identified with higher accuracy, and the type of the material and the estimation of the element types contained therein can be made possible.

[0047] In the above physical-based data augmentation process, specifically, the following various fluctuation-giving processes may be applied to the original data (I(θ) is the spectral value at the variable θ, and θ pr , θ ps etc. are variables that give peaks.). (i) Peak shift process (see Fig. 5(i)) - Randomly shift the position of the data with respect to the variable so that the variable at which the peak observed in the original data occurs shifts. I(θ) → I(θ ± Δ r ) [Δ r is the random shift amount] (ii) Peak intensity ratio change process (see Fig. 5(ii)) - Randomly change the intensity ratios of multiple peaks observed in the original data. I(θ pr ):I(θ ps )=α r :β r [α r :β r is the intensity ratio randomly set] (iii) Peak disappearance process (see Fig. 5(iii)) - Randomly eliminate the peaks observed in the original data. I(θ pr )→0 (iv) Peak splitting process (see Fig. 5(iv)) - Randomly split the peaks observed in the original data into two or more. I(θ pr )→I(θ pr1 )+I(θ pr2 ) The degree of giving each of the above variations may be appropriately set so as to be the same as the variations that can occur in actual measurements. Note that the process of giving the above variations may be executed in any combination for one existing data. The data generated by the physical-based data augmentation process is referred to as "extended data".

[0048] In principle, the more the above extended data, the higher the discrimination or classification accuracy. However, when learning is executed with a certain amount of data, the accuracy no longer changes. Therefore, it is useless to continue generating new extended data and learning processing from that state. Thus, in this embodiment, in the learning process, while verifying the accuracy, the cycle of generating new extended data and executing the learning process is repeated until almost no change in accuracy is observed. In one learning process cycle, several, for example, 5, pieces of extended data may be generated for each existing data and used as learning data together with the existing data.

[0049] Thus, in the configuration of the present embodiment described above, a MoE model composed of a plurality of expert models, a sparse gate circuit, and an MLP model, which are learned to be distinguishable for materials each containing a specific elemental species, is adopted. According to such a configuration, it is expected that it will be possible to identify a number of classes that cannot be identified by a discriminator consisting of a single CNN, so it will be possible to estimate the elemental species contained in a very large number of materials.

[0050] Although not shown, in the discriminator composed of the group of the above-mentioned expert models, a sparse gate circuit, and an MLP model, the MLP model may be configured to identify the type of the input material. In that case, the class identified in the output layer of the MLP model will be the type of material that can be identified by this system. In the learning process, as the correct label, the probability of the class of the material type corresponding to the input learning data is set to 1, and the probability of the other classes is set to 0.

[0051] Experimental example According to the teaching of the present embodiment described above, among the powder X-ray diffraction spectrum data provided by the ICSD provided by the National Institute of Standards and Technology, etc. in the United States, a discriminator for estimating the elemental species contained in the material from the spectrum data is constructed using the spectrum data of 6,800 lithium-containing materials, and the estimation accuracy of the elemental species contained in the material is detected to verify the effectiveness of the present embodiment. It should be understood that the following experimental examples illustrate the effectiveness of the present embodiment and do not limit the scope of the present invention.

[0052] In the configuration of the discriminator in the experiment, the expert model for each element has a configuration including a one-dimensional CNN processing unit and a distance learning processing unit. The distance learning processing unit is configured to execute the algorithm of hierarchical distance learning. The spectral data uses data in the range of 0 to 120° for 2θ, and the spectral intensity values at intervals of 0.02 degrees are input to each neuron of the input layer consisting of 6000 neurons in the one-dimensional CNN processing unit. The above 1D-RegNet was used for the algorithm of the one-dimensional CNN. The feature vector output by the one-dimensional CNN processing unit was 1024-dimensional. The distance learning processing unit of the expert model was configured to execute hierarchical distance learning that first executes the AdaCos algorithm and then executes the CosFace algorithm. In AdaCos, vectors of 1024 dimensions are input and output on both the input side and the output side. In CosFace, it receives a 1024-dimensional vector on the input side and outputs the logits of the classes of the number of material types containing each element on the output side. The calculation of the probability of each class from the CosFace logits was performed using the softmax function. And it was used as a feature model that passes the output of AdaCos of each expert model to the sparse gate circuit unit. The expert model was configured for each of a total of 75 element types up to element number 83 excluding hydrogen (H), oxygen (O), technetium (Tc), and noble gases (He, Ne, Ar, Kr, Xe).

[0053] In the determination of whether the MLP model part contains an element type, when the inclusion probability of each element type exceeded 0.5, it was determined that it contains it.

[0054] As learning data, existing data of each material obtained from ICSD and extended data generated from the existing data by the above-described physics-based data augmentation process were used. Five pieces of extended data were generated in each learning cycle and used as learning data. Each piece of extended data was generated by applying all of the above four variation-giving processes. And, 70% of the learning data was used as training data, 15% as validation data, and 15% as test data. The training data was used for updating the weight parameters by the error backpropagation method, the validation data was used for verifying the generality of the learned model and adjusting the hyperparameters, and the test data was used for detecting the identification accuracy. The learning cycle was executed for 150 epochs (the extended data was changed for each epoch).

[0055] As a result of the above computational experiment, the estimation accuracy was as follows.

Table 1

[0056] As understood from Table 1, it was shown that, with the configuration of the present embodiment, the element types contained in the material can be estimated with a high accuracy of about 77% from the spectral data of the material. Also, from the above results, it was shown that, in the expert model, when the distance learning process is executed, the estimation accuracy is remarkably improved, and that, as the configuration of the sparse gate circuit unit, when the weighted linear combination operation described above is adopted, the estimation accuracy is improved.

[0057] Thus, according to the above-described present embodiment, it becomes possible to estimate the element types contained in the material from the spectral data of the material.

[0058] The above description has been made in relation to embodiments of the present invention, but many modifications and changes are easily possible for those skilled in the art, and it is obvious that the present invention is not limited only to the embodiments illustrated above, and can be applied to various devices without departing from the concept of the present invention.

Claims

【Claim 1】 A system for estimating the elemental species contained in a material from the spectral data of the material, comprising: a data input layer into which the spectral data of the material is input, and a vector output layer that outputs an expert model feature vector of a predetermined dimension, the expert model being provided for each elemental species for calculating the expert model feature vector based on the spectral data; a sparse gate circuit unit that outputs a combined feature vector, which is a vector generated by a weighted linear combination of the expert model feature vectors output by the vector output layer of the expert model for each elemental species; a vector input layer into which the combined feature vector is input, and a determination output layer that outputs a determination of the presence or absence of the element in the material for each elemental species, the contained element determination multi-layer perceptron unit that outputs a determination as to whether or not each of the elemental species is contained in the material based on the combined feature vector; and the expert model has a one-dimensional CNN processing unit and a distance learning processing unit; the one-dimensional CNN processing unit is configured to calculate and output a first feature vector by an algorithm of a one-dimensional convolutional neural network based on the spectral data when the spectral data of the material is input; the distance learning processing unit is configured to calculate the expert model feature vector based on the first feature vector by a hierarchical distance learning algorithm that continuously executes the algorithm of deep distance learning twice when the first feature vector is input.

Citation Information

Patent Citations

  • Mixture of Expert Neural Networks

    JP2019537133A

  • Method, device, and program for estimating component of sample, method for learning, and learning program

    JP2021076411A

  • Data analysis system and data analysis method

    JP2021092467A

  • X-ray analysis device, x-ray analysis system, analysis method, and program

    WO2020121918A1