Learning method, discrimination method, learning device, discrimination device, and computer program

By dividing training data into groups and using multiple vector neural network models, the method addresses the inefficiencies in learning and classification times for large data sets, enhancing accuracy and efficiency.

JP7790123B2Active Publication Date: 2025-12-23SEIKO EPSON CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021200573
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-12-23
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

Conventional machine learning models face increased learning and classification times when dealing with large amounts of data.

Method used

Implement a method involving multiple vector neural network-type machine learning models with multiple vector neuron layers, where training data is divided into groups to enhance learning efficiency and classification accuracy.

Benefits of technology

This approach reduces learning and classification times while improving the accuracy of data classification by utilizing divided training data groups and feature spectra calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007790123000003
    Figure 0007790123000003
  • Figure 0007790123000004
    Figure 0007790123000004
  • Figure 0007790123000005
    Figure 0007790123000005
Patent Text Reader

Abstract

To provide a technique of suppressing increase of the learning time of a mechanical learning model or increase of time for inputting determination target data into the mechanical learning model and determining a class.SOLUTION: A method for learning includes: the step (a) of preparing plural pieces of learning data; (b) dividing the plural pieces of learning data into at least one data and generating at least one input learning data group; and (c) learning M number of mechanical learning models. The step (b) includes one of the step (b1) of dividing each of the plural pieces of input data into at least one region and generating the aggregate of first-type division input data after the division which belongs to the same region, as one input learning data group and the step (b2) of dividing plural pieces of learning data into at least one data and generating the aggregate of second-type division input data after the division, as one input learning data group.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technique for classifying data to be classified using a machine learning model. [Background technology]

[0002] Patent Documents 1 and 2 disclose a vector neural network type machine learning model that uses vector neurons, called a capsule network. A vector neuron is a neuron whose input and output are vectors. A capsule network is a machine learning model that uses vector neurons called capsules as network nodes. A vector neural network type machine learning model such as a capsule network can be used to classify input data. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] U.S. Patent No. 5,210,798 [Patent Document 2] International Publication No. 2019 / 083553 Summary of the Invention [Problem to be solved by the invention]

[0004] In conventional technology, when the amount of data used for learning or the data to be classified is large, the learning time of the machine learning model or the time required to input the data to be classified into the machine learning model and perform class classification may become long. [Means for solving the problem]

[0005] According to a first aspect of the present disclosure, there is provided a learning method for M machine learning models (M is an integer equal to or greater than 2) used to distinguish the class of data to be distinguished, the M machine learning models being of a vector neural network type having multiple vector neuron layers. This learning method includes the steps of: (a) preparing a plurality of training data sets each having input data and a prior label associated with the input data; (b) dividing the plurality of training data sets into one or more parts to generate the one or more input training data groups; and (c) training the M machine learning models by inputting the corresponding input training data groups into each of the M machine learning models so as to reproduce the correspondence between the input data and the prior label associated with the input data. The step (b) includes one of the steps of: (b1) dividing each of the plurality of input data sets into one or more regions and generating a set of first-type divided input data after division that belong to the same region as the input training data group; and (b2) dividing the plurality of training data sets belonging to one class into one or more parts and generating a set of second-type divided input data after division as the input training data group.

[0006] According to a second aspect of the present disclosure, there is provided a classification method for classifying a class of data to be classified using M (M is an integer equal to or greater than 2) machine learning models of a vector neural network type having a plurality of vector neuron layers.This discrimination method includes: (a) a step of preparing the M machine learning models trained using a plurality of training data each having input data and a prior label associated with the input data, wherein each of the M machine learning models is trained by dividing the plurality of training data into one or more input training data groups and using the input training data group corresponding to one of the one or more divided input training data groups; (b) a step of preparing M sets of known feature spectra associated with the M trained machine learning models, wherein the M sets of known feature spectra include known feature spectra obtained from the output of a specific layer among the plurality of vector neuron layers by inputting the input training data groups to the M trained machine learning models; and (c) a step of inputting discriminated data generated from the discriminated data into each of the M trained machine learning models, and obtaining, for each of the M machine learning models, individual data to be used for class discrimination of the discriminated data. and (d) performing class discrimination of the discriminated data using the M individual data obtained for each of the M machine learning models, wherein the individual data is generated using at least one of: (i) a similarity between the feature spectrum calculated from the output of the specific layer in response to the input of the discriminated data to the machine learning model and the group of known feature spectra; and (ii) an activation value corresponding to a judgment value of each class output from the output layer of the machine learning model in response to the input of the discriminated data; and (d) performing class discrimination of the discriminated data using the M individual data obtained for each of the M machine learning models, wherein the step (a) includes either: (a1) dividing each of the plurality of input data into one or more regions, and using a set of first-type divided input data after division that belongs to the same region as one of the input training data groups; or (a2) performing a division process to divide the plurality of training data belonging to one class into one or more parts, and using a set of second-type divided input data after the division process as one of the input training data groups.

[0007] According to a third aspect of the present disclosure, there is provided a learning device for M machine learning models (M is an integer greater than or equal to 2) used to distinguish the class of data to be distinguished, the M machine learning models being of a vector neural network type having multiple vector neuron layers. This learning device comprises a memory and a processor that performs learning of the M machine learning models, and the processor performs the following processes: dividing a plurality of training data sets, each having input data and a prior label associated with the input data, into one or more groups to generate the one or more input training data groups; and training the M machine learning models by inputting the corresponding input training data groups into each of the M machine learning models so as to reproduce the correspondence between the input data and the prior label associated with the input data. The process of generating the one or more input training data groups includes either a process of dividing each of the plurality of input data sets into one or more regions and generating a set of first-type divided input data after division that belong to the same region as the one input training data group; or a process of dividing the plurality of training data sets belonging to one class into one or more regions and generating a set of second-type divided input data after division as the one input training data group.

[0008] According to a fourth aspect of the present disclosure, there is provided a classification device that classifies the class of data to be classified using M (M is an integer equal to or greater than 2) machine learning models of a vector neural network type having multiple vector neuron layers.This discrimination device includes: a memory that stores M machine learning models trained using a plurality of training data each having input data and a prior label associated with the input data, wherein each of the M machine learning models is trained by dividing the plurality of training data into one or more input training data groups and using the input training data group corresponding to one of the divided one or more input training data groups; and a processor that inputs the data to be discriminated into the M machine learning models and performs class discrimination of the data to be discriminated, wherein the processor performs a process of generating M sets of known feature spectra associated with the trained M machine learning models, wherein the M sets of known feature spectra include known feature spectra obtained from the output of a specific layer out of the plurality of vector neuron layers by inputting the input training data groups to the trained M machine learning models; a process of inputting data and obtaining, for each of the M machine learning models, individual data to be used for class discrimination of the discriminated data, wherein the individual data is generated for each of the M machine learning models using at least one of: (i) a similarity between a feature spectrum calculated from the output of the specific layer in response to the input of the discriminated data to the machine learning model and the group of known feature spectra; and (ii) an activation value corresponding to a judgment value of each class output from the output layer of the machine learning model in response to the input of the discriminated data; and a process of performing class discrimination of the discriminated data using the M individual data obtained for each of the M machine learning models, wherein the input training data group is either a set of first-type divided input data after dividing each of the multiple input data into one or more regions and dividing the multiple training data belonging to one class into one or more regions, or a set of second-type divided input data after the division process.

[0009] According to a fifth aspect of the present disclosure, there is provided a computer program that causes a processor to execute training of M (M is an integer equal to or greater than 2) machine learning models used to classify data to be discriminated, the M machine learning models being of a vector neural network type having a plurality of vector neuron layers. The computer program includes: (a) a function of dividing a plurality of training data sets, each of the training data sets having input data and a prior label associated with the input data, into one or more groups to generate the one or more input training data groups; (b) A function of training the M machine learning models by inputting the corresponding input training data group into each of the M machine learning models so as to reproduce the correspondence between the input data and the prior label associated with the input data, wherein the function (a) includes either a function of dividing each of the multiple input data into one or more regions and generating a set of first-type divided input data after division that belong to the same region as the single input training data group, or a function of dividing the multiple training data belonging to one class into one or more regions and generating a set of second-type divided input data after division as the single input training data group.

[0010] According to a sixth aspect of the present disclosure, there is provided a computer program for causing a processor to identify a class of data to be identified using M (M is an integer greater than or equal to 2) machine learning models of a vector neural network type having multiple vector neuron layers.This computer program has the following functions: (a) a function for storing the M machine learning models trained using a plurality of training data each having input data and a prior label associated with the input data, wherein each of the M machine learning models is trained by dividing the plurality of training data into one or more input training data groups and using the input training data group corresponding to one of the one or more divided input training data groups; (b) a function for generating M sets of known feature spectra associated with the M trained machine learning models, wherein the M sets of known feature spectra include known feature spectra obtained from the output of a specific layer among the plurality of vector neuron layers by inputting the input training data groups to the M trained machine learning models; and (c) a function for inputting input discriminated data generated from the discriminated data to each of the M trained machine learning models, and generating the discriminated feature spectra for each of the M machine learning models. a function for obtaining individual data to be used for class discrimination of other data, wherein the individual data is generated, for each of the M machine learning models, using at least one of: (i) a similarity between the feature spectrum calculated from the output of the specific layer in response to input of the input discriminated data to the machine learning model and the group of known feature spectra; and (ii) an activation value corresponding to a judgment value of each class output from the output layer of the machine learning model in response to input of the input discriminated data; and (d) a function for performing class discrimination of the discriminated data using the M individual data obtained for each of the M machine learning models, wherein the input training data group is either a set of first-type divided input data after dividing each of the multiple input data into one or more regions and the divided input data belonging to the same region, or a set of second-type divided input data after performing a division process in which the multiple training data belonging to one class are divided into one or more regions. [Brief explanation of the drawings]

[0011] [Figure 1]FIG. 1 is a block diagram showing a class determination system according to a first embodiment. [Figure 2] FIG. 2 is a block diagram showing the functions of a discrimination device. [Figure 3] FIG. 1 is an explanatory diagram showing the configuration of a machine learning model. [Figure 4] FIG. 10 is an explanatory diagram showing another configuration of a machine learning model. [Figure 5] A flowchart showing the learning process for M machine learning models. [Figure 6] FIG. 10 is a diagram showing first data processing. [Figure 7] FIG. 10 is a diagram showing input learning data. [Figure 8] 10 is a flowchart showing a preparatory step for preparing a group of known characteristic spectra. [Figure 9] FIG. [Figure 10] FIG. 4 is an explanatory diagram showing how a group of known characteristic spectra is created. [Figure 11] FIG. 2 is an explanatory diagram showing the configuration of a group of known characteristic spectra. [Figure 12] 10 is a flowchart showing a class discrimination process for data to be discriminated. [Figure 13] 13 is a detailed flowchart of step S36 in FIG. 12. [Figure 14] Figure 1 explains the classification process. [Figure 15] Figure 2 explains the classification process. [Figure 16] FIG. 10 is a diagram for explaining another embodiment 1 of the class discrimination process. [Figure 17] FIG. 10 is a diagram for explaining another embodiment 2 of the class discrimination process. [Figure 18] FIG. 10 is a diagram for explaining another embodiment 3 of the class discrimination process. [Figure 19] FIG. 10 is a diagram for explaining another embodiment of the pre-determined class generation process. [Figure 20] FIG. 4 is a diagram for explaining another embodiment of the first embodiment. [Figure 21] 10 is a flowchart showing a learning process according to the second embodiment. [Figure 22]A conceptual diagram of clustering. [Figure 23] FIG. 10 is a diagram for explaining step S12a. [Figure 24] FIG. 10 is an explanatory diagram showing a first calculation method M1 of class-specific similarity. [Figure 25] FIG. 10 is an explanatory diagram showing a second calculation method M2 of class-specific similarity. [Figure 26] FIG. 10 is an explanatory diagram showing a third calculation method M3 of class-specific similarity. DETAILED DESCRIPTION OF THE INVENTION

[0012] A. First embodiment: FIG. 1 is a block diagram illustrating a discrimination system according to a first embodiment. The discrimination system 5 includes a discrimination device 20 and a spectrometer 30. The spectrometer 30 is capable of performing spectroscopic measurement of a target object 10 to obtain its spectral reflectance. In the present disclosure, the spectral reflectance is also referred to as "spectroscopic data." The spectrometer 30 includes, for example, a tunable interference spectral filter and a monochrome image sensor. The spectral data obtained by the spectrometer 30 is used as discrimination data to be input into a machine learning model (described later). The discrimination device 20 performs class discrimination processing on the spectral data using the machine learning model to determine which of multiple classes the target object 10 belongs to. The "class of the target object 10" refers to the type of the target object 10. The discrimination device 20 may output the determined type of the target object 10 to a display, which is an output unit. This allows the user to easily understand the type of the target object 10. The discrimination system according to the present disclosure can also be realized as a system other than the above, for example, a system that performs class discrimination using a discriminated image, one-dimensional data other than spectroscopic data, a spectroscopic image, time-series data, etc. as discriminated data.

[0013] 2 is a block diagram showing the functions of the discrimination device 20. The discrimination device 20 has a processor 110, a memory 120, an interface circuit 130, and an input device 140 and a display unit 150 connected to the interface circuit 130. The discrimination device 20 is, for example, a personal computer. The spectrometer 30 is also connected to the interface circuit 130. For example, but not limited to, the processor 110 not only has the function of executing the processes described in detail below, but also has the function of displaying data obtained by the processes and data generated during the processes on the display unit 150.

[0014] The processor 110 functions as a data generation unit 112 that generates data to be input to the machine learning model 200 from input data IM, such as each data in the input training data group IDG used for training the machine learning model 200 and the discriminated data IM, and a class discrimination processing unit 114 that executes class discrimination processing for the discriminated data IM. The data generation unit 112 and the class discrimination processing unit 114 are realized by the processor 110 executing a computer program stored in the memory 120.

[0015] The data generation unit 112 generates the input learning data group IDG by performing one of the following two data processes. <1> First data processing: Each of the plurality of input data IM is divided into M or more regions, and a set of divided first type divided input data IDa belonging to the same region is generated as one input learning data group IDG. <2> Secondary Data Processing: A division process is executed to divide a plurality of pieces of input data IM belonging to one class into M or more pieces, and a set of second-type divided input data IDb after division is generated as one input learning data group IDG.

[0016] The class discrimination processing unit 114 includes a similarity calculation unit 310 and a comprehensive judgment unit 320. The class discrimination processing unit 114 inputs discrimination target data IM to M machine learning models 200, and performs class discrimination of the discrimination target data IM using a plurality of individual data DD obtained for each of the M machine learning models 200. This will be described in detail later. Note that the symbol IM is attached to the data input to the machine learning model 200 regardless of the type of data.

[0017] In the above, at least some of the functions of the data generation unit 112 and the class discrimination processing unit 114 may be realized by a hardware circuit. The term "processor" in this specification also includes such hardware circuits. Furthermore, the processor that executes the class discrimination process may be a processor included in a remote computer connected to the discrimination device 20 via a network.

[0018] The memory 120 stores multiple machine learning models 200, a training data set TDG, multiple input training data sets IDG, and multiple known feature spectrum sets KSp. The machine learning models 200 are used in processing by the class discrimination processing unit 114. Each of the multiple machine learning models 200 is a vector neural network type machine learning model having multiple vector neuron layers. An example configuration and operation of the machine learning models 200 will be described later. When the number of machine learning models 200 is represented by M, M can be set to any integer equal to or greater than 2. In this embodiment, a case will be described in which five machine learning models 200 are used. When the five machine learning models 200 are used to distinguish them from one another, "_T (T is an integer from 1 to 5)" is added to the end of each model name. In other words, the five machine learning models 200 are machine learning models 200_1 to 200_5. The five machine learning models 200 are also referred to as a first model 200_1, a second model 200_2, a third model 200_3, a fourth model 200_4, and a fifth model 200_5.

[0019] The training data group TDG is a collection of training data TD, which are teacher data. In this embodiment, each training data TD in the training data group TDG has spectral data as input data and a priori label LB associated with the spectral data. In this embodiment, the priori label LB is a label indicating the type of the target object 10. In this embodiment, "label" and "class" have the same meaning. The input training data group IDG is a data group generated by the data generation unit 112 using the training data group TDG. M or more groups are generated as the input training data group IDG. The input training data group IDG is generated by dividing the multiple training data TD that make up the training data group TDG into M or more groups. In this embodiment, the number of input training data groups IDG is M, which is the same as the number of machine learning models 200. When the five input training data groups IDG are used to distinguish one another, "_T (T is an integer from 1 to 5)" is added to the end of the name. That is, the five input training data groups IDG are input training data groups IDG_1 to IDG_5.

[0020] The known feature spectrum group KSp is a set of feature spectra obtained when the training data group TDG is input to the trained machine learning model 200. The feature spectrum will be described later. The training data group TDG and the known feature spectrum group KSp that correspond to each machine learning model 200 are used.

[0021] 3 is an explanatory diagram showing the configuration of a machine learning model 200. This machine learning model 200 includes, in order from the input data IM side, a convolutional layer 210, a primary vector neuron layer 220, a first convolutional vector neuron layer 230, a second convolutional vector neuron layer 240, and a classification vector neuron layer 250, which is an output layer. Of these five layers 210 to 250, the convolutional layer 210 is the lowest layer, and the classification vector neuron layer 250 is the highest layer. In the following description, the layers 210 to 250 are also referred to as the "Conv layer 210," the "PrimeVN layer 220," the "ConvVN1 layer 230," the "ConvVN2 layer 240," and the "classVN layer 250," respectively.

[0022] In this embodiment, the input data IM is spectral data, and therefore is a one-dimensional array of data. For example, the input data IM is data in which 36 representative values ​​are extracted at intervals of 10 nm from spectral data in the range of 380 nm to 730 nm.

[0023] 3, two convolution vector neuron layers 230 and 240 are used, but the number of convolution vector neuron layers is arbitrary, and a convolution vector neuron layer may be omitted. However, it is preferable to use one or more convolution vector neuron layers.

[0024] The configuration of each of the layers 210 to 250 in FIG. 3 can be described as follows. <Description of the configuration of machine learning model 200> ·Conv layer 210: Conv[32,6,2] ·PrimeVN layer 220: PrimeVN[26,1,1] ·ConvVN1 layer 230:ConvVN1[20,5,2] ·ConvVN2 layer 240:ConvVN2[16,4,1] ·classVN layer 250:classVN[Nm,3,1] Vector dimension VD: VD=16 In the description of each of these layers 210 to 250, the character string before the parentheses is the layer name, and the numbers in the parentheses are, in order, the number of channels, the kernel surface size, and the stride. For example, the layer name of the Conv layer 210 is "Conv," the number of channels is 32, the kernel surface size is 1 x 6, and the stride is 2. In FIG. 3, these descriptions are shown below each layer. The hatched rectangles drawn in each layer represent the kernel surface size used when calculating the output vector of the adjacent higher layer. In this embodiment, since the input data IM is a one-dimensional array of data, the kernel surface size is also one-dimensional. Note that the parameter values ​​used in the description of each of the layers 210 to 250 are merely examples and can be changed as desired.

[0025] The Conv layer 210 is a layer composed of scalar neurons. The other four layers 220 to 250 are layers composed of vector neurons. A vector neuron is a neuron that uses vectors as input and output. In the above description, the dimension of the output vector of each vector neuron is constant at 16. In the following, the term "node" is used as a superordinate concept of scalar neurons and vector neurons.

[0026] FIG. 3 shows the first axis x and second axis y that define the planar coordinates of the node array for the Conv layer 210, and the third axis z that represents depth. It also shows that the sizes of the Conv layer 210 in the x, y, and z directions are 1, 16, and 32, respectively. The sizes in the x and y directions are called "resolution." In this embodiment, the resolution in the x direction is always 1. The size in the z direction is the number of channels. These three axes x, y, and z are also used in other layers as coordinate axes that indicate the position of each node. However, in FIG. 3, these axes x, y, and z are omitted from illustration in layers other than the Conv layer 210.

[0027] As is well known, the resolution W1 in the y direction after convolution is given by the following equation: W1=Ceil{(W0-Wk+1) / S} (1) Here, W0 is the resolution before convolution, Wk is the surface size of the kernel, S is the stride, and Ceil{X} is a function that rounds up the decimal point of X. The resolution of each layer shown in FIG. 3 is an example in which the resolution in the y direction of the input data IM is set to 36, and the actual resolution of each layer is changed appropriately depending on the size of the input data IM.

[0028] The classVN layer 250 has Nm channels. In the example of FIG. 3, Nm=3. Generally, Nm is an integer equal to or greater than 2 and represents the number of known classes that can be distinguished using the machine learning model 200. The number of distinguishable classes Nm can be set to a different value for each machine learning model 200. The total number of classes that can be distinguished by the M machine learning models 200 is represented by ΣNm. The three channels of the classVN layer 250 output judgment values ​​class1 to class3 for the three known classes. Typically, the class with the largest value among these judgment values ​​class1 to class3 is used as the class discrimination result for the data IM. Meanwhile, in this embodiment, the judgment value C output for each of the five machine learning models 200 constitutes one element of the individual data DD. The comprehensive judgment unit 320 then performs class discrimination of the data to be discriminated IM using the individual data DD for each machine learning model 200. This will be described in detail later. The judgment values ​​class1 to class3 are also referred to as activation values ​​a.

[0029] FIG. 3 also illustrates subregions Rn in each layer 210, 220, 230, 240, and 250. The subscript "n" in subregion Rn refers to the reference number of the layer. For example, subregion R210 indicates a subregion in the Conv layer 210. A "subregion Rn" is a region in each layer that is identified by a planar position (x, y) defined by the position of the first axis x and the position of the second axis y, and that includes multiple channels along the third axis z. The subregion Rn has dimensions of "Width" × "Height" × "Depth," corresponding to the first axis x, the second axis y, and the third axis z. In this embodiment, the number of nodes included in one "subregion Rn" is "1 × 1 × depth count," i.e., "1 × 1 × number of channels."

[0030] 3, a feature spectrum Sp_ConvVN1 (described later) is calculated from the output of the ConvVN1 layer 230 and input to the similarity calculation unit 310. Similarly, feature spectra Sp_ConvVN2 and Sp_classVN are calculated from the outputs of the ConvVN2 layer 240 and the classVN layer 250, respectively, and input to the similarity calculation unit 310. The similarity calculation unit 310 calculates a class-specific similarity Sclass (described later) using these feature spectra Sp_ConvVN1, Sp_ConvVN, and Sp_classVN and a known feature spectrum group KSp that has been created in advance.

[0031] In this disclosure, the vector neuron layer used to calculate the similarity is also referred to as the "specific layer." Any number of vector neuron layers, one or more, can be used as the specific layer. The configuration of the feature spectrum Sp and the method of calculating the similarity S using the feature spectrum Sp will be described later.

[0032] Figure 4 is an explanatory diagram showing another configuration of the machine learning model 200. This machine learning model 200 differs from the machine learning model 200 in Figure 3, which uses one-dimensional array data, in that the input data IM is two-dimensional array data. The configuration of each layer 210 to 250 in Figure 4 can be described as follows. <Description of each layer configuration> ·Conv layer 210: Conv[32,5,2] ·PrimeVN layer 220: PrimeVN[16,1,1] ·ConvVN1 layer 230:ConvVN1[12,3,2] ·ConvVN2 layer 240:ConvVN2[6,3,1] ·classVN layer 250:classVN[Nm,4,1] Vector dimension VD: VD=16

[0033] The machine learning model 200 shown in Fig. 4 can be used in, for example, a classification system that performs class classification on a target image. However, in the following description, the machine learning model 200 shown in Fig. 3 is used.

[0034] When the discrimination data IM is input to the five machine learning models 200, a feature spectrum Sp is calculated from a specific layer of each of the five machine learning models 200 and input to the similarity calculation unit 310. The similarity calculation unit 310 calculates a class-specific similarity Sclass, which is the similarity between the feature spectrum Sp and a known feature spectrum group KSp of the corresponding specific layer.

[0035] 5 is a flowchart showing the learning process of M machine learning models 200. In step S10, a plurality of pieces of training data TD are prepared, each of which includes spectral data serving as input data IM and a priori labels LB associated with the input data IM. That is, a spectrometer 30 performs spectroscopic measurement on a target object 10 whose type is known in advance to obtain spectral data, and this spectral data is used as input data IM. Furthermore, the priori labels LB corresponding to the known types are associated with the input data IM to generate training data TD.

[0036] In step S12, the data generation unit 112 executes a first data processing. Specifically, the data generation unit 112 divides the plurality of pieces of training data TD prepared in step S10 to generate M input training data groups IDG. An example of the data generation unit 112 executing step S12 through the first data processing will be described with reference to FIGS. 6 and 7.

[0037] FIG. 6 is a diagram illustrating the first data processing. FIG. 7 is a diagram illustrating input learning data ID. The data generation unit 112 divides a single piece of input data IM into five regions R1 to R5, thereby obtaining divided input data IMa_1 to IMa_5 corresponding to each region R1 to R5. The five regions R1 to R5 may be configured so as not to overlap with each other, or some elements of adjacent regions R1 to R5 may overlap. In this embodiment, the data generation unit 112 equally divides the wavelength range so that the wavelength λ range of 380 nm to 730 nm is of equal length. As shown in FIG. 7, each of the divided input data IMa_1 to IMa_5 is associated with the a priori label LB that was associated with the input data IM from which the data was divided, thereby obtaining first-type divided input data IDa_1 to IDa_5. Note that when the divided input data IMa_1 to IMa_5 are used without distinction, the divided input data IMa is used. Furthermore, when the first type divided input data IDa_1 to IDa_5 are used without distinction, the first type divided input data IDa is used.

[0038] The data generation unit 112 divides a single piece of input data IM for each piece of input data IM, thereby generating an input learning data group IDG, which is a set of first-type divided input data IDa that belong to the same region.

[0039] As shown in FIG. 5, in step S14 after step S12, the processor 110 trains the M machine learning models 200_1 to 200_5 by inputting corresponding input training data groups IDG_1 to IDG_5 to each of the M machine learning models 200_1 to 200_5. Specifically, the processor 110 trains the machine learning model 200 so as to reproduce the correspondence between the divided input data IMa as input data IM and the prior label LB associated with the divided input data IMa. Note that the machine learning model 200 and the input training data group IDG having the same last digit correspond to each other. In step S12, when the training of the machine learning model 200 is completed, the trained machine learning model 200 is stored in the memory 120.

[0040] 8 is a flowchart showing the advance preparation process for preparing a known feature spectrum group KSp. When a trained machine learning model 200 is stored in memory 120, first, in step S20, processor 110 generates a known feature spectrum group KSp by inputting the corresponding input training data group IDG used in training to M trained machine learning models 200_1 to 200_5. In step S22, processor 110 stores the known feature spectrum group KSp generated in step S20 in memory 120. The known feature spectrum group KSp is a collection of feature spectra Sp, which will be described below.

[0041] FIG. 9 is an explanatory diagram showing a feature spectrum Sp obtained by inputting arbitrary data into the trained machine learning model 200. Here, the feature spectrum Sp obtained from the output of the ConvVN1 layer 230 will be described. The horizontal axis in FIG. 9 represents the position of a vector element in the output vector of multiple nodes included in one subregion R230 of the ConvVN1 layer 230. The position of this vector element is represented by a combination of the element number ND of the output vector of each node and the channel number NC. In this embodiment, the vector dimension is 16, which is the number of elements of the output vector output by each node, so the element number ND of the output vector is 16, ranging from 0 to 15. Furthermore, since the ConvVN1 layer 230 has 20 channels, the channel number NC is 20, ranging from 0 to 19. In other words, this feature spectrum Sp is obtained by arranging multiple element values ​​of the output vector of each vector neuron included in one subregion R230 across multiple channels along the third axis z.

[0042] The vertical axis of Fig. 9 represents the feature value C at each spectral position. V In this example, the feature value C V is the value of each element of the output vector V ND Note that the feature value C V As the value of each element of the output vector V NDAlternatively, the normalization coefficient may be used as is. In the latter case, the feature value C included in the feature spectrum Sp is V The number of is equal to the number of channels, which is 20. The normalization coefficient is a value corresponding to the vector length of the output vector of the node.

[0043] The number of feature spectra Sp obtained from the output of the ConvVN1 layer 230 for one data item is six, because it is equal to the number of planar positions (x, y) of the ConvVN1 layer 230, i.e., the number of subregions R230. Similarly, three feature spectra Sp are obtained from the output of the ConvVN2 layer 240 for one data item, and one feature spectrum Sp is obtained from the output of the classVN layer 250.

[0044] When the divided input data IMa of the input training data group IDG is input again to the trained machine learning model 200, the similarity calculation unit 310 calculates the feature spectrum Sp shown in FIG. 9 and stores it in the memory 120 as a known feature spectrum group KSp.

[0045] FIG. 10 is an explanatory diagram showing how a known feature spectrum group KSp is created using an input training data group IDG. In this example, by inputting divided input data IMa, whose labels are 1 to 3, into the trained machine learning model 200, feature spectra KSp_ConvVN1, KSp_ConvVN2, and KSp_classVN associated with each label or class are obtained from the outputs of three vector neuron layers, namely, the ConvVN1 layer 230, the ConvVN2 layer 240, and the classVN layer 250. These feature spectra KSp_ConvVN1, KSp_ConvVN2, and KSp_classVN are stored in the memory 120 as a known feature spectrum group KSp. The similarity calculation unit 310 generates a known feature spectrum group KSp for each of the five trained machine learning models 200_1 to 200_5 and stores them in the memory 120.

[0046] Fig. 11 is an explanatory diagram showing the configuration of the known feature spectrum set KSp. In this example, the known feature spectrum set KSp_ConvVN1 obtained from the output of the ConvVN1 layer 230 of the machine learning model 200_1 is shown. The known feature spectrum set KSp_ConvVN2 obtained from the output of the ConvVN2 layer 240 and the known feature spectrum set KSp_ConvVN1 obtained from the output of the classVN layer 250 also have similar configurations, but are not shown in Fig. 11. Note that the known feature spectrum set KSp may be obtained from the output of at least one specific layer, which is a vector neuron layer.

[0047] Each record in the known feature spectrum group KSp_ConvVN1 includes a parameter m for distinguishing the M machine learning models 200, a parameter i indicating a label or class, a parameter j indicating a specific layer, a parameter k indicating a subregion Rn, a parameter q indicating a data number, and a known feature spectrum KSp associated with the various parameters i, j, k, and q. The known feature spectrum KSp is the same as the feature spectrum Sp in FIG. 9.

[0048] The class parameter i is classification information indicating which class the known feature spectrum KSp belongs to, and takes the same value as the label, 1 to 3. The specific layer parameter j takes a value of 1 to 3 indicating which of the three specific layers 230, 240, and 250 it belongs to. The subregion Rn parameter k takes a value indicating which of the multiple subregions Rn included in each specific layer it belongs to, i.e., which planar position (x, y) it belongs to. For the ConvVN1 layer 230, there are six subregions R230, so k = 1 to 6. The data number parameter q indicates the number of the post-division input data IMa with the same label, and takes a value of 1 to max1 for class 1, 1 to max2 for class 2, and 1 to max3 for class 3. A known feature spectrum KSp associated with the parameter i, which is classification information, is also referred to as a class-specific known feature spectrum KSp.

[0049] As described above, the known feature spectrum group KSp is obtained from the output of a specific layer by inputting the corresponding input training data group IDG to each of the M machine learning models 200_1 to 200_5.

[0050] The plurality of input training data groups IDG used in step S20 do not have to be the same as the plurality of input training data groups IDG used in step S14. However, if some or all of the input training data IDs used in step S14 are used in step S20 as well, there is an advantage in that there is no need to prepare new input training data IDs.

[0051] Fig. 12 is a flowchart showing the class discrimination process for discrimination target data IM. Fig. 13 is a detailed flowchart of step S36 in Fig. 12. Fig. 14 is a first diagram explaining the class discrimination process. Fig. 15 is a second diagram explaining the class discrimination process.

[0052] 12, in step S30, the data generation unit 112 generates input discriminated data IM_1 to IM_5 to be input to each of the machine learning models 200_1 to 200_5 from the discriminated data IM input to the discrimination device 20. Specifically, as shown in Fig. 14, the data generation unit 112 generates input discriminated data IM_1 to IM_5 from the discriminated data IM using a data processing method similar to that used to generate the input learning data ID. In this embodiment, the discriminated data IM is divided into five regions R1 to R5, and the data in each of the divided regions R1 to R5 becomes input discriminated data IM_1 to IM_5.

[0053] 12, in step S32, the processor 110 inputs the corresponding input discriminated data IM_1 to IM_5 generated from the discriminated data IM to each of the M trained machine learning models 200_1 to 200_5. Specifically, as shown in FIG. 14, the input discriminated data IM_1 to IM_5 of the same regions R1 to R5 as the regions R1 to R5 of the input training data ID used for training are input to the machine learning models 200_1 to 200_5. For example, the input discriminated data IM_1, which is data of region R1, is input to the machine learning model 200_1 trained using the divided input data IMa_1 for training region R1.

[0054] As shown in FIG. 12, in step S34, the similarity calculation unit 310 obtains individual data DD for each of the M trained machine learning models 200_1 to 200_5. As shown in FIGS. 14 and 15, the individual data DD has an activation value a and a similarity S corresponding to each class for each of the first model 200_1 to the fifth model 200_5. As described above, the activation value a is the judgment values ​​class1 to class3 output from the three channels of the classVN layer 250. The similarity S is an index indicating the degree to which the input discriminated data IM_1 to IM5 are similar to the divided input data IMa_1 to IMa_5 for each class. For example, the similarity S is a class-specific similarity Sclass corresponding to the class for which the activation value a is maximum for each of the first model 200_1 to the fifth model 200_5. When there are multiple specific layers, the similarity calculation unit 310 calculates a multiplication value for each of the multiple specific layers by multiplying a weighting coefficient set for each of the multiple specific layers by the similarity S corresponding to one of the multiple specific layers. The similarity calculation unit 310 then generates the sum of the multiple calculated multiplication values ​​as the similarity S for class discrimination. By setting weighting coefficients for multiple specific layers, it is possible to easily calculate the similarity S used for class discrimination even when there are multiple specific layers. The five individual data DD corresponding to the first model 200_1 to the fifth model 200_5 are referred to as symbols DD_1 to DD_5 when they are used separately. Note that the present invention is not limited to the above. For example, the similarity calculation unit 310 may generate the maximum or minimum value of the similarities S corresponding to each of the multiple specific layers as the similarity S used for class discrimination.

[0055] 12, in step S36, the overall judgment unit 320 integrates the five individual data DD_1 to DD_5 and performs class discrimination on the discriminated data IM based on the integration result. After step S36, the overall judgment unit 320 outputs the class discrimination result to the display unit 150.

[0056] 13, in step S40, the comprehensive judgment unit 320 executes an integration process to integrate the five individual data DD_1 to DD_5. Specifically, the comprehensive judgment unit 320 executes a first integration process to integrate the activation values ​​a of the five individual data DD_1 to DD_5, and a second integration process to integrate the similarities S of the five individual data DD_1 to DD_5.

[0057] In the first integration process, the overall determination unit 320 calculates a cumulative activation value by adding up the five activation values ​​a for each of the three classes. Then, as shown in FIG. 14 , the overall determination unit 320 calculates a discrimination activation value by applying an activation function to the cumulative activation values ​​for each of the three classes. In this embodiment, the overall determination unit 320 calculates a discrimination activation value for each class by applying a softmax function to the three cumulative activation values. Note that the calculation methods for the cumulative activation values ​​and the discrimination activation values ​​are not limited to those described above. For example, the discrimination activation value may be the cumulative activation value of each class as is. Furthermore, for example, the cumulative activation value may be the sum of values ​​obtained by setting weighting coefficients for the first model 200_1 to the fifth model 200_5 and multiplying each activation value a by the corresponding weighting coefficient.

[0058] In the second integration process, the comprehensive judgment unit 320 integrates the similarities S calculated for the first model 200_1 to the fifth model 200_5 to generate a discrimination similarity. In this embodiment, as shown in FIG. 15 , the comprehensive judgment unit 320 generates a discrimination similarity by multiplying the similarities S. Note that the method for generating the discrimination similarity is not limited to this. For example, the discrimination similarity may be a value obtained by dividing the sum of the similarities S by the number of machine learning models 200, or the maximum or minimum value of the similarities S.

[0059] As shown in FIG. 13, in step S41, the comprehensive judgment unit 320 determines the class with the highest discrimination activation value. In the example shown in FIG. 14, the comprehensive judgment unit 320 determines class 1 as the class with the highest discrimination activation value. Next, as shown in FIG. 13, in step S42, the comprehensive judgment unit 320 determines whether the discrimination similarity is equal to or greater than a predetermined threshold. If the discrimination similarity is equal to or greater than the predetermined threshold, the comprehensive judgment unit 320 determines the class with the highest discrimination activation value as the discrimination class in step S44. On the other hand, if the discrimination similarity is less than the predetermined threshold, the comprehensive judgment unit 320 determines the discrimination class as an unknown class regardless of the discrimination activation value. The unknown class is a class different from the class corresponding to the prior label, and in this embodiment, it is different from classes 1 to 3. As described above, the comprehensive judgment unit 320 can accurately determine the discrimination class using the discrimination similarity and the discrimination activation.

[0060] According to the first embodiment, as shown in FIGS. 5 to 7, five input training data groups IDG_1 to IDG_5, which are sets of first-type divided input data IDa_1 to IDa_5, are used to train five machine learning models 200_1 to 200_5. This reduces the amount of data in one input training data group IDG, thereby preventing the training time from becoming too long. Furthermore, since the data length of one input training data group IDG can be shortened, when calculating the output of each layer in the machine learning model 200, detailed features of the data can be utilized without being rounded up as a single feature. This allows for class discrimination that captures the features of the data in more detail. Furthermore, according to the first embodiment, as shown in FIGS. 12 to 15, individual data DD obtained from multiple machine learning models 200_1 to 200_5 can be integrated to facilitate class discrimination.

[0061] The comprehensive determination unit 320 may omit steps S42 and S46. That is, the comprehensive determination unit 320 may determine the class with the highest discrimination activation value as the discrimination class, regardless of the magnitude of the discrimination similarity. This allows the discrimination class to be easily determined using the discrimination activation without using the discrimination similarity.

[0062] B. Other embodiments of the classification process: The class determination process shown in Figures 12 and 13 is not limited to the above embodiment. Other embodiments of the class determination process will be described below.

[0063] B-1. Another embodiment of the classification process: Fig. 16 is a diagram for explaining another embodiment 1 of the class discrimination process. In this embodiment 1, steps S30 and S32 in Fig. 12 are the same as those performed, but the steps from step S34 onwards are different.

[0064] In step S34a, the similarity calculation unit 310 generates a pre-determined class as an element of the individual data DD, using the activation a corresponding to each class and the similarity S, as well as the activation and the similarity S. The similarity calculation unit 310 generates the pre-determined class by executing the following steps (1) and (2). ·Process (1): A process in which, for each of the first model 200_1 to the fifth model 200_5, when the similarity S is greater than or equal to a predetermined threshold, the class corresponding to the largest activation value a among the activation values ​​a corresponding to each class is determined as the pre-discrimination class. ·Process (2): A step of determining an unknown class different from the class corresponding to the pre-label as a pre-discrimination class for each of the first model 200_1 to the fifth model 200_5 when the similarity S is less than a predetermined threshold.

[0065] The threshold values ​​in the above steps (1) and (2) are set to values ​​that estimate that the input discrimination data IM_1 to IM5 are not similar to the divided input data IMa_1 to IMa_5 for each class. In this embodiment, the threshold value is set to 0.7.

[0066] As described above, step (1) allows the class corresponding to the largest activation value a to be determined as the pre-identified class. Furthermore, as described above, step (2) allows an unknown class to be set as the pre-identified class, enabling more accurate class determination of the data to be identified IM. In the present disclosure, the unknown class is represented as class 0.

[0067] In step S34a, the similarity calculation unit 310 executes the above steps (1) and (2) to generate class 1 for the first model 200_1, class 0 indicating an unknown class for the second model 200_2, class 1 for the third model 200_3, class 1 for the fourth model 200_4, and class 1 for the fifth model 200_5 as pre-identified classes.

[0068] Next, in step S48, the comprehensive determination unit 320 determines the most common class from among the pre-discrimination classes for each of the first model 200_1 to the fifth model 200_5 as the class of the discrimination target data IM. This makes it possible to easily perform class discrimination of the discrimination target data IM using the pre-discrimination classes.

[0069] In the above-described class discrimination process according to another embodiment 1, the similarity calculation unit 310 determines the pre-discrimination class taking into consideration a threshold value, but this is not limiting. For example, in step S34a, the similarity calculation unit 310 may generate, as the pre-discrimination class, a class corresponding to the largest activation value a among the activation values ​​a corresponding to each class for each of the first model 200_1 to the fifth model 200_5, regardless of the magnitude of the similarity S.

[0070] B-2. Alternative embodiment 2 of the classification process: 17 is a diagram for explaining another embodiment 2 of the class discrimination process. Step S48b shown in this embodiment 2 is executed in place of step S48 in FIG.

[0071] In step S48b, the comprehensive judgment unit 320 determines the class having the highest similarity S from among the pre-discrimination classes of the first model 200_1 to the fifth model 200_5 as the class of the discrimination target data IM. In the example shown in Fig. 17, class 1, which is the pre-discrimination class of the first model 200_1 having the highest similarity of 0.99, is determined as the discrimination class. This makes it possible to easily determine the class having the highest similarity S as the discrimination class of the discrimination target data IM without using the activation value a.

[0072] Note that the class discrimination process of another embodiment 2 is not limited to the above. For example, in step S48b, the comprehensive judgment unit 320 may calculate the sum or product of the similarities S of the individual data DD for each class having the same pre-discrimination class, and determine the pre-discrimination class with the largest calculated value as the discrimination class of the data to be discriminated IM. For example, the following example will be used to explain this. First model: (Pre-classification class = Class 1, similarity S = 0.8) Second model: (Pre-classification class = Class 1, similarity S = 0.7) Third model (pre-classification class = Class 3, similarity S = 0.7) Fourth model (pre-classification class = Class 2, similarity S = 0.9) Fifth model (pre-classification class = Class 2, similarity S = 0.8)

[0073] In the above case, the comprehensive judgment unit 320 calculates the sum of the similarities S of the individual data DD for each class having the same pre-discrimination class. In the sum, the sum of the similarities S in class 1 is 1.5, the sum of the similarities S in class 2 is 1.7, and the sum of the similarities S in class 3 is 0.7. Therefore, the comprehensive judgment unit 320 determines class 2, which has the largest sum of 1.7, as the discrimination class of the discrimination target data IM. This makes it possible to easily determine the discrimination class of the discrimination target data IM using the similarities S without using the activation value a.

[0074] B-3. ​​Other embodiment 3 of the classification process: FIG. 18 is a diagram for explaining another embodiment 3 of the class discrimination process. Step S48c shown in this embodiment 3 is executed instead of step S48b in FIG. 17. In this embodiment 3, weighting coefficients α1 to α5 are set for the first model 200_1 to the fifth model 200_5. In step S48c, the comprehensive judgment unit 320 multiplies the similarity S by the corresponding weighting coefficients α1 to α5 for each of the first model 200_1 to the fifth model 200_5 to calculate a reference value. Then, the comprehensive judgment unit 320 judges the class of the discrimination target data IM using the pre-discrimination class and the reference value. In detail, the comprehensive judgment unit 320 calculates the sum of the reference values ​​of the machine learning models 200 having the same pre-discrimination class. Then, the comprehensive judgment unit 320 determines the class with the largest sum as the discrimination class of the discrimination target data IM. Alternatively, the comprehensive determination unit 320 may determine the pre-discrimination class of the machine learning model 200 having the maximum or minimum reference value as the discrimination class of the discrimination target data IM. According to the above, the class of the discrimination target data IM can be determined in consideration of the weighting coefficients set for the machine learning models 200_1 to 200_5.

[0075] B-4. Other embodiment 4 of the determination process: In other embodiments 1 to 3 of the discrimination process, when one of the plurality of pre-discrimination classes corresponding to each of the plurality of machine learning models 200_1 to 200_5 indicates an unknown class, the comprehensive judgment unit 320 may determine the unknown class as the discrimination class of the discrimination target data IM regardless of the classes indicated by the other pre-discrimination classes. When one of the plurality of pre-discrimination classes indicates an unknown class, there is a possibility that the discrimination target data IM is unknown. Therefore, when one of the pre-discrimination classes indicates an unknown class, by determining the class of the discrimination target data IM as the unknown class, more accurate class discrimination can be performed.

[0076] C. Other embodiments of the pre-determined class generation process: According to another embodiment of the discrimination process described above, the similarity calculation unit 310 generates a pre-discrimination class using the similarity S and the activation value a, as shown in Figures 16 to 18. However, the present invention is not limited to this. Figure 19 is a diagram for explaining another embodiment of the pre-discrimination class generation process.

[0077] In another embodiment shown in FIG. 19, in step S34c, the similarity calculation unit 310 calculates a class-specific similarity Sclass, which is the similarity between the class-specific known feature spectrum KSp and the feature spectrum KSp of the input discrimination data IM_1 to IM_5, for each class of each of the machine learning models 200_1 to 200_5. The similarity calculation unit 310 performs statistical processing on the multiple similarities S calculated for each class to calculate a representative similarity, which is a representative value of the multiple similarities S for each class, as the class-specific similarity Sclass. The "representative value obtained by statistical processing" refers to the maximum value, median value, average value, or mode. This representative similarity is used to generate pre-discrimination classes, which will be described later. The method for calculating the class-specific similarity Sclass will be described in detail later.

[0078] Next, the similarity calculation unit 310 generates a class associated with the largest representative similarity among the representative similarities calculated for each class as the class-specific similarity Sclass as a pre-determined class, as one element of the individual data DD. This allows pre-determined classes to be easily generated using the class-specific similarity Sclass without using the activation value a. If the largest representative similarity is less than a predetermined threshold, an unknown class different from the class corresponding to the pre-determined label may be generated as a pre-determined class, as one element of the individual data, instead of the class associated with the representative similarity. This allows unknown classes to be generated as pre-determined classes, resulting in more accurate pre-determined classes. In this embodiment, the predetermined threshold is set to 0.7, as in steps (1) and (2) above.

[0079] D. Alternatives to the First Embodiment: FIG. 20 is a diagram illustrating another embodiment of the first embodiment. In the first embodiment, as shown in FIG. 14, one machine learning model 200_1 to 200_5 corresponds to each of the regions R1 to R5. However, multiple machine learning models 200 may be provided corresponding to each of the regions R1 to R5. In the example shown in FIG. 20, two machine learning models 200 are trained corresponding to each of the regions R1 to R5 and used for class discrimination. The two machine learning models 200_1 to 200_5 corresponding to each of the regions R1 to R5 are distinguished by adding "a" and "b" to the end of their names. In the training process, the data generation unit 112 divides multiple input training data groups IDG included in the input training data groups IDG_1 to IDG_5 corresponding to one of the regions R1 to R5 into two. For example, the data generation unit 112 divides the multiple input training data groups IDG into two groups so that the number of input training data groups IDG is equal. When training the machine learning models 200_1a and 200_1b, one of the two divided input training data groups IDG_1 is input to the machine learning model 200_1a, and the other is input to the machine learning model 200_1b, thereby performing training of the two machine learning models 200_1a and 200_1b.

[0080] In the class discrimination process, first, the similarity calculation unit 310 obtains individual data DD by inputting the input discrimination data IM_1 to IM_5 to the corresponding two machine learning models 200. Next, the similarity calculation unit 310 determines which individual data DD of the two machine learning models 200 corresponding to each of the regions R1 to R5 is to be used for class discrimination. Specifically, the similarity calculation unit 310 calculates a model reliability Rmodel that depends on the similarity S of the individual data DD, and determines the individual data DD obtained from the machine learning model 200 with the highest model reliability Rmodel as the data to be used for class discrimination. For the five individual data DD determined for each of the regions R1 to R5 for class discrimination, in the same manner as in the first embodiment, integrated processing is executed by the first integrated processing and the second integrated processing, thereby calculating a discrimination activation value and a discrimination similarity. Then, in the same manner as in the first embodiment, the comprehensive determination unit 320 determines a discrimination class from the discrimination activation and the discrimination similarity. Note that the determination method of determining the individual data DD obtained from the machine learning model 200 with the highest model reliability Rmodel as the data to be used for class discrimination can also be applied in other embodiments of the present disclosure. For example, in the first embodiment shown in FIG. 2, which individual data DD of the five machine learning models 200_1 to 200_5 is to be used for class discrimination is determined by the above determination method.

[0081] As the reliability function for obtaining the model reliability R from the similarity S, for example, any of the following can be used. Rmodel(i)=H1[S(i)]=S(i) (3a) Rmodel(i)=H2[S(i)]=Ac(i)×Wt+S(i)×(1-Wt) (3b) Rmodel(i)=H3[S(i)]=Ac(i)×S(i) (3c) Here, Ac(i) is an activation value corresponding to the determination value with the largest value in the output layer of the machine learning model 200, and Wt is a weight coefficient in the range of 0 < Wt < 1.

[0082] The reliability function H1 in the above equation (3a) is an identity function that uses the similarity S itself as the model reliability Rmodel. The reliability function H2 in the above equation (3b) is a function that calculates the model reliability Rmodel by taking a weighted average of the similarity S and the activation value Ac. The reliability function H3 in the above equation (3c) is a function that calculates the model reliability Rmodel by multiplying the similarity S by the activation value Ac. Reliability functions other than these may also be used. For example, a function that uses the power of the similarity S as the model reliability Rmodel may also be used. In this way, the model reliability Rmodel can be calculated as something that depends on the similarity S. Furthermore, it is preferable that the model reliability Rmodel has a positive correlation with the similarity S.

[0083] In the first embodiment, the data generation unit 112 performs the first data processing by dividing the input data into M or more regions and generating a set of first-type divided input data IDa belonging to the same region as a single input training data group IDG. However, the first data processing may also perform the first data processing by dividing the input data into one or more regions or two or more regions and generating a set of first-type divided input data IDa belonging to the same region as a single input training data group IDG. In this case, the same input training data group IDG may be input to at least two of the M machine learning models 200. Since the performance of the machine learning model 200 may change with each learning even with the same input training data, multiple machine learning models 200 trained with the same input training data may be used. Conventionally, when classifying data IM to be discriminated using a single machine learning model, the accuracy of classification may be reduced. However, according to the first embodiment and other embodiments of the first embodiment, the accuracy of classification can be improved by performing classification using multiple machine learning models 200 with different classification performance. Such an effect is also achieved in the second embodiment described later.

[0084] E. Second embodiment: FIG. 21 is a flowchart showing the learning process of the second embodiment. FIG. 22 is a conceptual diagram of clustering. FIG. 23 is a diagram for explaining step S12a. The second embodiment also uses a discrimination system 5 similar to that of the first embodiment. In the second embodiment, the number of machine learning models 200 shown in FIG. 2 is two. When distinguishing between the two machine learning models 200, the reference symbols 200_1 and 200_2 are used. The learning process of the second embodiment differs from the learning process of the first embodiment shown in FIG. 5 in step S12a. Therefore, in the learning process, the same steps in the first and second embodiments are denoted by the same reference symbols, and their description will be omitted.

[0085] As shown in FIG. 21 , in step S12a, the data generation unit 112 executes the second data processing. Specifically, the data generation unit 112 divides a plurality of pieces of training data TD belonging to one class into two by clustering, and generates an input training data group IDAG, which is a set of input training data IDA as the second type of divided input data after the division. In this embodiment, two input training data groups IDAG are generated. Note that the same input training data IDA may be classified across the two clusters after the division. One of the two input training data groups IDAG is also referred to as a first input training data group IDAG1, and the other is also referred to as a second input training data group IDAG2. The first input training data group IDAG1 is used for training the machine learning model 200_1. The second input training data group IDAG2 is used for training the machine learning model 200_2. For example, k-means is used for the clustering. When considering the case where the k-means method is used, the input training data groups IDAG1 and IDAG2 have representative points G1 and G2 that represent the input training data groups IDAG1 and IDAG2, respectively. These representative points G1 and G2 are, for example, centers of gravity. Note that instead of the clustering in step S12a above, the second data processing may be performed by randomly extracting multiple pieces of training data TD belonging to one class with sampling with replacement. By randomly extracting multiple pieces of training data TD with sampling with replacement, two groups that are sets of training data TD are generated.

[0086] 23, two groups A and B are generated for each of classes 1 to 3 by clustering or sampling with replacement. There is no particular limitation on whether these two groups A and B are assigned to the first input training data group IDAG1 or the second input training data group IDAG2. For example, for each of classes 1 to 3, the data generation unit 112 may randomly assign one group to the first input training data group IDAG1 and the other group to the second input training data group IDAG2. Furthermore, for example, the data generation unit 112 may assign, for classes 1 to 3, groups whose representative points G1 and G2 have similar Euclidean distances to the same input training data group IDAG. For example, if the Euclidean distance between the representative point G1 of group A of class 1 and the representative point G1 of group A of class 2 is closer than the Euclidean distance between the representative point G1 of group A of class 1 and the representative point G2 of group B of class 2, the data generation unit 112 assigns group A of class 1 and group A of class 2 to the same input training data group IDAG. The assignment method for class 3 is the same as for class 2. The index used for assigning to the same input training data group IDAG is not limited to the above-mentioned Euclidean distance, but may also be cosine similarity or Mahalanobis distance.

[0087] In the second embodiment, the preparation step shown in Fig. 8 and the discrimination step shown in Fig. 12, which were described in the first embodiment, are also executed. In this case, in step S30 of the discrimination step shown in Fig. 12, the discrimination target data IM is generated as input discrimination target data IM without being divided. Then, in steps S32 and S34, the processor 110 inputs one piece of input discrimination target data IM to two machine learning models 200_1 and 200_2, thereby obtaining individual data DD from the trained machine learning models 200_1 and 200_2.

[0088] According to the second embodiment, as shown in FIGS. 21 to 23, one machine learning model 200 is trained using an input training data group IDG, which is a collection of second-type divided input data IDb. This reduces the amount of data in one input training data group IDG, thereby preventing the training time of each machine learning model 200 from becoming long. Furthermore, according to the second embodiment, as in the first embodiment, the individual data DD obtained from multiple machine learning models 200 can be integrated to easily perform class discrimination. Furthermore, by integrating the individual data DD obtained from multiple machine learning models 200 to perform class discrimination, it is expected that each machine learning model 200 can recognize and discriminate different features, thereby enabling highly accurate class discrimination that takes into account more diverse features.

[0089] F. Alternatives to the second embodiment: In the second embodiment, the data generation unit 112 performs, as the second data processing, a division process for dividing multiple pieces of input data IM belonging to one class into M or more pieces, and generates a set of the resulting second-type divided input data IDb as a single input training data group IDG. However, the second data processing may also be a process for dividing multiple pieces of training data TD belonging to one class into one or more pieces or two or more pieces, and generating a set of the resulting second-type divided input data IDb as a single input training data group IDG. For example, multiple pieces of training data TD belonging to one class are designated as training data TD1, TD2, and TD3. In this case, the following seven input training data groups IDG are generated. For each of the machine learning models 200_1 and 200_2, one or more data groups are selected from the following seven generated input training data groups IDG and used for training. (1) First input training data group: This group is composed of training data TD1. (2) Second input training data group: This group is composed of training data TD2. (3) Third input training data group: This group is composed of training data TD3. (4) Fourth input training data group: Composed of training data TD1 and TD2. (5) Fifth input training data group: composed of training data TD1 and TD3. (6) Sixth input training data group: Composed of training data TD2 and TD3. (7) Seventh input training data group: Comprised of training data TD1, TD2, and TD3.

[0090] Note that, since the performance of machine learning models 200_1 and 200_2 may change for each learning even with the same input learning data, there may be multiple machine learning models 200 trained with the same input learning data. In other words, in the second data processing, regardless of the number of machine learning models 200, it is sufficient that multiple learning data TD belonging to one class are divided into one or more pieces.

[0091] G. Similarity calculation method: As a method for calculating the above-mentioned class-specific similarity Sclass, for example, one of the following three methods can be adopted. (1) A first calculation method M1 for calculating the class-specific similarity Sclass without considering the correspondence between the feature spectrum Sp and the subregion Rn in the known feature spectrum group KSp. (2) A second calculation method M2 for calculating the class-specific similarity Sclass between the feature spectrum Sp and the corresponding subregion Rn of the known feature spectrum group KSp. (3) A third calculation method M3 for calculating the class-specific similarity Sclass without considering the subregion Rn at all. Below, we will sequentially explain how to calculate the class-specific similarity Sclass_ConvVN1 from the output of the ConvVN1 layer 230 according to these three calculation methods M1, M2, and M3. Note that in the following explanation, the parameter m of the machine learning model 200 and the parameter q of the discriminated data IM are omitted.

[0092] Fig. 24 is an explanatory diagram showing a first calculation method M1 of class-specific similarity. In the first calculation method M1, first, a local similarity S(i,j,k) indicating the similarity to each class i for each subregion k is calculated from the output of the ConvVN1 layer 230, which is a specific layer. Then, one of the three types of class-specific similarity Sclass(i,j) shown on the right side of Fig. 24 is calculated from these local similarities S(i,j,k).

[0093] In the first calculation method M1, the local similarity S(i,j,k) is calculated using the following formula: S(i,j,k)=max[G{Sp(j,k), KSp(i,j,k=all,q=all)}] (c1) where: i is a parameter indicating the class, j is a parameter indicating a specific layer, k is a parameter indicating the subregion Rn, q is a parameter indicating the data number, G{a,b} is a function that calculates the similarity between a and b. Sp(j,k) is the feature spectrum obtained from the output of a specific subregion k of a specific layer j according to the data to be discriminated. KSp(i, j, k=all, q=all) is the known feature spectrum of all data numbers q in all subregions k of specific layer j associated with class i among the known feature spectrum group KSp shown in FIG. 11; max[X] is a logical operation that takes the maximum value of X. As the function G{a, b} for calculating the similarity, for example, an equation for calculating cosine similarity or an equation for calculating similarity according to distance can be used.

[0094] The three types of class-specific similarities Sclass(i,j) shown on the right side of FIG. 24 are calculated as representative similarities by statistically processing the local similarities S(i,j,k) for multiple partial regions k for each class i. The statistical processing is performed by taking the maximum, average, or minimum value of the multiple local similarities S(i,j,k). Although not shown, the class-specific similarity Sclass may also be obtained by taking the mode of the local similarities S(i,j,k) for multiple partial regions k. Whether the maximum, average, minimum, or mode is calculated depends on the purpose of the class discrimination process. For example, when the purpose is to discriminate objects using natural images, it is preferable to calculate the class-specific similarity Sclass(i,j) by taking the maximum value of the local similarities S(i,j,k) for each class i. Furthermore, when the purpose is to distinguish the type of target object 10 or to judge the quality of industrial products using images, it is preferable to calculate the class-specific similarity Sclass(i,j) by taking the minimum value of the local similarities S(i,j,k) for each class i. There may also be cases where it is preferable to calculate the class-specific similarity Sclass(i,j) by taking the average value of the local similarities S(i,j,k) for each class i. Which of these four types of calculations to use is determined in advance by the user experimentally or empirically.

[0095] As described above, in the first calculation method M1 of class-specific similarity, (1) Calculating a local similarity S(i,j,k) between a feature spectrum Sp obtained from the output of a specific subregion k of a specific layer j and all known feature spectra KSp associated with the specific layer j and each class i according to the discriminated data I M; (2) For each class i, the class-specific similarity Sclass(i,j) is calculated by taking the maximum, average, minimum or mode of the local similarities S(i,j,k) for multiple subregions k. According to this first calculation method M1, the class-specific similarity Sclass(i,j) can be calculated using relatively simple calculations and procedures.

[0096] 25 is an explanatory diagram showing a second calculation method M2 of class similarity. In the second calculation method M2, the local similarity S(i, j, k) is calculated using the following equation instead of the above-mentioned equation (c1). S(i,j,k)=max[G{Sp(j,k), KSp(i,j,k,q=all)}] (c2) where: KSp(i, j, k, q=all) is the known feature spectrum of all data numbers q in a specific partial region k of a specific layer j associated with class i, among the known feature spectrum group KSp shown in FIG.

[0097] While the first calculation method M1 described above uses known feature spectra KSp(i,j,k=all,q=all) in all partial regions k of a specific layer j, the second calculation method M2 uses only known feature spectra KSp(i,j,k,q=all) for the partial region k that is the same as the partial region k of the feature spectrum Sp(j,k). The other aspects of the second calculation method M2 are the same as those of the first calculation method M1.

[0098] In the second calculation method M2 of the class-specific similarity, (1) Calculating a local similarity S(i,j,k) between a feature spectrum Sp obtained from the output of a specific subregion k of a specific layer j and all known feature spectra Ksp associated with the specific subregion k of the specific layer j and each class i according to the discriminated data I M; (2) For each class i, the class-specific similarity Sclass(i,j) is calculated by taking the maximum, average, minimum, or mode of the local similarities S(i,j,k) for multiple subregions k. This second calculation method M2 also makes it possible to find the class-specific similarity Sclass(i,j) through relatively simple calculations and procedures.

[0099] 26 is an explanatory diagram showing a third calculation method M3 of class-specific similarity. In the third calculation method M3, the class-specific similarity Sclass(i,j) is calculated from the output of the ConvVN1 layer 230, which is a specific layer, without calculating the local similarity S(i,j,k).

[0100] The class-specific similarity Sclass(i,j) obtained by the third calculation method M3 is calculated using the following formula. Sclass(i,j)=max[G{Sp(j,k=all), KSp(i,j,k=all,q=all)}] (c3) where: Sp(j, k=all) is a feature spectrum obtained from the outputs of all partial regions k of a specific layer j according to the discriminated data IM.

[0101] As described above, in the third calculation method M3 of class-specific similarity, (1) The class-specific similarity Sclass(i,j) is calculated for each class, which is the similarity between all feature spectra Sp obtained from the output of specific layer j according to the discriminated data IM and all known feature spectra KSp associated with that specific layer j and each class i. According to the third calculation method M3, the class-specific similarity Sclass(i,j) can be calculated using even simpler calculations and procedures.

[0102] H. How to calculate the output vector of each layer of the machine learning model: The method for calculating the output of each layer in the machine learning model 200 shown in Figure 3 is as follows: The machine learning model 200 shown in Figure 4 is the same except for the values ​​of the individual parameters.

[0103] Each node in the PrimeVN layer 220 regards the scalar output of the 1x1x32 nodes in the Conv layer 210 as a 32-dimensional vector and obtains the vector output of that node by multiplying this vector by a transformation matrix. This transformation matrix is ​​an element of a kernel with a surface size of 1x1, and is updated by learning the machine learning model 200. Note that the processing of the Conv layer 210 and the PrimeVN layer 220 can also be integrated into one primary vector neuron layer.

[0104] When the PrimeVN layer 220 is referred to as the "lower layer L" and the ConvVN1 layer 230 adjacent to it on the upper side is referred to as the "upper layer L+1", the output of each node in the upper layer L+1 is determined using the following equation.

number

[0105] As the normalization function F(X), for example, the following formula (E3a) or (E3b) can be used.

number

[0106] In the above equation (E3a), the sum vector u j Norm of |u j The activation value a is normalized by the softmax function | j On the other hand, in equation (E3b), the sum vector u j Norm of |u j | is the norm |u j Activation value a by dividing by the sum of | j It should be noted that a function other than equation (E3a) or (E3b) may be used as the normalization function F(X).

[0107] The ordinal number i in the above equation (E2) is the output vector M of the jth node in the upper layer L+1. L+1 j The integer n is assigned for convenience to the nodes in the lower layer L used to determine the output vector M L+1 j is the number of nodes in the lower layer L used to determine . Thus, the integer n is given by n = Nk × Nc (E5) Here, Nk is the surface size of the kernel, and Nc is the number of channels in the lower layer, the PrimeVN layer 220. In the example of Fig. 3, Nk = 5 and Nc = 26, so n = 130.

[0108] One kernel used to calculate the output vector of the ConvVN1 layer 230 has 1 × 5 × 26 = 130 elements, with the kernel size 1 × 5 as the surface size and the number of channels in the lower layer being 26 as the depth. Each of these elements is a prediction matrix W L ij In addition, 20 sets of this kernel are required to generate output vectors for 20 channels of the ConvVN1 layer 230. Therefore, the prediction matrix W of the kernel used to obtain the output vector of the ConvVN1 layer 230 is L ij The number of prediction matrices W is 130 × 20 = 2600. L ij is updated by learning of the machine learning model 200.

[0109] As can be seen from the above equations (E1) to (E4), the output vector M of each node in the upper layer L+1 L+1 j is calculated by the following calculation: (a) Output vector M of each node in the lower layer L L i The prediction matrix W L ij Multiplying by the predicted vector v ij Seeking (b) Prediction vector v obtained from each node in the lower layer L ij The sum vector u is a linear combination of j Seeking (c) Sum vector u j Norm of |u j The activation value a is normalized by normalizing | j Seeking (d) Sum vector u j norm |u j Divide by | and then use the activation value a j Multiply by.

[0110] In addition, the activation value a j is the norm |u jis the normalization factor obtained by normalizing |. Therefore, the activation value a j can be considered as an index showing the relative output strength of each node among all nodes in the upper layer L+1. The norm used in equations (E3), (E3a), (E3b), and (4) is typically the L2 norm, which represents the vector length. In this case, the activation value a j is the output vector M L+1 j The activation value a corresponds to the vector length of j is only used in the above equations (E3) and (E4), and does not need to be output from the node. However, the activation value a j It is also possible to configure the upper layer L+1 so that it outputs

[0111] The configuration of a vector neural network is almost the same as that of a capsule network, and the vector neurons of a vector neural network correspond to the capsules of a capsule network. However, the calculations according to the above formulas (E1) to (E4) used in a vector neural network are different from the calculations used in a capsule network. The biggest difference between the two is that in a capsule network, the predicted vector v on the right side of the above formula (E2) ij are multiplied by weights, and the weights are searched by repeating dynamic routing multiple times. On the other hand, in the vector neural network of this embodiment, the output vector M is calculated by calculating the above-mentioned equations (E1) to (E4) once in order. L+1 j Therefore, there is no need to repeat dynamic routing, which has the advantage of allowing faster calculations. In addition, the vector neural network of this embodiment has the advantage that it requires less memory for calculations than a capsule network, and according to experiments by the inventors of this disclosure, it only requires about 1 / 2 to 1 / 3 of the memory required.

[0112] Vector neural networks are similar to capsule networks in that they use nodes that use vectors as input and output. Therefore, they share the advantages of using vector neurons with capsule networks. Furthermore, the multiple layers 210-250 are similar to conventional convolutional neural networks in that the higher layers represent features of larger areas and the lower layers represent features of smaller areas. Here, "feature" refers to a characteristic part contained in the input data to the neural network. Vector neural networks and capsule networks are superior to conventional convolutional neural networks in that the output vector of a node contains spatial information representing the spatial information of the feature represented by that node. That is, the vector length of a node's output vector represents the probability of the feature represented by that node, and the vector direction represents spatial information such as the direction and scale of the feature. Therefore, the vector direction of the output vectors of two nodes belonging to the same layer represents the relative positions of the respective features. Alternatively, the vector direction of the output vectors of the two nodes can be said to represent the variation of the feature. For example, for a node corresponding to the "eye" feature, the direction of the output vector can represent variations such as the narrowness of the eyes or the way they are lifted. In conventional convolutional neural networks, it is said that spatial information of features is lost due to the pooling process. As a result, vector neural networks and capsule networks have the advantage of being superior to conventional convolutional neural networks in terms of the performance of identifying input data.

[0113] The advantages of vector neural networks can also be considered as follows. In other words, the advantage of vector neural networks is that the output vectors of nodes represent the features of input data as coordinates in continuous space. Therefore, output vectors can be evaluated such that the closer the vector directions, the more similar the features. Another advantage is that even if the features contained in the input data are not covered by the training data, they can be determined by interpolation. On the other hand, conventional convolutional neural networks have the disadvantage that the features of input data cannot be represented as coordinates in continuous space due to the chaotic compression caused by the pooling process.

[0114] The outputs of each node in the ConvVN2 layer 240 and the ClassVN layer 250 are similarly determined using the above-mentioned equations (E1) to (E4), and therefore detailed explanations are omitted. The ClassVN layer 250, which is the top layer, has a resolution of 1x1 and a number of channels of n1.

[0115] The output of the ClassVN layer 250 is converted into a plurality of decision values ​​Class0 to Class2 for known classes. These decision values ​​are usually normalized by a softmax function. Specifically, for example, the decision value for each class can be obtained by performing the following operation: calculating the vector length of the output vector from the output vector of each node of the ClassVN layer 250, and then normalizing the vector length of each node by a softmax function. As described above, the activation value a obtained by the above formula (E3) is j is the output vector M L+1 j The activation value a at each node in the ClassVN layer 250 is a value corresponding to the vector length of j may be output and used as the judgment value for each class.

[0116] In the above-described embodiment, a vector neural network that determines an output vector by calculating the above equations (E1) to (E4) was used as the machine learning model 200, but instead, a capsule network disclosed in U.S. Pat. No. 5,210,798 or WO 2009 / 083553 may be used.

[0117] I. Other forms: The present disclosure is not limited to the above-described embodiments and can be realized in various forms without departing from the spirit thereof. For example, the present disclosure can also be realized in the following aspects. The technical features in the above embodiments corresponding to the technical features in each aspect described below can be appropriately replaced or combined to solve some or all of the problems of the present disclosure or to achieve some or all of the effects of the present disclosure. Furthermore, if a technical feature is not described as essential in this specification, it can be appropriately deleted.

[0118] (1) According to a first aspect of the present disclosure, there is provided a learning method for M machine learning models (M is an integer equal to or greater than 2) used to distinguish the class of data to be distinguished, the M machine learning models being of a vector neural network type having multiple vector neuron layers. This learning method includes the steps of: (a) preparing a plurality of training data sets each having input data and a prior label associated with the input data; (b) dividing the plurality of training data sets into one or more parts to generate the one or more input training data groups; and (c) training the M machine learning models by inputting the corresponding input training data groups into each of the M machine learning models so as to reproduce the correspondence between the input data and the prior label associated with the input data. The step (b) includes one of the steps of: (b1) dividing each of the plurality of input data sets into one or more regions and generating a set of first-type divided input data after division that belong to the same region as the input training data group; and (b2) dividing the plurality of training data sets belonging to one class into one or more parts and generating a set of second-type divided input data after division as the input training data group. According to this embodiment, a set of first-type divided input data can be used as one group of input training data for training one machine learning model, or a set of second-type divided input data can be used as one group of input training data for training one machine learning model. This reduces the amount of data used for training each machine learning model, thereby preventing the training time from becoming too long.

[0119] (2) According to a second aspect of the present disclosure, there is provided a method for classifying data to be classified using M (M is an integer equal to or greater than 2) machine learning models of a vector neural network type having multiple vector neuron layers.This discrimination method includes: (a) a step of preparing the M machine learning models trained using a plurality of training data each having input data and a prior label associated with the input data, wherein each of the M machine learning models is trained by dividing the plurality of training data into one or more input training data groups and using the input training data group corresponding to one of the one or more divided input training data groups; (b) a step of preparing M sets of known feature spectra associated with the M trained machine learning models, wherein the M sets of known feature spectra include known feature spectra obtained from the output of a specific layer among the plurality of vector neuron layers by inputting the input training data groups to the M trained machine learning models; and (c) a step of inputting discriminated data generated from the discriminated data into each of the M trained machine learning models, and obtaining, for each of the M machine learning models, individual data to be used for class discrimination of the discriminated data. and (d) performing class discrimination of the discriminated data using the M individual data obtained for each of the M machine learning models, wherein the individual data is generated using at least one of: (i) a similarity between the feature spectrum calculated from the output of the specific layer in response to the input of the discriminated data to the machine learning model and the group of known feature spectra; and (ii) an activation value corresponding to a judgment value of each class output from the output layer of the machine learning model in response to the input of the discriminated data; and (d) performing class discrimination of the discriminated data using the M individual data obtained for each of the M machine learning models, wherein the step (a) includes either: (a1) dividing each of the plurality of input data into one or more regions, and using a set of first-type divided input data after division that belongs to the same region as one of the input training data groups; or (a2) performing a division process to divide the plurality of training data belonging to one class into one or more parts, and using a set of second-type divided input data after the division process as one of the input training data groups. According to this embodiment, the set of first-type divided input data can be used as one group of input learning data for training one machine learning model, or the set of second-type divided input data can be used as one group of input learning data for training one machine learning model. This makes it possible to reduce the amount of data used for training each machine learning model, and therefore to discriminate the class of data to be discriminated using a machine learning model that prevents the learning time from becoming long.

[0120] (3) In the above embodiment, each of the M individual data may include an activation value corresponding to each class, and step (d) may determine, for each class, the class with the highest discrimination activation value calculated using a cumulative activation value obtained by adding up the activation values ​​of the M individual data. According to this embodiment, the discrimination class can be easily determined using the discrimination activation value.

[0121] (4) In the above embodiment, each of the M individual data includes the similarity, and step (d) may include: (d1) generating a discrimination similarity by integrating the similarities for each of the machine learning models; and (d2) determining, if the discrimination similarity is equal to or greater than a predetermined threshold, the class with the highest discrimination activation value as the discrimination class, and, if the discrimination similarity is less than the threshold, determining, regardless of the discrimination activation value, an unknown class different from the class corresponding to the prior label as the discrimination class. According to this embodiment, the discrimination similarity and the discrimination activation can be used to accurately determine the discrimination class.

[0122] (5) In the above embodiment, step (c) may include a step of generating, for each machine learning model, a class corresponding to the largest activation value among the activation values ​​corresponding to each class as a pre-discrimination class as one element of the individual data. According to this embodiment, by designating the class corresponding to the largest activation value as the pre-discrimination class, it is possible to generate classes that are candidates for the discrimination class.

[0123] (6) In the above embodiment, step (c) may include: generating, for each of the machine learning models, a class corresponding to the largest activation value among the activation values ​​corresponding to the classes as a pre-determined class, as one element of the individual data, when the similarity is equal to or greater than a predetermined threshold; and generating, for each of the machine learning models, the pre-determined class as an unknown class different from the class corresponding to the pre-label, as one element of the individual data, when the similarity is less than the threshold. According to this embodiment, an unknown class can be set as the pre-determined class, thereby enabling more accurate class determination of the data to be classified.

[0124] (7) In the above embodiment, the step (c) may include either (i) calculating a multiplication value for each of the plurality of specific layers by multiplying the weight coefficient set for each of the plurality of specific layers by the similarity corresponding to one of the plurality of specific layers, and setting the sum of the calculated multiplication values ​​as the similarity used for class discrimination, or (ii) setting the maximum or minimum value of the similarities corresponding to each of the plurality of specific layers as the similarity used for class discrimination. According to this embodiment, even when there are a plurality of specific layers, the similarity used for class discrimination can be easily calculated.

[0125] (8) In the above embodiment, when each known feature spectrum included in the group of known feature spectra prepared in step (b) is associated with class classification information indicating to which class the spectrum belongs, and the known feature spectrum associated with the class classification information is referred to as a class-specific known feature spectrum, step (c) may include the steps of: calculating, for each class, a class-specific similarity that is the similarity between the class-specific known feature spectrum and the feature spectrum; and generating, as a pre-identified class, one element of the individual data, the class associated with the class-specific similarity that has the largest value among the class-specific similarities calculated for the classes. According to this embodiment, the pre-identified class can be easily generated using the class-specific similarity.

[0126] (9) In the above-described embodiment, the step of calculating the class-specific similarities in step (c) may include: calculating, for each class, the similarities between the feature spectrum and each of the plurality of known class-specific characteristic spectra; and calculating a representative similarity of the plurality of similarities for each class by statistically processing the plurality of similarities calculated for each class. The generating step in step (c) may include generating, as one element of the individual data, the class associated with the representative similarity having the largest value among the representative similarities calculated for each class. According to this embodiment, the pre-identified class can be easily generated using the representative similarity.

[0127] (10) In the above-described embodiment, the statistical processing of the plurality of similarities may be performed by calculating the maximum value, median value, average value, or mode of the plurality of similarities as the representative similarity. According to this embodiment, the representative similarity can be used to easily generate pre-identified classes.

[0128] (11) In the above embodiment, step (c) may further include a step of generating, as the pre-determined class, an unknown class different from the class corresponding to the pre-label, as one element of the individual data, instead of the class associated with the class-specific similarity, when the largest value is less than a predetermined threshold. According to this embodiment, an unknown class can also be generated as the pre-determined class, thereby enabling the generation of pre-determined classes with higher accuracy.

[0129] (12) In the above embodiment, the step (d) may include a step of determining, as the class of the data to be classified, the most common class from among the pre-classified classes of the individual data for each of the plurality of machine learning models. According to this embodiment, the class of the data to be classified can be easily classified using the pre-classified classes.

[0130] (13) In the above-described embodiment, step (c) may include a step of generating a similarity between the feature spectrum calculated from the output of the specific layer and the group of known feature spectra as one element of the individual data, and step (d) may include either a step of (i) determining, from among the pre-discrimination classes of the individual data for each of the plurality of machine learning models, a class having the highest similarity as the class of the data to be discriminated, or a step of (ii) calculating, for each class having the same pre-discrimination class, the sum or product of the similarities of the individual data, and determining, as the class of the data to be discriminated, the pre-discrimination class having the largest calculated value. According to this embodiment, the discrimination class of the data to be discriminated can be easily determined using the similarity.

[0131] (14) In the above embodiment, step (c) may include a step of generating, as one element of the individual data, a similarity between the feature spectrum calculated from the output of the specific layer and the group of known feature spectra, and step (d) may include a reference value calculation step of calculating, for each of the plurality of machine learning models, a reference value using the similarity and a weighting factor preset for each of the plurality of machine learning models, and a class determination step of determining a class of the data to be discriminated using the pre-discrimination class and the calculated reference value. According to this embodiment, the class of the data to be discriminated can be determined taking into account the weighting factor set for each machine learning model.

[0132] (15) In the above aspect, the reference value calculation step may calculate the reference value for each of the plurality of machine learning models by multiplying the similarity by the weighting coefficient, and the class determination step may include either (i) determining the pre-discrimination class having the largest sum of the reference values ​​of the machine learning models having the same pre-discrimination class as the class of the data to be discriminated, or (ii) determining the pre-discrimination class of the machine learning model having the largest or smallest reference value as the class of the data to be discriminated. According to this aspect, the class of the data to be discriminated can be determined taking into account the weighting coefficient set for each machine learning model.

[0133] (16) In the above embodiment, in the step (d), when one of the plurality of pre-determined classes corresponding to the plurality of machine learning models indicates an unknown class different from the class corresponding to the pre-determined label, the unknown class may be determined as the class of the data to be discriminated, regardless of the classes indicated by the other pre-determined classes. According to this embodiment, when one of the pre-determined classes indicates an unknown class, the class of the data to be discriminated can be determined as the unknown class.

[0134] (17) In the above-described embodiment, the splitting process in step (a2) may be performed by (i) clustering the plurality of training data belonging to the one class, or (ii) randomly extracting the plurality of training data belonging to the one class by sampling with replacement. According to this embodiment, by clustering the plurality of training data or randomly extracting the plurality of training data by sampling with replacement, a set of second-type split input data can be easily generated.

[0135] (18) According to a third aspect of the present disclosure, there is provided a learning device for M machine learning models (M is an integer equal to or greater than 2) used to distinguish classes of data to be distinguished, the M machine learning models being of a vector neural network type having multiple vector neuron layers. This learning device comprises a memory and a processor that performs learning of the M machine learning models, and the processor performs the following processes: dividing a plurality of training data sets, each having input data and a prior label associated with the input data, into one or more groups to generate the one or more input training data groups; and training the M machine learning models by inputting the corresponding input training data groups into each of the M machine learning models so as to reproduce the correspondence between the input data and the prior label associated with the input data. The process of generating the one or more input training data groups includes either a process of dividing each of the plurality of input data sets into one or more regions and generating a set of first-type divided input data after division that belong to the same region as the one input training data group; or a process of dividing the plurality of training data sets belonging to one class into one or more regions and generating a set of second-type divided input data after division as the one input training data group. According to this embodiment, a set of first-type divided input data can be used as one group of input training data for training one machine learning model, or a set of second-type divided input data can be used as one group of input training data for training one machine learning model. This reduces the amount of data used for training each machine learning model, thereby preventing the training time from becoming too long.

[0136] (19) According to a fourth aspect of the present disclosure, there is provided a classification device that classifies the class of data to be classified using M (M is an integer equal to or greater than 2) machine learning models of a vector neural network type having multiple vector neuron layers.This discrimination device includes: a memory that stores M machine learning models trained using a plurality of training data each having input data and a prior label associated with the input data, wherein each of the M machine learning models is trained by dividing the plurality of training data into one or more input training data groups and using the input training data group corresponding to one of the divided one or more input training data groups; and a processor that inputs the data to be discriminated into the M machine learning models and performs class discrimination of the data to be discriminated, wherein the processor performs a process of generating M sets of known feature spectra associated with the trained M machine learning models, wherein the M sets of known feature spectra include known feature spectra obtained from the output of a specific layer out of the plurality of vector neuron layers by inputting the input training data groups to the trained M machine learning models; a process of inputting data and obtaining, for each of the M machine learning models, individual data to be used for class discrimination of the discriminated data, wherein the individual data is generated for each of the M machine learning models using at least one of: (i) a similarity between a feature spectrum calculated from the output of the specific layer in response to the input of the discriminated data to the machine learning model and the group of known feature spectra; and (ii) an activation value corresponding to a judgment value of each class output from the output layer of the machine learning model in response to the input of the discriminated data; and a process of performing class discrimination of the discriminated data using the M individual data obtained for each of the M machine learning models, wherein the input training data group is either a set of first-type divided input data after dividing each of the multiple input data into one or more regions and dividing the multiple training data belonging to one class into one or more regions, or a set of second-type divided input data after the division process. According to this embodiment, the set of first-type divided input data can be used as one group of input learning data for training one machine learning model, or the set of second-type divided input data can be used as one group of input learning data for training one machine learning model. This makes it possible to reduce the amount of data used for training each machine learning model, and therefore to discriminate the class of data to be discriminated using a machine learning model that prevents the learning time from becoming long.

[0137] (20) According to a fifth aspect of the present disclosure, there is provided a computer program that causes a processor to execute training of M (M is an integer equal to or greater than 2) machine learning models used to classify data to be classified, the M machine learning models being of a vector neural network type having a plurality of vector neuron layers. The computer program includes: (a) a function of dividing a plurality of training data sets, each of the training data sets having input data and a prior label associated with the input data, into one or more groups to generate the one or more input training data sets; (b) A function of training the M machine learning models by inputting the corresponding input training data group into each of the M machine learning models so as to reproduce the correspondence between the input data and the prior label associated with the input data, wherein the function (a) includes either a function of dividing each of the multiple input data into one or more regions and generating a set of first-type divided input data after division that belong to the same region as the single input training data group, or a function of dividing the multiple training data belonging to one class into one or more regions and generating a set of second-type divided input data after division as the single input training data group. According to this embodiment, a set of first-type divided input data can be used as one group of input training data for training one machine learning model, or a set of second-type divided input data can be used as one group of input training data for training one machine learning model. This reduces the amount of data used for training each machine learning model, thereby preventing the training time from becoming too long.

[0138] (21) According to a sixth aspect of the present disclosure, there is provided a computer program for causing a processor to execute the following: classifying data to be classified using M (M is an integer equal to or greater than 2) machine learning models of a vector neural network type having multiple vector neuron layers.This computer program has the following functions: (a) a function for storing the M machine learning models trained using a plurality of training data each having input data and a prior label associated with the input data, wherein each of the M machine learning models is trained by dividing the plurality of training data into one or more input training data groups and using the input training data group corresponding to one of the one or more divided input training data groups; (b) a function for generating M sets of known feature spectra associated with the M trained machine learning models, wherein the M sets of known feature spectra include known feature spectra obtained from the output of a specific layer among the plurality of vector neuron layers by inputting the input training data groups to the M trained machine learning models; and (c) a function for inputting input discriminated data generated from the discriminated data to each of the M trained machine learning models, and generating the discriminated feature spectra for each of the M machine learning models. a function for obtaining individual data to be used for class discrimination of other data, wherein the individual data is generated, for each of the M machine learning models, using at least one of: (i) a similarity between the feature spectrum calculated from the output of the specific layer in response to input of the input discriminated data to the machine learning model and the group of known feature spectra; and (ii) an activation value corresponding to a judgment value of each class output from the output layer of the machine learning model in response to input of the input discriminated data; and (d) a function for performing class discrimination of the discriminated data using the M individual data obtained for each of the M machine learning models, wherein the input training data group is either a set of first-type divided input data after dividing each of the multiple input data into one or more regions and the divided input data belonging to the same region, or a set of second-type divided input data after performing a division process in which the multiple training data belonging to one class are divided into one or more regions. According to this embodiment, the set of first-type divided input data can be used as one group of input learning data for training one machine learning model, or the set of second-type divided input data can be used as one group of input learning data for training one machine learning model. This makes it possible to reduce the amount of data used for training each machine learning model, and therefore to discriminate the class of data to be discriminated using a machine learning model that prevents the learning time from becoming long.

[0139] The present disclosure may be realized in various forms other than those described above, such as a non-transitory storage medium on which a computer program is recorded. [Explanation of symbols]

[0140] DD,DD_1~DD_5...individual data, G1,G2...representative points, ID...input training data, IDA...input training data, IDAG...input training data group, IDAG1...first input training data group, IDAG2...second input training data group, IDG...input training data group, IDG_1~IDG_5...input training data group, IDa,IDa_1~IDa_5...first type divided input data, IDb...second type divided input data, IM...input data, IM_1~IM_5...input discriminated data, IMa,IMa_1~IMa_5...post-division input data, KSp...known feature spectrum, TD...training data data, TDG... group of learning data, 5... discrimination system, 10... target object, 20... discrimination device, 30... spectrometer, 110... processor, 112... data generation unit, 114... class discrimination processing unit, 120... memory, 130... interface circuit, 140... input device, 150... display unit, 200, 200_1 to 200_5... machine learning model, 210... convolution layer, 220... primary vector neuron layer, 230... first convolution vector neuron layer, 240... second convolution vector neuron layer, 250... classification vector neuron layer, 310... similarity calculation unit, 320... overall judgment unit

Claims

1. A method for training M machine learning models (M is an integer equal to or greater than 2) used to classify data to be classified, the M machine learning models being of a vector neural network type having a plurality of vector neuron layers, the method being executed by a processor, the method comprising: (a) preparing a plurality of training data sets having input data sets and prior labels associated with the input data sets; (b) dividing the plurality of training data into a plurality of groups to generate the plurality of input training data groups; (c) inputting the corresponding group of input training data into each of the M machine learning models, thereby training the M machine learning models so as to reproduce the correspondence between the input data and the prior label associated with the input data; The step (b) (b1) dividing each of the plurality of input data into a plurality of regions, and generating a set of first type divided input data after division that belong to the same region as one of the input learning data groups; (b2) dividing the plurality of learning data belonging to one class into a plurality of pieces and generating a set of the second type divided input data after the division as one group of input learning data.

2. A classification method in which a processor classifies data to be classified using M (M is an integer of 2 or more) machine learning models of a vector neural network type having a plurality of vector neuron layers, the method comprising: (a) preparing the M machine learning models trained using a plurality of training data sets each having input data and a prior label associated with the input data, wherein each of the M machine learning models is trained using a corresponding one of the input training data sets obtained by dividing the plurality of training data sets into a plurality of input training data groups; (b) preparing M sets of known feature spectra associated with the M trained machine learning models, the M sets of known feature spectra including known feature spectra obtained from outputs of specific layers among the plurality of vector neuron layers by inputting the input training data set to the M trained machine learning models; (c) inputting input discriminated data generated from the discriminated data into each of the M trained machine learning models to obtain individual data to be used for class discrimination of the discriminated data for each of the M machine learning models, wherein the individual data is generated for each of the M machine learning models using at least one of: (i) a similarity between a feature spectrum calculated from the output of the specific layer in response to input of the input discriminated data to the machine learning model and the group of known feature spectra; and (ii) an activation value corresponding to a judgment value for each class output from the output layer of the machine learning model in response to input of the input discriminated data; (d) performing class discrimination of the data to be discriminated using the M individual data obtained for each of the M machine learning models; The step (a) (a1) dividing each of the plurality of input data into a plurality of regions, and using a set of first type divided input data after division that belongs to the same region as one of the input learning data groups; (a2) performing a division process of dividing the plurality of learning data belonging to one class into a plurality of pieces, and using a set of second-type divided input data after the division process as one group of input learning data; The step (b) is executed when the individual data is generated in the step (c) using the similarity in the step (i). , determination method.

3. The method of claim 2, Each of the M individual data includes the activation value corresponding to each of the classes; The step (d) is a discrimination method in which, for each class, the class with the highest discrimination activation value calculated using a cumulative activation value obtained by adding up the activation values ​​of the M individual data is determined as the discrimination class.

4. The determination method according to claim 3, Each of the M individual data includes the similarity; The step (d) (d1) generating a discrimination similarity by integrating the similarities for each machine learning model; (d2) A discrimination method in which, when the discrimination similarity is equal to or greater than a predetermined threshold, the class with the highest discrimination activation value is determined as the discrimination class, and, when the discrimination similarity is less than the threshold, the discrimination class is determined as an unknown class different from the class corresponding to the prior label regardless of the discrimination activation value.

5. The method of claim 2, The step (c) A discrimination method including a step of generating, for each machine learning model, a class corresponding to the largest activation value among the activation values ​​corresponding to each class as a pre-discrimination class and as one element of the individual data.

6. The method of claim 2, The step (c) For each of the machine learning models, when the similarity is equal to or greater than a predetermined threshold, generating a class corresponding to the largest activation value among the activation values ​​corresponding to each of the classes as a pre-discrimination class and as one element of the individual data; A discrimination method comprising: for each machine learning model, when the similarity is less than the threshold, generating the pre-discriminated class as an unknown class different from the class corresponding to the pre-label as one element of the individual data.

7. The determination method according to any one of claims 2 to 6, In the step (c), when there are a plurality of specific layers, (i) calculating a multiplication value obtained by multiplying a weighting coefficient set for each of the plurality of specific layers by the similarity corresponding to one of the plurality of specific layers, and setting the sum of the calculated multiplication values ​​as the similarity used for class discrimination; (ii) determining the maximum or minimum value of the similarities corresponding to the plurality of specific layers as the similarity used for class discrimination.

8. The method of claim 2, each known characteristic spectrum included in the group of known characteristic spectra prepared in step (b) is associated with class classification information indicating to which class the known characteristic spectrum belongs; When the known feature spectrum associated with the class classification information is referred to as a class-specific known feature spectrum, The step (c) calculating a class-specific similarity, which is the similarity between the class-specific known feature spectrum and the feature spectrum, for each class; generating the class associated with the class-specific similarity having the largest value among the class-specific similarities calculated for each class as a pre-discrimination class and as one element of the individual data.

9. The method of claim 8, The step (c) of calculating the class-specific similarity includes: calculating the similarity between each of the plurality of known feature spectra for each class and the feature spectrum for each class; and calculating a representative similarity of the plurality of similarities for each of the classes by statistically processing the plurality of similarities calculated for each of the classes, The generating step in step (c) comprises: A discrimination method comprising a step of generating the class associated with the representative similarity having the largest value among the representative similarities calculated for each class as the pre-discrimination class and as one element of the individual data.

10. The method of claim 9, The method for determining whether or not a plurality of similarities are statistically processed comprises calculating a maximum value, a median value, an average value, or a mode value of the plurality of similarities as the representative similarity.

11. The determination method according to any one of claims 8 to 10, The step (c) further comprises: A discrimination method including a step of generating an unknown class different from the class corresponding to the pre-label as the pre-discrimination class as one element of the individual data, instead of the class associated with the class-specific similarity, when the largest value is less than a predetermined threshold.

12. The method according to any one of claims 5, 6 and 8, The step (d) A discrimination method comprising a step of determining the most frequent class from among the pre-discrimination classes of the individual data for each of the plurality of machine learning models as the class of the data to be discriminated.

13. The method according to any one of claims 5, 6 and 8, the step (c) includes a step of generating a similarity between a feature spectrum calculated from an output of the specific layer and the group of known feature spectra as one element of the individual data; The step (d) (i) determining, from among the pre-discrimination classes of the individual data for each of the plurality of machine learning models, the class having the highest similarity as the class of the data to be discriminated; (ii) calculating the sum or product of the similarities of the individual data for each class having the same pre-discrimination class, and determining the pre-discrimination class with the largest calculated value as the class of the data to be discriminated.

14. The method according to any one of claims 5, 6 and 8, The step (c) generating a similarity between the feature spectrum calculated from the output of the specific layer and the group of known feature spectra as one element of the individual data, The step (d) a reference value calculation step of calculating a reference value for each of the plurality of machine learning models using the similarity and a weighting coefficient that is preset for each of the plurality of machine learning models; A discrimination method comprising a class determination step of determining a class of the data to be discriminated using the pre-discrimination class and the calculated reference value.

15. The method of claim 14, In the reference value calculation step, the reference value is calculated by multiplying the similarity by the weighting coefficient for each of the plurality of machine learning models; The class determination step includes: (i) determining the pre-discrimination class having the largest sum of the reference values ​​of the machine learning model having the same pre-discrimination class as the class of the data to be discriminated; (ii) determining the pre-identified class of the machine learning model for which the reference value is maximum or minimum as the class of the data to be identified.

16. The method according to any one of claims 5, 6 and 8, The step (d) A discrimination method in which, when one of the multiple pre-discrimination classes corresponding to each of the multiple machine learning models indicates an unknown class different from the class corresponding to the pre-label, the unknown class is determined to be the class of the data to be discriminated, regardless of the class indicated by the other pre-discrimination classes.

17. The method according to any one of claims 2 to 16, The dividing treatment in the step (a2) is (i) clustering the plurality of training data belonging to the one class; or (ii) A discrimination method that is performed by randomly extracting the plurality of learning data belonging to the one class with sampling with replacement.

18. A learning device for M machine learning models (M is an integer of 2 or more) used to classify data to be classified, the M machine learning models being of a vector neural network type having a plurality of vector neuron layers, Memory and a processor that executes training of the M machine learning models, The processor: A process of dividing a plurality of training data sets each having input data and a prior label associated with the input data into a plurality of sets to generate the plurality of input training data sets; a process of inputting the corresponding group of input training data into each of the M machine learning models, thereby training the M machine learning models so as to reproduce the correspondence between the input data and the prior label associated with the input data; The process of generating the plurality of input learning data groups includes: a process of dividing each of the plurality of input data into a plurality of regions and generating a set of first type divided input data after division that belong to the same region as one of the input learning data groups; a process of dividing the plurality of learning data belonging to one class into a plurality of pieces and generating a set of second-type divided input data after division as one group of input learning data.

19. A classification device that classifies a class of data to be classified using M (M is an integer of 2 or more) machine learning models of a vector neural network type having a plurality of vector neuron layers, comprising: a memory that stores the M machine learning models trained using a plurality of training data sets having input data and a prior label associated with the input data sets, wherein each of the M machine learning models is trained by dividing the plurality of training data sets into a plurality of input training data groups and using the input training data group of a corresponding one of the divided plurality of input training data groups; a processor that inputs the discriminated data into the M machine learning models to perform class discrimination of the discriminated data; The processor: a process of generating M sets of known feature spectra associated with the M trained machine learning models, the M sets of known feature spectra including known feature spectra obtained from outputs of specific layers among the plurality of vector neuron layers by inputting the input training data group to the M trained machine learning models; a process of inputting discriminated data generated from the discriminated data into each of the M trained machine learning models, and obtaining individual data to be used for class discrimination of the discriminated data for each of the M machine learning models, wherein the individual data is generated for each of the M machine learning models using at least one of: (i) a similarity between a feature spectrum calculated from an output of the specific layer in response to input of the discriminated data to the machine learning model and the group of known feature spectra; and (ii) an activation value corresponding to a judgment value for each class output from an output layer of the machine learning model in response to input of the discriminated data for input; performing a process of classifying the data to be discriminated using the M individual data obtained for each of the M machine learning models; The input learning data group is Each of the plurality of input data is divided into a plurality of regions, and the divided input data is a set of first type divided input data that belong to the same region; a division process is performed to divide the plurality of learning data belonging to one class into a plurality of pieces, and the resultant set is a set of second type divided input data after the division process; The process of generating the M sets of known characteristic spectra is executed when the individual data is generated using the similarity in (i) in the process of obtaining the individual data. , discriminator.

20. A computer program that causes a processor to execute learning of M machine learning models (M is an integer of 2 or more) used to classify data to be classified, the M machine learning models being of a vector neural network type having a plurality of vector neuron layers, the computer program comprising: (a) a function of dividing a plurality of training data sets, each of which has input data and a prior label associated with the input data, into a plurality of sets to generate the plurality of input training data sets; (b) a function of inputting the corresponding group of input training data into each of the M machine learning models, thereby training the M machine learning models so as to reproduce the correspondence between the input data and the prior label associated with the input data; The function (a) is a function of dividing each of the plurality of input data into a plurality of regions and generating a set of first type divided input data after division that belong to the same region as one of the input learning data groups; a function of dividing the plurality of learning data belonging to one class into a plurality of pieces and generating a set of second-type divided input data after division as one group of input learning data.

21. A computer program causing a processor to execute a process of classifying a class of data to be classified using M (M is an integer equal to or greater than 2) machine learning models of a vector neural network type having a plurality of vector neuron layers, the computer program comprising: (a) a function of storing the M machine learning models trained using a plurality of training data sets each having input data and a prior label associated with the input data, wherein each of the M machine learning models is trained by dividing the plurality of training data sets into a plurality of input training data groups and using a corresponding one of the divided input training data groups; (b) a function of generating M sets of known feature spectra associated with the M trained machine learning models, the M sets of known feature spectra including known feature spectra obtained from outputs of specific layers among the plurality of vector neuron layers by inputting the input training data group to the M trained machine learning models; (c) a function of inputting input discriminated data generated from the discriminated data into each of the M trained machine learning models, and obtaining individual data to be used for class discrimination of the discriminated data for each of the M machine learning models, wherein the individual data is generated for each of the M machine learning models using at least one of: (i) a similarity between a feature spectrum calculated from the output of the specific layer in response to input of the input discriminated data to the machine learning model and the group of known feature spectra; and (ii) an activation value corresponding to a judgment value for each class output from the output layer of the machine learning model in response to input of the input discriminated data; (d) a function of performing class discrimination of the data to be discriminated using the M individual data obtained for each of the M machine learning models, The input learning data group is Each of the plurality of input data is divided into a plurality of regions, and the divided input data is a set of first type divided input data that belong to the same region; a division process is performed to divide the plurality of learning data belonging to one class into a plurality of pieces, and the resultant set is a set of second type divided input data after the division process; The function (b) is executed when the individual data is generated in the function (c) using the similarity in the function (i). , computer programs.

Citation Information

Patent Citations

  • Capsule Neural Network

    JP2021501392A

  • Vector neural network for low signal-to-noise ratio detection of a target

    US5210798A

  • Capsule neural networks

    WO2019083553A1

  • Cloud platform-based automatic identification system and method for seven types of mass spectrograms of commonly used pesticides and chemical pollutants around the world

    WO2020191857A1