Training method, model generation method, training apparatus, training program, prediction method, prediction apparatus, and prediction program
The training method for neural networks with a softmax function in the final layer addresses convergence issues in regression problems by improving gradient balance and enhancing prediction accuracy.
Patent Information
- Application Number
- JP2023211524
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-26
AI Technical Summary
Neural networks (NNs) processing regression problems face convergence issues when the gradient of model parameters in one dimension becomes significantly larger than in other dimensions, making the training process difficult to converge.
The proposed solution involves a training method that uses a neural network model with a predetermined conversion function in the final layer, such as a softmax function, to improve convergence performance. This method inputs data into the neural network, obtains an output from the conversion function, and updates model parameters using a loss calculated based on prediction data and correct answer data.
The approach enhances the convergence performance of the training process for NNs processing regression problems, even when gradients in different dimensions are uneven, allowing for more effective model parameter updates and improved prediction accuracy.
Smart Images

Figure 2025095490000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a training method, a model generation method, a training device, a training program, a prediction method, a prediction device, and a prediction program.
Background Art
[0002] There is known a prediction technique that processes a regression problem using an NN (Neural Network) and outputs prediction data. In the case of an NN that processes a regression problem, in a case where the gradient of the model parameters in any dimension becomes larger than the gradient of the model parameters in other dimensions, the training process may be difficult to converge.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present disclosure improves the convergence performance in the training process for an NN that processes a regression problem.
Means for Solving the Problems
[0005] The training method according to one aspect of the present disclosure has, for example, the following configuration. That is, at least one processor inputs input data into a neural network model that processes a regression problem and has a predetermined conversion function in the final layer, and obtains an output from the predetermined conversion function, uses the output and a set of elements having the same number of elements as the number of dimensions of the output to obtain prediction data, Execute a process of updating the model parameters of the neural network model using the loss calculated based on the prediction data and the correct answer data.
Brief Description of Drawings
[0006]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 6A
Figure 6B
Figure 7A
Figure 7B
Figure 8A
Figure 8B
Figure 9A
Figure 9B
Figure 10A
Figure 10B
Embodiments for Carrying Out the Invention
[0007] Hereinafter, each embodiment will be described with reference to the accompanying drawings. In the present specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions are omitted.
[0008] [First Embodiment] <System Configuration of Information Processing System> First, the system configuration of an information processing system including a training device and a prediction device according to the first embodiment will be described. FIG. 1 is a diagram showing an example of the system configuration of the information processing system. As shown in FIG. 1, the information processing system 100 includes a training device 110 and a prediction device 120.
[0009] A training program is installed in the training device 110, and when the training program is executed, the training device 110 functions as a training unit 111.
[0010] The training unit 111 has an NN unit that processes regression problems. The training unit 111 performs supervised training processing on the NN unit using the training data stored in advance in the training data storage unit 112. Note that the training data storage unit 112 stores training data including input data and correct answer data.
[0011] A prediction program is installed in the prediction device 120. When the program is executed, the prediction device 120 functions as a prediction unit 121. The prediction unit 121 includes an NN unit that has been trained by the training unit 111. The prediction unit 121 processes a regression problem using the NN unit that has been trained. Specifically, the prediction unit 121 inputs input data to the NN unit that has been trained by the training unit 111 and outputs prediction data by executing the NN unit.
[0012] In the example of FIG. 1, in the information processing system 100, the training device 110 and the prediction device 120 are configured as separate entities, but the training device 110 and the prediction device 120 may be configured as one body. Also, in FIG. 1, an example is shown in which the training device 110 and the prediction device 120 are each configured by one device, but the training device 110 and the prediction device 120 may each be configured by a plurality of devices (that is, each may be configured as a system).
[0013] <Hardware Configuration of Training Device and Prediction Device> Next, the hardware configurations of the training device 110 and the prediction device 120 will be described. Since the hardware configuration of the training device 110 is the same as the hardware configuration of the prediction device 120, the hardware configuration of the training device 110 will be described here.
[0014] FIG. 2 is a diagram showing an example of the hardware configuration of the training device. The training device 110 includes, as components, a processor 201, a main storage device 202 (memory), an auxiliary storage device 203 (memory), a network interface 204, and a device interface 205. The training device 110 may be realized as a computer in which these components are connected via a bus 206. In the example of FIG. 2, the training device 110 is shown as having one of each component, but the training device 110 may have a plurality of the same components.
[0015] Various operations of the training device 110 may be executed in parallel using one or more processors. Also, various operations may be distributed to a plurality of arithmetic cores within the processor 201 and executed in parallel. Further, part or all of the processes, means, etc. of the present disclosure may be executed by an external device 230 (at least one of a processor and a storage device) provided on a cloud that can communicate with the training device 110 via the network interface 204.
[0016] The processor 201 may be an electronic circuit (processing circuit, Processing circuit, Processing circuitry, CPU, GPU, FPGA, or ASIC, etc.). Also, the processor 201 may be a semiconductor device including a dedicated processing circuit, etc. Note that the processor 201 is not limited to an electronic circuit using electronic logic elements, and may be realized by an optical circuit using optical logic elements. Further, the processor 201 may include an arithmetic function based on quantum computing.
[0017] The processor 201 performs various operations based on various data and instructions input from each device, etc. of the internal configuration of the training device 110, and outputs the operation results and control signals to each device, etc. The processor 201 controls each component included in the training device 110 by executing an OS (Operating System), an application, etc.
[0018] Also, the processor 201 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or devices. When using a plurality of electronic circuits, each electronic circuit may communicate wired or wirelessly.
[0019] The main memory device 202 is a storage device that stores instructions and various data executed by the processor 201, and various data stored in the main memory device 202 is read by the processor 201. The auxiliary storage device 203 is a storage device other than the main memory device 202. Note that these storage devices mean any electronic components capable of storing various data (for example, training data stored in the training data storage unit 112), and may be semiconductor memories. The semiconductor memory may be either a volatile memory or a non-volatile memory. The storage device for storing various data in the training device 110 may be realized by the main memory device 202 or the auxiliary storage device 203, or may be realized by a built-in memory built into the processor 201.
[0020] Also, a plurality of processors 201 may be connected (coupled) to one main memory device 202, or a single processor 201 may be connected. Alternatively, a plurality of main memory devices 202 may be connected (coupled) to one processor 201. When the training device 110 is composed of at least one main memory device 202 and a plurality of processors 201 connected (coupled) to this at least one main memory device 202, the configuration may include that at least one of the plurality of processors 201 is connected (coupled) to at least one main memory device 202.
[0021] The network interface 204 is an interface for connecting to the communication network 220, either wirelessly or by wire.
[0022] The device interface 205 is an interface such as USB that directly connects to the external device 240.
[0023] The external device 240 may be, for example, an input device. In the present embodiment, the input device is an electronic device such as a camera, a microphone, various sensors, a keyboard, a mouse, or a touch panel, and provides the acquired information to the training device 110.
[0024] Further, the external device 240 may be, for example, an output device. In the present embodiment, the output device may be, for example, a display device such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), or an organic EL (Electro Luminescence) panel, or may be a speaker or the like that outputs sound or the like.
[0025] Further, the external device 240 may be a storage device (memory). For example, the external device 240 may be a network storage or the like, and the external device 240 may be a storage such as an HDD.
[0026] Further, the external device 240 may be a device having some functions of the components of the training device 110. That is, the training device 110 may transmit and receive processing results to and from the external device 240.
[0027] <Functional Configuration of Training Device> Next, the functional configuration of the training device 110 according to the first embodiment will be described. FIG. 3A is a diagram showing an example of the functional configuration of the training device according to the first embodiment. As shown in FIG. 3A, the training unit 111 includes an NN unit 310, a conversion unit 320, and a loss function unit 330.
[0028] The NN unit 310 has an input layer, an intermediate layer, and a final layer, and performs forward propagation processing when the input data read from the training data storage unit 112 is input.
[0029] Specifically, in the NN unit 310, the input layer processes the input input data and outputs the processing result to the intermediate layer. There are a plurality of intermediate layers, and the processing result of the input layer is sequentially processed in each of the plurality of intermediate layers and output to the final layer. The final layer has a softmax function. Therefore, the output data output from the final layer is output data having the number of dimensions of x dimensions (x is an integer of 2 or more). In the example of FIG. 3A, the case where the NN unit 310 has a plurality of intermediate layers is shown, but the intermediate layer may be singular.
[0030] The conversion unit 320 has a product-sum operation unit. The product-sum operation unit has an array (x-dimensional array) having the same number of elements as the output dimensionality (x) of the output data output from the final layer of the NN unit 310. Note that the "array having the same number of elements as the output dimensionality" is a form of the "set of elements having the same number of elements as the output dimensionality". The product-sum operation unit performs a product-sum operation on the output data output from the final layer of the NN unit 310 using the x-dimensional array, and outputs the operation result to the loss function unit 330.
[0031] The loss function unit 330 reads the correct data from the training data storage unit 112. Also, the loss function unit 330 calculates the loss of the operation result obtained by performing the product-sum operation on the read correct data. Note that the loss calculated by the loss function unit 330 is input to the final layer of the NN unit 310, and backpropagation processing is performed. As a result, the model parameters included in the input layer, intermediate layer, and final layer in the NN unit 310 are updated (training processing for the NN unit 310 is performed).
[0032] In this way, when performing training processing on the NN unit 310 that processes a regression problem, the training device 110 according to the first embodiment arranges the softmax function in the final layer of the NN unit 310. Thereby, even in a case where the gradient of the model parameter in any dimension becomes larger than the gradient of the model parameter in other dimensions, the convergence performance in the training processing can be improved.
[0033] Also, the training device 110 according to the first embodiment performs a product-sum operation on the x-dimensional output data output from the final layer of the NN unit 310 using the x-dimensional array. Thereby, the x-dimensional output data output from the final layer can be converted into an arbitrary scalar value. As a result, according to the training device 110 according to the first embodiment, it becomes possible to convert and output data suitable for the processing result of the regression problem.
[0034] <Functional Configuration of Prediction Device> Next, the functional configuration of the prediction device 120 will be described. FIG. 3B is a diagram showing an example of the functional configuration of the prediction device according to the first embodiment. As shown in FIG. 3B, the prediction unit 121 includes an NN unit 311 and a conversion unit 320.
[0035] The NN unit 311 has the same configuration as the NN unit 310 of the training unit 111 shown in FIG. 3A. However, the model parameters of the input layer, intermediate layer, and final layer of the NN unit 311 have been updated by the training process.
[0036] When input data is input to the NN unit 311, the NN unit 311 performs forward propagation processing. Specifically, the input layer processes the input input data and outputs the processing result to the intermediate layer. Also, the processing result of the input layer is sequentially processed in each of the plurality of intermediate layers and output to the final layer. Since the final layer has a softmax function, x-dimensional output data is output from the final layer.
[0037] The conversion unit 320 is the same as the conversion unit 320 of the training unit 111 shown in FIG. 3A and has a product-sum operation unit. The product-sum operation unit performs a product-sum operation on the output data output from the final layer of the NN unit 311 using an x-dimensional array, and outputs prediction data suitable for the processing result of the regression problem.
[0038] <Summary> As is clear from the above description, the training device 110 according to the first embodiment · The final layer of the NN unit 310 has a softmax function. · It has a conversion unit 320 that performs a product-sum operation on the x-dimensional output data output from the final layer of the NN unit 310 using an x-dimensional array.
[0039] Accordingly, according to the training device 110 according to the first embodiment, even in a case where the gradient of the model parameter in any dimension becomes larger than the gradient of the model parameter in other dimensions, the convergence performance in the training process can be improved. Also, it becomes possible to convert and output data suitable for the processing result of the regression problem.
[0040] Also, the prediction device 120 according to the first embodiment · The final layer of the NN unit 311 that has undergone the training process by the training device 110 has a softmax function. · It has a conversion unit 320 that performs a dot product operation on the x-dimensional output data output from the final layer of the NN unit 311 that has undergone the training process by the training device 110, using an x-dimensional array.
[0041] Thus, according to the prediction device 120 according to the first embodiment, prediction data suitable for the processing result of the regression problem can be output.
[0042] [Second Embodiment] In the above first embodiment, the dot product unit of the conversion unit 320 is configured to have a predetermined x-dimensional array (that is, an x-dimensional array of fixed values). In contrast, in the second embodiment, the x-dimensional array included in the dot product unit of the conversion unit 320 is configured to be updated in the training process. Hereinafter, the second embodiment will be described centering on the differences from the first embodiment.
[0043] [Functional Configuration of the Training Device] First, the functional configuration of the training device 110 according to the second embodiment will be described. FIG. 4A is a diagram showing an example of the functional configuration of the training device according to the second embodiment. As shown in FIG. 4A, the training unit 111 includes an NN unit 310, a conversion unit 420, and a loss function unit 330.
[0044] Among these, since the NN unit 310 and the loss function unit 330 are the same as the NN unit 310 and the loss function unit 330 of the training unit 111 shown in FIG. 3A, the description thereof will be omitted here.
[0045] The conversion unit 420 includes a dot product unit. The dot product unit has an array (x-dimensional array) having the same number of elements as the number of output dimensions (x) of the output data output from the final layer of the NN unit 310, and performs a dot product operation on the output data output from the final layer of the NN unit 310 using the x-dimensional array.
[0046] Each element of the x-dimensional array included in the conversion unit 420 has a default element value, and is updated based on the loss output from the loss function unit 330 during the training process.
[0047] As described above, the training device 110 according to the second embodiment has the same configuration as the training device 110 according to the first embodiment, and further, each element of the x-dimensional array included in the product-sum calculation unit of the conversion unit 420 is configured to be updated by the training process. Accordingly, according to the second embodiment, while enjoying the same effects as the first embodiment, it is further possible to generate an optimal x-dimensional array.
[0048] <Functional Configuration of Prediction Device> Next, the functional configuration of the prediction device 120 according to the second embodiment will be described. FIG. 4B is a diagram showing an example of the functional configuration of the prediction device according to the second embodiment. As shown in FIG. 4B, the prediction unit 121 includes an NN unit 311 and a conversion unit 421.
[0049] Among these, since the NN unit 311 is the same as the NN unit 311 of the prediction unit 121 shown in FIG. 3B, the description thereof will be omitted here.
[0050] The conversion unit 421 is the same as the conversion unit 320 of the prediction unit 121 shown in FIG. 3B and includes a product-sum calculation unit. However, the x-dimensional array included in the product-sum calculation unit of the conversion unit 421 has been updated by the training process.
[0051] The product-sum calculation unit of the conversion unit 421 performs a product-sum calculation on the output data output from the final layer of the NN unit 311 using the x-dimensional array updated by the training process, and outputs prediction data suitable for the processing result of the regression problem.
[0052] <Summary> As is clear from the above description, the training device 110 according to the second embodiment · The final layer of the NN unit 310 has a softmax function. ·It has a conversion unit 420 that performs a dot product operation on the x-dimensional output data output from the final layer of the NN unit 310 using an x-dimensional array. ·The x-dimensional array used when the conversion unit 420 performs the dot product operation is updated by the training process.
[0053] Thus, according to the training device 110 according to the second embodiment, while enjoying the same effects as the first embodiment, it is further possible to generate an optimal x-dimensional array.
[0054] Also, the prediction device 120 according to the second embodiment ·The final layer of the NN unit 311 trained by the training device 110 has a softmax function. ·It has a conversion unit 421 that performs a dot product operation on the x-dimensional output data output from the final layer of the NN unit 311 trained by the training device 110 using the x-dimensional array updated by the training process.
[0055] Thus, according to the prediction device 120 according to the second embodiment, prediction data suitable for the processing result of the regression problem can be output.
[0056] [Third Embodiment] In the first embodiment above, the dot product unit of the conversion unit 320 has been described as having a predetermined x-dimensional array. In contrast, in the third embodiment, the x-dimensional array of the dot product unit of the conversion unit 320 is configured to be provided by an NN unit that calculates the x-dimensional array. Hereinafter, the third embodiment will be described centering on the differences from the first embodiment.
[0057] <Functional Configuration of Training Device> First, the functional configuration of the training device 110 according to the third embodiment will be described. FIG. 5A is a diagram showing an example of the functional configuration of the training device according to the third embodiment. As shown in FIG. 5A, the training unit 111 includes an NN unit 310, a conversion unit 520, a loss function unit 330, and an NN unit 510 (an example of a neural network model for outputting an element set).
[0058] Among these, since the NN unit 310 and the loss function unit 330 are the same as those of the NN unit 310 and the loss function unit 330 of the training unit 111 shown in FIG. 3A, the description thereof is omitted here.
[0059] The conversion unit 520 has a product-sum operation unit. The product-sum operation unit acquires an array (x-dimensional array) having the same number of elements as the output dimensionality (x) of the output data output from the final layer of the NN unit 310 from the NN unit 510. Further, the product-sum operation unit performs a product-sum operation on the output data output from the final layer of the NN unit 310 using the acquired x-dimensional array.
[0060] The NN unit 510 has an input layer, an intermediate layer, and a final layer (however, omitted in FIG. 5A). When the input data read from the training data storage unit 112 is input, forward propagation processing is performed, and an x-dimensional array is output. The x-dimensional array output by the NN unit 510 is provided to the product-sum operation unit of the conversion unit 520.
[0061] Note that the loss calculated by the loss function unit 330 is input to the final layer of the NN unit 510, and backpropagation processing is performed. Thereby, the model parameters included in the input layer, the intermediate layer, and the final layer in the NN unit 510 are updated (training processing for the NN unit 510 is performed).
[0062] As described above, the training device 110 according to the third embodiment has the same configuration as the training device 110 according to the first embodiment, and further, each element of the x-dimensional array included in the product-sum operation unit of the conversion unit 520 is configured to be provided by the NN unit 510. Thereby, according to the third embodiment, while enjoying the same effects as those of the first embodiment, it is further possible to generate an optimal x-dimensional array.
[0063] <Functional Configuration of Prediction Device> Next, the functional configuration of the prediction device 120 according to the third embodiment will be described. FIG. 5B is a diagram showing an example of the functional configuration of the prediction device according to the third embodiment. As shown in FIG. 5B, the prediction unit 121 includes an NN unit 311, a conversion unit 520, and an NN unit 511.
[0064] Among these, since the NN unit 311 is the same as the NN unit 311 of the prediction unit 121 shown in FIG. 3B, the description thereof will be omitted here. Also, since the conversion unit 520 is the same as the conversion unit 520 of the training unit 111 shown in FIG. 5A, the description thereof will be omitted here.
[0065] The NN unit 511 has the same configuration as the NN unit 510 of the training unit 111 shown in FIG. 5A. However, the model parameters of the input layer, intermediate layer, and final layer that the NN unit 511 has are updated by the training process.
[0066] When input data is input to the NN unit 511, forward propagation processing is performed, and when an x-dimensional array is output, the x-dimensional array is provided to the product-sum operation unit of the conversion unit 520.
[0067] <Summary> As is clear from the above description, the training device 110 according to the third embodiment · The final layer of the NN unit 310 has a softmax function. · It has a conversion unit 520 that performs a product-sum operation on the x-dimensional output data output from the final layer of the NN unit 310 using an x-dimensional array. · As the x-dimensional array used when the conversion unit 520 performs the product-sum operation, an x-dimensional array output by inputting input data to the NN unit 510 is used. · The NN unit 510 is subjected to a training process using the loss calculated by the loss function unit 330.
[0068] Accordingly, according to the training device 110 according to the third embodiment, while enjoying the same effects as those of the first embodiment, it is further possible to generate an optimal x-dimensional array.
[0069] In addition, the prediction device 120 according to the third embodiment · The final layer of the NN unit 311 that has undergone the training process by the training device 110 has a softmax function. · It has a conversion unit 520 that performs a dot product operation on the x-dimensional output data output from the final layer of the NN unit 311 that has undergone the training process by the training device 110, using an x-dimensional array. · As the x-dimensional array used when the conversion unit 520 performs the dot product operation, an x-dimensional array output by inputting input data to the trained NN unit 511 is used.
[0070] Accordingly, according to the training device 110 according to the third embodiment, prediction data suitable for the processing result of the regression problem can be output.
[0071] [Fourth Embodiment] In the above third embodiment, the NN unit 510 that calculates the x-dimensional array has been described as having the same configuration as the NN unit 310. In contrast, in the fourth embodiment, a case where the NN unit 510 that calculates the x-dimensional array has a configuration different from that of the NN unit 310 will be described. Hereinafter, the fourth embodiment will be described focusing on the differences from the above third embodiment.
[0072] [Functional Configuration of Training Device] First, the functional configuration of the training device 110 according to the fourth embodiment will be described. FIG. 6A is a diagram showing an example of the functional configuration of the training device according to the fourth embodiment. As shown in FIG. 6A, the training unit 111 includes an NN unit 310, a conversion unit 520, a loss function unit 330, and an NN unit 610 (an example of a neural network model for outputting an element set).
[0073] Among these, since the NN unit 310, the conversion unit 520, and the loss function unit 330 are the same as the NN unit 310, the conversion unit 520, and the loss function unit 330 of the training unit 111 shown in FIG. 5A, the description thereof will be omitted here.
[0074] The NN unit 610 does not have a layer corresponding to the input layer and part of the intermediate layer of the NN unit 310, and outputs an x-dimensional array with the output of the intermediate layer of the NN unit 310 as the input. The x-dimensional array output by the NN unit 610 is provided to the product-sum calculation unit of the conversion unit 520.
[0075] Note that the loss calculated by the loss function unit 330 is input to the final layer of the NN unit 610, and backpropagation processing is performed. As a result, the model parameters included in part of the intermediate layer and the final layer in the NN unit 610 are updated (training processing for the NN unit 610 is performed).
[0076] Thus, while the training device 110 according to the fourth embodiment has the same configuration as the training device 110 according to the third embodiment described above, the configuration of the NN unit 610 that outputs an x-dimensional array is simpler than the configuration of the NN unit 510 that outputs an x-dimensional array. Accordingly, according to the fourth embodiment, while enjoying the same effects as the third embodiment described above, it is possible to further reduce the training cost of the NN unit 610 that outputs an x-dimensional array.
[0077] <Functional Configuration of Prediction Device> Next, the functional configuration of the prediction device 120 according to the fourth embodiment will be described. FIG. 6B is a diagram showing an example of the functional configuration of the prediction device according to the fourth embodiment. As shown in FIG. 6B, the prediction unit 121 includes an NN unit 311, a conversion unit 520, and an NN unit 611.
[0078] Among these, since the NN unit 311 is the same as the NN unit 311 of the prediction unit 121 shown in FIG. 5B, the description thereof will be omitted here. Also, since the conversion unit 520 is the same as the conversion unit 520 of the training unit 111 shown in FIG. 6A, the description thereof will be omitted here.
[0079] The NN unit 611 has the same configuration as the NN unit 610 of the training unit 111 shown in FIG. 6A. However, part of the intermediate layer and the model parameters of the final layer included in the NN unit 611 have been updated by the training process.
[0080] Data output from a part of the intermediate layer of the NN unit 311 is input to the NN unit 611, and forward propagation processing is performed. By outputting an x-dimensional array, the x-dimensional array is provided to the product-sum operation unit of the conversion unit 520.
[0081] <Summary> As is clear from the above description, the training device 110 according to the fourth embodiment · The final layer of the NN unit 310 has a softmax function. · It has a conversion unit 520 that performs a product-sum operation on the x-dimensional output data output from the final layer of the NN unit 310 using an x-dimensional array. · As the x-dimensional array used when the conversion unit 520 performs the product-sum operation, an x-dimensional array output by inputting data output from a part of the intermediate layer in the NN unit 310 to the NN unit 610 is used. · Training processing is performed on the NN unit 610 using the loss calculated by the loss function unit 330.
[0082] Thus, according to the training device 110 according to the fourth embodiment, while enjoying the same effects as the third embodiment, it is possible to further reduce the training cost of the NN unit 610 that outputs an x-dimensional array.
[0083] Also, the prediction device 120 according to the fourth embodiment · The final layer of the NN unit 311 trained by the training device 110 has a softmax function. · It has a conversion unit 520 that performs a product-sum operation on the x-dimensional output data output from the final layer of the NN unit 311 trained by the training device 110 using an x-dimensional array. · As the x-dimensional array used when the conversion unit 520 performs the product-sum operation, an x-dimensional array output by inputting data output from a part of the intermediate layer in the trained NN unit 311 to the trained NN unit 611 is used.
[0084] Thus, according to the prediction device 120 according to the fourth embodiment, prediction data suitable for the processing result of the regression problem can be output.
[0085] [Fifth Embodiment] In the above fourth embodiment, the NN unit 610 that calculates the x-dimensional array is configured not to have a layer corresponding to a part of the input layer and the intermediate layer, thereby realizing a reduction in training cost. In contrast, in the fifth embodiment, a part of the model parameters of the NN unit that calculates the x-dimensional array is shared with a part of the model parameters of the NN unit 310, thereby realizing a reduction in training cost. Hereinafter, the fifth embodiment will be described centering on the differences from the above third embodiment.
[0086] [Functional Configuration of Training Device] First, the functional configuration of the training device 110 according to the fifth embodiment will be described. FIG. 7A is a diagram showing an example of the functional configuration of the training device according to the fifth embodiment. As shown in FIG. 7A, the training unit 111 includes an NN unit 310, a conversion unit 520, a loss function unit 330, and an NN unit 710 (an example of a neural network model for outputting an element set).
[0087] Among these, the NN unit 310, the conversion unit 520, and the loss function unit 330 are the same as the NN unit 310, the conversion unit 520, and the loss function unit 330 of the training unit 111 shown in FIG. 5A, and thus the description thereof will be omitted here.
[0088] The NN unit 710 has an input layer, an intermediate layer, and a final layer. When the input data read from the training data storage unit 112 is input, forward propagation processing is performed, and an x-dimensional array is output. The x-dimensional array output by the NN unit 710 is provided to the product-sum operation unit of the conversion unit 520.
[0089] Note that the loss calculated by the loss function unit 330 is input to the final layer of the NN unit 310, and backpropagation processing is performed. As a result, the model parameters included in the input layer, the intermediate layer, and the final layer in the NN unit 310 are updated (training processing for the NN unit 310 is performed).
[0090] The model parameters updated in the final layer and part of the intermediate layer within the NN unit 310 are notified to the NN unit 710 and set as the model parameters of the final layer and part of the intermediate layer within the NN unit 710. Thereby, the model parameters of the final layer and part of the intermediate layer within the NN unit 310 and the model parameters of the final layer and part of the intermediate layer within the NN unit 710 are shared.
[0091] Also, by performing backpropagation processing during the training process for the NN unit 310, the data output from part of the intermediate layer within the NN unit 310 is input to part of the intermediate layer within the NN unit 710, and backpropagation processing is performed.
[0092] Thereby, the model parameters included in part of the intermediate layer and the input layer within the NN unit 710 are updated (training processing for the NN unit 710 is performed).
[0093] Thus, the training device 110 according to the fifth embodiment has the same configuration as the training device 110 according to the third embodiment, while sharing part of the model parameters within the NN unit 710 with part of the model parameters within the NN unit 310. Thereby, according to the fifth embodiment, while enjoying the same effects as the third embodiment, it is further possible to reduce the training cost of the NN unit 710 that outputs an x-dimensional array.
[0094] <Functional Configuration of Prediction Device> Next, the functional configuration of the prediction device 120 according to the fifth embodiment will be described. FIG. 7B is a diagram showing an example of the functional configuration of the prediction device according to the fifth embodiment. As shown in FIG. 7B, the prediction unit 121 includes an NN unit 311, a conversion unit 520, and an NN unit 711.
[0095] Among these, since the NN unit 311 is the same as the NN unit 311 of the prediction unit 121 shown in FIG. 5B, the description thereof will be omitted here. Also, since the conversion unit 520 is the same as the conversion unit 520 of the training unit 111 shown in FIG. 7A, the description thereof will be omitted here.
[0096] The NN unit 711 has the same configuration as the NN unit 710 of the training unit 111 shown in FIG. 7A. However, a part of the intermediate layer and the model parameters of the final layer of the NN unit 711 are shared with a part of the intermediate layer and the model parameters of the final layer of the NN unit 311 of the prediction unit 121 shown in FIG. 7B. Also, a part of the intermediate layer and the model parameters of the input layer of the NN unit 711 are updated by the training process.
[0097] When input data is input to the NN unit 711, forward propagation processing is performed, and when an x-dimensional array is output, the x-dimensional array is provided to the product-sum calculation unit of the conversion unit 520.
[0098] <Summary> As is clear from the above description, the training device 110 according to the fifth embodiment · The final layer of the NN unit 310 has a softmax function. · It has a conversion unit 520 that performs a product-sum operation on the x-dimensional output data output from the final layer of the NN unit 310 using an x-dimensional array. · As the x-dimensional array used when the conversion unit 520 performs the product-sum operation, the x-dimensional array output by inputting input data to the NN unit 710 is used. · The NN unit 710 is trained using the loss calculated by the loss function unit 330. At this time, the model parameters of a part of the final layer and the intermediate layer of the NN unit 710 are shared with the model parameters of a part of the final layer and the intermediate layer of the NN unit 310.
[0099] Thus, according to the training device 110 according to the fifth embodiment, while enjoying the same effects as in the third embodiment, it is further possible to reduce the training cost of the NN unit 710 that outputs an x-dimensional array.
[0100] Also, the prediction device 120 according to the fifth embodiment · The final layer of the NN unit 311 trained by the training device 110 has a softmax function. ·The conversion unit 520 has a dot product operation using an x-dimensional array on the x-dimensional output data output from the final layer of the NN unit 311 for which training processing has been performed by the training device 110. ·As the x-dimensional array used when the conversion unit 520 performs the dot product operation, an x-dimensional array output by inputting input data to the NN unit 711 that has been trained and in which part of the final layer and intermediate layer of the NN unit 311 share model parameters with the NN unit 711 is used.
[0101] Accordingly, according to the prediction device 120 according to the fifth embodiment, prediction data suitable for the processing result of the regression problem can be output.
[0102] [Sixth Embodiment] In the first embodiment described above, the softmax function is arranged in the final layer of the NN unit 310. In contrast, in the sixth embodiment, a non-linear conversion function or a linear conversion function x other than the softmax function is arranged in the final layer of the NN unit 310. Hereinafter, the sixth embodiment will be described centering on the differences from the first embodiment.
[0103] <Functional Configuration of Training Device> First, the functional configuration of the training device 110 according to the sixth embodiment will be described. FIG. 8A is a diagram showing an example of the functional configuration of the training device according to the sixth embodiment. As shown in FIG. 8A, the training unit 111 includes an NN unit 810, a conversion unit 320, and a loss function unit 330.
[0104] Among these, since the conversion unit 320 and the loss function unit 330 are the same as the conversion unit 320 and the loss function unit 330 of the training unit 111 shown in FIG. 3A, the description thereof will be omitted here.
[0105] The NN unit 810 has an input layer, an intermediate layer, and a final layer, and forward propagation processing is performed when the input data read from the training data storage unit 112 is input.
[0106] Specifically, in the NN unit 810, the input layer processes the input input data and outputs the processing result to the intermediate layer. There are a plurality of intermediate layers, and the processing result of the input layer is sequentially processed in each of the plurality of intermediate layers and output to the final layer. The final layer has a non-linear transformation function or a linear transformation function other than the softmax function. Therefore, the output data output from the final layer is output data having the number of dimensions of x dimensions.
[0107] Note that the non-linear transformation function or the linear transformation function other than the softmax function is, for example, a function that · sets the minimum value to a predetermined value (for example, "0"), · sets the maximum value to a predetermined value (for example, "1"), · non-linearly transforms or linearly transforms so that the sum of all elements becomes a predetermined value. It may be a non-linear transformation or a linear transformation function.
[0108] Thus, when performing the training process on the NN unit 810 that processes the regression problem, the training device 110 according to the sixth embodiment is configured to arrange a non-linear transformation function or a linear transformation function other than the softmax function in the final layer of the NN unit 810.
[0109] Accordingly, according to the training device 110 according to the sixth embodiment, the same effects as those of the training device 110 according to the first embodiment can be enjoyed.
[0110] <Functional Configuration of Prediction Device> Next, the functional configuration of the prediction device 120 according to the sixth embodiment will be described. FIG. 8B is a diagram showing an example of the functional configuration of the prediction device according to the sixth embodiment. As shown in FIG. 8B, the prediction unit 121 includes an NN unit 811 and a conversion unit 320.
[0111] Among these, since the conversion unit 320 is the same as the conversion unit 320 of the prediction unit 121 shown in FIG. 3B, the description thereof will be omitted here.
[0112] The NN unit 811 has the same configuration as the NN unit 810 of the training unit 111 shown in FIG. 8A. However, the model parameters of the input layer, intermediate layer, and final layer of the NN unit 811 have been updated by the training process.
[0113] When input data is input to the NN unit 811, forward propagation processing is performed. Specifically, in the NN unit 811, the input layer processes the input input data and outputs the processing result to the intermediate layer. Also, the processing result of the input layer is sequentially processed in each of the plurality of intermediate layers and output to the final layer. Since the final layer has a non-linear transformation function or a linear transformation function other than the softmax function, x-dimensional output data is output from the final layer.
[0114] <Summary> As is clear from the above description, the training device 110 according to the sixth embodiment · The final layer of the NN unit 810 has a non-linear transformation function or a linear transformation function other than the softmax function. · It has a conversion unit 320 that performs a sum-of-products operation on the x-dimensional output data output from the final layer of the NN unit 810 using an x-dimensional array.
[0115] Accordingly, according to the training device 110 according to the sixth embodiment, even in a case where the gradient of the model parameters in any dimension becomes larger than the gradient of the model parameters in other dimensions, the convergence performance in the training process can be improved. Also, it becomes possible to convert and output to data suitable for the processing result of the regression problem.
[0116] Also, the prediction device 120 according to the sixth embodiment · The final layer of the NN unit 311 trained by the training device 110 has a non-linear transformation function or a linear transformation function other than the softmax function. · It has a conversion unit 320 that performs a sum-of-products operation on the x-dimensional output data output from the final layer of the NN unit 311 trained by the training device 110 using an x-dimensional array.
[0117] Accordingly, according to the prediction device 120 according to the sixth embodiment, prediction data suitable for the processing result of the regression problem can be output.
[0118] [Seventh Embodiment] In the second embodiment, when updating the x-dimensional array included in the product-sum calculation unit of the conversion unit 420 in the training process, a predetermined x-dimensional array is set as the initial value. On the other hand, in the seventh embodiment, the x-dimensional array output from the random number unit is set as the initial value. Hereinafter, the seventh embodiment will be described centering on the differences from the second embodiment.
[0119] [Functional Configuration of Training Device] First, the functional configuration of the training device 110 according to the seventh embodiment will be described. FIG. 9A is a diagram showing an example of the functional configuration of the training device according to the seventh embodiment. As shown in FIG. 9A, the training unit 111 includes an NN unit 310, a conversion unit 920, and a loss function unit 330.
[0120] Among these, since the NN unit 310 and the loss function unit 330 are the same as the NN unit 310 and the loss function unit 330 of the training unit 111 shown in FIG. 4A, the description thereof will be omitted here.
[0121] The conversion unit 920 includes a product-sum calculation unit and a random number unit. The random number unit determines each element of the x-dimensional array by a random number and sets it as the initial value in the product-sum calculation unit.
[0122] The product-sum calculation unit acquires, as the initial value, an array (x-dimensional array) having the same number of elements as the output dimension number (x) of the output data output from the final layer of the NN unit 310 from the random number unit. The conversion unit 920 performs a product-sum calculation on the output data output from the final layer of the NN unit 310 using the x-dimensional array.
[0123] Note that each element of the x-dimensional array set as the initial value in the product-sum calculation unit from the random number unit is updated based on the loss output from the loss function unit 330 during the training process.
[0124] Thus, while the training device 110 according to the seventh embodiment has the same configuration as the training device 110 according to the second embodiment described above, the x-dimensional array set as the initial value in the sum-of-products calculation unit is further configured to be determined by random numbers. According to this, according to the seventh embodiment, the same effects as those of the second embodiment can be enjoyed.
[0125] <Functional Configuration of Prediction Device> Next, the functional configuration of the prediction device 120 according to the seventh embodiment will be described. FIG. 9B is a diagram showing an example of the functional configuration of the prediction device according to the seventh embodiment. As shown in FIG. 9B, the prediction unit 121 includes an NN unit 311 and a conversion unit 921.
[0126] Among these, since the NN unit 311 is the same as the NN unit 311 of the prediction unit 121 shown in FIG. 4B, the description thereof will be omitted here.
[0127] The conversion unit 921 is the same as the conversion unit 920 of the training unit 111 in FIG. 9A and has a sum-of-products calculation unit. However, the x-dimensional array included in the sum-of-products calculation unit of the conversion unit 921 has an initial value set by random numbers and is updated by the training process.
[0128] The sum-of-products calculation unit of the conversion unit 921 performs a sum-of-products calculation on the output data output from the final layer of the NN unit 311 using the x-dimensional array updated by the training process, and outputs prediction data suitable for the processing result of the regression problem.
[0129] <Summary> As is clear from the above description, the training device 110 according to the seventh embodiment · The final layer of the NN unit 310 has a softmax function. · Has a conversion unit 920 that performs a sum-of-products calculation on the x-dimensional output data output from the final layer of the NN unit 310 using an x-dimensional array. · The x-dimensional array used when the conversion unit 920 performs the sum-of-products calculation has a value determined by random numbers as the initial value and is updated by the training process.
[0130] Accordingly, according to the training device 110 according to the seventh embodiment, the same effects as those of the second embodiment can be enjoyed.
[0131] Also, the prediction device 120 according to the seventh embodiment · The final layer of the NN unit 311 for which the training process has been performed by the training device 110 has a softmax function. · For the x-dimensional output data output from the final layer of the NN unit 311 for which the training process has been performed by the training device 110, using an x-dimensional array updated by performing the training process with a value determined by a random number as an initial value, it has a conversion unit 921 that performs a dot product operation.
[0132] Accordingly, according to the prediction device 120 according to the seventh embodiment, prediction data suitable for the processing result of the regression problem can be output.
[0133] [Eighth Embodiment] In the first embodiment, the dot product unit of the conversion unit 320 is configured to calculate a scalar value from the x-dimensional vector data by performing a dot product operation on the output data output from the final layer of the NN unit 310 using an x-dimensional array. On the other hand, in the eighth embodiment, a scalar value is output by selecting one value from the values of the number of output dimensions (x) included in the output data output from the final layer of the NN unit 310. Here, the "one value" refers to, for example, the value of the array in the dimension corresponding to the maximum element among the plurality of elements included in the output data (referred to as the value of the array corresponding to the maximum element). Hereinafter, the eighth embodiment will be described centering on the differences from the first embodiment.
[0134] <Functional Configuration of Training Device> First, the functional configuration of the training device 110 according to the eighth embodiment will be described. FIG. 10A is a diagram showing an example of the functional configuration of the training device according to the eighth embodiment. As shown in FIG. 10A, the training unit 111 includes an NN unit 310, a conversion unit 1020, and a loss function unit 330.
[0135] Among these, since the NN unit 310 and the loss function unit 330 are the same as those of the training unit 111 shown in FIG. 3A, the description thereof is omitted here.
[0136] The conversion unit 1020 has a selection unit. The selection unit outputs a scalar value by selecting one value (for example, the value of the array corresponding to the maximum element) from the values of the output dimension number (x) included in the output data output from the final layer of the NN unit 310.
[0137] As described above, the training device 110 according to the eighth embodiment has the same configuration as the training device 110 according to the first embodiment, and is configured such that the conversion unit outputs a scalar value without performing a dot product operation. Accordingly, according to the eighth embodiment, it is possible to enjoy the same effects as those of the first embodiment.
[0138] <Functional configuration of prediction device> Next, the functional configuration of the prediction device 120 according to the eighth embodiment will be described. FIG. 10B is a diagram showing an example of the functional configuration of the prediction device according to the eighth embodiment. As shown in FIG. 10B, the prediction unit 121 includes an NN unit 311 and a conversion unit 1020.
[0139] Among these, since the NN unit 311 is the same as the NN unit 311 of the prediction unit 121 shown in FIG. 3B, the description thereof is omitted here.
[0140] The conversion unit 1020 is the same as the conversion unit 1020 of the training unit 111 shown in FIG. 10A and has a selection unit. The selection unit of the conversion unit 1020 outputs prediction data suitable for the processing result of the regression problem by selecting one value (for example, the value of the array corresponding to the maximum element) from the values of the output dimension number (x) included in the output data output from the final layer of the NN unit 311.
[0141] <Summary> As is clear from the above description, the training device 110 according to the eighth embodiment · The final layer of the NN unit 310 has a softmax function. ·The conversion unit 1020 that outputs a scalar value is provided, which selects one value (for example, the value of the array corresponding to the maximum element) from the values of the number of output dimensions (x) included in the x-dimensional output data output from the final layer of the NN unit 310.
[0142] Accordingly, according to the training device 110 according to the eighth embodiment, it is possible to enjoy the same effects as those of the first embodiment.
[0143] Also, the prediction device 120 according to the eighth embodiment ·The final layer of the NN unit 311 that has been subjected to the training process by the training device 110 has a softmax function. ·The conversion unit 1020 that outputs a scalar value is provided, which selects one value (for example, the value of the array corresponding to the maximum element) from the values of the number of output dimensions (x) included in the x-dimensional output data output from the final layer of the NN unit 311 that has been subjected to the training process by the training device 110.
[0144] Accordingly, according to the prediction device 120 according to the eighth embodiment, prediction data suitable for the processing result of the regression problem can be output.
[0145] [Other Embodiments] In each of the above embodiments, the model to be trained is the NN unit, but the type of NN is not particularly limited, and for example, it may be a GNN (Graph Neural Network) or the like.
[0146] Also, in each of the above embodiments, as the functions of the conversion unit, the product-sum operation unit and the selection unit are given, but the functions of the conversion unit are not limited to these, and the conversion unit may have, for example, a normalization unit, a decoder unit, or the like.
[0147] In the first embodiment above, the method for determining the x-dimensional array was not mentioned. However, for the x-dimensional array, for example, the minimum value and the maximum value may be determined, and values obtained by equally dividing the range between the determined minimum value and maximum value into (x - 1) parts may be determined as the element values. Alternatively, values obtained by logarithmically dividing the range between the determined minimum value and maximum value into (x - 1) parts may be determined as the element values. Note that the minimum value and the maximum value may be determined according to the analysis target of the regression problem. For example, to predict the norm (in the range of 0 to 1) of the vector to be predicted, values obtained by dividing the range of -2 to 2 at intervals of 0.1 may be determined as the element values. Alternatively, to predict the variance, values obtained by dividing the range of 0.1 to 1.0 at intervals of 0.1 may be determined as the element values.
[0148] Also, in the sixth embodiment above, non-linear transformation functions or linear functions other than the softmax function were not described in detail. However, non-linear transformation functions or linear functions other than the softmax function are, for example, · Each dimension value included in the output x-dimensional vector is a positive value, · The sum of the output x-dimensional vector becomes a predetermined value (for example, "1"), and may be a function that transforms the input x-dimensional vector in such a manner.
[0149] For example, a non-linear transformation function other than the softmax function may include a relu function and a normalization unit that normalizes the magnitude of the x-dimensional vector output from the relu function to be "1".
[0150] Also, non-linear transformation functions or linear transformation functions other than the softmax function may include those obtained by modifying the softmax function and those obtained by modifying non-linear functions or linear transformation functions other than the softmax function. Specifically, it may include those obtained by taking the square root of the value calculated by the softmax function, those obtained by multiplying the value calculated by the softmax function by a constant, those obtained by dividing the value calculated by the relu function by the sum of the values calculated by the relu function, and the like.
[0151] In this specification (including the claims), when an expression such as "at least one (one) of a, b, and c" or "at least one (one) of a, b, or c" (including similar expressions) is used, it includes any one of a, b, c, a - b, a - c, b - c, or a - b - c. Also, for any element, multiple instances may be included, such as a - a, a - b - b, a - a - b - b - c - c, etc. Furthermore, it also includes adding other elements other than the enumerated elements (a, b, and c), such as having d as in a - b - c - d.
[0152] Also, in this specification (including the claims), when an expression such as "using data as input / based on data / according to data / in response to data" (including similar expressions) is used, unless otherwise specified, it includes cases where various data themselves are used as input, and cases where something obtained by performing some processing on various data (for example, data with noise added, normalized data, intermediate representations of various data, etc.) is used as input. Also, when it is described that a certain result is obtained "based on data / according to data / in response to data", it includes cases where the result is obtained based only on the said data, and also cases where the result is obtained under the influence of other data, factors, conditions, and / or states other than the said data. Also, when it is described that "data is output", unless otherwise specified, it includes cases where various data themselves are used as output, and cases where something obtained by performing some processing on various data (for example, data with noise added, normalized data, intermediate representations of various data, etc.) is output.
[0153] Also, in this specification (including the claims), when the terms "connected" and "coupled" are used, they are intended as non-limiting terms that include any of direct connection / coupling, indirect connection / coupling, electrical connection / coupling, communicative connection / coupling, operative connection / coupling, physical connection / coupling, etc. Although the terms should be appropriately interpreted according to the context in which they are used, connection / coupling forms that are not intentionally or naturally excluded should be interpreted non-limitingly as being included in the terms.
[0154] Also, in this specification (including the claims), when the expression "A configured to B" is used, it may include that the physical structure of element A has a configuration capable of performing operation B, and that a permanent or temporary setting / configuration of element A is set to actually perform operation B. For example, when element A is a general-purpose processor, it suffices that the processor has a hardware configuration capable of performing operation B and is set to actually perform operation B by a permanent or temporary program (instruction) setting. Also, when element A is a dedicated processor or a dedicated arithmetic circuit, etc., it suffices that the circuit structure of the processor is implemented to actually perform operation B regardless of whether control instructions and data are actually attached.
[0155] Also, in this specification (including the claims), when terms meaning containment or possession (such as "comprising / including" and "having") are used, they are intended as open-ended terms, including cases where the object contains or possesses things other than the object indicated by the object of the term. When the object of these terms meaning containment or possession does not specify a quantity or is an expression suggesting a singular form (an expression with "a" or "an" as the article), the expression should be construed as not being limited to a specific number.
[0156] Also, in this specification (including the claims), even if an expression such as "one or more" or "at least one" is used in one place and an expression that does not specify a quantity or suggests a singular form (an expression with "a" or "an" as the article) is used in another place, the latter expression is not intended to mean "one". Generally, an expression that does not specify a quantity or suggests a singular form (an expression with "a" or "an" as the article) should be construed as not necessarily being limited to a specific number.
[0157] Also, in this specification, if it is described that a specific effect (advantage / result) is obtained for a specific configuration of a certain embodiment, then, unless there are other reasons, it should be understood that the same effect can also be obtained for one or more other embodiments having the same configuration. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or states, etc., and the effect is not necessarily obtained by the configuration. The effect is only obtained by the configuration described in the embodiment when various factors, conditions, and / or states, etc., are satisfied, and the effect is not necessarily obtained in the invention according to the claim defining the configuration or a similar configuration.
[0158] In addition, in this specification (including the claims), when a plurality of hardware components perform a predetermined process, each hardware component may cooperate to perform the predetermined process, or some of the hardware components may perform all of the predetermined process. Also, some of the hardware components may perform a part of the predetermined process and another part of the hardware components may perform the remainder of the predetermined process. In this specification (including the claims), when expressions such as "one or more hardware components perform a first process and the one or more hardware components perform a second process" are used, the hardware components performing the first process and the hardware components performing the second process may be the same or different. That is, it is sufficient that the hardware components performing the first process and the hardware components performing the second process are included in the one or more hardware components. Note that the hardware may include an electronic circuit or a device including an electronic circuit, etc.
[0159] In addition, in this specification (including the claims), when a plurality of storage devices (memories) store data, each of the plurality of storage devices (memories) may store only a part of the data or may store all of the data.
[0160] As described above in detail with respect to the embodiments of the present disclosure, the present disclosure is not limited to the individual embodiments described above. Various additions, changes, replacements, and partial deletions, etc. are possible without departing from the conceptual ideas and spirit of the present invention derived from the content defined in the claims and their equivalents. For example, in all of the embodiments described above, the numerical values used in the description are shown as examples and are not limited thereto. Also, the order of each operation in the embodiments is shown as an example and is not limited thereto.
[0161] Note that in the disclosed technology, forms such as the following supplementary notes can be considered. (Supplementary Note 1) At least one processor A neural network model for processing regression problems, by inputting input data into a neural network model having a predetermined conversion function in a final layer, obtaining an output from the predetermined conversion function, using the output and a set of elements having the same number of elements as the number of dimensions of the output to obtain prediction data, updating model parameters of the neural network model using a loss calculated based on the prediction data and correct answer data. A training method for executing the process. (Appendix 2) The at least one processor obtains the prediction data by performing a dot product operation between the output and a set of elements having the same number of elements as the number of dimensions of the output. The training method according to Appendix 1, which executes the process. (Appendix 3) The at least one processor obtains the prediction data by selecting at least one value from the set of elements based on a plurality of values included in the output. The training method according to Appendix 1 or 2, which executes the process. (Appendix 4) The predetermined conversion function is a non-linear conversion function. The training method according to any one of Appendices 1 to 3. (Appendix 5) The predetermined conversion function is a softmax function. The training method according to Appendix 4. (Appendix 6) The values of the elements included in the set of elements having the same number of elements as the number of dimensions of the output are fixed values. The training method according to any one of Appendices 1 to 5. (Appendix 7) The values of the elements included in the set of elements having the same number of elements as the number of dimensions of the output are set based on the object of the regression problem. The training method according to Appendix 6. (Appendix 8) The at least one processor Updating the values of the elements included in the set of elements having the same number of elements as the dimension of the output using the loss. Executing a process, the training method according to any one of Appendices 1 to 7. (Appendix 9) The at least one processor Obtaining the prediction data based on the output from the neural network model for outputting a set of elements having the same number of elements as the dimension of the output and the output from the predetermined conversion function, Updating the model parameters of the neural network model for outputting a set of elements using the loss. Executing a process, the training method according to any one of Appendices 1 to 8. (Appendix 10) The at least one processor Inputting either the input data or the data output from an intermediate layer of the neural network model that processes the regression problem into the neural network model for outputting a set of elements. Executing a process, the training method according to Appendix 9. (Appendix 11) Generating the neural network model using the training method according to any one of Appendices 1 to 10. Model generation method. (Appendix 12) Comprising at least one processor, The at least one processor executes the training method according to any one of Appendices 1 to 10. Training device. (Appendix 13) Causing at least one processor to execute the training method according to any one of Appendices 1 to 10. Training program. (Appendix 14) The at least one processor A neural network model for processing regression problems, by inputting input data into a neural network model having a predetermined conversion function in the final layer, obtaining an output from the predetermined conversion function, A prediction method for executing a process of obtaining prediction data by using the output and a set of elements having the same number of elements as the number of dimensions of the output, The model parameters of the neural network model are updated using a loss calculated based on training prediction data and correct answer data. A prediction method. (Appendix 15) The at least one processor Obtaining the prediction data by performing a dot product operation between the output and a set of elements having the same number of elements as the number of dimensions of the output, Executing a process. The prediction method according to Appendix 14. (Appendix 16) The at least one processor Obtaining the prediction data by selecting at least one value from the set of elements based on a plurality of values included in the output, Executing a process. The prediction method according to Appendix 14 or 15. (Appendix 17) The predetermined conversion function is a non-linear conversion function. The prediction method according to any one of Appendices 14 to 16. (Appendix 18) The predetermined conversion function is a softmax function. The prediction method according to Appendix 17. (Appendix 19) The values of the elements included in the set of elements having the same number of elements as the number of dimensions of the output are fixed values. The prediction method according to any one of Appendices 14 to 18. (Appendix 20) The values of the elements included in the set of elements having the same number of elements as the number of dimensions of the output are set based on the target of the regression problem. The prediction method according to Appendix 19. (Appendix 21) The values of the elements included in the set of elements having the same number of elements as the dimension of the output are updated using the loss. The prediction method according to any one of Appendices 14 to 20. (Appendix 22) The at least one processor acquires the prediction data based on the output from the neural network model for outputting a set of elements having the same number of elements as the dimension of the output and the output from the predetermined conversion function. The neural network model for outputting a set of elements is updated using the loss. The prediction method according to any one of Appendices 14 to 21, which executes processing. (Appendix 23) The at least one processor inputs either the input data or the data output from an intermediate layer of the neural network model that processes the regression problem into the neural network model for outputting a set of elements. The prediction method according to Appendix 22, which executes processing. (Appendix 24) comprises at least one processor, The at least one processor executes the prediction method according to any one of Appendices 14 to 23. Prediction device. (Appendix 25) causes the at least one processor to execute the prediction method according to any one of Appendices 14 to 23. Prediction program.
Claims
1. At least one processor Inputs input data into a neural network model for processing regression problems, which has a predetermined transformation function in its final layer, and obtains an output from the predetermined transformation function; Obtains prediction data by using the output and a set of elements having the same number of elements as the number of dimensions of the output; Updates model parameters of the neural network model by using a loss calculated based on the prediction data and correct answer data. A training method for executing the process.
2. The at least one processor Obtains the prediction data by performing a dot product operation between the output and a set of elements having the same number of elements as the number of dimensions of the output. The training method according to claim 1, wherein the process is executed.
3. The at least one processor Obtains the prediction data by selecting at least one value from the set of elements based on a plurality of values included in the output. The training method according to claim 1, wherein the process is executed.
4. The predetermined transformation function is a non-linear transformation function. The training method according to claim 1.
5. The predetermined transformation function is a softmax function. The training method according to claim 4.
6. Values of elements included in a set of elements having the same number of elements as the number of dimensions of the output are fixed values. The training method according to claim 1.
7. Values of elements included in a set of elements having the same number of elements as the number of dimensions of the output are set based on an object of the regression problem. The training method according to claim 6.
8. The at least one processor Updates values of elements included in a set of elements having the same number of elements as the number of dimensions of the output by using the loss. The training method according to claim 1, wherein the process is executed.
9. The at least one processor Obtains the prediction data based on an output from a neural network model for outputting a set of elements that outputs a set of elements having the same number of elements as the number of dimensions of the output and the output from the predetermined transformation function, and Updates model parameters of the neural network model for outputting a set of elements by using the loss. The training method according to claim 1, wherein the process is executed.
10. The at least one processor Inputting either the input data or data output from an intermediate layer of the neural network model that processes the regression problem into the neural network model for outputting the element set. Executing the process, the training method according to claim 9.
11. Generating the neural network model by using the training method according to any one of claims 1 to 10. Model generation method.
12. Comprising at least one processor. The at least one processor executes the training method according to any one of claims 1 to 10. Training device.
13. Causing at least one processor to execute the training method according to any one of claims 1 to 10. Training program.
14. At least one processor Is a neural network model that processes a regression problem. By inputting input data into the neural network model having a predetermined conversion function in the final layer, an output from the predetermined conversion function is obtained. A prediction method for executing a process of obtaining prediction data by using the output and a set of elements having the same number of elements as the number of dimensions of the output. The model parameters of the neural network model are updated using a loss calculated based on prediction data and correct answer data for training. Prediction method.
15. The at least one processor Obtains the prediction data by performing a dot product operation between the output and a set of elements having the same number of elements as the number of dimensions of the output. Executing the process, the prediction method according to claim 14.
16. The at least one processor Obtains the prediction data by selecting at least one value from the set of elements based on a plurality of values included in the output. Executing the process, the prediction method according to claim 14.
17. The predetermined conversion function is a non-linear conversion function. The prediction method according to claim 14.
18. The predetermined conversion function is a softmax function. The prediction method according to claim 17.
19. The values of the elements included in the set of elements having the same number of elements as the number of dimensions of the output are fixed values. The prediction method according to claim 14.
20. The values of the elements included in the set of elements having the same number of elements as the number of dimensions of the output are set based on the object of the regression problem. The prediction method according to claim 19.
21. The value of an element included in a set of elements having the same number of elements as the number of dimensions of the output is updated using the loss. The prediction method according to claim 14. **Claim 22** The at least one processor acquires the prediction data based on an output from a neural network model for outputting a set of elements having the same number of elements as the number of dimensions of the output and the output from the predetermined conversion function; the neural network model for outputting the set of elements is updated using the loss. The prediction method according to claim 14, which executes processing. **Claim 23** The at least one processor inputs either the input data or data output from an intermediate layer of a neural network model that processes the regression problem into the neural network model for outputting the set of elements. The prediction method according to claim 22, which executes processing. **Claim 24** comprising at least one processor, wherein the at least one processor executes the prediction method according to any one of claims 14 to 23. Prediction device. **Claim 25** A prediction program for causing at least one processor to execute the prediction method according to any one of claims 14 to 23. Prediction program.
Citation Information
Patent Citations
Imaging apparatus
JP2017139646A