Information processing device, information processing method and program

By converting fully connected layers to convolutional layers and removing dimensionality conversion layers in deep learning models, the information processing apparatus achieves high-speed processing without accuracy loss, addressing the slow processing issue in models like the Transformer.

JP2025087518APending Publication Date: 2025-06-10PANASONIC AUTOMOTIVE SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023202225
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Deep learning models like the Transformer experience slow processing speeds due to the presence of dimensionality conversion and inverse-conversion layers, which are necessary for handling input information with different dimensionalities.

Method used

An information processing apparatus that converts fully connected layers in a deep learning model into convolutional layers and removes dimensionality conversion and inverse-conversion layers, resulting in a model with fewer layers that can process information more quickly without compromising accuracy.

Benefits of technology

This approach enables high-speed processing of deep learning models without sacrificing accuracy, making them more efficient for tasks like inference, especially in resource-constrained environments such as edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087518000001_ABST
    Figure 2025087518000001_ABST
Patent Text Reader

Abstract

To provide an information processing device that can generate a deep learning model allowing for high-speed processing without deterioration of accuracy.SOLUTION: An information processing device 10 comprises: an input unit 11 that acquires a first model m1 of deep learning; and a conversion unit 12 that selects a whole coupling layer L3 included in the first model m1, converts the selected whole coupling layer L3 to a convolution layer Lc, and deletes a dimensional conversion layer L2 and a dimensional inversion layer L4 included in the first model m1. The dimensional conversion layer L2 is configured to convert the number of dimensions of input information from a three-dimension to a two-dimension, and is the layer that outputs the input information represented by the two-dimension to the whole coupling layer L3. The dimensional inversion layer is configured to invert the number of dimensions of output information to be output from the whole coupling layer L3 from the two dimension to the three dimension.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus that performs processing related to machine learning and the like.

Background Art

[0002] In recent years, Transformer has been proposed (see, for example, Non-Patent Document 1). This Transformer is a deep learning model and is a network architecture in which an encoder and a decoder are connected only by Attention. Also, Transformer is used in ChatGPT (registered trademark).

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, there is a problem that processing using a deep learning model such as the Transformer of Non-Patent Document 1 above is slow.

[0005] Therefore, the present disclosure provides an information processing apparatus and the like that can generate a deep learning model capable of high-speed processing without deterioration in accuracy.

Means for Solving the Problems

[0006] An information processing apparatus according to an aspect of the present disclosure includes a processor and a memory connected to the processor. The processor uses the memory to acquire a first model of deep learning, select a fully connected layer included in the first model, convert the selected fully connected layer into a convolutional layer, delete a dimensional conversion layer and a dimensional inverse conversion layer included in the first model. The dimensional conversion layer is a layer that converts the number of dimensions of input information from three dimensions to two dimensions and outputs the input information represented in two dimensions to the fully connected layer. The dimensional inverse conversion layer is a layer that inversely converts the number of dimensions of output information output from the fully connected layer from two dimensions to three dimensions.

[0007] These general or specific aspects may be implemented in a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, a method, an integrated circuit, a computer program, and a recording medium. The recording medium may be a non-transitory recording medium.

Advantages of the Invention

[0008] The information processing apparatus of the present disclosure can generate a deep learning model capable of high-speed processing without deterioration in accuracy.

[0009] Furthermore, additional advantages and effects in an aspect of the present disclosure will be apparent from the specification and the drawings. Such advantages and / or effects are provided by some embodiments and the configurations described in the specification and the drawings, but not necessarily all configurations are required.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0011] (Findings on which the present disclosure is based) The present inventor has found that the following problems occur with respect to the Transformer of Non-Patent Document 1 described in the "Background Art" section.

[0012] The Transformer is applied to various tasks, as is also the case with ChatGPT. And the Transformer is evaluated to be fast and highly accurate compared to the LSTM (Long Short-Term Memory) network.

[0013] On the one hand, deep learning models such as Transformer include fully connected layers and other layers. In the fully connected layers and other layers, the dimensionality of the input information to be processed may be different. Therefore, in such deep learning models, processes for converting and inverse-converting the dimensionality of the input information are performed. When the dimensionality conversion and inverse-conversion are performed, the processing speed using the model becomes slow. Also, the more times the dimensionality conversion and inverse-conversion are performed, the slower the processing speed becomes. For example, Transformer includes a Multi-Head Attention layer, a Masked Multi-Head Attention layer, and a Feed Forward layer. And each of those layers further includes a fully connected layer and a dimensionality conversion layer and a dimensionality inverse-conversion layer for performing the above-described dimensionality conversion and inverse-conversion. Therefore, such a set of fully connected layers, dimensionality conversion layers, and dimensionality inverse-conversion layers is included in large numbers in Transformer. As a result, there is a problem that the processing using Transformer is slower than that of the LSTM network and also slower than that of the CNN (Convolutional Neural Network). Also, when Transformer is incorporated into an edge device such as a smartphone, the slowness of the processing becomes prominent.

[0014] To solve such problems, an information processing apparatus according to a first aspect of the present disclosure includes a processor and a memory connected to the processor. The processor uses the memory to acquire a first deep learning model, selects a fully connected layer included in the first model, converts the selected fully connected layer into a convolutional layer, deletes a dimensionality conversion layer and a dimensionality inverse-conversion layer included in the first model. The dimensionality conversion layer is a layer that converts the dimensionality of input information from three dimensions to two dimensions and outputs the input information represented in two dimensions to the fully connected layer. The dimensionality inverse-conversion layer is a layer that inverse-converts the dimensionality of output information output from the fully connected layer from two dimensions to three dimensions.

[0015] As a result, the fully-connected layer that performs processing on input information represented in two dimensions (i.e., two-dimensional processing) is converted into a convolutional layer that performs processing on input information represented in three dimensions (i.e., three-dimensional processing). Further, the dimension conversion layer and the inverse dimension conversion layer required for the two-dimensional processing of the fully-connected layer are deleted from the first model. By such conversion to the convolutional layer and deletion of the dimension conversion layer and the inverse dimension conversion layer, a second model equivalent to the first model can be generated. And since the number of layers included in the second model is less than that of the first model, the processing speed using the second model can be increased. As a result, a deep learning model capable of high-speed processing without deterioration in accuracy can be generated. Note that the processing can also be said to be inference. Therefore, a deep learning model capable of high-speed inference processing without deterioration in inference accuracy can be generated.

[0016] Also, in the information processing apparatus according to the second aspect, the processor may further perform machine learning on the second model generated by the conversion of the fully-connected layer to the convolutional layer and the deletion of the dimension conversion layer and the inverse dimension conversion layer. Note that the second aspect may be dependent on the first aspect.

[0017] Thereby, an inference result corresponding to the machine learning can be obtained. Also, an inference result equivalent to that of the pre-trained first model can be obtained at high speed.

[0018] Also, in the information processing apparatus according to the third aspect, the processor may further copy and insert the parameters of the first model on which machine learning has already been performed into the second model generated by the conversion of the fully-connected layer to the convolutional layer and the deletion of the dimension conversion layer and the inverse dimension conversion layer. Note that the third aspect may be dependent on the first aspect or the second aspect.

[0019] Thereby, even without performing machine learning on the second model, a model equivalent to the pre-trained second model can be generated, and the efficiency of model generation can be improved.

[0020] In the information processing apparatus according to the fourth aspect, the input information represented in three dimensions consists of a plurality of input values arranged according to the first axis, the second axis, and the channel axis, and the input information represented in two dimensions consists of a plurality of input values arranged according to the second axis and the channel axis. The fully connected layer has weight coefficients applied to the input information represented in two dimensions for each channel on the channel axis, and the convolutional layer has kernels applied to the input information represented in three dimensions for each channel on the channel axis. The size of the kernel for each channel is 1×1, and in the input information represented in three dimensions input to the convolutional layer, the number of input values arranged along the first axis may be 1. Note that the fourth aspect may be subordinate to any one of the first to third aspects. Also, the first axis, the second axis, and the channel axis are also respectively represented as the i-axis, the j-axis, and the Ci-axis.

[0021] Thereby, a deep learning model capable of high-speed processing without deterioration in accuracy can be appropriately generated.

[0022] (Embodiment) FIG. 1 is a block diagram showing an example of the configuration of the information processing apparatus in the present embodiment.

[0023] The information processing apparatus 10 in the present embodiment includes an input unit 11, a conversion unit 12, a storage unit 13, a learning unit 14, and a processing unit 15.

[0024] The input unit 11 acquires a first model m1 that is a deep learning model. Then, the input unit 11 outputs the first model m1 to the conversion unit 12. Note that the first model m1 may be a Transformer or another model. Also, the first model m1 may be a model in which machine learning has not been performed or a model in which machine learning has already been performed.

[0025] When the conversion unit 12 acquires the first model m1 from the input unit 11, it converts the first model m1 into the second model m2 and stores it in the storage unit 13.

[0026] The storage unit 13 is a recording medium for storing the second model m2 and the learned second model mk described later. For example, the storage unit 13 is a hard disk drive, RAM (Random Access Memory), ROM (Read Only Memory), or semiconductor memory. Note that such a storage unit 13 may be volatile or non-volatile.

[0027] The learning unit 14 generates a learned second model mk by performing machine learning on the second model m2 stored in the storage unit 13. Such a learned second model mk is stored in the storage unit 13.

[0028] The processing unit 15 acquires the information to be processed, and inputs the information to be processed to the learned second model mk stored in the storage unit 13, thereby acquiring the output information output from the learned second model mk. Then, the processing unit 15 outputs the output information to the outside of the information processing apparatus 10. Note that when the processing by the learned second model mk for the information to be processed is inference, the output information is the inference result for the information to be processed.

[0029] Note that the units including the input unit 11, the conversion unit 12, the learning unit 14, and the processing unit 15 may be configured as one or more processors such as a CPU (Central Processing Unit). Also, the unit may be realized, for example, by the processor executing a program stored in a memory. Note that the memory may be the storage unit 13. That is, the information processing apparatus 10 in the present embodiment may include a processor and a memory connected to the processor. In this case, the processor uses the memory to execute at least one of the processes of the input unit 11, the conversion unit 12, the learning unit 14, and the processing unit 15.

[0030] FIG. 2 is a diagram for explaining the processing operations of the information processing apparatus 10 according to the present embodiment for handling the first model m1 and the second model m2 in comparison with the processing operations of the apparatus of the comparative example.

[0031] The apparatus 90 of the comparative example includes an input unit 91, a learning unit 94, and a processing unit 95. When the input unit 91 acquires the first model m1, it outputs the first model m1 to the learning unit 94. The learning unit 94 generates a learned first model by performing machine learning on the first model m1. The processing unit 95 acquires output information output from the learned first model by inputting the information to be processed into the learned first model.

[0032] Here, in the apparatus 90 of the comparative example, machine learning is performed on the first model m1. As shown in FIG. 2, the first model m1 includes a front-stage three-dimensional processing layer L1, a dimensional conversion layer L2, a fully-connected layer L3, a dimensional inverse conversion layer L4, and a rear-stage three-dimensional processing layer L5. The front-stage three-dimensional processing layer L1 outputs output information represented in three dimensions by performing processing (i.e., three-dimensional processing) corresponding to the number of dimensions on the input information represented in three dimensions. For example, the front-stage three-dimensional processing layer L1 is a layer included in a Transformer and corresponds to, for example, batch normalization for normalizing a matrix. The dimensional conversion layer L2 treats the output information represented in three dimensions as input information, converts it into output information represented in two dimensions, and outputs it. The fully-connected layer L3 treats the output information represented in two dimensions as input information, and outputs output information represented in two dimensions by performing processing (i.e., two-dimensional processing) corresponding to the number of dimensions on the input information. The dimensional inverse conversion layer L4 treats the output information represented in two dimensions as input information, converts it into output information represented in three dimensions, and outputs it. The rear-stage three-dimensional processing layer L5 treats the output information represented in three dimensions as input information, and outputs output information represented in three dimensions by performing processing (i.e., three-dimensional processing) corresponding to the number of dimensions on the input information. For example, the rear-stage three-dimensional processing layer L5 is a layer included in a Transformer and corresponds to, for example, MatMul (i.e., Matrix Multiplication) for calculating a matrix product.

[0033] On the one hand, in the information processing apparatus 10 according to the present embodiment, the first model m1 acquired by the input unit 11 is converted into the second model m2 by the conversion unit 12. Then, in the learning unit 14, machine learning is performed on the second model m2. The conversion unit 12 converts the fully connected layer L3 included in the first model m1 into a convolutional layer Lc, and deletes the dimensionality conversion layer L2 and the dimensionality inverse conversion layer L4 before and after the fully connected layer L3. Thereby, the first model m1 is converted into the second model m2. The convolutional layer Lc of the second model m2 treats the output information represented in three dimensions output from the preceding three-dimensional processing layer L1 as input information, and outputs the output information represented in three dimensions by executing processing (i.e., three-dimensional processing) corresponding to the number of dimensions on the input information. That is, in the second model m2, since the fully connected layer L3 that performs two-dimensional processing is replaced with the convolutional layer Lc that performs three-dimensional processing, the dimensionality conversion layer L2 and the dimensionality inverse conversion layer L4 that perform conversion of the number of dimensions of the input information or the output information can be omitted.

[0034] FIG. 3 is a diagram for explaining the conversion of the number of dimensions.

[0035] The input information consists of a plurality of input values X arranged according to the i-axis, the j-axis, and the Ci-axis. That is, it can be said that the input information is represented in three dimensions of the i-axis, the j-axis, and the Ci-axis. In other words, each of the plurality of input values X is represented in three dimensions such as X(i, j, Ci) or X i,j,Ci as described above. Note that the output information also has the same configuration as the input information. The Ci-axis is also called the channel axis, and a plurality of input values X are arranged along the i-axis and the j-axis for each channel Ci on the channel axis. The preceding three-dimensional processing layer L1 and the subsequent three-dimensional processing layer L5 execute three-dimensional processing on the input information represented in such three dimensions.

[0036] Here, when there is only one input value X arranged along the i-axis, that is, when only i = 1, the input information represented in three dimensions of the i-axis, j-axis, and Ci-axis can be converted into input information represented in two dimensions of the j-axis and Ci-axis. That is, each of the input information, more specifically, each of the plurality of input values X included in the input information, is represented in two dimensions as X(j, Ci) or X j,Ci as shown. Conversely, the input information represented in two dimensions of the j-axis and Ci-axis can be converted into input information represented in three dimensions of the i-axis, j-axis, and Ci-axis. The dimension conversion layer L2 and the dimension inverse conversion layer L4 convert the number of dimensions of such input information. That is, the dimension conversion layer L2 converts the number of dimensions from three dimensions to two dimensions, and the dimension inverse conversion layer L4 converts the number of dimensions from two dimensions to three dimensions.

[0037] FIG. 4 is a diagram for explaining the processing operation of the fully connected layer L3.

[0038] The fully connected layer L3 has weight coefficients K Ci corresponding to each channel Ci. And the fully connected layer L3 performs two-dimensional processing on the input information represented in two dimensions using those weight coefficients K Ci . It should be noted that the input information can also be said to indicate a vector or feature amount for each channel Ci.

[0039] Specifically, the fully connected layer L3 multiplies and accumulates the input value X j,Ci at the position j on the j-axis included in each of the plurality of channels Ci by the weight coefficient K j,Ci of the channel Ci corresponding to the input value X Ci to output the output value Y j,C0 at that position j. That is, output information including the output value Y j,C0 is output. More specifically, the fully connected layer L3 outputs output information by performing the operation according to the following formula (1). In the following formula (1), p is an integer of 1 or more, indicating the number of channels, and Ci can take an integer from 1 to p.

[0040]

number

[0041] Fig. 5 is a diagram for explaining the processing operation of the convolutional layer Lc. Note that Fig. 5 shows the processing operation by the convolutional layer Lc for one channel Ci.

[0042] As shown in (a) of FIG. 5, the convolution layer Lc applies a filter to input information corresponding to a channel Ci, thereby outputting output information corresponding to the channel Ci. At this time, the convolution layer Lc outputs the output information while changing the position to which the filter is applied in the coordinate space of the input information. The input information and output information corresponding to the channel Ci are, for example, images. In the example of (a) of FIG. 5, the input information is made up of a 5×5 array of input values, and the output information is made up of a 3×3 array of output values. The filter is made up of a 3×3 array of filter coefficients, and is also called a kernel.

[0043] In a specific example, as shown in FIG. 5B, nine input values ​​X i+m,j+n For m = -1, 0, 1 and n = -1, 0, 1, there are 9 filter coefficients K m,n In the example of FIG. 5, i and j can take integers from 1 to 3. That is, the input value X i+m,j+n For filter coefficient K m,n The result is the output value Y 1,1 In the example in Figure 5, Y 1,1 =0 is calculated.

[0044] Next, the filter position is changed in the direction of the j-axis. As a result, for example, as shown in FIG. 5C, nine input values ​​X i+m,j+nFor nine filter coefficients K at nine positions indicated by m = -1, 0, 1 and n = -1, 0, 1 m,n are applied. That is, for the input value X i+m,j+n the filter coefficient K m,n is multiplied. As a result, the output value Y 1,2 at the position indicated by i = 1 and j = 2 is output. In the example of FIG. 5, Y 1,2 = 3 is output. In this way, the output value Y at each position is calculated while changing the position of the filter i,j .

[0045] That is, the convolutional layer Lc outputs output information, that is, the output value Y i,j by performing the operation according to FIG. 5(d) and the following formula (2).

[0046]

Equation

[0047] Therefore, when there are p channels Ci, the convolutional layer Lc outputs output information, that is, the output value Y i,j,C0 by performing the operation according to the following formula (3). In this way, the convolutional layer Lc outputs output information by performing three-dimensional processing on the input information represented by the three dimensions of the i-axis, j-axis, and channel axis.

[0048]

Equation

[0049] FIG. 6 is a diagram for explaining the processing operation of the convolutional layer Lc when the kernel size is 1×1. Note that the kernel size corresponds to the size of the filter in FIG. 5.

[0050] When the kernel size is 1×1, formula (3) is expressed as formula (4) below.

[0051]

Mathematics

[0052] Furthermore, when the number of input values X arranged along the i-axis is only one, that is, when i = 1 only, the processing operation of the convolutional layer Lc can omit the dimension of the i-axis as shown in FIG. 6. As a result, Equation (4) is expressed as the following Equation (5).

[0053]

Mathematics

[0054] The operation of the convolutional layer Lc expressed by such Equation (5) is equivalent to the operation of the fully connected layer L3 expressed by Equation (1) above.

[0055] Therefore, the conversion unit 12 of the information processing apparatus 10 in the present embodiment can convert the fully connected layer L3 included in the first model m1 into a convolutional layer Lc equivalent to the fully connected layer L3. Furthermore, the conversion unit 12 can delete the dimension conversion layer L2 and the dimension inverse conversion layer L4 that were inserted so that the two-dimensional processing by the fully connected layer L3 is performed in the three-dimensional processing by the first model m1.

[0056] FIG. 7 is a flowchart showing an example of the processing operation of the information processing apparatus 10 in the present embodiment.

[0057] First, the input unit 11 of the information processing apparatus 10 acquires the first model m1 (step S1). Then, the conversion unit 12 determines whether the acquired first model m1 includes the fully connected layer L3 (step S2). Here, when the conversion unit 12 determines that the fully connected layer L3 is included (Yes in step S2), it selects the fully connected layer L3 (step S3) and converts the fully connected layer L3 into a convolutional layer Lc (step S4). Furthermore, the conversion unit 12 deletes the dimension conversion layer L2 and the dimension inverse conversion layer L4 before and after the fully connected layer L3 (step S5).

[0058] After the process of step S5, the conversion unit 12 repeats the process from step S2 again. In step S2, when the conversion unit 12 determines that the fully connected layer L3 is not included in the first model m1 (No in step S2), the learning unit 14 executes the process of step S6. In step S6, the learning unit 14 generates a learned second model mk by performing machine learning on the second model m2 generated by the processes of steps S2 to S5 (step S6). Then, the processing unit 15 executes a process using the learned second model mk (step S7). That is, the processing unit 15 inputs the information to be processed into the learned second model mk, obtains the output information output from the learned second model mk, and outputs the output information to the outside of the information processing apparatus 10.

[0059] For example, when the first model m1 is a Transformer, the speed of the process using the learned second model mk, that is, the speed of the process in step S7, is about 20% faster than the speed of the process using the machine-learned first model m1.

[0060] In the example of FIG. 7, all the fully connected layers L3 included in the first model m1 are replaced with convolutional layers Lc, but the fully connected layer L3 may be left in the second model m2. That is, the second model m2 may be generated by performing the processes of steps S2 to S5 only on some of the fully connected layers L3 among all the fully connected layers L3 included in the first model m1.

[0061] As described above, the information processing apparatus 10 in the present embodiment acquires the first model m1 of deep learning, selects the fully connected layer L3 included in the first model m1, converts the selected fully connected layer L3 into a convolutional layer Lc, and deletes the dimension conversion layer L2 and the dimension inverse conversion layer L4 included in the first model m1. The dimension conversion layer L2 is a layer that converts the number of dimensions of the input information from three dimensions to two dimensions and outputs the input information represented in two dimensions to the fully connected layer L3. The dimension inverse conversion layer L4 is a layer that inversely converts the number of dimensions of the output information output from the fully connected layer L3 from two dimensions to three dimensions.

[0062] As a result, the fully-connected layer L3 that performs processing on input information represented in two dimensions (i.e., two-dimensional processing) is converted into a convolutional layer Lc that performs processing on input information represented in three dimensions (i.e., three-dimensional processing). Further, the dimension conversion layer L2 and the inverse dimension conversion layer L4 that are required for the two-dimensional processing of the fully-connected layer L3 are deleted from the first model m1. By such conversion to the convolutional layer Lc and deletion of the dimension conversion layer L2 and the inverse dimension conversion layer L4, a second model m2 equivalent to the first model m1 can be generated. And since the number of layers included in the second model m2 is smaller than that of the first model m1, the processing speed using the second model m2 can be increased. As a result, a deep learning model capable of high-speed processing without deterioration in accuracy can be generated. Note that the processing can also be said to be inference. Therefore, a deep learning model capable of high-speed inference processing without deterioration in inference accuracy can be generated.

[0063] Specifically, the input information represented in three dimensions consists of a plurality of input values arranged according to the first axis, the second axis, and the channel axis. Also, the input information represented in two dimensions consists of a plurality of input values arranged according to the second axis and the channel axis. The fully-connected layer L3 has a weight coefficient K Ci for each channel Ci on the channel axis. Also, the convolutional layer Lc has a kernel applied to the input information represented in three dimensions for each channel Ci on the channel axis. Here, the size of the kernel for each channel Ci is 1×1, and in the input information represented in three dimensions input to the convolutional layer Lc, the number of input values arranged along the first axis is 1. Note that the first axis, the second axis, and the channel axis respectively correspond to the above-described i axis, j axis, and Ci axis.

[0064] As a result, a deep learning model capable of high-speed processing without deterioration in accuracy can be appropriately generated.

[0065] Also, in the information processing apparatus 10 in the present embodiment, the processor realizes the function as the learning unit 14. That is, the processor performs machine learning on the second model m2 generated by converting the fully connected layer L3 into the convolutional layer Lc and deleting the dimensionality conversion layer L2 and the inverse dimensionality conversion layer L4.

[0066] Thereby, an inference result corresponding to the machine learning can be obtained. Also, an inference result equivalent to the inference result by the learned first model m1 can be obtained at high speed.

[0067] (Modification example) Here, in the above example, the learning unit 14 performs machine learning on the second model m2. On the other hand, in this modification example, if there is a first model m1 that has already been machine-learned, machine learning is not performed, and the parameters included in the learned first model m1 are copied to the second model m2.

[0068] FIG. 8 is a block diagram showing an example of the configuration of the information processing apparatus in the modification example of the present embodiment.

[0069] The information processing apparatus 10a in this modification example includes an input unit 11, a conversion unit 12, a storage unit 13, a parameter duplication unit 14a, and a processing unit 15. That is, the information processing apparatus 10a in this modification example has a configuration in which the learning unit 14 included in the information processing apparatus 10 shown in FIG. 1 is replaced by the parameter duplication unit 14a.

[0070] When the parameter duplication unit 14a acquires the already machine-learned first model m1, it duplicates the parameters included in the first model m1 and inserts them into the second model m2. If the first model m1 acquired by the input unit 11 is the already machine-learned first model m1, the parameter duplication unit 14a may duplicate the parameters included in the first model m1. The parameters to be duplicated are the parameters set by machine learning for the first model m1. Specifically, the parameters are the weight coefficients of the fully connected layer L3 included in the machine-learned first model m1.

[0071] FIG. 9 is a diagram for explaining the processing of the parameter replication unit 14a in a modification of the present embodiment.

[0072] For example, as shown in FIG. 9(a), after machine learning is performed on the first model m1, values corresponding to the machine learning are set as the weight coefficients for each channel Ci of the fully connected layer L3. For example, as the weight coefficients for channels Ci = 1 to p, values K 1 ’ to K p ’ are set respectively.

[0073] When the fully connected layer L3 is converted into the convolutional layer Lc by the conversion unit 12, the parameter replication unit 14a copies the weight coefficients K 1 ’ to K p ’ set by the machine learning of the fully connected layer L3 、 to the kernel of the convolutional layer Lc. That is, the weight coefficients K 1 ’ to K p ’ of the fully connected layer L3 are replicated and inserted into the second model m2 as the filter coefficients K 1 to K p of the convolutional layer Lc. Thereby, the convolutional layer Lc can be used as the convolutional layer Lc included in the learned second model mk. That is, the learned second model mk can be generated without performing machine learning on the second model m2.

[0074] As described above, in the information processing apparatus 10a in this modification, the processor replicates and inserts the parameters of the first model m1 on which machine learning has already been performed into the second model m2 generated by converting the fully connected layer L3 into the convolutional layer Lc and deleting the dimensionality conversion layer L2 and the inverse dimensionality conversion layer L4.

[0075] Thereby, the learned second model mk can be generated without performing machine learning on the second model m2, and the efficiency of model generation can be improved.

[0076] As described above, the information processing apparatus of the present disclosure has been described based on the above-described embodiments and their modifications. However, the present disclosure is not limited to those embodiments and modifications. As long as the gist of the present disclosure is not deviated from, various modifications conceived by those skilled in the art applied to the above-described embodiments and modifications may also be included in the present disclosure.

[0077] In the above-described embodiments, each component may be configured by dedicated hardware or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory. Here, the software that realizes the information processing apparatus 10 or 10a in each of the above-described embodiments is a computer program that causes a computer to execute each step of the flowchart shown in FIG. 7.

[0078] The following cases are also included in the present disclosure.

[0079] (1) Specifically, the at least one device described above is a computer system including a microprocessor, a ROM, a RAM, a hard disk unit, a display unit, a keyboard, a mouse, and the like. A computer program is stored in the RAM or the hard disk unit. By operating according to the computer program, the at least one device described above achieves its function. Here, the computer program is configured by combining a plurality of instruction codes indicating instructions to the computer in order to achieve a predetermined function.

[0080] (2) Some or all of the components constituting the at least one device described above may be configured from a single system LSI (Large Scale Integration). A system LSI is a super multifunctional LSI manufactured by integrating a plurality of components on a single chip, and specifically, is a computer system including a microprocessor, ROM, RAM, etc. A computer program is stored in the RAM. When the microprocessor operates according to the computer program, the system LSI achieves its function.

[0081] (3) Some or all of the components constituting the at least one device described above may be configured from an IC card or a single module detachable from the device. The IC card or module is a computer system composed of a microprocessor, ROM, RAM, etc. The IC card or module may include the above-mentioned super multifunctional LSI. When the microprocessor operates according to the computer program, the IC card or module achieves its function. This IC card or this module may have tamper resistance.

[0082] (4) The present disclosure may be the method described above. It may also be a computer program for realizing these methods by a computer, or a digital signal composed of a computer program.

[0083] Further, the present disclosure may be a computer program or a digital signal recorded on a computer-readable recording medium, such as a flexible disk, a hard disk, a CD (Compact Disc)-ROM, a DVD, a DVD-ROM, a DVD-RAM, a BD (Blu-ray (registered trademark) Disc), a semiconductor memory, etc. It may also be a digital signal recorded on these recording media.

[0084] Further, the present disclosure may also transmit a computer program or a digital signal via a telecommunication line, a wireless or wired communication line, a network represented by the Internet, data broadcasting, or the like.

[0085] Alternatively, it may be implemented by another independent computer system by recording and transferring the program or digital signal onto a recording medium or by transferring the program or digital signal via a network or the like.

Industrial Applicability

[0086] The information processing apparatus of the present disclosure can be applied to, for example, an apparatus or system that handles a deep learning model.

Explanation of Signs

[0087] 10, 10a Information processing apparatus 11 Input unit 12 Conversion unit 13 Storage unit 14 Learning unit 14a Parameter duplication unit 15 Processing unit 90 Apparatus 91 Input unit 94 Learning unit 95 Processing unit L1 Front-stage three-dimensional processing layer L2 Dimensional conversion layer L3 Fully-connected layer L4 Inverse dimensional conversion layer L5 Rear-stage three-dimensional processing layer Lc Convolution layer m1 First model m2 Second model mk Learned second model

Claims

1. A processor and, a memory connected to the processor, wherein the processor uses the memory to acquire a first deep learning model, select a fully connected layer included in the first model, convert the selected fully connected layer into a convolutional layer, delete a dimensionality conversion layer and a dimensionality inverse conversion layer included in the first model, wherein the dimensionality conversion layer is a layer that converts the dimensionality of input information from three dimensions to two dimensions and outputs the input information represented in two dimensions to the fully connected layer, and the dimensionality inverse conversion layer is a layer that inversely converts the dimensionality of output information output from the fully connected layer from two dimensions to three dimensions, an information processing apparatus.

2. The processor further performs machine learning on a second model generated by the conversion of the fully connected layer into the convolutional layer and the deletion of the dimensionality conversion layer and the dimensionality inverse conversion layer, the information processing apparatus according to claim 1.

3. The processor further copies and inserts parameters of the first model on which machine learning has already been performed with respect to the second model generated by the conversion of the fully connected layer into the convolutional layer and the deletion of the dimensionality conversion layer and the dimensionality inverse conversion layer, the information processing apparatus according to claim 1.

4. The input information represented in three dimensions consists of a plurality of input values arranged according to a first axis, a second axis, and a channel axis, the input information represented in two dimensions consists of a plurality of input values arranged according to the second axis and the channel axis, the fully connected layer has weight coefficients applied to the input information represented in two dimensions for each channel on the channel axis, the convolutional layer has a kernel applied to the input information represented in three dimensions for each channel on the channel axis, the size of the kernel for each channel is 1×1, in the input information represented in three dimensions input to the convolutional layer, the number of input values arranged along the first axis is 1, the information processing apparatus according to any one of claims 1 to 3.

5. An information processing method performed by a computer, comprising: acquiring a first deep learning model, selecting a fully connected layer included in the first model, converting the selected fully connected layer into a convolutional layer, deleting a dimensionality conversion layer and a dimensionality inverse conversion layer included in the first model, wherein the dimensionality conversion layer A layer that converts the dimensionality of input information from three dimensions to two dimensions and outputs the input information represented in two dimensions to the fully connected layer. The dimensionality inverse conversion layer is A layer that inverse-converts the dimensionality of the output information output from the fully connected layer from two dimensions to three dimensions. An information processing method. **Claim 6** Obtain a first deep learning model, Select the fully connected layer included in the first model, Convert the selected fully connected layer into a convolutional layer, Delete the dimensionality conversion layer and the dimensionality inverse conversion layer included in the first model. Cause a computer to execute the above, The dimensionality conversion layer is A layer that converts the dimensionality of input information from three dimensions to two dimensions and outputs the input information represented in two dimensions to the fully connected layer. The dimensionality inverse conversion layer is A layer that inverse-converts the dimensionality of the output information output from the fully connected layer from two dimensions to three dimensions. A program.