Inference device and inference method

The inference device addresses the computational inefficiencies in predicting and converting image resolutions by using a first neural network for prediction and a second for inference, sharing activation values to reduce computational load and enhance speed.

JP2025087005APending Publication Date: 2025-06-10NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023201348
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing techniques for predicting the minimum resolution for image inference require significant computational resources, and there is a need to reduce the inference time for one-dimensional data such as signal data.

Method used

An inference device comprising a resolution prediction module using a first neural network model to predict the minimum resolution for accurate inference, a resolution conversion module to adjust the image resolution, and an inference module using a second neural network model with partially shared activation values from the first model to infer the label efficiently.

Benefits of technology

This approach reduces the computational burden during resolution prediction and inference, enabling faster label inference for both images and one-dimensional data while maintaining accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087005000001_ABST
    Figure 2025087005000001_ABST
Patent Text Reader

Abstract

To provide an inference device capable of reducing an amount of calculation.SOLUTION: Resolution prediction means receives data of a certain resolution as input, and uses a first model including a plurality of layers to predict, among a plurality of resolution candidates, a minimum resolution at which a label of the data can be inferred with a predetermined accuracy. Resolution conversion means converts a resolution of the data to the predicted resolution. Inference means infers, using the resolution-converted data as input and a second model including a plurality of layers, a label of the data. At this time, the inference means infers, using a part of an activation value output from a predetermined layer of the first model, the label of the data.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an inference device, an inference method, and an inference program.

Background Art

[0002] When inferring an object shown in an image using a neural network model, the inference time can be shortened by reducing the resolution (size of the image) of the image. On the other hand, when the resolution of the image is reduced, the inference accuracy may decrease.

[0003] Non-Patent Document 1 describes a technique for predicting the minimum resolution that does not reduce the inference accuracy for a given image, converting the resolution of the given image to the predicted resolution, and inferring the object shown in the image with the converted resolution.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the technique described in Non-Patent Document 1, the amount of calculation required to predict the resolution for a given image is necessary. Therefore, it may take time to obtain the inference result.

[0006] Also, even when the given data is one-dimensional data such as signal data, it is preferable that the time required for inferring the label of the data can be reduced.

[0007] Therefore, an object of the present invention is to provide an inference device, an inference method, and an inference program that can reduce the amount of calculation when predicting the resolution for given data, converting the resolution of the given data to the predicted resolution, and inferring the label of the data with the converted resolution.

Means for Solving the Problems

[0008] The inference device according to the present disclosure includes a resolution prediction means for predicting, as an input, data with a certain resolution and using a first model including a plurality of layers, the minimum resolution among a plurality of resolution candidates that can infer the label of the data with a predetermined accuracy; a resolution conversion means for converting the resolution of the data to the predicted resolution; and an inference means for inferring the label of the data using a second model including a plurality of layers with the data having the converted resolution as an input, wherein the inference means infers the label of the data using a part of the activation values output from a predetermined layer of the first model.

[0009] The inference method according to the present disclosure is such that a computer predicts, as an input, data with a certain resolution and uses a first model including a plurality of layers, the minimum resolution among a plurality of resolution candidates that can infer the label of the data with a predetermined accuracy, converts the resolution of the data to the predicted resolution, infers the label of the data using a second model including a plurality of layers with the data having the converted resolution as an input, and uses a part of the activation values output from a predetermined layer of the first model when inferring the label of the data.

[0010] The inference program according to the present disclosure causes a computer to perform resolution prediction processing for predicting, using a first model including a plurality of layers, the minimum resolution among a plurality of resolution candidates that can infer the label of data with a predetermined accuracy, with data of a certain resolution as input; resolution conversion processing for converting the resolution of the data to the predicted resolution; and inference processing for inferring the label of the data, using a second model including a plurality of layers, with the data whose resolution has been converted as input, and causes the label of the data to be inferred using a part of the activation values output from a predetermined layer of the first model in the inference processing.

Advantages of the Invention

[0011] According to the present disclosure, it is possible to reduce the amount of calculation in the case of predicting the resolution for given data, converting the resolution of the given data to the predicted resolution, and inferring the label of the data whose resolution has been converted.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Mode for Carrying Out the Invention

[0013] Hereinafter, embodiments according to the present disclosure will be described with reference to the drawings. In each of the following embodiments, the case where an image is input as input data to the inference device will be described as an example. However, the input data is not limited to an image. One-dimensional data such as signal data may be input to the inference device. The inference device infers the label of the input data.

[0014] Embodiment 1. FIG. 1 is a block diagram showing a configuration example of an inference device according to the present disclosure. The inference device according to the present disclosure includes an input unit 1, a resolution prediction unit 2, a resolution conversion unit 3, an inference unit 4, an activation value size conversion unit 5, an output unit 6, and a learning unit 7.

[0015] The input unit 1 is an input interface to which an image with a certain resolution is input. For example, the input unit 1 is realized by a data reading device that reads an image with a certain resolution recorded on a data recording medium. However, the input unit 1 is not limited to such a data reading device.

[0016] An image with a certain resolution is input to the resolution prediction unit 2 via the input unit 1. Using a first model including a plurality of layers with the image of a certain resolution as input, the resolution prediction unit 2 predicts the minimum resolution among a plurality of resolution candidates at which an object depicted in the image can be inferred with a predetermined accuracy. The number of resolution candidates is, for example, three, but is not limited to three. Also, each resolution serving as a candidate is a resolution equal to or lower than the resolution of the input image.

[0017] The first model is held in the resolution prediction unit 2. The first model is, for example, a neural network model.

[0018] The resolution conversion unit 3 converts the resolution of the image input via the input unit 1 into the resolution predicted by the resolution prediction unit 2.

[0019] Using the image with the converted resolution as input, the inference unit 4 infers an object depicted in the image using a second model including a plurality of layers. At this time, the inference unit 4 infers an object depicted in the image using a part of the activation values output from a predetermined layer of the first model. Note that inferring an object depicted in the image can be said to be inferring the label of the image.

[0020] Data output from each layer of the model is called an activation value. The activation value has a channel, height (hereinafter referred to as h), and width (hereinafter referred to as w). In the drawings, the channel is described as ch. Also, when the input data is one-dimensional data, either h or w is a fixed value of "1".

[0021] The second model is held in the inference unit 4. The second model is, for example, a neural network model.

[0022] The layers from the first layer to a predetermined layer of the second model correspond to the layers from the first layer to a predetermined layer of the first model.

[0023] When learning the second model and the first model, first, the learning unit 7 tentatively learns the second model. Then, the learning unit 7 copies the weights from the first layer to the predetermined layer in the second model to the first layer to the predetermined layer in the first model. Among the weights of the copy source, the weights corresponding to a part of the activation values output from the predetermined layer of the first model described above are removed from the second model. Hereinafter, a case where a part of the activation values output from the predetermined layer of the first model is half of the activation values output from the predetermined layer of the first model will be described as an example. However, it is not limited to this example. In this case, among the weights of the copy source, half of the weights of each layer from the first layer to the predetermined layer are removed. Also, the ratio corresponding to a part of the activation values output from the predetermined layer of the first model is predetermined.

[0024] FIG. 2 is a schematic diagram showing an example of the second model tentatively learned and the first model in which the weights from the first layer to the predetermined layer in the second model are copied. Layers 51a and 51b are the first layers, and layers 54a and 54b are the predetermined layers. Layers 51a to 54a and layers 51b to 54b correspond to each other.

[0025] FIG. 3 is a schematic diagram showing an example of the second model in a state where half of the weights of the copy source are removed. Since half of the weights of the copy source are removed, the number of channels of layers 51a, 52a, and 54a shown in FIG. 3 is half of the number of channels of layers 51a, 52a, and 54a shown in FIG. 2, respectively.

[0026] Also, in the first model, the weights of the layers after the predetermined layer 54b are randomly determined. With this state as the initial state of the first model, the learning unit 7 learns the first model and the second model. These points are the same in the second embodiment described later. This learning (learning by the learning unit 7) will be described later.

[0027] Hereinafter, it will be described on the assumption that the first model and the second model have been learned. For the sake of convenience, even for the first model after learning and the second model after learning, layers will be described with reference numerals 51a to 54a and reference numerals 51b to 54b added.

[0028] Note that the end of the second model is one or more fully connected layers, and a GlobalAveragePooling layer is provided in front of the group of fully connected layers. As a result, even if the resolution of the image input to the second model is not constant, the inference unit 4 can derive an inference result for an image of each resolution.

[0029] Hereinafter, an operation in which the inference unit 4 infers an object shown in the image using a part (half) of the activation values output from a predetermined layer of the first model will be described.

[0030] FIG. 4 is a schematic diagram showing an example of activation values obtained when the resolution prediction unit 2 predicts the resolution. Hereinafter, in the drawings, layers are shown by solid lines, and inputs and activation values are shown by broken lines (see FIG. 4).

[0031] An activation value 61 is output from a predetermined layer 54b. It is assumed that a part (half) of this activation value 61 is used by the inference unit 4. The number of channels of the activation value 62 (half of the activation value 61) used by the inference unit 4 is half of the number of channels of the activation value 61 (see FIG. 4).

[0032] The inference unit 4 combines, in the channel direction, the activation value output from a predetermined layer 54a in the second model corresponding to the predetermined layer 54b with the above activation value 62 (see FIG. 4).

[0033] However, in order to combine activation values in the channel direction, it is necessary for h and w to match.

[0034] Also, according to the resolution of the image input to the inference unit 4, the size of the activation value output from each layer of the second model changes. In this example, assume that when the size (resolution) of the image is the largest, h = 256 and w = 256. Also, assume that when the size (resolution) of the image is the smallest, h = 64 and w = 64.

[0035] FIG. 5 is a schematic diagram showing an example of the size conversion of the activation value 62 and the combination of the activation value 62 after the size conversion and the activation value output from a predetermined layer 54a when the largest image is input to the inference unit 4 and when the smallest image is input to the inference unit 4.

[0036] For the activation value output from the predetermined layer 54a when the largest image is input to the inference unit 4, h = 128 and w = 128 (see the upper part of FIG. 5). In this case, the activation value size conversion unit 5 converts the size of the activation value 62 (half of the activation value 61) to the original size (h = 128, w = 128) (see the upper part of FIG. 5). In other words, the activation value size conversion unit 5 does not convert the size of the activation value 62.

[0037] Also, for the activation value output from the predetermined layer 54a when the smallest image is input to the inference unit 4, h = 32 and w = 32 (see the lower part of FIG. 5). In this case, the activation value size conversion unit 5 converts the size of the activation value 62 (half of the activation value 61) to h = 32 and w = 32 (see the lower part of FIG. 5).

[0038] In this way, the activation value size conversion unit 5 converts the size of the activation value 62 (half of the activation value 61) according to the resolution predicted by the resolution prediction unit 2 (in other words, the resolution after conversion by the resolution conversion unit 3). However, the resolution conversion unit 3 does not convert the number of channels of the activation value 62.

[0039] Since the activation value size conversion unit 5 converts the size of the activation value 62, the activation value 62 with the converted size can be combined in the channel direction with the activation value output from the predetermined layer 54a in the second model corresponding to the predetermined layer 54b. The inference unit 4 combines the activation value 62 with the converted size and the activation value output from the predetermined device 54a in the channel direction, and inputs the combination result 63 (see FIG. 5) into the next layer in the second model (the layer next to the predetermined layer 54a). Then, the inference unit 4 repeatedly inputs the output of the layer into the next layer to obtain an inference result.

[0040] The output unit 6 outputs the inference result obtained by the inference unit 4. For example, the output unit 6 causes the inference result to be displayed on a display device (not shown in FIG. 1) of the inference device. However, the output unit 6 may output the inference result in an output mode other than display.

[0041] The resolution prediction unit 2, the resolution conversion unit 3, the activation value size conversion unit 5, the inference unit 4, the output unit 6, and the learning unit 7 are realized, for example, by a CPU (Central Processing Unit) of a computer that operates according to an inference program. In this case, the CPU reads the inference program from a program recording medium such as a program storage device of the computer, and may operate as the resolution prediction unit 2, the resolution conversion unit 3, the activation value size conversion unit 5, the inference unit 4, the output unit 6, and the learning unit 7 according to the inference program.

[0042] Next, the process progress will be described. FIG. 6 is a flowchart showing an example of the process progress of the inference device according to the present disclosure. Regarding matters already described, the description will be omitted as appropriate.

[0043] First, the resolution prediction unit 2 uses the first model with an image of a certain resolution as input to predict the minimum resolution among a plurality of resolution candidates that can infer the object shown in the image with a predetermined accuracy (step S1).

[0044] Next, the resolution conversion unit 3 converts the resolution of the image to the predicted resolution (step S2).

[0045] Further, the activation value size conversion unit 5 converts the size of half of the activation value 61 (activation value 62) output from a predetermined layer 54b of the first model according to the predicted resolution (step S3).

[0046] Then, the inference unit 4 combines the activation value output from a predetermined layer 54a of the second model with the activation value 62 whose size has been converted, and inputs the combination result to the next layer. Then, the inference unit 4 repeatedly inputs the output of the layer to the next layer to obtain an inference result (step S4).

[0047] Then, the output unit 6 causes the inference result to be displayed on a display device (not shown in FIG. 1) (step S5).

[0048] According to the present embodiment, in the second model, some of the weights copied to the first model are removed, and the second model is learned. Therefore, the amount of calculation when the inference unit 4 makes an inference using the second model can be reduced by the amount of weights removed. And even if the weights are removed, the inference unit 4 makes an inference using the activation value 62 corresponding to the calculation result by the weights (half of the activation value 61 output from the predetermined layer 54b of the first model). That is, the activation value 62 and the activation value output from the predetermined layer 54a of the second model are combined in the channel direction, and subsequent inference operations are performed. Therefore, an appropriate inference result can be obtained while reducing the amount of calculation.

[0049] Embodiment 2. FIG. 7 is a block diagram showing a configuration example of an inference device according to the present disclosure. The inference device according to the present disclosure includes an input unit 1, a resolution prediction unit 2, a resolution conversion unit 3, an activation value size conversion unit 81, a layer identification unit 82, an inference unit 83, an output unit 6, and a learning unit 7.

[0050] The input unit 1, resolution prediction unit 2, resolution conversion unit 3, output unit 6, and learning unit 7 are the same as those in the first embodiment, and detailed descriptions thereof are omitted. Note that the learning by the learning unit 7 will be described later.

[0051] The activation value size conversion unit 81 converts the size of the activation value 62 (see FIG. 4) corresponding to a part (half) of the activation value 61 output from a predetermined layer 54b of the first model to a certain size according to the smallest resolution among a plurality of resolution candidates in the resolution prediction unit 2.

[0052] Specifically, the activation value size conversion unit 81 converts the size of the activation value 62 so that it has the same h and w as the activation value output from the predetermined layer 54a when the image with the smallest resolution among a plurality of resolution candidates is input to the second model. The activation value size conversion unit 81 does not convert the number of channels of the activation value 62. In this example, the activation value size conversion unit 81 converts the size of the activation value 62 from [number of channels 64, h = 128, w = 128] to [number of channels 64, h = 32, w = 32].

[0053] The layer identification unit 82 identifies a layer in the second model that outputs an activation value that can be combined with the activation value 62 after size conversion in the channel direction based on the resolution predicted by the resolution prediction unit 2 (in other words, the resolution after conversion by the resolution conversion unit 3).

[0054] Specifically, the layer identification unit 82 identifies a layer that outputs an activation value having the same h and w as the activation value 62 after size conversion when the image with the resolution predicted by the resolution prediction unit 2 is input to the second model.

[0055] The inference unit 83 infers an object shown in the image using the activation value 62 corresponding to a part (half in this example) of the activation value 61 output from a predetermined layer 54b of the first model.

[0056] In the second embodiment, the inference unit 83 combines, in the channel direction, the activation values output from the layer specified by the layer specifying unit 82 and the activation values 62 after size conversion, and inputs the combination result to the next layer in the second model.

[0057] FIG. 8 is a schematic diagram showing examples of combination when the largest image is input to the inference unit 83 and when the smallest image is input to the inference unit 83.

[0058] When the smallest image is input to the inference unit 83, the layer specifying unit 82 specifies a predetermined layer 54a. Accordingly, the inference unit 83 combines, in the channel direction, the activation values output from the predetermined layer 54a and the activation values 62 after size conversion, and inputs the combination result 63 to the next layer 55a (see the lower part of FIG. 8). Then, the inference unit 83 repeatedly inputs the output of the layer to the next layer to obtain an inference result. In FIG. 8, layers 55a to 58a after the layer 54a are also shown. The number of channels of the layer 55a and the layer 57a is 256 (see the lower part of FIG. 8).

[0059] When the largest image is input to the inference unit 83, the layer specifying unit 82 specifies a layer 58a that outputs activation values 65 where h = 32 and w = 32. Accordingly, the inference unit 83 combines, in the channel direction, the activation values 65 output from the layer 58a and the activation values 62 after size conversion, and inputs the combination result 66 to the next layer (the layer next to the layer 58a) (see the upper part of FIG. 8). Then, the inference unit 83 repeatedly inputs the output of the layer to the next layer to obtain an inference result.

[0060] As described above, the number of channels in layer 55a and layer 57a is 256 (see the lower part of FIG. 8). In the upper part of FIG. 8, showing the number of channels in layer 55a and layer 57a as 192 means that 192 out of the 256 channels are used. Therefore, when the largest image is input to the inference unit 83, no operation for 256 - 192 = 64 channels is performed in layer 55a and layer 57a. This unperformed operation is compensated by the activation value 62 after size conversion that is combined with the activation value 65. In this example, the amount of calculation is also reduced by not performing the operation for 64 channels.

[0061] The resolution prediction unit 2, the resolution conversion unit 3, the activation value size conversion unit 81, the layer identification unit 82, the inference unit 83, the output unit 6, and the learning unit 7 are realized, for example, by the CPU (Central Processing Unit) of a computer that operates according to an inference program. In this case, the CPU reads the inference program from a program recording medium such as a program storage device of the computer and operates as the resolution prediction unit 2, the resolution conversion unit 3, the activation value size conversion unit 81, the layer identification unit 82, the inference unit 83, the output unit 6, and the learning unit 7 according to the inference program.

[0062] Next, the process flow will be described. FIG. 9 is a flowchart showing an example of the process flow of the inference device according to the present disclosure. For matters already described, the description will be omitted as appropriate.

[0063] First, the inference device executes steps S1 and S2. Since steps S1 and S2 are the same processes as steps S1 and S2 shown in FIG. 6, the description will be omitted.

[0064] Next to step S2, the activation value size conversion unit 81 converts the size of the activation value 62 to a certain size according to the smallest resolution among a plurality of resolution candidates (step S11).

[0065] Next, the layer identification unit 82 identifies the layer in the second model that outputs an activation value that can be combined with the activation value 62 after size conversion (step S12).

[0066] Then, the inference unit 83 combines the activation value output from the layer identified in step S12 and the activation value 62 after size conversion in the channel direction, and inputs the combination result to the next layer. Then, the inference unit 83 repeats inputting the output of the layer to the next layer to obtain an inference result (step S13).

[0067] Then, the output unit 6 causes the inference result to be displayed on a display device (not shown in FIG. 7) (step S5). Step S5 is the same process as step S5 shown in FIG. 6.

[0068] Also in the second embodiment, as in the first embodiment, an appropriate inference result can be obtained while reducing the amount of calculation. Further, as described above, even if it exists as a channel, there may be a case where the calculation for that channel is not performed. In this case, the amount of calculation can be further reduced.

[0069] Next, in the first embodiment and the second embodiment, the operations for obtaining the first model and the second model by learning will be described.

[0070] FIG. 10 is a flowchart showing an example of the processing progress of the operations for obtaining the first model and the second model by learning.

[0071] First, the learning unit 7 learns the second model (step S21). The second model obtained in step S21 is a provisional second model.

[0072] Instead of step S21, a pre-trained second model may be prepared.

[0073] Next, the learning unit 7 copies the weights from the first layer to a predetermined layer of the second model obtained in step S21 to the first model, and randomly determines the weights of the layers after the predetermined layer of the first model (step S22). The weights of the first model determined in step S22 are the initial values of the weights of the first model.

[0074] Next, the resolution prediction unit 2 obtains the likelihood of each predicted resolution for the input image using the first model. The inference unit 4 (inference unit 83) obtains the likelihood of the inference result (label) obtained by inference for each resolution that can be predicted using the second model (step S23). The resolution prediction unit 2 sends each obtained likelihood to the learning unit 7. The inference unit 4 (inference unit 83) sends each obtained likelihood to the learning unit 7.

[0075] Then, the learning unit 7 calculates the loss function for the second model and obtains the loss. Also, the learning unit 7 calculates the loss function for the first model and obtains the loss (step S24).

[0076] When obtaining the loss function for the second model, the learning unit 7 uses the label of the input image as the teacher data. Also, when obtaining the loss function for the first model, the learning unit 7 uses the "value calculated using the loss of the inference unit 4 (inference unit 83) and the calculation amount of the inference unit 4 (inference unit 83)" at each resolution that can be predicted as the teacher data. For example, "the loss of the inference unit 4 (inference unit 83) + λ * the calculation amount of the inference unit 4 (inference unit 83)" is used as the teacher data. λ is a predetermined constant. When making the calculation amount smaller, the value of λ can be set to a large value. When maintaining the inference accuracy, the value of λ can be set to a small value.

[0077] The learning unit 7 sends the loss obtained for the second model to the inference unit 4 (inference unit 83) and updates the weights of the second model. Also, the learning unit 7 sends the loss obtained for the first model to the resolution prediction unit 2 and updates the weights of the first model (step S25).

[0078] Next, the learning unit 7 determines whether or not the number of executions of steps S23 to S25 has reached a predetermined number (step S26). If the number of executions of steps S23 to S25 has not reached the predetermined number (No in step S26), the operations after step S23 are repeated.

[0079] If the number of executions of steps S23 to S25 has reached the predetermined number (Yes in step S26), the learning unit 7 determines the first model and the second model at that time as the first model and the second model, respectively, and ends the process.

[0080] As described above, the first model is a learned model with the initial state being a state where the weights from the first layer to a predetermined layer are the same as the weights from the first layer to the predetermined layer of the second model.

[0081] FIG. 11 is a schematic block diagram showing a configuration example of a computer according to the inference device. The computer 2000 includes, for example, a CPU 2001, a main storage device 2002, an auxiliary storage device 2003, an interface 2004, a data reading device 2005, and a display device 2006. However, the configuration of the computer 2000 is not limited to the example shown in FIG. 11.

[0082] The inference device according to the present disclosure is realized, for example, by the computer 2000. The operations of the inference device are stored in the auxiliary storage device 2003 in the form of a program (inference program). The CPU 2001 reads the program from the auxiliary storage device 2003, expands the program in the main storage device 2002, and executes the processes described in the above embodiments according to the program.

[0083] The auxiliary storage device 2003 is an example of a non-transitory tangible medium. Other examples of non-transitory tangible media include magnetic disks, magneto-optical disks, CD-ROMs (Compact Disk Read Only Memories), DVD-ROMs (Digital Versatile Disk Read Only Memories), semiconductor memories, etc. connected via the interface 2004.

[0084] Also, the inference device may be realized by a plurality of communicably connected computers.

[0085] Next, an overview of the inference device according to the present disclosure will be described. FIG. 12 is a block diagram showing an overview of the inference device according to the present disclosure. The inference device includes a resolution prediction means 72, a resolution conversion means 73, and an inference means 74.

[0086] The resolution prediction means 72 (for example, the resolution prediction unit 2) takes data of a certain resolution as input and uses a first model including a plurality of layers to predict the minimum resolution among a plurality of resolution candidates that can infer the label of the data with a predetermined accuracy.

[0087] The resolution conversion means 73 (for example, the resolution conversion unit 3) converts the resolution of the data to the predicted resolution.

[0088] The inference means 74 (for example, the inference unit 4, the inference unit 83) takes the data with the converted resolution as input and uses a second model including a plurality of layers to infer the label of the data. At this time, the inference means 74 infers the label of the data using a part of the activation values output from a predetermined layer of the first model.

[0089] With such a configuration, for given data, it is possible to reduce the amount of calculation in the case of predicting the resolution, converting the resolution of the given data to the predicted resolution, and inferring the label of the data with the converted resolution.

[0090] Each of the above-described embodiments of the present invention can also be described as follows in the appended claims, but is not limited thereto.

[0091] (Appended Claim 1) Resolution prediction means for predicting, using a first model including a plurality of layers, the minimum resolution among a plurality of resolution candidates that can infer the label of the data with a predetermined accuracy, with data of a certain resolution as input; Resolution conversion means for converting the resolution of the data into the predicted resolution; Inference means for inferring the label of the data using a second model including a plurality of layers, with the data whose resolution has been converted as input, and the inference means infers the label of the data using a part of the activation values output from a predetermined layer of the first model An inference device characterized by this.

[0092] (Appended Claim 2) Activation value size conversion means for converting the size of a part of the activation values according to the predicted resolution, and the inference means combines, in the channel direction, the activation values output from the layer in the second model corresponding to the predetermined layer and a part of the activation values whose size has been changed, and inputs the combination result to the next layer in the second model The inference device according to Appended Claim 1.

[0093] (Appended Claim 3) Activation value size conversion means for converting the size of a part of the activation values into a certain size according to the smallest resolution among the plurality of resolution candidates, and Layer identification means for identifying a layer in the second model that outputs activation values that can be combined with a part of the activation values whose size has been changed, and the inference means combines, in the channel direction, the activation values output from the identified layer and a part of the activation values whose size has been changed, and inputs the combination result to the next layer in the second model The inference device according to Appended Claim 1.

[0094] (Appendix 4) The first model is a learned model with the weights from the first layer to a predetermined layer being the same as the weights from the first layer to the predetermined layer of the second model as the initial state. The inference device according to any one of Appendices 1 to 3.

[0095] (Appendix 5) The computer uses a first model including a plurality of layers with data of a certain resolution as input, predicts the minimum resolution among a plurality of resolution candidates that can infer the label of the data with a predetermined accuracy, converts the resolution of the data to the predicted resolution, uses a second model including a plurality of layers with the data whose resolution has been converted as input to infer the label of the data, and uses a part of the activation values output from a predetermined layer of the first model when inferring the label of the data. A method for inference, characterized by the above.

[0096] (Appendix 6) The computer converts the size of a part of the activation values according to the predicted resolution, combines, in the channel direction, the activation values output from the layer in the second model corresponding to the predetermined layer and the part of the activation values whose size has been changed, and inputs the combined result into the next layer in the second model. The inference method according to Appendix 5.

[0097] (Appendix 7) The computer converts the size of a part of the activation values to a certain size according to the smallest resolution among the plurality of resolution candidates, identifies the layer in the second model that outputs activation values that can be combined with the part of the activation values whose size has been changed, Combine the activation value output from the specified layer and a part of the activation value whose size has been transformed in the channel direction, and input the combination result into the next layer in the second model. The inference method according to Supplementary Note 5.

[0098] (Supplementary Note 8) Cause a computer to Using a first model including a plurality of layers with data of a certain resolution as input, perform a resolution prediction process of predicting the minimum resolution among a plurality of resolution candidates that can infer the label of the data with a predetermined accuracy. A resolution conversion process of converting the resolution of the data into the predicted resolution, and Using the data with the converted resolution as input, execute an inference process of inferring the label of the data using a second model including a plurality of layers. In the inference process, infer the label of the data using a part of the activation value output from a predetermined layer of the first model. An inference program for this purpose.

[0099] Part or all of the configurations described in Supplementary Notes 2 to 4, which are subordinate to Supplementary Note 1 described above, may be subordinate to Supplementary Notes 5 and 6 in the same subordinate relationship as Supplementary Notes 2 to 4. Furthermore, not limited to Supplementary Notes 1, 5, and 6, within the scope not departing from the above-described embodiments, similarly for various hardware, software, various recording means for recording software, or systems, part or all of the configurations described as supplementary notes may be made subordinate.

[0100] Although the present disclosure has been described with reference to the embodiments above, the present disclosure is not limited to the above-described embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. And each embodiment can be combined with other embodiments as appropriate.

Industrial Applicability

[0101] The present invention can be suitably applied to an inference device that infers an object shown in an image.

Description of Symbols

[0102] 1 Input section 2 Resolution prediction section 3 Resolution conversion section 4 Inference section 5 Activation value size conversion section 6 Output section 7 Learning section 81 Activation value size conversion section 82 Layer identification section 83 Inference section

Claims

1. Resolution prediction means for predicting, using a first model including a plurality of layers, the minimum resolution among candidates of a plurality of resolutions that can infer the label of the data with a predetermined accuracy, with data of a certain resolution as input; Resolution conversion means for converting the resolution of the data into the predicted resolution; Inference means for inferring the label of the data, using a second model including a plurality of layers, with the data whose resolution has been converted as input; The inference means infers the label of the data using a part of the activation values output from a predetermined layer of the first model An inference device characterized by this.

2. Activation value size conversion means for converting the size of a part of the activation values according to the predicted resolution; The inference means Combines, in the channel direction, the activation values output from the layer in the second model corresponding to the predetermined layer and a part of the activation values whose size has been changed, and inputs the combination result into the next layer in the second model The inference device according to Claim 1.

3. Activation value size conversion means for converting the size of a part of the activation values to a certain size according to the smallest resolution among the candidates of the plurality of resolutions; Layer identification means for identifying the layer in the second model that outputs activation values that can be combined with a part of the activation values whose size has been changed; The inference means Combines, in the channel direction, the activation values output from the identified layer and a part of the activation values whose size has been changed, and inputs the combination result into the next layer in the second model The inference device according to Claim 1.

4. The first model is a learned model with the weights from the first layer to a predetermined layer being the same as the weights from the first layer to a predetermined layer of the second model as the initial state. The inference device according to any one of Claims 1 to 3.

5. A computer Predicts, using a first model including a plurality of layers, the minimum resolution among candidates of a plurality of resolutions that can infer the label of the data with a predetermined accuracy, with data of a certain resolution as input; Converts the resolution of the data into the predicted resolution; Infers the label of the data, using a second model including a plurality of layers, with the data whose resolution has been converted as input; When inferring the label of the data, uses a part of the activation values output from a predetermined layer of the first model An inference method characterized by this.

6. The computer Convert the size of a part of the activation value according to the predicted resolution, Combine, in the channel direction, the activation value output from the layer in the second model corresponding to the predetermined layer and the part of the activation value whose size has been changed, and input the combination result into the next layer in the second model The inference method according to claim 5.

7. The computer Convert the size of a part of the activation value to a fixed size according to the smallest resolution among the plurality of resolution candidates, Identify the layer in the second model that outputs an activation value that can be combined with the part of the activation value whose size has been changed, Combine, in the channel direction, the activation value output from the identified layer and the part of the activation value whose size has been changed, and input the combination result into the next layer in the second model The inference method according to claim 5.

8. To the computer A resolution prediction process of predicting, using a first model including a plurality of layers, the smallest resolution among a plurality of resolution candidates that can infer the label of the data with a predetermined accuracy, with data of a certain resolution as input, A resolution conversion process of converting the resolution of the data to the predicted resolution, and An inference process of inferring the label of the data using a second model including a plurality of layers with the data whose resolution has been converted as input, In the inference process, use a part of the activation value output from a predetermined layer of the first model to infer the label of the data Inference program for