Image recognition apparatus and architecture searching method, and program

By using the block-level self-monitoring neural architecture search (BossNAS) method in image recognition equipment, the architectural search process of neural networks is optimized, and the problem of neural networks prone to redundant processing capabilities under the constraints of computing resources in the prior art is solved, achieving a balance between high accuracy and low computing resource consumption.

JP2025073883AActive Publication Date: 2025-05-13SOFTBANK CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023185025
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2025-05-13
Estimated Expiration
2043-10-27

AI Technical Summary

Technical Problem

The prior art generates redundant processing capabilities when generating neural networks that meet computing resource constraints, resulting in an increase in unnecessary processing load.

Method used

By introducing a block-level self-monitoring neural architecture search (BossNAS) method into the image recognition device, the architecture search process of the neural network is optimized, and the appropriate computing specification is selected to reduce the consumption of computing resources while ensuring a higher accuracy rate.

Benefits of technology

It realizes that while ensuring a high accuracy rate, it effectively reduces the computing resources consumed by neural networks in image recognition tasks and avoids redundant computing load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025073883000001_ABST
    Figure 2025073883000001_ABST
Patent Text Reader

Abstract

To make it possible to easily search for an architecture of a neural network which allows expectation of higher ratio of correct answer to some degree and sufficiently suppresses consumed calculation resources.SOLUTION: An image recognition apparatus according to the present invention includes: an evaluation block selecting unit for selecting an evaluation target block among a plurality of blocks; a setting unit for selecting, from among a plurality of candidates, a calculation specification defining calculation processing which the block executes, and setting the evaluation target block as an evaluation candidate calculation specification; a block evaluation unit for calculating an evaluation value of an output result of a neural network including the evaluation target block set with the evaluation candidate calculation specification; and a block calculation specification determination unit for determining a calculation specification of an evaluation target block based on a result of comparison between the evaluation value of the evaluation target block and an evaluation value of other block. The evaluation block selecting unit selects, from among the plurality of blocks, the evaluation target block in an order from a block with less pixel number of processing-target data to a block with more pixel number thereof.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image recognition device, an architecture search method, and a program that enable easy search for a neural network architecture that can be expected to have a relatively high accuracy rate and in which the computational resources consumed are sufficiently suppressed. [Background technology]

[0002] A neural network is generally a machine learning model that includes an input layer, a hidden layer, and an output layer, and predicts an output for an input. A neural network can have multiple hidden layers. In addition, in a neural network, model parameters are calculated by machine learning, and the model parameters are optimized.

[0003] A neural network has multiple layers connected before the output layer, and the output of each layer is used as the input to the next layer in the network. Each layer of the network calculates and generates output data from the input data according to the model parameters given to each layer.

[0004] In recent years, hybrid architectures consisting of various building blocks have been adopted as neural network structures (architectures). For example, in the field of image recognition, architectures based on convolutional neural networks (CNNs) are often used. Meanwhile, attention-based architectures have been highly evaluated alongside CNNs in recent years, and hybrid architectures that combine the two are sometimes used.

[0005] One technology that has attracted attention as an example of technology that can automatically optimize the structure of a neural network is Neural Architecture Search (NAS). With NAS, the structure of a neural network is optimized in a predefined search space prior to optimizing model parameters. For example, when optimizing a neural network related to image recognition using NAS, a structure is generated that combines multiple layers that process input data using operations such as convolution and attention.

[0006] Also, a technique has been proposed for determining the final architecture of a neural network based on the target computational resource usage of the final architecture (see, for example, Patent Document 1).

[0007] Furthermore, in NAS, block-wise self-supervised neural architecture search (BossNAS) has also been proposed as a technique that enables efficient search within a search space (see, for example, Non-Patent Document 1). [Prior art documents] [Patent documents]

[0008] [Patent Document 1] JP 2023-120204 A [Non-patent literature]

[0009] [Non-Patent Document 1] Changlin Li et al: BossNAS: Exploring Hybrid CNN-transformers with Block-wisely Self-supervised Neural Architecture Search (2021), The IEEE International Conference on ComputerVision(ICCV) Summary of the Invention [Problem to be solved by the invention]

[0010] A neural network can also be generated under a given constraint on the amount of calculation, as in Patent Document 1. However, even a neural network generated to satisfy the conditions imposed by such constraints has a problem in that it has redundant processing capabilities and may unnecessarily increase the processing load.

[0011] One aspect of the present invention aims to realize a technology that enables easy search for neural network architectures that can be expected to have a relatively high accuracy rate and consume sufficiently few computational resources. [Means for solving the problem]

[0012] An image recognition device according to one embodiment of the present invention is an image recognition device that generates a neural network in which a plurality of blocks corresponding to the number of pixels of the processing target data are combined, the plurality of blocks being connected to perform a predetermined arithmetic processing related to image feature extraction on processing target data including an image or feature map with a predetermined number of pixels, and outputting the result as output data with a predetermined number of dimensions. The image recognition device includes an evaluation block selection unit that selects an evaluation target block from among the plurality of blocks, a setting unit that selects an arithmetic specification that defines the arithmetic processing to be performed by the block from a plurality of candidates, and sets the evaluation target block as an evaluation candidate arithmetic specification, a block evaluation unit that calculates an evaluation value of the output result of the neural network including the evaluation target block to which the evaluation candidate arithmetic specification is set, and a block arithmetic specification determination unit that determines the arithmetic specification of the evaluation target block based on a comparison result between the evaluation value of the evaluation target block and the evaluation values ​​of other blocks, and the evaluation block selection unit selects an evaluation target block from among the plurality of blocks, starting from blocks with a smaller number of pixels of the processing target data toward blocks with a larger number of pixels.

[0013] An architecture exploration method according to one aspect of the present invention is a method for exploring architecture of a neural network in which a plurality of blocks corresponding to the number of pixels of the processing target data are combined, the blocks performing a predetermined calculation process related to image feature extraction on processing target data including an image or feature map with a predetermined number of pixels, and outputting the result as output data with a predetermined number of dimensions, the method comprising the steps of: selecting an evaluation target block from the plurality of blocks; selecting a calculation specification that defines the calculation process performed by the block from a plurality of candidates, and setting the evaluation target block as an evaluation candidate calculation specification; calculating an evaluation value of the output result of the neural network including the evaluation target block to which the evaluation candidate calculation specification is set; and determining the calculation specification of the evaluation target block based on a comparison result between the evaluation value of the evaluation target block and the evaluation values ​​of other blocks, and in the step of selecting the evaluation target block, evaluation target blocks are selected from the plurality of blocks starting from blocks with a smaller number of pixels of the processing target data toward blocks with a larger number of pixels.

[0014] Each aspect of the present invention may be realized by a computer. In this case, the present invention also includes a program for realizing the system on a computer by causing the computer to operate as each part (software element) of the above-mentioned system, and a computer-readable recording medium on which the program is recorded. Effect of the Invention

[0015] According to one aspect of the present invention, an object is to realize a technology that enables easy search for a neural network architecture that can be expected to have a relatively high accuracy rate and consumes sufficiently few computational resources. [Brief description of the drawings]

[0016] [Figure 1] 1 is a block diagram showing an example of the configuration of an image recognition device according to an embodiment; [Diagram 2]FIG. 2 is a diagram illustrating the calculation specifications of multiple blocks included in a prediction model. [Diagram 3] FIG. 2 is a diagram showing an example of an image recognized by an image recognition device. [Figure 4] FIG. 11 is a diagram showing another example of an image recognized by the image recognition device. [Diagram 5] 11 is a flowchart illustrating an example of the flow of an architecture exploration process. [Figure 6] 13 is a flowchart illustrating an example of the flow of a block evaluation process. [Figure 7] 13 is a flowchart illustrating an example of the flow of a computation specification process. [Figure 8] FIG. 1 is a diagram illustrating an example of the configuration of a computer that executes instructions of a program, which is software that realizes each function. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0017] First Embodiment (Image Recognition Device) Fig. 1 is a block diagram showing an example of the configuration of an image recognition device according to this embodiment. The image recognition device 10 shown in the figure is a device that predicts objects in an image included in input data using a prediction model, and outputs the prediction result as a recognition result for the image. In addition, the image recognition device 10 automatically generates a prediction model related to the above prediction. Here, the prediction model may be a neural network.

[0018] In other words, the image recognition device 10 performs a predetermined arithmetic process related to image feature extraction on processing target data including an image or feature map with a predetermined number of pixels, and generates a neural network in which a plurality of blocks corresponding to the number of pixels of the processing target data are connected to output data with a predetermined number of dimensions.

[0019] In the example of FIG. 1, the image recognition device 10 includes a search space setting unit 20, a prediction processing unit 30, and an architecture search unit 40.

[0020] (Search space setting part) The search space setting unit 20 sets a search space (or a search range) of a neural network that functions as a prediction model of the image recognition device 10. The search space setting unit 20 sets each block of a prediction model 32 (described later) and options and initial values ​​of the calculation specifications of each block. The search space setting unit 20 also sets the number of blocks of the prediction model 32 and the resolution of each block. The resolution of the processing target data of each block will be appropriately referred to as the block resolution. The processing target data of each block will be described later.

[0021] For example, the search space setting unit 20 sets the number of blocks to 5, and the resolution of each block to be 1 / 4 of the resolution of the model preceding the block. The number of blocks and the resolution of each block are arbitrary. The search space setting unit 20 sets the search space so that it includes multiple blocks, and the resolution of the subsequent blocks is lower than the resolution of the previous blocks.

[0022] The number of blocks and the resolution of each block may be set by the user, for example.

[0023] Furthermore, the search space setting unit 20 sets candidates selectable by the architecture search unit 40 for the computation specifications of each block. The computation specifications of each block are, for example, information specifying the type of computation executed in the block, the number of iterations of the computation, and the number of dimensions of the output data output by the block. The search space setting unit 20 sets, for example, a list of selectable candidates (values) for each of the above-mentioned types of computation, the number of iterations, and the number of dimensions of the output data.

[0024] The type of arithmetic processing is, for example, Convolution, Attention, etc. Furthermore, the number of repetitions is, for example, 1, 2, 4, etc. Furthermore, the number of dimensions of the output data is, for example, 128, 256, 512, etc. Furthermore, selectable candidates (values) for each of the arithmetic processing, the number of repetitions, and the number of dimensions of the output data may be set, for example, by the user.

[0025] (Prediction processing unit) The prediction processing unit 30 has an input unit 31 , a prediction model 32 , and an output unit 33 .

[0026] The input unit 31 accepts input of data. The input data input to the input unit 31 includes, for example, an image, a label of the image, etc. Alternatively, the input data may include an image obtained by performing a predetermined process such as a filter on an original image, data obtained by normalizing each pixel value after the filter process, etc.

[0027] The prediction model 32 learns model parameters for predicting an object in an image, for example, based on the image and label included in the data input via the input unit 31. After learning, the prediction model 32 predicts an object in an image included in the data input via the input unit 31.

[0028] The prediction model 32 is configured by a neural network, and the neural network includes a plurality of blocks set by the search space setting unit 20. In the example of Fig. 1, a block 110, a block 120, a block 130, a block 140, a block 150, ... are included in the prediction model 32. Also, each of the plurality of blocks is provided corresponding to a different resolution.

[0029] Each of the multiple blocks, except for the last block, is connected so that the output of each block becomes the input of the next block, and each block performs processing related to extracting image features and outputs output data to the block connected to it at the subsequent stage.

[0030] Furthermore, except for the first block, each block performs downsampling or the like on the input data to generate processing data for that block. Therefore, for example, the processing data for block 120 has fewer pixels than the processing data for block 110, and the processing data for block 130 has fewer pixels than the processing data for block 120. In this way, the blocks are provided so that the number of pixels of the processing data for the latter blocks is smaller (so that the resolution of the processing data is lower).

[0031] In the case of block 110, no downsampling process is performed on the input data, and the input data becomes the data to be processed.

[0032] The output unit 33 outputs a prediction result based on the output of the prediction model 32. For example, information indicating that an object in an image included in the input data is a cat is output as the prediction result. Alternatively, information indicating the probability that an object in an image included in the input data is a preset animal, tool, or the like may be output as the prediction result.

[0033] (Architecture Exploration Department 40) The architecture search unit 40 searches for optimal calculation specifications for each block of the prediction model 32. Here, the optimal calculation specifications refer to calculation specifications for each block that suppress the amount of calculation for the prediction processing unit 30 as a whole and enable the acquisition of prediction results with the highest possible accuracy under the amount of calculation.

[0034] That is, the architecture search unit 40 finds the optimal calculation specification by repeatedly selecting a calculation specification from the candidates for the calculation specification of each block set by the search space setting unit 20 and evaluating the prediction result of the prediction processing unit 30 corresponding to the selection. Identifying the calculation specification of each block included in the neural network in this way is also called searching for the architecture of the neural network.

[0035] In the example of FIG. 1, the architecture search unit 40 includes a block selection unit 41, a block processing setting unit 42, an evaluation unit 43, and a computation specification determination unit 44.

[0036] (Block selection section) The block selection unit 41 selects an evaluation target block from among a plurality of blocks included in the prediction model 32 .

[0037] Here, the evaluation target block refers to a block whose calculation specifications are changed when evaluating the prediction result by the prediction processing unit 30. For example, when evaluating the prediction result by the prediction processing unit 30 while changing calculation specifications such as the calculation process executed in the block 150, the number of repetitions of the calculation process, and the number of dimensions of the output data output by the block, the block 150 is the evaluation target block.

[0038] The block selection unit 41 first selects the block connected to the rearmost stage as the evaluation target block from among the multiple blocks included in the prediction model 32. For example, when the prediction model 32 includes blocks 110 to 150, the block selection unit 41 first selects block 150 as the evaluation target block, and then selects block 140 as the evaluation target block. Then, the block selection unit 41 selects blocks 130, 120, and 110 as the evaluation target blocks in this order.

[0039] In other words, the block selection unit 41 selects evaluation target blocks from blocks having low resolution (small number of pixels) of the processing target data toward blocks having high resolution (large number of pixels) of the processing target data.

[0040] (Block processing setting section) The block processing setting unit 42 sets a calculation specification that determines the calculation processing of the evaluation target block selected by the block selection unit 41. For example, when the evaluation target block is block 150, the block processing setting unit 42 selects a calculation specification from among the candidate calculation specifications set by the search space setting unit 20, and sets it as the evaluation candidate calculation specification for block 150.

[0041] For example, a list in which multiple options are described for each of the type of arithmetic processing, the number of iterations, and the number of dimensions of the output data is set as a candidate for the arithmetic specification by the search space setting unit 20. In the list, for example, numerical values ​​corresponding to each option are described, and the block processing setting unit 42 selects one option from the above-mentioned list and sets it as the setting value for each of the type of arithmetic processing, the number of iterations, and the number of dimensions of the output data.

[0042] As an example, the options for the type of computation process may be Convolution and Attention.

[0043] The options for the number of iterations may be 1, 2, and 4. The number of iterations does not necessarily mean the number of times the exact same operation is executed. For example, in the case of convolution processing using different kernels, where four types of convolution processing are executed once each, the type of operation is convolution, and the number of iterations is four.

[0044] Furthermore, the number of dimensions of the output data may be 128, 256, or 512. The number of dimensions of the output data may be the same as the number of channels of the output data.

[0045] The above options are merely examples, and it goes without saying that the options do not necessarily have to be included, or options other than those listed above may be included.

[0046] The block processing setting unit 42 sets multiple setting values ​​obtained by combining the options for the type of calculation processing, the number of repetitions, and the number of dimensions (of the output data) as evaluation candidate calculation specifications for evaluation of the prediction results by the evaluation unit 43 described later.

[0047] For example, when there are three options for each of the type of arithmetic processing, the number of repetitions, and the number of dimensions (of output data), the block processing setting unit 42 sets a setting value corresponding to a combination of the first arithmetic processing, the first number of repetitions, and the first number of dimensions as the first evaluation candidate arithmetic specification. The block processing setting unit 42 also sets a setting value corresponding to a combination of the second arithmetic processing, the first number of repetitions, and the first number of dimensions as the second evaluation candidate arithmetic specification. In this case, the block processing setting unit 42 sets 27 (=3×3×3) types of evaluation candidate arithmetic specifications.

[0048] In this manner, the block processing setting unit 42 selects a processing specification that defines the processing of the block from among a plurality of candidates, and sets the selected processing specification to the block to be evaluated as an evaluation candidate processing specification.

[0049] The setting values ​​of the calculation specifications of each block connected before the block to be evaluated are set to initial values ​​(initial calculation specifications). For example, if the block to be evaluated is block 150, the setting values ​​of blocks 110 to 140 are set to initial values, and if the block to be evaluated is block 140, the setting values ​​of blocks 110 to 130 are set to initial values.

[0050] The initial value may be, for example, a setting value that has the largest number of operations and the largest amount of calculation among the combinations of options. For example, the type of operation process may be Convolution, the number of iterations may be 4, and the number of dimensions of the output data may be 512. In other words, among the operations that can be executed in the block, the operation specification corresponding to the operation with the highest processing load is set as the initial operation specification. By setting the initial operation specification in the preceding block, it becomes possible to appropriately evaluate the feature detection performance in the following block.

[0051] Furthermore, among the evaluation candidate computation specifications set by the block processing setting unit 42, a minimum load computation specification that is an evaluation candidate computation specification with the lowest processing load (smallest amount of calculation) may be specified.

[0052] (Multiple block calculation specifications) 2 is a diagram for explaining the calculation specifications of a plurality of blocks included in the prediction model 32. In the diagram, blocks 110 to 150 are shown.

[0053] Block 110 is the first block among the multiple blocks included in prediction model 32, and image 111 included in the input data is input to block 110 as data to be processed. Image 111 is an RGB color image with 224×224 pixels, and is three-channel data. The number of channels can also be referred to as the number of dimensions.

[0054] In this example, the block 110 executes convolution processing as a computation process related to image feature extraction.

[0055] Also, the block 110 executes a total of two convolution processes by changing the kernel in the convolution process, etc. That is, the number of repetitions of the calculation process (convolution process) in the block 110 is two. In this case, the block 110 executes, for example, a convolution process CVL-1 using a first kernel, and executes a convolution process CVL-2 using a second kernel.

[0056] As a result of two convolution processes (for example, CVL-1 and CVL-2), tensors 112-1 and 112-2 are obtained, respectively. The output data of the block 110 has a pixel count of 224×224 and is 128-channel data. That is, the number of dimensions of the output data in the block 110 is 128.

[0057] The block 120 is the second block among the multiple blocks included in the prediction model 32, and data output from the block 110 is input to the block 120. In this example, the block 120 executes attention processing as a computation process related to image feature extraction.

[0058] Block 120 downsamples the data output from block 110 to generate a feature map 122 with 112×112 pixels and 128 channels, and executes one Attention process in total using this feature map as processing target data. That is, the number of times the calculation process in block 120 is repeated is one.

[0059] As a result of one Attention process, a tensor 122-1 is obtained. The output data of the block 120 has a pixel count of 112×112, and is 256-channel data. That is, the number of dimensions of the output data in the block 120 is 256.

[0060] Similarly, the type of calculation process, the number of iterations, and the number of dimensions of the output data are set for blocks 130, 140, and 150 as well.

[0061] 2, the blocks in the first stage have a larger number of pixels of the data to be processed (or a higher resolution), and the blocks in the second stage have a smaller number of pixels of the data to be processed (or a lower resolution). That is, the number of pixels of the feature map of block 130 is 56×56, the number of pixels of the feature map of block 140 is 28×28, and the number of pixels of the feature map of block 150 is 14×14. In this manner, the block with the second largest number of pixels of the data to be processed is combined with the block with the highest number of pixels of the data to be processed, the block with the third largest number of pixels of the data to be processed is combined with the block with the second largest number of pixels of the data to be processed, and so on.

[0062] That is, the neural network of the prediction model 32 further includes an image having a first number of pixels or a first block related to data to be processed including a feature map, an image connected after the first block and having a second number of pixels less than the first number of pixels, or a second block related to data to be processed including a feature map, and an image connected after the second block and having a third number of pixels less than the second number of pixels, or a third block related to data to be processed including a feature map.

[0063] In the example of FIG. 2, the resolution of block 110 is 224×224, whereas the resolution of block 120 is 112×112, which is 1 / 4 of the resolution. Similarly, for each of blocks 130 to 150, the resolution of the latter block is 1 / 4 of the resolution of the former block. However, the resolution of each block does not necessarily have to be set in such a regular manner. For example, the resolution of block 120 may be 1 / 4 of the resolution of block 110, and the resolution of block 130 may be 1 / 8 of the resolution of block 120. In short, it is sufficient that the resolution of the latter block is lower than the resolution of the former block.

[0064] (Evaluation Department) The evaluation unit 43 calculates an evaluation value of the output result of the neural network including the evaluation target block in which the evaluation candidate operation specification is set. That is, the evaluation unit 43 causes the prediction model 32 to execute a prediction process on the input data from the input unit 31, and evaluates the prediction result output from the output unit 33. At this time, the prediction model 32 causes the evaluation target block to execute a calculation according to the operation specification corresponding to each evaluation candidate operation specification set by the block processing setting unit 42.

[0065] The evaluation unit 43, for example, executes a computation specification corresponding to a certain candidate computation specification for evaluation in the block to be evaluated, calculates the probability that each of the prediction results output from the output unit 33 corresponding to a plurality of input data is correct, i.e., calculates the Accuracy indicating the accuracy rate, and sets it as an evaluation value corresponding to the candidate computation specification for evaluation. Whether each of the prediction results is correct or not can be determined, for example, by comparing the label included in the input data with the prediction result.

[0066] In this way, the evaluation unit 43 calculates an evaluation value corresponding to the evaluation candidate operation specification of the evaluation target block based on the accuracy rate obtained by comparing the prediction result output by the neural network including the evaluation target block, for which the evaluation candidate operation specification is set, in response to specified input data with the label.

[0067] In addition, by setting the initial calculation specification in the block in the preceding stage, the difference in the accuracy rate corresponding to each evaluation candidate calculation specification set in the block in the following stage becomes significant. In other words, it becomes easier to evaluate which calculation specification is optimal for the evaluation target block.

[0068] In addition, the evaluation unit 43 includes the calculation amount of the entire prediction processing unit 30 when the calculation specification corresponding to the evaluation candidate calculation specification is executed by the evaluation target block in the evaluation value corresponding to the evaluation candidate calculation specification. The calculation amount may be, for example, the number of floating-point calculations, and may be calculated as FLOPs.

[0069] Furthermore, for example, the difference (BA) between the calculation amount A of the entire prediction processing unit 30 when the evaluation target block is caused to execute a calculation according to the lowest load calculation specification and the calculation amount B of the entire prediction processing unit 30 when the evaluation target block is caused to execute a calculation according to the evaluation candidate calculation specification may be included in the evaluation value corresponding to the evaluation candidate calculation specification. The difference (BA) means an increase in the calculation amount when the evaluation target block is caused to execute a calculation according to the evaluation candidate calculation specification.

[0070] As described above, the lowest load computation specification is the evaluation candidate computation specification with the lowest processing load (the smallest amount of calculation) among the evaluation candidate computation specifications set by the block processing setting unit .

[0071] The evaluation value calculated by the evaluation unit 43 is stored in association with the evaluation target block and the evaluation candidate computation specification.

[0072] (Calculation specification determination section) The computation specification determination unit 44 determines the computation specification of the evaluation target block based on the result of comparing the evaluation value of the evaluation target block with the evaluation values ​​of the other blocks.

[0073] The computation specification determination unit 44 compares the Accuracy of the evaluation values ​​corresponding to the candidate evaluation computation specifications, thereby identifying the candidate evaluation computation specification with the highest accuracy rate, and stores this as the search result computation specification for the block to be evaluated.

[0074] Then, the computation specification determination unit 44 calculates the ratio of Accuracy (Ak) corresponding to the search result computation specification of the current evaluation target block (Kth evaluation target block) to Accuracy (Ak-1) of the previous evaluation target block (K-1th evaluation target block). Note that Accuracy (Ak-1) means the Accuracy corresponding to the search result computation specification when the search result computation specification is set for the K-1th evaluation target block, and means the Accuracy corresponding to the minimum load computation specification when the minimum load computation specification is set for the K-1th evaluation target block.

[0075] In this case, the accuracy rate ratio I is calculated using the following formula:

[0076] I=Ak / Ak-1 When the evaluation target block is the first block, that is, the block having the lowest resolution of the processing target data, the accuracy rate ratio I is always 1 since there is no previous evaluation target block.

[0077] The calculation specification determination unit 44 determines the calculation specification of the evaluation target block based on the accuracy rate ratio I and the calculation amount of the entire prediction processing unit 30.

[0078] As described above, the increase in the calculation amount of the entire prediction processing unit 30 when the evaluation target block is caused to execute a calculation according to a calculation specification corresponding to the evaluation candidate calculation specification is included in the evaluation value corresponding to the evaluation candidate calculation specification, so that the increase in the calculation amount when the evaluation target block is caused to execute a calculation according to the search result calculation specification can be obtained. For example, a ratio (or difference) F between the increase in calculation amount corresponding to the search result calculation specification and a predetermined target calculation amount is calculated. The calculation of the calculation amount ratio (difference) F may be performed by the evaluation unit 43 or may be performed by the calculation specification determination unit 44. The target calculation amount may be an upper limit of the calculation amount of the entire prediction processing unit 30.

[0079] Then, the computation specification determination unit 44 calculates the ratio between the accuracy rate ratio I and the calculation amount ratio (difference) F as the performance improvement contribution of the block to be evaluated. In this case, the performance improvement contribution C is calculated by the following formula.

[0080] C=I / F In other words, the performance improvement contribution C indicates how great the improvement in performance (correct answer rate) obtained by changing the calculation specifications of the block being evaluated is compared to the increase in the amount of calculation caused by changing the calculation specifications of the block being evaluated.

[0081] For example, the computation specification determination unit 44 compares the value of the performance improvement contribution degree C with a predetermined threshold value, and if the value is equal to or greater than the threshold value, determines the computation specification of the evaluation target block to be the search result computation specification. On the other hand, if the value of the performance improvement contribution degree C is less than the threshold value, the computation specification determination unit 44 determines the computation specification of the evaluation target block to be the minimum load computation specification.

[0082] In this way, the calculation specification determination unit 44 calculates a first ratio value calculated as the ratio between the accuracy rate corresponding to the search result calculation specification of the evaluation target block selected by the block selection unit 41 in the Kth time and the accuracy rate corresponding to the calculation specification determined by the block calculation specification determination unit for the evaluation target block selected by the block selection unit 41 in the K-1th time, and a third ratio value calculated as the ratio between the increase in the calculation amount of the neural network when the search result calculation specification is set for the evaluation target block selected by the block selection unit 41 in the Kth time and a predetermined target calculation amount, and determines the calculation specification of the evaluation target block by comparing the third ratio value with a threshold value.

[0083] Furthermore, the block operation specification determination unit 44 sets the operation specification of the evaluation target block to the search result operation specification when the third ratio value is equal to or greater than the threshold value, and sets the operation specification of the evaluation target block to a predetermined operation specification when the third ratio value is less than the threshold value. Here, the predetermined operation specification may be, for example, a minimum load operation specification.

[0084] To more simply determine the calculation specifications for a block, the calculation specification determination unit 44 may, for example, compare the value of the accuracy rate ratio I with a predetermined threshold, and if it is equal to or greater than the threshold, determine the calculation specifications for the block to be evaluated as the search result calculation specifications. In this case, if the accuracy rate ratio I is less than the threshold, the calculation specification determination unit 44 determines the calculation specifications for the block to be evaluated as the minimum load calculation specifications.

[0085] 3 is a diagram showing an example of an image recognized by image recognition device 10. Image 211 in the figure includes a spray can as a subject, and image 212 includes a bottle of hand soap as a subject. For example, in order to correctly recognize a spray can and a bottle of hand soap, a relatively high-resolution image or a feature map needs to be analyzed in detail.

[0086] For example, to extract features unique to a spray can that are not present in a hand soap bottle, attention must be paid to the details of the subject, and if the resolution of the image or feature map of the data to be processed is low, the features cannot be extracted properly. In such cases, it is necessary to improve the calculation specifications of blocks with high resolution of the data to be processed (e.g., blocks 110 and 120 in FIG. 2).

[0087] Fig. 4 is a diagram showing another example of an image recognized by the image recognition device 10. An image 221 in the figure includes an ostrich as a subject, and an image 222 includes a chicken as a subject. For example, in order to recognize an ostrich and a chicken without mistaking them for each other, it is not necessary to use a high-resolution image or to analyze the feature map in detail. It is considered that a relatively low-resolution image or to analyze the feature map in detail will be sufficient to recognize an ostrich and a chicken without mistaking them for each other.

[0088] For example, to extract features specific to an ostrich but not to a chicken, it is not necessary to focus on the details of the subject, but rather only on the general appearance of the subject. Therefore, even if the resolution of the image or feature map of the data to be processed is low, the features can be properly extracted. In such a case, it is necessary to improve the calculation specifications of the blocks with low resolution of the data to be processed (e.g., blocks 140 and 150 in FIG. 2).

[0089] For example, in a computation process in a block with high resolution of the processing target data, the number of pixels in the image and feature map is large, so the amount of calculation increases compared to a computation process in a block with low resolution of the processing target data, even if the computation process is the same type of computation process. In a block with high resolution of the processing target data, if the number of repetitions of the computation process or the number of dimensions of the output data is increased, the increase in the amount of calculation becomes even more remarkable.

[0090] In this way, enhancing the processing content in blocks with high resolution of the processing target data significantly increases the calculation amount of the entire prediction processing unit 30, but enhancing the processing content in blocks with low resolution of the processing target data results in a relatively small increase in the calculation amount of the entire prediction processing unit 30. Therefore, if complex arithmetic processing can be executed multiple times in blocks with low resolution of the processing target data to output output data with a large number of channels, and the processing load in blocks with high resolution of the processing target data can be reduced, it becomes possible to suppress the calculation amount of the entire prediction processing unit 30 and realize predictions with a relatively high degree of accuracy.

[0091] Moreover, the images recognized by the image recognition device 10 include a plurality of images as shown in Fig. 3, and also a plurality of images as shown in Fig. 4. For example, by repeating evaluation while changing the calculation specifications for blocks of low resolution data to be processed, it is possible to identify calculation specifications that can accurately recognize at least images as shown in Fig. 4. It is not realistic to have an image recognition device that can accurately recognize all images, and a relatively high accuracy rate often satisfies the performance requirements required of an image recognition device.

[0092] In such a case, it is possible to satisfy the performance requirements required for an image recognition device even if it is not possible to accurately recognize an image such as that shown in Figure 3. By repeating evaluation while changing the calculation specifications for blocks with low resolution of the data to be processed, it is possible to set calculation specifications that can accurately recognize at least an image such as that shown in Figure 4, thereby avoiding the generation of a redundant neural network.

[0093] In other words, by appropriately setting the threshold value to be compared with the performance improvement contribution C, it is possible to easily search for a neural network architecture that can expect a relatively high accuracy rate and in which the computational resources consumed are sufficiently suppressed.

[0094] (Neural network generation process) Next, an example of architecture search processing executed by the image recognition device 10 according to this embodiment will be described. Fig. 5 is a flowchart illustrating an example of the flow of the architecture search processing.

[0095] In step S101, the search space setting unit 20 sets a search space (or a search range) of a neural network that functions as a prediction model of the image recognition device 10.

[0096] The search space setting unit 20 sets each block of the prediction model 32 described later, and candidates and initial values ​​of the calculation specifications of each block. In addition, the search space setting unit 20 sets the number of models of the prediction model 32 and the resolution of each model.

[0097] In step S102, the block selection unit 41 selects an evaluation target block from among a plurality of blocks included in the prediction model 32. At this time, the block selection unit 41 selects the evaluation target block from a block having a low resolution (small number of pixels) of the processing target data toward a block having a high resolution (large number of pixels) of the processing target data.

[0098] In step S103, a block evaluation process is executed. At this time, the block process setting unit 42 sets an evaluation candidate calculation specification for the evaluation target block selected in step S102, and the evaluation unit 43 calculates an evaluation value corresponding to the evaluation candidate calculation specification.

[0099] In step S104, the computation specification determination unit 44 determines the computation specification of the evaluation target block selected in step S102.

[0100] In step S105, it is determined whether or not there is a next evaluation target block. If it is determined that there is a next evaluation target block, the process returns to step S102, and a new evaluation target block is selected by the block selection unit 41. At this time, a block with a higher resolution than the previously selected block is selected as the new evaluation target block.

[0101] In this manner, the block with the highest resolution is selected as the block to be evaluated, and the processes of steps S102 to S105 are repeatedly executed until the calculation specifications are determined.

[0102] In this manner, the neural network generation process is carried out.

[0103] (Block evaluation process) Next, a description will be given of an example of the block evaluation process in step S103 in Fig. 5. Fig. 6 is a flowchart illustrating an example of the flow of the block evaluation process.

[0104] In step S121, the block processing setting unit 42 sets the evaluation candidate calculation specification. At this time, for example, a list in which multiple options are described is used as a candidate calculation specification, and the block processing setting unit 42 selects one of the options in the list and sets it as the setting value for each of the calculation process, the number of repetitions, and the number of dimensions of the output data. The block processing setting unit 42 sets multiple setting values ​​obtained by combining each option as the evaluation candidate calculation specification.

[0105] In step S122, the prediction model 32 performs image recognition. At this time, data including, for example, an image and a label of the image is input as input data from the input unit 31, the prediction model 32 predicts an object in the image included in the input data, and the output unit 33 outputs a prediction result based on the output of the prediction model 32. For example, the input data includes a plurality of images, and a plurality of prediction results are output corresponding to each of the images.

[0106] In step S123, the evaluation unit 43 calculates an evaluation value of the output result of the neural network of the prediction model 32 obtained by the process of step S122. At this time, the evaluation unit 43 causes the evaluation target block to execute a calculation specification corresponding to a certain evaluation candidate calculation specification, and calculates Accuracy, which indicates the accuracy rate of each of the prediction results output from the output unit 33 corresponding to a plurality of input data.

[0107] Furthermore, the evaluation unit 43 includes in the evaluation value corresponding to the evaluation candidate computation specification an increase in the amount of calculation of the entire prediction processing unit 30 when the evaluation target block is caused to execute a computation according to the evaluation candidate computation specification.

[0108] In step S124, the evaluation unit 43 stores the evaluation value calculated in step S123 in association with the evaluation candidate calculation specification set in step S121.

[0109] In this manner, the block evaluation process is carried out.

[0110] (Calculation specification determination process) Next, a description will be given of an example of the computation specification determination process in step S104 in Fig. 5. Fig. 7 is a flowchart illustrating an example of the flow of the computation specification determination process.

[0111] In step S141, the computation specification determination unit 44 compares the Accuracy of the evaluation values ​​corresponding to the candidate evaluation computation specifications to identify the candidate evaluation computation specification with the highest accuracy rate.

[0112] At this time, the evaluation value of the evaluation candidate computation specification (search result computation specification) with the highest accuracy rate is also identified. When the current evaluation target block is the Kth evaluation target block, this evaluation value includes Accuracy (Ak).

[0113] In step S142, the computation specification determination unit 44 stores the evaluation candidate computation specification identified in step S141 as the search result computation specification of the evaluation target block.

[0114] In step S143, the computation specification determination unit 44 obtains the evaluation value of the previously evaluated block. When the current block to be evaluated is the Kth block to be evaluated, the evaluation value of the previously evaluated block includes the Accuracy (Ak-1) of the K-1th block to be evaluated. Note that, when the search result computation specification is set for the K-1th block to be evaluated, Accuracy (Ak-1) means the Accuracy corresponding to the search result computation specification, and when the minimum load computation specification is set for the K-1th block to be evaluated, it means the Accuracy corresponding to the minimum load computation specification.

[0115] Moreover, if the current evaluation target block is the first block, the process of step S143 is skipped since there is no previous evaluation target block.

[0116] In step S144, the computation specification determination unit 44 calculates the ratio I (=Ak / Ak-1) of the accuracy rates.

[0117] In step S145, the computation specification determination unit 44 calculates a ratio F between the increase in the computation amount and the target computation amount.

[0118] In step S146, the calculation specification determination unit 44 calculates the performance improvement contribution C (=I / F) of the block being evaluated, which is the ratio between the accuracy rate ratio I calculated in step S144 and the calculation amount ratio F calculated in step S145.

[0119] In step S147, the computation specification determination unit 44 judges whether the performance improvement contribution degree C calculated in step S146 is equal to or greater than the threshold value Th. If it is judged that the performance improvement contribution degree C is equal to or greater than the threshold value Th, the process of step S148 is executed. On the other hand, if it is judged that the performance improvement contribution degree C is less than the threshold value Th, the process of step S149 is executed.

[0120] In step S148, the computation specification determination unit 44 sets the search result computation specification for the evaluation target block (Kth evaluation target block).

[0121] In step S149, the computation specification determination unit 44 sets the minimum load computation specification for the evaluation target block (Kth evaluation target block).

[0122] In this manner, the computation specification determination process is carried out.

[0123] (Effects of the First Embodiment) As described above, according to the image recognition device 10 of this embodiment, the calculation specifications of each block of the prediction model 32 can be easily set, and the architecture of the neural network can be easily explored.

[0124] That is, the architecture search unit 40 selects an arithmetic specification from among the candidates for the arithmetic specification of each block set by the search space setting unit 20, and by repeatedly evaluating the prediction results by the prediction processing unit 30 based on the selection, the optimal arithmetic specification can be found.

[0125] In addition, since evaluation target blocks are selected from blocks with low resolution (small number of pixels in the data to be processed) toward blocks with high resolution (large number of pixels in the data to be processed), it is possible to prevent the generated neural network from becoming redundant. In other words, since the calculation specifications of the evaluation target blocks are determined based on the performance improvement contribution C, which indicates how large the degree of performance improvement obtained by changing the calculation specifications is relative to the increase in the amount of calculation caused by the change in the calculation specifications, it is possible to prevent the setting of calculation specifications that do not improve performance despite consuming large calculation resources.

[0126] In particular, because evaluation target blocks are selected in ascending order of resolution, the higher the resolution of the block, the smaller the performance improvement contribution C. Therefore, according to this embodiment, a relatively high accuracy rate can be expected, and it becomes possible to easily search for a neural network architecture that consumes less computational resources.

[0127] Second Embodiment Next, a second embodiment will be described. The configuration of an image recognition device 10 in the second embodiment is similar to that in the first embodiment.

[0128] In the first embodiment, the computation specifications of each block are, for example, the type of computation executed in the block, the number of repetitions of the computation, and the number of dimensions of the output data output by the block. Also, it was explained that the block processing setting unit 42 selects one option from the above-mentioned list and sets it as the setting value for each of the type of computation, the number of repetitions, and the number of dimensions of the output data. However, the computation specifications of each block and the setting of the evaluation candidate computation specifications by the block processing setting unit 42 may be performed as follows.

[0129] (element to be set) For example, the set values ​​for the type of arithmetic processing and the number of repetitions may be fixed to initial values, and the block processing setting unit 42 may set only the set value for the number of dimensions of the output data. Also, only the set value for the number of repetitions may be set, or only the set value for the type of arithmetic processing may be set.

[0130] In this way, only one of the three elements may be selected, and similarly, only two of the three elements may be selected.

[0131] (Skip repeats) For example, it is assumed that the type of calculation process to be executed in each block and the number of repetitions N of the calculation are set in advance by the search space setting unit 20. The block process setting unit 42 sets a calculation specification that skips the Kth (1≦K≦N) repetition as an evaluation candidate calculation specification. Note that there may be multiple repetitions to be skipped (for example, the Kth and Lth repetitions).

[0132] For example, assume that the type of arithmetic processing of block 120 is set to Convolution by the search space setting unit 20, and the number of iterations is set to 4. In this case, for example, from among the convolution processing CVL-1 to convolution processing CVL-4 using the first kernel to the fourth kernel, respectively, a convolution processing to be skipped is specified, and an evaluation candidate arithmetic specification is set.

[0133] For example, the computation specification in which the convolution processes CVL-2 and CVL-4 are skipped and the convolution processes CVL-1 and CVL-3 are executed is set as the evaluation candidate computation specification.

[0134] (Repeat order change) For example, it is assumed that the type of calculation process to be executed in each block and the number of repetitions N of the calculation are set in advance by the search space setting unit 20. The block process setting unit 42 sets, as a candidate calculation specification to be evaluated, a calculation specification in which the order of the Kth repetition is swapped with the Lth repetition.

[0135] For example, assume that the search space setting unit 20 sets the type of arithmetic processing of the block 120 to Convolution and the number of iterations to 4. In this case, for example, the order of the Convolution processing CVL-1 to Convolution processing CVL-4 using the first kernel to the fourth kernel, respectively, is swapped to set the evaluation candidate arithmetic specification.

[0136] For example, a computation specification in which convolution processing CVL-1, convolution processing CVL-3, convolution processing CVL-4, and convolution processing CVL-2 are executed in this order is set as an evaluation candidate computation specification.

[0137] <Third embodiment> Next, a third embodiment will be described. The configuration of an image recognition device 10 in the third embodiment is similar to that in the first embodiment.

[0138] In the above example, the block selection unit 41 selects, as evaluation target blocks, blocks connected to the most downstream side from among a plurality of blocks included in the prediction model 32. For example, when the prediction model 32 includes blocks 110 to 150, the block selection unit 41 first selects block 150 as the evaluation target block, and then selects block 140 as the evaluation target block. It has been described that the block selection unit 41 selects blocks 130, 120, and 110 as the evaluation target blocks in this order.

[0139] However, the computation specifications of some of the blocks included in the prediction model 32 may be fixed to the initial computation specifications. For example, it is assumed that the computation specifications of the blocks 150 and 120 are fixed to the initial computation specifications. In this case, the block selection unit 41 first selects the block 140 as the block to be evaluated, then selects the block 130 as the block to be evaluated, and thereafter selects the block 110 as the block to be evaluated.

[0140] However, even in this case, the block selection unit 41 still selects evaluation target blocks from blocks having low resolution (small number of pixels) of the processing target data toward blocks having high resolution (large number of pixels) of the processing target data.

[0141] <Example of software implementation> The image recognition device 10 described above is a program for causing a computer to function, and can be realized by a program for causing a computer to function as the image recognition device 10. In this case, the image recognition device 10 includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. An example of such a computer is shown in FIG. 8.

[0142] The computer 500 includes at least one processor 501 and at least one memory 502. The memory 502 stores a program 520 for causing the computer 500 to operate as the image recognition device 10. In the computer 500, the processor 501 reads and executes the program 520 from the memory 502, thereby implementing each function of the image recognition device 10.

[0143] The processor 501 may be, for example, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), a DSP (Digital Signal Processor), an MPU (Micro Processing Unit), an FPU (Floating point number Processing Unit), a PPU (Physics Processing Unit), a microcontroller, or a combination of these.

[0144] As the memory 502, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination of these can be used.

[0145] The computer 500 may further include a RAM (Random Access Memory) for expanding the program 520 during execution and for temporarily storing various data. The computer 500 may further include a communication interface for transmitting and receiving data to and from other devices. The computer 500 may further include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.

[0146] Furthermore, the program 520 for causing the computer 500 to operate as the image recognition device 10 can be recorded on a non-transitory tangible recording medium 530 that is readable by the computer 500. For example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit can be used as such a recording medium 530. The computer 500 can acquire the program 520 via such a recording medium 530.

[0147] In addition, the program 520 for operating the computer 500 as the image recognition device 10 can be transmitted via a transmission medium. For example, a communication network or a broadcast wave can be used as such a transmission medium. The computer 500 can also acquire the program 520 via such a transmission medium.

[0148] In addition, some or all of the functions of the image recognition device 10 can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the above control blocks is formed is also included in the scope of the present invention. In addition, the functions of each of the above control blocks can be realized by, for example, a quantum computer.

[0149] Furthermore, the effects of each aspect of the present invention described above contribute to the achievement of Goal 9 of the Sustainable Development Goals (SDGs) advocated by the United Nations, "Build resilient infrastructure, promote inclusive and sustainable industrialization," among others.

[0150] Furthermore, the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

[0151] 〔summary〕 An image recognition device according to a first aspect of the present invention is an image recognition device that generates a neural network in which a plurality of blocks corresponding to the number of pixels of the target data are combined, the block performing a predetermined arithmetic process related to image feature extraction on target data including an image or feature map with a predetermined number of pixels, and outputting the result as output data with a predetermined number of dimensions, the image recognition device comprising: an evaluation block selection unit that selects an evaluation target block from among the plurality of blocks; a setting unit that selects an arithmetic specification that defines the arithmetic process to be performed by the block from a plurality of candidates, and sets the evaluation target block as an evaluation candidate arithmetic specification; a block evaluation unit that calculates an evaluation value of the output result of the neural network including the evaluation target block to which the evaluation candidate arithmetic specification is set; and a block arithmetic specification determination unit that determines the arithmetic specification of the evaluation target block based on a comparison result between the evaluation value of the evaluation target block and the evaluation values ​​of other blocks, and the evaluation block selection unit selects an evaluation target block from among the plurality of blocks, starting from blocks with a smaller number of pixels of the target data toward blocks with a larger number of pixels.

[0152] In the image recognition device of aspect 2 of the present invention, in the above aspect 1, the block evaluation unit calculates an evaluation value corresponding to the evaluation candidate operation specification of the evaluation target block based on the accuracy rate obtained by comparing the prediction result output by the neural network including the evaluation target block, in which the evaluation candidate operation specification is set, in response to specified input data with a label.

[0153] In the image recognition device of aspect 3 of the present invention, in the above-mentioned aspect 2, the block evaluation unit further calculates the calculation amount of the neural network including the evaluation target block in which the evaluation candidate calculation specification is set, and the evaluation value further includes the calculation amount.

[0154] In the image recognition device according to aspect 4 of the present invention, in aspect 2 or 3 above, the block evaluation unit retains the evaluation candidate calculation specification corresponding to the evaluation value with the highest accuracy rate as the search result calculation specification for the evaluation target block.

[0155] In the image recognition device of aspect 5 of the present invention, in the above aspect 4, the block operation specification determination unit calculates a first ratio value calculated as the ratio between the accuracy rate corresponding to the search result operation specification of the evaluation target block selected by the evaluation block selection unit in the Kth time and the accuracy rate corresponding to the operation specification determined by the block operation specification determination unit in the evaluation target block selected by the evaluation block selection unit in the K-1th time, and a third ratio value calculated as the ratio between the increase in the calculation amount of the neural network when the search result operation specification is set to the evaluation target block selected by the evaluation block selection unit in the Kth time and a predetermined target calculation amount, and determines the operation specification of the evaluation target block by comparing the third ratio value with a threshold value.

[0156] In the image recognition device of aspect 6 of the present invention, in the above-mentioned aspect 5, the block operation specification determination unit sets the operation specification of the evaluation target block to a search result operation specification when the third ratio value is greater than or equal to the threshold value, and sets the operation specification of the evaluation target block to a predetermined operation specification when the third ratio value is less than the threshold value.

[0157] An image recognition device according to a seventh aspect of the present invention is, in the above-mentioned sixth aspect, characterized in that the predetermined calculation specification is a calculation specification corresponding to the smallest amount of calculation among the evaluation candidate calculation specifications.

[0158] An image recognition device according to aspect 8 of the present invention is any one of aspects 1 to 7 above, wherein the calculation specifications are setting values ​​indicating a combination of the type of the calculation process of the block, the number of repetitions of the calculation process, and the number of dimensions of the output data.

[0159] An image recognition device according to a ninth aspect of the present invention is any one of aspects 1 to 8 above, wherein the neural network further includes an image having a first number of pixels or a first block of processing target data including a feature map, an image having a second number of pixels less than the first number of pixels connected downstream of the first block, or a second block of processing target data including a feature map, and an image having a third number of pixels less than the second number of pixels connected downstream of the second block, or a third block of processing target data including a feature map.

[0160] An architecture exploration method according to aspect 10 of the present invention is a method for searching for an architecture of a neural network in which a plurality of blocks corresponding to the number of pixels of the data to be processed are connected, the blocks performing a predetermined calculation process related to image feature extraction on data to be processed including an image or feature map with a predetermined number of pixels, and outputting the result as output data with a predetermined number of dimensions, the method comprising the steps of: selecting an evaluation target block from the plurality of blocks; selecting a calculation specification that defines the calculation process performed by the block from a plurality of candidates, and setting the calculation specification as an evaluation candidate calculation specification for the evaluation target block; calculating an evaluation value of the output result of the neural network including the evaluation target block to which the evaluation candidate calculation specification is set; and determining the calculation specification for the evaluation target block based on a comparison result between the evaluation value of the evaluation target block and the evaluation values ​​of other blocks, and in the step of selecting the evaluation target block, evaluation target blocks are selected from the plurality of blocks, starting from blocks with a smaller number of pixels of the data to be processed toward blocks with a larger number of pixels.

[0161] A program according to an eleventh aspect of the present invention causes a computer to function as an image recognition device that generates a neural network in which a plurality of blocks corresponding to the number of pixels of the processing target data are combined, the plurality of blocks being blocks that perform a predetermined arithmetic process related to image feature extraction on processing target data including an image or feature map with a predetermined number of pixels, and output as output data with a predetermined number of dimensions, the program comprising: an evaluation block selection unit that selects an evaluation target block from among the plurality of blocks; a setting unit that selects an arithmetic specification that defines the arithmetic process to be performed by the block from a plurality of candidates, and sets the evaluation target block as an evaluation candidate arithmetic specification; a block evaluation unit that calculates an evaluation value of the output result of the neural network including the evaluation target block to which the evaluation candidate arithmetic specification is set; and a block arithmetic specification determination unit that determines the arithmetic specification of the evaluation target block based on a comparison result between the evaluation value of the evaluation target block and the evaluation values ​​of other blocks, the evaluation block selection unit causing a computer to function as an image recognition device that selects an evaluation target block from among the plurality of blocks, starting from blocks with a smaller number of pixels of the processing target data toward blocks with a larger number of pixels. [Explanation of symbols]

[0162] 10 Image Recognition Device 20 Search space setting section 30 Prediction processing unit 31 Input section 32 Predictive Models 33 Output section 40 Architecture Exploration Department 41 Block Selection Section 42 Block processing setting section 43 Evaluation Department 44 Calculation specification determination section 110~150 blocks

Claims

1. An image recognition device that generates a neural network in which a plurality of blocks corresponding to the number of pixels of the processing target data are connected, the blocks executing a predetermined arithmetic process related to image feature extraction for processing target data including an image or feature map with a predetermined number of pixels and outputting the result as output data with a predetermined number of dimensions, an evaluation block selection unit that selects an evaluation target block from the plurality of blocks; a setting unit that selects an operation specification that defines the operation processing to be executed by the block from among a plurality of candidates, and sets the selected operation specification as an evaluation candidate operation specification in the evaluation target block; a block evaluation unit that calculates an evaluation value of an output result of the neural network including the evaluation target block in which the evaluation candidate calculation specification is set; a block operation specification determination unit that determines operation specifications for the evaluation target block based on a comparison result between the evaluation value of the evaluation target block and the evaluation values ​​of other blocks; The evaluation block selection unit selects evaluation target blocks from among the plurality of blocks, starting from blocks with a smaller number of pixels of the processing target data toward blocks with a larger number of pixels. Image recognition device.

2. The block evaluation unit calculates an evaluation value corresponding to the evaluation candidate operation specification of the evaluation target block based on a rate of accuracy obtained by comparing a prediction result output by the neural network including the evaluation target block, in which the evaluation candidate operation specification is set, with a label in response to predetermined input data. The image recognition device according to claim 1 .

3. The block evaluation unit further calculates a calculation amount of the neural network including the evaluation target block for which the evaluation candidate calculation specification is set, and the evaluation value further includes the calculation amount. The image recognition device according to claim 2 .

4. The block evaluation unit holds the evaluation candidate calculation specification corresponding to the evaluation value with the highest accuracy rate as a search result calculation specification for the evaluation target block. The image recognition device according to claim 3 .

5. The block operation specification determination unit is a first ratio value calculated as a ratio between a correct answer rate corresponding to the search result operation specification of the evaluation target block selected by the evaluation block selection unit in the Kth iteration and a correct answer rate corresponding to the operation specification determined by the block operation specification determination unit in the evaluation target block selected by the evaluation block selection unit in the K-1th iteration; Calculating a third ratio value calculated by a ratio between an increase in the amount of calculation of the neural network when the search result calculation specification is set to the evaluation target block selected by the evaluation block selection unit for the Kth time and a second ratio value which is a ratio between the increase in the amount of calculation and a predetermined target amount of calculation; The third ratio value is compared with a threshold value to determine a calculation specification of the evaluation target block. The image recognition device according to claim 4.

6. The block operation specification determination unit If the third ratio value is equal to or greater than the threshold value, a calculation specification for the evaluation target block is set to a search result calculation specification; If the third ratio value is less than the threshold value, the calculation specification of the evaluation target block is set to a predetermined calculation specification. The image recognition device according to claim 5.

7. The predetermined calculation specifications are: Among the evaluation candidate calculation specifications, the calculation specification that corresponds to the smallest amount of calculation is The image recognition device according to claim 6.

8. The calculation specifications are set values ​​that indicate a combination of the type of calculation process of the block, the number of repetitions of the calculation process, and the number of dimensions of the output data. The image recognition device according to claim 1 .

9. The neural network includes a first block of processing target data including an image having a first number of pixels or a feature map, a second block of processing target data including an image having a second number of pixels less than the first number of pixels and connected to a subsequent stage of the first block, and a feature map; The method further includes a third block connected to a rear stage of the second block and relating to processing target data including an image or a feature map having a third number of pixels less than the second number of pixels. The image recognition device according to claim 5.

10. A neural network architecture search method in which a plurality of blocks corresponding to the number of pixels of processing target data are connected, the blocks executing a predetermined arithmetic process related to image feature extraction for processing target data including an image or feature map having a predetermined number of pixels, and outputting the result as output data having a predetermined number of dimensions, the method comprising: selecting an evaluation target block from the plurality of blocks; A step of selecting an operation specification that defines the operation processing to be executed by the block from among a plurality of candidates, and setting the operation specification as an evaluation candidate operation specification for the evaluation target block; calculating an evaluation value of an output result of the neural network including the evaluation target block in which the evaluation candidate operation specification is set; determining a calculation specification for the evaluation target block based on a comparison result between the evaluation value of the evaluation target block and the evaluation values ​​of other blocks; In the step of selecting a block to be evaluated, the block to be evaluated is selected from among the plurality of blocks in a sequence from a block having a smaller number of pixels of the data to be processed toward a block having a larger number of pixels. Architecture exploration methods.

11. Computer, An image recognition device that generates a neural network in which a plurality of blocks corresponding to the number of pixels of the processing target data are connected, the blocks executing a predetermined arithmetic process related to image feature extraction for processing target data including an image or feature map with a predetermined number of pixels and outputting the result as output data with a predetermined number of dimensions, an evaluation block selection unit that selects an evaluation target block from the plurality of blocks; a setting unit that selects an operation specification that defines the operation processing to be executed by the block from among a plurality of candidates, and sets the selected operation specification as an evaluation candidate operation specification in the evaluation target block; a block evaluation unit that calculates an evaluation value of an output result of the neural network including the evaluation target block in which the evaluation candidate calculation specification is set; a block operation specification determination unit that determines operation specifications for the evaluation target block based on a comparison result between the evaluation value of the evaluation target block and the evaluation values ​​of other blocks; The evaluation block selection unit functions as an image recognition device that selects evaluation target blocks from among the plurality of blocks, starting from blocks with a smaller number of pixels of the processing target data toward blocks with a larger number of pixels. program.

Citation Information

Patent Citations

  • Learning device, learning system, and learning method

    JP2021039640A

  • Integrated search device of artificial neural network and calculation accelerator structure, and method

    JP2023041579A

  • Neural network construction device, information processing device, neural network construction method, and program

    WO2019216404A1

  • Lightweight real-time facial alignment with one-shot neural architecture search

    WO2022184850A1

  • Compound model scaling for neural networks

    JP2023120204A