Learning device and method, inference device and method, program, and storage medium

By performing convolution processing in neural networks and updating parameters using weight operations, the problem of learning accuracy in Patent Document 1 is solved, achieving higher image processing quality and simplified calculations.

JP2025072174APending Publication Date: 2025-05-09CANON KK

Patent Information

Application Number
JP2023182751
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When using the binary neural network output structure described in Patent Document 1, learning accuracy may decrease due to simplification, resulting in poor quality of image processing.

Method used

By performing convolution processing in the learning device, the binary output image generated is used, weight operations and update parameters, to improve the learning accuracy and image quality of the neural network.

Benefits of technology

It realizes simplified neural network computing, while improving learning accuracy and image processing quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025072174000001_ABST
    Figure 2025072174000001_ABST
Patent Text Reader

Abstract

To simplify computations within a neural network for performing image processing and improve learning accuracy, thereby improving image quality as a result of the image processing.SOLUTION: Processing including convolution processing is performed on image data used for learning using a predetermined parameter to generate a plurality of output maps expressed in binary, each of the plurality of output maps generated is treated as image data representing one depth of bit depths, and a weighting computation is performed using a predetermined weighting value on bit strings at the same coordinates in the plurality of output maps. Then, the learning is performed by updating the parameter based on a comparison result obtained by comparing data obtained by the weighting computation with teacher data.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a learning device and method, an inference device and method, a program, and a storage medium. [Background technology]

[0002] In recent years, image processing devices that perform image processing using machine learning methods such as neural networks have been developed. However, in the learning of neural networks, image processing requires a larger number of patterns to be output by the neural network than in processing such as object detection, making the learning more complicated. Therefore, it has been considered to simplify the calculation by incorporating an output configuration that is binarized by the neural network. Patent Document 1 describes that the calculation is simplified and the circuit scale is reduced by outputting only sign bits, i.e., binary values, in the input layer, intermediate layer, or output layer of a neural network circuit. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2018-092377 A Summary of the Invention [Problem to be solved by the invention]

[0004] In the configuration described in Patent Document 1, the accuracy may deteriorate due to simplification of the learning.

[0005] The present invention has been made in consideration of the above problems, and aims to simplify the calculations within the neural network that performs image processing, increase the accuracy of learning, and improve the image quality as a result of image processing. [Means for solving the problem]

[0006] In order to achieve the above-mentioned object, the learning device of the present invention has a processing means that performs processing including convolution processing on image data used for learning using predetermined parameters, and generates a plurality of output maps expressed in binary, a calculation means that treats each of the plurality of output maps as image data representing one bit depth, and performs a weighting calculation using predetermined weighting values ​​on bit strings of the same coordinates in the plurality of output maps, and a learning means that updates the predetermined parameters based on a comparison result between the data obtained by the weighting calculation and teacher data. Effect of the Invention

[0007] According to the present invention, it is possible to simplify the calculations within the neural network that performs image processing, improve the learning accuracy, and improve the image quality as a result of image processing. [Brief description of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram showing a schematic functional configuration of a computer according to an embodiment of the present invention. [Diagram 2] 4A and 4B are diagrams showing an output map and weighting process of a neural network in the embodiment; [Diagram 3] 5 is a flowchart showing a learning process in the first embodiment. [Figure 4] 13 is a flowchart showing a process of updating a weight array in the second and third embodiments. [Diagram 5] 13 is a flowchart showing an inference process according to the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0010] ●Computer 10 Configuration The present invention will be described below by taking as an example a case where the present invention is implemented on a computer 10. Fig. 1 is a block diagram showing a computer 10. As shown in Fig. 1, the computer 10 includes a CPU 11, a memory 12, a storage device 13, a communication unit 14, a display unit 15, an input control unit 16, a GPU 17 (Graphics Processing Unit), and an internal bus 100.

[0011] The CPU 11 executes a computer program stored in the storage device 13 to control the operation of each part of the computer 10 via the internal bus 100. In addition, a GPU 17 that assists the operation of the CPU 11 may be used for calculations according to the computer program.

[0012] The memory 12 is a rewritable volatile memory. The memory 12 temporarily stores a computer program for controlling the operation of each part of the computer 10, information on each operation of the computer 10, information before and after processing by the CPU 11, and the like. The memory 12 has a sufficient storage capacity for temporarily storing these. Furthermore, the memory 12 stores a computer program describing the processing contents of the neural network, and learned coefficient parameters such as weighting coefficients and bias values. The weighting coefficient is a value for indicating the strength of the connection between nodes in the neural network, and the bias is a value for providing an offset to the integrated value of the weighting coefficient and input data.

[0013] The storage device 13 is an electrically erasable and recordable memory, and may be, for example, a hard disk or a solid state drive (SSD), etc. The storage device 13 stores computer programs that control each part of the computer 10, processing results temporarily stored in the memory 12, and other information.

[0014] The communication unit 14 communicates with peripheral devices such as external devices and recording media, and transmits information recorded in the storage device 13. As a communication method, a method conforming to a wireless communication standard such as USB (Universal Serial Bus) or IEEE802.11 is used, and the communication method is not particularly limited.

[0015] The display unit 15 may be, for example, a liquid crystal display, an organic EL display, or the like, and displays an image based on an image signal sent from the GPU 17 or the like. The GPU 17 processes image signals to be sent to the display unit 15, and performs calculations of computer programs, etc. Since the GPU 17 has a circuit capable of high-speed processing in parallel calculation processing, depending on the calculation, the GPU 17 performs the processing instead of the CPU 11.

[0016] The input control unit 16 controls input from an input device (not shown). Examples of the input device include a keyboard, a mouse, a touch panel, etc. The input control unit 16 converts the input from the input device into an electrical signal and transmits the input signal to each part of the computer 10. The above-mentioned components are connected to each other via an internal bus 100 so as to be accessible to each other.

[0017] <First embodiment> First, the learning of the neural network in the computer 10 will be described. Note that the neural network is, for example, a CNN (Convolutional Neural Network) or one configured with a fully connected layer, but the present invention is not limited to these. In addition, the present invention is not limited to learning by the computer 10. Learning refers to optimizing the processing by bringing coefficient parameters such as the weighting coefficients and bias values ​​of the neural network closer to appropriate values ​​for the target processing result.

[0018] In this embodiment, a neural network is used to perform image processing to improve image quality by, for example, removing noise information and distortion information contained in the image. Note that the uses of image processing are not limited to these.

[0019] In image processing using a neural network, learning is performed so that the difference between the calculation result obtained from the input learning data and the teacher data becomes small. For example, in a neural network that sharpens an image containing noise information, the teacher data is data that does not contain noise information, and the learning data is data that combines the teacher data and data that contains noise information, and learning is performed using these data.

[0020] When an image containing noise information is input as learning data, a calculation process is performed using this learning data and the coefficient parameters of the neural network. This process is repeated to obtain a calculation result. A loss value that indicates the difference (comparison result) between the output calculation result and the teacher data is calculated using a loss function, and learning is performed so that this loss value becomes smaller. In machine learning, when it is desired to learn image processing, which is the application target of this embodiment, a mean squared error (MSE) is generally used as a loss function that calculates the difference between the output of the neural network to be learned and the teacher data. The definition formula of MSE is shown in Equation (1). However, it is not necessary to be limited to MSE. In formula (1), x indicates the error of the data, and n indicates the number of data. In this case, the data corresponds to the pixels of the image.

[0021] By using the coefficient parameters obtained by learning in this way, an input image including noise information is arithmetically processed by a neural network, thereby improving the image quality, and thus image processing can be realized. Note that the present invention is not limited to supervised learning, and reinforcement learning, etc. may also be used.

[0022] Here, we will explain the relationship between the number of patterns output by CNN and the difficulty of learning, that is, the accuracy of CNN. Here, we will use ResNet, one of the most representative CNN models, as an example. However, the present invention is not limited to CNN models. ResNet is a CNN model that has shown revolutionary progress in accuracy from previous network configurations, and its developers have shown that the error rate of ImageNet Classification, a 1000-class classification, is 25.03% with ResNet-34. On the other hand, in 10-class classification by CIFAR-10, it has been shown that the error rate is 7.51% with ResNet-32, which has a network scale similar to ResNet-34 and a similar amount of calculation.

[0023] In class classification, the fewer the number of classifications, i.e., the number of output patterns, the lower the error rate and the more accurate the model. This means that the fewer the number of output patterns, the easier it is to learn, and if the scale of calculations for the CNN model is the same, the fewer the number of output patterns, the more accurate the model can be. This idea is also true for CNN models that perform image processing. Generally, class classification calculates one classification result for one input image, whereas image processing generates one output image for one input image through calculations. Therefore, in the case of image processing using CNN, the number of possible output patterns for each pixel of the output image is the number of bits in the image. As mentioned above, reducing the number of output patterns of CNN leads to the accuracy of the inference output, which is the learning result, so it is desirable to reduce the number of output patterns of CNN in image processing as well.

[0024] Next, the processing in the neural network of this embodiment will be described. In this embodiment, the neural network that performs image processing is controlled so as to output the same number of output maps as the bit depth of the teacher data or learning data.

[0025] As an example, the configuration of an output map of a neural network for teacher data and learning data as shown in Fig. 2(a) is shown here. In the example shown in Fig. 2(a), the horizontal data length W and vertical data length H of the teacher data, learning data, and each output map are the same. However, the data lengths of the teacher data, learning data, and output map of a neural network that performs image processing are not limited to this, and the data lengths of the teacher data and learning data may be different.

[0026] In this embodiment, each of the multiple output maps output from the neural network is expressed as a binary value such as "0" or "1". As a method for expressing the value of each element constituting each output map as a binary value, for example, a threshold may be used to assign each element a value of 1 if the value of the element is equal to or greater than the threshold, and a value of 0 if the value is less than the threshold. Furthermore, the binary expression is not limited to "0" or "1", and may be, for example, "-1" and "1". Then, as shown in FIG. 2(a), the multiple output maps are treated as data representing each of the most significant bit (MSB) to the least significant bit (LSB).

[0027] In this embodiment, as shown in FIG. 2(b), elements at the same coordinates of each output map are treated as one data array from the most significant bit (MSB) to the least significant bit (LSB), and the value obtained by multiplying and adding the data array and a weight array consisting of the same number of elements as the bit depth is set as the final calculation result of the neural network. The calculation result thus obtained is compared with the teacher data, and the MSE is calculated to proceed with learning. In this case, too, the greater the difference between the calculation result obtained from the multiple output maps and the weight array and the teacher data, the greater the calculated error amount MSE, and the smaller the difference, the smaller the calculated error amount MSE.

[0028] Here, the multiplication and accumulation operation using the output map and the weight array will be explained in more detail. FIG. 2(b) shows the output map of the neural network in which the multiple output maps shown in FIG. 2(a) are arranged in the depth direction, and indicates that each output map corresponds to the MSB to the LSB. Also, the frame surrounding the output map showing the MSB to the part of the output map showing the LSB in FIG. 2(b) indicates that the coordinate (0,0) of each output map is selected. The data array (bit string) surrounded by this frame represents one pixel of the same coordinate, and a multiplication and accumulation operation is performed between the data array representing this one pixel and the weight array. Note that the weight array length shown in FIG. 2(b) and the data array length representing one pixel correspond to the bit depth to be output, and have the same data length. Then, as a result of the multiplication and accumulation operation, the value of one pixel corresponding to the coordinate (0,0) is obtained. By performing this process for all the coordinates of the output map, the weighting operation result is obtained.

[0029] Here, we compare the number of output patterns when multiple output maps output from a neural network are each converted to binary values ​​and subjected to product-and-sum operations with a weight array as in this embodiment, with the number of output patterns when a single output map is output from a neural network. In the case shown in FIG. 2, as an example, the weight array is all 1 from the MSB to the LSB. For example, when a pixel with a bit depth of 10 bits is expressed as "1010101010", the operation result obtained by the multiplication and accumulation operation with the weight array is (1×1+0×1+1×1+0×1+1×1+0×1+1×1+0×1+1×1+0×1+0×1)=5. Thus, in the case of 10 bits, the possible values ​​that each pixel can take as a result of the multiplication and accumulation operation are 11 types from 0 to 10. Therefore, the number of output patterns that can be calculated by the multiplication and accumulation operation of the horizontal data length W of the output map and the vertical data length H of the output map is 11 to the power of (H×W).

[0030] In contrast, when one output map is output from a neural network as in the past, if the bit depth of the output map is 10 bits as in Fig. 2, each pixel can take a value from 0 to 1023 (2 to the power of 10). Therefore, the number of output patterns that can be generated as the output map is 2 to the power of (10 × H × W).

[0031] In this way, the output format of the neural network in the present invention makes it possible to significantly reduce the number of patterns output by the neural network compared to a normal output map configuration by binarizing each output map and then performing a product-sum operation using a weight array. Although the bit depth is described as 10 bits as an example here, the present invention is not limited to 10 bits.

[0032] Moreover, by making all the weight arrays from MSB to LSB 1, the number of output patterns is minimized. However, it is not necessary that all the values ​​constituting the weight array are equal. As an example, the case where the weights for some bits are made equal is shown in formula (2). In formula (2), A indicates the number of patterns that can be output in the output format using product-sum calculation in this embodiment, and B indicates the number of patterns that a normal neural network can output one output map as described above. In addition, in formula (2), the number of bits that make the weights equal in the weight array is M, the bit depth of the data output by the neural network is N, the horizontal data length of the output map is W, and the vertical data length of the output map is H. As can be seen from formula (2), in the output format of the neural network in this embodiment, when the number of bits M that make the weights equal in the weight array is greater than 1, the product of the number of output patterns that can be output and the number of output elements is reduced compared to a normal neural network. Note that when the number of bits M that make the weights equal in the weight array is 1, it indicates that there are no bits with equal weights, and the number of output patterns is the same as in the conventional method. TIFF2025072174000003.tif17117

[0033] In this way, according to the first embodiment, the number of output patterns that the neural network can output can be reduced compared to a normal neural network. This makes it possible to perform image processing with the same accuracy with a smaller learning amount than that of a conventional neural network. In other words, it is possible to generate a neural network that can obtain a more accurate image processing result with the same learning amount.

[0034] Next, the weight array in this embodiment will be described. Equation (3) shows an example of a weight array that performs a multiply-and-accumulate operation with a plurality of output maps output from a neural network. n-1 represents the weighting value for the output map representing the most significant bit, and w0 represents the weighting value for the output map representing the least significant bit. n-1 The number of elements in the weight array ~w0 corresponds to the number of output maps of the neural network, i.e., the bit depth you want to express as the output of the neural network. n-1 The magnitude relationship of the value of ~w0 is as shown in formula (3). When the values ​​of each element are different, w n-1 will be the maximum value in the weight array and w0 will be the minimum value. w n-1 ≧w n-2 ≧…≧w1≧w0…(3)

[0035] Also, the weight array w n-1 An example of the weight array w is shown in equation (4). n-1 By making ~w0 a power of 2, the process becomes the same as binary-to-decimal conversion, and it becomes possible to carry out learning within the same range of output values ​​as a normal neural network that outputs a single map as the calculation result. 2 n-1 =w n-1 ≧w n-2 ≧ … ≧ w1 ≧ w0 = 1 … (4)

[0036] Also, the weight array w n-1An example of ~w0 expressed in a non-binary relationship is shown in equation (5). In equation (5), β represents a value other than 2. By defining the weight array as in equation (5), it becomes possible to proceed with learning even in power relationships when the base is not 2. β n-1 =w n-1 ≧w n-2 ≧ … ≧ w1 ≧ w0 = 1 … (5)

[0037] In this way, the weight array may be defined as in formula (3), formula (4), or formula (5). Note that formula (4) and formula (5) do not limit formula (3).

[0038] Next, the learning process of this embodiment using the neural network in the CPU 11 will be described with reference to FIG. 3 is executed by loading a computer program stored in storage device 13 into memory 12 and having CPU 11 read out the computer program from memory 12. CPU 11 also loads learning data, teacher data, and weight arrays corresponding to each output map of the neural network, which are stored in storage device 13, into memory 12. The processing executed by CPU 11 may be performed using GPU 17.

[0039] In S301, the CPU 11 inputs a learning image (image data) to be subjected to image processing, which is learning data expanded in the memory 12, to the neural network. Next, in S302, the CPU 11 performs a calculation based on the learning image input to the neural network and the parameters of the neural network.

[0040] In S303, the CPU 11 performs weighting calculations (product-sum calculations) using the multiple output maps output from the neural network in S302 and the weight array expanded in the memory 12, as described with reference to FIG.

[0041] Next, in S304, the CPU 11 uses the result of the weighting calculation in S303 and the teaching data expanded in the memory 12 to perform calculations using a loss function to obtain a loss value. In S305, the CPU 11 uses the calculation result obtained in S304 to update the coefficient parameters of the neural network so that the loss value due to the loss function becomes smaller, and then ends this process.

[0042] One process shown in FIG. 3 corresponds to one neural network learning process, and this process is repeated until learning is completed.

[0043] As described above, according to the first embodiment, the output format based on the neural network's coefficient parameters is configured to output binary output maps with the same number as the bit depth of the teacher data or learning data, and the output multiple output maps are weighted with a predetermined weight array. Then, a loss function is performed using the result of the weighting calculation obtained, and learning is performed based on the obtained loss value. This reduces the number of output patterns that the neural network can output, making learning simpler than neural networks that perform normal image processing, and thus improving the accuracy of learning, and as a result, improving the accuracy of image processing.

[0044] <Second embodiment> Next, a second embodiment of the present invention will be described. In the above-described first embodiment, the output format based on the coefficient parameters of the neural network that performs image processing is configured to output a binary output map with the same number of bits as the bit depth of the teacher data or learning data, and the result of the weighting calculation based on a predetermined weight array is used to calculate a loss function, and the coefficient parameters are updated based on the obtained loss value.

[0045] In contrast, in the second embodiment, as learning progresses, not only the coefficient parameters but also the values ​​of each element of the weight array are updated, thereby making the calculation results between multiple output maps and the weight array closer to the training data.

[0046] In the second embodiment, the computer 10 described with reference to FIG. 1 in the first embodiment can be used, and therefore a description thereof will be omitted here. Also in the second embodiment, as an example, an output map of a neural network for teacher data and learning data having the configuration shown in Fig. 2 will be described. However, the data lengths of the teacher data, learning data, and output maps of the neural network performing image processing, and the number of output maps are not limited to these, and the data lengths of the teacher data and learning data may be different. Moreover, the flow of the learning process and the weighting process of the output map of the neural network in the second embodiment are similar to the process described in the first embodiment with reference to FIGS.

[0047] In the second embodiment, the weight array w n-1 The value of 〜w0 is updated as the learning progresses. This update process will be described with reference to FIG. 4. Note that this update process is performed separately from the learning of the neural network that performs image processing. The process shown in FIG. 4 is carried out by loading a computer program stored in storage device 13 into memory 12, and CPU 11 reading and executing the computer program from memory 12. Note that the process carried out by CPU 11 may be carried out using GPU 17.

[0048] In S401, the CPU 11 determines whether the number of times of learning by the learning process described with reference to Fig. 3 has reached a predetermined number of times for updating the weight array. The predetermined number of times may be determined in advance before learning, or may be determined from the convergence of learning. Note that the method of determining the predetermined number of times is not limited to these methods. If the number of times of learning has reached the predetermined number of times (Yes in S401), the CPU 11 proceeds with the process from S401 to S402, and if the number of times of learning has not reached the predetermined number of times (No in S401), the CPU 11 ends this process.

[0049] In S402, the CPU 11 calculates the weight array w n-1 It is determined whether or not ∼w0 is a target for updating. This determination may be made based on whether or not the loss value in the immediately preceding predetermined learning period is greater than a predetermined value. The predetermined learning period may be determined in advance before learning, or may be determined from the convergence of the learning. The predetermined value for determining the loss value may also be determined in advance before learning, or may be determined from the convergence of the loss value associated with learning. Note that the method for determining whether or not the weight array is a target for updating is not limited to this, and it may be set as to whether or not the weight array is a target for updating.

[0050] If the weight array is to be updated (Yes in S402), the CPU 11 advances the process to S403, and if it is not to be updated (No in S402), the CPU 11 ends this process. In steps S403 to S408, the CPU 11 performs a process of updating the weight array.

[0051] First, in S403, the CPU 11 inputs an evaluation target image during update processing expanded in the memory 12 to the neural network, and obtains a plurality of output maps as shown in Fig. 2. At this time, the evaluation target image may be one or more.

[0052] Next, in S404, the CPU 11 acquires a loss value based on the loss function before updating the weight array from the result of a product-sum operation obtained from the multiple output maps obtained in S403 and the weight array before the update, and the teacher data of the image to be evaluated.

[0053] Next, in S405, the CPU 11 updates the weight array.

[0054] In S406, the CPU 11 acquires a loss value based on the loss function after the weight array is updated, from the result of the product-sum operation obtained from the multiple output maps obtained in S403 and the updated weight array, and the teacher data of the image to be evaluated.

[0055] In S407, CPU 11 determines whether the updating of the weight array has been completed. Here, if the updating of the weight array has been performed a predetermined number of times, it is determined that the updating of the weight array has been completed. Note that the predetermined number of times to update the weight array may be determined in advance before updating the weight array, or may be a number less than the predetermined number of times. If the updating of the weight array has been completed (Yes in S407), CPU 11 proceeds to S408, and if not (No in S507), CPU 11 returns to S405 and repeats the above process until it is determined that the updating of the weight array has been completed.

[0056] In S408, CPU 11 determines the value of the weight array. Here, the weight array is determined from the loss value by the loss function in the weight array before update obtained in S404 and the loss value by the loss function in each weight array after update obtained in S406. When determining the weight array, the minimum value or the local minimum value of the obtained loss value is used, but is not limited to the minimum value or the local minimum value. When the values ​​of the weight array are finalized in S408, the weight array update process ends. The above process is repeated until the learning of the neural network is completed.

[0057] The processes in S403 to S407 may be realized by a neural network different from the neural network that performs the image processing. Also, the means for implementing S403 to S407 is not limited to a neural network.

[0058] As described above, according to the second embodiment, as the learning progresses, the values ​​of the weighting array that weights each output map of the neural network that performs image processing are also updated, and the weighting array is updated to a more appropriate value along with the learning, thereby further improving the accuracy of image processing by the neural network.

[0059] <Third embodiment> Next, a third embodiment of the present invention will be described. In the first and second embodiments, the output format of the neural network that performs image processing is configured to output the same number of output maps as the bit depth of the teacher data or learning data, and the loss value is calculated using the result weighted by a predetermined weight array to proceed with learning.

[0060] In the third embodiment, the weight array is made different for each scene, such as a portrait or a night scene. This makes it possible to further improve the image quality in each scene when image processing calculation (inference) is performed using a trained model in which learning is completed and the coefficient parameters of the neural network are optimized. Note that the scene is not limited to a portrait or a night scene.

[0061] In the third embodiment, the computer 10 described with reference to FIG. 1 in the first embodiment can be used, and therefore a description thereof will be omitted here. Also in the third embodiment, as an example, an output map of a neural network for the teacher data and learning data shown in Fig. 2 will be described. However, the data lengths and the number of output maps of the teacher data, learning data, and output map of the neural network performing image processing are not limited to these, and the data lengths of the teacher data and learning data may be different. The same applies to the weighting process of the output map of the neural network described in the first embodiment with reference to FIG.

[0062] In the second embodiment, the update process of the weight array that weights the output map of the neural network during learning was described with reference to Fig. 4, and the method of determining the weight array for each scene in the third embodiment is the same process except for a part of S403. Therefore, here, the differences in the process between the third embodiment and the second embodiment will be mainly described with reference to the flowchart in Fig. 4. Note that, in the third embodiment, too, a computer program stored in the storage device 13 is loaded into the memory 12, and the CPU 11 reads out and executes the computer program loaded into the memory 12. Note that the process performed by the CPU 11 may be performed using the GPU 17.

[0063] In S403, the CPU 11 inputs the evaluation target image expanded in the memory 12 to the neural network and obtains a plurality of output maps. At this time, the evaluation target image may be one or more. In this embodiment, in order to determine the weighting arrangement for each scene, when determining the weighting arrangement for a particular scene, the evaluation target image is limited to the particular scene and input to the neural network.

[0064] The processes other than S403 are similar to those described in the second embodiment, and therefore will not be described.

[0065] Also, in the third embodiment, the processes corresponding to S405 to S407 may be realized by a neural network other than the neural network that performs image processing. Also, the means for implementing S405 to S407 is not limited to a neural network. Then, the processes in S403 to S408 are executed for each scene, thereby obtaining a different weight array for each scene.

[0066] Of the various scenes, in the case of landscape scenes, the relationship between the elements of the weight array corresponding to the MSB to LSB shown in formula (3) is thought to tend to be smaller than 2 when considered as a base number, compared to a state in which the base number shown in formula (4) can be expressed as a power of 2, where the base number is 2. This is because in landscape scenes, which tend to be complex images, the base number of the elements of the weight array is thought to be smaller than 2. On the other hand, in the case of night scenes with large differences in brightness, the relationship between the elements of the weight array is thought to tend to be a weight array with a base number greater than 2.

[0067] Next, the inference processing in the CPU 11 or the inference device will be described with reference to Fig. 5. In addition, the inference processing shown in Fig. 5 is performed by expanding a computer program stored in the storage device 13 into the memory 12, and the CPU 11 reads out and executes the computer program expanded into the memory 12. Note that the processing performed by the CPU 11 may be performed using the GPU 17. In addition, when the inference processing is performed by an inference device (not shown), the processing by the CPU 11 is performed by a neural network processing unit, and the memory 12 is performed by a recording unit.

[0068] During inference, a weighting array is selected according to the scene of the target image input to the neural network, and the output map of the neural network is weighted with a weighting array that differs for each scene, and an inference image that is the result of the inference process is output. In this embodiment, image processing is performed, so an image that has been subjected to image processing is output.

[0069] First, in S601, the CPU 11 inputs an image to be subjected to image processing, which is expanded in the memory 12, to the neural network.

[0070] Next, in S602, the CPU 11 performs calculations on the image input to the neural network using parameters of the neural network.

[0071] In S603, the CPU 11 determines the scene of the image input to the neural network, and selects a weighting array corresponding to the determined scene from among weighting arrays expanded in the memory 12 and corresponding to various scenes.

[0072] In S604, the CPU 11 performs a weighting operation (product-sum operation) using each output map output by the neural network and the weight array selected in S603. This operation process is the same as the weighting process of the output map of the neural network described in the first embodiment with reference to Fig. 2. After performing S604, the process ends.

[0073] One process shown in FIG. 6 corresponds to one inference process of the neural network, and this process is repeated if there are other images to be inferred.

[0074] As described above, according to the third embodiment, in a neural network that performs image processing, a different weight array is determined for each scene during learning, and an output map is weighted using a weight array according to the scene during inference, which can improve the accuracy of inference compared to the case where the same weight array is used for different scenes.

[0075] <Other embodiments> The present invention may be applied to a system made up of a plurality of devices, or to an apparatus made up of a single device.

[0076] The present invention can also be realized by supplying a program for implementing one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0077] <Summary> The disclosure of this embodiment includes the following configuration.

[0078] (Item 1) A processing means for performing processing including convolution processing on image data used for learning using predetermined parameters to generate a plurality of output maps expressed in binary form; A calculation means for performing a weighting calculation on bit strings at the same coordinates in the plurality of output maps by using a predetermined weighting value, with each of the plurality of output maps being image data representing one bit depth; a learning means for updating the predetermined parameters based on a comparison result obtained by comparing the data obtained by the weighting calculation with teacher data; A learning device comprising: (Item 2) 2. The learning device according to item 1, wherein each of the plurality of output maps has data having at least the same horizontal data length and vertical data length as the image data or the teacher data. (Item 3) 3. The learning device according to item 1 or 2, wherein the weighting operation is a product-sum operation. (Item 4) 4. The learning device according to any one of items 1 to 3, wherein the weighting value is a power of 2 or a predetermined number. (Item 5) 5. The learning device according to any one of items 1 to 4, wherein, among the weighting values, a value corresponding to a most significant bit is equal to or greater than a value corresponding to a least significant bit. (Item 6) The method further includes: changing the weighting value based on the comparison result; 6. The learning device according to any one of items 1 to 5, wherein the calculation means uses the weighting value changed by the change means. (Item 7) 6. The learning device according to any one of items 1 to 5, wherein the weighting values ​​are set to different values ​​depending on the scene of the image data used for the learning. (Item 8) A storage unit that stores parameters updated by the learning unit in the learning device according to any one of items 1 to 6; An acquisition means for acquiring image data to be used for inference; a processing means for performing processing including a convolution process on the image data using the parameters stored in the storage means, and generating a plurality of output maps expressed in binary form; A calculation means for performing a weighting calculation using the weighting value on bit strings at the same coordinates in the plurality of output maps, with each of the plurality of output maps being image data representing one bit depth; An inference device comprising: (Item 9) A storage means for storing parameters updated by the learning means in the learning device according to item 7 and weighting values ​​that vary depending on the scene; An acquisition means for acquiring image data to be used for inference; a processing means for performing processing including a convolution process on the image data using the parameters stored in the storage means, and generating a plurality of output maps expressed in binary form; A discrimination means for discriminating a scene of image data used for the inference; a calculation means for performing a weighting calculation on bit strings at the same coordinates in the plurality of output maps by using a weighting value corresponding to the determined scene among the weighting values ​​stored in the storage means, with each of the plurality of output maps being image data representing one bit depth; An inference device comprising: (Item 10) A processing step of performing processing including convolution processing on image data used for learning using predetermined parameters to generate a plurality of output maps expressed in binary form; a calculation step of performing a weighting calculation on bit strings at the same coordinates in the plurality of output maps by using a predetermined weighting value, with each of the plurality of output maps being image data representing one bit depth; a learning process for updating the predetermined parameters based on a comparison result between the data obtained by the weighting calculation and teacher data; A learning method comprising: (Item 11) A storage step of storing the parameters updated by the learning means in the learning device according to any one of items 1 to 6 in a storage means; An acquisition step of acquiring image data to be used for inference; a processing step of performing processing including a convolution process on the image data using the parameters stored in the storage means to generate a plurality of output maps expressed in binary; a calculation step of performing a weighting calculation using the weighting value on bit strings at the same coordinates in the plurality of output maps, with each of the plurality of output maps being image data representing one bit depth; 13. An inference method comprising: (Item 12) A storage step of storing the parameters updated by the learning means in the learning device according to item 7 and weighting values ​​that vary depending on the scene in a storage means; An acquisition step of acquiring image data to be used for inference; a processing step of performing processing including a convolution process on the image data using the parameters stored in the storage means to generate a plurality of output maps expressed in binary; a discrimination step of discriminating a scene of image data used for the inference; a calculation step of performing a weighting calculation on bit strings at the same coordinates in the plurality of output maps by using a weighting value corresponding to the determined scene among the weighting values ​​stored in the storage means, with each of the plurality of output maps being image data representing one bit depth; 13. An inference method comprising: (Item 13) A program for causing a computer to function as each of the means of the learning device described in any one of items 1 to 7. (Item 14) A program for causing a computer to function as each of the means of the inference device described in item 8 or 9. (Item 15) 15. A computer-readable storage medium storing the program according to item 13 or 14.

[0079] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0080] 10: Computer, 11: CPU, 12: Memory, 13: Storage device, 14: Communication unit, 15: Display unit, 16: Input control unit, 17: GPU, 100: Internal bus

Claims

1. A processing means for performing processing including convolution processing on image data used for learning using predetermined parameters to generate a plurality of output maps expressed in binary form; a calculation means for performing a weighting calculation on bit strings at the same coordinates in the plurality of output maps by using a predetermined weighting value, with each of the plurality of output maps being image data representing one bit depth; a learning means for updating the predetermined parameters based on a comparison result obtained by comparing the data obtained by the weighting calculation with teacher data; A learning device comprising:

2. 2. The learning device according to claim 1, wherein each of the plurality of output maps has data having at least the same horizontal data length and vertical data length as the image data or the teacher data.

3. 2. The learning device according to claim 1, wherein the weighting operation is a multiply-and-accumulate operation.

4. 2. The learning device according to claim 1, wherein the weighting value is a power of 2 or a predetermined number.

5. 2. The learning device according to claim 1, wherein the weighting value has a value corresponding to the most significant bit equal to or greater than a value corresponding to the least significant bit.

6. The method further comprises: changing means for changing the weighting value based on the comparison result; 2. The learning device according to claim 1, wherein said calculation means uses the weighting values ​​changed by said change means.

7. 2. The learning device according to claim 1, wherein the weighting values ​​are set to different values ​​depending on the scene of the image data used for the learning.

8. A storage means for storing parameters updated by the learning means in the learning device according to any one of claims 1 to 6; An acquisition means for acquiring image data to be used for inference; a processing means for performing processing including a convolution process on the image data using the parameters stored in the storage means, and generating a plurality of output maps expressed in binary form; a calculation means for performing a weighting calculation using the weighting value on bit strings at the same coordinates in the output maps, with each of the output maps being image data representing one bit depth; An inference device comprising:

9. a storage means for storing parameters updated by said learning means in the learning device according to claim 7 and weighting values ​​which vary depending on the scene; An acquisition means for acquiring image data to be used for inference; a processing means for performing processing including a convolution process on the image data using the parameters stored in the storage means, and generating a plurality of output maps expressed in binary form; A discrimination means for discriminating a scene of image data used for the inference; a calculation means for performing a weighting calculation on bit strings at the same coordinates in the plurality of output maps by using a weighting value corresponding to the determined scene among the weighting values ​​stored in the storage means, with each of the plurality of output maps being image data representing one bit depth; An inference device comprising:

10. A processing step of performing processing including a convolution process on image data used for learning using predetermined parameters to generate a plurality of output maps expressed in binary form; a calculation step of performing a weighting calculation on bit strings at the same coordinates in the plurality of output maps by using a predetermined weighting value, with each of the plurality of output maps being image data representing one bit depth; a learning process for updating the predetermined parameters based on a comparison result between the data obtained by the weighting calculation and teacher data; A learning method comprising:

11. a storage step of storing the parameters updated by the learning means in a storage means in the learning device according to any one of claims 1 to 6; An acquisition step of acquiring image data to be used for inference; a processing step of performing processing including a convolution process on the image data using the parameters stored in the storage means to generate a plurality of output maps expressed in binary; a calculation step of performing a weighting calculation using the weighting value on bit strings at the same coordinates in the plurality of output maps, with each of the plurality of output maps being image data representing one bit depth; 13. An inference method comprising:

12. a storage step of storing the parameters updated by the learning means and weighting values ​​that vary depending on the scene in a storage means in the learning device according to claim 7; An acquisition step of acquiring image data to be used for inference; a processing step of performing processing including a convolution process on the image data using the parameters stored in the storage means to generate a plurality of output maps expressed in binary; a discrimination step of discriminating a scene of image data used for the inference; a calculation step of performing a weighting calculation on bit strings at the same coordinates in the plurality of output maps by using a weighting value corresponding to the determined scene among the weighting values ​​stored in the storage means, with each of the plurality of output maps being image data representing one bit depth; 13. An inference method comprising:

13. A program for causing a computer to function as each of the means of the learning device according to any one of claims 1 to 7.

14. A program for causing a computer to function as each of the means of the inference device according to claim 8.

15. A program for causing a computer to function as each of the means of the inference device according to claim 9.

16. A computer-readable storage medium storing the program according to claim 13.

17. A computer-readable storage medium storing the program according to claim 14.

18. A computer-readable storage medium storing the program according to claim 15.

Citation Information

Patent Citations

  • Neural network circuit device, neural network, neural network processing method and neural network executing program

    JP2018092377A

Cited By

  • Game machine

    JP2025105806A

  • Game machine

    JP2025105808A