Image processing apparatus, image capturing apparatus, control method and program

By eliminating the repeated calculation of cross pixel areas in the CNN algorithm in the control unit of the image processing device, the problems of poor edge processing after image segmentation and increased calculation time are solved, and efficient convolution operation and accurate image processing results are achieved.

JP2025072175APending Publication Date: 2025-05-09CANON KK

Patent Information

Application Number
JP2023182752
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When using the CNN algorithm for image processing, due to the buffer size, the image needs to be divided into small tiles for processing, resulting in poor processing results at the edge of the image, and the calculation time is increased due to the need to include crossed pixel areas.

Method used

An image processing device is designed to perform convolution operations on only the unique pixel areas in each tile by excluding repeated calculations of crossed pixel areas in the control unit, thereby improving processing efficiency.

Benefits of technology

It realizes efficient convolution operation under the finite buffer size, which reduces the calculation time, and the processing results are the same as those of the unsegmented images, maintaining accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025072175000001_ABST
    Figure 2025072175000001_ABST
Patent Text Reader

Abstract

To efficiently execute convolutional computation processing with respect to an image.SOLUTION: An image processing apparatus includes: computation means for executing convolutional computation processing in a neural network with respect to input data; obtainment means for obtaining a plurality of tiles that corresponds to respective partial regions in an image; and control means for performing control so as to cause the computation means to execute the convolutional computation processing while using each of the plurality of tiles as the input data. The control means controls the computation means so that, with respect to some of the plurality of tiles, overlapping pixels included in the tiles and corresponding to the same region in the image as another tile are excluded from a target of the convolutional computation processing.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image processing device, an imaging device, a control method and a program, and more particularly to an image processing technique that utilizes a neural network. [Background technology]

[0002] In recent years, deep learning technology using neural networks has been used in a wide range of technical fields. In particular, convolutional neural networks (CNNs) are widely used in the field of image processing. Convolutional neural networks can achieve highly accurate learning by using image features obtained by recursively performing convolution operations (Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2019-125128 A Summary of the Invention [Problem to be solved by the invention]

[0004] Meanwhile, deep learning processing using CNN is in wide demand, and in recent years, it has become possible to execute it in real time not only on servers with a certain computing power but also on various edge devices. For example, hardware specialized for CNN (hereinafter referred to as CNN computing device) is also available in the form of AIChip and IP, and by incorporating this hardware, it is possible to execute deep learning processing on edge devices.

[0005] The CNN calculator has a buffer for holding the data to be calculated, but the size of the buffer is finite. In recent years, the number of pixels in image sensors has increased, and the data size of images handled by digital cameras and the like has increased. In order to process an image with the CNN calculator, the image must be divided into tiles smaller than the buffer size, and the CNN calculator must process the image on a tile-by-tile basis.

[0006] On the other hand, if an image is simply divided into regions to form tiles, suitable processing results may not be obtained at the edge of the boundary between tiles. This is because a convolution operation is performed for one pixel of the image by referring to the pixel and the pixels distributed around it. In other words, pixels that are not included in the tile, i.e., pixels that are not held in the buffer of the CNN calculator, cannot be referenced, so the processing results of the convolution operation differ from those obtained without dividing the image into tiles.

[0007] Therefore, in order to obtain a suitable processing result for the entire image, each tile input to the CNN calculator must be configured to include not only an area into which the image is simply divided, but also its surrounding pixels. In other words, each of the tiles that are input sequentially to the CNN calculator and subjected to the convolution calculation is configured to include pixel areas that overlap with other tiles.

[0008] However, forming tiles that include overlapping pixel areas in this way results in a longer processing time than when a convolution operation is performed on an image without division.

[0009] The present invention has been made in consideration of the above-mentioned problems, and has an object to provide an image processing device, an imaging device, a control method, and a program that efficiently execute convolution calculation processing on an image. [Means for solving the problem]

[0010] In order to achieve the above-mentioned object, the image processing device of the present invention comprises a calculation means for performing a convolution calculation process in a neural network on input data, an acquisition means for acquiring a plurality of tiles, each of which corresponds to a partial area in an image, and a control means for controlling the calculation means to perform a convolution calculation process using each of the plurality of tiles as input data, and is characterized in that the control means controls the calculation means so that, for at least a portion of the plurality of tiles, overlapping pixels contained in the tile and having the same corresponding area in the image as other tiles are excluded from the convolution calculation process. Effect of the Invention

[0011] With this configuration, the present invention makes it possible to efficiently execute convolution processing on an image. [Brief description of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram showing an example of a hardware configuration of an image processing device 100 according to an embodiment and a modification of the present invention. [Diagram 2] FIG. 1 is a diagram for explaining a plurality of tiles generated from a processing target image in the CNN calculation according to the embodiment and the modified example of the present invention; [Diagram 3] FIG. 1 is a diagram for explaining a convolution calculation process executed by a product-sum calculation processing unit 103 according to an embodiment and a modification of the present invention; [Figure 4] FIG. 1 is a diagram illustrating an example of memory space allocation of data related to CNN calculation according to an embodiment and a modification of the present invention; [Diagram 5] 1 is a flowchart illustrating a hierarchical calculation process executed in the image processing device 100 according to an embodiment and a modification of the present invention. [Figure 6] 1 is a flowchart illustrating a generation process executed in an image processing device 100 according to an embodiment and a modification of the present invention. [Figure 7] FIG. 11 is a diagram for explaining the receptive field of CNN calculation according to the second embodiment of the present invention; [Figure 8]FIG. 11 is a diagram for explaining a plurality of tiles generated from a processing target image in the CNN calculation according to the third embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] [Embodiment 1] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0014] In the embodiment described below, an example of the present invention is applied to an image processing device configured to be able to execute convolution processing in CNN on a captured image, as an example of an image processing device. However, the present invention is applicable to any device capable of executing image processing including convolution processing on an image input as multiple tiles.

[0015] Configuration of the image processing device 1 is a block diagram showing a hardware configuration of an image processing device 100 according to this embodiment. The image processing device 100 is a device configured to be able to execute image processing using a neural network used in deep learning or the like. In this embodiment, the image processing device 100 is configured to be able to execute various operations related to a neural network on a captured image. In the following, it will be described that the image processing device 100 is capable of executing image processing including various operations related to a convolutional neural network (hereinafter referred to as CNN operations) that is one aspect of a neural network and is mainly targeted at images.

[0016] The CNN calculation unit 101 executes CNN calculation in the image processing device 100. As shown in the figure, in this embodiment, the CNN calculation unit 101 includes a CPU 102, a product-sum calculation processing unit 103, and a shared memory 105.

[0017] The CPU 102 is a control device that controls the operation of each block of the image processing device 100. The CPU 102 is equipped with, for example, an internal ROM and RAM, and can control the operation of each block by reading out the operation program of each block from the ROM, expanding it in the RAM, and executing it. Although details will be described later, the CPU 102 controls the operation by supplying various parameters required for executing CNN calculation to each block. Note that, in the example of FIG. 1, the CPU 102 is included in the CNN calculation unit 101, but it goes without saying that the CPU 102 may be provided outside the CNN calculation unit 101.

[0018] The multiply-accumulate processor 103 executes convolution processing, which is the core of CNN processing. The multiply-accumulate processor 103 includes a plurality of MACs 104, which are multiply-accumulation (MAC) cores. In the convolution processing, the MACs 104 are controlled to repeatedly perform MAC calculations on input data input to the multiply-accumulate processor 103.

[0019] The shared memory 105 is a storage device configured to be accessible by the CPU 102, the product-sum calculation processing unit 103, and the interconnect 106. The shared memory 105 can store, for example, parameters used in the convolution calculation process and the results of the convolution calculation process.

[0020] The interconnect 106 is an interface that realizes data communication inside and outside the CNN calculation unit 101. More specifically, the interconnect 106 realizes interconnection between the CPU 102, the product-sum calculation processing unit 103, the shared memory 105, the tile generation unit 107, and the external memory 108, and performs data communication based on a predetermined protocol. In this embodiment, the data communication between the blocks is described as being realized via the interconnect 106, but it should be understood that the implementation of the present invention is not limited to this.

[0021] The tile generation unit 107 generates data of tiles to be subjected to the convolution operation. The product-sum operation processing unit 103 of this embodiment does not perform the convolution operation on the entire region of the image to be subjected to the CNN operation at once, but performs the convolution operation on each tile extracted from the image. That is, the tiles generated by the tile generation unit 107 are processed, and the CNN operation in the CNN operation unit 101 of this embodiment is performed on a tile-by-tile basis generated by the tile generation unit 107. The tile generation unit 107 generates data of tiles by extracting pixels included in each partial region for a plurality of partial regions set for the image.

[0022] The external memory 108 is a storage device that stores an image to be subjected to CNN calculation (sometimes referred to as a processing target image) and parameters used in the convolution calculation process. In general, the external memory 108 is configured as a storage device that is slower and has a larger capacity than the shared memory 105.

[0023] <Tile Generation> Next, tile generation based on the processing target image by the tile generation unit 107 will be described with reference to Fig. 2. Note that in this embodiment, the tile generation unit 107 will be described as generating four tiles based on the processing target image.

[0024] Fig. 2(a) illustrates an example of dividing an image to be processed into four tiles by dividing the image to be processed into two equal parts horizontally and vertically. That is, the example illustrated in Fig. 2(a) illustrates an example of generating tiles by simply dividing the image to be processed into regions. In the figure, the four tiles are identified by labeling the tile located at the top left of the image to be processed as "t0", the tile located at the top right as "t1", the tile located at the bottom left as "t2", and the tile located at the bottom right as "t3".

[0025] Although details will be described later, in the convolution operation performed in the product-sum operation processing unit 103, for each pixel of the input data, a filter process is performed that refers to the pixel values ​​of the pixel and the surrounding pixels arranged around the pixel. In this type of filter process, for pixels located at the edge of an image, the surrounding pixels are not included in the image, so the operation is performed with the pixel values ​​of non-existent surrounding pixels set to 0. Therefore, even if the image to be processed is divided into tiles in the manner shown in Fig. 2(a) and a convolution operation is performed for each tile, pixels not included in the tile will be referenced in the operation as shown in Fig. 2(b).

[0026] Fig. 2(b) illustrates a pixel group 202 that is referenced when performing a filter process based on a 3 x 3 filter kernel on pixel 201 at the top left corner of tile "t3" in Fig. 2(a). As illustrated, five surrounding pixels of pixel group 202 are not included in the tile t3, so in the filter process for pixel 201, the pixel values ​​of these surrounding pixels are replaced with 0, etc., and calculations are performed.

[0027] On the other hand, when a filter process is performed on the entire image to be processed without dividing it into tiles, since there are pixels around the pixel 201, the significant pixel values ​​of those pixels are referred to in the filter process to obtain the calculation result. Therefore, when tiles are generated in the manner shown in FIG. 2(a) and a convolution calculation process is performed on each tile, the calculation result obtained for the pixels distributed at the edge of the tile is different from that obtained when the calculation is performed without dividing the tile. In other words, in the manner shown in FIG. 2(a), the pixel values ​​of the pixels that could have been originally referred to in the convolution calculation process cannot be used for the calculation for some of the tiles, so that a suitable calculation result cannot be obtained. As a result, the accuracy of the image processing performed by the CNN calculation unit 101 by dividing the image to be processed into tiles in the manner shown in FIG. 2(a) is lower than that when the image processing is performed without dividing the tile.

[0028] For this reason, in the image processing device 100 of this embodiment, the tile generating unit 107 generates four tiles based on the image to be processed in a manner including overlapping areas between the tiles as shown by hatching in FIG. 2(c). More specifically, the tile generating unit 107 sets four partial areas related to the tiles to be generated for the image to be processed so that they overlap at least with adjacent partial areas, and generates each tile by extracting pixels included in each partial area. The solid lines shown in the image to be processed in FIG. 2(c) are shown for comparison with the embodiment in FIG. 2(a), and indicate a line that bisects the image to be processed in the horizontal direction and a line that bisects the image to be processed in the vertical direction. The ends of each tile are shown by dashed lines in FIG. 2(c), and when each tile is separated, it becomes an embodiment as shown in FIG. 2(d). That is, the tiles generated by the tile generating unit 107 of this embodiment have an area larger than the tiles that simply divide the image to be processed into four equal parts as shown in FIG. 2(a). Furthermore, the tiles generated by the tile generating unit 107 include pixels that are not included in tiles (tiles with the same label) that have the same relative position in the image to be processed in the mode of FIG. 2(a).

[0029] In other words, the tiles generated by the tile generating unit 107 of this embodiment correspond to four partial regions set in the image to be processed, each of which includes at least a region overlapping with other adjacent partial regions. In the following description, pixels included in a region where adjacently set partial regions overlap will be referred to as overlapping pixels. That is, in the aspects of FIGS. 2(c) and (d), the tile "t0" includes at least overlapping pixels with the tile "t1" and overlapping pixels with the tile "t2". The tile "t1" also includes at least overlapping pixels with the tile "t0" and overlapping pixels with the tile "t3". The tile "t2" also includes at least overlapping pixels with the tile "t0" and overlapping pixels with the tile "t3". Similarly, the tile "t3" includes at least overlapping pixels with the tile "t1" and overlapping pixels with the tile "t2".

[0030] FIG. 2(e) illustrates a tile with a label "t3" generated based on the partial region set in the mode of FIG. 2(c). As illustrated, the tile "t3" generated by the tile generating unit 107 of this embodiment includes all of the peripheral pixels of the pixel 201 located at the upper left corner of the tile t3 in the mode of FIG. 2(a). Therefore, the calculation result of the filter process for the pixel 201 is the same as the calculation result when the filter process is performed on the same pixel without dividing the processing target image. As a result, the processing accuracy of the image processing performed by the CNN calculation unit 101 by generating the tile of the mode of FIG. 2(c) from the processing target image is not reduced compared to the case where the tile is not divided.

[0031] CNN Calculation One aspect of neural networks used for image recognition and the like is a convolutional neural network. The CNN calculation unit 101 can execute CNN calculations that mainly include convolution calculation processing in a convolutional neural network. By executing the convolution calculation processing on an image, for example, it is possible to derive the feature amount of the image. The feature amount derived in this way can be used for deep learning and various image analyses.

[0032] The CNN computation executed in the CNN computation unit 101 will be outlined below with reference to FIG.

[0033] As described above, in the convolution operation included in the CNN operation, a filter process is performed. In this embodiment, a filter kernel having a kernel size of 3 pixels x 3 pixels is used for the filter process. Here, the product-sum operation processing unit 103 of this embodiment is configured not to change the size of the image between the input data (input image) and the output data (output image). Therefore, the product-sum operation processing unit 103 adds 0 padding of one pixel width to each of the top, bottom, left and right of the input image prior to the filter process. That is, the pixels of each tile generated by the tile generation unit 107 are added with 0 padding by adding a horizontal line of one pixel to each of the top and bottom of the tile and a vertical line of one pixel to each of the left and right. That is, the convolution operation process in the product-sum operation processing unit 103 is performed on an image obtained by further adding pixels related to 0 padding to the pixel group of the tile including overlapping pixels, which is the input image, as the subject of the convolution operation.

[0034] The filter process related to the convolution operation is performed for each position while moving a filter kernel of a predetermined size from the top left of the input image in raster order. The filter kernel specifies the pixel group to be referenced in the filter process, and the filter process obtains an operation result corresponding to the pixel located at the center of the filter kernel. In other words, the filter process related to the convolution operation performs an operation by referring to the pixel values ​​of pixels included in an area of ​​the kernel size (3×3) centered on the pixel of interest while changing the position of the pixel of interest in raster order, and obtains an output value (operation result) related to the pixel of interest.

[0035] Here, in the convolution operation process for the pixel of interest x, if the filter coefficient w of the filter kernel is as shown in FIG. 3(a), the operation result (output value Out(x)) can be derived by the following equation (1).

[0036] Out(x)=α(w00 x00+w01 x01+w02 x02 +w10 x10+w11 x11+w12 x12 +w20 x20+w21 x21+w22 x22 +b) (1) Here, x indicates the pixel value of the image to be subjected to the convolution operation, and the subscript number corresponds to the position of the filter coefficient with the same subscript. That is, the pixel value of the pixel of interest x is x11, which is multiplied by the filter coefficient w11 corresponding to the center of the filter kernel in formula (1). Also, a is an activation function, and for example, Relu (Rectified Linear Unit) can be used. Also, b is a bias value. Hereinafter, the filter coefficient and the bias value are collectively referred to as "model parameters". The model parameters are parameters used in the convolution operation process, and are stored in advance in, for example, the external memory 108, and are read out by the CPU 102 and applied to the product-sum operation processing unit 103. The CNN operation unit 101 can be used for deep learning applications by selecting the model parameters.

[0037] 3(b), in the CNN calculation performed in the CNN calculation unit 101 of this embodiment, the convolution calculation process is repeated multiple times on the input image I0. More specifically, the CNN calculation repeatedly performed as shown in the figure is performed hierarchically in such a manner that the calculation result (output data) of the convolution calculation process is used as input data for the next convolution calculation process a predetermined number of times.

[0038] In the example of FIG. 3(b), three stages of convolution operation processing (CNN0, CNN1, CNN2) are performed on the input image I0 to obtain the operation result (output O2) of the CNN operation. Here, CNN0, CNN1, and CNN2 in the CNN operation are called layers (hierarchies), and correspond to the convolution layers in the convolutional neural network. That is, in the CNN operation, the output data of the previous layer is used as the input data of the next layer. That is, the output O0 of CNN0 becomes the input I1 of CNN1, and the output O1 of CNN2 becomes the input I2 of CNN2. In addition, CNN0 may be referred to as the input layer, meaning that the image to be processed is input, and CNN1 and CNN2 may be referred to as the subsequent layers, meaning that the output data of the previous layer is used. In addition, the data between each hierarchy may be referred to as an intermediate feature image.

[0039] Therefore, the CPU 102 instructs the product-sum calculation processing unit 103 on the hierarchical relationship and the model parameters to be used according to the contents of the image processing to be performed as the CNN calculation. Based on the instruction, the product-sum calculation processing unit 103 executes the convolution calculation process using the MAC 104 and stores the output data of the calculation result in the shared memory 105, and the CPU 102 performs activation and the like to realize a series of CNN calculations.

[0040] <<Processing for overlapping pixels>> Incidentally, in a mode in which each of the multiple tiles generated from the processing target image by the tile generation unit 107 includes overlapping pixels with other tiles as described above, the processing time is longer by the amount of overlapping pixels than when the processing target image is not divided into tiles and CNN operation is performed. That is, the total number of pixels of the multiple tiles generated by the tile generation unit 107 is greater than the total number of pixels of the processing target image by the amount of overlapping pixels of each tile, so the number of times that the convolution operation process is performed is increased accordingly, resulting in a longer processing time for the CNN operation.

[0041] In the CNN calculation unit 101 of this embodiment, in order to reduce such an increase in processing time, a process is performed to limit pixels for which convolution calculation processing is to be performed by the product-sum calculation processing unit 103. More specifically, for at least some of the multiple tiles, the CPU 102 excludes overlapping pixels for which convolution calculation processing has already been performed in other tiles, and sets these as input to the product-sum calculation processing unit 103.

[0042] 4 illustrates an example of the arrangement of data of each tile used in the convolution operation in the product-sum operation processing unit 103 in the memory space 400. Here, in order to clearly identify various data, hereinafter, a tile identifies the position of a partial area set for the image to be processed, and may be described using a label (t0, t1, t2, t3) assigned to the tile as necessary. Furthermore, "tile input data" refers to data input to the product-sum operation processing unit 103 as a target for the convolution operation (data referenced by the product-sum operation processing unit 103 from the shared memory 105). On the other hand, "tile output data" refers to data output as the operation result of the convolution operation executed for the tile.

[0043] The memory space 400 is, for example, the shared memory 105, and stores data used in the CNN calculation process. The memory space 400 stores input data 401 of a plurality of tiles and output data 411, which are the results of the respective convolution calculation processes, for the convolution calculation process performed by the product-sum calculation processing unit 103 in one layer.

[0044] To facilitate understanding of the invention, in the example of Fig. 4, the memory space 400 has a width equal to the number of horizontal pixels of a tile image generated by the tile generation unit 107. In the example of Fig. 4, four tiles of input data (I0_t0, I0_t1, I0_t2, I0_t3) are arranged as input data 401 related to the input layer (first layer CNN0). The CPU 102 reads the tile input data from the memory space 400 in the order of t0, t1, t2, t3, and inputs it to the product-sum calculation processing unit 103 to perform convolution calculation processing.

[0045] Here, when the product-sum calculation processor 103 performs convolution calculation processing on all pixels included in the input data of each tile, output data 402 of four tiles is stored in memory space 400 as shown in Fig. 4(a). O0_t0 indicates output data of t0, O0_t1 indicates output data of t1, O0_t2 indicates output data of t2, and O0_t3 indicates output data of t3. These output data are input to CNN1 in the next layer, and therefore are also treated as input data of the four tiles related to the next layer (I1_t0, I1_t1, I1_t2, I1_t3).

[0046] However, the input data of tiles related to the input layer has corresponding partial areas that overlap, and the same calculation is performed on the overlapping pixels included in the overlapping areas in each of the convolution calculation processes for the multiple tiles, and the same calculation result may be output. More specifically, if the overlapping pixels and all of their surrounding pixels are included in the image of the tile, the calculation result of the convolution calculation process will not change. Therefore, when the product-sum calculation processing unit 103 performs a convolution calculation process on the overlapping pixels included in any tile, the output data of another tile that has previously performed a convolution calculation process may already contain the calculation result for the overlapping pixels. In other words, the output data of multiple tiles may contain the same calculation result for the overlapping pixels.

[0047] For example, as shown in FIG. 2(c), the tile t3 has overlapping pixels because the partial regions overlap with the tile t1 distributed above and with the tile t2 distributed to the left. That is, when the input data I0_t3 of t3 is focused on, the data 411 distributed at the bottom end of the input data I0_t1 of t1 also contains the same overlapping pixels, as shown by the polka dot pattern in FIG. 4(b). Similarly, the data 412 distributed at the right end of the input data I0_t2 of t2 also contains the same overlapping pixels as the input data I0_t3 of t3. Therefore, the calculation results of the convolution calculation process performed on these overlapping pixels by referring to the same surrounding pixels are the same. Therefore, when the output data O0_t3 of t3 is focused on, the output data O0_t1 of t1 and the output data O0_t2 of t2 contain the calculation results 413 and 414 on the overlapping pixels, as shown by the shading.

[0048] For this reason, in the image processing device 100 of this embodiment, a partial area including overlapping pixels is set in the processing target image to generate input data for a plurality of tiles, but control is performed to avoid overlapping calculations in the process of sequentially performing convolution calculation processing on the input data for the tiles. In the following example, in order to facilitate understanding of the invention, for input data t3 corresponding to a partial area set in the lower right of the processing target image, a pixel area not including overlapping pixels is input to the product-sum calculation processing unit 103, so that the overlapping pixels are excluded from the target of the convolution calculation processing. That is, in the CNN calculation unit 101 of this embodiment, as shown by hatching in FIG. 4(c), the upper end pixel of the input data t3 that overlaps with the input data t1 and the left end pixel that overlaps with the input data t2 are excluded from the target of the convolution calculation processing. In other words, among the pixel areas shown as the input data of t3 in FIG. 4(c), only the plain pixel area 421 not hatched is the target of the convolution calculation processing for the input data of t3.

[0049] The exclusion from the target of the convolution operation can be realized, for example, by transmitting information on the memory addresses of the pixel area excluding overlapping pixels together with an execution command related to the input data of t3 transmitted from the CPU 102 to the product-sum operation processing unit 103. The product-sum operation processing unit 103 receives the information, reads the corresponding data from the shared memory 105, and executes the convolution operation.

[0050] In this way, by limiting the processing target of the product-sum calculation processor 103 as the input data of t3 in the input layer, the output data O0_t3 (422) of t3 has a size smaller than the output data related to other partial regions as shown in FIG. 4(c). At this time, in order to make the input data I1_t3 of the next layer the same size as the input data related to other tiles, for example, the tile generator 107 may use the calculation results obtained by processing other tiles for the output data O0_t3 of t3 for the overlapping pixels. That is, as shown in the figure, pixel information 423 and 424 indicating the calculation results related to the overlapping pixels are obtained from the output data O0_t1 of t1 and the output data O0_t2 of t2. Then, by combining the obtained information with the output data O0_t3 (422) of t3, the input data I1_t3 (425) of t3 in the next layer can be generated. As a result, for the convolution calculation process related to the input layer (CNN0), output data of the same size as the input data can be obtained for each partial region.

[0051] 4 shows the control of input data for the convolution operation processing related to the input layer (CNN0) of the CNN operation, but such control can be performed not only in the input layer but also in each layer. That is, when making the product-sum operation processing unit 103 perform the convolution operation processing related to CNN1, for example, the CPU 102 can also perform processing to limit the processing target of the product-sum operation processing unit 103 for the input data I1_t3 of t3 corresponding to the lower right partial area.

[0052] In addition, the output data of multiple tiles obtained by the convolution operation processing for each layer does not need to be synthesized into an image of the same size as the image to be processed after the output data of all tiles of each layer is obtained. On the other hand, the output data of each tile is the same size as the input data, and the corresponding partial area in the image to be processed is unchanged. That is, regardless of the input data or output data of any layer, they each correspond to the partial area related to the input data of multiple tiles in the input layer of the CNN operation. In other words, for one tile, regardless of the input data or output data of any layer, the corresponding partial area is the same, so the pixel positions and number of overlapping pixels are also unchanged.

[0053] In this way, in the CNN calculation unit 101 of this embodiment, when the processing target image is separated into multiple tiles and CNN calculation is performed, the amount of calculation and calculation time can be reduced by limiting the targets of the convolution calculation process and reusing the calculation results obtained for the other tiles.

[0054] Hierarchical Calculation Processing Hereinafter, the hierarchical calculation process executed for each layer of the CNN calculation of the processing target image in the CNN calculation unit 101 of this embodiment will be described in detail with reference to the flowchart of Fig. 5. The process corresponding to the flowchart can be realized by the CPU 102 reading out the corresponding processing program stored in, for example, the external memory 108, and expanding and executing the program in the shared memory 105. This hierarchical calculation process will be described as being started when processing each layer of the CNN calculation of the processing target image is performed, for example.

[0055] It is assumed that, prior to the start of this hierarchical calculation process, input data for multiple tiles for which convolution calculation processing is to be performed in the corresponding layer is stored in the shared memory 105. In the input layer, the input data for the multiple tiles is generated by setting a predetermined number of partial regions in the image to be processed by the tile generation unit 107 and extracting pixels related to the partial regions. In the subsequent layer, the input data for the multiple tiles is generated from output data output by the convolution calculation processing for each of the multiple tiles by the product-sum calculation processing unit 103 in the previous layer, or by combining the output data.

[0056] In S501, the CPU 102 selects a tile to be subjected to the convolution operation process (target tile) based on a predetermined order.

[0057] In S502, CPU 102 judges whether or not the pixels included in the input data of the target tile include pixels for which convolution calculation processing of the same layer has already been performed for another tile. The judgment in this step is made based on whether or not overlapping pixels exist in corresponding partial areas between the target tile and a previously processed tile. If CPU 102 judges that the pixels included in the input data of the target tile include pixels for which convolution calculation processing of the same layer has already been performed for another tile, CPU 102 shifts the process to S503, and if CPU 102 judges that the pixels do not include overlapping pixels, CPU 102 shifts the process to S504.

[0058] In S503, the CPU 102 excludes pixels (overlapping pixels) for which the convolution calculation has already been completed from the input data of the target tile, and determines a pixel area (rectangular data) for which the product-sum calculation processor 103 is to execute the convolution calculation.

[0059] On the other hand, if it is determined in S502 that the input data of the target tile does not include any pixels on which convolution calculation has already been performed, then in S504 the CPU 102 determines the entire input data of the target tile as the pixel area on which convolution calculation is to be performed.

[0060] In S505, the CPU 102 causes the product-sum operation processing unit 103 to execute a convolution operation process for the target tile. At this time, the product-sum operation processing unit 103 executes a convolution operation process for the pixel region of the target tile determined in S503 or S504, and stores the operation result in the shared memory 105 as output data of the target tile.

[0061] In S506, the CPU 102 determines whether or not any tiles on which convolution calculation processing has not been performed for the current layer are present in the shared memory 105. If the CPU 102 determines that any tiles on which convolution calculation processing has not been performed are present, the process returns to S501, and if the CPU 102 determines that any tiles on which convolution calculation processing has not been performed are not present, the process ends this layer calculation processing.

[0062] In this way, in the image processing device 100 of this embodiment, when the processing target image is separated into multiple tiles and CNN calculations are performed, the number of pixels to be subjected to convolution calculation processing at each layer can be reduced, thereby reducing the amount of calculations.

[0063] <<Generation Process>> Here, when the pixel area where the convolution operation process is performed for the input data of a specific tile is limited, the output data of the tile has fewer pixels than the output data of other tiles. Therefore, in order to make the input data of the next layer have the same number of pixels for all tiles (partial areas), the generation process of the input data of the next layer executed in the image processing device 100 will be described in detail using the flowchart of FIG. 6. The process corresponding to the flowchart can be realized by the CPU 102 reading out the corresponding processing program stored in, for example, the external memory 108, and expanding and executing it in the shared memory 105. This generation process will be described as being started when, for example, the convolution operation process for one target tile in the above-mentioned hierarchical operation process is completed and the output data of the target tile is stored in the shared memory 105.

[0064] In S601, the tile generation unit 107, under the control of the CPU 102, determines whether the number of pixels of the output data of the target tile is different from the number of pixels of the input data of the target tile. If the tile generation unit 107 determines that the number of pixels of the output data of the target tile is different from the number of pixels of the input data, it proceeds to S602, and if it determines that they are the same, it ends this generation process.

[0065] In S602, the tile generation unit 107, under the control of the CPU 102, acquires information corresponding to pixel positions of overlapping pixels from output data of other tiles on which convolution calculation processing has been performed prior to the target tile. The tile generation unit 107 then adds the acquired information to the output data of the target tile to generate input data of the next layer for the target tile. The input data of the next layer for the target tile has the same number of pixels as the input data of the target tile and the input data of the next layer for the other tiles.

[0066] In this way, the amount of calculation for the convolution calculation process for multiple tiles at each layer of the CNN calculation can be reduced, while input data for multiple tiles having a common number of pixels can be easily generated for the next layer.

[0067] As described above, according to the image processing device 100 of this embodiment, it is possible to efficiently execute convolution calculation processing on an image.

[0068] In this embodiment, the tile generation unit 107 is described as setting a partial region for the image to be processed and generating multiple tiles, but the implementation of the present invention is not limited to this. That is, a configuration for generating multiple tiles is not essential for implementing the present invention, and the present invention is also applicable to a mode in which multiple tiles generated by another device are obtained and loaded into the shared memory 105. In this case, information on pixel positions related to overlapping pixels between tiles is also obtained in a similar manner.

[0069] In addition, in the present embodiment, the present invention has been described as being applied to the convolution layer of a CNN operation, but the implementation of the present invention is not limited to this, and the present invention can also be applied to aspects including, for example, a fully connected layer in a neural network.

[0070] In this embodiment, in order to facilitate understanding of the invention, four tiles t0, t1, t2, and t3 are provided, and only for t3, overlapping pixels with adjacent tiles are excluded from the convolution operation processing. However, the implementation of the present invention is not limited to this, and the present invention can be applied to an embodiment in which a similar exclusion control is performed on the condition that overlapping pixels are included with other tiles and convolution operation processing has already been completed for the other tiles. In other words, the present invention can be applied to an embodiment in which convolution operation processing is performed using each of a plurality of tiles as input data, and overlapping pixels in the same area in the image to be processed as other tiles are excluded from the target for at least some tiles.

[0071] In this embodiment, a total of four partial regions, two in the horizontal direction and two in the vertical direction, are set to include overlapping pixels in the image to be processed, and CNN calculations are performed on the four tiles corresponding to these partial regions. However, the implementation of the present invention is not limited to this, and various modes can be adopted for setting the partial regions as long as they have overlapping regions at least at the boundaries with other tiles. In addition, a predetermined number of partial regions may be set for the image to be processed, or the number may be determined so that the data size of each tile is below a predetermined capacity.

[0072] [Variation 1] In the above-described embodiment, a mode has been described in which convolution calculation processing relating to overlapping pixels is avoided by having the CPU 102 specify a partial pixel area of ​​the input data of a tile and have the product-sum calculation processing unit 103 execute the convolution calculation processing. However, the present invention is not limited to this embodiment, and for example, the CPU 102 may input input data including overlapping pixels to the product-sum calculation processing unit 103, and the product-sum calculation processing unit 103 may perform the convolution calculation processing while excluding the overlapping pixels based on past calculation results.

[0073] [Embodiment 2] In the above-described embodiment and modified example, in all layers of the CNN operation, overlapping pixels between partial regions set for the processing target image are excluded from the target of the convolution operation process. On the other hand, in a case where a filter kernel such as 3×3 is applied, the deeper the layer, the wider the range (receptive field) of pixels in the original processing target image that contribute to the convolution operation process of one pixel becomes.

[0074] For example, in the filter processing of the input layer for the region in which overlapping pixels are distributed in the input data of t3 as shown in FIG. 7(a), eight peripheral pixels included in a filter region 702 centered on a pixel 701 in the region are processed as shown in FIG. 7(b). Since all of the eight peripheral pixels for pixel 701 are included in the overlapping pixels, the calculation result is obtained in the convolution calculation processing for the input data of t1, and as a result, it can be excluded from the target of the convolution calculation processing for the input data of t3. On the other hand, in the next layer of the input layer, as shown in FIG. 7(c), eight peripheral pixels included in a filter region 712 centered on a pixel 711 at the same pixel position are also all included in the overlapping pixel region, but the pixel positions affected are different in terms of the receptive field. That is, the pixel value of each of the eight peripheral pixels included in the filter region 712 is derived by referring to the peripheral pixels in the input layer, and the region 713 shown in FIG. 7(c) is the receptive field. That is, some of the pixels included in the filter region 712 (the three pixels aligned at the bottom) are pixels that have been generated by referring to pixels padded with 0 in the convolution operation process on the input data of t1.

[0075] Therefore, in the convolution operation process of the lower layer, the operation results may differ between the output data of the tiles even for overlapping pixels. More specifically, in the next input layer, the pixel values ​​of the horizontal row of the bottommost overlapping pixels in the output data of t1 are not derived with reference to all the pixels of the image to be processed contained in the receptive field. Similarly, in the next input layer, the pixel values ​​of the vertical column of the rightmost overlapping pixels in the output data of t2 are not derived with reference to all the pixels of the image to be processed contained in the receptive field.

[0076] Therefore, since the size of the receptive field related to the convolution operation processing becomes larger at a deeper layer, the CNN operation unit 101 of this embodiment performs control to reduce the number of pixels to be excluded from the target of the convolution operation processing as overlapping pixels at a deeper layer. In other words, depending on which layer of the CNN operation the product-sum operation processing unit 103 is to perform the convolution operation processing, the CPU 102 specifies overlapping pixels by differentiating the pixels that can reuse the previous operation result based on the receptive field, and excludes the overlapping pixels from the operation target.

[0077] In the example of FIG. 7(c), for the convolution operation process of the second layer (CNN1), among the pixels included in the input data of t3, the CPU 102 may identify overlapping pixels between the adjacent input data of t1 and t2, for example, as follows. The CPU 102 identifies overlapping pixels between t1 and t3 by excluding the bottommost row of one pixel among the pixels overlapping between the corresponding partial regions set in the processing target image. The CPU 102 also identifies overlapping pixels between t2 and t3 by excluding the rightmost column of one pixel among the pixels overlapping between the corresponding partial regions set in the processing target image. It is easy to understand that such identification of overlapping pixels for each layer changes depending on the size of the filter kernel applied in the filter process of the convolution operation process and the depth of the layer.

[0078] In this way, it is possible to improve the accuracy of CNN calculations while reducing the calculation time.

[0079] [Embodiment 3] In the above-described embodiment and modified example, a method was described in which the output data of the convolution operation process of at least some tiles (t3) is generated by combining the operation result obtained by excluding overlapping pixels from the input data with the operation result obtained by other tiles for the overlapping pixels. That is, in the aspect shown in FIG. 4(c), in order to generate output data (next-layer input data I1_t3) corresponding to the input data I0_t3 of t3, the operation result of other tiles is diverted to O0_t3 obtained by the convolution operation process excluding overlapping pixels. In this case, in order to generate I1_t3, it is necessary to access the memory addresses shown by hatching in O0_t1 and O0_t2 in the shared memory 105 in FIG. 4(b).

[0080] Incidentally, when tiles are generated by setting partial regions in a manner that divides the image to be processed horizontally as in the above-described embodiment and modified example, there are two types of other tiles that reuse the output data. Also, reading pixel values ​​of a column of n pixels at the right end of the memory space, such as O0_t2 shown in FIG. 4(b), requires cumbersome memory access to discontinuous memory addresses. The discontinuous memory access requires a corresponding amount of processing time, and the start of the convolution operation process of the next layer is delayed. In contrast, reading pixel values ​​of a row of n pixels at the bottom end of the memory space, such as O0_t1 shown in FIG. 4(b), can be completed by memory access to consecutive memory addresses in raster order.

[0081] For this reason, in this embodiment, the tile generating unit 107 generates multiple tiles by setting partial regions so as to divide the processing target image only vertically, without dividing it horizontally, as shown in Fig. 8. That is, by generating multiple tiles as partial regions as shown in Fig. 8, memory access when reusing output data of other tiles can be memory access to consecutive memory addresses in raster order, similar to O0_t1 in Fig. 4(b). By adopting such a tile generation mode, memory access to non-consecutive memory addresses can be avoided, and as a result, it is possible to realize high-speed processing related to CNN calculation.

[0082] The size of the partial regions into which the image to be processed is divided vertically and the number of partial regions to be set may be determined based on the size of the memory space prepared for the input data of one tile. Here, the memory space prepared for the input data of one tile is configured to be able to store data after pixels are added around the tile generated by the tile generating unit 107 by zero padding or mirroring. Therefore, the size and number of partial regions may be determined based on the number of the added pixels.

[0083] [Variation 2] In the above-mentioned third embodiment, a mode has been described in which tiles are generated by setting partial regions in a mode in which the image to be processed is divided only in the vertical direction, thereby making memory access more efficient when generating input data for the next layer, but the present invention is not limited to this mode. For example, even in a mode in which tiles are generated by setting partial regions in which the image to be processed is also divided in the horizontal direction, the same effect can be obtained by selecting tiles adjacent in the vertical direction, rather than in raster order, from the upper left tile in the order of convolution calculation processing. In this case, the CPU 102 repeatedly selects a tile for which the next convolution calculation processing is to be performed, the corresponding partial region of which is adjacent in the vertical direction to the partial region corresponding to the tile for which calculation was performed most recently.

[0084] [Variation 3] In the above-mentioned embodiment and modified example, it has been described that the number of pixels of the input and output data does not change in the convolution operation process of each layer of the CNN operation, but the implementation of the present invention is not limited to this. For example, the convolution operation process is performed without zero padding or the like on the input data, and the output data becomes data with a smaller number of pixels than the input data. In this case, each input data of the next layer also becomes data with a smaller number of pixels than the input data of the previous layer.

[0085] In such an embodiment, it is easily understood that the generation process for generating the input data for the next layer changes the judgment criteria in S601 and the number of pixels that reuse the output data of other tiles. For example, in an embodiment in which a 3×3 filter kernel is applied, the number of horizontal pixels and the number of vertical pixels of the output data by the convolution operation process are each 2 less than the number of the input data.

[0086] [Variation 4] In the above-described embodiment and modified example, a plurality of tiles are generated from the image to be processed by setting a plurality of equal-sized partial regions such that the number of overlapping pixels between the partial regions is equal, but the present invention is not limited to this. It will be easily understood that the present invention is also applicable even if the number of overlapping pixels between the partial regions differs for each tile.

[0087] [Variation 5] In the above-described embodiment and modified example, the product-sum operation processing unit 103 has been described as executing the product-sum operation related to the filter processing as the operation of the convolution layer related to the CNN operation, but the product-sum operation processing unit 103 may be configured to perform activation function operation or the like. In addition, the operation result (output data) by the product-sum operation processing unit 103 has been described as being stored in the shared memory 105, but the output data may also be stored in another storage device such as the external memory 108. In such an embodiment, when reusing the output data related to the tile on which the operation was previously performed, a process of reading the corresponding output data into the shared memory 105 is also required.

[0088] [Variation 6] In the above-described embodiment and modified example, the image processing device 100 has been described as a standalone device that acquires a captured image and performs CNN calculations. Such an image processing device 100 can be incorporated into, for example, an imaging device, and may be used to generate multiple tiles for a captured image obtained by imaging and perform CNN calculations in sequence.

[0089] [Variation 7] In the above-described embodiment and modified examples, the target of the CNN calculation by the CNN calculation unit 101 is described as a two-dimensional image, but the present invention can also be applied to data other than two-dimensional images (such as three-dimensional data).

[0090] [Other embodiments] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0091] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention.

[0092] [Summary of the embodiment and modifications] The disclosure of this specification includes the following image processing device, imaging device, control method, and program. (Item 1) A calculation means for performing a convolution calculation process in a neural network on input data; an acquisition means for acquiring a plurality of tiles, each of the tiles corresponding to a sub-region in an image; a control means for controlling the calculation means to execute the convolution calculation process using each of the plurality of tiles as the input data; having The control means controls the calculation means to exclude overlapping pixels included in at least a part of the tiles, which overlap a corresponding area in the image with another tile, from the target of the convolution calculation process. 13. An image processing device comprising: (Item 2) 2. The image processing device according to item 1, wherein the control means excludes the overlapping pixels from the convolution operation processing by the calculation means by using an image constructed by excluding the overlapping pixels from at least some of the tiles as the input data. (Item 3) the plurality of tiles are images corresponding to any of a predetermined number of partial regions set in the image, the predetermined number of partial regions being set so as to overlap at least one of the other partial regions, The overlapping pixels correspond to an area where adjacent partial areas set among the predetermined number of partial areas overlap. 3. The image processing device according to item 1 or 2. (Item 4) The control means selecting each of the plurality of tiles in turn as the input data, and causing the calculation means to execute the convolution calculation process; On condition that the convolution operation process for the other tiles has been completed, the calculation means is controlled to exclude the overlapping pixels included in the at least some of the tiles from the convolution operation process. 4. The image processing device according to any one of items 1 to 3, (Item 5) 5. The image processing device according to item 4, wherein the control means selects, as the next input data, a tile corresponding to a partial area vertically adjacent in the image to a partial area corresponding to a tile most recently selected as the input data. (Item 6) further comprising an output unit that generates output data for each of the plurality of tiles based on a result of the convolution operation performed by the operation unit; The output means generates the output data for the at least some of the tiles based on a result of the convolution operation process for the tile and a result of the convolution operation process for the overlapping pixels for the other tiles. 6. The image processing device according to any one of items 1 to 5, (Item 7) the neural network is a convolutional neural network that includes a plurality of layers as convolutional layers and repeatedly performs the convolution operation process on the image in the plurality of layers; the control means causes the calculation means to execute the convolution calculation process for each of the plurality of tiles for each layer of the convolutional neural network; The acquisition means includes: For an input layer of the convolutional neural network, images included in a plurality of partial regions set in the image are acquired as the plurality of tiles; For a subsequent layer of the convolutional neural network, a plurality of pieces of output data generated by the output means for a previous layer are obtained as the plurality of tiles of the layer. 7. The image processing device according to item 6, (Item 8) 8. The image processing device according to item 7, wherein the control means changes the pixels to be excluded as the overlapping pixels depending on the hierarchy of the convolutional neural network for the at least some of the tiles corresponding to the same partial region. (Item 9) the calculation means performs the convolution calculation process for each pixel included in the input data in each layer of the convolutional neural network by using pixel values ​​of pixels included in a predetermined area defined based on the pixel; The size of the receptive field in the image that contributes to the result of the convolution operation for one pixel is determined according to the layer of the convolutional neural network that executed the convolution operation, The control means determines the pixels to be excluded as the overlapping pixels based on the size of the receptive field in the convolution operation process to be executed by the operation means. 9. The image processing device according to item 8, (Item 10) The size of the receptive field is larger in a deeper layer of the convolutional neural network, 10. The image processing device according to item 9, wherein the control means is configured to exclude fewer pixels as overlapping pixels for at least a portion of tiles corresponding to the same partial region as the hierarchy of the convolutional neural network becomes deeper. (Item 11) 11. The image processing device according to item 9 or 10, characterized in that in the convolution operation process for the other tiles, the control means does not exclude pixels whose receptive fields include pixels that are not included in the other tiles as the overlapping pixels. (Item 12) generating means for generating the plurality of tiles based on the image; The generating means generates the plurality of tiles by setting partial regions so as not to divide the image in the horizontal direction. 12. The image processing device according to any one of items 1 to 11, (Item 13) An imaging means; An image processing device according to any one of items 1 to 12, An imaging device having The acquisition means acquires the plurality of tiles based on the image captured by the imaging means. 1. An imaging device comprising: (Item 14) A calculation step of performing a convolution calculation process in a neural network on the input data; acquiring a plurality of tiles, each of the tiles corresponding to a sub-region in the image; a control step of controlling the calculation step to execute the convolution calculation process using each of the plurality of tiles as the input data; having In the control step, the calculation step is controlled so that, for at least a part of the plurality of tiles, overlapping pixels included in the tile and corresponding to other tiles in the image in the same area are excluded from the convolution calculation process. 23. A method for controlling an image processing apparatus comprising the steps of: (Item 15) 13. A program for causing a computer to function as each of the means of the image processing device according to any one of items 1 to 12. [Explanation of symbols]

[0093] 100: image processing device, 101: CNN calculation unit, 102: CPU, 103: product-sum calculation processing unit, 104: MAC, 105: shared memory, 107: tile generation unit

Claims

1. A calculation means for performing a convolution calculation process in a neural network on input data; an acquisition means for acquiring a plurality of tiles, each of the tiles corresponding to a sub-region in an image; a control means for controlling the calculation means to execute the convolution calculation process using each of the plurality of tiles as the input data; having The control means controls the calculation means to exclude overlapping pixels included in at least a part of the tiles, which overlap a corresponding area in the image with another tile, from the target of the convolution calculation process.

13. An image processing device comprising:

2. 2. The image processing device according to claim 1, wherein the control means excludes the overlapping pixels from the convolution operation processing by the calculation means by using an image constructed by excluding the overlapping pixels from at least some of the tiles as the input data.

3. the plurality of tiles are images corresponding to any of a predetermined number of partial regions set in the image, the predetermined number of partial regions being set so as to overlap at least one of the other partial regions, The overlapping pixels correspond to an area where adjacent partial areas set among the predetermined number of partial areas overlap.

2. The image processing device according to claim 1,

4. The control means selecting each of the plurality of tiles in turn as the input data, and causing the calculation means to execute the convolution calculation process; On condition that the convolution operation process for the other tiles has been completed, the calculation means is controlled to exclude the overlapping pixels included in the at least some of the tiles from the convolution operation process.

2. The image processing device according to claim 1,

5. 5. The image processing device according to claim 4, wherein the control means selects, as the next input data, a tile corresponding to a partial area vertically adjacent in the image to a partial area corresponding to a tile most recently selected as the input data.

6. further comprising an output unit that generates output data for each of the plurality of tiles based on a result of the convolution operation performed by the operation unit; The output means generates the output data for the at least some of the tiles based on a result of the convolution operation process for the tile and a result of the convolution operation process for the overlapping pixels for the other tiles.

2. The image processing device according to claim 1,

7. the neural network is a convolutional neural network that includes a plurality of layers as convolutional layers and repeatedly performs the convolution operation process on the image in the plurality of layers; the control means causes the calculation means to execute the convolution calculation process for each of the plurality of tiles for each layer of the convolutional neural network; The acquisition means includes: For an input layer of the convolutional neural network, images included in a plurality of partial regions set in the image are acquired as the plurality of tiles; For a subsequent layer of the convolutional neural network, a plurality of pieces of output data generated by the output means for a previous layer are obtained as the plurality of tiles of the layer.

7. The image processing device according to claim 6,

8. The image processing device according to claim 7 , wherein the control means causes pixels to be excluded as the overlapping pixels to differ depending on a hierarchical level of the convolutional neural network for the at least some of the tiles corresponding to the same partial region.

9. the calculation means performs the convolution calculation process for each pixel included in the input data in each layer of the convolutional neural network by using pixel values ​​of pixels included in a predetermined area defined based on the pixel; The size of a receptive field in the image that contributes to the result of the convolution operation for one pixel is determined according to a layer of the convolutional neural network that executed the convolution operation, The control means determines the pixels to be excluded as the overlapping pixels based on the size of the receptive field in the convolution operation process to be executed by the operation means.

9. The image processing device according to claim 8,

10. The size of the receptive field is larger in a deeper layer of the convolutional neural network, The image processing device according to claim 9 , wherein the control means is configured to exclude fewer pixels as the overlapping pixels for at least the part of tiles corresponding to the same partial region as the hierarchy of the convolutional neural network becomes deeper.

11. The image processing device according to claim 9 , wherein the control means does not exclude, as the overlapping pixels, pixels whose receptive fields include pixels not included in the other tiles in the convolution operation process for the other tiles.

12. generating means for generating the plurality of tiles based on the image; The generating means generates the plurality of tiles by setting partial regions so as not to divide the image in the horizontal direction.

2. The image processing device according to claim 1,

13. An imaging means; An image processing device according to any one of claims 1 to 12, An imaging device having The acquisition means acquires the plurality of tiles based on the image captured by the imaging means.

1. An imaging device comprising:

14. A calculation step of performing a convolution calculation process in a neural network on the input data; acquiring a plurality of tiles, each of the tiles corresponding to a sub-region in the image; a control step of controlling the calculation step to execute the convolution calculation process using each of the plurality of tiles as the input data; having In the control step, the calculation step is controlled so that, for at least a part of the plurality of tiles, overlapping pixels included in the tile and corresponding to other tiles in the image are the same, are excluded from the target of the convolution calculation process.

23. A method for controlling an image processing apparatus comprising the steps of:

15. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Information processing device, control method and program

    JP2019125128A

Cited By

  • Game machine

    JP2025105806A

  • Game machine

    JP2025105807A