Processing apparatus

The processing apparatus efficiently performs convolutional and pooling operations in neural networks with a reduced circuit scale by utilizing shared hardware components, addressing the challenge of large circuit scales in existing methods.

JP2025089119APending Publication Date: 2025-06-12CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023204128
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing methods for neural network processing, such as those using Convolutional Neural Networks (CNNs), require dedicated circuits for convolutional and pooling processing, leading to a large circuit scale due to the different configurations of these circuits.

Method used

A processing apparatus that includes a multiplication circuit, a memory, an addition circuit, a comparison circuit, and a selection circuit, which collectively enable efficient convolutional and pooling processing while maintaining a reduced circuit scale by sharing hardware components.

Benefits of technology

The proposed solution allows for efficient performance of convolutional and pooling processing in neural networks without increasing the circuit scale, thereby improving processing efficiency and supporting various neural network structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025089119000001_ABST
    Figure 2025089119000001_ABST
Patent Text Reader

Abstract

To efficiently perform convolution operation and pooling operation in neural network processing, while preventing an increase in the circuit scale for performing the processing.SOLUTION: A processing apparatus comprises: a multiplier circuit that sequentially outputs a multiplication result of each data of a plurality of pieces of data and a corresponding coefficient; a memory; an adding circuit that adds the multiplication result output from the multiplier circuit and data held in the memory, and outputs an adding result; a comparison circuit that compares the multiplication result output from the multiplier circuit with the data held in the memory, and outputs one of the multiplication result output from the multiplier circuit and the data held in the memory; and a selection circuit that outputs, to the memory, one of the output from the adding circuit and the output from the comparison circuit, so that the memory holds one of the outputs.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a processing device, and particularly to neural network processing.

Background Art

[0002] Neural networks including Convolutional Neural Networks (CNNs) are used in deep learning. Processing in neural networks often includes convolutional operations and pooling processing.

[0003] It is required to shorten the overall processing time by efficiently performing processing in neural networks. In particular, in processing using CNNs, many operations are performed, so it is expected that the processing time can be effectively shortened by improving the efficiency of operations. In particular, when applying a neural network to an embedded system such as a mobile terminal or in-vehicle device, improvement of operation efficiency is strongly required. For such purposes, improving the efficiency of convolutional processing and pooling processing has been studied.

[0004] For example, Patent Document 1 and Patent Document 2 propose configurations for efficiently performing convolutional processing and pooling processing by using dedicated hardware or dedicated circuits. In the method of Patent Document 1, a matrix operation unit performs convolutional processing. Also, a vector calculation unit performs pooling processing according to a specified stride. In the method of Patent Document 2, a matrix operation device performs convolutional processing. Also, a pooling unit having an aligner for aligning the output of convolutional processing and a pooler for applying a pooling operation performs pooling processing according to a specified stride.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] The methods described in Feature Document 1 and Feature Document 2 require two types of dedicated circuits including a dedicated circuit for performing convolution processing and a dedicated circuit for performing pooling processing. On the other hand, since the circuit for performing pooling processing has a different configuration from the circuit for performing convolution processing, it is necessary to increase the circuit scale. For this reason, the methods described in Feature Document 1 and Feature Document 2 had the problem that the circuit scale of the hardware used for processing was large.

[0007] An object of the present invention is to efficiently perform convolution processing and pooling processing while suppressing an increase in the circuit scale for performing these processes in neural network processing.

Means for Solving the Problems

[0008] A processing apparatus according to an embodiment of the present invention includes the following configuration. That is, a multiplication circuit that sequentially outputs multiplication results of each data among a plurality of data and corresponding coefficients, a memory, an addition circuit that adds the multiplication result output by the multiplication circuit and the data held in the memory and outputs the addition result, a comparison circuit that compares the multiplication result output by the multiplication circuit and the data held in the memory and outputs one of the multiplication result output by the multiplication circuit and the data held in the memory, and a selection circuit that outputs one of the output of the addition circuit and the output of the comparison circuit to the memory so that the memory holds it.

Effects of the Invention

[0009] In neural network processing, it is possible to efficiently perform convolution processing and pooling processing while suppressing an increase in the circuit scale for performing these processes.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

MODE FOR CARRYING OUT THE INVENTION

[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.

[0012] <Configuration Example of Data Parallel Processing Apparatus> FIG. 3 is a block diagram showing a configuration example of an information processing apparatus according to an embodiment of the present invention. Note that the information processing apparatus may have various other components, which are omitted here for explanation.

[0013] The input unit 301 acquires an instruction or data from the user. The input unit 301 can be a keyboard, a pointing device, a button, or the like.

[0014] The data storage unit 302 is a recording medium for storing data. The data storage unit 302 can store image data, programs, or other data. The data storage unit 302 can be, for example, a hard disk, a floppy disk, a CD-ROM, a CD-R, a DVD, a memory card, a CF card, a smart media, an SD card, a memory stick, an xD picture card, or a USB memory, etc. A part of the RAM 308 described later may be used as the data storage unit 302.

[0015] The display unit 304 displays an image. The display unit 304 can display an image before or after image processing, or an image such as a GUI. The display unit 304 can be a CRT or a liquid crystal display, etc.

[0016] Also, the display unit 304 and the input unit 301 may be realized by the same device. For example, a touch screen device can be used as the display unit 304 and the input unit 301. In this case, the input on the touch screen can be treated as an input to the input unit 301.

[0017] The convolution processing unit 305 performs convolution processing as described later. For example, convolution processing can be performed on the image stored in the RAM 308. Specifically, the convolution processing unit 305 can perform processing (S101 to S114) according to the flowchart of FIG. 1 described later. That is, the convolution processing unit 305 can perform convolution neural network processing including a sum-of-products operation on the result of image processing stored in the RAM 308. Also, the convolution processing unit 305 outputs the result obtained by the processing to the data storage unit 302 (or the RAM 308). The convolution processing unit 305 can function as a convolution processing device according to an embodiment. Note that the information processing device may have a plurality of convolution processing units 305 in order to perform the processing in the neural network in parallel.

[0018] The CPU 306 controls the operation of the entire device. Note that FIG. 3 shows a configuration in which the information processing device has only one CPU (CPU 306). However, the information processing device may have a plurality of CPUs, GPUs (Graphics Processing Units), NPUs (Neural Processing Units), or the like.

[0019] The ROM 307 and the RAM 308 provide programs, data, work areas, etc. necessary for processing by the CPU 306. When a program is stored in the data storage unit 302 or the ROM 307, the program is first loaded into the RAM 308. Then, the program on the RAM 308 is executed. Further, the information processing device may receive a program via the communication unit 303. In this case, the program is first recorded in the data storage unit 302 and then loaded into the RAM 308. Alternatively, the program is directly loaded from the communication unit 303 into the RAM 308.

[0020] Also, the CPU 306 can generate the result of image processing or image recognition based on the processing result by the convolution processing unit 305. In one embodiment, the convolution processing unit 305 outputs a confidence map representing the probability that a detection target object exists for each position or region of the input image. In this case, the CPU 306 can generate and output information indicating the position of a specific subject in the image according to the confidence map. For example, the CPU 306 can determine that a subject exists at the peak position of the value in the confidence map. Then, the CPU 306 can superimpose information indicating the position of the subject determined on the input image.

[0021] Note that the convolution processing unit 305 may perform processing on each of a plurality of frames of a moving image. In this case, the CPU 306 can generate the result of image processing or image recognition for the moving image. The CPU 306 can store the result of image processing or image recognition in the RAM 308.

[0022] The image processing unit 309 performs image processing on images. For example, the image processing unit 309 can perform image processing on the image data written in the data storage unit 302 according to the command received from the CPU 306, and write the result into the RAM 308. The type of image processing is not particularly limited. The image processing may be, for example, a range adjustment process of pixel values.

[0023] The communication unit 303 is an interface for performing communication between devices. In FIG. 3, it is illustrated that the information processing apparatus includes elements such as the input unit 301, the data storage unit 302, and the display unit 304. However, these elements may be connected via a communication path. For example, the display unit 304 may be a display device outside the information processing apparatus connected via a cable or the like. Also, the data storage unit 302 may be virtually configured. For example, a storage device connected via the communication unit 303 may be used as the data storage unit 302. Thus, the information processing apparatus according to one embodiment is configured by a plurality of devices connected via a communication path.

[0024] <Processing target network> As described above, the information processing apparatus according to the present embodiment can perform processing in a neural network. This neural network can include a layer in which convolution processing is performed and a layer in which pooling processing is performed. In one embodiment, the neural network is a CNN.

[0025] In the processing in the neural network, convolution processing using filter coefficients (weight coefficients) determined by learning and pixel values (feature data) of a feature image is performed for each spatial local (window). The convolution process is a product-sum operation and includes a plurality of multiplication processes and cumulative addition processes.

[0026] In addition, in the processing of a neural network, pooling processing is also performed. Pooling processing is a process of outputting a representative value (maximum value, minimum value, average value, etc.) for each spatial local area (window). Stride is a parameter of the pooling processing and indicates the movement width of the window. When the stride is 2, the feature image is reduced to half the size in both the vertical and horizontal directions by the pooling processing.

[0027] For this reason, the information processing apparatus has a convolutional processing apparatus or can function as a convolutional processing apparatus. This convolutional processing apparatus can perform convolutional processing and pooling processing on an image.

[0028] In the present embodiment, by combining the filter processing and the stride processing, the pooling processing according to any stride as described above is realized. In the filter processing, a representative value for each window is calculated by applying a filter. The feature image obtained by the filter processing has the representative value thus calculated as feature data for each pixel. In the present embodiment, the stride of the filter processing is 1. As the filter, a maximum value filter, a minimum value filter, an average value filter, or the like can be used. The filter processing corresponds to the pooling processing with a stride of 1. As will be described later, the convolutional processing apparatus can perform two or more types of pooling processing.

[0029] Also, in the stride process, the image is reduced according to a predetermined stride. Specifically, a part of the feature data constituting the feature image is extracted at intervals according to the stride. The stride process corresponds to a pooling process with a window size of 1×1. Also, the stride process corresponds to a pooling process that uses the value at a predetermined position of the window as a representative value. Note that when the stride is 1, the feature image is the same before and after the stride process. In the present embodiment, the stride of the stride process may be greater than 1. Also, the stride of the stride process can be set independently of the window size of the pooling process. That is, the stride of the stride process and the window size of the pooling process (or the filter process) may be different.

[0030] Figure 2 shows an example of a CNN used by a convolutional processing apparatus according to an embodiment. In this CNN, a plurality of layers are hierarchically connected. And the filters and a plurality of feature data are also hierarchically configured. As shown in Figure 2, the network structure of the neural network is defined by the information of each layer (the connection relationship between layers, the filter structure, the bit width of the weight coefficient, the size of the feature image, the bit width, and the number of sheets, etc.). In the example of Figure 2, the number of layers is 4 (input layer and layers 1 to 3). Also, each layer has one or a plurality of feature images. Specifically, there are 4 feature images in layers 1 to 2. There is 1 feature image in layer 3. A plurality of feature data are included in one feature image. The plurality of feature images correspond to a plurality of channels.

[0031] The network shown in Figure 2 outputs a confidence map (feature image 204) representing the probability that a detection target object exists for each position or region of the input image. In the input layer, a sum-of-products operation is performed using a plurality of input images 201 and weight coefficients according to Equation (1). Thus, a plurality of feature images 202 of layer 1 are generated. The number of input images input to the input layer is 3. Each of the 3 input images corresponds to the R (red), G (green), or B (blue) channel.

[0032] In layer 1, filtering is performed using a plurality of feature images 202 according to Equation (2). In this example, the window size (filter size) is 3×3. In layer 1, further, striding is performed according to Equation (4). Thus, a plurality of feature images 203 of layer 2 are generated. In this example, the stride is 2. The size (number of pixels) of the feature image 203 is 1 / 4 of that before the striding process. The combination of the process of layer 1 and the process of layer 2 corresponds to the pooling process. However, the window size of the pooling process is not limited to 3×3. The window size of the pooling process can be any size. Also, the stride of the pooling process is not limited to 2. The stride can be any positive integer.

[0033] In layer 2, a multiplication-accumulation operation is performed using a plurality of feature images 203 and weight coefficients according to Equation (1). Thus, the feature image 204 of layer 3 is generated.

[0034] As described above, the feature image of the target layer is generated by a convolution process using the feature image of the previous layer and the weight coefficient corresponding to the previous layer. To generate one feature image of the target layer, a plurality of feature images of the previous layer are used. Equation (1) shows an example of a calculation formula for the convolution process.

Equation

[0035] Equations (2) and (3) show examples of calculation formulas for filter processing. Equation (2) represents the filter processing for max pooling. Also, Equation (3) represents the filter processing for average pooling.

Number

Number

[0036] Equation (4) shows an example of a calculation formula for stride processing.

Number

[0037] In the processing in the neural network, after performing the convolution processing or the filter processing, based on the network structure, further activation processing, filter processing, or stride processing, etc. can be performed. For example, further activation processing can be performed on the processing result obtained as described above. In the above example, for the sum-of-products operation result O i,j (n), by performing the activation processing, the feature images of layers 1 and 3 are generated. The pixel (i, j) of the nth feature image can have the feature data obtained by the activation processing for the sum-of-products operation result O i,j (n).

[0038] Figure 6 shows an example of convolution processing. Convolution processing is performed on the feature data extracted from the same positions of the four feature images 601 in layer 2. Activation processing is performed on the sum-of-products operation result obtained by the convolution processing. The result of the activation processing is used as the feature data at the same position of the feature image 602 in layer 3.

[0039] Such neural network learning can be performed based on the error between the output of the neural network for the input image for learning and the corresponding teacher image. As a specific learning method, the error backpropagation method can be mentioned. As the teacher image, for example, an image in which pixel values following a Gaussian distribution are set centered on the position of the subject to be detected in the input image for learning can be used.

[0040] <Configuration of Convolution Processing Unit> Figure 4 shows a configuration example of the convolution processing unit 305. The convolution processing unit 305 includes a control unit 401, a feature data holding unit 402, a coefficient holding unit 404, a coefficient setting unit 405, a multiplication processing unit 406, and an addition / comparison processing unit 403. The convolution processing unit 305 may also include an activation / stride processing unit 407 and a data holding unit 408. Note that the convolution processing unit 305 may include a plurality of multiplication processing units 406 and a plurality of addition / comparison processing units 403 for parallel processing. The number of other processing units is not limited to one.

[0041] The control unit 401 controls the overall operation of the convolution processing unit 305. The control unit 401 may have, for example, a CPU, a sequencer, an ASIC, or an FPGA for control. On the other hand, the function of the control unit 401 may be realized by the CPU 306.

[0042] The data holding unit 408 temporarily holds the feature image, the weight coefficient, and the network structure information. For example, the data holding unit 408 can hold one or more feature images and the weight coefficients used for convolution processing. The data holding unit 408 is, for example, a memory. However, the RAM 308 or the data storage unit 302 may be used as the data holding unit 408.

[0043] The coefficient holding unit 404 temporarily holds the weight coefficients used in the convolution process. For example, the coefficient holding unit 404 holds the weight coefficient C x,y (m,n) for the sum-of-products operation for calculating the feature data of one pixel of the n-th feature image of the target layer, which can be read from and held by the data holding unit 408. x and y indicate the relative pixel positions within the window of the convolution process or the filter process. The coefficient holding unit 404 is, for example, a memory such as DRAM or SRAM.

[0044] The feature data holding unit 402 temporarily holds at least a part of the feature image to be subjected to the convolution process or the filter process. The feature data holding unit 402 can read and hold at least a part of the feature data of the feature image I(m) of the previous layer from the data holding unit 408. The feature data holding unit 402 may hold the feature data of each pixel within the window of the convolution process or the filter process. The feature data holding unit 402 is, for example, a memory such as DRAM or SRAM.

[0045] The coefficient setting unit 405, the multiplication processing unit 406, and the addition / comparison processing unit 403 calculate the result of the convolution process or the filter process.

[0046] The coefficient setting unit 405 sets a plurality of coefficients used by the multiplication processing unit 406. Here, the coefficient setting unit 405 can set the coefficients corresponding to each of the plurality of data according to whether the convolution operation or the pooling process is to be performed. In one embodiment, the coefficient setting unit 405 sets the weight coefficient C held by the coefficient holding unit 404 x,yAdjust (m, n). When performing convolution processing, the weight coefficients are the same before and after adjustment. In this case, the weight coefficients may vary according to the values of x and y. Also, when performing pooling processing, the weight coefficients after adjustment vary depending on the type of pooling processing. The weight coefficients are set so that they can be calculated by the multiplication processing unit 406 in the same way as in convolution processing. When performing pooling processing, the coefficient setting unit 405 can set constant weight coefficients regardless of the values of x and y. For example, when performing max pooling processing, the weight coefficient may be set to 1. The coefficient setting unit 405 may have a CPU, ASIC, or FPGA for processing. On the other hand, the function of the coefficient setting unit 405 may be realized by the control unit 401 or the CPU 306.

[0047] The multiplication processing unit 406 sequentially outputs the multiplication results of each data among a plurality of data and the corresponding coefficients. The plurality of data can be the feature data of a plurality of pixels within the window of convolution processing or filter processing in the feature image I(m) of the previous layer. Therefore, the multiplication processing unit 406 can sequentially output the multiplication results of each data among the plurality of data included in the window of convolution processing or the pooling processing set on the image and the corresponding coefficients. Specifically, the multiplication processing unit 406 outputs the multiplication results of these data and the corresponding weight coefficient C x,y (m, n) for each combination of different x and y. In this embodiment, the multiplication processing unit 406 is arithmetic hardware. In this specification, the multiplication processing unit 406 may be referred to as a multiplication circuit.

[0048] The addition and comparison processing unit 403 integrates the multiplication results sequentially output by the multiplication processing unit 406 or selects one of the multiplication results sequentially output by the multiplication processing unit 406 according to the control signal from the control unit 401.

[0049] FIG. 5 shows a configuration example of the addition and comparison processing unit 403. The addition and comparison processing unit 403 includes an addition unit 501, a comparison unit 502, a selection unit 503, and a result holding unit 504. The addition unit 501 adds the multiplication result output by the multiplication processing unit 406 and the data held by the result holding unit 504, and outputs the addition result. In the present embodiment, the addition unit 501 is arithmetic hardware. In this specification, the addition unit 501 may be referred to as an addition circuit.

[0050] The comparison unit 502 compares the multiplication result output by the multiplication processing unit 406 and the data held by the result holding unit 504. Then, the comparison unit 502 outputs one of the multiplication result output by the multiplication processing unit 406 and the data held by the result holding unit 504. For example, the comparison unit 502 can compare the output of the multiplication processing unit 406 and the output of the result holding unit 504, and output the larger (or smaller) one. In the present embodiment, the comparison unit 502 is arithmetic hardware. In this specification, the comparison unit 502 may be referred to as a comparison circuit.

[0051] The selection unit 503 outputs to the result holding unit 504 so that the result holding unit 504 holds one of the output of the addition unit 501 and the output of the comparison unit 502. In this way, the selection unit 503 can select one of the output of the addition unit 501 and the output of the comparison unit 502. The selection unit 503 can select the output of the addition unit 501 or the comparison unit 502 according to the control of the control unit 401. For example, the control unit 401 can control the selection operation by the selection unit 503 according to the network structure information. That is, the selection unit 503 can output to the result holding unit 504 one of the output of the addition unit 501 and the output of the comparison unit 502 selected for each layer of the neural network. Also, the selection unit 503 can output to the result holding unit 504 one of the output of the addition unit 501 and the output of the comparison unit 502 selected according to the type of processing performed in the layer. Specifically, when performing convolution processing or average pooling processing, the selection unit 503 selects the output of the addition unit 501. Otherwise (for example, when performing max pooling processing), the selection unit 503 selects the output of the comparison unit 502. In the present embodiment, the selection unit 503 is arithmetic hardware. In this specification, the selection unit 503 may be referred to as a selection circuit.

[0052] The result holding unit 504 is a memory. The result holding unit 504 can hold one of the output of the addition unit 501 and the output of the comparison unit 502 selected by the selection unit 503. As described above, the result holding unit 504 can output the held data to the addition unit 501 and the comparison unit 502.

[0053] According to the above configuration, the selection unit 503 repeats the operation of outputting, to the result holding unit 504, one of the output of the addition unit 501 and the output of the comparison unit 502 for each of a plurality of data to be processed by the multiplication processing unit 406. After this processing, the data held by the result holding unit 504 is the result of the convolution processing or the result of the filter processing for the pooling processing. Therefore, the data held by the result holding unit 504 is output as the processing result for the plurality of data. The data thus output from the result holding unit 504 corresponds to the pixel data of the image indicating the result of the convolution processing or the filter processing for the image of the previous layer.

[0054] For example, when the selection unit 503 selects the output of the addition unit 501, the addition unit 501 adds the multiplication results sequentially output by the multiplication processing unit 406 to the data held by the result holding unit 504. In this way, the addition and comparison processing unit 403 can perform cumulative addition processing. Finally, the data held by the result holding unit 504 is the integrated result of the respective multiplication results output by the multiplication processing unit 406. Therefore, the result holding unit 504 can output the result of the convolution processing. Also, by appropriately setting the weight coefficient, the addition and comparison processing unit 403 can also output the average value of the plurality of multiplication results.

[0055] On the other hand, when the selection unit 503 selects the output of the comparison unit 502, the multiplication result selected from the multiplication results sequentially output by the multiplication processing unit 406 is stored in the result holding unit 504. For example, when the comparison unit 502 outputs the larger of the two outputs, the maximum value among the multiplication results sequentially output by the multiplication processing unit 406 is stored in the result holding unit 504. In this way, the result holding unit 504 can output the maximum value of the plurality of multiplication results. In other words, the result holding unit 504 can output the result of the maximum value filter processing for the maximum value pooling processing.

[0056] Thus, in one embodiment, the result holding unit 504 can hold the result of the convolution process, and can also hold the result of the filter process for the pooling process. Therefore, according to this embodiment, the memory for holding the result of the convolution process can be made common with the memory for holding the result of the filter process for the pooling process. For this reason, the circuit scale of the convolution processing unit 305 and the result holding unit 504 can be suppressed.

[0057] Also, in this embodiment, the addition and comparison processing unit 403 can perform the filter process for the max pooling process or the average pooling process in the same manner as the convolution process. Therefore, the maximum window size of the pooling process can be made the same as the maximum window size of the convolution process. For example, when the maximum value of the window size X×Y of the convolution process is 7×7, the maximum value of the window size X×Y of the pooling process can also be made 7×7. In one embodiment, the size of the window in the case of performing the pooling process is larger than 1×1. Thus, since the pooling process according to various window sizes can be performed, the flexibility of the pooling process using the convolution processing unit 305 can be enhanced. Further, in this embodiment, for the pooling process, the filter process and the stride process are performed separately. Therefore, the window size and the stride of the pooling process performed by the convolution processing unit 305 can be set separately. Thus, the convolution processing unit 305 can cope with the processing using networks having various structures.

[0058] The activation / stride processing unit 407 performs activation processing or stride processing. The activation / stride processing unit 407 can perform activation processing on the data output by the addition / comparison processing unit 403 (i.e., the data held by the result holding unit 504). For example, the activation / stride processing unit 407 can perform activation processing on the result of the convolution processing output by the addition / comparison processing unit 403. Also, the activation / stride processing unit 407 can perform stride processing on the result of the comparison processing output by the addition / comparison processing unit 403 without performing activation processing.

[0059] Also, the activation / stride processing unit 407 can perform stride processing on the data output by the addition / comparison processing unit 403 (i.e., the data held by the result holding unit 504). For example, the activation / stride processing unit 407 can perform stride processing on the result of the filter processing output by the addition / comparison processing unit 403. As described above, the data output from the result holding unit 504 corresponds to the pixel data of the image indicating the result of the convolution processing or the filter processing on the image of the previous layer. According to the stride processing, a part of the data output from the result holding unit 504 is extracted so that the image is reduced according to a predetermined stride.

[0060] The stride in the stride processing can be changed. For example, the activation / stride processing unit 407 may perform processing using two or more types of strides (e.g., R = 1 and any value other than R = 1). In this case, the stride may be set for each layer of the neural network. For example, in one embodiment, the stride in the case of performing pooling processing is greater than 1. Also, in one embodiment, even when performing convolution processing, stride processing by the activation / stride processing unit 407 is performed. On the other hand, in one embodiment, the stride in the stride processing in the case of performing convolution processing is 1.

[0061] In this way, the activation / stride processing unit 407 can output the result of the convolution process or the pooling process based on the data stored in the result holding unit 504. In this embodiment, the activation / stride processing unit 407 is a hardware circuit. In this specification, the activation / stride processing unit 407 may be referred to as an activation processing circuit or a stride processing circuit. However, instead of the activation / stride processing unit 407, a combination of a circuit that performs the activation process and a circuit that performs the stride process may be used. Also, at least one of the activation process and the stride process may be performed by the control unit 401 or the CPU 306.

[0062] Next, a processing method performed by the convolution processing unit 305 will be described with reference to FIG. 1. In S101, the control unit 401 reads a plurality of input images, weight coefficients used in the convolution process, and network structure information from the RAM 308 and stores them in the data holding unit 408. The control unit 401 may read feature data of a feature image such as a feature image obtained by processing in the neural network from the RAM 308 and store it in the data holding unit 408.

[0063] In S102, a loop for each layer starts. In the process according to FIG. 1, first, the first layer (layer 1) can be selected. In this case, layer 1 is the target layer to be processed.

[0064] In S103, the control unit 401 sets a convolution process or a filter process according to the network structure information held in the data holding unit 408. In this example, the network structure information instructs to perform a convolution process to generate a feature image of layer 1. Therefore, when the target layer is layer 1, the control unit 401 can set a convolution process. Also, the network structure information instructs to perform a pooling process to generate a feature image of layer 2. Therefore, when the target layer is layer 2, the control unit 401 can set a filter process.

[0065] In S104, a loop for each pixel of the output feature image starts. The output feature image is each of a plurality of feature images of the target layer. Hereinafter, a process for calculating the feature data of the pixel (i, j) of the n-th feature image of the target layer will be described. By repeating such a process, the feature data for each pixel of the output feature image is calculated in order. Also, by performing such a process for each feature image of the target layer, a plurality of feature images of the target layer are generated.

[0066] In S105, the control unit 401 initializes the processing result held in the addition and comparison processing unit 403. The control unit 401 can perform the initialization by setting the value held by the result holding unit 504, which will be described later, to zero.

[0067] In S106, a loop for each input feature image starts. The input feature image is a feature image of the previous layer of the target layer used to calculate the feature data of each pixel of the output feature image. When performing convolution processing, an output of one channel can be generated based on inputs of a plurality of channels. Therefore, when convolution processing is set in S103, in order to calculate the feature data of each pixel of the output feature image (for example, feature image (2, 1)), each of a plurality of feature images of the previous layer (for example, feature images (1, 1) to (1, 3)) can be used. For this reason, a loop using each of a plurality of feature images of the previous layer is repeated. On the other hand, in the case of pooling processing (for example, max pooling processing or average pooling processing), an output result of one channel is generated based on an input of one channel. Therefore, when pooling processing is set in S103, in order to calculate the feature data of each pixel of the output feature image (for example, feature image (3, 1)), one feature image of the previous layer (for example, feature image (2, 1)) is used. For this reason, there is one input feature image and the loop is performed once. Hereinafter, a case where a process using the m-th feature image of the previous layer is performed will be described.

[0068] In S107, the control unit 401 reads out partial feature data of the input feature image used in the process in S108 from the data holding unit 408 and transfers it to the feature data holding unit 402. For example, the control unit 401 can transfer the feature data (I i-(X-1) / 2,j-(Y-1) / 2 (m) to I i+(X-1) / 2,j+(Y-1) / 2 (m)) within the window of the convolution process. Also, the control unit 401 reads out some weight coefficients used in the process in S108 from the data holding unit 408 and transfers them to the coefficient holding unit 404. For example, the control unit 401 can transfer the weight coefficient C x,y (m,n) to the coefficient holding unit 404.

[0069] In S108, the coefficient setting unit 405 sets a plurality of coefficients used by the addition / comparison processing unit 403. The coefficient setting unit 405 can set coefficients according to a control signal from the control unit 401. The coefficient setting unit 405 can set the weight coefficient C’ x,y (m,n) according to Equation (5). In the present embodiment, the coefficient setting unit 405 sets the weight coefficient C’ according to whether maximum value pooling processing (Max pooling), average value pooling processing (Average pooling), or convolution processing (Convolution) is performed. As described above, the coefficient setting unit 405 may set the weight coefficient C’ x,y (m,n) held by the coefficient holding unit 404 by adjusting the weight coefficient C x,y (m,n).

Equation

[0070] In addition, when performing pooling processing, the weight coefficient C’ x,y (m,n) may be set as follows. That is, when n = m, the weight coefficient C’ x,y (m,n) = 1 (max pooling processing) or 1 / XY (average pooling processing). Also, when n ≠ m, the weight coefficient C’ x,y (m,n) = 0. By using such weight coefficients, it is possible to unify the loop processing for each input feature image starting from S106 when pooling processing is set and when convolution processing is set. That is, in this case, even when pooling processing is set, the loop that uses each of the plurality of feature images of the previous layer is repeated. Also, when n ≠ m, in order to reduce the processing time, the loop processing can be omitted.

[0071] In S109, the multiplication processing unit 406 performs a multiplication process on the input feature data I(m) held by the feature data holding unit 402 and the weight coefficient C’ x,y (m,n) held by the coefficient holding unit 404. Here, the multiplication processing unit 406 sequentially outputs the multiplication results of each of the plurality of feature data and the corresponding weight coefficients. Specifically, the multiplication processing unit 406 can output the product of the feature data I i+x-(X-1) / 2,j+y-(Y-1) / 2 (m) and the corresponding weight coefficient C’ x,y (m,n).

[0072] Also, the addition / comparison processing unit 403 integrates the multiplication results sequentially output by the multiplication processing unit 406 or selects one of the multiplication results sequentially output by the multiplication processing unit 406 according to the control signal from the control unit 401. The output of the addition / comparison processing unit 403 can be represented, for example, by Equation (6).

Equation

[0073] As already described, in response to an instruction to perform a convolution process to generate a feature image, the multiplication processing unit 406 sequentially outputs the multiplication results of each of the plurality of data within the window of the convolution process and the corresponding convolution coefficient. This convolution coefficient is the weight coefficient C x,y (m,n). Also, in this case, the selection unit 503 outputs the output of the addition unit 501 to the result holding unit 504.

[0074] Also, in response to an instruction to perform an average pooling process to generate a feature image, the multiplication processing unit 406 sequentially outputs the multiplication results of each of the plurality of data within the window of the average pooling process and the same coefficient. This coefficient is 1 / XY. Also, in this case, the selection unit 503 outputs the output of the addition unit 501 to the result holding unit 504.

[0075] Also, in response to an instruction to perform a max pooling process to generate a feature image, the multiplication processing unit 406 sequentially outputs the multiplication results of each of the plurality of data within the window of the max pooling process and the same convolution coefficient. This coefficient is 1. Also, in this case, the selection unit 503 outputs the output of the comparison unit 502 to the result holding unit 504.

[0076] The process of S109 can be performed according to S115 to S121. In S115, a loop for each piece of feature data starts. This loop can be performed for each of the feature data (I i-(X-1) / 2,j-(Y-1) / 2 (m) to I i+(X-1) / 2,j+(Y-1) / 2 (m)) of the input feature image. Specifically, the following loop can be performed for each combination of x (0 to X - 1) and y (0 to Y - 1).

[0077] In S116, the multiplication processing unit 406 outputs the multiplication result of the feature data I i+x-(X-1) / 2,j+y-(Y-1) / 2 (m) and the corresponding weight coefficient C’ x,y (m,n).

[0078] In S117, the control unit 401 selects either cumulative addition processing or comparison processing based on the network structure information held in the data holding unit 408. When performing convolution processing or average pooling processing to generate the feature image of the target layer, the control unit 401 selects cumulative addition processing. In this case, the process proceeds to S118. Otherwise, the control unit 401 selects comparison processing. In this case, the process proceeds to S119.

[0079] In S118, the addition / comparison processing unit 403 performs addition processing. That is, the addition unit 501 outputs the sum of the output of the multiplication processing unit 406 and the output of the result holding unit 504 as described above. Also, the control unit 401 controls the selection unit 503 to select the output of the addition unit 501.

[0080] In S119, the addition / comparison processing unit 403 performs comparison processing. In the present embodiment, the comparison unit 502 selects the larger of the output of the multiplication processing unit 406 and the output of the result holding unit 504. Also, the control unit 401 controls the selection unit 503 to select the output of the comparison unit 502.

[0081] In S120, the processing result held in the result holding unit 504 is replaced with the processing result selected by the selection unit 503.

[0082] In S121, the loop for each piece of feature data ends. As a result, when the control unit 401 selects cumulative addition processing, the multiplication result output in S116 is accumulated in the result holding unit 504. Also, when the control unit 401 selects comparison processing, the maximum of the multiplication results output in S116 after initialization of the result holding unit 504 is stored in the result holding unit 504.

[0083] In S110, the control unit 401 determines whether the loop for each input feature image has ended. When the processing for all the input feature images used to calculate the feature data of each pixel of the output feature image has been completed, the process proceeds to S111. Otherwise, the process returns to S107. Thereafter, the processing for the next input feature image starts.

[0084] When S110 ends, the result holding unit 504 of the addition and comparison processing unit 403 stores the result of the convolution process or the filtering process corresponding to the pixel (i, j) of the n-th feature image of the target layer. Specifically, the result holding unit 504 stores the value O i,j (n). The variable M indicates the number of input feature images used to calculate the feature data of each pixel of the output feature image. When performing the convolution process, M is an arbitrary positive integer. In this case, the result of Equation (7) is the same as the result of Equation (1). Also, when performing the pooling process, since the weight coefficient with m≠n is zero, it is the same as the result of substituting n with m in Equation (6). In this case, the result of Equation (7) is the same as the result of Equation (2) or (3). Thus, according to S106~S110, the result of the convolution process or the filtering process can be obtained.

Number

[0085] In S111, the activation and stride processing unit 407 can perform processing on the data output by the addition and comparison processing unit 403 according to the control signal from the control unit 401. For example, the activation and stride processing unit 407 can perform activation processing on the convolution processing result held in the result holding unit 504 according to the network structure information. Specifically, the activation and stride processing unit 407 can perform activation processing according to Equation (8).

Number

[0086] In addition, the activation / stride processing unit 407 can perform stride processing on the convolution processing result held in the result holding unit 504. The control unit 401 can control the stride of the stride processing according to the network structure information. For example, the control unit 401 can set the stride of the stride processing to 1 in response to an instruction to perform convolution processing to generate a feature image. Also, the control unit 401 can set the stride of the stride processing to 2 or more in response to an instruction to perform pooling processing to generate a feature image. The activation / stride processing unit 407 may extract the pixels of the feature image at equal intervals in the spatial direction according to the network structure information. For example, the activation / stride processing unit 407 can adjust the size of the output feature image by calculating the stride processing result according to Equation (4). Specifically, when the pixel (i, j) is not a pixel to be extracted, the activation / stride processing unit 407 may discard the convolution processing result held in the result holding unit 504.

[0087] Note that it is not essential for the activation / stride processing unit 407 to perform activation processing or stride processing. The activation / stride processing unit 407 can perform these processes as necessary. Also, the activation / stride processing unit 407 may perform both activation processing and stride processing in one layer. Also, the activation processing or the stride processing may be performed after the loop for all the pixels of one output feature image has ended, or after the loop for all the pixels of all the output feature images has ended.

[0088] In S112, the control unit 401 stores the processing result of the activation / stride processing unit 407 (or the output from the addition / comparison processing unit 403) in the feature data holding unit 402. The data stored in the feature data holding unit 402 can be treated as the feature data of the pixel (i, j) of the feature image m of the layer of interest.

[0089] In S113, the control unit 401 determines the end of the loop for each pixel of the output feature image. When the processing for all pixels of all output feature images is completed, the process proceeds to S114. At this time, a plurality of feature images of the target layer are stored in the feature data holding unit 402. If not, the process returns to S105. In this case, the processing for the pixels of the next output feature image starts.

[0090] In this example, for each of the plurality of windows set on the feature image of the previous layer so that the stride is 1, the result of the convolution process or the result of the filter process used in the pooling process is obtained. The control unit 401 controls the multiplication processing unit 406 and the selection unit 503 so that such a result is obtained. Specifically, the control unit 401 can control the supply of the feature data to the multiplication processing unit 406 so that such a result is obtained.

[0091] In S114, the control unit 401 determines the end of the loop for each layer. When the processing for all layers is completed, the control unit 401 writes the feature image of the final layer held by the feature data holding unit 402 to the RAM 308. Thus, the processing in the neural network ends. If not, the process returns to S103. In this case, the processing target layer is changed, and the processing for the next layer starts.

[0092] According to the above embodiment, the pooling process is divided into a filter process and a stride process. Then, the convolution process and the filter process are performed using a common arithmetic circuit (convolution processing unit 305). Therefore, the processing can be efficiently performed while suppressing an increase in the circuit scale.

[0093] Thus, in one embodiment, the result holding unit 504 holds the output result of the addition unit 501 or the comparison unit 502. Therefore, it is possible to reduce the circuit scale of the convolution processing device that can perform both the convolution process and the pooling process. Also, in one embodiment, the common feature data holding unit 402 holds the data used for the convolution process and the pooling process. Therefore, it is possible to improve the efficiency of the convolution process and the pooling process with a small circuit scale as compared with a configuration in which the convolution process and the pooling process are performed by different arithmetic units. Furthermore, according to the present embodiment, it becomes possible to set different values for the window size and the stride of the pooling process. Therefore, the convolution processing device according to the present embodiment has an advantage that there are many variations of networks that can be supported.

[0094] In the above embodiment, the coefficient setting unit 405 sets the coefficient according to Equation (5). However, before performing the neural network process, the coefficient according to Equation (5) may be prepared. For example, in S101, the control unit 401 may read the weight coefficient from the RAM 308, adjust the weight coefficient according to Equation (5), and store the adjusted weight coefficient in the data holding unit 408. Also, before the start of the process according to FIG. 1, the adjusted weight coefficient may be stored in the RAM 308. For example, the CPU 306 may store the weight coefficient according to Equation (5) in the RAM 308. In this case, in S107, the control unit 401 can read the adjusted weight coefficient from the data holding unit 408 and transfer it to the feature data holding unit 402. Also, the process of S108 can be omitted. According to such a configuration, in this case, the convolution processing unit 305 may not have the coefficient setting unit 405. Therefore, the circuit scale of the convolution processing unit 305 can be further reduced.

[0095] (Modification example) The above-described convolution processing unit 305 can perform a filter process different from the convolution process and the pooling process on the image. By controlling the weight coefficient supplied to such a multiplication processing unit 406 and the operation of the selection unit 503, various filter processes can be realized.

[0096] Using the above convolution processing unit 305, post-processing on the result of the neural network processing may be performed. That is, the convolution processing unit 305 can perform further filter processing on the image obtained by the convolution processing or the pooling processing on the image. For example, as described above, the information processing apparatus can generate a reliability map (feature image 204) using the convolution processing unit 305. Further, the information processing apparatus may perform peak value detection processing on the feature image 204. For example, the information processing apparatus can set all pixel values other than the peak value in the feature image 204 to zero. In this way, the information processing apparatus can delete overlapping detection results in the feature image 204. In this modification example, such peak value detection processing is efficiently performed by using the convolution processing unit 305 used in the neural network processing.

[0097] The convolution processing unit 305 may perform a part of the post-processing. In the following example, the convolution processing unit 305 performs a part of the peak value detection processing. Specifically, the convolution processing unit 305 can apply a maximum value filter to the feature image A901 according to Equation (2). Also, another processing unit such as the CPU 306 performs at least a part of the processing using the image obtained by the convolution processing or the pooling processing on the image. In this example, the CPU 306 performs the remaining part of the peak value detection processing.

[0098] FIG. 9 shows an example of the structure of a neural network used for performing peak value detection processing. The number of layers of this network is 1. This network has a layer A. There is one feature image in layer A. This one feature image includes a plurality of feature data (for example, reliability) as pixel values. In layer A, a feature image B902 is generated by applying a maximum value filter to the feature image A901. Thereafter, by referring to the feature image A901 and the feature image B902, a feature image C903 indicating the result of the peak detection processing is generated.

[0099] FIG. 8 shows a flowchart of the post-processing in this modified example. S801 to S810 correspond to a part of the processing shown in FIG. 1. This means that the processing shown in FIG. 8 can be performed using the convolution processing unit 305. The feature image A901 can be generated by the processing described with reference to FIG. 1.

[0100] In S801, the control unit 401 reads out the feature image A901, the weight coefficient, and the network structure information from the RAM 308 and stores them in the data holding unit 408. However, the weight coefficient read from the RAM 308 is not used for the post-processing described here. Therefore, the control unit 401 can store an arbitrary weight coefficient in the data holding unit 408.

[0101] In S802, the control unit 401 sets the filter processing according to the network structure information held in the data holding unit 408.

[0102] In S803, the control unit 401 initializes the processing result held in the addition / comparison processing unit 403 in the same manner as S105.

[0103] In S804, a loop for each pixel of the output feature image starts. In this example, the output feature image is the feature image B902. Hereinafter, the process of calculating the feature data of the pixel (i, j) of the n-th feature image of the target layer will be described. By repeating such a process, the feature data for each pixel of the output feature image is calculated in order.

[0104] In S805, the control unit 401 reads out a part of the feature data of the input feature image used for the processing in S806 from the data holding unit 408 and transfers it to the feature data holding unit 402 in the same manner as S107. Also, the control unit 401 reads out the weight coefficient from the data holding unit 408 and transfers it to the coefficient holding unit 404.

[0105] In S806, the coefficient setting unit 405 sets a plurality of coefficients used by the addition / comparison processing unit 403. In this example, the coefficient setting unit 405 adjusts the weight coefficient according to Equation (9). [Number]

[0106] In S807, the multiplication processing unit 406, similar to S109, multiplies the input feature data I(m) held by the feature data holding unit 402 by the weight coefficient C’ x,y (m,n) held by the coefficient holding unit 404. Further, the addition / comparison processing unit 403, similar to S109, outputs the maximum value among the multiplication results sequentially output by the multiplication processing unit 406. That is, the selection unit 503 outputs the output of the comparison unit 502 to the result holding unit 504. In this example, the window size is 3×3.

[0107] In S808, the control unit 401 stores the output from the addition / comparison processing unit 403 in the feature data holding unit 402. The data stored in the feature data holding unit 402 can be treated as the feature data of the pixel (i,j) of the feature image B902.

[0108] In S809, the control unit 401 determines the end of the loop for each pixel of the output feature image. When the processing for all pixels of the output feature image is completed, the process proceeds to S810. At this time, the feature image B902 is stored in the feature data holding unit 402 as the output feature image. Otherwise, the process returns to S805. In this case, the processing for the next pixel of the output feature image starts.

[0109] In S810, the control unit 401 writes the feature image B902 held in the feature data holding unit 402 to the RAM 308.

[0110] In S811, the CPU 306 reads out the feature image A901 and the feature image B902 held in the RAM 308. Then, the CPU 306 calculates the feature image C903 according to Equation (10). In S811, different from the convolution process and the filter process, the general-purpose CPU 306 performs the process. [Number] In Equation (10), IC represents the feature image C903. IA represents the feature image A901 before filter processing. The variable IB represents the feature image B902 after filter processing. The variables i and j represent the coordinates of the feature image. Finally, the CPU 306 stores the obtained feature image C903 in the RAM 308. Thus, the post-processing is completed.

[0111] Figures 7(A) to (C) show examples of post-processing. In Figures 7(A) to (C), pixels having zero values are represented in black. The feature image A901 has an 8×8 image size. The feature image B902 obtained as described above corresponds to the result of applying a 3×3 maximum value filter to the feature image A901. And in the feature image A901, two pixels (values are 10 and 15) having pixel values larger than surrounding pixels are shown in the feature image C903. These pixels are treated as peaks. The peak here means a pixel having the maximum value within a 3×3 pixel range.

[0112] In this modified example, part of the peak detection process is performed using the convolution processing unit 305. Therefore, the processing speed can be increased compared to the case where the CPU 306 performs all of the peak detection process. On the other hand, by using the convolution processing unit 305 used for the convolution process in the post-processing, the circuit scale can be reduced compared to the case of providing a dedicated circuit for the post-processing.

[0113] (Other Embodiments) In the above embodiment, the stride is represented by one variable R. However, it is not essential to represent the stride by one variable. For example, as in Equation (11), the stride may be expressed by two variables. [Number] In Equation (11), the variable R is the horizontal stride, and the variable S is the vertical stride. In this case, the size (number of pixels) of the feature image after the stride process becomes 1 / (R × S) of that before the process. In this specification, the fact that the stride is greater than 1 means that at least one of the variables R and S is greater than 1.

[0114] Also, the type of the pooling process is not limited to the maximum value pooling process and the average value pooling process. For example, the convolutional processing unit 305 may perform the minimum value pooling process. Equation (12) represents the minimum value pooling process (Min Pooling). The window size is X × Y.

Equation

[0115] When performing the minimum value pooling process, the comparison unit 502 of the addition / comparison processing unit 403 outputs the smaller value instead of the larger value between the output of the multiplication processing unit 406 and the output of the result holding unit 504. In S105, the control unit 401 can perform initialization by setting the value held by the result holding unit 504 to the theoretical maximum value of the output of the multiplication processing unit 406 (2 X −1 in the case of unsigned X bits). Other processes are the same as those in the case of performing the maximum value pooling process.

[0116] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiment to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0117] The disclosure of this specification includes the following processing apparatuses. (Item 1) A multiplication circuit that sequentially outputs multiplication results of each data among a plurality of data and corresponding coefficients, A memory, An adder circuit that adds the multiplication result output by the multiplication circuit and the data held by the memory, and outputs the addition result. A comparison circuit that compares the multiplication result output by the multiplication circuit and the data held by the memory, and outputs one of the multiplication result output by the multiplication circuit and the data held by the memory. A selection circuit that outputs one of the output of the adder circuit and the output of the comparison circuit to the memory so that the memory holds it. A processing device, characterized by comprising the above. (Item 2) For each of the plurality of data, the selection circuit repeatedly performs an operation of outputting one of the output of the adder circuit and the output of the comparison circuit to the memory. The processing device according to item 1, characterized in that after the selection circuit repeatedly performs the operation of outputting for each of the plurality of data, the data held by the memory is output as a processing result for the plurality of data. (Item 3) The processing device according to any one of items 1 to 2, further comprising an activation processing circuit that performs activation processing on the data held by the memory. (Item 4) The processing device can perform convolution processing and pooling processing on an image. The multiplication circuit sequentially outputs multiplication results of each data among a plurality of data included in a window of the convolution processing or the pooling processing set on the image and corresponding coefficients. For each of the plurality of data, the selection circuit repeatedly performs an operation of outputting one of the output of the adder circuit and the output of the comparison circuit to the memory. After the selection circuit repeatedly performs the operation of outputting for each of the plurality of data, the memory outputs the data held by the memory as a result of convolution processing for the window or a result of filter processing for the pooling processing. The processing device according to any one of items 1 to 3, characterized by the above. (Item 5) The processing device according to item 4, characterized in that the processing device can perform two or more types of pooling processing. (Item 6) The data output from the memory corresponds to the pixel data of an image indicating the result of the convolution processing or the filter processing on the image, The processing device according to any one of items 4 to 5, further comprising a stride processing circuit that performs a stride processing of extracting a part of the data output from the memory so that the image is reduced according to a predetermined stride. (Item 7) The processing device according to item 6, characterized in that the stride can be changed. (Item 8) The processing device according to any one of items 6 to 7, characterized in that when performing the pooling processing, the size of the window is larger than 1×1 and the stride is larger than 1. (Item 9) The processing device performs processing in a neural network, The processing device according to any one of items 6 to 8, characterized in that the stride is set for each layer of the neural network. (Item 10) Further comprising control means, The control means controls the multiplication circuit and the selection circuit so as to output the result of the convolution processing or the result of the filter processing for the pooling processing for each of a plurality of windows set on the image such that the stride becomes 1. The control means sets the stride of the stride processing to 1 in response to an instruction to perform the convolution processing. The control means sets the stride of the stride processing to 2 or more in response to an instruction to perform the pooling processing. The processing device according to any one of items 6 to 9, characterized in that. (Item 11) The processing device according to any one of items 4 to 10, further comprising setting means for setting coefficients corresponding to the respective data according to whether convolution processing or pooling processing is to be performed on the image. (Item 12) In response to an instruction to perform the convolution processing, the multiplication circuit sequentially outputs multiplication results of each of the plurality of data within the window of the convolution processing and the corresponding convolution coefficient, the selection circuit outputs the output of the addition circuit to the memory The processing device according to any one of items 4 to 11, characterized in that. (Item 13) In response to an instruction to perform average pooling processing, the multiplication circuit sequentially outputs multiplication results of each of the plurality of data within the window of the average pooling processing and the same coefficient, the selection circuit outputs the output of the addition circuit to the memory The processing device according to any one of items 4 to 12, characterized in that. (Item 14) In response to an instruction to perform maximum or minimum pooling processing, the multiplication circuit sequentially outputs multiplication results of each of the plurality of data within the window of the maximum or minimum pooling processing and the same coefficient, the selection circuit outputs the output of the comparison circuit to the memory The processing device according to any one of items 4 to 13, characterized in that. (Item 15) The processing device can perform a filter process different from the convolution process and the pooling process on the image, The processing device according to any one of items 4 to 14, characterized in that the filter process is performed on the image obtained by the convolution process or the pooling process on the image. (Item 16) The processing device further includes a processing means, and the processing means performs at least a part of processing using an image obtained by the convolution processing or the pooling processing on the image, and is the processing device according to any one of items 4 to 15. (Item 17) The processing device according to item 16, wherein the processing using the image is peak detection processing. (Item 18) The processing device performs processing in a neural network, The selection circuit outputs, to the memory, one of the output of the addition circuit and the output of the comparison circuit selected for each layer of the neural network, and is the processing device according to any one of items 1 to 17. (Item 19) The processing device according to item 18, wherein the selection circuit outputs, to the memory, one of the output of the addition circuit and the output of the comparison circuit selected according to the type of processing performed in the layer. (Item 20) A processing device capable of performing convolution processing and pooling processing, Setting means for setting a plurality of coefficients according to whether convolution processing or pooling processing is performed; A multiplication circuit that sequentially outputs multiplication results of each data among a plurality of data and corresponding coefficients; A memory that holds an integrated result of each multiplication result output by the multiplication circuit or data indicating a multiplication result selected from each multiplication result; Output means for outputting a result of the convolution processing or the pooling processing based on data stored in the memory; A processing device characterized by including the above.

[0118] The invention is not limited to the above embodiments, and various changes and modifications are possible without departing from the spirit and scope of the invention. Therefore, claims are attached to disclose the scope of the invention.

Explanation of Reference Numerals

[0119] 401: Control Unit, 402: Feature Data Holding Unit, 403: Addition / Comparison Processing Unit, 404: Coefficient Holding Unit, 405: Coefficient Setting Unit, 406: Multiplication Processing Unit, 407: Activation / Stride Processing Unit, 408: Data Holding Unit, 501: Addition Unit, 502: Comparison Unit, 503: Selection Unit, 504: Result Holding Unit

Claims

1. A multiplication circuit that sequentially outputs the multiplication results of each of a plurality of data and corresponding coefficients; A memory; An addition circuit that adds the multiplication result output by the multiplication circuit and the data held by the memory and outputs the addition result; A comparison circuit that compares the multiplication result output by the multiplication circuit and the data held by the memory and outputs one of the multiplication result output by the multiplication circuit and the data held by the memory; A selection circuit that outputs one of the output of the addition circuit and the output of the comparison circuit to the memory so that the memory holds it; A processing device comprising the same.

2. The selection circuit repeats the operation of outputting one of the output of the addition circuit and the output of the comparison circuit to the memory for each of the plurality of data, After the selection circuit repeats the operation of outputting for each of the plurality of data, the data held by the memory is output as a processing result for the plurality of data. The processing device according to claim 1.

3. The processing device according to claim 1, further comprising an activation processing circuit that performs activation processing on the data held by the memory.

4. The processing device can perform convolution processing and pooling processing on an image, The multiplication circuit sequentially outputs the multiplication results of each of a plurality of data included in a window of the convolution processing or the pooling processing set on the image and corresponding coefficients, The selection circuit repeats the operation of outputting one of the output of the addition circuit and the output of the comparison circuit to the memory for each of the plurality of data, After the selection circuit repeats the operation of outputting for each of the plurality of data, the memory outputs the data held by the memory as a result of convolution processing for the window or a result of filter processing for the pooling processing. The processing device according to claim 1.

5. The processing device according to claim 4, characterized in that the processing device can perform two or more types of pooling processing.

6. The data output from the memory corresponds to the pixel data of an image indicating the result of the convolution processing or the filter processing on the image. The processing device according to claim 4, further comprising a stride processing circuit that performs a stride processing for extracting a part of the data output from the memory so that an image is reduced according to a predetermined stride.

7. The processing device according to claim 6, wherein the stride is changeable.

8. The processing device according to claim 6, wherein, when performing the pooling processing, the size of the window is larger than 1×1, and the stride is larger than 1.

9. The processing device performs processing in a neural network, The processing device according to claim 6, wherein the stride is set for each layer of the neural network.

10. Further comprising control means, The control means controls the multiplication circuit and the selection circuit so as to output the result of the convolution processing or the result of the filter processing for the pooling processing for each of a plurality of windows set on the image such that the stride becomes 1. The control means sets the stride of the stride processing to 1 in response to an instruction to perform the convolution processing. The control means sets the stride of the stride processing to 2 or more in response to an instruction to perform the pooling processing. The processing device according to claim 6, characterized in that.

11. The processing device according to claim 4, further comprising setting means for setting a coefficient corresponding to each data according to whether convolution processing or pooling processing is performed on an image.

12. In response to an instruction to perform the convolution processing, The multiplication circuit sequentially outputs multiplication results of each data among a plurality of data in the window of the convolution processing and the corresponding convolution coefficient. The selection circuit outputs the output of the addition circuit to the memory. The processing device according to claim 4, characterized in that.

13. In response to an instruction to perform average value pooling processing, The multiplication circuit sequentially outputs multiplication results of each data among a plurality of data in the window of the average value pooling processing and the same coefficient. The selection circuit outputs the output of the addition circuit to the memory. The processing device according to claim 4, characterized in that.

14. In response to an instruction to perform maximum value or minimum value pooling processing, The multiplication circuit sequentially outputs the multiplication results of each of a plurality of data within the window of the maximum value or minimum value pooling process and the same coefficient. The selection circuit outputs the output of the comparison circuit to the memory. The processing device according to claim 4, characterized in that.

15. The processing device can perform a filtering process different from the convolution process and the pooling process on the image. The processing device according to claim 4, characterized in that the filtering process is performed on the image obtained by the convolution process or the pooling process on the image.

16. The processing device further includes processing means, and the processing means performs at least a part of the process using the image obtained by the convolution process or the pooling process on the image. The processing device according to claim 4, characterized in that.

17. The processing device according to claim 16, characterized in that the process using the image is a peak detection process.

18. The processing device performs processing in a neural network. The selection circuit according to claim 1, characterized in that one of the output of the addition circuit and the output of the comparison circuit selected for each layer of the neural network is output to the memory.

19. The selection circuit according to claim 18, characterized in that one of the output of the addition circuit and the output of the comparison circuit selected according to the type of process performed in the layer is output to the memory.

20. A processing device capable of performing a convolution process and a pooling process, Setting means for setting a plurality of coefficients according to whether to perform the convolution process or the pooling process; A multiplication circuit that sequentially outputs the multiplication results of each of a plurality of data and the corresponding coefficients; A memory that holds the integrated result of each multiplication result output by the multiplication circuit or data indicating the multiplication result selected from each multiplication result; Output means for outputting the result of the convolution process or the pooling process based on the data stored in the memory; A processing device characterized by comprising.

Citation Information

Patent Citations

  • Vector computation unit in neural network processor

    JP2020017281A

  • Systems and methods for hardware-based pooling

    JP2021509747A