Image processing device

By combining the first processor and the second processor in the image processing device, image processing is performed using the results of the recurrent neural network computing in the cache, the problems of long waiting time and high cost in the image processing of CNN are solved, and a low-cost and efficient image processing effect is achieved.

CN113409182BActive Publication Date: 2025-06-20KK TOSHIBA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010893923.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-17
Filing Date
2020-08-31
Publication Date
2025-06-20
Estimated Expiration
2040-08-31

AI Technical Summary

Technical Problem

In the prior art, convolutional neural networks (CNNs) have a large waiting time and high cost in image processing, making it difficult to achieve low-cost and efficient image processing.

Method used

An image processing device is designed, and a first processor is combined with a second processor. A cache is provided in the first processor, and the second processor uses multiple pixel data of the image data and the recurrent neural network calculation results in the cache to perform recurrent neural network calculation.

Benefits of technology

By converting image data into stream data and performing RNN operations sequentially, neural network computing processing with small waiting time and low cost is realized, which improves the efficiency and cost-effectiveness of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113409182B_ABST
    Figure CN113409182B_ABST
Patent Text Reader

Abstract

An embodiment provides an image processing apparatus with a small waiting time and capable of being implemented at low cost. The image processing apparatus according to the embodiment includes: an image processing processor (11) to which image data is input; a state buffer (21) provided in the image processing processor (11); and a recurrent neural network processor (22) that performs a recurrent neural network operation using at least one of a plurality of pixel data of the image data and an operation result of a recurrent neural network operation stored in the state buffer (21).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority based on Japanese Patent Application No. 2020-46914 (filing date: March 17, 2020), the entire content of which is incorporated herein by reference. Technical Field

[0002] Embodiments of the present invention relate to an image processing apparatus. Background Art

[0003] There is a technique for performing recognition processing and the like on image data by a neural network. For example, in a kernel operation in a Convolutional Neural Network (CNN; hereinafter referred to as CNN), after holding the entire image data of an image in a frame buffer in an off-chip memory such as a DRAM, a window of a specified size is slid relative to the entire held image data while performing the operation.

[0004] Therefore, since it takes time to save the entire image data to the off-chip memory and to access the off-chip memory for writing and reading the feature map for each kernel operation, the latency of CNN operations is large. In a device such as an Image Signal Processor, it is preferable that the latency is small.

[0005] In order to reduce the latency of CNN operations, a line buffer smaller than the size of the frame buffer may be used. However, since accesses to the line buffer for kernel operations occur frequently, a memory capable of high-speed access needs to be used in the line buffer, and the cost of the image processing apparatus becomes high. Summary of the Invention

[0006] An object of the present invention is to provide an image processing apparatus with a small latency and capable of being implemented at low cost.

[0007] The image processing apparatus according to the aspect of the invention includes: a first processor to which image data is input; a buffer provided in the first processor; and a second processor that performs the recurrent neural network operation using at least one of a plurality of pixel data of the image data and an operation result of the recurrent neural network operation stored in the buffer. Brief Description of the Drawings

[0008] Figure 1 It is a block diagram of an image processing apparatus according to an embodiment.

[0009] Figure 2 It is a diagram for explaining the processing content of an image processing processor according to an embodiment.

[0010] Figure 3It is a block diagram showing the structure of an image processing processor according to an embodiment.

[0011] Figure 4 It is a structural diagram of a recurrent neural network unit processor according to an embodiment.

[0012] Figure 5 It is a diagram for explaining the transformation from input image data to stream data according to an embodiment.

[0013] Figure 6 It is a diagram for explaining the processing order of the recurrent neural network unit for multiple pixel values included in the input image data according to an embodiment.

[0014] Figure 7 It is a diagram for explaining the processing order of the line end unit for the output value of the final column of each row according to Modification 1.

[0015] Figure 8 It is a diagram for explaining the processing order of the recurrent neural network unit for multiple pixel values included in the input image data according to Modification 2.

[0016] Figure 9 It is a diagram for explaining the receptive field of a convolutional neural network.

[0017] Figure 10 It is a diagram for explaining the receptive field of an embodiment.

[0018] Figure 11 It is a diagram for explaining the difference in the range of the receptive field between a convolutional neural network and a recurrent neural network.

[0019] Figure 12 It is a diagram for explaining the input stride of the recurrent neural network unit according to Modification 2.

[0020] Figure 13 It is a diagram for explaining the set range of the receptive field according to Modification 2. Detailed Embodiment

[0021] Hereinafter, an embodiment will be described with reference to the drawings.

[0022] (Structure)

[0023] Figure 1 It is a block diagram of an image processing apparatus according to the present embodiment. An image processing system 1 using the image processing apparatus of the present embodiment processes image data from a camera device, performs processing such as image recognition, and outputs information on the processing result.

[0024] The image processing system 1 includes an Image Signal Processor (ISP; hereinafter referred to as ISP) 11, an off-chip memory 12, and a processor 13.

[0025] The ISP 11 is connected to a camera device (not shown) through an interface that follows standards such as MIPI (Mobile Industry Processor Interface) CSI (Camera Serial Interface). The ISP 11 receives a captured signal from the image sensor 14 of the camera device, performs prescribed processing on the captured signal, and outputs the result data of the prescribed processing. That is, for the ISP 11 as a processor, multiple pixel data of image data are sequentially input. Here, the ISP 11 takes the captured signal (hereinafter referred to as input image data) IG from the image sensor 14 as a imaging element as input, and outputs image data (hereinafter referred to as output image data) OG as result data. For example, the ISP 11 performs noise removal, etc. on the input image data IG, and outputs output image data OG without noise, etc.

[0026] In addition, all of the input image data IG from the image sensor 14 can be input to the ISP 11, and the following-described RNN operation can be performed on all of the input image data IG, or the following-described RNN operation can be performed on a part of the input image data IG.

[0027] The ISP 11 includes a status buffer 21 and an RNN unit processor 22 that repeatedly performs a prescribed operation based on a Recurrent Neural Network (RNN; hereinafter referred to as RNN). The structure of the ISP 11 will be described later.

[0028] The off-chip memory 12 is a memory such as a DRAM. The output image data OG generated in the ISP 11 and output from the ISP 11 is stored in the off-chip memory 12.

[0029] The processor 13 performs recognition processing, etc. based on the output image data OG stored in the off-chip memory 12. The processor 13 outputs result data RD obtained from the recognition processing, etc. Thus, the ISP 11, the off-chip memory 12, and the processor 13 constitute, for example, an image recognition device (indicated by the dashed line of Figure 1 ) 2 that performs image recognition processing, etc. on an image.

[0030] Figure 2 is a diagram for explaining the processing content of the ISP 11. As Figure 2As shown, the ISP 11 uses an RNN cell processor 22 (described later) to perform a prescribed process such as noise removal on the input image data IG from the image sensor 14, and generates output image data OG.

[0031] For example, when the image recognition device 2 performs recognition processing or the like by the processor 13 based on the output image data OG, since the output image data OG is data from which noise has been removed, an improvement in the accuracy of recognition processing or the like in the processor 13 can be expected.

[0032] Figure 3 is a block diagram showing the structure of the ISP 11. Figure 4 is a structural diagram of the RNN cell processor 22. The ISP 11 includes a state buffer 21, an RNN cell processor 22, and a pixel stream decoder 23. The pixel stream decoder 23 is a circuit that transforms the input image data IG into stream data SD and outputs it to the RNN cell processor 22.

[0033] Figure 5 is a diagram for explaining the transformation from the input image data IG to the stream data SD. Here, to simplify the explanation, in Figure 5 the image of the input image data IG is composed of 6 rows of image data. Each row contains a plurality of pixel data. That is, the image is composed of multiple rows (here 6 rows) and multiple columns of pixel data.

[0034] If the pixel stream decoder 23 receives the input image data IG from the image sensor 14, it transforms the plurality of pixel data of the received input image data IG into stream data SD in a prescribed order.

[0035] The pixel stream decoder 23 generates, from the input image data IG, stream data SD composed of row data L1 from the pixel at the first column of the first row (i.e., the leftmost pixel in the uppermost row) to the pixel at the final column of the first row (i.e., the rightmost pixel in the uppermost row), followed by row data L1, row data L2 from the pixel at the first column of the second row (i.e., the leftmost pixel in the second row from the top) to the pixel at the final column of the second row (i.e., the rightmost pixel in the second row),..., and data column LL from the pixel at the first column of the sixth row as the final row (i.e., the leftmost pixel in the lowermost row) to the pixel at the final column of the sixth row (i.e., the rightmost pixel in the lowermost row), and outputs it.

[0036] Thus, the pixel stream decoder 23 is a circuit that transforms the input image data IG into stream data SD and outputs it to the RNN cell processor 22.

[0037] As Figure 4As shown, the RNN cell processor 22 is a processor including one RNN cell 31. The RNN cell 31 is a simple RNN cell, which is a hardware circuit that outputs the hidden state obtained by performing a specified operation on two input values IN1 and IN2 as two output values OUT1 and OUT2.

[0038] In addition, although the RNN cell processor 22 includes one RNN cell 31 here, it may also include two or more RNN cells 31. Alternatively, the number of RNN cells 31 may also be the same as the number of layers described later.

[0039] The input value IN1 of the RNN cell 31 is i l,t . l represents the layer and t represents the step. The input value IN2 of the RNN cell 31 is the hidden state h l,t-1 . The output value OUT1 of the RNN cell 31 is the hidden state h l,t , which is the input value IN1 (i.e., i l+1,t ) of the next step t of the next layer (l + 1). The output value OUT2 of the RNN cell 31 is the hidden state h l,t , which is the input value IN2 of the RNN cell 31 of the next step (t + 1) of the same layer.

[0040] The step t is also called the time step. It is a number that increases whenever one sequence data is input to the RNN and the hidden state is updated. It is assigned as an index for the hidden state and input / output, and is a virtual unit that is not necessarily the same as the actual time.

[0041] As Figure 3 shown, the RNN cell 31 can read various parameters (represented by dotted lines) used in the RNN operation from the off-chip memory 12 and hold them inside the RNN cell 31. The parameters include the weight parameter w and the bias value b of each RNN operation of each layer described later, etc.

[0042] In addition, the RNN cell 31 can also be implemented by software executed by a central processing unit (CPU).

[0043] The RNN cell 31 performs corresponding actions according to each layer described later. In the first layer (the first layer), the streaming data SD is sequentially input as the input value IN1 of the RNN cell 31. The RNN cell 31 performs a specified operation, generates the hidden state h l,t as the output values OUT1 and OUT2, and outputs them to the state buffer 21.

[0044] Each output value OUT1, OUT2 obtained in each layer is stored in a specified storage area in the state buffer 21. The state buffer 21 is, for example, a line buffer.

[0045] Since the state buffer 21 is provided in the ISP11, the RNN unit 31 can write data to and read data from the state buffer 21 at high speed. The RNN unit 31 stores the hidden state h obtained by performing a specified operation in the state buffer 21. The state buffer 21 is an SRAM including a line buffer and is a buffer that stores at least the number of streaming data.

[0046] The RNN unit 31 can perform multiple layer operations. Here, the RNN unit 31 can perform a first layer operation that performs a specified operation with the streaming data SD as an input, a second layer operation that performs a specified operation with the hidden state h as the operation result in the first layer as an input, and a third layer operation that performs a specified operation with the hidden state h as the operation result in the second layer as an input, and so on.

[0047] The specified operation in the RNN unit 31 will be described. In the l (letter "L")-th layer operation, the RNN unit 31, at a certain time step t, takes the input value IN1 as the pixel data i, uses the activation function tanh, which is a non-linear function, as the specified operation, and outputs the output values OUT1, OUT2. The output values OUT1, OUT2 are the hidden state ht. Here, as Figure 4 shown, the hidden state h l,t is calculated by the following equation (1).

[0048] h l,t = tanh(w l,ih i l,t + w l,hh h l,t-1 + b l )…(1)

[0049] Here, w l,ih and w l,hh are weight parameters represented by the following equations (2) and (3), respectively.

[0050] w l,ih ∈ R e×d …(2)

[0051] w l,h h ∈ R e×e …(3)

[0052] Here, R e×d and R e×e are spaces formed by real number matrices of e rows and d columns and e rows and e columns, respectively, and both represent matrices formed by real numbers.

[0053] In addition, the input value (pixel data i l,t ) and the output value (hidden state h l,t ) are respectively represented by the following formulas (4) and (5).

[0054] i l,t ∈R d …(4)

[0055] h l,t ∈R e …(5)

[0056] Here, R d represents a d-dimensional real number space, and R e represents an e-dimensional real number space, both representing vectors formed by real numbers.

[0057] The values of the respective weight parameters of the above non-linear function are optimized through the learning of the RNN.

[0058] The pixel data i l,t is an input vector, for example, a three-dimensional vector when an RGB image is input, and the number of its channels (channels) in the case of an intermediate feature map. The hidden state h l,t is an output vector. D and e respectively represent the dimensions of the input vector and the output vector. l is the layer number and is the index of the sequence data. B is the bias value.

[0059] In addition, in Figure 4 , the RNN unit 31 generates and outputs two output values OUT1 and OUT2 with the same value based on the input value IN1 and taking the output value from the previous pixel as the input value IN2, but the RNN unit 31 can also output two mutually different output values OUT1 and OUT2.

[0060] In the second-layer operation, the RNN unit 31 takes the input value IN1 as the output value OUT1 of the first layer, and uses the activation function tanh, which is a non-linear function as a specified operation, to output the output values OUT1 and OUT2.

[0061] When performing the layer operations of the third, fourth, etc. that follow the second-layer operation, in the layer operations of the third, fourth, etc., similar to the second-layer operation, the RNN unit 31 sets the input value IN1 as the output value OUT1 of the previous layer, and uses the activation function tanh, which is a non-linear function as a specified operation, to output the output values OUT1 and OUT2.

[0062] (Function)

[0063] Next, the operation of ISP11 will be described. Here, an example with three layers will be described. As described above, the pixel stream decoder 23 outputs the stream data SD( Figure 5 ), and this stream data SD arranges the input image data IG in the order of a plurality of pixel data from the pixel at the left end to the pixel at the right end of the first row L1, a plurality of pixel data from the pixel at the left end to the pixel at the right end of the second row L2,..., a plurality of pixel data from the pixel at the left end to the pixel at the right end of the data column LL (i.e., L6) of the last row (the order indicated by the arrow A).

[0064] In the first layer, the initial input value IN1 to the RNN unit 31 is the initial data of the stream data SD (i.e., the pixel at the first column of the first row of the input image data IG), and the input value IN2 is a prescribed default value.

[0065] In the first layer, if the RNN unit 31 is input with two input values IN1 and IN2 in the initial step t1, it performs a prescribed operation and outputs the output values OUT1 and OUT2. The output values OUT1 and OUT2 are stored in a prescribed storage area in the state buffer 21. The output value OUT1 of the first layer at the step t1 is read out from the state buffer 21 in the initial step t1 of the next second layer and used as the input value IN1 of the RNN unit 31. In the first layer, the output value OUT2 at the step t1 is used as the input value IN2 in the next step t2.

[0066] Similarly hereinafter, in the first layer, the output value OUT1 in each subsequent step is read out from the state buffer 21 in the corresponding step in the subsequent second layer and used as the input value IN1 of the RNN unit 31. In the first layer, the output value OUT2 in each subsequent step is read out from the state buffer 21 in the next step and used as the input value IN2 of the RNN unit 31.

[0067] If the prescribed operation for each pixel data of the stream data SD in the first layer ends, the processing of the second layer is executed.

[0068] If the prescribed operation for the first pixel data in the first layer ends, the processing corresponding to the first pixel of the second layer is executed.

[0069] In the second layer, a plurality of output values OUT1 obtained from the first to the last step in the first layer are used as the input value IN1 and sequentially input to the RNN unit 31. Similar to the processing in the first layer, in the order from the first step to the last step of the first layer, in the second layer, the RNN unit 31 performs a prescribed operation.

[0070] If the specified operations for each output value OUT1 of the first layer in the second layer are completed, the processing of the third layer is executed.

[0071] If the specified operations for the first pixel data in the second layer are completed, the processing corresponding to the first pixel in the third layer is executed.

[0072] In the third layer, multiple output values OUT1 obtained from the first to the last step in the second layer are sequentially input as input values IN1 to the RNN cell 31. Similar to the processing in the second layer, in the third layer, the RNN cell 31 executes the specified operations in the order from the first step to the last step of the second layer.

[0073] Figure 6 It is a diagram for explaining the processing order of the RNN cell 31 for multiple pixel values included in the input image data IG. Figure 6 It represents the flow of input values IN1, IN2 input to the RNN cell 31 and output values OUT1, OUT2 output from the RNN cell 31 in multiple steps. In the first layer, the RNN cell 31 is represented as the RNN cell (RNNCell) 1, in the second layer, the RNN cell is represented as the RNN cell 2, and in the third layer, the RNN cell is represented as the RNN cell 3.

[0074] In Figure 6 only the flow of processing for the pixel data of column x and its preceding columns (x - 1), (x - 2) in row y of the input image data IG is shown.

[0075] As Figure 6 shown, the input value IN1 of the RNN cell 1 in column (x - 2) of the first layer (layer 1) is the pixel data input in step t k The input value IN2 of the RNN cell 1 in column (x - 2) of the first layer is the output OUT2 of the RNN cell 1 in column (x - 3) of the first layer. The output value OUT1 of the RNN cell 1 in column (x - 2) of the first layer is the input value IN1 of the RNN cell 2 in column (x - 2) of the second layer. The output value OUT2 of the RNN cell 1 in column (x - 2) of the first layer is the input value IN2 of the RNN cell 1 in column (x - 1) of the first layer.

[0076] Similarly, the input value IN1 of the RNN cell 1 in column (x - 1) of the first layer is the pixel data input in step t (k+1)The pixel data input therein. The input value IN2 of the RNN unit 1 in column (x - 1) of the first layer is the output OUT2 of the RNN unit 1 in column (x - 2) of the first layer. The output value OUT1 of the RNN unit 1 in column (x - 1) of the first layer is the input value IN1 of the RNN unit 2 in column (x - 1) of the second layer. The output value OUT2 of the RNN unit 1 in column (x - 1) of the first layer is the input value IN2 of the RNN unit 1 in column (x) of the first layer.

[0077] The input value IN1 of the RNN unit 1 in column (x) of the first layer is the pixel data input at step t (k+2) The input value IN2 of the RNN unit 1 in column (x) of the first layer is the output OUT2 of the RNN unit 1 in column (x - 1) of the first layer. The output value OUT1 of the RNN unit 1 in column (x) of the first layer is the input value IN1 of the RNN unit 2 in column (x) of the second layer. The output value OUT2 of the RNN unit 1 in column (x - 1) of the first layer is used as the input value IN2 of the RNN unit l in the next time step.

[0078] As described above, the RNN unit 31 of the RNN processor 22 performs RNN operations on the input multiple pixel data in sequence, and saves the information of the hidden state into the state buffer 21. The hidden state is the output of the RNN unit 31.

[0079] The input value IN1 of the RNN unit 2 in column (x - 2) of the second layer (layer 2) is the output value OUT1 of the RNN unit 1 in column (x - 2) of the first layer. The input value IN2 of the RNN unit 2 in column (x - 2) of the second layer is the output OUT2 of the RNN unit 2 in column (x - 3) of the second layer. The output value OUT1 of the RNN unit 2 in column (x - 2) of the second layer is the input value IN1 of the RNN unit 3 in column (x - 2) of the third layer. The output value OUT2 of the RNN unit 2 in column (x - 2) of the second layer is the input value IN2 of the RNN unit 2 in column (x - 1) of the second layer.

[0080] Similarly, the input value IN1 of the RNN unit 2 in column (x - 1) of the second layer is the output value OUT1 of the RNN unit 1 in column (x - 1) of the first layer. The input value IN2 of the RNN unit 2 in column (x - 1) of the second layer is the output OUT2 of the RNN unit 2 in column (x - 3) of the second layer. The output value OUT1 of the RNN unit 2 in column (x - 1) of the second layer is the input value IN1 of the RNN unit 3 in column (x - 1) of the third layer. The output value OUT2 of the RNN unit 2 in column (x - 1) of the second layer is the input value IN2 of the RNN unit 2 in column (x) of the second layer.

[0081] The input value IN1 of the RNN cell 2 in column (x) of the second layer is the output value OUT1 of the RNN cell 1 in column (x) of the first layer. The input value IN2 of the RNN cell 2 in column (x) of the second layer is the output OUT2 of the RNN cell 2 in column (x - 1) of the second layer. The output value OUT1 of the RNN cell 2 in column (x) of the second layer is the input value IN1 of the RNN cell 3 in column (x) of the third layer. The output value OUT2 of the RNN cell 2 in column (x) of the second layer is used as the input value IN2 of the RNN cell 2 in the next time step.

[0082] The input value IN1 of the RNN cell 3 in column (x - 2) of the third layer (layer 3) is the output value OUT1 of the RNN cell 2 in column (x - 2) of the second layer. The input value IN2 of the RNN cell 3 in column (x - 2) of the third layer is the output OUT2 of the RNN cell 3 in column (x - 3) of the third layer. The output value OUT1 of the RNN cell 3 in column (x - 2) of the third layer is input to the softmax layer here, and the output image data OG is output from the softmax layer. The output value OUT2 of the RNN cell 3 in column (x - 2) of the third layer is the input value IN2 of the RNN cell 3 in column (x - 1) of the third layer.

[0083] Similarly, the input value IN1 of the RNN cell 3 in column (x - 1) of the third layer is the output value OUT1 of the RNN cell 2 in column (x - 1) of the second layer. The input value IN2 of the RNN cell 3 in column (x - 1) of the third layer is the output OUT2 of the RNN cell 3 in column (x - 2) of the third layer. The output value OUT1 of the RNN cell 3 in column (x - 1) of the third layer is input to the softmax layer here, and the output image data OG is output from the softmax layer. The output value OUT2 of the RNN cell 3 in column (x - 1) of the third layer is the input value IN2 of the RNN cell 3 in column (x) of the third layer.

[0084] The input value IN1 of the RNN cell 3 in column (x) of the third layer is the output value OUT1 of the RNN cell 2 in column (x) of the second layer. The input value IN2 of the RNN cell 3 in column (x) of the third layer is the output OUT2 of the RNN cell 3 in column (x - 1) of the third layer. The output value OUT1 of the RNN cell 3 in column (x) of the third layer is input to the softmax layer here, and the output image data OG is output from the softmax layer. The output value OUT2 of the RNN cell 3 in column (x) of the third layer is used as the input value IN2 of the RNN cell 3 in the next time step.

[0085] Thus, the output of the third layer is the data of multiple output values OUT1 obtained in multiple steps. The output of the third layer is input to the softmax layer. The output of the softmax layer is transformed into image data of y rows and x columns and saved as output image data OG in the off-chip memory 12.

[0086] As described above, the RNN cell processor 22 performs a recurrent neural network operation using at least one of the multiple pixel data of the image data and the hidden state that is the operation result of the RNN operation stored in the state buffer 21. The RNN processor 22 can execute multiple layers that are processing units for performing the RNN operation multiple times. The multiple layers include a first processing unit (first layer) that performs the RNN operation with multiple pixel data as input, and a second processing unit (second layer) that performs the RNN operation with the data of the hidden state obtained in the first processing unit (first layer) as input.

[0087] In addition, as described above, the values of the weight parameters of the non-linear function in the RNN operation are optimized through the learning of the RNN.

[0088] As described above, according to the above-described embodiment, the RNN is used instead of the CNN to perform a prescribed process on the image data.

[0089] Thus, different from the method of performing a kernel operation while sliding a window of a prescribed size relative to the entire image data after holding the image data in the off-chip memory 12, since the image data is transformed into the stream data SD and the RNN operation is sequentially performed in the image processing apparatus of the present embodiment, the neural network operation process can be performed with a small waiting time and at low cost.

[0090] (Modification Example 1)

[0091] In the above-described embodiment, the image data composed of multiple pixels in multiple rows and multiple columns is transformed into the stream data SD, and the pixel values from the pixel value of the first row and first column to the pixel value of the final row and final column are sequentially input as the input value IN1 of one RNN cell processor 31.

[0092] However, in the case of image data, the trend of the feature amount is different between the pixel value of the pixel in the first column of each row and the pixel value of the final column of the previous row.

[0093] Therefore, in this Modification Example 1, instead of using the output value OUT2 of the final column of each row as the initial input value IN2 of the next row as it is, a line end unit that is set as the initial input value IN2 of the RNN cell 31 of the next row after being changed to a prescribed value is added.

[0094] As the line end unit, the RNN unit 31 can be used by changing the execution content of the RNN unit 31 to perform operations of a non-linear function different from the above non-linear function, or as shown by the dashed line in Figure 3 , a line end unit 31a provided in the RNN unit processor 22 and used as an operation unit different from the RNN unit 31 can be used.

[0095] The values of the respective weight parameters of the non-linear function of the line end unit are also optimized through the learning of the RNN.

[0096] Figure 7 FIG. is a diagram for explaining the processing order of the line end unit 31a for the output value OUT2 of the final column of each row. Here, each row of the image data has W pixel values. That is, the image data has W columns.

[0097] As shown in Figure 7 , after the RNN unit 31 performs a prescribed operation on the pixel data of the final column (W−1) when the first column is set to 0, the output value OUT2 is input to the line end unit 31a.

[0098] As shown in Figure 7 , the line end unit 31a processes the output value OUT2 of the RNN unit 31 of the final column (W−1) of each row for each layer. In Figure 7 , the line end unit 31a in the first layer is represented as the line end unit (LineEndCell) 1, the line end unit 31a in the second layer is represented as the line end unit 2, and the line end unit 31a in the third layer is represented as the line end unit 3.

[0099] In the first layer, the line end unit 31a of the y-th row takes the output value OUT2 (h 1(W-1,y) ) of the RNN unit l of the final column of the y-th row in the first layer as an input, and takes the hidden state h 1(line) of the output value as the operation result as the input value IN2 of the RNN unit 1 of the subsequent (y + 1)-th row.

[0100] Similarly, in the second layer, the line end unit 31a of the y-th row also takes the output value OUT2 (h 2(W-1,y) ) of the RNN unit 2 of the final column of the y-th row in the second layer as an input, and takes the hidden state h 2(line) of the output value as the operation result as the input value IN2 of the RNN unit 2 of the subsequent (y + 1)-th row.

[0101] Similarly, in the third layer, the line end unit 31a of the y-th row also takes the output value OUT2 (h 3(W-1,y) ) of the RNN unit 3 of the final column of the y-th row in the third layer as an input, and takes the hidden state h 3(line)As the input value IN2 of the RNN unit 3 of the subsequent (y + 1)-th row.

[0102] As described above, when the image data is composed of pixel data of n rows and m columns, the RNN unit processor 22 has the line-end unit 31a that performs a prescribed operation on the hidden state between two adjacent rows.

[0103] Thus, the line-end unit 31a is provided at the change point of the row in each layer. And the line-end unit 31a performs a process of changing the input output value OUT2, and uses the changed output value as the input value IN2 of the RNN unit 31 when processing the next row.

[0104] As described above, by changing the output value OUT2 of the final column of each row by the line-end unit 31a, it is possible to eliminate the influence of the difference in the tendency of the feature amount between the final pixel value of each row and the initial pixel value of the next row, and furthermore, an improvement in accuracy such as noise removal can be expected.

[0105] (Modification Example 2)

[0106] In the above-described embodiment, the input value IN1 of the RNN unit 31 is obtained at a consistent step size between all layers. In contrast, in this Modification Example 2, the input value IN1 of the RNN unit 31 is not obtained at a layer-consistent step size, but is obtained with an offset delay so that the RNN operation has the same receptive field as the receptive field of the CNN. In other words, the image processing apparatus of this Modification Example 2 is configured to perform the RNN operation with an offset between layers.

[0107] Figure 8 It is a diagram for explaining the processing sequence of the RNN unit 31 for a plurality of pixel values included in the input image data IG in relation to this Modification Example 2.

[0108] As Figure 8 shown, the pixel data i of the stream data SD is sequentially processed in the first layer. However, in the second layer, as the input value IN1 of the RNN unit 2, the output value OUT1 of the RNN unit 1 is used with a delay offset u1 in the x direction of the image and a delay offset v1 in the y direction of the image. In addition, the offset information is written into the off-chip memory 12 and written from the off-chip memory 12 to the RNN unit processor 22 as a parameter.

[0109] In Figure 8 , the input value IN1 of the RNN unit 2 is represented by the following formula (6).

[0110] i 2(x-u1,y-v1) =h 1(x-u1,y-v1) …(6)

[0111] Furthermore, in the third layer, the input value IN1 of the RNN unit 3 uses the output value OUT1 of the RNN unit 1 with a delay offset of (u1 + u2) in the x direction of the image and a delay offset of (v1 + v2) in the y direction of the image. That is, in Figure 8 the input value IN1 of the RNN unit 3 is represented by the following equation (7).

[0112]

[0113] The output value OUT1 of each RNN unit 3 in the third layer is represented by the following equation (8).

[0114]

[0115] Figure 9 is a diagram for explaining the receptive field of a CNN. The receptive field is the range of input values that affect the kernel operation. By performing a CNN operation on the input image data IG through the layer LY1, the output image data OG is generated. In this case, the range R2 wider than the kernel size R1 of the layer LY1 affects the output value P1 of the output image data. Thus, in the case of a CNN, if the CNN operation is repeated, the receptive field, which is the range of input values directly or indirectly referred to in order to obtain the output value, becomes larger.

[0116] In contrast, in the above-described embodiment, since the RNN operation is performed, the range of the results of the RNN operation performed earlier in the operation step for each layer can be referred to as the receptive field.

[0117] Figure 10 is a diagram for explaining the receptive field of the above-described embodiment. Figure 11 is a diagram for explaining the difference in the range of the receptive fields of a CNN and an RNN. If the RNN unit 31 performs an RNN operation on the stream data SD of the input image data IG in the layer LY11, then in Figure 10 the range R12 indicated by the dotted line in the input image data IG is the receptive field. The receptive field of the output value P1 of the layer LY11 is the range R11 of the operation results of the steps earlier than the operation step of the output value P1.

[0118] Therefore, in the above-described embodiment, in Figure 9 the operation results of the pixel values around the output value P1 as in the CNN are not used in the RNN operation. As Figure 11 shown, the receptive field RNNR of the RNN is different from the receptive field CNNR of the CNN.

[0119] Therefore, in the above-described embodiments, similar to CNN, in order to perform RNN operations considering the receptive field, the RNN unit 31 offsets the range of the input value IN1 read from the state buffer 32 so that the input value IN1 of the RNN unit 31 used in a certain step in a certain layer becomes the hidden state h (output value) of the RNN unit 31 in a step different from this step in the previous layer. That is, the data of the hidden state obtained in the first layer as the first processing unit is given from the state buffer 21 to the RNN processor 22 in the second layer as the second processing unit at a step with a set offset delay.

[0120] As Figure 8 shown, in the second layer, the input value IN1 of the RNN unit 2 becomes the output value OUT1 at the pixel position offset by u1 in the x direction and v1 in the y direction. That is, in the second layer, the output value OUT1 of the RNN operation at the pixel position where the RNN unit 2 is shifted by a specified value (u1, v1) in the horizontal and vertical directions of the image data becomes the input value IN1 of the RNN unit 2 in the second layer.

[0121] In addition, in the third layer, the input value IN1 of the RNN unit 3 becomes the output value OUT1 offset by (u1 + u2) in the x direction and (v1 + v2) in the y direction in the output image of the second layer.

[0122] Moreover, the output value OUT1 of the RNN unit 3 becomes the output value offset by (u1 + u2 + u3) in the x direction and (v1 + v2 + v3) in the y direction in the output image of the second layer.

[0123] Figure 12 is a diagram for explaining the input step of the RNN unit 31. As Figure 12 shown, the output value OUT1 of the RNN unit 1 with the initial pixel data i1(0, 0) as the input value IN1 is used as the input value IN1 in the second layer at the step t a corresponding to the offset value. The offset value in the second layer is the step difference relative to the acquisition step of the pixel data of the streaming data SD in the first layer. Here, the offset value is a value corresponding to the step difference from the position (0, 0) of the pixel in the first row and first column to the pixel position (u1, v1) of the u1-th row and v1-th column.

[0124] Thus, in the first step t a of the second layer, the input value IN1 of the RNN unit 2 becomes the output value OUT1 at the step offset by the offset value from the first step t b in the first layer.

[0125] Furthermore, the offset value can also be the same between layers, but here it is different for each layer. As Figure 12 shown, the output value OUT1 of the RNN unit 31 with the step size t a in the third layer is offset by the value of the pixel position (u11, v11) to become the input value IN1 of the RNN unit 31 in the third layer.

[0126] Figure 13 is a diagram for explaining the setting range of the receptive field of this Modification 2. When setting the offset value of the input value IN of the layer LY21, a predetermined area AA is added to the input image data IG by padding. And, as Figure 13 shown, the output value P1 is output under the influence of the input value P2 within the receptive field RNNR. Thus, the output value P1 is affected by the output value of the receptive field RNNR of the layer LY21, and the receptive field RNNR of the layer LY21 is affected by the input value of the receptive field RNNR of the input image data IG. The output value PE is affected by the input value P3 of the added area AA.

[0127] As described above, by setting the offset amount of the input step of the input value IN1 in each RNN operation for each layer, in image processing using an RNN, it is also possible to set the receptive field in the same way as a CNN.

[0128] As described above, according to the above-described embodiments and each modification, it is possible to provide an image processing apparatus with a small waiting time and capable of being implemented at low cost.

[0129] In addition, the above-described RNN unit 31 is a simple RNN, but it may also have a structure such as an LSTM (Long Short Term Memory) network or a GRU (Gated Recurrent Unit).

[0130] The above has described several embodiments of the present invention, but these embodiments are exemplified as examples and are not intended to limit the scope of the invention. These new embodiments can be implemented in various other forms, and various omissions, substitutions, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are included in the invention described in the claims and its equivalent scope.

Claims

1. An image processing apparatus, characterized in that, comprising: a first processor to which image data is input; a cache provided in the first processor; and a second processor provided in the first processor, which performs the recurrent neural network operation using at least one of a plurality of pixel data of the image data and an operation result of the recurrent neural network operation stored in the cache, the plurality of pixel data are sequentially input to the second processor; the second processor sequentially performs the recurrent neural network operation on the input plurality of pixel data and stores the operation result in the cache, the second processor is capable of executing a plurality of layers, which are processing units that perform the recurrent neural network operation multiple times, the plurality of layers include a first processing unit that inputs the plurality of pixel data and performs the recurrent neural network operation, and a second processing unit that inputs the operation result obtained in the first processing unit and performs the recurrent neural network operation, the operation result obtained in the first processing unit is delayed by an offset amount u1 in the x direction of the image data and an offset amount v1 in the y direction of the image data, and is input to the second processing unit. Offset information is written into the cache and provided to the second processor from the cache.

2. The image processing apparatus according to claim 1, characterized in that, the operation result of the recurrent neural network operation is a hidden state.

3. The image processing apparatus according to claim 1, characterized in that, the image data is composed of pixel data of n rows and m columns; the second processor performs a prescribed operation on the operation results between two adjacent rows.

Citation Information

Patent Citations

  • Contact media using touch screen

    JP2020046914A

  • Artificial intelligence system and methods for performing image analysis

    WO2019177639A1