Systems and methods for image convolution
The Teplitz-like kernel matrix for image convolution addresses the efficiency and power consumption issues in image processing systems by optimizing the convolution process, enabling efficient and robust image processing in resource-constrained environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ルカス サントス フェレイラ
- Filing Date
- 2021-11-05
- Publication Date
- 2026-05-01
AI Technical Summary
Existing image processing systems in autonomous vehicles and virtual reality applications face challenges in performing computationally intensive tasks efficiently due to limited battery capacity and physical size, particularly in extracting prominent points from images, which consume significant processing time and power.
The use of a Teplitz-like kernel matrix for image convolution, where the kernel is padded with zeros and processed using a processing element array, allowing for efficient multiplication of image data without approximation, reducing power consumption and processing time.
This approach enables fast and robust image processing with reduced power consumption, effectively handling large image datasets with minimal computational overhead.
Smart Images

Figure 0007854436000004 
Figure 0007854436000005 
Figure 0007854436000006
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system and method for convolution of data. More specifically, the present invention relates to a system and method for convolution of images by matrix multiplication with Teplitz-like kernel matrices. [Background technology]
[0002] With the current increase in autonomous vehicles, drones, and virtual reality applications, the need for robust, fast, and efficient image processing has become critical. These image processing systems typically have to perform several challenging and computationally intensive tasks in real time, within a limited energy balance and physical size. This is particularly evident in the case of autonomous drones and virtual reality applications, where both limited battery capacity and physical size are problematic. For example, a 5kg battery and a full-sized graphics card cannot be fitted into a small indoor drone, considering the expected flight conditions.
[0003] What all these applications have in common is that they use images to estimate the 3D structure of the world, as well as the position and orientation of things within it. To achieve this, special points in the image, called prominent points or feature points, are extracted. These points have special properties, such as high texture and uniqueness, which allow them to be consistently extracted in other images as well. By aligning these features across several images, the 3D structure of the world, as well as the orientation and position of the camera when the image was taken, can be estimated. This allows onboard systems to track where drones, vehicles, and / or people are in the environment.
[0004] To find these prominent points in an image, various patterns around each point must be extracted to evaluate its characteristics. The number of pixels in an image is usually very large, varying from hundreds of thousands to billions. The higher the resolution, the more accurately the solution can be estimated. This means that feature extraction is a very computationally and power-intensive operation. Typically, this step contributes a considerable amount of processing time and power consumption to the system.
[0005] Therefore, improved image processing is needed. [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] In consideration of the foregoing, an object of the present invention is to overcome at least partially one or more of the limitations of the prior art identified above. In particular, an object is to have improved image processing systems and methods. [Means for solving the problem]
[0007] According to the first aspect, the image convolution accelerator system comprises a processing element having kernel elements corresponding to a generated Teplitz-like kernel, wherein the generated Teplitz-like kernel is padded with zeros based on a desired convolution kernel, and an image controller configured to activate the processing elements to multiply the kernel elements by image data when image data from the same row of the image is assigned to all processing elements corresponding to non-zero rows of the Teplitz-like kernel.
[0008] According to a second embodiment, the imaging apparatus comprises an imaging unit configured to collect an input image, and an image convolution accelerator system comprising a processing element having kernel elements corresponding to a generated Teplitz-like kernel, wherein the generated Teplitz-like kernel is padded with zeros based on a desired convolution kernel, and an image controller configured to activate the processing elements to multiply the kernel elements by image data when image data from the same row of the image is assigned to all processing elements corresponding to non-zero rows of the Teplitz-like kernel.
[0009] According to a third aspect, the image processing method comprises the steps of generating a Teplitz-like kernel, wherein the generated Teplitz-like kernel is padded with zeros based on a desired convolution kernel; and activating processing elements having kernel elements corresponding to the generated Teplitz-like kernel such that when image data from the same row of the image is assigned to all processing elements corresponding to the non-zero rows of the Teplitz-like kernel, the kernel elements are multiplied by the image data.
[0010] According to a fourth aspect, the program causes a computer to perform the following steps: generating a Teplitz-like kernel, wherein the generated Teplitz-like kernel is padded with zeros based on a desired convolution kernel; and activating processing elements having kernel elements corresponding to the generated Teplitz-like kernel such that when image data from the same row of the image is assigned to all processing elements (PEs) corresponding to non-zero rows of the Teplitz-like kernel, the kernel elements are multiplied by the image data.
[0011] In this application, the Teplitz-like kernel should be understood as a doubly blocked cyclic matrix kernel.
[0012] Further examples of the present disclosure are defined in the dependent claims, and features for the fourth and subsequent embodiments of the present disclosure are similar to features for the first to third embodiments with necessary modifications.
[0013] Some examples of this disclosure provide a way to aggregate partial results using processing and / or storage elements.
[0014] Some examples of this disclosure provide a way to store partial results.
[0015] Some examples of this disclosure provide timing of image data between processing elements.
[0016] In general, all terms used in the claims should be interpreted according to their ordinary meanings in the art unless otherwise expressly defined herein. All references to “a / an / the [element, device, component, means, step, etc.]” should be broadly interpreted as referring to at least one instance of such element, device, component, means, step, etc. unless otherwise expressly stated. The steps of any method disclosed herein do not need to be performed in the exact order disclosed unless otherwise expressly stated.
[0017] The above-mentioned and additional objects, features and advantages of the present invention will be better understood by the following exemplary and non-limiting detailed description of the invention with reference to the accompanying drawings, and the same reference numerals will be used for similar elements. [Brief explanation of the drawing]
[0018] [Figure 1] This is a schematic diagram of an image convolution accelerator system. [Figure 2] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 3]This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 4] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 5] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 6] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 7] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 8] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 9] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 10] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 11] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 12] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 13] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 14] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 15] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 16] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 17] This is a schematic diagram of an image convolution accelerator system that convolves an image using a Teplitz-like kernel obtained by unfolding a 3x3 kernel. [Figure 18] This is a top view of a drone equipped with an image convolution accelerator system and a camera. [Modes for carrying out the invention]
[0019] Next, the present invention will be described more thoroughly below with reference to the accompanying drawings illustrating current preferred examples of the invention. However, the present invention can be illustrated in many different forms and should not be construed as being limited to the examples described herein; rather, these examples are provided for thoroughness and completeness and will fully convey the scope of the invention to those skilled in the art.
[0020] Figure 1 shows an image convolution accelerator system 100. For simplicity, the accelerator system 100 is shown with processing elements PE arranged in rows and columns. The processing elements PE may be arranged in ways other than rows and columns.
[0021] The processing element PE corresponds to the generated Teplitz-like kernel. The generated Teplitz-like kernel is padded with zeros based on the desired kernel. The generated Teplitz-like kernel is loaded into the corresponding processing element PE, which is, for example, one of the neighboring k processing elements PE. 00As shown by the accelerator system, the generated Teplitz-like kernel k is multiplied by the generated Teplitz-like kernel k so that the image is summed with it. XY The system further comprises a controller 10 configured to activate processing element PEs based on non-zero kernel elements within the image, and when non-zero kernel elements are assigned to all processing element PEs in a single row. In this case, the image may be convolved with a Teplitz-like kernel. The convolution also has less loss because there is no approximation in the kernel.
[0022] The above Teplitz-like kernel k XY Implementations based on this approach benefit from the simple control logic required to perform convolution. Furthermore, because the processing element PE is only active when performing an operation, the system has very low power consumption compared to other convolutional systems.
[0023] Teplitz-sama Kernel K XY This can be generated based on the desired 3x3 kernel. [Table 1] Next, Teplitz-sama kernel k XY This is generated based on the desired kernel and can be padded with zeros, which results in the following: [Table 2] Next, convolution can be performed by multiplying the row-stack flattened input image by a Teplitz-like kernel. Row-stack flattening should be understood as the rows of the image being staggered within a single row. Teplitz-like kernel k XY The input image can also be transposed and multiplied by the row stack flattened input image. The input image may be streamed or read directly in vector form in some examples. For example, the input image may be read directly on the scan line as produced by the image sensor. Image convolution accelerators are advantageous because they eliminate the need for arbitrary input buffers.
[0024] An example of convolution by the accelerator system 100 is shown in FIGS. 2 to 17. In the illustrated accelerator system, the first row performs the calculation of the first diagonal K 00 , K 01 , K 02 , and the second row performs the calculation of the second diagonal K 10 , K 11 , K 12 , and the third row performs the calculation of the third diagonal K 20 , K 21 , K 22 . The diagonals take into account the valid partial sums.
[0025] In this example, the input image is a 4×4 image.
Table 3
[0026] Next, the convolution of the image by the Toeplitz-like kernel k XY in the accelerator system 100 is described in a step process for each cycle.
[0027] Clock cycle 0}7] Starting in FIG. 2, the input value i 00The '' represents image data, which is clocked in the first row of the processing element PE. Figure 2 shows timing logic configured to time input values and image data between rows and / or columns of the processing element PE. The timing logic could be, for example, a simple counter that blocks the input values.
[0028] During this clock cycle, all processing elements PE are inactive to conserve power, and these are indicated by shaded PE symbols. Activated processing elements are indicated by unshaded PE symbols.
[0029] Input value i 00 The second and third rows of the processing element PE are skipped.
[0030] Below the accelerator system 100 and a description of how convolution proceeds within the system, the corresponding Teplitz-like kernel k XY And Teplitz-like kernel k at different stages throughout the entire convolution XY Examples of its application include [examples of application].
[0031] Clock cycle 1 In Figure 3, the next input value i of the image 01 This is clocked in the first row. During this clock cycle, all processing elements PE are deactivated to conserve power.
[0032] Clock cycle 2 Figure 4, here the third input value i in the image 02 However, it is clocked to the first row. This means that all three processing elements PE of the first row are assigned and the processing elements PE are activated. Each processing element PE is assigned to the respective Teplitz-like first row kernel element k 00 , k 01 , k 02 The image data i assigned to 00 i 01 i 02The partial result is calculated by multiplying by . The partial result is combined by an arithmetic logic unit shown as an adder. The combined result is stored in the first memory element shown as a FIFO.
[0033] Teplitz-sama Kernel K XY It can also be seen that the first row of the first diagonal is applied for the convolution.
[0034] Clock cycle 3 Figure 5, the following input value i 03 This is clocked in the first row. In this case as well, all three processing elements PE of the first row are assigned and the processing elements PE are activated. Each processing element PE is assigned to the Teplitz-like first row kernel element k 00 , k 01 , k 02 The image data i assigned to 01 i 02 i 03 The partial results are calculated by multiplying by . The partial results are combined, and the combined result is stored in a memory element.
[0035] Furthermore, Teplitz-like kernel k XY The second row of the first diagonal is currently being applied for convolution.
[0036] Clock cycle 4 In Figure 6, the next input value i in the second row of the image is 10 However, it is clocked in the first and second rows. Next, the skip logic is executed because the conditions of the skip logic are met, so the input value i 10 The input is passed from the first row to the second row. The condition is that this number of input values is skipped based on the number of columns in the image.
[0037] In this case as well, all processing elements PE are deactivated.
[0038] Clock cycle 5 In Figure 7, the following input value i for the second row of the image 11However, it is clocked in the first and second rows. Still, all processing elements PE are deactivated.
[0039] Clock cycle 6 Figure 8, here the third input value i of the image 12 However, the first and second rows are clocked. This means that all three processing element PEs of the first and second rows are assigned, and the processing element PEs of the first and second rows are activated. Each processing element PE is assigned to the Teplitz-like first row kernel element k 00 , k 01 , k 02 The image data i assigned to 10 i 11 i 12 Multiply by the second row kernel element k 10 , k 11 , k 12 The image data i assigned to 10 i 11 i 12 The partial result is calculated by multiplying by . The partial result of the first row is combined and stored in the first memory element. The partial result of the second row is combined and added together with the first stored data in the first memory element, which represents the first combined result of the first row of the image. This addition is then stored in the second memory element, which is also shown as FIFO.
[0040] As shown below, Teplitz-like kernel k XY The third row of the first diagonal is now a Teplitz-like kernel k for convolution. XY This is applied similarly to the first row of the second diagonal.
[0041] Clock cycle 7 Figure 9, the following input value i 13 This is clocked in the first and second rows. In this case as well, all three processing elements PE in the first and second rows are assigned and activated.
[0042] The partial results of the first row are combined and stored in the first memory element. The partial results of the second row are combined and added together with the second stored data in the first memory element, which represents the second combined result of the first row of the image. This addition is then stored in the second memory element.
[0043] Currently, Teplitz-sama Kernel K XY The fourth row of the first diagonal is a Teplitz-like kernel k for convolution. XY This is applied similarly to the second row of the second diagonal.
[0044] Clock cycle 8 In Figure 10, the input value i for the third row of the image is shown. 20 However, it is clocked in the second and third rows. The first diagonal k 00 , k 01 , k 02 Since the calculation was completed in the previous cycle 7, it is now possible to load new kernel elements for the new second convolution into the first row of processing element PE using a pipeline approach. All processing element PEs are deactivated.
[0045] Clock cycle 9 In Figure 11, the input value i for the third row of the image is shown. 21 However, the second and third rows are clocked. All processing elements PE remain deactivated.
[0046] Clock cycle 10 Figure 12, the third input value i in the third row of the image. 22 However, the second and third rows are clocked. In this case as well, all three processing element PEs of the second and third rows are assigned, and the processing element PEs of the second and third rows are activated. Each processing element PE is assigned to the Teplitz-like second row kernel element k 10 , k 11 , k 12 The image data i assigned to 20 i 21 i 22Multiply by the third row kernel elements k 20 k 21 k 22 and the assigned image data i 20 i 21 i 22 to calculate partial results by multiplication.
[0047] The partial result of the second row is combined and added with the third stored data in the first storage element representing the first combined result of the second row of the image. This addition is then stored in the second storage element. The partial result of the third row is combined and added with the first stored data in the second storage element representing the first combined results of the first and second rows of the image.
[0048] As shown below, the third element of the second diagonal of the Toeplitz-like kernel k XY is applied in the same way as the first element of the third diagonal of the Toeplitz-like kernel k XY for convolution.
[0049] Clock cycle 11 FIG. 13, the fourth input value i 23 of the third row of the image is clocked into the second and third rows. Again, all three processing elements PE of the second and third rows are assigned and the processing elements PE of the second and third rows are activated. The processing elements PE multiply the assigned image data i 10 i 11 i 12 by the respective Toeplitz-like second row kernel elements k 21 i 22 i 23 and multiply the assigned image data i 20 i 21 i 22 by the third row kernel elements k 21 i 22 i 23 to calculate partial results by multiplication.
[0050] The partial result of the second row is combined and added with the fourth stored data in the first storage element representing the second combined result of the second row of the image. This addition is then stored in the second storage element. The partial result of the third row is combined and added with the second stored data in the second storage element representing the second combined results of the first and second rows of the image.
[0051] As shown below, the fourth row of the second diagonal of the Toeplitz-like kernel k XY is applied in the same way as the second row of the third diagonal of the Toeplitz-like kernel k XY for convolution.
[0052] Clock cycle 12 In FIG. 14, the input value i 30 for the fourth row of the image is clocked in the third row. Since the calculations of the second diagonals k 10 , k 11 , k 12 were completed in the previous cycle 11, it is now possible to load the new kernel elements of the second convolution, which started in cycle 8, into the second row of the processing element PE in a pipelined manner. All processing elements PE are deactivated.
[0053] Clock cycle 13 In FIG. 15, the input value i 31 for the fourth row of the image is clocked in the third row. All processing elements PE are still deactivated.
[0054] Clock cycle 14 In FIG. 16, the input value i 32 for the fourth row of the image is clocked in the third row. This means that all three processing elements PE of the third row are assigned and the processing elements PE of the third row are activated.
[0055] The processing element PE has the respective Toeplitz-like third row kernel elements k 20 , k 21, k 22 The image data i assigned to 30 i 31 i 32 The partial result is calculated by multiplying by . The partial result of the third row is combined and added together with the first stored data in the second memory element, which represents the first combined result of the second and third rows of the image.
[0056] Teplitz-sama Kernel K XY The third row of the third diagonal is currently being used for the fold.
[0057] Clock cycle 15 In Figure 17, the last input value i for the fourth row of the image is 33 However, it is clocked in the third row. This means that all three processing elements PE in the third row are assigned and the processing elements PE in the third row are activated.
[0058] The processing element PE is the third row kernel element k of the Teplitz-like form. 20 , k 21 , k 22 The image data i assigned to 31 i 32 i 33 The partial result is calculated by multiplying by . The partial result of the third row is combined and added together with the second stored data in the second memory element, which represents the second combined result of the second and third rows of the image.
[0059] Teplitz-sama Kernel K XY The fourth row of the third diagonal is currently being used for the fold.
[0060] Therefore, 16×4 Teplitz-like kernel k XY The complete convolution of a 4x4 image is achieved in 16 clock cycles.
[0061] In this example, two clock cycles are used as interruptions to fill the processing element PE, and then two clock cycles are used for calculating the result. The two clock cycles are due to two additional zeros in the Teplitz-like kernel.
[0062] In some examples, the generated Teplitz-like kernel (k XY The arrangement of processing elements (PEs) corresponding to ) is scaled for other kernels, such as a 5x5 kernel with 5 rows, each having 5 processing elements (PEs), or a 7x7 kernel with 7 rows, each having 7 processing elements.
[0063] It is also possible to combine 7x7 kernels, each having 7 rows with 7 PEs, and add one additional PE to produce 7x7+1PE. This allows for one convolution by the 7x7 kernel, two parallel convolutions by the 5x5 kernel, or five parallel convolutions by the 3x3 kernel (with 5 PEs idle and potentially deactivated). Thus, there are many different ways to arrange the processing element PEs to correspond to the generated Teplitz-like kernel. It is also possible to use logic that can be configured or reconfigured to adapt to the different kernel solutions described above, such as software solutions with processing elements or objects.
[0064] When a larger processing element arrangement is used, and it is also configured to implement smaller kernel convolutions in parallel, the number of memory elements must be only as many as are required for the smaller parallel kernels. For example, if you want to use a 7x7 kernel and also be able to perform five 3x3 kernel convolutions in parallel, you will need six memory elements. Or, if five 3x3 kernels are desired, you will need 5*2 memory elements = 10 memory elements. In general, the number of memory elements is calculated as the maximum dimension of the supported kernel minus 1.
[0065] It is also possible to divide the memory elements to accommodate more memory elements of smaller sizes. In this way, instead of having several memory elements of different sizes, or FIFOs, for each convolution of different sizes, greater utilization of the total memory can be achieved. The memory elements or FIFOs can then be configured to store at least the number of columns in the image - the dimension of the kernel + 1, for the computation of the processing elements.
[0066] In some examples, the imaging device 400 comprises an imaging unit 410 configured to collect input images, as shown in Figure 18, and an image convolution accelerator system 100, where a drone is shown. Such an imaging device could be an autonomous vehicle or an augmented reality device. In general, the image convolution accelerator system 100 can be used in any kind of convolution-based machine learning system, which can be used for, for example, cancer detection, object detection, segmentation, depth prediction, classification, image reconstruction, compression, big data processing, data anonymization, recognition, etc.
[0067] The functions and operations of the image convolution accelerator system 100 may be embodied in the form of executable logical routines (e.g., lines of code, software programs, etc.) stored in a non-temporary computer-readable medium such as memory. These logical routines may be executed by a control circuit such as a processor. Furthermore, the functions and operations of the image convolution accelerator system 100 may be a standalone software application or form part of a software application that performs additional tasks related to the image convolution accelerator system 100. The described functions and operations may be considered as ways in which the corresponding device is configured to perform them. While the described functions and operations may be implemented in software, such functions may also be performed through dedicated hardware or firmware, or any combination of hardware, firmware, and / or software.
[0068] The software is a Teplitz-style kernel k XY A step to generate a Teplitz-like generated kernel k XY However, based on the desired kernel, the steps are to generate a zero-padding, and the generated Teplitz-like kernel k XY The steps include assigning processing elements PE arranged in corresponding rows and columns to the input image, and generating a Teplitz-like kernel k XY The generated Teplitz-like kernel k is multiplied by the result. XY This could be a software program that causes a computer to perform the steps of: activating processing element PEs based on non-zero kernel elements within a row, and activating the processing element PEs when all processing element PEs in a row have been assigned non-zero kernel elements. The software program may be stored in a non-temporary storage medium.
[0069] For even more improved image processing performance, the convolutional accelerator system 100 may be combined with a parallel memory system that supports the simultaneous writing and reading of many different data patterns. The parallel memory system comprises multiple memory banks, a method for tagging images and distributing them to different memory banks, and auxiliary functions / circuits for enabling flexible data access with high implementation efficiency. The parallel memory system is described in more detail in an application titled "SYSTEM AND METHOD FOR HIGH-THROUGHPUT IMAGE PROCESSING" filed on the same day by the same inventors as this application.
[0070] It will be understood that the present invention is not limited to the embodiments shown. Several modifications and variations are therefore conceivable within the scope of the invention, as exclusively defined by the appended claims.
Claims
1. A processing element (PE) having kernel elements corresponding to a generated Teplitz-like kernel which is a double-block cyclic matrix, wherein the generated Teplitz-like kernel is padded with zeros based on a desired convolution kernel, An image controller (10) is configured to activate the processing elements (PEs) to multiply the kernel elements by the image data when image data from the same row of the image is assigned to all processing elements (PEs) corresponding to non-zero rows of the Teplitz-like kernel, At least one storage element configured to store a value obtained by subtracting the dimension of the desired convolution kernel + 1 from the number of columns of the image, An image convolution accelerator system equipped with [the following features].
2. The arithmetic logic unit further comprises the processing element (PE) and / or memory element configured to add the multiplication, The image convolution accelerator system according to claim 1.
3. The number of memory elements is calculated as the dimension of the desired convolution kernel minus 1. The image convolution accelerator system according to claim 1 or claim 2.
4. The system further includes timing logic configured to time image data between the aforementioned processing elements. The image convolution accelerator system according to claim 1.
5. An imaging unit configured to collect input images, An image convolution accelerator system according to any one of claims 1 to 4, An imaging device equipped with the following features.
6. A step of generating a Teplitz-like kernel which is a double-block cyclic matrix, wherein the generated Teplitz-like kernel is padded with zeros based on a desired convolution kernel. When image data from the same row of the image is assigned to all processing elements (PEs) corresponding to non-zero rows of the Teplitz-like kernel, the steps include activating the processing elements comprising the kernel elements corresponding to the generated Teplitz-like kernel so that the kernel elements are multiplied by the image data, The steps include storing in at least one memory element a value obtained by subtracting the dimension of the desired convolution kernel + 1 from the number of rows in the image, An image processing method comprising:
7. A step of generating a Teplitz-like kernel which is a double-block cyclic matrix, wherein the generated Teplitz-like kernel is padded with zeros based on a desired convolution kernel. When image data from the same row of the image is assigned to all processing elements (PEs) corresponding to non-zero rows of the Teplitz-like kernel, the steps include activating the processing elements comprising the kernel elements corresponding to the generated Teplitz-like kernel so that the kernel elements are multiplied by the image data, The steps include storing in at least one memory element a value obtained by subtracting the dimension of the desired convolution kernel + 1 from the number of rows in the image, A program that causes a computer to execute something.
8. A non-temporary storage medium storing the program described in claim 7.
Citation Information
Patent Citations
Image processing apparatus, image processing method, and program
JP2009087252A
Image processing apparatus, image processing method, and program
US20090087118A1