Convolution operation system and method

By dividing the pulse data into sub-matrices and calculating the convolution sum by column, the problem of high resource usage of convolution operation is solved, and more efficient resource utilization and computing efficiency are achieved.

CN120561434BActive Publication Date: 2025-09-26INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511058512.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-09-26
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

Convolution operations occupy a lot of hardware resources, especially memory and computing resources, resulting in low computational efficiency.

Method used

By dividing the pulse data into sub-matrices and calculating the convolution sum by column, the storage and computing resources are reduced. The storage controller and operation circuit are used to process the number of columns of the sub-matrix respectively, and the convolution operation process is optimized.

Benefits of technology

It effectively reduces the usage of memory and computing resources, improves the efficiency of convolution operations, and reduces the demand for hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561434B_ABST
    Figure CN120561434B_ABST
Patent Text Reader

Abstract

The present application discloses a convolution operation system and method, which relates to the field of computer technology, including: first, the storage space of the memory only needs to store a preset number of rows of pulse data, and the preset number of rows is equal to the side length of the first preset convolution kernel, and the side length of the first preset convolution kernel is smaller than the width of a frame of pulse data, so that memory resources can be saved. Secondly, by generating a sub-matrix and determining the number of columns of the sub-matrix, the original large number of convolution kernels are replaced. Moreover, the sub-matrix includes valid data, and the operation process is performed by column, which can reduce the occupation of computing resources. In summary, by reducing the occupation of computing resources and storage resources, the hardware resources occupied by convolution operations can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a convolution operation system and method. Background Art

[0002] In the field of computer technology, a pulse camera can simply send pulse data to a computer device, which then processes the pulse data to determine whether it can be used by a neural network model to predict or identify relevant information in an image. The pulse data typically includes the coordinates of pixels whose brightness changes. The computer device can include a central processing unit (CPU) and an accelerator.

[0003] After receiving a frame of pulse data from the pulse camera, the central processing unit can first decode the pulse data. Specifically, based on the coordinates in the pulse data, the values ​​at the locations corresponding to these coordinates in the memory are set to a first preset value (for example, it can be 1), and the data at other locations in the preset matrix are set to a second preset value (for example, it can be 0) to obtain the target matrix. The central processing unit can then send the target matrix to the accelerator. Furthermore, the accelerator can gradually move a window of a preset convolution kernel size within the target matrix to obtain multiple convolution kernels. For each convolution kernel, the convolution kernel can be convolved with the preset weight matrix to obtain the convolution sum corresponding to the convolution kernel. This process requires memory to store a complete frame of pulse data, and the number of convolution kernels is relatively large. Each convolution kernel requires a convolution sum calculation, which consumes a lot of hardware resources. Summary of the Invention

[0004] The present application provides a convolution operation system, method, electronic device, storage medium, and program product to solve the problem that convolution operation occupies more hardware resources.

[0005] The present application provides a convolution operation system, which includes a storage controller, a memory, and an operation circuit;

[0006] A storage controller is configured to read a preset number of rows of pulse data from a memory; determine at least one submatrix based on a first preset convolution kernel side length equal to the preset number of rows, a second preset convolution kernel side length, the pulse data, and a preset weight matrix; determine the number of columns of each submatrix in the at least one submatrix; and send each submatrix and the number of columns of each submatrix to an operation circuit;

[0007] The operation circuit is used to determine the convolution sum of the target submatrix according to the target submatrix and the number of columns of the target submatrix, wherein the target submatrix is ​​any one of the at least one submatrix.

[0008] The present application provides a convolution operation method, which is applied to a convolution operation system. The convolution operation system includes a storage controller, a memory, and an operation circuit. The method includes:

[0009] In the current round, the storage controller reads a preset number of rows of pulse data from the memory; determines at least one submatrix based on a first preset convolution kernel side length equal to the preset number of rows, a second preset convolution kernel side length, the pulse data, and a preset weight matrix; determines the number of columns of each submatrix in the at least one submatrix; and sends each submatrix and the number of columns of each submatrix to the operation circuit;

[0010] The operation circuit determines the convolution sum of the target submatrix according to the target submatrix and the number of columns of the target submatrix, wherein the target submatrix is ​​any one of the at least one submatrix.

[0011] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any one of the above-mentioned convolution operation methods when executing the computer program.

[0012] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned convolution operation methods are implemented.

[0013] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned convolution operation methods when executed by a processor.

[0014] Through the present application, the storage space of the memory only needs to store a preset number of rows of pulse data. The preset number of rows is equal to the side length of the first preset convolution kernel, and the side length of the first preset convolution kernel is smaller than the width of a frame of pulse data. Therefore, memory resources can be saved. Furthermore, after the storage controller reads the preset number of rows of pulse data, it can determine at least one submatrix based on the first preset convolution kernel side length, the second preset convolution kernel side length, the pulse data, and the preset weight matrix, and determine the number of columns of each submatrix, and transmit each submatrix and the corresponding number of columns to the operation circuit. In the related art, the size of each convolution kernel is equal, and the number of convolution kernels is large. The convolution calculation process of each convolution kernel requires a large amount of computing resources. However, some convolution kernels are all zero data, and participating in the operation will waste a lot of resources. In this solution, the submatrix including valid data is determined. Since the number of columns of different submatrices is different, the number of columns of each submatrix can be determined. In order to save computing resources, the operation circuit can calculate by column, and accordingly, the storage controller can send each sub-matrix and the number of columns of each sub-matrix to the operation circuit. In this way, the operation circuit can accurately identify how many columns constitute a sub-matrix, and summarize the results of the column calculation to determine the convolution sum of the target sub-matrix. Since the size of the sub-matrix is ​​less than or equal to the size of the convolution kernel, and the operation process is summed by column, rather than summing the values ​​of each coordinate in the convolution kernel separately as in the related art, that is, the number of objects to be summed is reduced, which can reduce the computing resource usage. In summary, by reducing the usage of computing resources and storage resources, the hardware resources occupied by the convolution operation can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0016] Figure 1 A schematic diagram of a pulse data processing process provided in an embodiment of the present application;

[0017] Figure 2 A schematic diagram of another pulse data processing process provided in an embodiment of the present application;

[0018] Figure 3 A schematic diagram of a convolution operation provided in an embodiment of the present application;

[0019] Figure 4 A schematic diagram of an adder performing a convolution operation according to an embodiment of the present application;

[0020] Figure 5A schematic diagram of the architecture of a convolution operation system provided in an embodiment of the present application;

[0021] Figure 6 A schematic diagram of the architecture of another convolution operation system provided in an embodiment of the present application;

[0022] Figure 7 A schematic diagram of the architecture of another convolution operation system provided in an embodiment of the present application;

[0023] Figure 8 A schematic diagram of another pulse data processing process provided in an embodiment of the present application;

[0024] Figure 9 A schematic diagram of pulse data stored in a memory provided in an embodiment of the present application;

[0025] Figure 10 A schematic diagram of pulse data stored in another memory provided in an embodiment of the present application;

[0026] Figure 11 A schematic diagram of determining a candidate convolution kernel according to an embodiment of the present application;

[0027] Figure 12 A schematic diagram of a preset weight matrix provided in an embodiment of the present application;

[0028] Figure 13 A schematic diagram of a generator matrix provided in an embodiment of the present application;

[0029] Figure 14 A schematic diagram of a convolution operation performed by a computing circuit provided in an embodiment of the present application;

[0030] Figure 15 A schematic diagram of updating a convolution sum by a pulse data updater provided in an embodiment of the present application;

[0031] Figure 16 A schematic diagram of a convolution operation method provided in an embodiment of the present application;

[0032] Figure 17 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0034] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0035] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0036] After generating a frame of pulse data, the pulse camera can send the pulse data to the central processing unit in the computer device. The central processing unit can then parse the pulse data to obtain coordinate information. Based on the coordinate information, the central processing unit determines the storage locations corresponding to the coordinate information in the memory and sets the values ​​of these storage locations to a first preset value (for example, 1). In addition, the central processing unit can set the values ​​of other storage locations in the memory to a second preset value (for example, 0), thereby completing the decoding operation of the pulse data.

[0037] like Figure 1 As shown, the pulse camera can obtain the coordinate information of the pixel whose brightness changes, generate pulse data and send it to the central processor, for example, [2, 2], [2, 8], [3, 6], [4, 1]. After receiving this pulse data, the central processor can decode the pulse data and obtain Figure 1 The first matrix in .

[0038] Furthermore, the specific process of calculating the convolution sum of the first matrix may be:

[0039] First, if Figure 2 As shown, the central processing unit can construct a sliding window according to the preset convolution kernel size (for example, it can be 3×3), and move step by step in the first matrix to obtain multiple convolution kernels. For each convolution kernel, the central processing unit can input the convolution kernel and the preset weight matrix into the accelerator. Then, the accelerator can perform a convolution operation on the convolution kernel and the preset weight matrix to obtain the convolution sum corresponding to the convolution kernel. The accelerator can include multiple acceleration devices, each of which can perform a convolution operation on a convolution kernel and a preset weight matrix, and multiple acceleration devices can perform convolution operations in parallel. For example, Figure 2 The accelerator in the black bold box can perform convolution operations on the preset weight matrix and the convolution kernel indicated by the sliding window.

[0040] When the first preset value is 1 and the second preset value is 0, all valid data required for convolution calculation are multiplication and addition with data 1 or 0, and no multiplication calculation is required. Figure 3 As shown. Accordingly, the convolution calculation process can be as follows Figure 4 As shown, no multiplier is required. In the accelerator, the first preset value and the second preset value serve as control signals for data selection, controlling whether the weighted data or 0 is involved in the calculation. Figure 4 As shown, when the value corresponding to the same coordinate is 0 and the weight value is w1, the value of the input adder is 0; when the value corresponding to the same coordinate is 1 and the weight value is w3, the value of the input adder is w3, and so on. Figure 4 The input data on the right can then be used to calculate the convolution sum of a convolution kernel through multiple adders in the next stage.

[0041] In the above-mentioned process of calculating the convolution sum, the central processing unit needs to receive a complete frame of pulse data before calculating the convolution sum. The required storage space is large, and the number of adders used in the operation process is large, which occupies more computing resources. For example, in the above-mentioned Figure 3 In the example, an accelerator can be used to perform convolution operations on one or more convolution kernels. Each accelerator can include the above Figure 4 There are nine adders in , which requires a large number of adders.

[0042] In order to solve the above problems, the embodiment of the present application provides a convolution operation system, such as Figure 5 As shown, the convolution operation system may include a storage controller 110 , a memory 120 , and an operation circuit 130 .

[0043] The memory 120 may be a random access memory (RAM).

[0044] The storage controller 110 may be configured to read a preset number of rows of pulse data from the memory 120. Based on a first preset convolution kernel side length equal to the preset number of rows, a second preset convolution kernel side length, the pulse data, and a preset weight matrix, the memory controller 110 may determine at least one submatrix and the number of columns of each submatrix within the at least one submatrix. The memory controller 110 may also transmit each submatrix and the number of columns of each submatrix to the computing circuit 130.

[0045] The operation circuit 130 may be configured to determine the convolution sum of the target sub-matrix according to the target sub-matrix and the number of columns of the target sub-matrix.

[0046] The target sub-matrix is ​​any sub-matrix in the at least one sub-matrix.

[0047] Specifically, the storage controller 110 can detect whether the number of rows of pulse data stored in the memory 120 is equal to a preset number of rows. When it is determined that the number of rows of pulse data stored in the memory 120 is equal to the preset number of rows, the storage controller 110 can read the preset number of rows of pulse data from the memory 120. After reading the preset number of rows of pulse data, the storage controller 110 can process the pulse data according to the first preset convolution kernel side length, the second preset convolution kernel side length, and the preset weight matrix to obtain at least one submatrix. In addition, the storage controller 110 also needs to determine the number of columns of each submatrix. Further, each submatrix and the number of columns of each submatrix can be transmitted to the operation circuit 130. For any submatrix, the operation circuit 130 can use the number of columns of the submatrix as a statistical threshold and calculate the sub-sum value of each column in turn. When it is determined that the number of calculated sub-sum values ​​is equal to the statistical threshold, it means that the sub-sum values ​​of all columns of the submatrix have been calculated. The summation operation can be directly performed to obtain the convolution sum of the submatrix. Similarly, the convolution sum of each sub-matrix can be determined.

[0048] In the convolution operation system of the embodiment of the present application, the storage space of the memory 120 only needs to store a preset number of rows of pulse data. The preset number of rows is equal to the side length of the first preset convolution kernel, and the side length of the first preset convolution kernel is smaller than the width of a frame of pulse data. Therefore, storage resources can be saved. Furthermore, after reading the preset number of rows of pulse data, the storage controller 110 can determine at least one submatrix based on the first preset convolution kernel side length, the second preset convolution kernel side length, the pulse data, and the preset weight matrix, and determine the number of columns of each submatrix. Each submatrix and the corresponding column number are transmitted to the operation circuit 130. In the related art, the size of each convolution kernel is equal, and the number of convolution kernels is large. The convolution calculation process of each convolution kernel requires a large amount of computing resources. However, some convolution kernels contain all zero data, which wastes a lot of resources. In this solution, the submatrix containing valid data is determined. Because the number of columns of different submatrices varies, the number of columns of each submatrix can be determined. In order to save computing resources, the operation circuit 130 can calculate by column, and accordingly, the storage controller 110 can send each sub-matrix and the number of columns of each sub-matrix to the operation circuit 130. In this way, the operation circuit 130 can accurately identify how many columns are a sub-matrix, and summarize the results of the column calculation to determine the convolution sum of the target sub-matrix. Since the size of the sub-matrix is ​​less than or equal to the size of the convolution kernel, and the operation process is to sum by column, rather than summing the values ​​of each coordinate in the convolution kernel separately as in the related art, that is, the objects to be summed are reduced, which can reduce the occupation of computing resources. In summary, by reducing the occupation of computing resources and storage resources, the hardware resources occupied by convolution operations can be reduced.

[0049] The specific functions of the storage controller 110 and the operation circuit 130 are described below.

[0050] First, the storage controller 110 processes the pulse data according to the first preset convolution kernel side length, the second preset convolution kernel side length, and the preset weight matrix to obtain at least one submatrix. The specific process may include:

[0051] Step 1: Determine multiple convolution kernels according to the first preset convolution kernel side length, the second preset convolution kernel side length, and the pulse data.

[0052] Step 2: Select the convolution kernels to be sorted that meet the target conditions from multiple convolution kernels.

[0053] Step 3: After screening out at least one convolution kernel to be sorted, each convolution kernel to be sorted is eliminated to obtain a candidate convolution kernel corresponding to each convolution kernel to be sorted.

[0054] Step 4: Generate a submatrix corresponding to each candidate convolution kernel based on each candidate convolution kernel and the preset weight matrix.

[0055] In step 1, the storage controller 110 may construct a sliding window based on the first and second preset convolution kernel side lengths, and sequentially slide the window from left to right within the pulse data. After each sliding operation, the pulse data falling within the sliding window is determined as a convolution kernel. In this way, multiple convolution kernels that match the first and second preset convolution kernel side lengths can be determined within the pulse data.

[0056] In step 2, the storage controller 110 can analyze each convolution kernel according to the target condition and determine whether each convolution kernel meets the target condition. If so, the convolution kernel is determined to be to be sorted; if not, the convolution kernel is eliminated.

[0057] The target condition may be that the value of the center coordinate is a first preset value, for example, the first preset value may be 1. Accordingly, the storage controller 110 may determine the convolution kernel whose center coordinate value is the first preset value as the convolution kernel to be sorted that meets the target condition. Since the value of the center coordinate is a second preset value, indicating that the pulse data in the convolution kernel is invalid data, pre-empting the invalid data can reduce computing resources wasted in subsequent calculations of the invalid data, thereby improving the efficiency of the convolution operation.

[0058] In step three, for each convolution kernel to be sorted, the columns that do not need to participate in the calculation can be eliminated in the convolution kernel to be sorted, so as to reduce the occupation of computing resources and the number of calculation steps, thereby improving the calculation efficiency of the convolution sum.

[0059] Taking the target convolution kernel to be sorted (any one of the at least one convolution kernel to be sorted) as an example, the storage controller 110 may perform a culling operation on the target convolution kernel to be sorted through the following specific steps:

[0060] Step 1: Determine whether the values ​​of all coordinates in the i-th column of the target convolution kernel to be sorted are all second preset values.

[0061] Wherein, i is a positive integer greater than 0 and less than or equal to the side length of the second preset convolution kernel.

[0062] Step 2: When it is determined that the values ​​of all coordinates in the i-th column are all the second preset values, the i-th column is removed from the target convolution kernel to be sorted. Alternatively, when it is determined that the value of any coordinate in the i-th column is not the second preset value, the i-th column is retained in the target convolution kernel to be sorted.

[0063] Step 3: After completing the removal operation or retention operation of each column in the target convolution kernel to be sorted, a candidate convolution kernel corresponding to the target convolution kernel to be sorted is obtained.

[0064] The storage controller 110 can traverse each column of the target convolution kernel to be sorted, and each time it traverses a column, it can determine whether the values ​​of all coordinates in the column are all second preset values ​​(for example, it can be 0). If so, the column can be removed from the target convolution kernel to be sorted. If not, the column data can be retained and the next column in the target convolution kernel to be sorted can be traversed. After traversing each column in the target convolution kernel to be sorted and completing the corresponding removal operation or retention operation, a candidate convolution kernel corresponding to the target convolution kernel to be sorted can be obtained. For each convolution kernel to be sorted, the removal operation can be performed in the above manner to obtain candidate convolution kernels corresponding to all convolution kernels to be sorted.

[0065] Since the subsequent calculation process is performed in columns, if all the values ​​in the column are the second preset values, it means that the data in this column is invalid data. The data in this column can be eliminated to reduce the occupancy of computing resources and the number of calculation steps, thereby improving the calculation efficiency of the convolution sum.

[0066] In step 4, for each candidate convolution kernel, it is not necessary to directly perform a convolution operation with the preset weight matrix as in the related art. Instead, the preset weight matrix is ​​used as a reference to directly replace the values ​​of the coordinates included in the candidate convolution kernel. Accordingly, the storage controller 110 can perform the replacement operation according to the following specific steps:

[0067] Step 1: Get the coordinates of the target candidate convolution kernel to be replaced.

[0068] Among them, the target candidate convolution kernel is any candidate convolution kernel.

[0069] Step 2: According to the coordinates to be replaced, determine the target weight value corresponding to the coordinates to be replaced in the preset weight matrix.

[0070] Among them, the target candidate convolution kernel is any candidate convolution kernel, and the coordinates to be replaced are the coordinates in the target candidate convolution kernel whose values ​​are not the second preset values.

[0071] Step 3: Replace the values ​​of the coordinates to be replaced in the target candidate convolution kernel with the target weight values ​​to obtain the submatrix corresponding to the target candidate convolution kernel.

[0072] The storage controller 110 can traverse the numerical value of each coordinate of the target candidate convolution kernel, and each time it traverses to the numerical value of a coordinate, it determines whether the numerical value of the coordinate is a first preset numerical value. If not, it can continue to traverse the numerical value of the next coordinate in the target candidate convolution kernel. If so, the coordinate is determined as the coordinate to be replaced. Accordingly, the storage controller 110 can determine the target weight value corresponding to the coordinate to be replaced in the preset weight matrix based on the coordinate to be replaced, and then replace the numerical value of the coordinate to be replaced in the target candidate convolution kernel with the target weight value. And so on, until all the numerical values ​​of the coordinates in the target candidate convolution kernel are traversed, the submatrix corresponding to the target candidate convolution kernel can be obtained. For each candidate convolution kernel, the above method can be used to determine the corresponding submatrix. In this way, by replacing the numerical value of the coordinate to be replaced in the candidate convolution kernel with the corresponding weight value in advance, there is no need to subsequently perform addition calculation on the preset weight matrix and the convolution kernel through an adder, which can reduce the number of adders and further reduce the occupation of hardware resources.

[0073] Second, the operation circuit 130 may determine the convolution of the target sub-matrix according to the target sub-matrix and the number of columns of the target sub-matrix, and the specific process may include:

[0074] Step 1: sum the values ​​of all coordinates of each column in the target submatrix to obtain the sub-sum value of each column in the target submatrix.

[0075] Step 2: After determining that the number of obtained sub-sum values ​​is equal to the number of columns of the target sub-matrix, the sub-sum values ​​of each column in the target sub-matrix are summed to obtain the convolution sum of the target sub-matrix.

[0076] Taking the target submatrix as an example, when the storage controller 110 sends the target submatrix, it can send pulse data from left to right or from right to left in the column direction. Each time the operation circuit 130 receives a column of pulse data, it can sum the values ​​of all coordinates included in the column of pulse data to obtain the sub-sum value of the column. From the perspective of the operation circuit 130, the operation circuit 130 does not know the number of columns of the target submatrix. Therefore, the storage controller 110 needs to send the number of columns of the target submatrix to the operation circuit 130 in advance. In this way, the operation circuit 130 can understand which columns' sub-sum values ​​are the sum values ​​of a submatrix, thereby accurately calculating the convolution sum.

[0077] Each time the computation circuit 130 obtains a sub-sum value for a column, it may add that sub-sum value to a previously accumulated value (the previously calculated accumulated value of the sub-sum values ​​for other columns in the target sub-matrix). For example, the sub-sum value for the first column calculated for the sub-matrix may be stored. When the sub-sum value for the second column is obtained, it may be added to the previously calculated sub-sum value and stored. When the sub-sum value for the third column is obtained, the previously accumulated sum value and the newly calculated sub-sum value may be added to the accumulated sum value and stored. Simultaneously, the computation circuit 130 may count the number of calculated sub-sum values. When the number of sub-sum values ​​is determined to be equal to the number of columns in the sub-matrix, the accumulated sum value is determined as the convolution sum of the sub-matrix. The stored accumulated sum value and the counted sub-sum values ​​are cleared to zero, thereby recording the sub-sum values ​​generated during the convolution operation for the next sub-matrix and the number of counted sub-sum values. In this way, the computation circuit 130 completes the convolution sum calculation process for a sub-matrix.

[0078] Storage controller 110 continues to send subsequent sub-matrices and sub-matrix column numbers in a similar manner, and computation circuit 130 completes the convolution and computation process for each sub-matrix. Because the number of coordinates in each column of a sub-matrix is ​​much smaller than the number of coordinates in a convolution kernel, fewer elements need to be calculated during the convolution and computation, saving computing resources and improving computational efficiency.

[0079] The structure of the operation circuit 130 will be described in detail below.

[0080] like Figure 6 As shown, the arithmetic circuit 130 may include a plurality of adders and accumulation circuits 130 a , and the plurality of adders may include a primary adder 130 b and a secondary adder 130 c .

[0081] The primary adder 130b may be configured to sum the values ​​of the first coordinate and the second coordinate in the j-th column of the target submatrix to obtain a primary sum value, and output the sum value to the secondary adder 130c.

[0082] Wherein, the first coordinate is any coordinate in the j-th column, the second coordinate is any coordinate in the j-th column except the first coordinate, the j-th column is any column in the target submatrix, and j is a positive integer greater than 0 and less than or equal to the number of columns of the target submatrix.

[0083] The secondary adder 130 c may be configured to sum the primary sum value and the values ​​of the coordinates other than the first coordinate and the second coordinate in the j-th column to obtain a sub-sum value of the j-th column and transmit the sub-sum value to the accumulator circuit 130 a .

[0084] Accumulation circuit 130a may be configured to obtain a target accumulated value stored in a preset storage location. The target accumulated value and the sub-sum value of the j-th column are summed to obtain an updated accumulated value. After determining that the number of received sub-sum values ​​equals the number of columns of the target sub-matrix, the updated accumulated value is determined as the convolution sum of the target sub-matrix.

[0085] The storage controller 110 may input the values ​​of the first coordinate and the second coordinate in the j-th column of the target submatrix into a primary adder 130b. The primary adder 130b may sum the values ​​of the first coordinate and the second coordinate to obtain a primary sum. The primary adder 130b may input the primary sum into a secondary adder 130c. The storage controller 110 may also input the values ​​of the other coordinates in the j-th column of the target submatrix, except for the first coordinate and the second coordinate, into a secondary adder 130c. The secondary adder 130c may sum the primary sum input by the primary adder 130b and the values ​​of the other coordinates input by the storage controller 110 to obtain a sub-sum value for the j-th column and transmit it to the accumulation circuit 130a. The accumulation circuit 130a may sum the target accumulated value stored in the preset storage location with the sub-sum value of the j-th column to obtain an updated accumulated value. At the same time, the accumulation circuit 130a can count whether the number of currently calculated sub-sum values ​​is equal to the number of columns of the target sub-matrix. If so, it means that the convolution sum of the target sub-matrix has been calculated, and the updated accumulated value can be determined as the convolution sum of the target sub-matrix. If not, it means that the convolution sum of the target sub-matrix has not been calculated yet, and the accumulation calculation can be repeated after the sub-sum value of the next column is calculated. The accumulation calculation is performed again until the number of currently calculated sub-sum values ​​is counted to be equal to the number of columns of the target sub-matrix. The final accumulated value is then determined as the convolution sum of the target sub-matrix. In this way, the operation circuit 130 only needs a small number of adders to perform the sum calculation of the values ​​in one column of the sub-matrix, greatly reducing the number of required adders.

[0086] The number of secondary adders 130c included in the operation circuit 130 may be one or more. When there are multiple coordinates other than the first coordinate and the second coordinate in the j-th column, there are multiple secondary adders 130c. Accordingly, the storage controller 110 may input the value of each of the other coordinates into the secondary adder 130c corresponding to the row where the coordinate is located, and the secondary adder 130c sums the sum values ​​input from the other adders and the value of the coordinate input from the storage controller 110.

[0087] Specifically, the primary adder 130b can be connected to the first secondary adder 130c among the multiple secondary adders 130c. Accordingly, the primary adder 130b can input the primary sum value into the first secondary adder 130c. The storage controller 110 can input the value of the coordinate corresponding to the first secondary adder 130c in the j-th column into the first secondary adder 130c. The first secondary adder 130c can sum the values ​​of the two input coordinates to obtain a secondary sum value, and output it to the next secondary adder 130c connected to the first secondary adder 130c. Similarly, the storage controller 110 can input the values ​​of the coordinates corresponding to the next secondary adder 130c in the j-th column into the next secondary adder 130c, and the next secondary adder 130c can sum the values ​​of the two input coordinates. Similarly, the sub-sum calculated by the last sub-adder 130c is the sub-sum of the j-th column. The last sub-adder 130c is connected to the accumulator 130a and can input the sub-sum of the j-th column into the accumulator 130a, which then performs subsequent operations. This allows for flexible use of the corresponding number of adders for convolution kernels of different sizes.

[0088] In some optional embodiments, when it is determined that the number of received sub-sum values ​​is equal to the number of columns of the target sub-matrix, the accumulation circuit 130a can also be used to reset the register to clear the value stored in the register. In this way, the accumulation circuit 130a can record the convolution sum generated during the convolution operation of the next sub-matrix and perform the accumulation operation.

[0089] In some optional embodiments, when it is determined that the number of received sub-sum values ​​is equal to the number of columns of the target sub-matrix, the accumulation circuit 130a may be further configured to send a completion signal to the memory controller 110. After receiving the completion signal, the memory controller 110 may be configured to read the updated accumulation value from the register and transmit it to the pulse data updater, so that the pulse data updater determines whether to output pulse data corresponding to the target sub-matrix.

[0090] In some optional embodiments, such as Figure 7 As shown, the accumulation circuit 130a may include an accumulator adder 130a1, a resetter 130a2, and a register 130a3. The preset storage location may be register 130a3. Accordingly, the accumulator adder 130a1 may be configured to accumulate the sub-sum value of the j-th column input by the last secondary adder 130c and the target accumulated value stored in register 130a3. The resetter 130a2 may be configured to increment a count value by one each time it detects that the last adder inputs a sub-sum value to the accumulator adder 130a1. When the count value equals the number of columns of the target sub-matrix, the updated accumulated value stored in register 130a3 is determined as the convolution sum of the target sub-matrix.

[0091] In addition, when the resetter 130a2 determines that the count value is equal to the number of columns of the target sub-matrix, it can also send a convolution and calculation completion notification to the storage controller 110. Upon receiving the convolution and calculation completion notification, the storage controller 110 reads the updated accumulated value from the register 130a3 as the convolution sum of the target sub-matrix and inputs it to the pulse data updater. Furthermore, after completing the operation of outputting the convolution sum of the target sub-matrix, the storage controller 110 can send a reset notification to the resetter 130a2, and the resetter 130a2 can send a reset signal to the register 130a3 to clear the value stored in the register 130a3.

[0092] By designing the cumulative adder 130a1, the resetter 130a2 and the register 130a3 in the accumulation circuit 130a, each column of the submatrix can be accumulated step by step, which greatly reduces the number of required adders and thus reduces the occupation of hardware resources.

[0093] In some optional embodiments, the pulse data updater may include a convolution sum adder, a target memory, and a comparator. The storage controller 110 may read the first convolution sum corresponding to the first center coordinate from the target memory based on the center coordinate of the target submatrix (hereinafter referred to as the first center coordinate) (this convolution sum may be the convolution sum calculated based on the pulse data of the previous frame). The storage controller 110 may input the convolution sum of the target submatrix and the first convolution sum into the convolution sum adder, and the convolution sum adder may sum the two input values ​​and input the obtained target convolution sum into the comparator. The comparator may compare the target convolution sum with a preset threshold value. When it is determined that the target convolution sum is greater than the preset threshold value, the pulse data of the center coordinate is output, or, when it is determined that the target convolution sum is less than or equal to the preset threshold value, the target convolution sum may be overwritten and written to the storage location corresponding to the center coordinate in the target memory.

[0094] After completing the output operation of the convolution sum of each submatrix, the storage controller 110 may delete the first row of pulse data from the preset number of rows of pulse data in the memory 120 to store the next row of pulse data. The storage controller 110 may then recognize that the preset number of rows of pulse data is stored in the memory 120, read the data, and perform the next round of convolution operations. This process can be repeated in this manner to complete the convolution operations for a frame of pulse data.

[0095] In some optional embodiments, the maximum storage space of memory 120 can store a preset number of rows plus n rows of pulse data (for example, n can be 1). While completing the convolution operation on all sub-matrices corresponding to the preset number of rows of pulse data, memory controller 110 can continue writing the next row of pulse data. In this way, the clock frequency of operation circuit 130 only needs to be no less than the data input frequency. That is, while the memory controller 110 is receiving a row of pulse data, the convolution operation system can complete the convolution operation on all sub-matrices corresponding to the preset number of rows of pulse data.

[0096] In some optional embodiments, the convolution operation system may further include a pulse data generator (e.g., a pulse camera). Accordingly, the storage controller 110 may be further configured to, upon receiving coordinate information sent by the pulse data generator, set the value of the storage location corresponding to the coordinate information in the memory 120 to a first preset value based on the coordinate information.

[0097] The coordinate information may be coordinate information corresponding to a line of pulse data, and the coordinate information may include coordinate information of pixels whose brightness changes.

[0098] Specifically, after identifying the locations where the brightness changes, the pulse data generator can determine the coordinate information corresponding to these locations and send the coordinate information to the storage controller 110 row by row. Each time the storage controller 110 receives the coordinate information, it can determine the storage locations corresponding to the coordinate information in the memory 120 based on the coordinate information, set the values ​​of these storage locations to the first preset value, and set the values ​​of other storage locations in the memory 120 to the second preset threshold.

[0099] Taking a frame of pulse data as an example, when the storage controller 110 receives a preset number of rows of pulse data, it can immediately start processing and perform a convolution operation. During this process, the storage controller 110 can continue to receive the next row of coordinate information in the frame of pulse data and write the corresponding pulse data into the memory 120 (that is, the above-mentioned operation of setting the first preset value and the second preset value). After completing the convolution operation of the first preset number of rows of pulse data, the first row of data in the preset number of rows of pulse data will be deleted. At this time, a new row of pulse data has been written. Further, the storage controller 110 can re-read the new preset number of rows of pulse data for convolution operation processing. And so on, until all convolution operations for the frame of pulse data are completed. In this way, there is no need for the central processing unit to perform decoding operations. Instead, the storage controller 110 can directly set the pulse data in the memory 120, which can reduce the hardware resources occupied by the central processing unit.

[0100] The above-mentioned convolution operation system can be an accelerator. After the above-mentioned improvements, the area of ​​the accelerator is reduced. That is to say, the accelerator can be installed in computer equipment, such as computers, servers, etc., or directly installed in pulse cameras, which can speed up the processing efficiency of pulse data.

[0101] The following describes in detail the execution process of the convolution operation method performed by the above-mentioned convolution operation system using a specific example.

[0102] When the convolution operation system is set as an accelerator in the pulse camera, the pulse data processing process can be done without the participation of the central processor, such as Figure 8 shown.

[0103] The image size can be 8×5, and the preset convolution kernel size can be 3×3, that is, the side length of the first preset convolution kernel and the side length of the second preset convolution kernel can both be 3. The number of memories can be 4. For example, 4 single-port RAMs are set in the accelerator. The length of the RAM can be the same as the length of the image, 8. The initial value of the 8 coordinates included in each RAM can be 0. Accordingly, it can be as follows Figure 9 Thus, setting the number of memories to be 1 more than the side length of the second preset convolution kernel allows caching the fourth row of pulse data while processing the first three rows of pulse data, thereby improving data processing efficiency.

[0104] When the pulse camera detects a change in light, it determines the coordinates of the brightness change and inputs these coordinates, row by row, into the accelerator. The accelerator then generates pulse data based on these coordinates and performs a convolution operation on the pulse data. For example, the coordinates might include [2, 2], [2, 8], [3, 6], [4, 1], and so on.

[0105] When sending coordinate information, the pulse camera can send it by line. Figure 10 As shown in the figure, when the accelerator receives the coordinate information of the first row ([1, 2], [1, 5], [1, 7]), the accelerator sets the value of the storage location corresponding to the coordinate information in the first row of RAM to 1. Similarly, the setting operation of each row of pulse data can be completed. Since the preset convolution kernel size is 3×3, after completing the setting operation of three rows of pulse data, the data in the memory can be as follows Figure 11 As shown on the left, the accelerator can determine the second matrix based on the three rows of pulse data stored in the memory. At the same time, the accelerator can continue to receive the coordinate information of the fourth row and perform a value setting operation on the RAM of the fourth row.

[0106] When determining the second matrix, the accelerator can slide the convolution kernel window from left to right to obtain multiple convolution kernels. Among these multiple convolution kernels, there are only two convolution kernels whose center coordinate value is 1, that is, Figure 11 The two convolution kernels indicated by the bold box (i.e., the convolution kernels to be sorted). For the first convolution kernel, each column has a coordinate value of 1, and no culling operation is required. For the second convolution kernel, all the coordinate values ​​of the first column are 0, so the column can be culled. After completing the column culling operation or the retention operation, the second matrix can be obtained (as shown in Figure 11 The second matrix includes two sub-matrices. In this way, sparse data is eliminated without affecting the calculation results, and the number of addition operations can be reduced, which can improve the calculation speed and save computing resources.

[0107] The preset weight matrix can be Figure 12 The accelerator can use the preset weight matrix to replace the positions where the coordinate values ​​in the above two sub-matrices are 1 to obtain a third matrix, as shown in Figure 13 As shown. When performing the replacement operation, the coordinates of the submatrix with a value of 1 can be determined, and based on this coordinate, the corresponding weight value can be determined in the preset weight matrix, and then the determined weight value can be replaced with the corresponding position of the submatrix. For example, the value of the coordinate [1, 1] of the first submatrix is ​​replaced by w1, the value of the coordinate [2, 2] of the first submatrix is ​​replaced by w5, and the value of the coordinate [3, 3] of the first submatrix is ​​replaced by w9.

[0108] The accelerator can calculate the sum of each column from right to left according to the pipeline calculation method of the third matrix. Figure 14As shown, the accelerator can first input w3 and 0 into the first adder (which can be the primary adder 130b described above) to obtain the sum 1. This sum 1 and w9 are then input into the second adder (which can be the secondary adder 130c described above) to obtain the sum 2. This sum 2 is the calculated sub-sum of the first column (the first column from right to left). The second adder outputs sum 2 to the third adder. The third adder (which can be the accumulator 130a1 described above) adds sum 2 and the 0 in the register to obtain sum 3 (which can be the updated accumulated value described above) and transfers it to the register. When the resetter 130a2 detects that the second adder has transmitted the first sum, it increments the counter by one.

[0109] Furthermore, the accelerator inputs 0 and w5 in the second column of the first matrix from right to left into the first adder, obtaining a sum of 4. This sum and 0 are then input into the second adder, obtaining a sum of 5. This sum is the calculated sub-sum of the second column. The second adder inputs the sum 5 into the third adder, which then adds the sum 5 to the sum 4 in the register to obtain a sum of 6. When resetter 130a2 detects that the second adder has transmitted the second sum, it increments the count by one. This count now equals the number of columns in the second sub-matrix, indicating that the convolution sum calculation of the second sub-matrix is ​​complete and that the sum 6 can be processed further.

[0110] like Figure 15 As shown, the accelerator can read the second convolution sum from the storage location corresponding to the center coordinate in the cache value stored in the target memory (which can be a cache) according to the center coordinate of the second sub-matrix, and the storage controller can input the convolution sum of the second sub-matrix and the second convolution sum into the convolution sum adder. The convolution sum adder sums the two input values ​​and inputs the obtained total convolution sum into the comparator. The comparator can compare the total convolution sum with a preset threshold. When it is determined that the total convolution sum is greater than the preset threshold, the pulse data of the center coordinate (for example, 1) is output, or, when it is determined that the total convolution sum is less than or equal to the preset threshold, the total convolution sum can be overwritten and written to the storage location corresponding to the center coordinate in the target memory.

[0111] The embodiment of the present application provides a convolution operation method, which can be applied to the above-mentioned convolution operation system, such as Figure 16 As shown, the specific processing steps of the convolution operation method may include:

[0112] Step S1601: In the current round, the memory controller reads a preset number of rows of pulse data from the memory.

[0113] In step S1602, the storage controller determines at least one sub-matrix based on a first preset convolution kernel side length equal to a preset number of rows, a second preset convolution kernel side length, pulse data, and a preset weight matrix.

[0114] Step S1603: The storage controller determines the number of columns of each sub-matrix in at least one sub-matrix.

[0115] In step S1604 , the storage controller sends each sub-matrix and the number of columns of each sub-matrix to the operation circuit.

[0116] In step S1605 , the operation circuit determines the convolution sum of the target sub-matrix according to the target sub-matrix and the number of columns of the target sub-matrix.

[0117] The target sub-matrix is ​​any sub-matrix in the at least one sub-matrix.

[0118] The specific processing of steps S1601 to S1605 can refer to the above introduction to the convolution operation system and will not be repeated here.

[0119] In the convolution operation method of the embodiment of the present application, the storage space of the memory only needs to store a preset number of rows of pulse data. The preset number of rows is equal to the side length of the first preset convolution kernel, and the side length of the first preset convolution kernel is smaller than the width of a frame of pulse data. Therefore, memory resources can be saved. Furthermore, after the storage controller reads the preset number of rows of pulse data, it can determine at least one submatrix based on the side length of the first preset convolution kernel, the side length of the second preset convolution kernel, the pulse data, and the preset weight matrix, and determine the number of columns of each submatrix, and transmit each submatrix and the corresponding column number to the operation circuit. In the related art, the size of each convolution kernel is equal, and the number of convolution kernels is large. The convolution calculation process of each convolution kernel requires a large amount of computing resources. However, some convolution kernels are all zero data, and participating in the operation will waste a lot of resources. In this solution, the submatrix containing valid data is determined. Since the number of columns of different submatrices varies, the number of columns of each submatrix can be determined. In order to save computing resources, the operation circuit can calculate by column, and accordingly, the storage controller can send each sub-matrix and the number of columns of each sub-matrix to the operation circuit. In this way, the operation circuit can accurately identify how many columns constitute a sub-matrix, and summarize the results of the column calculation to determine the convolution sum of the target sub-matrix. Since the size of the sub-matrix is ​​less than or equal to the size of the convolution kernel, and the operation process is summed by column, rather than summing the values ​​of each coordinate in the convolution kernel separately as in the related art, that is, the number of objects to be summed is reduced, which can reduce the computing resource usage. In summary, by reducing the usage of computing resources and storage resources, the hardware resources occupied by the convolution operation can be reduced.

[0120] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0121] The embodiment of the present application also provides an electronic device, such as Figure 17 As shown, the electronic device includes a first memory 10 and a processor 20. The first memory 10 stores a computer program, and the processor 20 is configured to run the computer program to perform the steps of any of the above-mentioned convolution operation method embodiments. The electronic device can be one of the above-mentioned accelerators, pulse cameras, and computer devices.

[0122] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned convolution operation method embodiments when run.

[0123] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0124] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned convolution operation method embodiments are implemented.

[0125] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned convolution operation method embodiments are implemented.

[0126] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0127] The above is a detailed introduction to the convolution operation system, method, electronic device, storage medium, and program product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.

Claims

1. A convolution operation system, characterized in that: The convolution operation system includes a storage controller, a memory, and an operation circuit; The memory controller is configured to read a preset number of rows of pulse data from the memory; Determining a plurality of convolution kernels according to a first preset convolution kernel side length equal to the preset number of rows, a second preset convolution kernel side length, and the pulse data; and screening a convolution kernel to be sorted that meets a target condition from the plurality of convolution kernels; In the case of screening out at least one of the convolution kernels to be sorted, when determining whether the values ​​of all coordinates in the i-th column in the target convolution kernel to be sorted are all second preset values, wherein the target convolution kernel to be sorted is any one of the at least one convolution kernel to be sorted, and i is a positive integer greater than 0 and less than or equal to the side length of the second preset convolution kernel; when it is determined that the values ​​of all coordinates in the i-th column are all the second preset values, the i-th column is removed from the target convolution kernel to be sorted; after completing the removal operation or retention operation of each column in the target convolution kernel to be sorted, a candidate convolution kernel corresponding to the target convolution kernel to be sorted is obtained; according to each of the candidate convolution kernels and the preset weight matrix, a submatrix corresponding to each of the candidate convolution kernels is generated; the number of columns of each of the submatrices in at least one submatrix is ​​determined respectively; each of the submatrixes and the number of columns of each of the submatrixes is sent to the operation circuit; The operation circuit is used to determine the convolution sum of the target submatrix according to the target submatrix and the number of columns of the target submatrix, wherein the target submatrix is ​​any one of the at least one submatrix.

2. The convolution operation system according to claim 1, wherein: The storage controller is specifically configured to: The convolution kernel whose center coordinate value among the multiple convolution kernels is a first preset value is determined as the convolution kernel to be sorted.

3. The convolution operation system according to claim 1 or 2, characterized in that: The storage controller is further configured to: When it is determined that the value of any coordinate in the i-th column is not the second preset value, the i-th column is retained in the target convolution kernel to be sorted.

4. The convolution operation system according to claim 2, wherein: The storage controller is specifically configured to: Obtaining the coordinates of the target candidate convolution kernel to be replaced, wherein the target candidate convolution kernel is any one of the candidate convolution kernels; Determining, according to the coordinates to be replaced, a target weight value corresponding to the coordinates to be replaced in the preset weight matrix, wherein the coordinates to be replaced are coordinates in the target candidate convolution kernel whose values ​​are the first preset values; The values ​​of the coordinates to be replaced in the target candidate convolution kernel are replaced with the target weight values ​​to obtain a submatrix corresponding to the target candidate convolution kernel.

5. The convolution operation system according to claim 1 or 2, characterized in that: The arithmetic circuit includes a plurality of adders and accumulator circuits, wherein the plurality of adders include primary adders and secondary adders; The primary adder is configured to sum the values ​​of the first coordinate and the second coordinate in the j-th column of the target submatrix to obtain a primary sum value, and output the sum value to the secondary adder, wherein the first coordinate is any coordinate in the j-th column, the second coordinate is any coordinate in the j-th column except the first coordinate, the j-th column is any column in the target submatrix, and j is a positive integer greater than 0 and less than or equal to the number of columns of the target submatrix; The secondary adder is configured to sum the primary sum and the values ​​of the coordinates in the j-th column except the first coordinate and the second coordinate to obtain a sub-sum value in the j-th column and transmit the sub-sum value to the accumulator circuit; The accumulator circuit is configured to obtain a target accumulated value stored in a preset storage location; sum the target accumulated value and the sub-sum value of the j-th column to obtain an updated accumulated value; and after determining that the number of received sub-sum values ​​is equal to the number of columns of the target sub-matrix, determine the updated accumulated value as the convolution sum of the target sub-matrix.

6. The convolution operation system according to claim 5, characterized in that: When it is determined that the number of received sub-sum values ​​is equal to the number of columns of the target sub-matrix, the accumulation circuit is further configured to perform a reset operation on the preset storage location to clear the updated accumulated value stored in the preset storage location.

7. The convolution operation system according to claim 5, wherein: In the case of determining that the number of received sub-sum values ​​is equal to the number of columns of the target sub-matrix, the accumulating circuit is further configured to send a completion signal to the storage controller; The storage controller is used to read the updated accumulated value from the preset storage location after receiving the completion signal, and transmit it to the pulse data updater, so that the pulse data updater determines whether to output the pulse data corresponding to the target sub-matrix.

8. The convolution operation system according to claim 2, wherein: The convolution operation system also includes a pulse data generator; the storage controller is also used to set the value of the storage position corresponding to the coordinate information in the memory to the first preset value according to the coordinate information after receiving the coordinate information sent by the pulse data generator.

9. A convolution operation method, characterized in that: The convolution operation method is applied to a convolution operation system, which includes a storage controller, a memory, and an operation circuit; the method includes: In the current round, the storage controller reads a preset number of rows of pulse data from the memory; determines a plurality of convolution kernels according to a first preset convolution kernel side length equal to the preset number of rows, a second preset convolution kernel side length, and the pulse data; filters out convolution kernels to be sorted that meet the target conditions from the plurality of convolution kernels; when at least one convolution kernel to be sorted is screened out, determines whether the values ​​of all coordinates of the i-th column in the target convolution kernel to be sorted are all the second preset values, wherein the target convolution kernel to be sorted is any one of the at least one convolution kernel to be sorted, and i is greater than 0 and less than A positive integer equal to the side length of the second preset convolution kernel; when it is determined that the values ​​of all coordinates in the i-th column are all the second preset values, the i-th column is removed from the target convolution kernel to be sorted; after completing the removal operation or retention operation of each column in the target convolution kernel to be sorted, a candidate convolution kernel corresponding to the target convolution kernel to be sorted is obtained; according to each candidate convolution kernel and a preset weight matrix, a submatrix corresponding to each candidate convolution kernel is generated; the number of columns of each submatrix in at least one submatrix is ​​determined respectively; each submatrix and the number of columns of each submatrix are sent to the operation circuit; The operation circuit determines the convolution sum of the target submatrix according to a target submatrix and the number of columns of the target submatrix, wherein the target submatrix is ​​any one of the at least one submatrix.

Citation Information

Patent Citations

  • Sparse convolutional neural network sorting method, operation method and device and equipment

    CN112200295A

  • Convolution acceleration method and convolution accelerator for spiking neural network

    CN116720551A