Convolution calculation method and device based on brain-like chip hardware optimization

By mapping the convolution kernel in a semi-folded form on the brain-like chip and building a delay layer and a convolutional computing layer, the problems of large resource overhead and long calculation time in the existing technology are solved, and efficient convolutional computing and resource utilization are achieved.

CN119988804APending Publication Date: 2025-05-13PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411811284.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When deploying convolution operators to brain-like chips, the resource overhead is large and the calculation time is long, limiting the deployable convolution scale and the overall scale of neural networks, resulting in a degradation in the performance of inference tasks.

Method used

The convolution kernel is mapped to the cross-array architecture of brain-like chips in a semi-folded form. By building a delay layer and a convolutional calculation layer, convolutional calculation is performed using time division multiplexing to reduce the use of hardware resources and improve the computing speed.

Benefits of technology

While ensuring computing efficiency, it significantly reduces resource occupancy and power consumption, improves the speed and resource utilization of convolutional computing, and supports real-time inference of large-scale neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988804A_ABST
    Figure CN119988804A_ABST
Patent Text Reader

Abstract

The invention provides a convolution calculation method and device based on brain-like chip hardware optimization, and belongs to the technical field of computers.The method comprises the steps that a delay layer and a convolution calculation layer are constructed, the convolution calculation layer comprises at least one convolution kernel, and the convolution kernels are mapped to a cross array architecture of a brain-like chip in a semi-folding mode; the number of the delay layers is determined based on the width of the convolution kernel; and performing alignment and convolution calculation on the received input feature map based on the delay layer and the convolution calculation layer, splitting the feature map output by the delay layer in the row / column dimension by the convolution calculation layer based on a convolution kernel, and calculating the split feature map in the column / row dimension by adopting a time division multiplexing mode to obtain an output feature map. According to the method, a semi-folding mapping mode is adopted, and by folding in the row / column dimension, required hardware resources are reduced, calculation is carried out in the column / row dimension, the calculation parallelism is ensured, the calculation speed is increased, and the resource occupancy rate and power consumption can be effectively reduced while the calculation efficiency is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a convolution calculation method and device based on brain-like chip hardware optimization. Background Art

[0002] Brain-like chips are the carriers of brain-like computing. They use a multi-core distributed parallel computing architecture and contain a large number of independently running processing cores. A single processing core has full operating capabilities and has an array of neurons and synapses inside. Different processing cores are connected and communicated through a specific hardware routing structure.

[0003] Convolution is one of the most commonly used operators in neural network models and is used in most neural network models. In order to improve the applicability of brain-like chips, it is necessary to implement efficient deployment and operation of convolution operators on them. The current means of deploying convolution to brain-like chips is to flatten the convolution kernel and determine the connection matrix between the input feature map and the output feature map based on the calculation principle of convolution. After being deployed to the brain-like chip, the input feature map and the output feature map will be mapped to one or more processing core neurons respectively, and the connection matrix will be mapped to the synaptic array between the two. The matrix represents the weight attribute. The convolution operation expressed in the above form is completely equivalent to the conventional form of convolution operation.

[0004] Since brain-like chips use an in-memory architecture, the SRAM storage space where the synaptic array is located needs to be written and determined before a computing task begins, and it cannot be read and written repeatedly to modify its data during the calculation. The existing fully expanded operators need to flatten the convolution kernel, and the size of the final expanded connection matrix is ​​proportional to the size of the input feature map and the output feature map. The fully expanded operator will cause huge resource overhead. The fully folded operator completely reuses the convolution kernel, and completes the convolution operation across multiple sliding windows in a cycle-by-cycle manner. The final calculation time is proportional to the size of the output feature map. The convolution calculation implemented in the fully folded form has a long calculation cycle.

[0005] The shortcomings of existing technologies limit the scale of deployable convolutions, resulting in the overall scale of neural networks that can be deployed on brain-like chips being limited, or causing the convolution calculations to run too long, resulting in reduced performance of reasoning tasks, hindering the widespread development and application of brain-like chips. Summary of the invention

[0006] The present invention provides a convolution calculation method and device based on brain-like chip hardware optimization, which can effectively reduce resource occupancy and power consumption while ensuring calculation efficiency.

[0007] The present invention provides a convolution calculation method based on brain-like chip hardware optimization, comprising: Constructing a delay layer and a convolution calculation layer, wherein the convolution calculation layer includes at least one convolution kernel, and the convolution kernel is mapped to the cross array architecture of the brain-like chip in a semi-folded form, and the number of the delay layers is determined based on the width of the convolution kernel; Based on the delay layer and the convolution calculation layer, the received input feature map is aligned and convolutionally calculated respectively. The convolution calculation layer is used to split the feature map output by the delay layer in the row dimension based on the convolution kernel, and calculate the split feature map in the column dimension by using time division multiplexing, or split the feature map output by the delay layer in the column dimension, and calculate the split feature map in the row dimension by using time division multiplexing, so as to obtain the output feature map.

[0008] As an embodiment, the step of constructing the convolutional computing layer includes: Determining the size of the output feature map according to the size of the input feature map and the properties of the convolution kernel; Determining a calculation period of the convolution calculation layer according to the height / width of the input feature map; According to the height / width of the output feature map, determining the number of column / row reuses of the convolution kernel to construct the convolution calculation layer; Correspondingly, the delay layer is used to cooperate with the input cache mechanism of the brain-like chip to perform delay alignment on the input data input to the convolution kernel in each calculation cycle.

[0009] As an embodiment, the delay layer is used to cooperate with the input cache mechanism of the brain-like chip to perform delay alignment on the input data input to the convolution kernel in each calculation cycle, including: Setting a delay parameter for each of the calculation cycles based on the delay layer; The input cache address of the brain-like chip is determined according to the delay parameter so that the convolution kernel can access the input data of the input cache address at the time corresponding to the delay parameter, thereby realizing delay alignment.

[0010] As an embodiment, the convolution kernel includes a convolution calculation kernel corresponding to each of the calculation cycles; the feature map output by the delay layer is split in the row dimension based on the convolution kernel, and the split feature map is calculated in the column dimension by time division multiplexing, or the feature map output by the delay layer is split in the column dimension, and the split feature map is calculated in the row dimension by time division multiplexing to obtain the output feature map, including: If the split feature map is calculated in a time division multiplexing manner in the column dimension, the input feature map is divided into at least one periodic input feature map according to the column dimension of the input feature map; or, if the split feature map is calculated in a time division multiplexing manner in the row dimension, the input feature map is divided into at least one periodic input feature map according to the row dimension of the input feature map; Perform convolution calculation according to the convolution calculation kernel and the periodic input feature map, if delay alignment is not achieved, use the data output by the convolution calculation kernel as invalid data, if delay alignment is achieved, use the data output by the convolution calculation kernel as valid data; The convolution calculation is repeated until all the valid data are obtained, and the output feature map is obtained according to the valid data.

[0011] As an embodiment, the number of the convolution kernels is multiple, and correspondingly, the output feature map is also used as the periodic input feature map required by the next convolution kernel.

[0012] As an embodiment, the output interval of the valid data corresponding to the convolution calculation layer is determined based on the sliding step size of the convolution kernel.

[0013] The present invention also provides a convolution computing device based on brain-like chip hardware optimization, comprising: A construction module, used to construct a delay layer and a convolution calculation layer, wherein the convolution calculation layer includes at least one convolution kernel, and the convolution kernel is mapped to the cross array architecture of the brain-like chip in a semi-folded form, and the number of the delay layers is determined based on the width of the convolution kernel; A convolution calculation module is used to align and perform convolution calculation on the received input feature maps based on the delay layer and the convolution calculation layer, respectively. The convolution calculation layer is used to split the feature map output by the delay layer in the row dimension based on the convolution kernel, and calculate the split feature map in the column dimension using a time division multiplexing method, or split the feature map output by the delay layer in the column dimension, and calculate the split feature map in the row dimension using a time division multiplexing method to obtain an output feature map.

[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a convolution calculation method based on brain-like chip hardware optimization as described in any one of the above is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the convolution calculation method based on brain-like chip hardware optimization as described in any one of the above is implemented.

[0016] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements a convolution calculation method based on brain-like chip hardware optimization as described in any one of the above.

[0017] The present invention provides a convolution calculation method and device based on hardware optimization of a brain-like chip, which constructs a delay layer and a convolution calculation layer, wherein the convolution calculation layer includes at least one convolution kernel, which is mapped to the cross array architecture of the brain-like chip in a semi-folded form, and the number of the delay layers is determined based on the width of the convolution kernel; the received input feature map is aligned and convolutionally calculated based on the delay layer and the convolution calculation layer, respectively, and the convolution calculation layer is used to split the feature map output by the delay layer in the row dimension based on the convolution kernel, and calculate the split feature map in the column dimension using a time-division multiplexing method, or split the feature map output by the delay layer in the column dimension, and calculate the split feature map in the row dimension using a time-division multiplexing method to obtain an output feature map. The present invention adopts a semi-folded mapping method, which reduces the required hardware resources by folding in the row dimension, and expands the calculation in the column dimension to ensure the parallelism of the calculation, improve the calculation speed, and can effectively reduce resource occupancy and power consumption while ensuring calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 It is a flowchart of a convolution calculation method based on brain-like chip hardware optimization provided by the present invention.

[0020] Figure 2 It is a schematic diagram of the input feature map, convolution kernel and output feature map provided by the present invention.

[0021] Figure 3 It is a structural schematic diagram of the half-folded convolution operator provided by the present invention.

[0022] Figure 4 It is a schematic diagram of the process of constructing a convolutional computing layer provided by the present invention.

[0023] Figure 5 It is a schematic diagram of data arrangement of the computing core input cache of the brain-like chip provided by the present invention.

[0024] Figure 6It is a schematic diagram of data arrangement when the data arranged in the input cache provided by the present invention exceeds the input cache depth of the computing core.

[0025] Figure 7 It is a schematic diagram of the convolution calculation layer provided by the present invention performing convolution calculation on the input feature map.

[0026] Figure 8 It is a schematic diagram of convolution calculation provided by the present invention when the convolution kernel sliding step size is 2.

[0027] Fig. 9 It is a schematic diagram of the connection structure between the half-folded convolution operator provided by the present invention and the subsequent fully connected layer.

[0028] Fig.10 yes Fig. 9 Schematic diagram of the data arrangement of the input cache of the fully connected computing layer.

[0029] Fig.11 It is a structural schematic diagram of a convolution computing device based on brain-like chip hardware optimization provided by the present invention.

[0030] Fig.12 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] It should be noted that all actions of acquiring signals, information or data in the present invention are performed in compliance with the corresponding data protection laws and policies of the location and with the authorization given by the owner of the corresponding device.

[0033] There are currently two methods to deploy convolution operators on brain-like chips based on a crossbar architecture, one is the fully expanded form, and the other is the fully folded form. Both methods have failed to find an ideal balance between computing speed and resource consumption, especially in application environments that have strict restrictions on resources and energy consumption, such as real-time visual processing tasks for drones.

[0034] The present invention provides a semi-folded expansion form to deploy the convolution operator on a brain-like chip. The semi-folded expansion form is applicable to convolution layers, pooling layers and fully connected layers. It can significantly reduce the resource consumption in the calculation process while ensuring a certain level of inference speed, thereby achieving a better balance between speed and resource consumption.

[0035] Figure 1 It is a flowchart of a convolution calculation method based on brain-like chip hardware optimization provided by the present invention, such as Figure 1 As shown, the present invention provides a convolution calculation method based on brain-like chip hardware optimization, including steps S100 to S200.

[0036] Step S100, construct a delay layer and a convolution calculation layer, the convolution calculation layer includes at least one convolution kernel, the convolution kernel is mapped to the cross array architecture of the brain-like chip in a semi-folded form, and the number of delay layers is determined based on the width of the convolution kernel. The semi-folded form includes folding the row dimension and expanding the column dimension and folding the column dimension and expanding the row dimension. The row dimension folding and column dimension expansion refers to dividing the input feature map into multiple periodic input feature maps according to the row dimension, and performing parallel calculations on the periodic input feature maps in the column dimension. The column dimension folding and row dimension expansion refers to dividing the input feature map into multiple periodic input feature maps according to the column dimension, and performing parallel calculations on the periodic input feature maps in the row dimension. The semi-folded form combines the advantages of the fully expanded form and the fully folded form.

[0037] Step S200, based on the delay layer and the convolution calculation layer, respectively align and convolutionally calculate the received input feature map, the convolution calculation layer is used to split the feature map output by the delay layer in the row dimension based on the convolution kernel, and calculate the split feature map in the column dimension using time division multiplexing, or split the feature map output by the delay layer in the column dimension, and calculate the split feature map in the row dimension using time division multiplexing to obtain the output feature map.

[0038] It should be noted that the input feature map can be a feature map obtained by extracting features from the input image or a feature map output by the previous convolutional calculation layer.

[0039] It can be understood that the present invention adopts a semi-folding mapping method, which reduces the required hardware resources by folding in the row / column dimension, expands the calculation in the column / row dimension, ensures the parallelism of the calculation, improves the calculation speed, and can effectively reduce resource occupancy and power consumption while ensuring calculation efficiency.

[0040] Figure 2 It is a schematic diagram of the input feature map, convolution kernel and output feature map provided by the present invention, such as Figure 2As shown, based on the above embodiment, as an optional embodiment, the steps of constructing the convolution calculation layer include the following steps.

[0041] Step S110, determining the size of the output feature map according to the size of the input feature map and the properties of the convolution kernel.

[0042] For a given size The input feature map of the convolution kernel size is , sliding step length , fill size , then the output feature map size is ,in: ; ; In the formula, the number of channels of the input feature map is ,high ,width , the number of output channels of the convolution kernel , Number of input channels ,high ,width , sliding step length With fill size , the number of channels of the output feature map ,high ,width .

[0043] like Figure 2 As shown in the figure, it is assumed that the input feature map size is 5×5, the number of channels is 1, the convolution kernel size is 3×3, the number of input and output channels are both 1, the sliding step is 1, the padding size is 0, the output feature map size is 3×3, and the number of channels is 1.

[0044] Step S120, determining the calculation period of the convolution calculation layer according to the height / width of the input feature map.

[0045] In the embodiment of the present invention, a semi-folded form of folding the row dimension and expanding the column dimension is taken as an example, and the calculation period of the convolution calculation layer is determined according to the width of the input feature map, that is, the input feature map is in the width Fold in dimension to form The feature map of size, that is, the input at each time step is A column of data of channels, the entire input feature map is passed The computation cycle of the convolutional layer is equal to the width of the input feature map.

[0046] Step S130, determining the number of column reuses of the convolution kernel according to the height / width of the output feature map to construct the convolution calculation layer.

[0047] Taking the semi-folded form of folding the row dimension and expanding the column dimension as an example, the number of column reuses of the convolution kernel is determined according to the width of the output feature map, and the width of the output feature map is used as the number of column reuses of the convolution kernel. The width of the output feature map is , then the number of column reuses of the convolution kernel is .For example Figure 2 In the example, the width of the output feature map is 3, then one column of the convolution kernel , and The number of reuses of is 3, and the number of reuses of the other two columns is also 3.

[0048] After determining the calculation cycle of the convolution calculation layer and the number of column reuses of the convolution kernel, a weight matrix is ​​formed and mapped to the cross array architecture of the brain-like chip. When the convolution calculation layer performs convolution calculations, it performs calculations based on the cross array architecture of the brain-like chip.

[0049] For two-dimensional images, row dimension folding and column dimension folding are equivalent. The calculation method of the calculation cycle and the number of reuse times corresponding to the semi-folded form of folding the column dimension and unfolding the row dimension can refer to the above embodiment and will not be repeated here.

[0050] It can be understood that the embodiment of the present invention provides a technical solution for constructing a convolutional computing layer, which greatly reduces the hardware resource occupation of the cross array by folding resources in the row dimension and expanding the calculation in the column dimension, or folding resources in the column dimension and expanding the calculation in the row dimension. By reusing computing resources, the waste of hardware resources caused by the reuse of a large amount of data is avoided, especially in large-scale networks. In addition, while maintaining a high degree of parallelism, the calculation cycle of the convolution operator is greatly shortened by means of pipeline operation and time division multiplexing technology. Compared with the mapping method of the full folding form, the calculation cycle is reduced by hundreds of times. The present invention not only provides support for real-time reasoning of large-scale neural networks, but also significantly improves the utilization rate of computing resources by optimizing the use of hardware resources in each computing cycle, and effectively avoids the waste of resources caused by hardware idleness.

[0051] Based on the above embodiments, as an optional embodiment, the present invention provides a method for deploying a convolution operator in a semi-folded form on a brain-like chip with a cross array architecture. In other embodiments, the convolution calculation can also be achieved by setting a delay layer on the input side of the convolution calculation layer to align the feature map and then perform calculations. The following is an example of an embodiment of setting a delay layer.

[0052] A delay layer is also provided on the input side of the convolution calculation layer. The number of the delay layers is determined based on the width of the convolution kernel. The delay layer is used to cooperate with the input cache mechanism of the brain-like chip to perform delay alignment on the input data input to the convolution kernel in each calculation cycle.

[0053] The present invention maps the semi-folded form to the convolution operator of the brain-like chip as the semi-folded convolution operator. Figure 3 : is a schematic diagram of the structure of the half-folded convolution operator provided by the present invention, such as Figure 3 As shown in the figure, the semi-folded convolution operator includes an input layer, multiple delay layers and a convolution calculation layer connected in sequence, and the number of delay layers corresponds to the width of the convolution kernel. The data of the input layer will be passed to multiple delay layers to achieve the effect of replication. The input data of the convolution calculation layer is composed of multiple output data from the delay layer. The delay layer uses delay neurons to arrange the data from the input layer with different offsets into the input buffer before the convolution calculation layer, thereby realizing the alignment of the overall input data of the convolution calculation layer, ensuring that the convolution kernel obtains the required input data in the correct calculation cycle.

[0054] Taking the semi-folded form of folding the row dimension and expanding the column dimension as an example, the flowchart of constructing the convolution calculation layer provided by the present invention is as follows: Figure 4 As shown, the following steps are included.

[0055] Determine the properties of the half-folded convolution operator: for a given size The input feature map of the convolution kernel size is , sliding step length , fill size , then the output feature map size is .

[0056] The width of each output channel to collapse dimension, input a column of input feature maps in each calculation cycle, The input is completed within calculation cycles, and the total calculation time is cycle.

[0057] structure Delay layers, each delay layer contains output neurons; the output delays of each layer are , ,……,1.

[0058] Construct a convolutional calculation layer, where each channel convolution kernel of the convolutional calculation layer appears Second-rate.

[0059] According to the above process, the convolutional computing layer needs to receive The input data is divided into columns, and only one column is input in each computing cycle, so the input data must be aligned. To this end, the present invention realizes delay alignment by constructing a delay layer and utilizing the input cache mechanism of the computing core of the brain-like chip. Each computing core of the brain-like chip is equipped with an input cache, and each cache address corresponds to the input of a computing cycle.

[0060] Optionally, the delay layer is used to cooperate with the input cache mechanism of the brain-like chip to perform delay alignment on the input data input to the convolution kernel in each calculation cycle, including the following steps.

[0061] Delay parameters are set for each of the calculation cycles based on the delay layer. A delay layer is provided with a delay parameter, and the delay parameters of each delay layer are respectively , , ..., 1. Among them, the delay parameter equal to 1 means that the output data of the delay layer is arranged at the 1-1=0 address of the convolution calculation layer input cache, and equal to 2 means that it is arranged at the 1st address, and so on. At this moment, the convolutional computing layer input buffer simultaneously obtains the 1st to Column input data.

[0062] The input buffer address of the brain-like chip is determined according to the delay parameter so that the convolution calculation layer can access the input data of the input buffer address at the time corresponding to the delay parameter to achieve delay alignment. , which means that the output data of this layer will be arranged in the first order of the input cache of the computing core where its target layer is located. Then, the computing core where the target layer is located will be Access the data at this address at all times. In layman's terms, assuming that the output of delay layer 1 is provided with calculation layer 2, then delay layer 1 sets the delay parameter , which affects the data output from delay layer 1 to computing layer 2 and places it at the address of the input cache of the computing core where layer 2 is located, thereby achieving data delay alignment.

[0063] Figure 5 Schematic diagram of data arrangement of the computing core input cache of the brain-like chip provided by the present invention, such as Figure 5 As shown, ~ They correspond to the 1st to 4th columns of the input feature map. The arrangement of each column of the input feature map in the input buffer is adjusted by the delay parameter of the delay neuron to achieve correct convolution calculation. The horizontal dimension is the depth dimension of the input buffer, and the vertical dimension is the width dimension of the input buffer. 0 indicates that this area is a 0-value matrix of the corresponding size. Since the input buffer has an upper depth limit, it is recorded as , the computing core reads data from index 0 of the input buffer after starting to work, and reads to index After processing the data, it reads again from index 0 in a loop.

[0064] Figure 6 Schematic diagram of data arrangement when the data arranged in the input cache provided by the present invention exceeds the input cache depth of the computing core, such as Figure 6 As shown in the figure, if the data arrangement exceeds the input buffer depth, the arrangement will be continued from the input buffer index 0. For example, when the input buffer depth is 3, the input data in the 4th column is arranged. way.

[0065] It can be understood that the present invention, by setting a delay layer, is conducive to delay alignment of the input data input to the convolution kernel in each calculation cycle, ensuring that the convolution calculation layer obtains the required input data in the correct calculation cycle and realizes correct convolution calculation.

[0066] Figure 7 It is a schematic diagram of the convolution calculation layer provided by the present invention performing convolution calculation on the input feature map, such as Figure 7 As shown, based on the above embodiments, as an optional embodiment, the convolution kernel includes a convolution calculation kernel corresponding to each of the calculation cycles; the feature map output by the delay layer is split in the row dimension based on the convolution kernel, and the split feature map is calculated in the column dimension using time division multiplexing, or the feature map output by the delay layer is split in the column dimension, and the split feature map is calculated in the row dimension using time division multiplexing to obtain the output feature map, including the following steps.

[0067] Step S210, if the split feature map is calculated in a time-division multiplexing manner in the column dimension, the input feature map is divided into at least one periodic input feature map according to the column dimension of the input feature map, or if the split feature map is calculated in a time-division multiplexing manner in the row dimension, the input feature map is divided into at least one periodic input feature map according to the row dimension of the input feature map.

[0068] Step S220, performing convolution calculation according to the convolution calculation kernel and the periodic input feature map, if delay alignment is not achieved, the data output by the convolution calculation kernel is used as invalid data, if delay alignment is achieved, the data output by the convolution calculation kernel is used as valid data.

[0069] Step S230, repeat the convolution calculation until all the valid data are obtained, and obtain the output feature map according to the valid data.

[0070] Taking the semi-folded form in which the row dimension is folded and the column dimension is expanded as an example, the convolution calculation kernel receives the input of the first cycle for calculation. The input cache of the first calculation cycle lacks the second and third columns of the input feature map required for the convolution calculation, and outputs a column of invalid data.

[0071] The convolution calculation kernel receives the input of the second cycle for calculation. The input cache of the second calculation cycle lacks the third column of the input feature map required for the convolution calculation, and outputs a column of invalid data.

[0072] The convolution calculation kernel receives the input of the third cycle for calculation. The input cache of the third calculation cycle aligns the correct input data required for the convolution calculation and outputs the first column of the output feature map.

[0073] The convolution calculation core continues to receive the input of the fourth cycle for calculation. The input cache of the fourth calculation cycle aligns the correct input data required for the convolution calculation and outputs the second column of the output feature map. The convolution calculation core receives the input of the fifth cycle for calculation. The input cache of the fifth calculation cycle aligns the correct input data required for the convolution calculation and outputs the third column of the output feature map. All valid data are output.

[0074] It is understandable that the calculation result of the convolution operator of the present invention includes both valid and invalid data. After all data calculations are completed, the valid data will be concatenated to obtain the final result, which is equal to the original convolution calculation result.

[0075] Based on the above embodiment, as an optional embodiment, the number of the convolution kernels is multiple, and correspondingly, the output feature map is also used as the periodic input feature map required by the next convolution kernel.

[0076] The output feature map obtained by convolution calculation in the form of step S210-step S230 can be used as the input data map of the subsequent convolution kernel using the same calculation method, because the output feature map has been divided into multiple cycles of data in the column / row dimension. Based on this feature, multiple convolution kernels using such a calculation method can be connected.

[0077] It can be understood that the present invention also uses the output feature map as the periodic input feature map required by the next convolution kernel, which reflects the scalability and versatility of the calculation method provided by the present invention.

[0078] Based on the above embodiment, as an optional embodiment, the output interval of the valid data corresponding to the convolution calculation layer is determined based on the sliding step size of the convolution kernel.

[0079] Convolution kernel sliding step parameter It will affect the delay parameters of the delay layer, the calculation period of the effective output and the arrangement of the weights. Every time the data passes through a half-folded convolution operator, the output interval of the effective output data is multiplied by the operator's value.

[0080] Figure 8 is a schematic diagram of convolution calculation provided by the present invention when the convolution kernel sliding step is 2, Figure 8 A first layer having Figure 2 The input feature map and convolution kernel of the same size, but the convolution kernel sliding step size At this time, the number of channels of the output feature map is 1, and the size becomes 2×2. The specific calculation logic is as follows: In the first and second calculation cycles, the input cache lacks the data required for the convolution calculation, and a column of invalid data is output. In the third calculation cycle, the input cache is aligned with the correct data required for the convolution calculation, and the first column of the feature map is output. In the fourth calculation cycle, the input cache is not aligned with the correct data required for the convolution calculation, and a column of invalid data is output. In the fifth calculation cycle, the input cache is aligned with the correct data required for the convolution calculation, and the second column of the feature map is output. It takes a total of 5 calculation cycles until all valid data is output, and the output interval of the valid output data is 2.

[0081] For those with A neural network with consecutive half-folded convolution operators, in the The output interval of the effective output data of the layer Sliding step length of the convolution kernel of this layer and The output interval of the effective output of the layer related: ; The first The output interval of the effective output data of the layer It can be expressed as: ; It can be deduced that The output interval of the effective output data of the layer For the front The product of the sliding step of the layer convolution kernel: ; For layer 1, the previous layer is external input and its sliding step size is , the output interval of effective output data =1. Therefore, the above formula can be simplified to: ; When the first and second layers of half-folded convolution operators When , the delay parameters of the delay layers in the first and second layers of semi-folded convolution operators are multiplied by 2 accordingly, that is, the output interval of the effective output of the first layer is 2, and the output interval of the effective output of the second layer is 2×2=4.

[0082] The calculation method of the effective output cycle of the convolution calculation layer is: calculate the first effective output cycle of the layer, and its value is equal to the product of the first effective output cycle of the previous layer plus the difference between the convolution kernel size of this layer minus 1 and the output interval of the effective output data of the previous layer. For the first layer, its previous layer represents external input data, so its first effective output cycle is 0. With this as the initial value, each subsequent effective output cycle of this layer is linearly related to the output interval of the effective output data of this layer.

[0083] For those with A neural network with consecutive half-folded convolution operators. For the first layer, the first valid output cycle of the previous layer . No. The convolution kernel size is , the output interval is , the output size is , the calculation of the first valid output cycle of this layer is as follows: ; in, For the The first valid output cycle of the layer half-folded convolution operator, For the The first valid output cycle of the layer half-folded convolution operator, For the The convolution kernel width of the layer half-folded convolution operator, For the The output interval of the effective output data of the layer half-folded convolution operator.

[0084] For Layer Valid output cycles , it can be expressed as: ; in, For the The first layer of the semi-folded convolution operator Valid output cycles, For the The width of the output feature map corresponding to the layer half-folded convolution operator.

[0085] After determining all the cycles where the valid output of the convolution calculation layer is located, the complete convolution calculation result can be obtained.

[0086] In addition to the half-folded convolution operator, the half-folded mapping method proposed in the present invention is also applicable to the average pooling operator and the maximum pooling operator. The structures of the average pooling operator and the maximum pooling operator are similar to those of the half-folded convolution operator. They are composed of a delay layer for data alignment and a pooling layer for pooling calculation. The weight matrices finally deployed are generated by expanding the pooling kernel in the same form. The convolution kernel sliding step size parameter is The effect on the output interval of the valid output is the same.

[0087] Fig. 9 A schematic diagram of the connection structure between the half-folded convolution operator and the subsequent fully connected layer is shown, where the output feature map size of the half-folded convolution operator is 5×5. For the fully connected layer after the half-folded operator, since the half-folded operator outputs feature maps in multiple calculation cycles, it is necessary to add a delay layer to align the output feature maps of multiple calculation cycles before calculation.

[0088] Fig.10 yes Fig. 9 The data arrangement diagram of the input cache of the fully connected computing layer in the figure further illustrates Fig. 9 The data arrangement of the input cache of the fully connected computing layer in the diagram. Fig. 9 When the first valid output cycle of the convolution calculation layer of the half-folded convolution operator in is 1, its valid output is output over a total of 6 calculation cycles, which are: invalid data, , ,……, ,in, Represents the output feature map According to the behavior of the five delay layers, the output data of the convolution calculation layer will be Fig.10 The arrangement is arranged in the input cache of the computing core where the fully connected computing layer is located. In the first five computing cycles of the convolution computing layer, the input data of the fully connected computing layer is not fully aligned, and the output data calculated at this time is invalid data; in the sixth computing cycle of the convolution computing layer, all the input data is obtained, and the output data calculated at this time is the correct fully connected computing result.

[0089] In summary, the present invention will greatly reduce resource usage. Compared with the traditional full expansion and full folding mapping methods, the half-folding mapping method greatly reduces the hardware resource usage of the cross array by folding resources in the row dimension and expanding calculations in the column dimension. Specifically, on the CIFAR-10 and CIFAR-100 datasets, under the VGG16 network, the number of computing cores required using the half-folding mapping method is reduced to 2.4%. This method reuses computing resources to avoid the waste of hardware resources caused by the cross array architecture arranging a large number of zero-value matrices due to the full expansion of the convolution kernel, and the effect is particularly significant in large-scale networks.

[0090] The present invention can be applied in application scenarios with limited hardware resources, such as embedded systems or edge computing devices. The semi-folded mapping method can still support the deployment of large-scale and complex neural networks, ensuring that the system can achieve a faster reasoning speed even under limited computing resources.

[0091] A convolution computing device based on brain-like chip hardware optimization provided by the present invention is described below. The convolution computing device based on brain-like chip hardware optimization described below and the convolution computing method based on brain-like chip hardware optimization described above can refer to each other.

[0092] Fig.11 is a structural schematic diagram of a convolution computing device based on brain-like chip hardware optimization provided by the present invention, such as Fig.11 As shown, the present invention also provides a convolution computing device based on brain-like chip hardware optimization, including the following modules.

[0093] A construction module 1110 is used to construct a delay layer and a convolution calculation layer, wherein the convolution calculation layer includes at least one convolution kernel, and the convolution kernel is mapped to the cross array architecture of the brain-like chip in a semi-folded form, and the number of the delay layers is determined based on the width of the convolution kernel; The convolution calculation module 1120 is used to align and perform convolution calculation on the received input feature map based on the delay layer and the convolution calculation layer, respectively. The convolution calculation layer is used to split the feature map output by the delay layer in the row dimension based on the convolution kernel, and calculate the split feature map in the column dimension using a time division multiplexing method, or split the feature map output by the delay layer in the column dimension, and calculate the split feature map in the row dimension using a time division multiplexing method to obtain an output feature map.

[0094] As an embodiment, the building module 1110 is further used for: Determining the size of the output feature map according to the size of the input feature map and the properties of the convolution kernel; Determining a calculation period of the convolution calculation layer according to the height / width of the input feature map; According to the height / width of the output feature map, determining the number of column / row reuses of the convolution kernel to construct the convolution calculation layer; Correspondingly, the delay layer is used to cooperate with the input cache mechanism of the brain-like chip to perform delay alignment on the input data input to the convolution kernel in each calculation cycle.

[0095] As an embodiment, the building module 1110 is further used for: Setting a delay parameter for each of the calculation cycles based on the delay layer; The input cache address of the brain-like chip is determined according to the delay parameter so that the convolution kernel can access the input data of the input cache address at the time corresponding to the delay parameter, thereby realizing delay alignment.

[0096] As an embodiment, the convolution kernel includes a convolution calculation kernel corresponding to each of the calculation cycles; the convolution calculation module 1120 is further used for: If the split feature map is calculated in a time division multiplexing manner in the column dimension, the input feature map is divided into at least one periodic input feature map according to the column dimension of the input feature map; or, if the split feature map is calculated in a time division multiplexing manner in the row dimension, the input feature map is divided into at least one periodic input feature map according to the row dimension of the input feature map; Perform convolution calculation according to the convolution calculation kernel and the periodic input feature map, if delay alignment is not achieved, use the data output by the convolution calculation kernel as invalid data, if delay alignment is achieved, use the data output by the convolution calculation kernel as valid data; The convolution calculation is repeated until all the valid data are obtained, and the output feature map is obtained according to the valid data.

[0097] As an embodiment, the number of the convolution kernels is multiple, and correspondingly, the output feature map is also used as the periodic input feature map required by the next convolution kernel.

[0098] As an embodiment, the output interval of the valid data corresponding to the convolution calculation layer is determined based on the sliding step size of the convolution kernel.

[0099] It should be noted that the convolution computing device based on brain-like chip hardware optimization provided by the present invention can execute a convolution computing method based on brain-like chip hardware optimization described in any of the above embodiments during specific operation, and has the technical effect corresponding to the method, which will not be elaborated in this embodiment.

[0100] Fig.12 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.12As shown, the electronic device may include: a processor (processor) 1210 , a communication interface (Communications Interface) 1220 , a memory (memory) 1230 and a communication bus 1240 , wherein the processor 1210 , the communication interface 1220 , and the memory 1230 communicate with each other via the communication bus 1240 . The processor 1210 can call the logic instructions in the memory 1230 to execute a convolution calculation method based on the hardware optimization of the brain-like chip, the method including: constructing a delay layer and a convolution calculation layer, the convolution calculation layer including at least one convolution kernel, the convolution kernel is mapped to the cross array architecture of the brain-like chip in a half-folded form, and the number of the delay layers is determined based on the width of the convolution kernel; based on the delay layer and the convolution calculation layer, respectively aligning and convolution calculation are performed on the received input feature map, the convolution calculation layer is used to split the feature map output by the delay layer in the row dimension based on the convolution kernel, and calculate the split feature map in the column dimension using time division multiplexing, or split the feature map output by the delay layer in the column dimension, and calculate the split feature map in the row dimension using time division multiplexing to obtain the output feature map.

[0101] In addition, the logic instructions in the above-mentioned memory 1230 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0102] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a convolution calculation method based on brain-like chip hardware optimization provided by the above methods, the method including: constructing a delay layer and a convolution calculation layer, the convolution calculation layer including at least one convolution kernel, the convolution kernel being mapped to the cross array architecture of the brain-like chip in a semi-folded form, and the number of the delay layers being determined based on the width of the convolution kernel; aligning and convolutionally calculating the received input feature maps based on the delay layer and the convolution calculation layer, respectively, the convolution calculation layer being used to split the feature map output by the delay layer in the row dimension based on the convolution kernel, and calculating the split feature map in the column dimension using time division multiplexing, or splitting the feature map output by the delay layer in the column dimension, and calculating the split feature map in the row dimension using time division multiplexing to obtain an output feature map.

[0103] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute a convolution calculation method based on brain-like chip hardware optimization provided by the above methods, the method comprising: constructing a delay layer and a convolution calculation layer, the convolution calculation layer comprising at least one convolution kernel, the convolution kernel being mapped to the cross array architecture of the brain-like chip in a semi-folded form, and the number of the delay layers being determined based on the width of the convolution kernel; aligning and convolutionally calculating the received input feature map based on the delay layer and the convolution calculation layer, respectively, the convolution calculation layer being used to split the feature map output by the delay layer in the row dimension based on the convolution kernel, and calculating the split feature map in the column dimension using time division multiplexing, or splitting the feature map output by the delay layer in the column dimension, and calculating the split feature map in the row dimension using time division multiplexing, to obtain an output feature map.

[0104] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0105] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A convolution calculation method based on brain-like chip hardware optimization, characterized in that: include: Constructing a delay layer and a convolution calculation layer, wherein the convolution calculation layer includes at least one convolution kernel, and the convolution kernel is mapped to the cross array architecture of the brain-like chip in a semi-folded form, and the number of the delay layers is determined based on the width of the convolution kernel; Based on the delay layer and the convolution calculation layer, the received input feature map is aligned and convolutionally calculated respectively. The convolution calculation layer is used to split the feature map output by the delay layer in the row dimension based on the convolution kernel, and calculate the split feature map in the column dimension by using time division multiplexing, or split the feature map output by the delay layer in the column dimension, and calculate the split feature map in the row dimension by using time division multiplexing, so as to obtain the output feature map.

2. The convolution calculation method based on brain-like chip hardware optimization according to claim 1 is characterized in that: The steps of constructing the convolutional computing layer include: Determining the size of the output feature map according to the size of the input feature map and the properties of the convolution kernel; Determining a calculation period of the convolution calculation layer according to the height / width of the input feature map; According to the height / width of the output feature map, determining the number of column / row reuses of the convolution kernel to construct the convolution calculation layer; Correspondingly, the delay layer is used to cooperate with the input cache mechanism of the brain-like chip to perform delay alignment on the input data input to the convolution kernel in each calculation cycle.

3. The convolution calculation method based on brain-like chip hardware optimization according to claim 2 is characterized in that: The delay layer is used to cooperate with the input cache mechanism of the brain-like chip to perform delay alignment on the input data input to the convolution kernel in each calculation cycle, including: Setting a delay parameter for each of the calculation cycles based on the delay layer; The input cache address of the brain-like chip is determined according to the delay parameter so that the convolution kernel can access the input data of the input cache address at the time corresponding to the delay parameter, thereby realizing delay alignment.

4. The convolution calculation method based on brain-like chip hardware optimization according to claim 3 is characterized in that: The convolution kernel includes a convolution calculation kernel corresponding to each of the calculation cycles; the feature map output by the delay layer is split in the row dimension based on the convolution kernel, and the split feature map is calculated in the column dimension by using a time division multiplexing method, or the feature map output by the delay layer is split in the column dimension, and the split feature map is calculated in the row dimension by using a time division multiplexing method to obtain an output feature map, including: If the split feature map is calculated in a time division multiplexing manner in the column dimension, the input feature map is divided into at least one periodic input feature map according to the column dimension of the input feature map; or, if the split feature map is calculated in a time division multiplexing manner in the row dimension, the input feature map is divided into at least one periodic input feature map according to the row dimension of the input feature map; Perform convolution calculation according to the convolution calculation kernel and the periodic input feature map, if delay alignment is not achieved, use the data output by the convolution calculation kernel as invalid data, if delay alignment is achieved, use the data output by the convolution calculation kernel as valid data; The convolution calculation is repeated until all the valid data are obtained, and the output feature map is obtained according to the valid data.

5. The convolution calculation method based on brain-like chip hardware optimization according to claim 4 is characterized in that: There are multiple convolution kernels, and correspondingly, the output feature map is also used as the periodic input feature map required by the next convolution kernel.

6. A convolution calculation method based on brain-like chip hardware optimization according to claim 4 or 5, characterized in that: The output interval of the valid data corresponding to the convolution calculation layer is determined based on the sliding step size of the convolution kernel.

7. A convolution computing device based on brain-like chip hardware optimization, characterized in that: include: A construction module, used to construct a delay layer and a convolution calculation layer, wherein the convolution calculation layer includes at least one convolution kernel, and the convolution kernel is mapped to the cross array architecture of the brain-like chip in a semi-folded form, and the number of the delay layers is determined based on the width of the convolution kernel; A convolution calculation module is used to align and perform convolution calculation on the received input feature map based on the delay layer and the convolution calculation layer respectively. The convolution calculation layer is used to split the feature map output by the delay layer in the row dimension based on the convolution kernel, and calculate the split feature map in the column dimension using a time division multiplexing method, or split the feature map output by the delay layer in the column dimension, and calculate the split feature map in the row dimension using a time division multiplexing method to obtain an output feature map.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements a convolution calculation method based on brain-like chip hardware optimization as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a convolution calculation method based on brain-like chip hardware optimization as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, a convolution calculation method based on brain-like chip hardware optimization as described in any one of claims 1 to 6 is implemented.