A full-add convolution method and device
By using the method of accumulating instead of multiplication in artificial intelligence chips, the binarization characteristics of pulse data of the input image are solved, and the problem of high consumption of computing resources and storage resources in the prior art is achieved, and the effect of reducing power consumption and reducing area is achieved.
Patent Information
- Application Number
- CN202010653816.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-07-08
AI Technical Summary
During the addition operation of the existing artificial intelligence chips, the input neuron activation values need to meet high-precision requirements, resulting in an increase in the consumption of computing resources and storage resources, and the power consumption and area also increase.
By utilizing the binarization characteristics of the pulse data of the input image, cumulative is used instead of multiplication and addition, the calculation amount and calculation time are reduced, the chip power consumption is reduced, and the chip area is reduced.
In the same convolution task, the time required to calculate using the accumulation operation is greatly reduced, which can save half of the time, reduce chip power consumption, and reduce chip area.
Smart Images

Figure CN111860778B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a full-add convolution method and device. Background Art
[0002] In the process of calculating the addition operation of the neural network, the input neuron activation value of the existing artificial intelligence chip must meet high-precision requirements. Multipliers and adders are used in the calculation process, resulting in the consumption of computing resources and storage resources, and increasing the power consumption and area of the artificial intelligence chip. Summary of the invention
[0003] To solve the above problems, the purpose of the present invention is to provide a full-add convolution method and device, which utilizes the binary characteristics of the pulse data of the input image, adopts accumulation instead of multiplication and addition, reduces the amount of calculation and calculation time, reduces chip power consumption, and reduces the chip area.
[0004] The present invention provides a full-add convolution method, the method comprising:
[0005] According to the pulse data of the input image, each target weight address for accumulation operation is determined; the accumulation operation is performed on the weights corresponding to the target weight addresses to obtain a full-add convolution value, and the full-add convolution value is determined as a feature vector of the input image, wherein the pulse data of the input image is pulse data of 0 and 1.
[0006] As a further improvement of the present invention, the target weight addresses for performing the accumulation operation are determined according to the pulse data of the input image, including:
[0007] Generating a weight address corresponding to each pulse data according to the pulse data of the input image;
[0008] The pulse data of the input image are traversed, and the weight address corresponding to the pulse data with a value of 1 is determined as the target weight address.
[0009] As a further improvement of the present invention, the target weight addresses for the accumulation operation are determined according to the pulse data of the input image, including:
[0010] Generate weight addresses corresponding to pulse data with a value of 1 according to the pulse data of the input image;
[0011] Each weight address corresponding to the pulse data having a value of 1 is determined as the target weight address.
[0012] As a further improvement of the present invention, the target weight addresses for the accumulation operation are determined according to the pulse data of the input image, including:
[0013] Generating a weight address corresponding to each pulse data according to the pulse data of the input image;
[0014] The weight address corresponding to each pulse data is determined as the target weight address.
[0015] As a further improvement of the present invention, performing an accumulation operation on the weights corresponding to the target weight addresses to obtain a full-add convolution value includes:
[0016] Traversing the impulse data of the input image;
[0017] When the pulse data is 0, 0 is determined as the weight to be accumulated;
[0018] When the pulse data is 1, the weight corresponding to the reference weight address is determined as the weight to be accumulated, wherein the reference weight address is the target weight address corresponding to the pulse data having a value of 1;
[0019] An accumulation operation is performed on each of the weights to be accumulated to obtain a full-add convolution value.
[0020] As a further improvement of the present invention, performing an accumulation operation on the weights corresponding to the target weight addresses to obtain a full-add convolution value includes:
[0021] Traversing the impulse data of the input image;
[0022] When the pulse data is 0, the accumulation operation is not performed;
[0023] When the pulse data is 1, the weight corresponding to the reference weight address is determined as the weight to be accumulated, wherein the reference weight address is the target weight address corresponding to the pulse data having a value of 1;
[0024] An accumulation operation is performed on each of the weights to be accumulated to obtain a full-add convolution value.
[0025] As a further improvement of the present invention, the method further includes: determining the membrane potential information at time step t based on the feature vector of the input image and the membrane potential information at time step t-1.
[0026] As a further improvement of the present invention, determining the membrane potential information at time step t according to the feature vector of the input image and the membrane potential information at time step t-1 includes:
[0027] The feature vector of the input image and the membrane potential information of the t-1 time step are added element by element to obtain the membrane potential information of the t time step.
[0028] The present invention also provides a full-add convolution device, the device comprising:
[0029] An address generation module, used to determine each target weight address for accumulation operation according to the pulse data of the input image;
[0030] An input data buffer module, used for inputting the pulse data of the input image;
[0031] A full-add convolution module is used to perform an accumulation operation on the weights corresponding to the target weight addresses to obtain a full-add convolution value, and determine the full-add convolution value as a feature vector of the input image;
[0032] The pulse data of the input image is pulse data of 0 and 1.
[0033] As a further improvement of the present invention, the address generation module is used to generate a weight address corresponding to each pulse data according to the pulse data of the input image; traverse the pulse data of the input image, and determine the weight address corresponding to the pulse data with a value of 1 as the target weight address.
[0034] As a further improvement of the present invention, the address generation module is used to generate each weight address corresponding to the pulse data with a value of 1 according to the pulse data of the input image; and determine each weight address corresponding to the pulse data with a value of 1 as the target weight address.
[0035] As a further improvement of the present invention, the address generation module is used to generate a weight address corresponding to each pulse data according to the pulse data of the input image; and determine the weight address corresponding to each pulse data as the target weight address.
[0036] As a further improvement of the present invention, the device further comprises:
[0037] A judgment module is used to judge the pulse data of the input image and determine the pulse data with a value of 1.
[0038] The address generation module is used to generate each weight address corresponding to the pulse data with a value of 1.
[0039] As a further improvement of the present invention, the address generation module is also used to judge the pulse data of the input image and determine the pulse data with a value of 1.
[0040] As a further improvement of the present invention, the full-add convolution module includes a multiplexer and an accumulator;
[0041] Traversing the impulse data of the input image;
[0042] When the pulse data is 0, the multiplexer determines 0 as the weight to be accumulated and outputs it;
[0043] When the pulse data is 1, the multiplexer determines the weight corresponding to the reference weight address as the weight to be accumulated and outputs it, wherein the reference weight address is the target weight address corresponding to the pulse data with a value of 1;
[0044] The accumulator performs an accumulation operation on each of the weights to be accumulated output by the multiplexer to obtain a full-add convolution value.
[0045] As a further improvement of the present invention, the full-add convolution module is an enabled accumulator;
[0046] Traversing the impulse data of the input image;
[0047] When the pulse data is 0, the enable accumulator is enabled to 0 and no accumulation operation is performed;
[0048] When the pulse data is 1, the enable accumulator is enabled to 1, and the weight corresponding to the reference weight address is determined as the weight to be accumulated, wherein the reference weight address is the target weight address corresponding to the pulse data with a value of 1;
[0049] The enabling accumulator performs an accumulation operation on each of the weights to be accumulated to obtain a full-add convolution value.
[0050] The present invention also provides an electronic device, comprising a memory and a processor, characterized in that the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the full-add convolution method.
[0051] The present invention also provides a computer-readable storage medium on which a computer program is stored, characterized in that the computer program is executed by a processor to implement the full-add convolution method.
[0052] The beneficial effects of the present invention are: utilizing the binary characteristics of pulse data of the input image, adopting accumulation instead of multiplication and addition, reducing the amount of calculation and calculation time, reducing chip power consumption, and reducing chip area. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor.
[0054] Figure 1 A schematic diagram of a flow chart of a full-add convolution method according to an exemplary embodiment of the present disclosure;
[0055] Figure 2 It is a schematic diagram of a multiplication-addition convolution device in the prior art;
[0056] Figure 3 A schematic diagram of a full-add convolution device according to an exemplary embodiment of the present disclosure;
[0057] Figure 4 A schematic diagram of a full-add convolution device according to another exemplary embodiment of the present disclosure;
[0058] Figure 5 It is a flowchart of a zero-jumping operation according to an exemplary embodiment of the present disclosure;
[0059] Figure 6 This is a calculation example of a zero-jumping operation according to an exemplary embodiment of the present disclosure;
[0060] Figure 7 A schematic diagram of a pulse full-add convolutional layer according to an exemplary embodiment of the present disclosure;
[0061] Figure 8 The figure is a schematic diagram of a pulse full-convolution network according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0063] It should be noted that if the embodiments of the present disclosure involve directional indications (such as up, down, left, right, front, back, etc.), such directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0064] In addition, in the description of the present disclosure, the terms used are only for illustrative purposes and are not intended to limit the scope of the present disclosure. The terms "include" and / or "comprise" are used to specify the existence of elements, steps, operations and / or components, but do not exclude the existence or addition of one or more other elements, steps, operations and / or components. The terms "first", "second" and the like may be used to describe various elements, do not represent the order, and do not limit these elements. In addition, in the description of the present disclosure, unless otherwise specified, the meaning of "multiple" is two and more than two. These terms are only used to distinguish one element from another element. In conjunction with the following drawings, these and / or other aspects become apparent, and it is easier for those of ordinary skill in the art to understand the description of the embodiments of the present disclosure. The accompanying drawings are used to describe the embodiments of the present disclosure for illustrative purposes only. Those skilled in the art will easily recognize from the following description that, without departing from the principles of the present disclosure, alternative embodiments of the structures and methods shown in the present disclosure can be adopted.
[0065] A full-add convolution method described in the embodiment of the present disclosure is as follows: Figure 1 As shown, according to the pulse data of the input image, the target weight addresses for the accumulation operation are determined;
[0066] Performing an accumulation operation on the weights corresponding to the target weight addresses to obtain a full-add convolution value, and determining the full-add convolution value as a feature vector of the input image;
[0067] The pulse data of the input image is pulse data of 0 and 1.
[0068] like Figure 2 As shown, in the prior art, artificial intelligence chips based on pulse neural networks use multipliers and adders in the calculation process. The address generation module (AGU) generates each weight address, and the input data buffer module (IN-Buffer) generates input data corresponding to each weight address. The multiplier performs corresponding element multiplication operations on the input X and the weight W, and the output is added to the result generated by the multiplier of the previous address to finally obtain the convolution value. However, the input of the pulse neural network is a binary input (0 means no pulse, 1 means pulse), that is, the data output by the IN-Buffer has only two values 0 and 1. Therefore, in the convolution process, the multiplication operation is a redundant operation, which will increase the computing resources, increase the power consumption of the chip and the area of the chip.
[0069] The present invention utilizes the binary characteristics of the pulse data of the input image, and adopts the accumulation operation with only full addition function to replace the addition operation of the multiplication and addition function in the prior art. Since the pulse data of the input image is a binary input, this characteristic is utilized to determine the target weight addresses that need to be accumulated according to the pulse data of the input image, and the weights corresponding to each target weight address are accumulated. In the same convolution task, the time required for calculation using the accumulation operation is greatly reduced compared with the traditional multiplication and addition operation, which can save half the time and reduce the chip power consumption. For example, the multiplication and addition of fp32 precision consumes 3.7+0.9=4.6 (pJ) energy, while the addition of fp32 precision only consumes 0.9 (pJ) energy. It can be seen that the use of accumulation instead of multiplication and addition can significantly reduce the power consumption of the convolution process. After removing the multiplier, the chip can save computing space, and the chip area will be greatly reduced, which is conducive to improving the integration of the chip.
[0070] The method described in the present disclosure can be used in pulse neural networks, and can also be used in a fusion network of a pulse neural network and an artificial neural network, etc. The method described in the present disclosure does not impose any specific limitation on the network used.
[0071] In an optional implementation, determining each target weight address for performing an accumulation operation according to the pulse data of the input image includes:
[0072] Generating a weight address corresponding to each pulse data according to the pulse data of the input image;
[0073] The pulse data of the input image are traversed, and the weight address corresponding to the pulse data with a value of 1 is determined as the target weight address.
[0074] In an optional implementation, determining each target weight address for the accumulation operation according to the pulse data of the input image includes:
[0075] Generate weight addresses corresponding to pulse data with a value of 1 according to the pulse data of the input image;
[0076] Each weight address corresponding to the pulse data having a value of 1 is determined as the target weight address.
[0077] In an optional implementation, determining each target weight address for the accumulation operation according to the pulse data of the input image includes:
[0078] Generating a weight address corresponding to each pulse data according to the pulse data of the input image;
[0079] The weight address corresponding to each pulse data is determined as the target weight address.
[0080] The method disclosed in the present invention, for example, can first generate weight addresses corresponding to all pulse data, and then determine the weight addresses corresponding to the pulse data with a value of 1 as the target weight addresses for the accumulation operation. For example, the weight addresses corresponding to all pulse data with a value of 1 can be directly generated, and these weight addresses can be determined as the target weight addresses for the accumulation operation. For example, the weight addresses corresponding to all pulse data can be generated, and these weight addresses can be determined as the target weight addresses for the accumulation operation.
[0081] In an optional implementation, performing an accumulation operation on the weights corresponding to the target weight addresses to obtain a full-add convolution value includes:
[0082] Traversing the impulse data of the input image;
[0083] When the pulse data is 0, 0 is determined as the weight to be accumulated;
[0084] When the pulse data is 1, the weight corresponding to the reference weight address is determined as the weight to be accumulated, wherein the reference weight address is the target weight address corresponding to the pulse data having a value of 1;
[0085] An accumulation operation is performed on each of the weights to be accumulated to obtain a full-add convolution value.
[0086] For example, let X be the pulse data of the input image, W be the weight of the convolution, and traverse the pulse data of the input image. For each pulse data X in X i have:
[0087] If the pulse data X i =0, then 0 is determined as the weight to be accumulated, 0 is output, S=S+0, and i=i+1 is executed to enter the next pulse data;
[0088] If the pulse data X i =1, then W i Determine as the weight to be accumulated, and W i Put it into the accumulator, S = S + W i , and execute i=i+1 to enter the next pulse data;
[0089] The output full convolution value S is used as the feature vector of the input image. After passing through the dynamic equation of the pulse neural network, a new binary input is generated, which can be used as the input of the next layer in the pulse neural network.
[0090] In an optional implementation, performing an accumulation operation on the weights corresponding to the target weight addresses to obtain a full-add convolution value includes:
[0091] Traversing the impulse data of the input image;
[0092] When the pulse data is 0, the accumulation operation is not performed;
[0093] When the pulse data is 1, the weight corresponding to the reference weight address is determined as the weight to be accumulated, wherein the reference weight address is the target weight address corresponding to the pulse data having a value of 1;
[0094] An accumulation operation is performed on each of the weights to be accumulated to obtain a full-add convolution value.
[0095] For example, let X be the pulse data of the input image, W be the weight of the convolution, and traverse the pulse data of the input image. For each pulse data X in X i have:
[0096] If the pulse data X i =0, then the corresponding weight W i Do not perform the accumulation operation, directly execute i=i+1 and enter the next pulse data;
[0097] If the pulse data X i =1, then the corresponding weight W i Determine as the weight to be accumulated, and W i Put it into the enable accumulator, S=S+W i , and execute i=i+1 to enter the next pulse data;
[0098] The output full convolution value S is used as the feature vector of the input image. After passing through the dynamic equation of the pulse neural network, a new binary input is generated, which can be used as the input of the next layer in the pulse neural network.
[0099] This embodiment can be understood as using a zero-skipping operation in the calculation process, that is, when the pulse data of the input image is a "0" value, the weight address corresponding to the "0" value is skipped, and the next pulse data is directly calculated. The weight W is a dense matrix, and the input X is a sparse matrix. Due to the sparsity of X, only a very small number of rows of W need to be integrated, so the valid data of W can be indexed by the weight address index. The weight address index can be calculated by using the value 1 in the pulse data X of the input image as a flag, thereby realizing direct calculation by skipping "0". Using the zero-skipping operation in the convolution process can further improve the calculation efficiency and reduce power consumption.
[0100] In an optional embodiment, according to the feature vector Add_Conv(X, W) of the input image and the membrane potential information U at the t-1 time step t-1 , determine the membrane potential information U at time step t t .
[0101] In an optional implementation, determining the membrane potential information at time step t according to the feature vector Add_Conv(X, W) of the input image and the membrane potential information at time step t-1 includes:
[0102] Add_Conv(X, W) to the feature vector of the input image and the membrane potential information U at the t-1 time step t-1 Add element by element to get the membrane potential information U at the time step t t , where U t =Add_Conv(X,W)+U t-1 .
[0103] A full-add convolution device according to an embodiment of the present disclosure includes:
[0104] An address generation module, used to determine each target weight address for accumulation operation according to the pulse data of the input image;
[0105] An input data buffer module, used for inputting the pulse data of the input image;
[0106] A full-add convolution module is used to perform an accumulation operation on the weights corresponding to the target weight addresses to obtain a full-add convolution value, and determine the full-add convolution value as a feature vector of the input image;
[0107] The pulse data of the input image is pulse data of 0 and 1.
[0108] As mentioned above, in the convolution process, the prior art performs redundant multiplication operations, which will increase the computing resources, power consumption and area of the chip. The full addition convolution device disclosed in the present invention utilizes the binary characteristics of the pulse data of the input image, and adopts the accumulation operation with only the full addition function to replace the addition operation of the multiplication and addition function in the prior art. Since the pulse data of the input image is a binary input, this characteristic is utilized to determine the target weight addresses that need to be accumulated according to the pulse data of the input image, and the weights corresponding to each target weight address are accumulated. In the same convolution task, the time required for calculation using the accumulation operation is greatly reduced compared with the traditional multiplication and addition operation, which can save half the time and reduce the chip power consumption. For example, the fp32 precision multiplication and addition requires 3.7+0.9=4.6 (pJ) energy, while the fp32 precision addition only consumes 0.9 (pJ) energy. It can be seen that the use of accumulation instead of multiplication and addition can significantly reduce the power consumption of the convolution process. After removing the multiplier, the chip can save computing space, the chip area will be greatly reduced, and it is beneficial to improve the integration of the chip.
[0109] The device described in the present disclosure can be used in a pulse neural network, and can also be used in a fusion network of a pulse neural network and an artificial neural network, etc. The device described in the present disclosure does not impose any specific restrictions on the network used.
[0110] In an optional embodiment, the address generation module is used to generate a weight address corresponding to each pulse data according to the pulse data of the input image; traverse the pulse data of the input image, and determine the weight address corresponding to the pulse data with a value of 1 as the target weight address.
[0111] In an optional embodiment, the address generation module is used to generate each weight address corresponding to the pulse data with a value of 1 based on the pulse data of the input image; and determine each weight address corresponding to the pulse data with a value of 1 as the target weight address.
[0112] In an optional implementation, the address generation module is used to generate a weight address corresponding to each pulse data according to the pulse data of the input image; and determine the weight address corresponding to each pulse data as the target weight address.
[0113] The address generation module described in the present disclosure can, for example, first generate weight addresses corresponding to all pulse data, and then determine the weight addresses corresponding to the pulse data with a value of 1 as the target weight addresses for the accumulation operation. For example, the weight addresses corresponding to all pulse data with a value of 1 can be directly generated, and these weight addresses can be determined as the target weight addresses for the accumulation operation. For example, the weight addresses corresponding to all pulse data can be generated, and these weight addresses can be determined as the target weight addresses for the accumulation operation.
[0114] In an optional embodiment, the device further comprises:
[0115] A judgment module is used to judge the pulse data of the input image and determine the pulse data with a value of 1.
[0116] The address generation module is used to generate each weight address corresponding to the pulse data with a value of 1.
[0117] In an optional implementation, the address generation module is further used to judge the pulse data of the input image and determine the pulse data with a value of 1.
[0118] The judgment module disclosed in the present invention may be provided separately, or the function of the judgment module may be realized by an address generation module.
[0119] In an optional implementation, the full-add convolution module includes a multiplexer and an accumulator;
[0120] Traversing the impulse data of the input image;
[0121] When the pulse data is 0, the multiplexer determines 0 as the weight to be accumulated and outputs it;
[0122] When the pulse data is 1, the multiplexer determines the weight corresponding to the reference weight address as the weight to be accumulated and outputs it, wherein the reference weight address is the target weight address corresponding to the pulse data with a value of 1;
[0123] The accumulator performs an accumulation operation on each of the weights to be accumulated output by the multiplexer to obtain a full-add convolution value.
[0124] like Figure 3 As shown, the address generation module is, for example, an AGU module, which is responsible for determining the target weight addresses for the accumulation operation according to the pulse data of the input image; the input data buffer module is, for example, an IN-buffer module, which is responsible for inputting the pulse data of the input image, and W is the convolution weight.
[0125] For example, let X be the pulse data of the input image of the spiking neural network, W be the weight of the convolution, and traverse the pulse data of the input image. For each pulse data X in X i have:
[0126] If the pulse data X i =0, the multiplexer determines 0 as the weight to be accumulated, the multiplexer outputs 0, S=S+0, and executes i=i+1 to enter the next pulse data;
[0127] If the pulse data X i =1, the multiplexer will W i Determine as the weight to be accumulated, and W i Put it into the accumulator, and the accumulator performs the accumulation operation S=S+W i , the multiplexer executes i=i+1 and enters the next pulse data;
[0128] The output full convolution value S is used as the feature vector of the input image. After passing through the dynamic equation of the pulse neural network, a new binary input is generated, which can be used as the input of the next layer in the pulse neural network.
[0129] The device described in this embodiment replaces the traditional multiplier with a multiplexer. When the pulse data corresponding to the target weight address generated by the address generation module (AGU) in the input data buffer module (IN-Buffer) is 0, the multiplexer (Sel) outputs 0. When the pulse data corresponding to the target weight address generated by the address generation module (AGU) in the input data buffer module (IN-Buffer) is 1, the multiplexer outputs the weight in W. After the multiplexer is executed, the weight output by the multiplexer is added to the calculation result of the previous target weight address calculated by the accumulator, and the accumulation operation is finally completed. In the same convolution task, the time required for calculation using the accumulation operation is greatly reduced compared to the traditional multiplication and addition operation, which can save half the time, reduce chip power consumption, and reduce chip area.
[0130] In an optional implementation, the full-add convolution module is an enabled accumulator;
[0131] Traversing the impulse data of the input image;
[0132] When the pulse data is 0, the enable accumulator is enabled to 0 and no accumulation operation is performed;
[0133] When the pulse data is 1, the enable accumulator is enabled to 1, and the weight corresponding to the reference weight address is determined as the weight to be accumulated, wherein the reference weight address is the target weight address corresponding to the pulse data with a value of 1;
[0134] The enabling accumulator performs an accumulation operation on each of the weights to be accumulated to obtain a full-add convolution value.
[0135] like Figure 4 As shown, the address generation module is, for example, an AGU module, which is responsible for determining the target weight addresses for the accumulation operation according to the pulse data of the input image; the input data buffer module is, for example, an IN-buffer module, which is responsible for inputting the pulse data of the input image, and W is the convolution weight.
[0136] For example, let X be the pulse data of the input image, W be the weight of the convolution, and traverse the pulse data of the input image. For each pulse X in X i have:
[0137] If the pulse data X i = 0, the accumulator is enabled to determine the corresponding weight W i Do not perform the accumulation operation, enable the accumulator to 0, and directly execute i=i+1 to determine the weight corresponding to the next pulse data;
[0138] If the pulse data X i=1, then the accumulator is enabled to 1, and the corresponding weight W is determined i is the weight to be accumulated, and W i Put it into the enabled accumulator, and enable the accumulator to perform the accumulation operation S=S+W i , and execute i=i+1 to determine the weight corresponding to the next pulse data;
[0139] The output full convolution value S is used as the feature vector of the input image. After passing through the dynamic equation of the pulse neural network, a new binary input is generated, which can be used as the input of the next layer in the pulse neural network.
[0140] The full-add convolution device described in this embodiment replaces the multiplier and the adder with an enable accumulator, such as Figure 5 As shown, the accumulator is enabled to use a zero-skipping operation during the calculation process, that is, when the pulse data of the input image is a "0" value, the weight address corresponding to the "0" value is skipped to calculate the next pulse data. The weight W is a dense matrix, and the input X is a sparse matrix. Due to the sparsity of X, only a very small number of rows of W need to be integrated, so the valid data can be indexed by the weight address index. The weight address index can be calculated by using the value 1 in the pulse data X of the input image as a flag, thereby realizing direct calculation by skipping "0". The full-add convolution device can further improve the calculation efficiency, reduce chip power consumption, and reduce the chip area by performing a zero-skipping operation during the convolution process.
[0141] For example, Figure 6 As shown, when the pulse data X of the input image in the IN-Buffer module corresponding to the target weight addresses generated by the AGU module are 0, 1, 1, 0, ..., 1 respectively, the weights corresponding to the target weight addresses are W00, W001, W010, W011, ..., W311 respectively. When the pulse data is 0, the enable accumulator is enabled to 0, and the weight W corresponding to the address does not perform the accumulation operation. When the pulse data is 1, the enable accumulator is enabled to 1, and the weight corresponding to the weight address performs the accumulation operation, and the full convolution value is: W001+W010+,,,+W311.
[0142] The full-add convolution device described in the embodiment of the present disclosure obtains a full-add convolution value Add_Conv(X, W), i.e., a feature vector of the input image, according to the pulse data of the input image, where X represents the input sparse matrix and W represents the weight;
[0143] The membrane potential module converts the membrane potential value U at time step t-1 into t-1 Add the feature vector Add_Conv(X, W) of the input image element by element to get the membrane potential value U at time step t t , U t =Add_Conv(X,W)+Ut-1 ;
[0144] The issuing module determines the output value of time step t according to the membrane potential value of time step t as the time series convolution vector, and performs time step and enable updates on the membrane potential module.
[0145] The full convolution device may, for example, use the address generation module, input data buffer module, and full convolution module described in the above embodiment. The full convolution module may, for example, use the multiplexer and accumulator described in the above embodiment, or may use the enable accumulator described in the above embodiment, which will not be described in detail here.
[0146] For example, Figure 7 As shown, the full-add zero-jumping convolution module in the figure can be understood as an implementation method in which the full-add convolution module uses a zero-jumping operation, that is, the pulse full-add convolution layer uses an enable accumulator. Since the zero-jumping operation is used in the convolution process, the pulse data of the input image passes through the full-add zero-jumping convolution module to obtain a full-add convolution value. The full-add convolution value and the membrane potential information of the previous time step (t-1 time step) in the membrane potential module are added element by element to obtain the membrane potential information of the current time step (t time step). The issuing module processes the membrane potential information of the current time step, obtains the output vector as the time series convolution vector, and updates the time step update enable signal to update the membrane potential module. The pulse full-add convolution layer is a convolution process with time series information. The time series convolution vector contains the time dimension. In specific applications, the pulse full-add convolution layer can connect the time series information between feature maps. At the same time, the zeroing operation also reduces the amount of calculation, saves calculation time, reduces chip power consumption, and reduces chip area.
[0147] The pulse full-add convolution network described in the embodiment of the present disclosure includes at least one pulse full-add convolution layer, and the pulse data of the input image is subjected to at least one time series convolution process through at least one pulse full-add convolution layer to obtain a time series convolution vector. The pulse full-add convolution layer is as described in the above embodiment and will not be described in detail here.
[0148] In an optional implementation, performing at least one temporal convolution process on the pulse data of the input image through at least one pulse full-add convolution layer to obtain a temporal convolution vector includes:
[0149] Performing a first time-series convolution process on the pulse data of the input image through a first pulse full-add convolution layer to obtain a first vector;
[0150] When performing a time series convolution process, the first vector is determined as a time series convolution vector;
[0151] When performing n-time convolution processing, the first vector is pooled to obtain a first intermediate vector, and the first intermediate vector is subjected to a second time convolution processing through a second pulse full-add convolution layer to obtain a second vector, and the second vector is pooled to obtain a second intermediate vector, and so on, and the nth vector is determined as a time convolution vector, where n is an integer greater than 1.
[0152] The process of performing time series convolution processing on the pulse data of the input image is as described in the above embodiment and will not be described in detail here.
[0153] For example, Figure 8 As shown, the pulse full-add convolution network includes two pulse full-add convolution layers. The first pulse full-add convolution layer performs a first time series convolution processing on the pulse data of the input image to obtain a first vector, and the first vector is pooled to obtain a first intermediate vector. The second pulse full-add convolution layer performs a second time series convolution processing on the first intermediate vector to obtain a second vector. The second vector is determined as a time series convolution vector as the input of the next layer.
[0154] An artificial intelligence chip described in an embodiment of the present disclosure includes the pulse full-addition convolution device. The pulse full-addition convolution device is as described in the previous embodiment and will not be described in detail here. During the calculation process, the chip removes the redundant multiplication operations in the existing convolution, greatly reducing the amount of calculation and reducing the power consumption of the chip. For example, fp32 precision multiplication and addition requires 3.7+0.9=4.6 (pJ) energy, while fp32 precision addition only requires 0.9 (pJ) energy. It can be seen that using addition instead of multiplication and addition can significantly reduce the power consumption of the convolution process. After removing the multiplier, the chip can save computing space, and the chip area will be greatly reduced, which is conducive to improving the integration of the chip.
[0155] The present disclosure also relates to an electronic device, including a server, a terminal, etc. The electronic device includes: at least one processor; a memory connected to the at least one processor; and a communication component connected to the storage medium, the communication component receiving and sending data under the control of the processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to implement the full-add convolution method in the above embodiment.
[0156] In an optional embodiment, the memory, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The processor executes various functional applications and data processing of the device by running the non-volatile software programs, instructions and modules stored in the memory, that is, implementing the full-add convolution method.
[0157] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store a list of options, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to an external device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0158] One or more modules are stored in the memory, and when executed by one or more processors, the full-add convolution method in any of the above method embodiments is executed.
[0159] The above-mentioned product can execute the full-add convolution method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the full-add convolution method provided in the embodiment of the present application.
[0160] The present disclosure also relates to a computer-readable storage medium for storing a computer-readable program, wherein the computer-readable program is used for a computer to execute part or all of the above-mentioned full-add convolution method embodiments.
[0161] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0162] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present disclosure can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.
[0163] In addition, it will be understood by those skilled in the art that although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present disclosure and form different embodiments. For example, in the claims, any one of the claimed embodiments may be used in any combination.
[0164] It will be appreciated by those skilled in the art that, although the present disclosure has been described with reference to exemplary embodiments, various changes may be made and elements thereof may be replaced with equivalents without departing from the scope of the present disclosure. In addition, many modifications may be made to adapt specific circumstances or materials to the teachings of the present disclosure without departing from the substantive scope of the present disclosure. Therefore, the present disclosure is not limited to the specific embodiments disclosed, but the present disclosure will include all embodiments falling within the scope of the appended claims.
Claims
1. A full-add convolution method, characterized in that: The method comprises: According to the pulse data of the input image, each target weight address for the accumulation operation is determined, wherein when the pulse data of the input image is 0, the weight address corresponding to the 0 value is skipped and the calculation of the next pulse data is performed, and the pulse data of the input image is the pulse data generated by the pulse neural network; An accumulation operation is performed on the weights corresponding to the target weight addresses to obtain a full-add convolution value, and the full-add convolution value is determined as a feature vector of the input image, wherein the pulse data of the input image is pulse data of 0 and 1.
2. The method of claim 1, wherein: Determining each target weight address for performing an accumulation operation according to the pulse data of the input image includes: Generating a weight address corresponding to each pulse data according to the pulse data of the input image; The pulse data of the input image are traversed, and the weight address corresponding to the pulse data with a value of 1 is determined as the target weight address.
3. The method of claim 1, wherein: According to the pulse data of the input image, the target weight addresses for the accumulation operation are determined, including: Generate weight addresses corresponding to pulse data with a value of 1 according to the pulse data of the input image; Each weight address corresponding to the pulse data having a value of 1 is determined as the target weight address.
4. The method according to claim 1, wherein determining each target weight address for performing an accumulation operation according to the pulse data of the input image comprises: Generating a weight address corresponding to each pulse data according to the pulse data of the input image; The weight address corresponding to each pulse data is determined as the target weight address.
5. The method of claim 4, wherein: Performing an accumulation operation on the weights corresponding to the target weight addresses to obtain a full-add convolution value includes: Traversing the impulse data of the input image; When the pulse data is 0, 0 is determined as the weight to be accumulated; When the pulse data is 1, the weight corresponding to the reference weight address is determined as the weight to be accumulated, wherein the reference weight address is the target weight address corresponding to the pulse data having a value of 1; An accumulation operation is performed on each of the weights to be accumulated to obtain a full-add convolution value.
6. The method of claim 4, wherein: Performing an accumulation operation on the weights corresponding to the target weight addresses to obtain a full-add convolution value includes: Traversing the impulse data of the input image; When the pulse data is 0, the accumulation operation is not performed; When the pulse data is 1, the weight corresponding to the reference weight address is determined as the weight to be accumulated, wherein the reference weight address is the target weight address corresponding to the pulse data having a value of 1; An accumulation operation is performed on each of the weights to be accumulated to obtain a full-add convolution value.
7. The method according to any one of claims 1 to 6, further comprising: The membrane potential information at time step t is determined based on the feature vector of the input image and the membrane potential information at time step t-1.
8. The method of claim 7, wherein determining the membrane potential information at time step t based on the feature vector of the input image and the membrane potential information at time step t-1 comprises: The feature vector of the input image and the membrane potential information of the t-1 time step are added element by element to obtain the membrane potential information of the t time step.
9. A full-add convolution device, characterized in that: The device comprises: An address generation module, used to determine each target weight address for accumulation operation according to the pulse data of the input image, wherein when the pulse data of the input image is 0, the weight address corresponding to the 0 value is skipped and the calculation of the next pulse data is performed, and the pulse data of the input image is the pulse data generated by the pulse neural network; An input data buffer module, used for inputting the pulse data of the input image; A full-add convolution module is used to perform an accumulation operation on the weights corresponding to the target weight addresses to obtain a full-add convolution value, and determine the full-add convolution value as a feature vector of the input image; The pulse data of the input image is pulse data of 0 and 1.
10. The device of claim 9, wherein: The address generation module is used to generate a weight address corresponding to each pulse data according to the pulse data of the input image; traverse the pulse data of the input image, and determine the weight address corresponding to the pulse data with a value of 1 as the target weight address.
11. The device of claim 9, wherein: The address generation module is used to generate each weight address corresponding to the pulse data with a value of 1 according to the pulse data of the input image; and determine each weight address corresponding to the pulse data with a value of 1 as the target weight address.
12. The device of claim 9, wherein: The address generation module is used to generate a weight address corresponding to each pulse data according to the pulse data of the input image; and determine the weight address corresponding to each pulse data as the target weight address.
13. The device of claim 11, wherein: The device also includes: A judgment module is used to judge the pulse data of the input image and determine the pulse data with a value of 1. The address generation module is used to generate each weight address corresponding to the pulse data with a value of 1.
14. The device of claim 11, wherein: The address generation module is also used to judge the pulse data of the input image and determine the pulse data with a value of 1.
15. The device according to claim 9, wherein: The full-add convolution module includes a multiplexer and an accumulator; Traversing the impulse data of the input image; When the pulse data is 0, the multiplexer determines 0 as the weight to be accumulated and outputs it; When the pulse data is 1, the multiplexer determines the weight corresponding to the reference weight address as the weight to be accumulated and outputs it, wherein the reference weight address is the target weight address corresponding to the pulse data with a value of 1; The accumulator performs an accumulation operation on each of the weights to be accumulated output by the multiplexer to obtain a full-add convolution value.
16. The device of claim 9, wherein: The full-add convolution module is an enabled accumulator; Traversing the impulse data of the input image; When the pulse data is 0, the enable accumulator is enabled to 0 and no accumulation operation is performed; When the pulse data is 1, the enable accumulator is enabled to 1, and the weight corresponding to the reference weight address is determined as the weight to be accumulated, wherein the reference weight address is the target weight address corresponding to the pulse data with a value of 1; The enabling accumulator performs an accumulation operation on each of the weights to be accumulated to obtain a full-add convolution value.
17. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the full-add convolution method according to any one of claims 1 to 8.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the full-add convolution method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Velocity measurement device and velocity measurement method
CN105008934A
SCNN reasoning device based on asynchronous circuit and PE unit, processor and computer equipment thereof
CN110378469A