Neural Network-Based Feature Data Processing Method and Apparatus
By splitting multi-bit feature data into unit-bit layers for neural network chips, the method addresses the inflexibility and complexity of existing designs, enhancing computational speed and accuracy while reducing power consumption.
Patent Information
- Application Number
- CN202111406741.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-24
AI Technical Summary
In the prior art, the design flexibility of neural network chips is poor, resulting in insufficient computing speed and accuracy, and the quantization process leads to increased computing accuracy loss and power consumption.
In a neural network chip, multi-bit feature data is split into unit bit data of multiple feature layers, and multiplication and addition operations are performed to simplify the calculation process and avoid quantization and inverse quantization processes.
It improves the computing speed and accuracy of neural networks, reduces the computing power consumption of hardware, simplifies the computing complexity, and avoids the loss of computing accuracy.
Smart Images

Figure CN114169498B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and in particular, to a method and device for processing feature data based on a neural network. Background Art
[0002] With the large-scale application of neural networks, dedicated chips for neural networks have emerged, mainly used to complete the computational processing process of neural networks to improve the computational speed of neural networks.
[0003] Currently, the chip design is usually changed based on the neural network structure to make the chip adapt to the neural network structure, so as to achieve the purpose of accelerating the computational speed and quantifying the model.
[0004] In related technologies, the chip is often independently designed according to the computational process executed by the neural network. Taking matrix calculation as an example, for the multiply-accumulate operation of the core computational function matrix in deep learning, in the Volta architecture, a dedicated execution module for matrix operations (i.e., Tensor Core) is introduced to execute matrix operations faster and improve the computational speed. However, this design method has poor flexibility and high development difficulty.
[0005] Therefore, there is an urgent need to provide a technical solution to overcome the above technical problems. Summary of the Invention
[0006] One of the technical problems solved by this application is to provide a method and device for processing feature data based on a neural network to improve the computational speed of the neural network and the accuracy of the computational results, and enhance the hardware processing efficiency.
[0007] According to an embodiment of the first aspect of this application, a method for processing feature data based on a neural network is provided. This method is applied to implement the multiply-accumulate operation in a neural network. This method is executed by a neural network chip, and the neural network chip includes a bit convolution module and an activation module. This method includes:
[0008] Obtain multi-bit feature data from the neural network;
[0009] Split the multi-bit feature data into multiple feature layers to obtain unit-bit feature data in the multiple feature layers. Among them, the number of feature layers corresponds to the data volume of the multi-bit feature data, the multiple feature layers correspond to multiple weight parameters, the feature data volume of each feature layer is a unit bit, the unit bit is a preset value, and the preset value is 1 bit;
[0010] According to the unit-bit feature data in each feature layer and the weight parameter corresponding to each feature layer, output the multiply-accumulate operation results of the multiple feature layers.
[0011] According to an embodiment of the second aspect of the present application, a feature data processing device based on a neural network is provided. The device is applied to implement the multiplication and addition operations in the neural network. The device is arranged in a neural network chip and includes:
[0012] An acquisition module configured to acquire multi-bit feature data from the neural network;
[0013] A bit matrix unit includes a bit convolution module and an activation module. The activation module is configured to split the multi-bit feature data into multiple feature layers to obtain unit-bit feature data in the multiple feature layers. Wherein, the number of feature layers of the multiple feature layers corresponds to the data volume of the multi-bit feature data. The multiple feature layers correspond to multiple weight parameters. The feature data volume of each feature layer is a unit bit, and the unit bit is a preset value, and the preset value is 1 bit. The bit convolution module is configured to receive the multiplication and addition operation results of the multiple feature layers output based on the unit-bit feature data in each feature layer and the weight parameters corresponding to each feature layer.
[0014] According to an embodiment of the third aspect of the present application, an electronic device is provided, which includes a processor and a memory. Wherein, an executable code is stored on the memory. When the executable code is executed by the processor, the processor can at least implement the feature data processing method based on the neural network in the first aspect.
[0015] According to an embodiment of the fourth aspect of the present application, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by an electronic device, the electronic device can be enabled to execute at least the feature data processing method based on the neural network in the first aspect.
[0016] According to the fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program / instructions. When the computer program / instructions are executed by a processor, the feature data processing method based on the neural network in the first aspect is implemented.
[0017] In the embodiments of the present application, the neural network chip includes a bit convolution module and an activation module. In the process of the neural network chip implementing the multiplication and addition operations in the neural network, multi-bit feature data from the neural network is obtained and split into multiple feature layers to obtain unit-bit feature data in the multiple feature layers. Since the number of feature layers of the multiple feature layers corresponds to the data volume of the multi-bit feature data, the multiple feature layers correspond to multiple weight parameters, and the feature data volume of each feature layer is a unit bit, and the unit bit is a preset value of 1 bit. Through the above splitting process, the high-order multiplication and addition operations can be split into relatively simple selective addition operations, greatly reducing the operation complexity and improving the hardware operation speed. That is, according to the unit-bit feature data in each feature layer and the weight parameter corresponding to each feature layer, the multiplication and addition operation results of the multiple feature layers are output.
[0018] In the embodiments of the present application, the multiplication and addition operations in the neural network can be simplified, the computational amount brought by the quantization / anti-quantization process in the neural network can be reduced, the operation speed of the hardware (such as a chip) can be improved, and the computational power consumption of the hardware can be reduced. In addition, since there is no need to go through the quantization / anti-quantization process, the loss of computational accuracy brought by the above process can be avoided, and the accuracy of the calculation result can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:
[0020] Figure 1 is a schematic flowchart of a method for processing feature data based on a neural network according to an embodiment of the present application;
[0021] Figure 2 is a schematic diagram of the principle of a method for processing feature data based on a neural network according to an embodiment of the present application;
[0022] Figures 3 to 4 is a schematic diagram of the principle of a storage method according to an embodiment of the present application;
[0023] Figure 5 is a schematic diagram of the working principle of a chip according to an embodiment of the present application;
[0024] Figure 6 is a schematic diagram of the effect of a method for processing feature data based on a neural network according to an embodiment of the present application;
[0025] Figure 7 is a schematic diagram of a training image according to an embodiment of the present application;
[0026] Figure 8 is a schematic diagram of the principle of model training according to an embodiment of the present application;
[0027] Figure 9 It is a schematic diagram of the effect of model training according to an embodiment of the present application;
[0028] Figure 10 It is a schematic structural diagram of a neural network-based feature data processing device according to an embodiment of the present application. Detailed implementation manners
[0029] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0030] The computer device includes a user device and a network device. Among them, the user device includes, but is not limited to, a computer, a smart phone, a PDA, etc.; the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing (Cloud Computing), where cloud computing is a type of distributed computing and consists of a super virtual computer formed by a group of loosely coupled computer sets. Among them, the computer device can run alone to implement the present application, or can be connected to the network and implement the present application through interaction with other computer devices in the network. Among them, the network where the computer device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, etc.
[0031] It should be noted that the user device, the network device, the network, etc. are only examples, and other existing or future possible computer devices or networks that are applicable to the present application should also be included within the protection scope of the present application and are hereby incorporated by reference.
[0032] The methods discussed later (some of which are illustrated by flowcharts) can be implemented by hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segments for implementing the necessary tasks can be stored in a machine or computer-readable medium (such as a storage medium). One or more processors can implement the necessary tasks.
[0033] The specific structural and functional details disclosed herein are merely representative and are for the purpose of describing exemplary embodiments of the present application. However, the present application can be embodied in many alternative forms and should not be construed as limited only to the embodiments set forth herein.
[0034] It should be understood that although the terms "first", "second", etc. may be used herein to describe various modules, these modules should not be limited by these terms. These terms are only used to distinguish one module from another. For example, without departing from the scope of the exemplary embodiments, the first module may be referred to as the second module, and similarly the second module may be referred to as the first module. The term "and / or" used herein includes any and all combinations of one or more of the associated listed items.
[0035] It should be understood that when a module is referred to as being "connected" or "coupled" to another module, it can be directly connected or coupled to the other module, or there may be intermediate modules. In contrast, when a module is referred to as being "directly connected" or "directly coupled" to another module, there are no intermediate modules. Other words used to describe the relationship between modules should be interpreted in a similar manner (e.g., "between" compared to "directly between", "adjacent to" compared to "directly adjacent to", etc.).
[0036] The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the exemplary embodiments. Unless the context clearly dictates otherwise, the singular forms "a", "an" used herein are also intended to include the plural. It should also be understood that the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, modules, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, modules, components, and / or combinations thereof.
[0037] With the large-scale application of neural networks, dedicated chips for neural networks have emerged, mainly used to complete the computational processing process of neural networks to improve the computational speed of neural networks.
[0038] Currently, it is usually based on the neural network structure to change the chip design to make the chip adapt to the neural network structure, so as to achieve the purpose of accelerating the computational speed and quantifying the model.
[0039] In related technologies, the chip is often independently designed according to the calculation process performed by the neural network. Taking matrix calculation as an example, for the multiply-accumulate operation of the core calculation function matrix in deep learning, in the Volta architecture, a dedicated execution module for matrix operations (i.e., Tensor Core) is introduced to execute matrix operations faster and improve the calculation speed. However, this design method has poor flexibility and high development difficulty.
[0040] In addition, during the process of quantizing convolution in existing neural networks, it is necessary to appropriately reduce the precision to match the quantization model. Specifically, the input data type of the neural network is generally fp16 (16-bit), fp32 (32-bit), or int8 (8-bit), etc. The initial result of int32 type is obtained through layer-by-layer processing. The sum of the initial result and the offset (bias) needs to be further quantized into int8 type data, and the value obtained by passing the int8 type data through the Rectified Linear Unit (ReLU) is dequantized into the final output result. Therefore, multiple quantizations and dequantizations are required during the neural network calculation process, resulting in loss of calculation precision, affecting the accuracy of the final calculation result, and increasing the power consumption of the dedicated execution module.
[0041] Therefore, there is an urgent need to provide a technical solution to overcome the above technical problems.
[0042] To address at least one of the above technical problems, the present application proposes a method and device for processing feature data based on a neural network to improve the calculation speed and accuracy of the neural network and enhance the hardware processing efficiency.
[0043] The core principle of this technical solution is: when implementing the multiply-accumulate operation in the neural network on the neural network chip, multi-bit feature data from the neural network is obtained; the multi-bit feature data is split into multiple feature layers to obtain unit-bit feature data in the multiple feature layers. Since the number of feature layers in the multiple feature layers corresponds to the data volume of the multi-bit feature data, the multiple feature layers correspond to multiple weight parameters, and the feature data volume of each feature layer is a unit bit, and the unit bit is 1 bit. Through the above splitting process, the high-order multiply-accumulate operation can be split into relatively simple selective addition operations, thereby reducing the operation complexity and improving the hardware operation speed, that is, according to the unit-bit feature data in each feature layer and the weight parameters corresponding to each feature layer, the multiply-accumulate operation results of the multiple feature layers are output.
[0044] Through the above method, the multiplication and addition operations in the neural network can be simplified, the computational complexity brought by the quantization / anti-quantization process in the neural network can be reduced, the operation speed of the hardware (such as a chip) can be improved, and the computational power consumption of the hardware can be reduced. In addition, since there is no need to go through the quantization / anti-quantization process, the loss of computational accuracy brought by the above process can be avoided, and the accuracy of the calculation result can be improved.
[0045] Based on the above core principle, the technical solution provided by the present disclosure can be applied to various neural networks. Among them, the applicable neural networks include, but are not limited to, any one of Deep Neural Networks, Recurrent Neural Networks models, Recursive Neural Networks, Convolutional Neural Networks, Graph Convolutional Networks (GCNs), or models obtained by transforming one or more of the above deep learning models. This embodiment is not limited.
[0046] The feature data processing solution based on the neural network provided by the present disclosure can be executed by an electronic device, which can be a terminal device such as a smart phone, a tablet computer, a PC, a laptop computer, etc. In an alternative embodiment, the electronic device can also be implemented as a service device for executing the technical solution provided by this application, such as a server, a server cluster, a cloud server, etc.
[0047] After introducing the core principle, application scenarios, and execution devices, the following will introduce a method, apparatus, device, and medium provided by the present disclosure in combination with specific embodiments.
[0048] Figure 1 is a flowchart of a method for processing feature data based on a neural network shown according to an exemplary embodiment. It can be understood that this method is applied to implement the multiplication and addition operations in the neural network, especially the multiplication and addition operations of high-order feature data. This method can be executed by a neural network chip, which includes a bit convolution module and an activation module. As Figure 1 shown, the method includes the following steps:
[0049] In step 101, multi-bit feature data from the neural network is obtained.
[0050] In step 102, the multi-bit feature data is split into multiple feature layers to obtain unit-bit feature data in the multiple feature layers. This step can be implemented by the activation module. For the specific introduction of the activation module, please refer to the following.
[0051] For example, after the neural network obtains the data to be processed, it performs corresponding processing on the data to be processed and extracts the feature data. For example, the data to be processed can be one or more pictures, or a video, or text. Or other types of data. In practical applications, the data to be processed can be obtained and input into the neural network model according to the specific scenario. The relevant parameters and specific acquisition methods of the data to be processed are not limited in this application.
[0052] Specifically, in 101, the feature data extracted by the neural network is received. Generally, the feature data is a high-order number, such as 16bit or 32bit. Furthermore, in 102, the feature data is split according to multiple feature layers. Among them, the number of feature layers of the multiple feature layers corresponds to the data volume of the multi-bit feature data, and the multiple feature layers correspond to multiple weight parameters (weight). It should be particularly noted that the feature data volume of each feature layer is a unit bit, and the unit bit is a preset value, and the preset value is 1 bit (bit). For example, the n-bit feature data output by the neural network is split into n unit-bit feature data. Assuming that the unit bit is 1bit, based on this, n feature layers are extracted from the n-bit feature data, and the data volume of each feature layer is 1bit.
[0053] Furthermore, for the feature data of the multiple feature layers, in 103, according to the unit-bit feature data in each feature layer and the weight parameter corresponding to each feature layer, the multiplication-addition operation results of the multiple feature layers are output, so that the multiplication-addition operation results are applied to the subsequent processing process. This step can be implemented through a bit convolution module. For the specific introduction of the bit convolution module, please refer to the following text.
[0054] Specifically, in an optional embodiment, different feature layers have different weight parameters. Assume that the type of the weight parameter is Int8 (including data of 8bit) or UInt8 type, where the integer range of Int8 is [-128:127], and the integer range of UInt8 is [0:255]. Based on this, in 103, the unit-bit feature data in each feature layer is multiplied by the weight parameter corresponding to each feature layer to obtain the products of the multiple feature layers. Furthermore, the above products are added together to obtain the multiplication-addition operation results of the multiple feature layers. In the entire calculation process of the neural network, since the data volume in the feature layer is 1bit and the value of the unit-bit feature data is 0 or 1, in this case, splitting the high-order feature data into multiple feature layers and performing multiplication-addition operations with the weight parameter (Int8) can greatly reduce the calculation amount and complexity of the multiplication-addition operation, improve the operation processing efficiency, and reduce the power consumption required for the operation.
[0055] Of course, in another embodiment, specifically, in 103, the bit convolution module obtains the product of the unit bit feature data in each feature layer and the weight parameter corresponding to each feature layer, and adds up the products corresponding to the multiple feature layers respectively to obtain the multiplication and addition operation result of the multiple feature layers; through the activation module in 102, the multiplication and addition operation result is compared with the offset to obtain the multiplication and addition operation result of unit bits.
[0056] Optionally, the value of the offset can be set according to the neural network. For example, assume that the feature layer and its corresponding weight parameter are respectively: the feature layer 0 and its corresponding weight parameter is 21, the feature layer 1 and its corresponding weight parameter is 35, the feature layer 1 and its corresponding weight parameter is -15, and the feature layer 0 and its corresponding weight parameter is 100. Based on the above assumption, the following formula is used in 103 to determine the multiplication and addition operation result of the final output:
[0057] S = bias + [0 * 20 + 1 * 35 + 1 * (-15) + 0 * 100 + ……]
[0058] S = bias + 20
[0059] In the above formula, the offset bias can be obtained through the neural network. When S > 0, the output multiplication and addition operation result is 1; when S <= 0, the multiplication and addition operation result is 0. Or, according to the specific situation, it can also be set that when S >= 0, the multiplication and addition operation result is 1; when S < 0, the multiplication and addition operation result is 0. Among them, when bias is not required, the value of bias can be set to 0.
[0060] In the embodiments of the present application, the weight parameter corresponding to the feature layer is of floating-point type before quantization (Quantize). The quantization method is introduced in combination with the following embodiments, that is:
[0061] In an alternative embodiment, before 102, the maximum absolute value among multiple weight parameters is obtained; according to the maximum absolute value and the weight parameter corresponding to each feature layer, the compensation coefficient corresponding to each feature layer is determined. Based on this, in 102, the implementation manner of comparing the operation result of 103 with the offset to obtain the multiplication and addition operation result of unit bits may include: subtracting the product of the compensation coefficient and the offset from the intermediate operation result to obtain the quantization value of the multiplication and addition operation result.
[0062] For example, assume that there are multiple weight parameters to be quantized. The parameter with the largest absolute value can be selected from these weight parameters as the range. Based on this, the process of quantizing to int8 can be expressed as weight = int(weight * 127 / range). Simply put, it is to directly multiply weight by the compensation coefficient a = 127 / range. Furthermore, when calculating the 1-bit convolution (conv), when subtracting the above product from the offset (bias), only multiply the bias by the compensation coefficient a.
[0063] For example, assume that the weight parameters to be quantized are [+10.5, +11.1, +12.3, +15.7, +18.9]. According to the above method, the parameter with the largest absolute value is 18.9. The quantization results obtained by using weight = int(weight * 127 / range) are [71, 75, 83, 105, 127].
[0064] However, referring to the above example, it is not difficult to find that the quantization process changes the parameter distribution from [-127, +127] to [71, 127], as Figure 2 shown. In Figure 2 it can be clearly seen the precision difference between the two parameter distributions. Therefore, the quantization process in the above example will obviously cause precision loss, which in turn affects the accuracy of the multiplication and addition operation results in the neural network.
[0065] Therefore, in another alternative embodiment, a quantization method is also provided. Before 102, according to the weight parameters and bias parameters corresponding to each feature layer, determine the compensation coefficient corresponding to each feature layer. Based on this, in 102, when comparing the intermediate operation result with the offset to obtain the implementation method of the unit-bit multiplication and addition operation result, it may include: obtaining the number of unit-bit feature data with a preset data value for multiple feature layers; subtracting the product of the compensation coefficient and the offset from the intermediate operation result, and adding the difference to the product of the number and the bias parameter (offset) to obtain the quantization value of the multiplication and addition operation result.
[0066] For example, assume that the weight parameters to be quantized are [-1.8, +0.1, +0.5, +2.3, +2.5]. Assume that the preset data value is 1. Based on the above assumptions, the difference between this quantization method and the previous embodiment is that during the calculation of the 1-bit conv, count the number of unit-bit feature data with a value of 1 in the unit-bit feature data of multiple feature layers. Assume the number is n. When subtracting the product of weight and the compensation coefficient a from the offset (bias), it is also necessary to add n * offset. Although this quantization method is more cumbersome than the previous one, it can effectively improve the precision of weight.
[0067] ByFigure 1 The provided feature processing method can simplify the multiplication and addition operations in a neural network, reduce the computational load brought about by the quantization / anti-quantization process in the neural network, improve the hardware operation speed, and reduce the hardware computing power consumption. In addition, since there is no need to go through the quantization / anti-quantization process, it is possible to avoid the loss of computational accuracy brought about in the above process and improve the accuracy of the calculation results.
[0068] In the above or following embodiments, optionally, the multiplication and addition operation results are stored in a vertical depth storage manner. Specifically, the basic unit for storing data is a byte. Assuming the input data type is 1 bit, and assuming that 8-bit data is stored as a group during storage. Two storage (pack) methods are provided in this application:
[0069] The first storage method is hierarchical storage, that is, a group of data is stored in each layer, and each group of data is 8 bits, as Figure 3 shown. The implementation of this storage method is relatively simple, but the speed and efficiency are not good.
[0070] The second storage method is position-based storage. That is, the unit-bit feature data in different feature layers is stored in the same position. The vertical depth storage method. This storage method is because during the process of calculating 1-bit conv, n 1-bit feature layers are multiplied and added with the weight, and then stored in depth of 8 bits, as Figure 4 shown. In this way, during convolution calculation, the 8-bit feature data can be directly multiplied and added with the corresponding weight, thereby further improving the multiplication and addition operation speed and efficiency.
[0071] In the above or following embodiments, the neural network chip is a bitmatrix chip unit. This bitmatrix chip unit includes two basic module structures: The first module is the bit convolution (bit conv) module. Taking Figure 5 as an example, the input of this bit conv module is 1 bit, and in this bit conv module, multiplication and addition operations can be performed through 1 bit and weight (int8), and finally a 16-bit or 32-bit multiplication and addition operation result is obtained. The second module is the activation module. Taking Figure 5 as an example, the input feature data is compared with the offset (bias) to obtain a 1-bit multiplication and addition operation result, and the storage result is obtained through the pack method (such as the above two storage methods). It should be noted that the above two basic modules can be called jointly or separately.
[0072] Specifically, the following provides the calling methods for the above two basic modules:
[0073] The first calling method is to call the bit conv module alone. Specifically, the input 1-bit feature data is output as int16 or int32 output data through the bit conv module; further, it is quantized to int8 output data for performing other operations in the neural network.
[0074] The second calling method is to call the activation module alone. Specifically, after the convolution operation of two int8-type feature data in the neural network, the multiplication and addition operation result can be converted into 1-bit feature data through the activation module and then stored in a packed manner.
[0075] The third calling method is the combined call of the bit conv module and the activation module. Specifically, the 1-bit feature data is input into the bit conv module to output int16 or int32 output data; further, 1-bit feature data (i.e., the unit bit feature data introduced above) is output through the activation module and then stored in a packed manner. Of course, the 1-bit multiplication and addition operation result can also be used as the input of other bit conv modules, which depends on the design of the neural network itself and is not limited in this application.
[0076] Through the above three calling methods, the mutual conversion and connection of the internal modules of the bitmatrix chip unit can be achieved, and at the same time, the connection and intercommunication of other functional modules in the chip can also be achieved through the above calling methods.
[0077] It should be noted that there is another advantage in using the bit conv module. Specifically, for a convolutional neural network, during the network training process, the more the number of weights, the better the training result. However, in actual applications, the more the number of weights, the slower the calculation speed.
[0078] Take Figure 6 as an example. Suppose the input data volume to be processed is 1*128*64*641, where 1 is the number, 128 is the number of channels, and 64*64 is the feature layer size. After one convolution calculation, a calculation result of 1*256*64*64 can be obtained. At this time, the corresponding weight should be 256*128*3*3, where 256 is the number, 128 is the number of channels, and 3*3 is the feature layer size. As the number of weights increases, the bit conv module provided by this application can further improve the calculation speed in the case of more weight parameters.
[0079] In the above or following embodiments, it is assumed that the CIFAR-10 / CIFAR-100 image classification model is used for training. Among them, the classification of the CIFAR-10 dataset is an open benchmark problem in machine learning. The specific task objective is to classify a group of 32x32 RGB images, and this dataset covers 10 categories: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck, as Figure 7 shown.
[0080] The goal of this test is to build a relatively small convolutional neural network for image recognition. In the related art, CIFAR-10 is selected because its complexity is sufficient to test most functions in TensorFlow and can be extended to a larger model. At the same time, due to the small size of the model, the training speed is very fast, which is more suitable for testing new ideas and verifying new technologies.
[0081] Taking the CIFAR-10 image classification as an example, the first layer is a conventional convolutional neural network, using the activation function sigmoid(x), and this sigmoid function is continuously differentiable everywhere. From the second to the fourth time, the complete bit-conv and activation are executed. Furthermore, the final output result is obtained after the fifth bit-conv, and no activation is required before this output result. The implementation methods of the above bit-conv and activation can be seen in the previous text and will not be elaborated here.
[0082] During the training process, x in sigmoid(x*p) is multiplied by the parameter p, and then the value of the parameter p is gradually increased to make the curve of the sigmoid function closer to the real curve (such as Figure 9 the changing trend of the solid curve in), so that the target model can be obtained through multiple (such as millions of times) model trainings. When verifying the neural network model, the input data of the bitmatrix chip unit can be controlled by a hardware switch. After passing through the neural network model matched by the bitmatrix chip unit, if the obtained execution result is consistent with the execution result obtained by the related art, it means passing the verification.
[0083] The embodiment of the present application also provides a neural network-based feature data processing device corresponding to the above method. This device is applied to implement the multiplication and addition operations in the neural network, and this device is set in the neural network chip. As Figure 10 shown in is the schematic structural diagram of the device. This device mainly includes:
[0084] Acquisition module 10, which is configured to acquire multi-bit feature data from a neural network; the acquisition module 10 receives the feature data extracted by the neural network. For example, the neural network extraction can extract the feature data to be further processed from pictures, videos, texts, and voices.
[0085] Bit array chip unit 11, including a bit convolution module and an activation module, where
[0086] The activation module is configured to split the multi-bit feature data into multiple feature layers to obtain unit-bit feature data in the multiple feature layers;
[0087] The bit convolution module is configured to receive the multiplication and addition operation results of the multiple feature layers output based on the unit-bit feature data in each feature layer and the weight parameters corresponding to each feature layer.
[0088] Among them, the number of feature layers of the multiple feature layers corresponds to the data volume of the multi-bit feature data. The multiple feature layers correspond to multiple weight parameters, and the feature data volume of each feature layer is a unit bit. The unit bit is a preset value, and the preset value is 1 bit.
[0089] Optionally, when the bit array chip unit 11 outputs the multiplication and addition operation results of the multiple feature layers according to the unit-bit feature data in each feature layer and the weight parameters corresponding to each feature layer, it is specifically configured as:
[0090] The bit convolution module obtains the product of the unit-bit feature data in each feature layer and the weight parameters corresponding to each feature layer, and adds the products corresponding to the multiple feature layers respectively to obtain the multiplication and addition operation results of the multiple feature layers;
[0091] The activation module compares the intermediate operation result with an offset to obtain the unit-bit multiplication and addition operation result.
[0092] Optionally, the value of the offset can be set according to the neural network.
[0093] Optionally, the bit array chip unit 11 is further configured to: obtain the maximum absolute value among the multiple weight parameters; determine the compensation coefficient corresponding to each feature layer according to the maximum absolute value and the weight parameters corresponding to each feature layer.
[0094] When the bit array chip unit 11 compares the intermediate operation result with an offset to obtain the multiplication and addition operation result, it is specifically configured as: subtracting the product of the compensation coefficient and the offset from the intermediate operation result to obtain the quantization value of the multiplication and addition operation result.
[0095] Optionally, the bit matrix chip unit 11 is further configured to determine a compensation coefficient corresponding to each feature layer according to the weight parameter and the bias parameter corresponding to each feature layer.
[0096] When the bit matrix chip unit 11 compares the intermediate operation result with the offset to obtain the multiplication and addition operation result, it is specifically configured as follows:
[0097] Obtain the number of unit bit feature data of the multiple feature layers that are preset data values; subtract the product of the compensation coefficient and the offset from the intermediate operation result, and add the difference to the product of the number and the bias parameter to obtain the quantization value of the multiplication and addition operation result.
[0098] Optionally, the multiplication and addition operation result is stored in a vertical depth storage manner.
[0099] In the embodiments of the present application, the multiplication and addition operations in the neural network can be simplified, the computational complexity brought by the quantization / anti-quantization process in the neural network can be reduced, the hardware operation speed can be improved, and the hardware computational power consumption can be reduced. In addition, since there is no need to go through the quantization / anti-quantization process, the computational accuracy loss brought by the above process can be avoided, and the accuracy of the calculation result can be improved.
[0100] It should be noted that the present application can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application specific integrated circuit (ASIC), a general purpose computer, or any other similar hardware device. In one embodiment, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. In addition, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with a processor to execute each step or function.
[0101] In addition, a part of the present application can be applied as a computer program product, such as computer program instructions. When executed by a computer, through the operation of the computer, the methods and / or technical solutions according to the present application can be called or provided. The program instructions for calling the method of the present application may be stored in a fixed or removable recording medium, and / or transmitted through a data stream in a broadcast or other signal-bearing medium, and / or stored in the working memory of a computer device that runs according to the program instructions. Here, an embodiment according to the present application includes a device, which includes a memory for storing computer program instructions and a processor for executing the program instructions. When the computer program instructions are executed by the processor, the device is triggered to run based on the methods and / or technical solutions according to the foregoing multiple embodiments of the present application.
[0102] For those skilled in the art, it is obvious that the present application is not limited to the details of the above-mentioned exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present application. Any reference signs in the claims should not be construed as limiting the claims concerned. In addition, it is obvious that the term "comprising" does not exclude other modules or steps, and the singular does not exclude the plural. The multiple modules or devices stated in the system claims can also be implemented by one module or device through software or hardware. The words "first", "second", etc. are used to denote names and do not denote any particular order.
Claims
1. A method for processing feature data based on a neural network, characterized in that Applied to implement the multiply-accumulate operation in a neural network. The method is executed by a neural network chip, and the neural network chip includes a bit convolution module and an activation module. The method includes: Obtain multi-bit feature data from the neural network; Split the multi-bit feature data into multiple feature layers to obtain unit-bit feature data in the multiple feature layers. Wherein, the number of feature layers of the multiple feature layers corresponds to the data volume of the multi-bit feature data. The multiple feature layers correspond to multiple weight parameters. The feature data volume of each feature layer is a unit bit, and the unit bit is a preset value, and the preset value is 1 bit; Obtain the product of the unit-bit feature data in each feature layer and the weight parameter corresponding to each feature layer; Add the products corresponding to the multiple feature layers respectively to obtain an intermediate operation result of the multiple feature layers; Compare the intermediate operation result with an offset to obtain a multiply-accumulate operation result; Wherein, the comparing the intermediate operation result with the offset to obtain the multiply-accumulate operation result includes: determining a compensation coefficient corresponding to each feature layer according to the weight parameter and the bias parameter corresponding to each feature layer; obtaining the number of unit-bit feature data in the multiple feature layers that are preset data values; subtracting the product of the compensation coefficient and the offset from the intermediate operation result, and adding the difference to the product of the number and the bias parameter to obtain a quantization value of the multiply-accumulate operation result.
2. The method according to claim 1, characterized in that, The value of the offset can be set according to the neural network.
3. The method according to claim 1, wherein It further includes: Obtain the maximum absolute value among the multiple weight parameters; Determine the compensation coefficient corresponding to each feature layer according to the maximum absolute value and the weight parameter corresponding to each feature layer; The comparing the intermediate operation result with the offset to obtain the multiply-accumulate operation result includes: Subtract the product of the compensation coefficient and the offset from the intermediate operation result to obtain a quantization value of the multiply-accumulate operation result.
4. The method according to any one of claims 1 to 3, characterized in that Store the multiply-accumulate operation result in a vertical depth storage manner.
5. A feature data processing device based on a neural network, characterized in that, Applied to implement the multiply-accumulate operation in a neural network. The device is arranged in a neural network chip, and the device includes An acquisition module configured to obtain multi-bit feature data from the neural network; A bit matrix unit includes a bit convolution module and an activation module. The activation module is configured to split the multi-bit feature data into multiple feature layers to obtain unit-bit feature data in the multiple feature layers. Wherein, the number of feature layers in the multiple feature layers corresponds to the data volume of the multi-bit feature data, the multiple feature layers correspond to multiple weight parameters, the feature data volume of each feature layer is a unit bit, the unit bit is a preset value, and the preset value is 1 bit; the bit convolution module is configured to receive and obtain the product of the unit-bit feature data in each feature layer and the weight parameter corresponding to each feature layer, add the products corresponding to the multiple feature layers respectively to obtain an intermediate operation result of the multiple feature layers, and determine a compensation coefficient corresponding to each feature layer according to the weight parameter and the bias parameter corresponding to each feature layer; obtain the number of unit-bit feature data in the multiple feature layers that are preset data values; subtract the product of the compensation coefficient and the offset from the intermediate operation result, and add the sum of the difference and the product of the number and the bias parameter to obtain a quantization value of the multiply-accumulate operation result.
6. An electronic device, characterized in that, Comprising: A memory and a processor; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor executes the neural network-based feature data processing method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by an electronic device, the electronic device is enabled to execute the neural network-based feature data processing method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method and apparatus for hardware simulation and emulation during running, and device and storage medium
CN113228056A
Method and system for bit quantization of artificial neural network
CN113396427A