Neural network quantization method and apparatus, electronic device, medium
By expanding and compressing the static data of recurrent neural networks in the time dimension, the problem of limited degrees of freedom of static data is solved, achieving high-precision quantization and improved storage efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-26
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies for quantizing recurrent neural networks suffer from limited degrees of freedom in static data, making quantization difficult and leading to the accumulation of quantization errors that cause the neural network to fail.
By unfolding the neural network units along the time dimension, determining the static data at multiple time steps, quantizing them separately, and then compressing them to obtain shared data, the degree of freedom of the static data is increased and the difficulty of quantization is reduced.
This improves the accuracy of the quantized neural network, reduces storage requirements, and avoids neural network failure caused by excessive quantization errors.
Smart Images

Figure CN114912584B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a neural network quantization method and apparatus, electronic device, and medium. Background Technology
[0002] A Recurrent Neural Network (RNN) consists of one or more basic RNN units, each containing multiple neurons. When processing sequential data, such as text classification, speech recognition, and text translation, RNN units are used cyclically in the time dimension, meaning that computations for a certain number of time steps are performed on the same RNN unit. Summary of the Invention
[0003] This disclosure provides a neural network quantization method, apparatus, electronic device, and medium.
[0004] In a first aspect, this disclosure provides a method for quantizing a neural network, the neural network comprising at least one neural network unit, each neural network unit comprising multiple time steps, the method comprising:
[0005] The neural network units to be quantized in the neural network are quantized one by one to obtain the target neural network unit;
[0006] Based on the target neural network unit, determine the quantized target neural network;
[0007] The step of quantizing the neural network units to be quantized in the neural network includes:
[0008] For any neural network unit to be quantized, the neural network unit is expanded in the time dimension to determine the static data of the neural network unit at multiple time steps;
[0009] The static data of each of the multiple time steps are quantized to obtain the corresponding quantized data;
[0010] The quantized data of multiple time steps are compressed to obtain shared data for multiple time steps; wherein the amount of shared data is less than the amount of static data.
[0011] The neural network unit corresponding to the shared data is used as the target neural network unit.
[0012] In a second aspect, this disclosure provides a neural network quantization device, the neural network including at least one neural network unit, each neural network unit including multiple time steps, the device comprising:
[0013] The quantization module is used to quantize the neural network units to be quantized in the neural network to obtain the target neural network units;
[0014] The determining module is used to determine the quantized target neural network based on the target neural network unit;
[0015] The quantization module includes:
[0016] An unfolding unit is used to unfold any neural network unit to be quantized in the time dimension to determine the static data of the neural network unit at multiple time steps.
[0017] A quantization unit is used to quantize the static data of multiple time steps respectively to obtain the corresponding quantized data;
[0018] A compression unit is used to compress the quantized data of multiple time steps to obtain shared data of multiple time steps; wherein the amount of shared data is less than the amount of static data.
[0019] A determining unit is used to identify the neural network unit corresponding to the shared data as the target neural network unit.
[0020] Thirdly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the neural network quantization method described above.
[0021] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor / processing core, implements the above-described neural network quantization method.
[0022] The embodiments provided in this disclosure, for any neural network unit to be quantized, unfold the neural network unit in the time dimension to determine the static data of the neural network unit at multiple time steps; quantize the static data of multiple time steps respectively to obtain the corresponding quantized data; then compress the quantized data of multiple time steps to obtain shared data of multiple time steps; finally, take the neural network unit corresponding to the shared data as the target neural network unit, and determine the quantized target neural network based on the target neural network unit. This method can adaptively share data, improve the degree of freedom of static data, reduce the difficulty of quantization, and at the same time, improve the accuracy of the neural network while meeting the storage requirements of the neural network, avoiding the failure of the neural network due to excessive quantization error.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0024] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0025] Figure 1 This is a schematic diagram showing the unfolding of a neural network unit along the time dimension.
[0026] Figure 2 A flowchart of a neural network quantization method provided in this embodiment of the disclosure;
[0027] Figure 3 This is a flowchart illustrating the quantization process for neural network units to be quantized in an embodiment of this disclosure.
[0028] Figure 4 This is a schematic diagram illustrating how a neural network unit, after being unfolded in the time dimension, independently stores static data according to an embodiment of this disclosure.
[0029] Figure 5 A flowchart of a compression process provided in this embodiment of the disclosure;
[0030] Figure 6 This is a schematic diagram illustrating the clustering of quantified data in an embodiment of this disclosure;
[0031] Figure 7 This is a flowchart illustrating the retraining of the neural network in an embodiment of this disclosure;
[0032] Figure 8 A flowchart of the neural network quantization method provided in the embodiments of this disclosure;
[0033] Figure 9 A block diagram of a neural network quantization device provided in an embodiment of this disclosure;
[0034] Figure 10 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0035] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0036] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0037] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0038] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0039] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0040] The neural network quantization method according to embodiments of this disclosure can be executed by electronic devices such as terminal devices or servers. Terminal devices can be in-vehicle devices, user equipment (UE), mobile devices, user terminals, terminals, cellular phones, cordless phones, personal digital assistants (PDAs), handheld devices, computing devices, in-vehicle devices, wearable devices, etc. The method can be implemented by a processor calling computer-readable program instructions stored in memory. Alternatively, the method can be executed by a server.
[0041] Figure 1 This is a schematic diagram illustrating the unfolding of a neural network unit along the time dimension. For example... Figure 1 As shown, the neural network units are used cyclically in the time dimension. Except for the first time step, the inputs of the second, third, and fourth time steps include both the sequence data corresponding to the current time step and the intermediate results (output results) of the neural network units in the previous time step. The neural network units share static data W0 (including weight data and bias data) at different time steps, that is, the neural network reuses static data W0 in time.
[0042] To improve the efficiency of neural networks, it is necessary to compress their size. Quantization is a common method used in the process of compressing neural networks. However, when quantizing a neural network, the use of the same static data across all time steps restricts the degrees of freedom of the static data, making quantization difficult and making it hard to obtain static data suitable for all time steps. Moreover, quantization errors accumulate over time, potentially causing the quantized neural network to fail.
[0043] This disclosure provides a neural network quantization method that can increase the degrees of freedom of static data at each time step, reduce the difficulty of quantization, and avoid neural network failure due to excessive quantization error.
[0044] Figure 2 This is a flowchart illustrating a neural network quantization method provided in an embodiment of this disclosure. The neural network involved in this method includes at least one neural network unit, each neural network unit is used cyclically in the time dimension, and each neural network unit includes multiple time steps.
[0045] like Figure 2 As shown, the neural network quantization method includes:
[0046] Step S201: Quantize the neural network units to be quantized in the neural network to obtain the target neural network units.
[0047] Quantization involves converting neural network data from high precision to low precision to compress the network's size. For example, converting the numerical format of static data such as weights and biases from FP32 (32-bit floating-point, single precision) to FP16 (half-precision floating-point) or INT8 (8-bit fixed-point integer) can compress the size of the neural network. A common neural network compression method is to convert FP32 to INT8, and the compressed target neural network unit stores static data in INT8 format.
[0048] Step S202: Determine the quantized target neural network based on the target neural network unit.
[0049] The target neural network unit is a neural network unit that can meet both accuracy and storage requirements. The target neural network determined by the target neural network unit can also meet both accuracy and storage requirements at the same time.
[0050] For example, the neural network includes a first neural network unit and a second neural network unit. The first neural network unit and the second neural network unit are quantized in step S201 to obtain a first target neural network unit and a second target neural network unit. In other words, the first neural network unit and the second neural network unit are quantized to obtain the corresponding first target neural network unit and the second target neural network unit. The neural network determined by the first target neural network unit and the second target neural network unit is the target neural network.
[0051] Figure 3 This is a flowchart illustrating the quantization process for neural network units to be quantized in an embodiment of this disclosure. For example... Figure 3 As shown, step S201, which involves quantizing the neural network units to be quantized in the neural network, includes:
[0052] Step S301: For any neural network unit to be quantized, expand the neural network unit in the time dimension to determine the static data of multiple time steps of the neural network unit.
[0053] In a brain-inspired many-core architecture, the time loop of a neural network can be expanded in the time dimension, that is, expanded from time into space, mapping time steps to computation kernels, with each time step corresponding to one computation kernel, and then executing computation in a pipelined manner to improve the throughput of the neural network. After the neural network unit is expanded in the time dimension, processing the neural network in the brain-inspired many-core architecture is equivalent to multiplying the number of time steps with the number of neural network units.
[0054] In some embodiments, any neural network unit to be quantized is unfolded along the time dimension to determine the static data of the neural network unit at multiple time steps. If the number of time steps in the neural network is K, then K copies of the static data need to be stored independently. This allows each neural network to have different static data, reducing the quantization difficulty and improving the accuracy of the quantized neural network. Initially, the initial value of the static data is W0, i.e., W i =W0, where i = 1, 2, ..., K, and K is an integer greater than 1.
[0055] Figure 4 This is a schematic diagram illustrating how the neural network units of this disclosure are unfolded in the time dimension and then independently stored as static data, according to an embodiment of the present disclosure. Figure 4As shown, the first, second, third, and fourth time steps store the first static data W1, the second static data W2, the third static data W3, and the fourth static data W4, respectively. These four static data can change independently, increasing the degrees of freedom of the static data and reducing the errors in the neural network caused by quantization.
[0056] It should be noted that, initially, the first static data W1, the second static data W2, the third static data W3, and the fourth static data W4 of the neural network unit can be set as static data W0, and perturbation values can be added to the static data W0, that is, the static data W0 at each time step can be modified, but the magnitude of the modification is within the allowable range.
[0057] Step S302: Quantize the static data of multiple time steps respectively to obtain the corresponding quantized data.
[0058] In the embodiments of this disclosure, a general method for quantizing static data can be used, and this disclosure does not limit the specific method of quantization.
[0059] For example, a quantization function Q is used for static data W. i Quantization is performed to obtain quantized data W. i ′, i.e. W i ′=Q(W i ), where i = 1, 2, ..., m.
[0060] After static data quantization, the corresponding quantized data W is obtained at each time step. i For example, the first static data W1, the second static data W2, the third static data W3, and the fourth static data W4 correspond to the first quantized data W1′, the second quantized data W2′, the third quantized data W3′, and the fourth quantized data W4′, respectively.
[0061] In this embodiment of the disclosure, since each time step has its own static data, the degree of freedom of the static data is increased, and each static data can be quantized separately, which reduces the difficulty of neural network quantization and improves the quantization accuracy. For example, the quantized data changes from floating point to fixed point data, which reduces the storage and computation requirements of the neural network.
[0062] Step S303: Compress the quantized data from multiple time steps to obtain shared data from multiple time steps.
[0063] The amount of shared data is less than that of static data. Compression processing involves merging or decomposing quantized data to obtain shared data, which can then be used by multiple time steps to reduce the storage requirements of the neural network.
[0064] Step S304: The neural network unit corresponding to the shared data is taken as the target neural network unit.
[0065] In this embodiment, for any neural network unit to be quantized, the neural network unit is expanded in the time dimension to determine the static data of multiple time steps of the neural network unit; the static data of multiple time steps is quantized separately to obtain the corresponding quantized data; the quantized data of multiple time steps is then compressed to obtain shared data of multiple time steps; finally, the neural network unit corresponding to the shared data is taken as the target neural network unit, and the quantized target neural network is determined based on the target neural network unit. This method can adaptively confirm the shared data, increase the degree of freedom of static data, reduce the difficulty of quantization, improve the accuracy of the neural network while meeting the storage requirements of the neural network, and avoid the failure of the neural network due to excessive quantization error.
[0066] In some embodiments, quantized data from multiple time steps are compressed using clustering to reduce the amount of quantized data, thereby reducing the storage space requirements of the neural network.
[0067] In some embodiments, step S302 involves quantizing the static data at multiple time steps to obtain corresponding quantized data, and then retraining the neural network based on the quantized data to improve the accuracy of the neural network. After improving the accuracy of the neural network, step S303 is then performed to reduce the storage space requirements of the neural network.
[0068] Figure 5 A flowchart illustrating a compression process provided in an embodiment of this disclosure. Figure 5 As shown, step S303 involves compressing the quantized data from multiple time steps to obtain shared data from multiple time steps, including:
[0069] Step S501: Set the number of clusters Kn for the nth clustering process.
[0070] Where n≥1, Kn≥2, and n and Kn are integers.
[0071] In some embodiments, the compression process may go through multiple loops, each loop adjusting the number of clusters to cluster quantized data across multiple time steps.
[0072] Step S502: Perform the nth clustering process on the quantized data of multiple time steps to obtain Kn first cluster center data.
[0073] In some embodiments, the K-means clustering algorithm is used to cluster the quantified data to obtain Kn first cluster centers. It should be noted that other clustering algorithms may also be used to cluster the quantified data, and this disclosure does not limit the clustering algorithm.
[0074] Step S503: Determine whether the neural network constructed from the Kn first cluster center data meets the preset termination condition. If the termination condition is met, proceed to step S504; otherwise, proceed to step S501 to set the number of clusters to be processed in the next loop.
[0075] In some embodiments, the termination condition includes at least one of the neural network's storage requirements and the neural network's accuracy requirements.
[0076] Step S504: If the neural network constructed from the Kn first cluster center data satisfies the preset termination condition, the Kn first cluster center data are used as shared data for multiple time steps.
[0077] Figure 6 This is a schematic diagram illustrating the clustering of quantified data according to an embodiment of this disclosure. For example... Figure 6 As shown, clustering is performed on the quantized data at four time steps, specifically on the first quantized data W1′, the second quantized data W2′, the third quantized data W3′, and the fourth quantized data W4′, with a cluster size Kn of 2. After clustering these four data points, two first cluster center values are obtained. Second cluster center value First cluster center value Replace the first quantized data W1′ and the third quantized data W3′, so that the first time step and the third time step share the value of the first cluster center. Second cluster center value Replace the second quantized data W2′ and the fourth quantized data W4′, so that the second time step and the fourth time step share the value of the second cluster center. Based on the values of the first cluster centers Second cluster center value If the constructed neural network meets the preset termination conditions, the value of the first cluster center will be... Second cluster center value This serves as shared data across four time steps. The data is derived from the values of the first cluster centers. Second cluster center value If the constructed neural network does not meet the preset termination condition, return to step S401 and set the number of clusters to be processed in the next loop.
[0078] This disclosure compresses quantized data through clustering, requiring each neural network unit to store only Kn first cluster center data, thus reducing the amount of static data and lowering the storage requirements of the neural network.
[0079] In some embodiments, such as Figure 7 As shown, after performing the nth clustering process on the quantized data from multiple time steps to obtain Kn first cluster center data, step S502 further includes:
[0080] Step S701: Retrain the neural network constructed from the Kn first cluster center data to obtain the neural network after the nth retraining.
[0081] After obtaining the first cluster center data through the nth clustering process, the first cluster center data can be optimized to enhance the accuracy of the neural network. Therefore, this disclosure retrains the neural network constructed from Kn first cluster center data to optimize the first cluster center data, such as optimizing the numerical values of the first cluster centers. Second cluster center value
[0082] Step S702: Quantize the static data of the neural network units in the neural network after the nth retraining at multiple time steps to obtain the corresponding nth quantized data.
[0083] In step S702, the neural network after the nth retraining is quantized again. That is, the static data of multiple time steps of the neural network units in the nth retrained neural network are quantized respectively to obtain the corresponding nth quantized data. The quantization method is the same as that used in step S201, and will not be described again here.
[0084] Step S703: Perform clustering processing on the nth quantization data of multiple time steps to obtain Kn second cluster center data.
[0085] Step S704: If the neural network constructed from the Kn second cluster center data satisfies the preset termination condition, the Kn second cluster center data are used as shared data for multiple time steps.
[0086] For example, in step S702, the static data of the neural network units in the neural network after the nth retraining are quantized at four time steps to obtain the corresponding nth quantized data, namely, the first quantized data W1″, the second quantized data W2″, the third quantized data W3″ and the fourth quantized data W4″.
[0087] In step S703, the first quantized data W1″, the second quantized data W2″, the third quantized data W3″, and the fourth quantized data W4″ are clustered to obtain four second cluster center data, which are the values of the first cluster centers. Second cluster center value
[0088] In step S704, if the neural network constructed from the data of the two second cluster centers satisfies the preset termination condition, the value of the first cluster center is... Second cluster center value As shared data across the four time steps, the first and third time steps share the value of the first cluster center. The second and fourth time steps share the second cluster center values.
[0089] This disclosure improves the accuracy of the neural network by retraining the neural network constructed from the first cluster center data to obtain the retrained neural network after the nth retraining. The static data of the neural network units in the retrained neural network at multiple time steps are quantized to obtain the corresponding quantized data at the nth time step. Then, the quantized data at the nth time step are clustered to obtain Kn second cluster center data. When the neural network constructed from the Kn second cluster center data meets the preset termination condition, the Kn second cluster center data are used as shared data for multiple time steps.
[0090] In this embodiment of the disclosure, if the neural network constructed from Kn first cluster center data does not meet the preset termination condition, the process returns to step S501 for the next compression process, and the number of clusters in the clustering process is adjusted during the next compression process.
[0091] In some embodiments, the termination condition includes the storage requirements of the neural network. In step S503, if the neural network constructed from Kn first cluster center data does not meet the preset storage requirements, the number of clusters in the next clustering process is reduced to decrease the storage requirements of the neural network.
[0092] For example, when the value of the first cluster center... Second cluster center value If the constructed neural network does not meet the preset storage requirements, the number of clusters in the next clustering process is reduced to lower the storage requirements of the neural network.
[0093] In some embodiments, the termination condition includes the accuracy requirement of the neural network. In step S503, if the neural network constructed from Kn first cluster center data does not meet the preset accuracy requirement, the number of clusters in the next clustering process is increased to improve the accuracy of the neural network.
[0094] For example, when the value of the first cluster center... Second cluster center value If the constructed neural network does not meet the preset accuracy requirements, the number of clusters in the next clustering process is increased to improve the accuracy of the neural network.
[0095] In some embodiments, the quantized data of multiple time steps is compressed by decomposition. Step S303, compressing the quantized data of multiple time steps to obtain shared data of multiple time steps, includes:
[0096] The quantized data at each time step is decomposed to obtain shared data, which includes fixed data and sparse data. That is, each quantized data is decomposed into two parts: fixed data and sparse data.
[0097] In this system, fixed data is shared across multiple time steps, while sparse data is stored independently for each time step. Fixed data occupies more storage space, while sparse data occupies less. Although fixed data occupies more storage space, sharing it across multiple time steps saves storage space. The number of sparse matrices is relatively large, but they occupy less storage space. Overall, decomposing quantized data into fixed and sparse data reduces the storage requirements of the neural network.
[0098] In some embodiments, for each quantized data W i The data is decomposed to obtain fixed data W0 and sparse data ΔW. i Where i = 1, 2, ..., m. For example, the quantized data is a matrix, with fixed data W0 as a dense matrix and sparse data ΔW... i It is a sparse matrix. A sparse matrix contains a large number of "0" elements, which reduces storage requirements and can reduce the storage requirements of neural networks.
[0099] When decomposing quantized data, sparse data ΔW can be optimized by training a neural network. i The sparsity of the data is ensured, and a regularization term can be added during training to guarantee the sparsity of the data ΔW. i Sparsity.
[0100] In some embodiments, static data includes weight data and / or bias data. For example, any neural network unit to be quantized is unfolded along the time dimension, the weight data for each time step is determined, and then the weight data for each time step is quantized separately to obtain the corresponding quantized weight data. The quantized weight data is then compressed to obtain shared weight data for multiple time steps, and the neural network corresponding to the shared weight data is used as the target neural network unit. The bias data is the same as the weight data and will not be described further here.
[0101] To better understand the neural network quantization method disclosed herein, the following explanation uses clustering to quantize multiple time steps of a neural network unit as an example.
[0102] Figure 8 A flowchart illustrating the neural network quantization method provided in this embodiment of the disclosure. Figure 8 As shown, neural network quantization methods include;
[0103] Step S801: Expand the neural network unit in the time dimension to determine the static data of multiple time steps of the neural network unit.
[0104] Step S802: Quantize the static data of multiple time steps respectively to obtain the corresponding quantized data.
[0105] Step S803: Set the number of clusters Kn for clustering.
[0106] Step S804: Perform clustering processing on the quantized data of multiple time steps to obtain Kn first cluster center data.
[0107] Step S805: Retrain the neural network constructed from Kn first cluster center data to obtain the retrained neural network. Then, quantize the static data of the neural network units in the retrained neural network at multiple time steps to obtain the corresponding quantized data. Cluster the quantized data at multiple time steps to obtain Kn second cluster center data.
[0108] Step S806: Determine if the neural network constructed from the Kn second cluster center data meets the storage requirements. If it does, proceed to step 807; otherwise, proceed to step S803 and reduce the number of clusters Kn for the next clustering process.
[0109] Step S807: Determine whether the neural network constructed from the Kn second cluster center data meets the accuracy requirements. If it does not meet the accuracy requirements, proceed to step S803 and increase the number of clusters Kn for the next clustering process. If it meets the accuracy requirements, proceed to step S808.
[0110] Step S808: Use the Kn second cluster center data as shared data for multiple time steps, and use the neural network units corresponding to the shared data as target neural network units.
[0111] In some embodiments, after step S804, the process can proceed directly to step S806, where it is determined that the neural network constructed from the Kn first cluster center data meets the storage requirements. That is, the judgment object in step S806 changes from the neural network constructed from the Kn second cluster center data to the neural network constructed from the Kn first cluster center data. The other steps are the same and will not be described again here.
[0112] It is understood that the neural network quantization method provided in this disclosure improves the degree of freedom of static data, reduces the difficulty of quantization, and takes into account both storage resources and the accuracy of the quantized neural network. While meeting the storage requirements of the neural network, it improves the accuracy of the neural network and avoids the failure of the neural network due to excessive quantization error.
[0113] In some embodiments, the neural network is used to perform any one of image processing tasks, speech processing tasks, text processing tasks, and video processing tasks.
[0114] It should be noted that the neural network quantization method and apparatus provided in this disclosure are not only used for the quantization of RNNs, but also for the quantization of spiking neural networks (SNNs) involving time-sharing mechanisms, such as the quantization of SNNs involving sequence data.
[0115] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0116] In addition, this disclosure also provides a neural network quantization device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any of the neural network quantization methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the relevant section of the method and will not be repeated here.
[0117] Figure 9 This is a block diagram of a neural network quantization device provided in an embodiment of the present disclosure.
[0118] like Figure 9 As shown, this disclosure provides a neural network quantization device. The neural network includes at least one neural network unit, and each neural network unit includes multiple time steps. The neural network quantization device 900 includes:
[0119] Quantization module 91 is used to quantize the neural network units to be quantized in the neural network to obtain the target neural network units;
[0120] The determination module 92 is used to determine the quantized target neural network based on the target neural network unit.
[0121] The quantization module 92 includes:
[0122] The unfolding unit 921 is used to unfold any neural network unit to be quantized in the time dimension to determine the static data of the neural network unit at multiple time steps.
[0123] The quantization unit 922 is used to quantize the static data of multiple time steps to obtain the corresponding quantized data.
[0124] Compression unit 923 is used to compress quantized data from multiple time steps to obtain shared data from multiple time steps; wherein the amount of shared data is less than the amount of static data.
[0125] The determining unit 924 is used to identify the neural network unit corresponding to the shared data as the target neural network unit.
[0126] The neural network quantization device provided in this embodiment includes an unfolding unit that unfolds any neural network unit to be quantized along the time dimension to determine the static data of the neural network unit at multiple time steps; a quantization unit that quantizes the static data at multiple time steps to obtain corresponding quantized data; a compression unit that compresses the quantized data at multiple time steps to obtain shared data at multiple time steps; and a determination unit that uses the neural network unit corresponding to the shared data as the target neural network unit. The determination module determines the quantized target neural network based on the target neural network unit. This device can adaptively share data, increase the degree of freedom of static data, reduce the difficulty of quantization, improve the accuracy of the neural network while meeting the storage requirements of the neural network, and avoid the failure of the neural network due to excessive quantization error.
[0127] Figure 10 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0128] Reference Figure 10This disclosure provides an electronic device comprising: at least one processor 1001; a memory 1002 communicatively connected to the at least one processor 1001; and one or more I / O interfaces 1003 connected between the processor 1001 and the memory 1002, configured to enable information interaction between the processor 1001 and the memory 1002; wherein the memory 1002 stores one or more computer programs executable by the at least one processor 1001, and the one or more computer programs are executed by the at least one processor 1001 to enable the at least one processor 1001 to execute the aforementioned neural network quantization method.
[0129] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor / processor core, implements the aforementioned neural network quantization method. The computer-readable storage medium may be volatile or non-volatile.
[0130] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device executes the above-described neural network quantization method.
[0131] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0132] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0133] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0134] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0135] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0136] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0137] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0138] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0140] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A neural network quantization method, characterized in that, The method is applied to a brain-like many-core architecture, the neural network comprises at least one neural network unit, each neural network unit comprises a plurality of time steps, and the method comprises the following steps: respectively quantizing neural network units to be quantized in the neural network to obtain target neural network units; determining a quantized target neural network according to the target neural network units; obtaining a task execution result by using the quantized target neural network units during execution of the neural network, wherein the task comprises any one of an image processing task, a speech processing task, a text processing task and a video processing task; wherein the step of respectively quantizing the neural network units to be quantized in the neural network to obtain the target neural network units comprises: for any neural network unit to be quantized, unfolding the neural network unit in a time dimension to determine static data of a plurality of time steps of the neural network unit, wherein the static data of each time step is independently stored in a storage unit, and each time step corresponds to a computing core; quantizing the static data of the plurality of time steps in a pipeline execution mode to obtain corresponding quantized data; compressing the quantized data of the plurality of time steps to obtain shared data of the plurality of time steps, wherein a data amount of the shared data is less than a data amount of the static data; taking the neural network unit corresponding to the shared data as a target neural network unit.
2. The neural network quantization method of claim 1, wherein, The step of compressing the quantized data of the plurality of time steps to obtain the shared data of the plurality of time steps comprises: setting a cluster group number Kn of an n-th clustering process, n≥1, Kn≥2 and n and Kn are integers; performing an n-th clustering process on the quantized data of the plurality of time steps to obtain Kn first cluster center data; in a case where the neural network constructed by the Kn first cluster center data satisfies a preset termination condition, taking the Kn first cluster center data as shared data of the plurality of time steps.
3. The neural network quantization method of claim 2, wherein, After the step of performing the n-th clustering process on the quantized data of the plurality of time steps to obtain the Kn first cluster center data, the method further comprises the following steps: retraining the neural network constructed by the Kn first cluster center data to obtain an n-th retrained neural network; quantizing the static data of the plurality of time steps of the neural network unit in the n-th retrained neural network to obtain corresponding n-th quantized data; performing a clustering process on the n-th quantized data of the plurality of time steps to obtain Kn second cluster center data; in a case where the neural network constructed by the Kn second cluster center data satisfies a preset termination condition, taking the Kn second cluster center data as shared data of the plurality of time steps.
4. The neural network quantization method of claim 2, wherein, The termination condition comprises a neural network storage requirement; in a case where the neural network constructed by the Kn first cluster center data after the n-th clustering process does not satisfy a preset neural network storage requirement, reducing a cluster group number of a next clustering process.
5. The neural network quantization method of claim 2, wherein, The termination condition comprises a neural network precision requirement; In a case where the neural network constructed by the Kn first clustering center data after the n-th clustering processing does not meet the preset accuracy requirement, the number of clustering groups of the next clustering processing is increased.
6. The neural network quantization method of claim 1, wherein, The compression processing on the quantized data of the plurality of time steps obtains shared data of the plurality of time steps, including: The quantized data of each time step is decomposed to obtain shared data, including fixed data and sparse data; wherein the fixed data is shared data of the plurality of time steps, and the sparse data is independently stored data of each time step.
7. The neural network quantization method of any one of claims 1-6, wherein, The static data includes weight data and / or bias data.
8. A neural network quantization apparatus, characterized by, Applied to a brain-like many-core architecture, the neural network includes at least one neural network unit, each neural network unit includes a plurality of time steps, and the device includes: A quantization module is configured to quantize each neural network unit to be quantized in the neural network to obtain a target neural network unit. A determination module is configured to determine a quantized target neural network according to the target neural network unit, and obtain a task execution result through the quantized target neural network unit during execution of the neural network, wherein the task includes any one of an image processing task, a speech processing task, a text processing task, and a video processing task. The quantization module includes: An unfolding unit is configured to unfold each neural network unit to be quantized in a time dimension to determine static data of a plurality of time steps of the neural network unit, wherein the static data of each time step is independently stored in a storage unit, and each time step corresponds to a computing core. A quantization unit is configured to calculate and quantize the static data of the plurality of time steps in a pipeline execution manner to obtain corresponding quantized data. A compression unit is configured to compress the quantized data of the plurality of time steps to obtain shared data of the plurality of time steps, wherein the data amount of the shared data is less than the data amount of the static data. A determination unit is configured to determine the shared data corresponding to the neural network unit as a target neural network unit.
9. An electronic device, comprising: The device includes: At least one processor; and A memory connected in communication with the at least one processor; wherein The memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to execute the neural network quantization method of any one of claims 1-7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the neural network quantization method of any one of claims 1-7.
Citation Information
Patent Citations
Three-dimensional monitoring system and monitoring method thereof
CN111726562A
Causal modeling method and device for event sequence
CN112069227A