Data fixed-point quantification method and device, computer storage medium and electronic equipment

Through the fixed-point quantization method of data, the decimal point position is changed using preset quantization parameters to realize the change of the fixed-point size and convert it into a decimal number, which solves the problem of inefficient quantization into floating-point numbers in the existing quantization method, and improves the calculation speed and efficiency.

CN119990202APending Publication Date: 2025-05-13INTERNATIONAL INNOVATION CENTER OF TSINGHUA UNIVERSITY SHANGHAI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510071477.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing quantization method, the inverse quantization is used to calculate it again, resulting in low operating efficiency.

Method used

A fixed-point quantization method for data is proposed. By changing the position of the decimal point by preset quantization parameters, the size of the fixed-point number is changed, and the size of the fixed-point number is reduced to a decimal number, and the multiplication, shift and addition operations are directly calculated.

Benefits of technology

It solves the disadvantages of inverse quantization into floating point numbers and then calculates them, improves the calculation speed and efficiency, and is suitable for terminal equipment with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990202A_ABST
    Figure CN119990202A_ABST
Patent Text Reader

Abstract

The invention discloses a data fixed-point quantification method and device, a computer storage medium and electronic equipment. The method comprises the steps that to-be-processed data and preset offset data are acquired; carrying out fixed-point position alignment on the to-be-processed data and the preset offset data, and then adding to obtain first data; and performing shift rounding processing on the first data according to a preset quantization parameter to obtain target output data, the preset quantization parameter being used for representing a data length of a fractional part of the target output data. Therefore, according to the method, the size of the corresponding fixed point number is changed by changing the position of the decimal point based on the preset quantization parameter, quantization of the to-be-processed data is completed, and the to-be-processed data can be restored into the corresponding decimal number; therefore, in a subsequent calculation process, network calculation after quantization can be directly realized through multiplication, shift and addition operations based on a binary form of a fixed-point number, and the defect that calculation is performed after inverse quantization is performed into a floating-point number in a related quantization method is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data fixed-point quantization method, a computer-readable storage medium, a data fixed-point quantization device, and an electronic device. Background Art

[0002] As these neural network models are continuously optimized and upgraded, more and more networks are deployed in terminal devices. In order to support the operation of complex neural network models, terminal devices need to have stronger computing power, higher energy efficiency and more optimized algorithm framework. Quantizing neural networks into integer calculations is an effective optimization method that can improve the performance and efficiency of the model in a resource-limited environment, while adapting to the deployment requirements of various terminal devices.

[0003] There are various quantization tools and methods for neural networks, such as PyTorch, TensorFlow Lite, and Cafferistretto. However, the quantization methods in related technologies have the disadvantage of dequantizing to floating-point numbers and then performing calculations, which affects the operating efficiency. Summary of the invention

[0004] The present application aims to solve at least one of the technical problems in the related art to a certain extent. To this end, the first purpose of the present application is to propose a data fixed-point quantization method, which changes the size of the corresponding fixed-point number by changing the position of the decimal point based on a preset quantization parameter, and can restore it to the corresponding decimal number, so that in the subsequent calculation process, the quantized network calculation can be directly realized by multiplication, shift, and addition operations based on the binary form of the fixed-point number, thereby solving the drawback of the related quantization method of dequantizing to a floating-point number and then calculating.

[0005] A second object of the present application is to provide a computer-readable storage medium.

[0006] The third objective of the present application is to provide a data fixed-point quantization device.

[0007] The fourth objective of the present application is to provide an electronic device.

[0008] To achieve the above-mentioned purpose, the first aspect of the present application proposes a data fixed-point quantization method, which includes: obtaining data to be processed and preset bias data; aligning the data to be processed and the preset bias data at fixed points and adding them to obtain first data; shifting and rounding the first data according to a preset quantization parameter to obtain target output data, wherein the preset quantization parameter is used to characterize the data length of the decimal part of the target output data.

[0009] According to the data fixed-point quantization method of the embodiment of the present application, the data to be processed and the preset bias data are obtained, the data to be processed and the preset bias data are aligned at fixed points and then added to obtain the first data, and the first data is shifted and rounded according to the preset quantization parameter to obtain the target output data, wherein the preset quantization parameter is used to characterize the data length of the decimal part of the target output data. Thus, the method changes the size of the corresponding fixed-point number by changing the position of the decimal point based on the preset quantization parameter, and can restore it to the corresponding decimal number, so that in the subsequent calculation process, the quantized network calculation can be directly realized by multiplication, shift, and addition operations based on the binary form of the fixed-point number, thereby solving the drawback of the related quantization method of dequantizing to a floating-point number and then calculating.

[0010] In addition, the data fixed-point quantization method according to the above embodiment of the present application may also have the following additional technical features:

[0011] According to one embodiment of the present application, a shift and rounding process is performed on first data according to a preset quantization parameter, including: adding one to the preset quantization parameter to obtain a target shift parameter; performing a first shift on the first data according to the target shift parameter to obtain second data, wherein the target shift parameter is used to characterize the data length of the decimal part of the second data; rounding the last bit of the second data according to the data type of the second data and the value of the last bit of the second data; and shifting the rounded second data for a second time to obtain target output data.

[0012] According to one embodiment of the present application, the last bit of the second data is rounded according to the data type of the second data and the value of the last bit of the second data, including: when the second data is a positive number and the last bit of the second data is 1, the last bit of the second data is rounded; when the second data is a negative number, the last bit of the second data is rounded according to a preset rounding mode and the value of the last bit of the second data.

[0013] According to an embodiment of the present application, the data fixed-point quantization method further includes: performing data overflow processing on the target output data based on a preset data range.

[0014] According to one embodiment of the present application, applied to a target neural network model, the data fixed-point quantization method also includes: determining a preset quantization parameter based on a parameter distribution range of different layers of the target neural network model, wherein the parameter distribution range is positively correlated with the preset quantization parameter.

[0015] According to an embodiment of the present application, the data fixed-point quantization method further includes: performing activation processing on the target output data based on a preset activation function.

[0016] According to one embodiment of the present application, the data fixed-point quantization method also includes: determining a preset network model in combination with a target neural network model; and determining a hardware circuit model based on the trained preset network model.

[0017] To achieve the above-mentioned purpose, a second aspect of the present application proposes a computer-readable storage medium on which a data fixed-point quantization program is stored. When the data fixed-point quantization program is executed by a processor, the above-mentioned data fixed-point quantization method is implemented.

[0018] According to the computer-readable storage medium of the embodiment of the present application, the above-mentioned data fixed-point quantization method is implemented when the data fixed-point quantization program stored thereon is executed by the processor. Based on the above-mentioned data fixed-point quantization method, the disadvantage of dequantizing to floating-point numbers and then calculating in related quantization methods is solved, thereby improving the calculation speed.

[0019] To achieve the above-mentioned purpose, the third aspect of the present application proposes a data fixed-point quantization device, which includes: an acquisition module, used to acquire the data to be processed and the preset bias data; a calculation module, used to align the data to be processed and the preset bias data at fixed points and then add them to obtain first data; a processing module, used to shift and round the first data according to a preset quantization parameter to obtain target output data, wherein the preset quantization parameter is used to characterize the data length of the decimal part of the target output data.

[0020] According to the data fixed-point quantization device of the embodiment of the present application, the acquisition module acquires the data to be processed and the preset bias data, and the calculation module performs fixed-point position alignment on the data to be processed and the preset bias data and then adds them to obtain the first data, and the processing module performs shift and rounding processing on the first data according to the preset quantization parameter to obtain the target output data, wherein the preset quantization parameter is used to characterize the data length of the decimal part of the target output data. Thus, the device changes the size of the corresponding fixed-point number by changing the position of the decimal point based on the preset quantization parameter, and can restore it to the corresponding decimal number, so that in the subsequent calculation process, the quantized network calculation can be directly realized by multiplication, shift, and addition operations based on the binary form of the fixed-point number, thereby solving the drawbacks of the related quantization method of dequantizing to a floating-point number and then calculating.

[0021] To achieve the above-mentioned objectives, the fourth aspect embodiment of the present application proposes an electronic device, including a memory, a processor, and a data fixed-point quantization program stored in the memory and executable on the processor. When the processor executes the data fixed-point quantization program, the above-mentioned data fixed-point quantization method is implemented.

[0022] According to the electronic device of the embodiment of the present application, when the processor executes the data fixed-point quantization program, the above-mentioned data fixed-point quantization method is implemented. Based on the above-mentioned data fixed-point quantization method, the disadvantage of dequantizing to floating-point numbers and then calculating in related quantization methods is solved, thereby improving the operation speed.

[0023] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flow chart of a data fixed-point quantization method according to an embodiment of the present application;

[0025] Figure 2 A schematic diagram of a data fixed-point quantization method according to a specific embodiment of the present application;

[0026] Figure 3 A flow chart is provided for constructing a circuit according to one embodiment of the present application;

[0027] Figure 4 is a block diagram of a data fixed-point quantization circuit according to an embodiment of the present application;

[0028] Figure 5 It is a flowchart of a data fixed-point quantization method according to a specific embodiment of the present application;

[0029] Figure 6 is a connection diagram of a data fixed-point quantization device according to an embodiment of the present application;

[0030] Figure 7 Schematic diagram of a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0031] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0032] The following describes the data fixed-point quantization method, computer-readable storage medium, data fixed-point quantization device and electronic device proposed in the embodiments of the present application with reference to the accompanying drawings.

[0033] Figure 1 The figure is a flow chart of a data fixed-point quantization method according to an embodiment of the present application.

[0034] like Figure 1 As shown, the data fixed-point quantization method of the embodiment of the present application may include:

[0035] S1, obtaining data to be processed and preset bias data;

[0036] S2, performing fixed-point position alignment on the data to be processed and the preset offset data and adding them together to obtain first data;

[0037] S3, performing shift and rounding processing on the first data according to a preset quantization parameter to obtain target output data, wherein the preset quantization parameter is used to characterize the data length of the decimal part of the target output data.

[0038] Specifically, Figure 2 For example, the data to be processed is data mac, and the preset bias data is data bias, wherein data mac and data bias are both 8 bits, fl of data mac=8, that is, the data length of the decimal part of data mac is 8 bits, and fl of data bias is 3, that is, the data length of the decimal part of data bias is 3 bits. The vertical dotted lines in the accompanying drawings are used to indicate the fixed point positions of each data.

[0039] First, shift the data bias left by 5 bits to get the data bias of fl=8, align the data bias with the fixed-point position of the data mac, and add the data bias of fl=8 to the data mac of fl=8 to get the first data, i.e., data tmp of fl=8. This step essentially adds a bias parameter to the data to be processed. For example, when the quantization method is applied to a neural network, the data to be processed is the intermediate result of multiplication and accumulation in the convolution layer of the neural network model. The bias parameter can help the model to make more refined adjustments to the data, so that the neural network model can better fit the data. Among them, the preset bias data can be set according to the actual situation.

[0040] Then, the first data is shifted and rounded according to the preset quantization parameter. Assuming that the preset quantization parameter is fl=3, the first data obtained by addition, that is, data tmp with fl=8, is shifted and rounded to obtain data out with fl=3. The preset quantization parameter can be set according to actual conditions.

[0041] This embodiment can effectively change the size of the corresponding fixed-point number by changing the position of the decimal point fl through the preset quantization parameter, and can restore it to the corresponding decimal number, which means that the quantized network calculation can be directly implemented by multiplication, shift, and addition operations based on the binary form of the fixed-point number, thereby avoiding the disadvantages of other quantization methods of dequantizing to floating-point numbers and then performing calculations.

[0042] In one embodiment of the present application, shift and rounding processing is performed on first data according to a preset quantization parameter, including: adding one to the preset quantization parameter to obtain a target shift parameter; performing a first shift on the first data according to the target shift parameter to obtain second data, wherein the target shift parameter is used to characterize the data length of the decimal part of the second data; rounding the last bit of the second data according to the data type of the second data and the value of the last bit of the second data; and shifting the rounded second data for a second time to obtain target output data.

[0043] Specifically, continue to combine Figure 2 As shown, taking the data to be processed as data mac, the preset bias data as data bias, and the preset quantization parameter as fl=3 as an example, the target shift parameter is fl=4.

[0044] After calculating and obtaining the first data, i.e., data tmp with fl=8, the data tmp is shifted for the first time to the position of fl=4. In order to determine the rounding mode, the first shift is actually 1 bit less than the final required shift, i.e., the second data is one bit more than the target output data to be obtained. Then, the rounding processing mode for the last bit of the second data is determined according to the data type of the second data and the value of the last bit of the second data. For example, when the second data is positive, when the last bit of the second data is 1, a carry processing is performed; otherwise, no carry processing is performed. After the second data is rounded, the second data is shifted for the second time, i.e., 1 bit, to obtain the final required fixed-point position, and the data out with fl=3 is output.

[0045] In one embodiment of the present application, the last bit of the second data is rounded according to the data type of the second data and the value of the last bit of the second data, including: when the second data is a positive number and the last bit of the second data is 1, the last bit of the second data is rounded; when the second data is a negative number, the last bit of the second data is rounded according to a preset rounding mode and the value of the last bit of the second data.

[0046] That is to say, if the second data is a positive number, as long as the lowest bit is 1, it must be a carry, so an add operation is required; otherwise, no carry operation is performed. If the second data is a negative number, it is necessary to further determine whether to carry based on the preset rounding mode and the low-order data before the shift. Among them, the preset rounding mode can be configured as symmetric rounding or asymmetric rounding.

[0047] In one embodiment of the present application, the data fixed-point quantization method further includes: performing data overflow processing on the target output data based on a preset data range.

[0048] That is to say, when the data calculation is completed, overflow judgment is required, and the value exceeding the preset data range is represented by the maximum value and the minimum value. That is, positive overflow is represented by the maximum value of the preset data range, and negative overflow is represented by the minimum value of the preset data range, to ensure that the final target output data result is within the valid range and is finally truncated to the target data length.

[0049] exist Figure 2 In the embodiment shown, the data length is 8 bits, where the fl of the data mac is 8, which means that the 8-bit data is all decimals; and the fl of the data bias is 3, which means that the decimal part is 3 bits and the integer part is 5 bits. After the decimal points are aligned, the data mac and the data bias are added together, and finally truncated to 8 bits of data and output to the next network layer for calculation.

[0050] Data quantization can improve storage and computing efficiency, but it may also introduce precision loss problems. This precision loss may cause the performance of the neural network model to decline, especially in certain tasks and applications, when the accuracy requirements of the neural network model are high. In addition, in terminal applications, since only a single rounding mode quantization method is supported, the network precision loss will further increase as the network calculation propagates layer by layer. To this end, in one embodiment of the present application, the data fixed-point quantization method is applied to the target neural network model, and further includes: determining a preset quantization parameter based on the parameter distribution range of different layers of the target neural network model, wherein the parameter distribution range is positively correlated with the preset quantization parameter.

[0051] Specifically, the weights and biases of the neural network model can effectively compress the model size, and the inter-layer feature mapping results in the network calculation process also need to be quantized. The distribution ranges of the network's convolution parameters, full connection parameters, and intermediate layer result data are very different. Most of the network's parameters are between 2 -10 To 2 10 The distribution range is very wide. Among them, the result data of the intermediate layer as the result of the convolution operation can even reach 2 14 Aiming at the problem of large differences in data distribution in different network structures, the network is quantized and optimized based on the dynamic fixed-point quantization method combined with the idea of ​​fixed-point quantization.

[0052] This embodiment finds a suitable preset quantization parameter fl for each layer for quantization according to the parameter distribution range of different layers of the neural network model, which can achieve a better quantization effect. In the target neural network model, each forward feedback network layer, convolution layer and fully connected layer has its own corresponding quantization parameter, and it is necessary to configure the corresponding parameter, i.e., the preset quantization parameter, before each quantization calculation. In addition, the parameter may also include a preset data range, which is not specifically limited.

[0053] In one embodiment of the present application, the data fixed-point quantization method further includes: performing activation processing on the target output data based on a preset activation function.

[0054] In other words, an additional simple ReLU (Rectified Linear Unit) activation function is added, and the calculation of ReLU and Leaky ReLU (Leaky Rectified Linear Unit) is realized by adding one more selector and shift.

[0055] When deploying and executing quantized neural network models on terminal devices, the terminal's CPU (Central Processing Unit) is currently used for calculations due to development cycle and hardware resource limitations. Although using CPU to deploy neural networks in terminal devices is a common approach, it also has some disadvantages. The computing power of the CPU is relatively limited, especially when processing complex neural network models, the inference speed may be slow and cannot meet real-time requirements. At the same time, when performing neural network calculations, the CPU cannot fully utilize its multi-core architecture and other features, resulting in low resource utilization. The design goal of the CPU is not specifically for neural network calculations, so it will be bottlenecked when performing neural network inference, resulting in a slower inference speed.

[0056] For applications with high performance requirements and sensitive to power consumption, it is usually considered to use a dedicated neural network acceleration circuit to replace the CPU for neural network calculations to obtain better performance and efficiency. However, the development of a dedicated neural network acceleration circuit is also a complex and challenging task, involving multiple technical and design considerations. Comprehensive consideration and optimization are required in terms of architecture design, software and hardware co-design, performance optimization, flexibility, etc. to achieve efficient and high-performance neural network calculation acceleration.

[0057] In one embodiment of the present application, the data fixed-point quantization method also includes: determining a preset network model in combination with a target neural network model; and determining a hardware circuit model based on the trained preset network model.

[0058] That is to say, in view of the complexity and debugging difficulties of the current design and deployment of neural network acceleration circuits, this embodiment proposes a fixed-point quantization method, modeling, and software-hardware co-design paradigm, which can be used to quickly iterate and optimize the original fixed-point quantization algorithm to obtain a fixed-point quantization circuit implementation. Figure 3 As shown, the process of establishing the hardware circuit may include the following steps:

[0059] S201, network analysis: determine the network model and determine the quantization method. In order to determine the quantized network model based on the determined network model and quantization method, for example, the network model is a Resnet network and the above-mentioned dynamic fixed-point quantization method.

[0060] S202, modeling implementation: using C / C++ language to model the quantization network.

[0061] Specifically, after training the quantized network model, and using the quantized network model that meets the requirements, high-level modeling can be implemented using C / C++ language. This process requires continuous iterations, and ultimately a modeling model of the simulated hardware circuit, namely a hardware circuit model, is obtained to guide subsequent hardware circuit design.

[0062] S203, Circuit Implementation: Design and Verification of Dynamic Fixed-Point Quantization Circuit.

[0063] This stage is based on the hardware circuit model and uses hardware description languages ​​such as Verilog to perform RTL (Register-Transfer Level, an abstract level of digital circuit design) design, simulation and synthesis of the quantization circuit.

[0064] S204, cross-validation: software and hardware collaboratively iterate to obtain the final implementation.

[0065] This phase is a cross-validation of software and hardware. With high-level modeling as the reference model, the circuit design is continuously optimized and iterated to obtain a circuit implementation that meets the functional and performance requirements. Specifically, the modeling model of the analog hardware circuit can be adjusted according to the hardware circuit, and then the hardware design can be carried out based on the adjusted modeling model, which effectively improves the circuit energy efficiency and computing efficiency.

[0066] Therefore, this embodiment takes into account the complexity of designing neural network circuits and the difficulty of debugging when deployed on hardware, and proposes an efficient method and process for hardware accelerated circuit modeling and design. Through modeling, designers can perform high-level modeling and verification on the server. This debugging method is more efficient than circuit simulation and verification, and helps to discover and solve problems before the actual circuit design, thereby reducing the cost of prototype design and testing. Through modeling, designers can perform detailed analysis and optimization of circuits to improve performance, reduce power consumption, or improve other indicators, thereby designing more efficient circuits. At the same time, the implementation of high-level modeling can also serve as a reference model for circuit design, which is conducive to rapid problem location and debugging iterations.

[0067] In addition, this embodiment can effectively improve circuit energy efficiency, reduce power consumption and improve computing efficiency by executing quantized neural network algorithms in a software-hardware collaborative manner. By adopting dynamic quantization technology and acceleration design, the circuit can reduce the demand for hardware resources while ensuring computing performance. This helps to reduce hardware costs, improve the utilization of hardware resources, and make the circuit more suitable for scenarios such as resource-constrained embedded devices.

[0068] The circuit may include configuration modules, shift alignment, adders, rounding judgment, overflow judgment, ReLU type activation function calculation, and registers. Figure 4 As shown, the circuit is interconnected with the processor system through a standard bus through a configuration register to perform information configuration and state reading operations to configure the corresponding parameters before each quantization calculation, wherein the configuration register is composed of a multiplexer and multiple registers; the calculation path is performed in a pipeline manner, divided into two stages of pipelines, the first stage is mainly for alignment and addition, and the second stage is mainly for rounding and overflow judgment. Specifically, when the configuration is completed, the input data is aligned by shifting and then added together, saved by registers, and then rounding and overflow judgment are performed to obtain the target output data within the preset data range. By adding a multiplexer and shifting, the calculation of ReLU and Leaky ReLU is realized, and finally the calculated result is saved through a register for subsequent calculations.

[0069] As a specific embodiment of the present application, continuing to apply the target neural network model as an example, the data fixed-point quantization method can be as follows Figure 5 As shown, the following steps are included:

[0070] S101, obtaining data to be processed and preset bias data.

[0071] S102, performing fixed-point position alignment on the data to be processed and the preset offset data and then adding them together to obtain first data.

[0072] S103, determining a preset quantization parameter according to the parameter distribution range of different layers of the target neural network model.

[0073] S104, adding one to the preset quantization parameter to obtain a target shift parameter.

[0074] S105, performing a first shift on the first data according to the target shift parameter to obtain second data.

[0075] S106, determine whether the second data is a positive number. If yes, execute step S107; if no, execute step S120.

[0076] S107, determine whether the last bit of the second data is 1. If yes, execute step S108; otherwise, execute step S109.

[0077] S108, performing carry processing on the last bit of the second data.

[0078] S109, no carry processing is performed on the last bit of the second data.

[0079] S110, according to a preset rounding mode and the last bit of the second data.

[0080] S111, performing a second shift on the rounded second data to obtain target output data.

[0081] S112, performing activation processing on the target output data based on a preset activation function for use in neural network calculations.

[0082] Therefore, this data quantization method can realize neural network quantization. Converting neural network parameters from floating point numbers to integers can significantly reduce the size of the model, thereby reducing storage and transmission costs, speeding up reasoning, and reducing power consumption, so that terminal devices consume less energy and reduce the model's memory usage to a certain extent, so as to achieve the purpose of optimizing the neural network model.

[0083] In summary, according to the data fixed-point quantization method of the embodiment of the present application, the data to be processed and the preset bias data are obtained, the data to be processed and the preset bias data are aligned at fixed points and then added to obtain the first data, and the first data is shifted and rounded according to the preset quantization parameter to obtain the target output data, wherein the preset quantization parameter is used to characterize the data length of the decimal part of the target output data. Therefore, the method changes the size of the corresponding fixed-point number by changing the position of the decimal point based on the preset quantization parameter, and can restore it to the corresponding decimal number, so that in the subsequent calculation process, the quantized network calculation can be directly realized by multiplication, shift, and addition operations based on the binary form of the fixed-point number, thereby solving the drawbacks of the related quantization method of dequantizing to floating-point numbers and then calculating.

[0084] Corresponding to the above embodiments, the present application also proposes a computer-readable storage medium.

[0085] The computer-readable storage medium of the embodiment of the present application stores a data fixed-point quantization program, which implements the above-mentioned data fixed-point quantization method when executed by a processor.

[0086] According to the computer-readable storage medium of the embodiment of the present application, the above-mentioned data fixed-point quantization method is implemented when the data fixed-point quantization program stored thereon is executed by the processor. Based on the above-mentioned data fixed-point quantization method, the disadvantage of dequantizing to floating-point numbers and then calculating in related quantization methods is solved, and the running speed is improved.

[0087] Corresponding to the above embodiments, the present application also proposes a data fixed-point quantization device.

[0088] like Figure 6 As shown, the data fixed-point quantization device of the embodiment of the present application may include: an acquisition module 10, a calculation module 20 and a processing module 30.

[0089] The acquisition module 10 is used to acquire the data to be processed and the preset bias data. The calculation module 20 is used to perform fixed-point position alignment on the data to be processed and the preset bias data and then add them to obtain the first data. The processing module 30 is used to perform shift rounding processing on the first data according to the preset quantization parameter to obtain the target output data, wherein the preset quantization parameter is used to characterize the data length of the decimal part of the target output data.

[0090] According to one embodiment of the present application, the processing module 30 performs shift and rounding processing on the first data according to a preset quantization parameter, specifically for: adding one to the preset quantization parameter to obtain a target shift parameter; performing a first shift on the first data according to the target shift parameter to obtain second data, wherein the target shift parameter is used to characterize the data length of the decimal part of the second data; rounding the last bit of the second data according to the data type of the second data and the value of the last bit of the second data; and shifting the rounded second data for a second time to obtain target output data.

[0091] According to one embodiment of the present application, the processing module 30 rounds the last bit of the second data according to the data type of the second data and the value of the last bit of the second data, and is specifically used for: when the second data is a positive number and the last bit of the second data is 1, performing a carry process on the last bit of the second data; when the second data is a negative number, rounding the value of the last bit of the second data according to a preset rounding mode and the value of the last bit of the second data.

[0092] According to an embodiment of the present application, the processing module 30 is further configured to: perform data overflow processing on the target output data based on a preset data range.

[0093] According to one embodiment of the present application, applied to the target neural network model, the processing module 30 is also used to: determine the preset quantization parameter according to the parameter distribution range of different layers of the target neural network model, wherein the parameter distribution range is positively correlated with the preset quantization parameter.

[0094] According to an embodiment of the present application, the processing module 30 is further used to: perform activation processing on the target output data based on a preset activation function.

[0095] According to one embodiment of the present application, the data fixed-point quantization device also includes a circuit design module for determining a preset network model in combination with a target neural network model, and determining a hardware circuit model based on the trained preset network model.

[0096] It should be noted that for details not disclosed in the data fixed-point quantization device of the embodiment of the present application, please refer to the details disclosed in the data fixed-point quantization method of the above embodiment of the present application, and the details will not be repeated here.

[0097] According to the data fixed-point quantization device of the embodiment of the present application, the acquisition module acquires the data to be processed and the preset bias data, and the calculation module performs fixed-point position alignment on the data to be processed and the preset bias data and then adds them to obtain the first data, and the processing module performs shift and rounding processing on the first data according to the preset quantization parameter to obtain the target output data, wherein the preset quantization parameter is used to characterize the data length of the decimal part of the target output data. Thus, the device changes the size of the corresponding fixed-point number by changing the position of the decimal point based on the preset quantization parameter, and can restore it to the corresponding decimal number, so that in the subsequent calculation process, the quantized network calculation can be directly realized by multiplication, shift, and addition operations based on the binary form of the fixed-point number, thereby solving the drawbacks of the related quantization method of dequantizing to a floating-point number and then calculating.

[0098] Corresponding to the above embodiment, the present application also proposes an electronic device.

[0099] like Figure 7 As shown, the electronic device 100 of the embodiment of the present application includes a memory 110, a processor 120, and a data fixed-point quantization program stored in the memory 110 and executable on the processor 120. When the processor 110 executes the data fixed-point quantization program, the above-mentioned data fixed-point quantization method is implemented. The electronic device 100 includes but is not limited to smart phones, smart home devices, wearable devices, etc., which integrate advanced neural network models to achieve more intelligent and personalized services. The development of this edge computing makes artificial intelligence applications more popular and convenient, while also reducing dependence on cloud computing resources and improving real-time performance and privacy protection.

[0100] According to the electronic device of the embodiment of the present application, when the processor executes the data fixed-point quantization program, the above-mentioned data fixed-point quantization method is implemented. Based on the above-mentioned data fixed-point quantization method, the disadvantage of dequantizing to floating-point numbers and then calculating in related quantization methods is solved, and the operating speed is improved.

[0101] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0102] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0103] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0104] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0105] In this application, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0106] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A data fixed-point quantization method, characterized in that: The method comprises: Obtaining data to be processed and preset bias data; Performing fixed-point position alignment on the data to be processed and the preset offset data and then adding them together to obtain first data; The first data is shifted and rounded according to a preset quantization parameter to obtain target output data, wherein the preset quantization parameter is used to characterize the data length of the decimal part of the target output data.

2. The data fixed-point quantization method according to claim 1, characterized in that: The performing shift and rounding processing on the first data according to a preset quantization parameter includes: Adding one to the preset quantization parameter to obtain a target shift parameter; Performing a first shift on the first data according to the target shift parameter to obtain second data, wherein the target shift parameter is used to represent the data length of the decimal part of the second data; rounding the last digit of the second data according to the data type of the second data and the value of the last digit of the second data; The second data after the rounding process is shifted for the second time to obtain the target output data.

3. The data fixed-point quantization method according to claim 2, characterized in that: The rounding the last bit of the second data according to the data type of the second data and the value of the last bit of the second data includes: When the second data is a positive number and the last bit of the second data is 1, performing a carry process on the last bit of the second data; When the second data is a negative number, the value of the last digit of the second data is rounded according to a preset rounding mode and the value of the last digit of the second data.

4. The data fixed-point quantization method according to claim 1, characterized in that: The method further comprises: The target output data is subjected to data overflow processing based on a preset data range.

5. The data fixed-point quantization method according to claim 1, characterized in that: Applied to the target neural network model, the method further comprises: The preset quantization parameter is determined according to the parameter distribution range of different layers of the target neural network model, wherein the parameter distribution range is positively correlated with the preset quantization parameter.

6. The data fixed-point quantization method according to claim 5, characterized in that: The method further comprises: The target output data is activated based on a preset activation function.

7. The data fixed-point quantization method according to claim 6, characterized in that: The method further comprises: Determining a preset network model in combination with the target neural network model; The hardware circuit model is determined based on the trained preset network model.

8. A computer-readable storage medium, characterized in that: A data fixed-point quantization program is stored thereon, and when the data fixed-point quantization program is executed by a processor, the data fixed-point quantization method according to any one of claims 1-7 is implemented.

9. A data fixed-point quantization device, characterized in that: The device comprises: An acquisition module, used to acquire data to be processed and preset bias data; A calculation module, used for performing fixed-point position alignment on the data to be processed and the preset bias data and then adding them together to obtain first data; A processing module is used to perform shift and rounding processing on the first data according to a preset quantization parameter to obtain target output data, wherein the preset quantization parameter is used to characterize the data length of the decimal part of the target output data.

10. An electronic device, characterized in that: The invention comprises a memory, a processor and a data fixed-point quantization program stored in the memory and executable on the processor. When the processor executes the data fixed-point quantization program, the data fixed-point quantization method according to any one of claims 1 to 7 is implemented.