Neural network quantitative deployment method, device, equipment, medium and program product

By performing network splitting and layer fusion of deep learning models on an embedded platform and generating fixed-point convolution IP cores, the problem of difficult deployment of deep learning models on embedded devices in the existing technology is solved, and efficient inference speed and low-power calculations are achieved.

CN119939126AInactive Publication Date: 2025-05-06UNIV OF SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510428805.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing deep learning models are difficult to deploy on resource-constrained embedded devices, with high computing resources, high power consumption and difficult to improve inference speed.

Method used

Fixed-point quantization is achieved through collaboratively on the processing system side and the programmable logic side, including network splitting and layer fusion of deep learning models, generation of fixed-point convolution IP cores, performing convolutional inference, and completing target data output in combination with floating-point operations.

Benefits of technology

It effectively reduces the storage space and computing resource requirements of the model, improves the speed and efficiency of convolutional inference, and ensures the accuracy of the model output results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939126A_ABST
    Figure CN119939126A_ABST
Patent Text Reader

Abstract

The invention provides a neural network quantitative deployment method which can be applied to the technical field of deep learning. The method is applied to an embedded platform, the embedded platform comprises a programmable logic side and a processing system side, and the method comprises the following steps: loading a target deep learning model on the processing system side; preprocessing the input data; transmitting the preprocessed data to an editable logic side, and calling a fixed-point convolution IP core to carry out convolution reasoning so as to obtain feature map data; the feature map data is transmitted back to the processing system side, floating point calculation operation is completed by the processing system side so as to output target data, the target deep learning model is obtained through fixed-point quantization, and the fixed-point quantization is achieved based on cooperation of the programmable logic side and the processing system side. The invention further provides a neural network quantitative deployment device and equipment, a storage medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of deep learning technology, and more specifically to a neural network quantitative deployment method, device, equipment, medium and program product. Background Art

[0002] In the field of computer vision and embedded systems, SLAM (simultaneous localization and mapping) technology is widely used in mobile robots, drone mapping, autonomous driving and other fields. With the development of deep learning, feature extraction methods based on convolutional neural networks are mainly applied to SLAM systems, which significantly improves the accuracy and robustness of feature detection and description. However, these deep learning models are usually based on floating-point operations, which consume a lot of computing resources and are difficult to deploy directly on resource-constrained embedded devices, such as FPGA field-programmable gate arrays (FPGAs). In addition, the existing quantization deployment methods are not sufficient in terms of fixed-point quantization and network structure optimization, resulting in more hardware resources, high power consumption and difficulty in improving the inference speed. Summary of the invention

[0003] In view of the above problems, the present invention provides a neural network quantization deployment method, device, equipment, medium and program product that improves reasoning speed and reduces power consumption.

[0004] According to a first aspect of the present invention, a neural network quantization deployment method is provided, which is applied to an embedded platform, wherein the embedded platform includes a programmable logic side and a processing system side, including: loading a target deep learning model on the processing system side; preprocessing input data; transmitting the preprocessed data to the editable logic side, calling a fixed-point convolution IP core to perform convolution inference to obtain feature map data; transmitting the feature map data back to the processing system side, and completing floating-point operations on the processing system side to output target data, wherein the target deep learning model is obtained through fixed-point quantization, and the fixed-point quantization is implemented collaboratively based on the programmable logic side and the processing system side.

[0005] According to an embodiment of the present invention, fixed-point quantization based on the collaboration between the programmable logic side and the processing system side includes: preprocessing the deep learning model, wherein the preprocessing includes network splitting and network layer fusion; and fixed-point quantization of the network parameters of the preprocessed deep learning model based on the collaboration between the programmable logic side and the processing system side.

[0006] According to an embodiment of the present invention, network splitting and network layer fusion specifically include: implementing shared encoder and partial decoder convolution layers on the programmable logic side, and implementing floating-point operations on the processing system side; and fusing the batch normalization layer and the convolution layer, and merging the activation function with the truncation of the quantization range.

[0007] According to an embodiment of the present invention, the network parameters of the pre-processed deep learning model are fixed-point quantized based on the programmable logic side and the processing system side, including: using a deep learning framework to train the neural network model and save the model weight parameters; offline counting the maximum activation value of each layer to determine the range of the quantization parameter; fusing the batch normalization layer and the convolution layer of each layer to determine the fused convolution weight; calculating the quantization parameters of each layer based on the quantization strategy; merging the convolution layer, the batch normalization layer and the activation layer, and generating a fixed-point convolution IP core according to the quantization parameters; integrating the fixed-point convolution IP core on the programmable logic side; and completing the descriptor interpolation and feature point decoding operations on the processing system side.

[0008] A second aspect of the present invention provides a neural network quantization deployment device, which is applied to an embedded platform, wherein the embedded platform includes a programmable logic side and a processing system side, and the programmable logic side pre-integrates a fixed-point convolution IP core, including:

[0009] A quantized model loading module, used to load the target deep learning model on the processing system side;

[0010] An input data preprocessing module, used for preprocessing input data;

[0011] The convolution inference module is used to transfer the preprocessed data to the editable logic side and call the fixed-point convolution IP core for convolution inference to obtain feature map data;

[0012] A computing module is used to transmit the feature map data back to the processing system side, and the processing system side performs floating-point computing operations to output target data.

[0013] Among them, the target deep learning model is obtained by fixed-point quantization after network splitting, and the fixed-point quantization is implemented based on the collaboration between the programmable logic side and the processing system side.

[0014] According to an embodiment of the present invention, the device further includes:

[0015] A model preprocessing module, used to preprocess the deep learning model, wherein the preprocessing includes network splitting and network layer fusion; and

[0016] The fixed-point quantization module is used to perform fixed-point quantization on the network parameters of the preprocessed deep learning model based on the collaboration between the programmable logic side and the processing system side.

[0017] According to an embodiment of the present invention, the model preprocessing module includes a network splitting submodule and a network layer fusion submodule.

[0018] a network splitting submodule for implementing shared encoder and partial decoder convolutional layers on the programmable logic side and floating point operations on the processing system side; and

[0019] The network layer fusion submodule is used to fuse the batch normalization layer and the convolution layer, and to combine the activation function with the truncation of the quantization range.

[0020] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0021] The fourth aspect of the present invention further provides a computer-readable storage medium on which a computer program or instruction is stored, and the steps of the above method are implemented when the above computer program or instruction is executed by a processor.

[0022] The fifth aspect of the present invention also provides a computer program product, including a computer program or instructions, which implement the steps of the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0024] Figure 1 A flowchart of a neural network quantization deployment method according to an embodiment of the present invention is schematically shown;

[0025] Figure 2A A schematic diagram of a deep learning model network structure according to an embodiment of the present invention is shown;

[0026] Figure 2B A schematic diagram of the structure of a deep learning model network after splitting according to an embodiment of the present invention is shown;

[0027] Figure 3 A flowchart of a neural network quantization deployment method according to another embodiment of the present invention is schematically shown;

[0028] Figure 4 A flowchart of a neural network quantization deployment method according to another embodiment of the present invention is schematically shown;

[0029] Figure 5A schematic diagram of a structure of a neural network quantization deployment device according to an embodiment of the present invention is shown;

[0030] Figure 6 A block diagram of an electronic device suitable for implementing a neural network quantization deployment method according to an embodiment of the present invention is schematically shown. DETAILED DESCRIPTION

[0031] Below, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of concepts of the present invention.

[0032] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0033] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0034] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0035] First, the terms appearing in the embodiments of the present invention are explained:

[0036] SLAM algorithm: A simultaneous localization and mapping algorithm used by autonomous systems such as robots and drones to build maps and determine their own positions in real time in unknown environments.

[0037] SuperPoint Network: A deep learning model for feature point detection and description that can simultaneously output probability maps and descriptors of feature points, sharing encoders to reduce computation and parameter count.

[0038] Fixed-point quantization: Convert model parameters and activation values ​​represented by floating-point numbers into fixed-point numbers to reduce the model's storage space and computing resource requirements and improve operating efficiency on embedded devices.

[0039] BN layer (Batch Normalization Layer): Batch normalization layer is used in deep learning models to accelerate the training process and improve model stability and performance by normalizing the input of each layer and reducing internal covariate shift.

[0040] ReLU operation (Rectified Linear Unit): Rectified linear unit, a commonly used activation function used to introduce nonlinearity.

[0041] FP32 weights: Neural network weights represented as 32-bit single-precision floating point numbers, which have high precision and are suitable for model training and reasoning tasks with high precision requirements.

[0042] INT8 quantization strategy: A quantization method that converts a floating-point model into an 8-bit integer model. The quantization parameters are determined by offline statistics of the activation value range of each layer, which reduces the model size and computing resource consumption and improves the inference speed.

[0043] AXI bus (Advanced eXtensible Interface): A bus protocol used for FPGA chips such as Zynq, connecting the PS end and the PL end to achieve data transmission and control signal transmission.

[0044] PS (Processing System, PS): The processing system side, the ARM processor part in the FPGA, is responsible for running the operating system and high-level applications, managing resources and task scheduling.

[0045] PL (Programmable Logic, PL): Programmable logic side, the programmable logic part in the FPGA, used to implement custom hardware acceleration functions, such as convolution operations.

[0046] SLAM (simultaneous localization and mapping) has been widely used in the fields of mobile robots, Mars exploration, and drone mapping. Traditional SLAM uses a manually designed feature extraction algorithm to detect and calculate feature points and descriptors separately, which easily leads to computational redundancy and the inability to fully share the computational process. In recent years, with the development of deep learning, feature detection and descriptor generation based on convolutional neural networks have significantly outperformed traditional algorithms in terms of accuracy and robustness. Taking the SuperPoint network as an example, as an end-to-end feature point and descriptor extraction model, the SuperPoint network can share the encoder and output feature point probabilities and descriptors at the same time, reducing engineering complexity and cumulative errors. However, the floating-point calculations it uses are still very expensive in resource-constrained embedded scenarios (especially FPGAs), and floating-point operations require a lot of hardware overhead. At present, most of the research on the deployment of deep learning models such as SuperPoint on embedded FPGAs has only done network pruning or lightweighting, and the optimization of fixed-point quantization and BN fusion, ReLU fusion, and other FPGA adaptation is not sufficient, resulting in more hardware resources, high power consumption, and difficulty in further improving the inference speed.

[0047] Based on the above technical problems, an embodiment of the present invention provides a neural network quantization deployment method, which is applied to an embedded platform, wherein the embedded platform includes a programmable logic side and a processing system side, and the method includes: loading a target deep learning model on the processing system side; preprocessing the input data; transmitting the preprocessed data to the editable logic side, calling a fixed-point convolution IP core for convolution inference to obtain feature map data; transmitting the feature map data back to the processing system side, and the processing system side completes floating-point operations to output target data, wherein the target deep learning model is obtained through fixed-point quantization, and the fixed-point quantization is implemented collaboratively based on the programmable logic side and the processing system side.

[0048] Figure 1 The flowchart of the neural network quantization deployment method according to an embodiment of the present invention is schematically shown.

[0049] The neural network quantization deployment method provided in the embodiment of the present invention can be applied to an embedded platform, which includes a programmable logic side and a processing system side. The programmable logic side is the FPGA reprogrammable logic part, which is used to implement custom hardware acceleration functions, such as convolution operations; the processing system side is the ARM processor part in the FPGA, which is responsible for running the operating system and application programs, managing resources and task scheduling. Figure 1 As shown, the neural network quantization deployment method of this embodiment includes operations S210 to S240.

[0050] In operation S210 , a target deep learning model is loaded on the processing system side.

[0051] According to an embodiment of the present invention, the target deep learning model is obtained by fixed-point quantization after network splitting, and the fixed-point quantization is implemented based on the collaboration between the programmable logic side and the processing system side.

[0052] In one example, on the processing system side of the embedded platform, a C++ or Python program is used to load a deep learning model that has been fixed-point quantized. The specific process of fixed-point quantization can be referred to operation S310 and operation S320. In an embodiment of the present invention, the deep learning model can be, for example, a SuperPoint model. The SuperPoint network is an end-to-end feature point and descriptor extraction model that can share an encoder and output feature point probabilities and descriptors at the same time. The deep learning model of the embodiment of the present invention integrates a batch normalization layer and an activation function. After training, the weights and activation values ​​of the model have been converted to INT8 format so that the model can run on a resource-constrained embedded platform, reducing the storage space and computing resource requirements of the model.

[0053] In operation S220, the input data is preprocessed.

[0054] In one example, the input SLAM data is preprocessed. The input data may be, for example, image data from a camera. This includes operations such as adjusting the image size and normalizing pixel values ​​to make it meet the input requirements of the deep learning model, thereby improving the input data quality of the model and helping to improve the performance and accuracy of the model.

[0055] In operation S230, the preprocessed data is transmitted to the editable logic side, and the fixed-point convolution IP core is called to perform convolution inference to obtain feature map data.

[0056] In one example, the pre-processed data is transferred to the editable logic side through the AXI bus, and the pre-integrated fixed-point convolution IP core is called for convolution reasoning. The editable logic side uses the fixed-point convolution accelerator to perform the fused convolution layer operation in the integer domain, including pipeline parallel processing, to quickly obtain feature map data. Hardware acceleration is achieved, the speed and efficiency of convolution reasoning are improved, and the parallel computing capability of FPGA is fully utilized. The use of fixed-point convolution accelerators and pipeline parallel processing further improves the efficiency of convolution operations and reduces computing latency.

[0057] In operation S240, the feature map data is transmitted back to the processing system side, and the processing system side performs floating point operation to output target data.

[0058] In one example, complex floating-point operations are completed on the processing system side, ensuring the accuracy of the model output results while leveraging the CPU's advantages in complex operations. Ultimately, accurate SLAM results are obtained to achieve the positioning and map building functions of robots or autonomous systems.

[0059] Through the embodiments of the present invention, by loading the quantization model on the embedded platform, calling the FPGA accelerator for convolution reasoning, and completing subsequent floating-point operations on the processing system side, the storage space and computing resource requirements of the model are effectively reduced, the speed and efficiency of convolution reasoning are improved, and the accuracy of the model output results is guaranteed, thereby realizing efficient positioning and map building functions of robots or autonomous systems.

[0060] Figure 2A The schematic diagram of the deep learning model network structure according to an embodiment of the present invention is shown schematically. Figure 2B A schematic diagram of the network structure of a deep learning model after splitting according to an embodiment of the present invention is shown schematically. Figure 3 The flowchart of the neural network quantization deployment method according to another embodiment of the present invention is schematically shown. Figure 4 The flowchart of a neural network quantization deployment method according to another embodiment of the present invention is schematically shown.

[0061] According to an embodiment of the present invention, the implementation of fixed-point quantization based on the programmable logic side and the processing system side includes operations S310 and S320.

[0062] In operation S310, the deep learning model is preprocessed, and the preprocessing includes network splitting and network layer fusion.

[0063] According to an embodiment of the present invention, network splitting and network layer fusion specifically include: implementing shared encoder and partial decoder convolution layers on the programmable logic side, and implementing floating-point operations on the processing system side; and fusing the batch normalization layer and the convolution layer, and merging the activation function with the truncation of the quantization range.

[0064] In one example, due to the limited computing resources of the embedded platform, the embodiment of the present invention preprocesses the deep learning model, splits the network and fuses the network layers. Specifically, Figure 2A and Figure 2BAs shown in the figure, the convolutional layer, BN layer and some activation functions of the SuperPoint network are merged and divided to the FPGA side for implementation, and complex mathematical operations such as interpolation, softmax, and L2 normalization are retained on the processing system side for execution. Shared encoder and some decoder convolutional layers are implemented on the programmable logic side (FPGA side), such as the convolutional layer and maximum pooling layer in the SuperPoint network; floating-point operations such as descriptor interpolation (bicubic interpolation) and feature point decoding (softmax) are implemented on the processing system side (PS side). The batch normalization layer and the convolutional layer are fused and simplified into a "convolution + bias" form to avoid repeated resource consumption of the batch normalization layer in the hardware; the activation function and the truncation of the quantization range are merged and processed, and the positive value range counted after ReLU is used to determine the quantization scaling factor and zero point, so that Conv+ReLU can directly complete the activation truncation in the fixed-point domain. Through network splitting and network layer fusion, the computing tasks are reasonably allocated, and the convolution operations suitable for hardware acceleration are placed on the FPGA side, while the complex floating-point operations are left on the PS side, giving full play to the advantages of both sides and improving the overall computing efficiency. Network layer fusion reduces the redundant layers in the network, reduces the computing complexity and resource consumption, while maintaining the performance and accuracy of the model. This optimizes the model structure, reduces the amount of computing and resource usage, while maintaining the performance of the model.

[0065] In operation S320 , fixed-point quantization is performed on the network parameters of the preprocessed deep learning model based on the collaboration between the programmable logic side and the processing system side.

[0066] like Figure 4 As shown, fixed-point quantization of the network parameters of the preprocessed deep learning model based on the collaboration between the programmable logic side and the processing system side includes operations S321 to S327.

[0067] In operation S321, a deep learning framework is used to train the neural network model and save the model weight parameters.

[0068] In operation S322, the maximum activation value of each layer is counted offline to determine the range of the quantization parameter; and the batch normalization layer and the convolution layer of each layer are fused to determine the fused convolution weight.

[0069] In operation S323 , quantization parameters of each layer are calculated based on the quantization strategy.

[0070] In operation S324, the convolution layer, the batch normalization layer, and the activation layer are merged, and a fixed-point convolution IP core is generated according to the quantization parameter.

[0071] In operation S325 , the fixed-point convolution IP core is integrated on the programmable logic side.

[0072] In operation S326, the descriptor interpolation and feature point decoding operations are completed on the processing system side.

[0073] In one example, a deep learning framework such as PyTorch can be used on any PC to train a neural network model, and the FP32 weight parameters of the model are saved after training. The model is run using a representative data set, and the maximum and minimum activation values ​​of each layer are counted offline to determine the range of quantization parameters. The batch normalization layer and convolution layer of each layer are fused to determine the fused convolution weight, which simplifies the network structure, improves computational efficiency, and reduces computational workload and resource usage. Based on a quantization strategy, such as an INT8 quantization strategy, the quantization parameters of each layer are calculated, including scaling factors and zero points. The convolution layer, batch normalization layer, and activation layer are merged, and a fixed-point convolution IP core is generated based on the quantization parameters, and the IP core is integrated on the programmable logic side (FPGA side). Descriptor interpolation and feature point decoding operations are implemented on the processing system side, such as using C++ or Python programs to perform floating-point operations such as bicubic interpolation, softmax, and L2 normalization.

[0074] The network fixed-point quantization method provided by the embodiment of the present invention and the collaborative splitting strategy based on the programmable logic side and the processing system side give full play to the flexibility of the processing system side and the efficiency of the programmable logic side, realize the efficient quantization of the model, and improve the deployment efficiency and operation performance of the model on the embedded platform. It can effectively realize the efficient reasoning of SLAM feature extraction on the FPGA platform, with significant acceleration effect, and maintain the key points and descriptor accuracy similar to the original network.

[0075] Based on the above neural network quantitative deployment method, the present invention also provides a neural network quantitative deployment device. Figure 5 The device is described in detail.

[0076] Figure 5 The structural block diagram of a neural network quantization deployment device according to an embodiment of the present invention is schematically shown.

[0077] like Figure 5 As shown, the neural network quantization deployment device 500 of this embodiment includes a quantization model loading module 510, an input data preprocessing module 520, a convolutional reasoning module 530 and an operation module 540.

[0078] The quantization model loading module 510 is used to load the target deep learning model on the processing system side. In one embodiment, the quantization model loading module 510 can be used to perform the operation S210 described above, which will not be repeated here.

[0079] The input data preprocessing module 520 is used to preprocess the input data. In one embodiment, the input data preprocessing module 520 can be used to perform the operation S220 described above, which will not be described in detail here.

[0080] The convolution inference module 530 is used to transmit the preprocessed data to the editable logic side, and call the fixed-point convolution IP core to perform convolution inference to obtain feature map data. In one embodiment, the convolution inference module 530 can be used to perform the operation S230 described above, which will not be repeated here.

[0081] The operation module 540 is used to transmit the feature map data back to the processing system side, and the processing system side completes the floating point operation to output the target data. In one embodiment, the operation module 540 can be used to perform the operation S240 described above, which will not be repeated here.

[0082] According to an embodiment of the present invention, the device further includes:

[0083] A model preprocessing module, used to preprocess the deep learning model, wherein the preprocessing includes network splitting and network layer fusion; and

[0084] The fixed-point quantization module is used to perform fixed-point quantization on the network parameters of the preprocessed deep learning model based on the collaboration between the programmable logic side and the processing system side.

[0085] The model preprocessing module includes a network splitting submodule and a network layer fusion submodule.

[0086] a network splitting submodule for implementing shared encoder and partial decoder convolutional layers on the programmable logic side and floating point operations on the processing system side; and

[0087] The network layer fusion submodule is used to fuse the batch normalization layer and the convolution layer, and to combine the activation function with the truncation of the quantization range.

[0088] The fixed-point quantization module is specifically used to train the neural network model using a deep learning framework and save the model weight parameters; offline statistics of the maximum activation value of each layer to determine the range of the quantization parameter; the batch normalization layer and the convolution layer of each layer are fused to determine the fused convolution weight; the quantization parameters of each layer are calculated based on the quantization strategy; the convolution layer, the batch normalization layer and the activation layer are merged, and a fixed-point convolution IP core is generated according to the quantization parameters; the fixed-point convolution IP core is integrated on the programmable logic side; and the descriptor interpolation and feature point decoding operations are implemented on the processing system side.

[0089] According to an embodiment of the present invention, any multiple modules of the quantization model loading module 510, the input data preprocessing module 520, the convolution reasoning module 530 and the operation module 540 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the quantization model loading module 510, the input data preprocessing module 520, the convolution reasoning module 530 and the operation module 540 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in a suitable combination of any of them. Alternatively, at least one of the quantization model loading module 510, the input data preprocessing module 520, the convolutional reasoning module 530 and the operation module 540 can be at least partially implemented as a computer program module, and when the computer program module is executed, the corresponding function can be performed.

[0090] Figure 6 A block diagram of an electronic device suitable for implementing a neural network quantization deployment method according to an embodiment of the present invention is schematically shown.

[0091] like Figure 6 As shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 to a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include an onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0092] In RAM 603, various programs and data required for the operation of electronic device 600 are stored. Processor 601, ROM 602 and RAM 603 are connected to each other via bus 604. Processor 601 performs various operations of the method flow according to the embodiment of the present invention by executing the program in ROM 602 and / or RAM 603. It should be noted that the program can also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 can also perform various operations of the method flow according to the embodiment of the present invention by executing the program stored in the one or more memories.

[0093] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage portion 608 as needed.

[0094] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.

[0095] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 602 and / or RAM 603 described above and / or one or more memories other than ROM 602 and RAM 603.

[0096] The embodiment of the present invention also includes a computer program product, which includes a computer program, and the computer program includes a program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the neural network quantization deployment method provided by the embodiment of the present invention.

[0097] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when it is executed by the processor 601. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0098] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0099] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.

[0100] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0101] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0102] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention may be combined and / or combined in various ways. All of these combinations and / or combinations fall within the scope of the present invention.

[0103] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination advantageously. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A neural network quantization deployment method, applied to an embedded platform, wherein the embedded platform includes a programmable logic side and a processing system side, characterized in that: The method comprises: Load the target deep learning model on the processing system side; Preprocess the input data; The preprocessed data is transferred to the editable logic side, and the fixed-point convolution IP core is called for convolution inference to obtain feature map data; The feature map data is transmitted back to the processing system side, and the processing system side performs floating point operations to output target data. Among them, the target deep learning model is obtained by fixed-point quantization after network splitting, and the fixed-point quantization is implemented based on the collaboration between the programmable logic side and the processing system side.

2. The method according to claim 1, characterized in that The implementation of fixed-point quantization based on the programmable logic side and the processing system side includes: Preprocessing the deep learning model, wherein the preprocessing includes network splitting and network layer fusion; and The network parameters of the preprocessed deep learning model are fixed-point quantized based on the collaboration between the programmable logic side and the processing system side.

3. The method according to claim 2, characterized in that Network splitting and network layer fusion specifically include: Implementing shared encoder and partial decoder convolutional layers on the programmable logic side and implementing floating point operations on the processing system side; and The batch normalization layer is fused with the convolutional layer, and the activation function is combined with the truncation of the quantization range.

4. The method according to claim 2, characterized in that: The method of performing fixed-point quantization on the network parameters of the preprocessed deep learning model based on the collaboration between the programmable logic side and the processing system side includes: Use a deep learning framework to train the neural network model and save the model weight parameters; Offline statistics of the maximum activation value of each layer are used to determine the range of quantization parameters; Fuse the batch normalization layer and convolution layer of each layer to determine the fused convolution weights; Calculate the quantization parameters of each layer based on the quantization strategy; Combine the convolution layer, batch normalization layer, and activation layer, and generate a fixed-point convolution IP core based on the quantization parameters; Integrate the fixed-point convolution IP core on the programmable logic side; and Descriptor interpolation and feature point decoding operations are completed on the processing system side.

5. A neural network quantization deployment device, applied to an embedded platform, the embedded platform includes a programmable logic side and a processing system side, the programmable logic side pre-integrates a fixed-point convolution IP core, characterized in that: The device comprises: A quantized model loading module, used to load the target deep learning model on the processing system side; An input data preprocessing module, used for preprocessing input data; The convolution inference module is used to transfer the preprocessed data to the editable logic side and call the fixed-point convolution IP core for convolution inference to obtain feature map data; A computing module is used to transmit the feature map data back to the processing system side, and the processing system side performs floating-point computing operations to output target data. Among them, the target deep learning model is obtained by fixed-point quantization after network splitting, and the fixed-point quantization is implemented based on the collaboration between the programmable logic side and the processing system side.

6. The device according to claim 5, characterized in that The device also includes: A model preprocessing module, used to preprocess the deep learning model, wherein the preprocessing includes network splitting and network layer fusion; and The fixed-point quantization module is used to perform fixed-point quantization on the network parameters of the preprocessed deep learning model based on the collaboration between the programmable logic side and the processing system side.

7. The device according to claim 6, characterized in that The model preprocessing module includes a network splitting submodule and a network layer fusion submodule. A network splitting submodule for implementing shared encoder and partial decoder convolutional layers on the programmable logic side and floating point operations on the processing system side; as well as The network layer fusion submodule is used to fuse the batch normalization layer and the convolution layer, and to combine the activation function with the truncation of the quantization range.

8. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 4.

9. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Method and system for efficiently quantizing FPGA (Field Programmable Gate Array) hardware acceleration of long and short-term memory network and detecting abnormal electroencephalogram signals of FPGA hardware acceleration

    CN118228789A

  • Heterogeneous hardware accelerator for lightweight deep neural network and design method

    CN118520917A