Digital Recognition Optimization Method, System and Medium Based on High-Level Synthesis Tools
Through high-level comprehensive tools, the loop body of the LeNet-5 convolutional neural network is optimized, converted into RTL code and implemented on FPGA, solving the problem of low development efficiency of real-time digital recognition algorithms on FPGAs, and improving hardware development efficiency and execution speed.
Patent Information
- Application Number
- CN202211570725.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-08
AI Technical Summary
In the prior art, the hardware development efficiency of real-time digital recognition algorithms on FPGAs is low, and the system execution speed is slow, mainly due to insufficient optimization processing of the loop body, which leads to high computing power requirements and slow execution speed.
A high-level comprehensive tool is used to build a LeNet-5 convolutional neural network, model training is performed through training data sets, high-level language representation is used and simulated and verified, optimize the loop body of each layer, transform it into RTL code and export it to IP core, and combine FPGA to realize image data input and recognition.
The hardware development efficiency and system execution speed of digital recognition algorithms are improved, and the rapid implementation of digital recognition algorithms on FPGAs is realized.
Smart Images

Figure CN115830415B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular to a digital recognition optimization method, system and medium based on a high-level synthesis tool. Background Art
[0002] Convolutional neural networks have extremely important research significance and application value in the fields of image classification and processing, video surveillance, and machine vision. Taking the original image as the input, it automatically processes the input to avoid the preprocessing link of the image, especially the preprocessing link of the image involving manual participation. This is also one of the advantages of convolutional neural networks compared to traditional image processing methods. In 1989, LeCun constructed a convolutional neural network applied to computer vision problems, which is an early version of LeNet, including two convolutional layers and two fully connected layers with a total of more than 60,000 parameters. Its structure is similar to that of modern convolutional neural network models, and the concept of "convolution" was pioneered, so convolutional neural networks got their name. In 1998, LeCun constructed a more complete convolutional neural network LeNet5 and applied it to handwritten font recognition. A pooling layer was added on the basis of the original LeNet, and the recognition accuracy of the model on the MNIST dataset reached more than 98%.
[0003] To implement a real-time digital recognition algorithm on an FPGA, a hardware description language is required, and only hardware engineers can perform hardware design. Software engineers cannot complete such work, which greatly limits the development efficiency of real-time digital recognition algorithms. During the FPGA development process, the execution of the loop body is often the part that consumes the most system time. Therefore, the optimization of the loop body has become the top priority of the design. Moreover, there are many loop body structures in each layer of the digital recognition algorithm, which requires high system computing power and significantly slows down the system execution speed. Summary of the Invention
[0004] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.
[0005] To this end, an object of an embodiment of the present invention is to provide a digital recognition optimization method based on a high-level synthesis tool, which improves the hardware development efficiency and system execution speed of the digital recognition algorithm.
[0006] Another object of an embodiment of the present invention is to provide a digital recognition optimization system based on a high-level synthesis tool.
[0007] In order to achieve the above technical object, the technical solutions adopted in the embodiments of the present invention include:
[0008] On the one hand, an embodiment of the present invention provides a digital recognition optimization method based on a high-level synthesis tool, including the following steps:
[0009] Construct a LeNet-5 convolutional neural network, and train the LeNet-5 convolutional neural network according to a preset training data set to obtain a digital recognition model;
[0010] Represent the digital recognition model in a high-level language, and perform simulation verification on the digital recognition model through a high-level synthesis tool;
[0011] Optimize each layer loop body in the digital recognition model through the high-level synthesis tool;
[0012] Convert the optimized digital recognition model into RTL code and export it as an IP core;
[0013] Establish a connection between the IP core and a camera module through an FPGA, so that the camera module inputs the acquired image data into the IP core for digital recognition.
[0014] Further, in an embodiment of the present invention, the step of constructing the LeNet-5 convolutional neural network specifically includes:
[0015] Determine the input image format and input image size of the LeNet-5 convolutional neural network according to the test samples of the training data set;
[0016] Determine the receptive field size, activation function, and output connection method of the LeNet-5 convolutional neural network according to the target function;
[0017] Construct the LeNet-5 convolutional neural network according to the input image format, the input image size, the receptive field size, the activation function, and the output connection method.
[0018] Further, in an embodiment of the present invention, the step of training the LeNet-5 convolutional neural network according to a preset training data set to obtain a digital recognition model specifically includes:
[0019] Input the training data set into the LeNet-5 convolutional neural network to obtain a digital recognition result;
[0020] Determine the loss value of the LeNet-5 convolutional neural network according to the digital recognition result and the sample label of the training data set;
[0021] Update the model parameters of the LeNet-5 convolutional neural network according to the loss value through the backpropagation algorithm, and return the step of inputting the training data set into the LeNet-5 convolutional neural network;
[0022] When the loss value reaches a preset first threshold or the number of model iterations reaches a preset second threshold, stop training to obtain a trained digital recognition model.
[0023] Further, in an embodiment of the present invention, the step of representing the digital recognition model in a high-level language and simulating and verifying the digital recognition model through a high-level synthesis tool specifically includes:
[0024] Write the model code of the digital recognition model in C language;
[0025] Perform C language simulation on the model code through a high-level synthesis tool, and verify the accuracy of the digital recognition model according to the simulation results.
[0026] Further, in an embodiment of the present invention, the step of optimizing each layer loop body in the digital recognition model through the high-level synthesis tool includes at least one of the following steps:
[0027] Unroll the loop body into multiple sub-loop bodies with the same structure through the unroll optimization instruction of the high-level synthesis tool;
[0028] Shorten the instruction trigger interval between loop bodies through the pipeline optimization instruction of the high-level synthesis tool;
[0029] Enable the loop body to execute in parallel through the dataflow optimization instruction of the high-level synthesis tool.
[0030] Further, in an embodiment of the present invention, the step of converting the optimized digital recognition model into RTL code and exporting it as an IP core specifically includes:
[0031] Obtain the C language expression form of the optimized digital recognition model;
[0032] Convert the C language expression form into RTL code through the high-level synthesis tool;
[0033] Export the RTL code as an IP core callable by the FPGA.
[0034] Further, in an embodiment of the present invention, the step of establishing a connection between the IP core and the camera module through the FPGA, so that the camera module inputs the acquired image data into the IP core for digital recognition, specifically includes:
[0035] Converting the image data acquired by the camera module into first image data in RGB format through the FPGA;
[0036] Performing filtering processing on the first image data through the FPGA to obtain second image data that conforms to the input image format and input image size;
[0037] Inputting the second image data into the IP core through the FPGA and outputting the corresponding digital recognition result through an HDMI transmission line.
[0038] On the other hand, an embodiment of the present invention provides a digital recognition optimization system based on a high-level synthesis tool, including:
[0039] A model training module, configured to construct a LeNet-5 convolutional neural network and train the LeNet-5 convolutional neural network according to a preset training data set to obtain a digital recognition model;
[0040] A model simulation module, configured to represent the digital recognition model in a high-level language and perform simulation verification on the digital recognition model through a high-level synthesis tool;
[0041] A loop body optimization module, configured to optimize each layer of loop body in the digital recognition model through the high-level synthesis tool;
[0042] An IP core export module, configured to convert the optimized digital recognition model into RTL code and export it as an IP core;
[0043] A digital recognition module, configured to establish a connection between the IP core and the camera module through the FPGA, so that the camera module inputs the acquired image data into the IP core for digital recognition.
[0044] On the other hand, an embodiment of the present invention provides a digital recognition optimization device based on a high-level synthesis tool, including:
[0045] At least one processor;
[0046] At least one memory, configured to store at least one program;
[0047] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned digital recognition optimization method based on a high-level synthesis tool.
[0048] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the above-mentioned digital recognition optimization method based on a high-level synthesis tool when executed by the processor.
[0049] The advantages and beneficial effects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention:
[0050] An embodiment of the present invention provides a digital recognition optimization method based on a high-level synthesis tool. First, a LeNet-5 convolutional neural network is constructed and trained using a preset training data set to obtain a digital recognition model. Then, the digital recognition model is represented by a high-level language, and the digital recognition model is simulated and verified by a high-level synthesis tool. Then, each layer loop body in the digital recognition model is optimized by the high-level synthesis tool, and the optimized digital recognition model is converted into RTL code and exported as an IP core. Furthermore, the IP core is connected to a camera module through an FPGA, so that the camera module inputs the collected image data into the IP core for digital recognition. The embodiment of the present invention uses a high-level synthesis tool to optimize the loop body of the digital recognition model described by a high-level language, and converts the optimized digital recognition model into RTL code and then exports it as an IP core, so that the digital recognition algorithm can be quickly implemented on an FPGA, improving the hardware development efficiency and system execution speed of the digital recognition algorithm. Description of the Drawings
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the drawings required to be used in the embodiments of the present invention. It should be understood that the drawings introduced below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a flowchart of the steps of a digital recognition optimization method based on a high-level synthesis tool provided by an embodiment of the present invention;
[0053] Figure 2 It is a schematic structural diagram of a LeNet-5 convolutional neural network provided by an embodiment of the present invention;
[0054] Figure 3 It is a schematic diagram of the convolution operation code from the input layer to the convolution layer of the digital recognition model provided by an embodiment of the present invention;
[0055] Figure 4Schematic diagram of resource occupation after implementing a digital recognition algorithm on an FPGA provided by an embodiment of the present invention;
[0056] Figure 5 Schematic diagram of the interface of the serial port control filter provided by an embodiment of the present invention;
[0057] Figure 6 Block diagram of a digital recognition optimization system based on a high-level synthesis tool provided by an embodiment of the present invention;
[0058] Figure 7 Block diagram of a digital recognition optimization device based on a high-level synthesis tool provided by an embodiment of the present invention. Detailed implementation manners
[0059] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0060] In the description of the present invention, the meaning of "a plurality of" is two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention.
[0061] Referring to Figure 1 , an embodiment of the present invention provides a digital recognition optimization method based on a high-level synthesis tool, which specifically includes the following steps:
[0062] S101. Construct a LeNet-5 convolutional neural network, and train the LeNet-5 convolutional neural network according to a preset training data set to obtain a digital recognition model.
[0063] Further as an optional implementation manner, the step of constructing a LeNet-5 convolutional neural network specifically includes:
[0064] S1011. Determine the input image format and input image size of the LeNet-5 convolutional neural network according to the test samples of the training data set;
[0065] S1012. Determine the receptive field size, activation function, and output connection method of the LeNet-5 convolutional neural network according to the target function;
[0066] S1013. Construct the LeNet-5 convolutional neural network according to the input image format, input image size, receptive field size, activation function, and output connection method.
[0067] Specifically, the embodiment of the present invention uses the MINST dataset as the training dataset, and determines the picture format and size requirements of the input layer according to the test samples in the MINST dataset, so as to constrain the digital recognition model. As Figure 2 shown in the structural schematic diagram of the LeNet-5 convolutional neural network provided by the embodiment of the present invention, analyze the functions of the 2 convolutional layers, 2 pooling layers, and 3 fully connected layers of LeNet5, and determine information such as its receptive field size, stride, boundary padding, and activation function; determine the connection method of the output layer of the digital recognition algorithm, so that the recognition result can be displayed on the display screen in real time via the HDMI transmission line in the subsequent process.
[0068] The purpose of the convolution operation is to extract different features of the input. The first convolutional layer may only be able to extract some low-level features such as edges, lines, and corners, etc. More layers of the network can iteratively extract more complex features from the low-level features. The pooling layer is actually a kind of downsampling, and there are various forms of non-linear pooling functions, and the max pooling and average sampling are the most common among them. The role of the pooling layer is equivalent to converting a picture with a higher resolution into a picture with a lower resolution. The pooling layer can further reduce the number of nodes in the last fully connected layer, so as to achieve the purpose of reducing the parameters in the entire neural network. The fully connected layer is generally in the last few layers of the algorithm and is responsible for extracting the features after convolution and pooling.
[0069] The last fully connected layer of LeNet5 is the output layer of the algorithm, and the network connection method of the radial basis function (RBF) is adopted. This layer has a total of 10 nodes, which respectively represent the numbers 0 to 9. When the algorithm recognizes the result, the output result will be saved in the pointer, and then the ARM core calls this pointer to take out the recognition result and output it via HDMI.
[0070] Further as an optional implementation manner, the step of training the LeNet-5 convolutional neural network according to the preset training dataset to obtain the digital recognition model specifically includes:
[0071] S1014. Input the training dataset into the LeNet-5 convolutional neural network to obtain the digital recognition result;
[0072] S1015. Determine the loss value of the LeNet-5 convolutional neural network according to the digital recognition result and the sample label of the training dataset;
[0073] S1016. Update the model parameters of the LeNet-5 convolutional neural network according to the loss value through the backpropagation algorithm, and return to the step of inputting the training data set into the LeNet-5 convolutional neural network;
[0074] S1017. When the loss value reaches a preset first threshold or the number of model iterations reaches a preset second threshold, stop the training to obtain a trained digital recognition model.
[0075] S102. Represent the digital recognition model in a high-level language, and perform simulation verification on the digital recognition model through a high-level synthesis tool.
[0076] Specifically, high-level synthesis (HLS) is abbreviated as HLS, which refers to the process of automatically converting the logical structure described in a high-level language into a circuit model described in a low-level abstraction language. The HLS tool can reduce the design time of hardware engineers and also enable software engineers to complete hardware design. High-level languages, including C, C++, etc., usually have a high degree of abstraction and usually do not have the concept of timing, while so-called low-level languages, such as Verilog, VHDL, SystemVerilog, etc., are usually used to describe register transfer models and usually have the concept of clock. Step S102 specifically includes the following steps:
[0077] S1021. Write the model code of the digital recognition model in C language;
[0078] S1022. Perform C language simulation on the model code through a high-level synthesis tool, and verify the accuracy rate of the digital recognition model according to the simulation results.
[0079] Specifically, the model code of the digital recognition model written in C language in the embodiment of the present invention has been trained on the MNIST data set and has relevant weight information for digital recognition; perform C language simulation through a high-level synthesis tool to verify the correctness of the algorithm logic.
[0080] As Figure 3 shown is a schematic diagram of the convolution operation code from the input layer to the convolution layer of the digital recognition model provided by the embodiment of the present invention. According to the determined parameter information of the LeNeT5 network, write the convolution operation of the convolution layer, and the operation algorithms of the other layers are also written in this way respectively.
[0081] In the embodiment of the present invention, a simulation program is written in a high-level synthesis tool, and the test set in MNIST is imported for simulation recognition. After verification, the accuracy rate of digital recognition reaches 98%.
[0082] S103. Optimize each loop body in the digital recognition model through a high-level synthesis tool.
[0083] Specifically, there are many loop bodies in the digital recognition algorithm. At the same time, most loop bodies are perfect loops. Only the innermost loop has the main content. There is no specified logic between loop statements. The loop bounds are constant, and the result of the next iteration is independent of the previous iteration in loop iterations. Therefore, a large number of optimization operations can be performed on the loop body. This situation is particularly suitable for using high-level synthesis technology to perform related design optimizations.
[0084] The embodiment of the present invention uses high-level synthesis technology to optimize the loop body in the digital recognition model, so that two adjacent operations in the loop are realized with the minimum time delay. Step S103 includes at least one of the following steps:
[0085] S1031. Expand the loop body into multiple sub-loop bodies with the same structure through the unroll optimization instruction of the high-level synthesis tool;
[0086] S1032. Shorten the instruction trigger interval between loop bodies through the pipeline optimization instruction of the high-level synthesis tool;
[0087] S1033. Make the loop body execute in parallel through the dataflow optimization instruction of the high-level synthesis tool.
[0088] Specifically, the unroll optimization instruction can expand the loop to increase data access and throughput. The Unroll optimization instruction optimizes in the code area of the for loop. This instruction does not contain the concept of pipeline execution. It simply expands a long loop body into sub-loop bodies with the same structure and uses more hardware resources to implement it, ensuring that parallel loop bodies are independent of each other during scheduling; the role of the pipeline optimization instruction is to shorten the instruction trigger interval within a C function or C loop. It can be used at both the loop and function levels. By increasing repeated operation instructions (such as increasing resource usage, etc.) to reduce the initialization interval, after optimization, the FPGA can process a large amount of data asynchronously as possible; the dataflow optimization instruction is a task-level pipeline instruction that enables loops or functions to execute in parallel from a higher task level, aiming to reduce latency and increase throughput, thereby improving the overall design interval. In the embodiment of the present invention, the loop information and corresponding optimization instructions of each function layer loop body in the digital recognition model in the high-level synthesis tool are shown in Table 1 below.
[0089]
[0090] Table 1
[0091] In the convolutional layer and the pooling layer, the loop body is a perfect loop and the number of loops is small, so the selected optimization method is HLS PIPELINE; while in the fully connected layer, since all features need to be integrated in the algorithm of the fully connected layer, the number of loops is large. Although the HLS PIPELINE optimization method can shorten the instruction trigger interval between loops, its optimization effect is not as high as that of the HLS UNROLL optimization method. HLS UNROLL can unroll the loop, which means replicating the same circuit structure multiple times on the FPGA and using more hardware resources to implement the loop, making it more suitable for optimizing loop bodies with a large number of loops; in the Top function, there are assignment operations inside the loop body, which is not a perfect loop body, so pipeline optimization cannot be performed, but dataflow optimization (i.e., dataflow) can be applied, which enables the loop and the function to execute in parallel to the greatest extent, thereby improving the overall design interval.
[0092] In the embodiments of the present invention, after optimizing the loop bodies of the Top function, the convolutional layer, the pooling layer, and the fully connected layer respectively through corresponding optimization instructions, the optimized digital recognition model can be obtained. The comparison of the latency and interval before and after optimizing the digital recognition model using a high-level synthesis tool is shown in Table 2 below.
[0093] solution2 Solution3 Latency(cycles) min 14392845 6125761 max 14392845 6125761 Latency(absolute) min 96.432ms 41.043ms max 96.432ms 41.043ms Interval(cycles) min 14392846 6125762 max 14392846 6125762
[0094] Table 2
[0095] In Table 2, solution2 represents the result before optimization, and solution3 represents the result after optimization. It can be seen that the latency and interval are significantly reduced after optimization, improving the speed of the algorithm for processing data. The time for the algorithm to process a set of data in the ideal state is 41.043 milliseconds.
[0096] S104. Convert the optimized digital recognition model into RTL code and export it as an IP core.
[0097] Specifically, in order to connect the algorithm and the hardware, it is necessary to use a high-level synthesis tool to convert the algorithm function implemented by a high-level language such as C language into a function IP core that can be recognized by a low-level language, so that different functional modules can be connected into a system through the FPGA. Step S104 specifically includes the following steps:
[0098] S1041. Obtain the C language expression form of the optimized digital recognition model;
[0099] S1042. Convert the C language expression form into RTL code through a high-level synthesis tool;
[0100] S1043. Export the RTL code as an IP core that can be called by the FPGA.
[0101] S105. Establish a connection between the IP core and the camera module through the FPGA, so that the camera module inputs the acquired image data into the IP core for digital recognition.
[0102] Specifically, the FPGA has the characteristics of high parallelism and fast operation speed, and can easily complete data processing such as image format type conversion and filtering of image data. In the development tool of the FPGA, it is allowed to encapsulate the RTL code in the library file into a module, which is convenient for each module to be clearly and intuitively connected to the IP core graphically. Step S105 specifically includes the following steps:
[0103] S1051. Convert the image data collected by the camera module through the FPGA into the first image data in RGB format;
[0104] S1052. Perform filtering processing on the first image data through the FPGA to obtain the second image data that meets the input image format and input image size;
[0105] S1053. Input the second image data into the IP core through the FPGA, and output the corresponding digital recognition result through the HDMI transmission line.
[0106] Specifically, the model of the camera module in the embodiment of the present invention is OV5640, which is an image optical sensor chip with a target surface size of 1 / 4 inch. The data output format of OV5640 supports MIPI (Mobile Industry Processor Interface) and DVP (Digital Video Parallel), while the input of the digital recognition algorithm is generally picture data. In order to be able to realize the function of digital recognition in real time, it is necessary to convert and process the image data collected by OV5640.
[0107] Convert the image data collected by the camera module OV5640 into YUV format through the FPGA, and then convert it into RGB format; in the FPGA, the real-time image data in RGB format is processed by the filter module fir2d, and the image information is converted into the format required by the input layer of the digital recognition algorithm; the IP core of the digital recognition model recognizes the digital information in the processed real-time image and displays it on the screen in real time through the HDMI transmission line.
[0108] The fir2d module is a module with image data processing function encapsulated by RTL code. Four register addresses are set inside it, which are 0x00, 0x04, 0x08, and 0x0C respectively. Among them, the registers at the first three addresses are the parameters of the filter input data stream, and information such as resolution and image format size can be modified according to these three addresses; the fourth register is the filter parameter register, and different filter models can be replaced by modifying the register parameter values.
[0109] High-Definition Multimedia Interface (HDMI) is an interface standard with the ability to transmit high-definition digital video and digital audio. It is small in size, high in transmission rate, large in transmission bandwidth, and good in compatibility. It can transmit uncompressed audio and video signals simultaneously without the need for digital-to-analog or analog-to-digital conversion before signal transmission. It is a dedicated digital interface very suitable for image transmission. In the embodiment of the present invention, it is used as the output port of the camera module OV5640 and the digital recognition result, and can clearly display the data acquisition information and the recognition result on the screen in real time.
[0110] As Figure 4 shown is the schematic diagram of resource occupancy after implementing the digital recognition algorithm on the FPGA provided by the embodiment of the present invention. It can be seen that after the optimized digital recognition algorithm is synthesized and implemented on the FPGA, the resource occupancy does not exceed the resource limit of the selected system-on-chip.
[0111] As Figure 5 shown is the schematic diagram of the interface of the serial port control filter provided by the embodiment of the present invention. In the embodiment of the present invention, the upper computer and the system-on-chip are interactively controlled through the serial port communication protocol. The module fir2d internally has different filter models, such as no filtering, Gaussian filtering, high-boost filtering, and Sobel operator filtering, etc. Gaussian filtering is a linear smoothing filter, which is a process of weighted averaging of the entire image and is suitable for removing Gaussian noise in the image; high-boost filtering can sharpen the image, enhance the contrast by increasing the local gray difference, and is suitable for scenes with relatively dim light; the Sobel operator is a discrete difference operator used to calculate the approximate value of the gradient of the image brightness function and is mainly used for edge detection. The filter model can be modified in real time through serial port control to adapt to digital recognition in different light scenes and improve the recognition accuracy.
[0112] The method steps of the embodiment of the present invention are described above. It can be understood that the embodiment of the present invention uses a high-level synthesis tool to optimize the loop body of the digital recognition model described in a high-level language, and converts the optimized digital recognition model into RTL code and then exports it as an IP core, so that the digital recognition algorithm can be quickly implemented on the FPGA, improving the hardware development efficiency and system execution speed of the digital recognition algorithm.
[0113] Referring to Figure 6 , the embodiment of the present invention provides a digital recognition optimization system based on a high-level synthesis tool, including:
[0114] A model training module, configured to construct a LeNet-5 convolutional neural network, and train the LeNet-5 convolutional neural network according to a preset training data set to obtain a digital recognition model;
[0115] A model simulation module, which is used to represent a digital recognition model in a high-level language and perform simulation verification on the digital recognition model through a high-level synthesis tool;
[0116] A loop body optimization module, which is used to optimize each layer of loop bodies in the digital recognition model through a high-level synthesis tool;
[0117] An IP core export module, which is used to convert the optimized digital recognition model into RTL code and export it as an IP core;
[0118] A digital recognition module, which is used to establish a connection between the IP core and a camera module through an FPGA, so that the camera module inputs the collected image data into the IP core for digital recognition.
[0119] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0120] Referring to Figure 7 , an embodiment of the present invention provides a digital recognition optimization device based on a high-level synthesis tool, including:
[0121] At least one processor;
[0122] At least one memory, which is used to store at least one program;
[0123] When the above at least one program is executed by the above at least one processor, the above at least one processor implements the above digital recognition optimization method based on a high-level synthesis tool.
[0124] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0125] An embodiment of the present invention also provides a computer-readable storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to execute the above digital recognition optimization method based on a high-level synthesis tool when executed by the processor.
[0126] A computer-readable storage medium according to an embodiment of the present invention can execute a digital recognition optimization method based on a high-level synthesis tool provided by an embodiment of the method of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.
[0127] Embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.
[0128] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the above-mentioned blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0129] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More precisely, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0130] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0131] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.
[0132] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which the above program can be printed, because the above program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways when necessary, and then storing it in a computer memory.
[0133] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0134] In the foregoing description of this specification, descriptions with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0135] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
[0136] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A digital recognition optimization method based on a high-level synthesis tool, characterized in that, It includes the following steps: Construct a LeNet-5 convolutional neural network, and train the LeNet-5 convolutional neural network according to a preset training dataset to obtain a digital recognition model; Represent the digital recognition model in a high-level language, and perform simulation verification on the digital recognition model through a high-level synthesis tool; Optimize each layer loop body in the digital recognition model through the high-level synthesis tool; Convert the optimized digital recognition model into RTL code and export it as an IP core; Establish a connection between the IP core and a camera module through an FPGA, so that the camera module inputs the collected image data into the IP core for digital recognition.
2. The digital recognition optimization method based on a high-level synthesis tool according to claim 1, characterized in that The step of constructing the LeNet-5 convolutional neural network specifically includes: Determine the input image format and input image size of the LeNet-5 convolutional neural network according to the test samples of the training dataset; Determine the receptive field size, activation function, and output connection method of the LeNet-5 convolutional neural network according to the target function; Construct the LeNet-5 convolutional neural network according to the input image format, the input image size, the receptive field size, the activation function, and the output connection method.
3. A digital recognition optimization method based on a high-level synthesis tool according to claim 1, characterized in that The step of training the LeNet-5 convolutional neural network according to a preset training dataset to obtain a digital recognition model specifically includes: Input the training dataset into the LeNet-5 convolutional neural network to obtain a digital recognition result; Determine the loss value of the LeNet-5 convolutional neural network according to the digital recognition result and the sample label of the training dataset; Update the model parameters of the LeNet-5 convolutional neural network through the backpropagation algorithm according to the loss value, and return to the step of inputting the training dataset into the LeNet-5 convolutional neural network; When the loss value reaches a preset first threshold or the model iteration times reach a preset second threshold, stop training to obtain a trained digital recognition model.
4. A digital recognition optimization method based on a high-level synthesis tool according to claim 1, characterized in that The step of representing the digital recognition model in a high-level language and performing simulation verification on the digital recognition model through a high-level synthesis tool specifically includes: Write the model code of the digital recognition model in C language; Perform C language simulation on the model code through a high-level synthesis tool, and verify the accuracy of the digital recognition model according to the simulation result.
5. A digital recognition optimization method based on a high-level synthesis tool according to claim 1, characterized in that The step of optimizing each layer loop body in the digital recognition model through the high-level synthesis tool includes at least one of the following steps: Unroll the loop body into multiple sub-loop bodies with the same structure through the unroll optimization instruction of the high-level synthesis tool; Shorten the instruction trigger interval between loop bodies through the pipeline optimization instruction of the high-level synthesis tool; Make the loop body execute in parallel through the dataflow optimization instruction of the high-level synthesis tool.
6. A digital recognition optimization method based on a high-level synthesis tool according to claim 1, characterized in that The step of converting the optimized digital recognition model into RTL code and exporting it as an IP core specifically includes: Obtain the C language expression form of the optimized digital recognition model; Convert the C language expression form into RTL code through the high-level synthesis tool; Export the RTL code as an IP core callable by the FPGA.
7. A digital recognition optimization method based on a high-level synthesis tool according to any one of claims 1 to 6, characterized in that, The step of establishing a connection between the IP core and the camera module through the FPGA so that the camera module inputs the acquired image data into the IP core for digital recognition specifically includes: Convert the image data acquired by the camera module into first image data in RGB format through the FPGA; Perform filtering processing on the first image data through the FPGA to obtain second image data that conforms to the input image format and input image size; Input the second image data into the IP core through the FPGA and output the corresponding digital recognition result through the HDMI transmission line.
8. A digital recognition optimization system based on a high-level synthesis tool, characterized in that, Including: A model training module for constructing a LeNet-5 convolutional neural network and training the LeNet-5 convolutional neural network according to a preset training data set to obtain a digital recognition model; A model simulation module for representing the digital recognition model in a high-level language and performing simulation verification on the digital recognition model through a high-level synthesis tool; A loop body optimization module for optimizing each layer of loop body in the digital recognition model through the high-level synthesis tool; An IP core export module for converting the optimized digital recognition model into RTL code and exporting it as an IP core; A digital recognition module for establishing a connection between the IP core and the camera module through the FPGA so that the camera module inputs the acquired image data into the IP core for digital recognition.
9. A digital recognition optimization device based on a high-level synthesis tool, characterized in that Including: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a digital recognition optimization method based on a high-level synthesis tool as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to execute a digital recognition optimization method based on a high-level synthesis tool as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Arm architecture-based NumPy operation acceleration optimization method
CN112783503A
Deep learning model optimization method and system based on high-level synthesis tool
CN113780553A