A method for realizing handwritten digit recognition

By implementing accelerated calculation of CNN inference process on the FPGA platform, using fixed-point conversion and parallel computing, the shortcomings of handwritten numeric recognition algorithm in recognition speed and accuracy are solved, and efficient hardware acceleration and simplified network structure are achieved.

CN114299514BActive Publication Date: 2025-05-13NANJING UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111488462.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-05-13
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

The existing handwritten numeric recognition algorithms have shortcomings in recognition speed and recognition accuracy, especially when dealing with diverse handwritten characters, the recognition accuracy is not high, and the model cannot be deployed across platforms, making it poor in practicality.

Method used

Through the collaboration of software and hardware, some CNN inference processes are placed in FPGA to accelerate calculations, and fixed-point transformation and parallel computing are used to optimize the structure and calculation process of convolutional neural networks to achieve efficient hardware acceleration for handwritten digit recognition.

Benefits of technology

While taking into account the recognition accuracy, it speeds up the recognition speed, simplifies the convolutional neural network structure, makes it easier to implement on the hardware platform, gives full play to the advantages of FPGA, improves computing efficiency and reduces power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299514B_ABST
    Figure CN114299514B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for realizing handwritten digit recognition, and belongs to the field of image recognition. The present invention mainly deploys a convolutional neural network on a ZYNQ embedded hardware platform, and realizes handwritten digit recognition through software and hardware collaborative acceleration. The present invention first grays and binarizes the input image and matches the recognition frame with the size of the data set image, and then stores the recognition frame image in a BRAM storage unit; then completes convolution operation, activation function, and pooling operation acceleration on the recognition frame image data at the PL end; then constructs the camera timing with the pooled image data and transmits it to the PS end DDR; finally, completes the hidden layer and output layer operations at the PS end, and transmits the recognition results to the PL end to display the recognition results. The present invention can accelerate the reasoning operation of some neural networks and quickly recognize handwritten digits in the picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of image recognition, and in particular relates to a method for realizing handwritten digit recognition. Background Art

[0002] At present, the recognition of human handwriting has become a research hotspot. This technology has been widely used in tax form processing, mail classification, and computer grading. In these applications, handwritten digit recognition algorithms are usually required to have high recognition speed and recognition accuracy as well as high reliability and stability. Although there are only ten types of handwritten digits and the strokes are simple, there are still great difficulties in their recognition. Some existing test results have shown that the correct recognition rate of digits is not as high as the recognition rate of printed Chinese characters, and even lower than the recognition rate of online handwritten Chinese characters. The main reason is that the fonts are not much different, and everyone writes differently, which makes recognition more difficult. How to use equipment to automatically, intelligently, and efficiently recognize numbers and characters to improve work efficiency has become a research problem that needs to be solved urgently.

[0003] The concept of convolutional neural network was proposed by Yann Le Cun of New York University in 1998. It is a neural network that can be successfully applied to many image classifications such as handwritten character recognition and license plate recognition. The essence of convolutional neural network is a kind of multi-layer perceptron. The reason why convolutional neural network can be successfully applied in many aspects is mainly because the sparse connection and parameter sharing of convolutional neural network can greatly reduce the number of parameters in the network, making the computational complexity of training the entire convolutional neural network greatly reduced and practical.

[0004] Handwritten digit recognition is a classic problem in the field of CNN, targeting the Minist dataset, which is widely used in the field of handwritten recognition. Compared with printed digits, handwritten digits are more random and have higher uncertainty, but the recognition of handwritten digits has a wide range of applications in life, such as bank bill recognition, automatic mail sorting, and mobile phone number recognition. Due to the huge number of parameters and computational complexity of CNN, the model cannot be deployed across platforms and has poor practicality. On the one hand, various lightweight and miniaturized network models have been designed. On the other hand, due to the data independence of convolution operations, parallel acceleration of convolution operations becomes possible. Therefore, it is of great significance to implement hardware implementation of handwritten digit recognition.

[0005] The implementation of CNN model is divided into two parts: training and reasoning. At present, the acceleration of model training stage is carried out by CPU+GPU, and good results have been achieved. Therefore, this paper mainly studies the hardware acceleration of model reasoning stage. In the reasoning stage of the model, FPGA platform has become a research platform for CNN hardware acceleration due to its advantages of high performance, low power consumption, reconfigurability, and the ability to implement domain-specific architecture on existing devices without the need to develop new chips. When designing hardware, based on the large amount of convolution operations and the large number of parameters, by analyzing the hardware structure of the convolution acceleration unit, the storage and data transmission characteristics, the relationship between the acceleration performance of the model and the resources and bandwidth of the acceleration platform, the relationship between the power consumption of the model and the data flow between the modules in the model, and the relationship between the accelerator structure, performance and power consumption are comprehensively considered.

[0006] However, since CNN calculations involve many different types of operations, such as two-dimensional convolution operations, nonlinear activation function operations, pooling (subsampling) operations, and fully connected layer operations, and these operations often involve a large amount of data access and storage of intermediate result data, it is still a challenging task to use FPGA to implement such a complex and computationally intensive CNN. At the same time, a large number of floating-point calculations will cause accuracy problems. Therefore, how to implement a convolutional neural network system for handwritten digit recognition on FPGA with high performance has important theoretical research significance and practical value for the development of the image recognition field. Summary of the invention

[0007] The purpose of the present invention is to propose a method for realizing handwritten digit recognition, which uses the collaborative method of software and hardware to put part of the CNN reasoning process into FPGA to accelerate the calculation, while taking into account the recognition accuracy and speeding up the recognition speed.

[0008] The technical solution to achieve the purpose of the present invention is: a method for realizing handwritten digit recognition, comprising the following steps:

[0009] Step 1: Obtain image data, complete video data acquisition at the PL end, and complete image preprocessing in the digital acquisition area;

[0010] Step 2: Store the preprocessed data in the BRAM storage unit at the PL end;

[0011] Step 3: Perform fixed-point conversion on the convolutional neural network parameters and store them in the ROM IP unit in the embedded platform ZYNQ that integrates ARM and FPGA;

[0012] Step 4: Construct a corresponding convolution matrix according to the convolution kernel size of the convolutional neural network to complete the convolution operation of the preprocessed data and the fixed-point parameters;

[0013] Step 5: Activate the convolutional layer operation results and perform maximum pooling operation on the activation results;

[0014] Step 6: Construct the corresponding video timing for the pooled data, and use the Video In to AXI4-Stream IP core and VDMA IP core to transmit the pooling results to the PS end DDR;

[0015] Step 7. The PS side completes the hidden layer and output layer operations of the convolutional neural network, and transmits the recognition results to the PL side through AXI-lite. The PL side drives the display to realize the display function of the recognition results on the display screen.

[0016] Preferably, the image preprocessing performed on the digital acquisition area in step 1 includes: grayscale conversion, smoothing and noise reduction, binarization, and downsampling processing.

[0017] Preferably, in step 3, the weight parameter values ​​of the convolutional neural network are processed as floating-point numbers using fixed-point operations, and the fixed-point numbers are obtained by expanding the weight parameters by a certain integer multiple.

[0018] Preferably, 5 ROM IP cores are used to store weight parameters, the size of the convolution kernel is 5*5, and each ROM stores 150 weight parameters, which are arranged in a 5*5 matrix.

[0019] Preferably, for the convolution kernel size of the convolutional neural network, a corresponding convolution matrix is ​​constructed to complete the convolution operation of the preprocessed data and the fixed-point parameters, including: the input structure of the convolution kernel and the input structure of the image data, which are respectively specifically:

[0020] Convolution kernel input structure: set the parameter storage arrangement order, complete the convolution kernel arrangement through ROM IP parallel splicing; when performing convolution operation, start reading the input convolution kernel weight in ROM, stop reading weight after the corresponding clock cycle, cache the read weight in the register unit, and ensure that the weight value of a single convolution kernel remains unchanged; after the convolution operation traversal of the entire input image, output the corresponding feature map; read other convolution kernel weights and perform convolution operation with the input image;

[0021] The steps of constructing the input of image data are as follows: constructing a convolution window through a shift register cache, shifting and delaying the input image data to obtain a pixel matrix window of the input image data, multiplying and accumulating the weight parameters of the corresponding position of the convolution kernel template with the input pixel matrix, and using the result as the output result of the window center to complete a convolution operation and obtain the first output element of the first feature map; sliding the matrix window of the input pixel row by row in the input image area, and performing a convolution operation with the first layer of convolution kernel to obtain the first feature map; using the convolution kernel to complete the convolution operation in sequence to obtain the corresponding feature map.

[0022] Preferably, in step 5, the result of the convolution layer is processed using the activation function: ReLu(x)=max(0,x); the pooling layer is designed using 2×2 maximum pooling; the design of the pooling window is consistent with the convolution window, and both use a register delay method; the convolution layer and the activation function have a total pipeline delay of 5 clocks, and the image data is numbered in rows and columns before entering the convolution layer, and after a delay of 5 clocks, it enters the pooling layer synchronously with the data output by the activation function; the row and column numbers of the counting delay are respectively shifted right by one position to obtain the pooling index number, and the pooling window is pooled with this number, and a pooling operation is performed every two clocks.

[0023] Compared with the prior art, the present invention has the following significant advantages: (1) The present invention optimizes the hardware implementation method for handwritten digit recognition, and accelerates the recognition speed while taking into account the recognition accuracy through the coordination of software and hardware; (2) The present invention simplifies the original handwritten digit recognition algorithm, thereby making the entire convolutional neural network structure simpler and easier to implement on the hardware platform, giving full play to the advantages of FPGA, and at the same time utilizing the parallelism of FPGA operations to speed up the operation speed and reduce power consumption.

[0024] Implementation The present invention is further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A block diagram of a handwritten digit recognition implementation method provided by the present invention.

[0026] Figure 2 This is a calculation flow chart of the convolutional neural network of the present invention.

[0027] Figure 3 This is the CNN flow chart adopted by the present invention.

[0028] Figure 4 This is a flow chart of the convolution module operation of the present invention.

[0029] Figure 5 The logic for implementing the ReLu function of the present invention.

[0030] Figure 6 This is the operation process of the pooling module of the present invention.

[0031] Figure 7 This is a handwritten digit recognition effect diagram of the present invention. DETAILED DESCRIPTION

[0032] In one embodiment, a method for implementing handwritten digit recognition mainly includes the following steps:

[0033] Step 1: Acquire image data, complete video data acquisition at the PL end, and complete image preprocessing in the digital acquisition area.

[0034] Preprocessing is a series of processing work on the pictures containing numbers. The primary goal of the entire preprocessing is to discard the invalid information in the picture, and at the same time standardize and unify the image, which will facilitate the subsequent processing. It usually includes: grayscale conversion, smoothing and noise reduction, binarization, and downsampling. In accordance with the data size requirements of the convolutional neural network input layer, the image of the acquisition area is compressed to a size of 28*28;

[0035] Step 2: Store the preprocessed data in the BRAM storage unit at the PL end;

[0036] Step 3: Perform fixed-point conversion on the convolutional neural network parameters and store them in the ROM storage unit of the embedded platform ZYNQ that integrates ARM and FPGA.

[0037] Specifically, since the weight parameters of the convolutional network layer have a decimal part, floating-point operations are required. A large number of floating-point operations will consume a lot of computing time. The present invention uses fixed-point operations to process the weight parameter values ​​of the convolutional neural network in floating-point form. Fixed-point numbers are obtained by expanding the weight parameters by a certain integer multiple, and then these fixed-point parameters are stored in a ROM memory unit. When the input image data is convolved, these fixed-point parameters are read from the ROM to perform convolution operations with the input image. In this way, floating-point operations can be converted into fixed-point operations between integers, which speeds up the calculation process. By reasonably setting the expansion multiple, the balance between calculation accuracy and resource consumption can be effectively controlled.

[0038] The present invention uses 5 ROM IP cores to store weight parameters. The size of the convolution kernel is 5*5, and a total of 30 convolution kernels are constructed. Each convolution kernel has 25 weight parameters, so a total of 750 weight parameters need to be stored. Each ROM stores 150 weight parameters, and the arrangement order is arranged in a 5*5 matrix, that is, the first ROM stores the first row of parameters of each convolution kernel, and each convolution kernel matrix has 5 weight parameters in one row; and so on to the last ROM to store the fifth row of weight parameters of each convolution kernel.

[0039] Step 4: Perform convolution operation on the data. According to the convolution kernel size of the convolutional neural network, 30 convolution matrices of size 5*5 are constructed to complete the convolution operation;

[0040] In this embodiment, according to the convolution kernel size of the convolutional neural network, a 5*5 convolution matrix is ​​constructed, the reading timing of the preprocessed data stored in the BRAM and the weight parameters stored in the ROM is controlled, and the convolution operation of 30 convolution kernels is completed. It is specifically divided into the following two steps: the input construction of the convolution kernel and the input of the image data.

[0041] Input structure of convolution kernel. The size of the convolution kernel is 5*5, that is, 25 data, corresponding to 25 weights, and it remains unchanged throughout the process of the same input image. In order to ensure that the preprocessed data stored in the BRAM and the weight parameters stored in the ROM correspond one to one when reading, the present invention requires the parameter storage arrangement order in step 4, and the 30 5*5 convolution kernels are arranged in parallel through 5 ROM IPs. When performing convolution operations, start reading the input convolution kernel weights in 5 ROMs, and stop reading the weights after 5 clock cycles. The 25 read weights are cached in the register unit, and the weight value of a single convolution kernel is guaranteed to remain unchanged. After the convolution operation is traversed for a whole 28*28 input image, a 24*24 feature map is output. After that, other convolution kernel weights are read and convolved with the input image.

[0042] In the process of constructing the convolution kernel input, two points must be ensured. First, the order of the convolution kernel input process must be correct, and second, the position of the input convolution kernel must be correct (neither exceeded nor not reached). Therefore, during the convolution kernel input process, a convolution kernel valid signal is counted at the same time, represented by FLAG, so that the effective time of a convolution kernel is exactly 24*24 clock cycles. Within the effective time, the convolution kernel just reaches the correct position, and then the FLAG is invalidated to ensure that the convolution kernel remains unchanged.

[0043] The image signal is input serially row by row. When performing the convolution operation step, it is necessary to use the grayscale values ​​of multiple neighboring pixels for calculation at the same time, so a 5*5 matrix window matching the size of the convolution kernel is constructed for the preprocessed image data. Specific construction steps: Construct the convolution window through the shift register cache. The size of the shift register is 28 pixels, and a total of four shift registers are required. The input image data is shifted and delayed to obtain a 5*5 pixel matrix window of the input image data. The weight parameters at the corresponding position of the convolution kernel template are multiplied and accumulated with the input 5*5 pixel matrix. The result is used as the output result of the window center, thereby completing a convolution operation and obtaining the first output element of the first feature map. The 5*5 matrix window of the input pixel is slid row by row in the 28*28 input image area, and convolution operation is performed with the first layer of convolution kernel, and finally the first 24*24 feature map is obtained. Considering the balance between area and speed, if all operations will consume a large number of multipliers, 30 convolution kernels are used to complete the convolution operation in sequence. Finally, 30 feature maps of size 24*24 are obtained.

[0044] Although there are many multipliers used in the convolution layer operation module, in terms of time consumption, the method proposed in the present invention utilizes the characteristics of FPGA parallel operation, and can operate on 25 pixels in the same clock. At the same time, in order to avoid combinatorial logic delay, a four-level pipeline is used to split the data stream during timing design. The input of each level of pipeline operation comes from the corresponding upper-level output and does not affect each other. The final output is only delayed by 4 clock times compared to the pixel input. The overall implementation can output the convolution result in each clock cycle, greatly improving the overall operation efficiency of the system.

[0045] Step 5: Activate the convolutional layer operation results and perform maximum pooling operation on the activation results;

[0046] In this embodiment, the 30 24*24 matrix feature map results obtained after the convolution layer operation are activated, and the activated results are subjected to 2*2 maximum pooling. The result of the convolution layer is processed using the activation function: ReLu(x)=max(0,x). Choose to design the pooling layer using 2×2 maximum pooling. Since pooling also processes serial data streams, the design of the pooling window is consistent with the convolution window, and both use register delay. The convolution layer and the activation function have a total pipeline delay of 5 clocks. In order to ensure the continuity of the data flow, the image data is numbered in rows and columns before the image enters the convolution layer, and after a delay of 5 clocks, it enters the pooling layer synchronously with the data output by the activation function. The row and column numbers of the counting delay are shifted right by one position to obtain the index number of the pooling. The pooling window is pooled with this number, and the pooling operation is performed every two clocks.

[0047] Step 6: Construct the corresponding video timing for the pooled data, and use the Video In toAXI4-Stream IP core and VDMA IP core to transmit the pooled results to the PS end DDR;

[0048] Step 7: The PS side completes the hidden layer and output layer operations of the convolutional neural network, and finally transmits the recognition results to the PL side through AXI-lite. The PL side drives the display to realize the display function of the recognition results on the display screen.

Claims

1. A method for realizing handwritten digit recognition, characterized in that: The following steps are involved: Step 1: Obtain image data, complete video data acquisition at the PL end, and complete image preprocessing in the digital acquisition area; Step 2: Store the preprocessed data in the BRAM storage unit at the PL end; Step 3: Perform fixed-point conversion on the convolutional neural network parameters and store them in the ROM IP unit in the embedded platform ZYNQ that integrates ARM and FPGA; Step 4: Construct the corresponding convolution matrix according to the convolution kernel size of the convolution neural network, and complete the convolution operation of the preprocessed data and the fixed-point parameters, including: the input structure of the convolution kernel and the input structure of the image data, which are respectively: Convolution kernel input structure: set the parameter storage arrangement order, complete the convolution kernel arrangement through ROM IP parallel splicing; when performing convolution operation, start reading the input convolution kernel weight in ROM, stop reading weight after the corresponding clock cycle, cache the read weight in the register unit, and ensure that the weight value of a single convolution kernel remains unchanged; after the convolution operation traversal of the entire input image, output the corresponding feature map; read other convolution kernel weights and perform convolution operation with the input image; The steps of constructing the input of image data are as follows: constructing a convolution window through a shift register cache, shifting and delaying the input image data to obtain a pixel matrix window of the input image data, multiplying and accumulating the weight parameters of the corresponding position of the convolution kernel template with the input pixel matrix, and using the result as the output result of the window center to complete a convolution operation and obtain the first output element of the first feature map; sliding the matrix window of the input pixel row by row in the input image area, and convolving it with the first layer of convolution kernel to obtain the first feature map; using the convolution kernel to complete the convolution operation in sequence to obtain the corresponding feature map; Step 5: Activate the convolutional layer operation results and perform maximum pooling operation on the activation results; Step 6: Construct the corresponding video timing for the pooled data, and use the Video In to AXI4-Stream IP core and VDMA IP core to transfer the pooled results to the PS end DDR; Step 7. The PS side completes the hidden layer and output layer operations of the convolutional neural network, and transmits the recognition results to the PL side through AXI-lite. The PL side drives the display to realize the display function of the recognition results on the display screen.

2. The method for realizing handwritten digit recognition according to claim 1, characterized in that: The image preprocessing performed on the digital acquisition area in step 1 includes: grayscale conversion, smoothing and noise reduction, binarization, and downsampling processing.

3. The method for realizing handwritten digit recognition according to claim 1, characterized in that: Step 3 uses fixed-point operations to perform floating-point processing on the weight parameter values ​​of the convolutional neural network, and obtains fixed-point numbers by expanding the weight parameters by a certain integer multiple.

4. The method for realizing handwritten digit recognition according to claim 1, characterized in that: Use 5 ROM IP cores to store weight parameters. The size of the convolution kernel is 5*5. Each ROM stores 150 weight parameters, which are arranged in a 5*5 matrix.

5. The method for realizing handwritten digit recognition according to claim 1, characterized in that: In step 5, the result of the convolution layer is processed using the activation function: ReLu(x)=max(0,x); the 2×2 maximum value pooling is selected to design the pooling layer; the design of the pooling window is consistent with the convolution window, and both use the register delay method; the convolution layer and the activation function have a total pipeline delay of 5 clocks, and the image data is numbered in rows and columns before entering the convolution layer, and after a delay of 5 clocks, it enters the pooling layer synchronously with the data output by the activation function; the row and column numbers of the counting delay are shifted right by one position to obtain the pooling index number, and the pooling window is pooled with this number, and the pooling operation is performed every two clocks.

Citation Information

Patent Citations

  • Hand-printed character recognition method based on FPGA+ARM multilayer convolution neural network

    CN106250939A