FPGA-based Real-time Classification System for Defective Pills

Through the real-time classification system of tablet residues based on FPGA, the MobileNetV2 network is deployed and storage and calculation is optimized, which solves the problems of slow tablet classification speed and poor real-time performance, and achieves low power consumption and efficient real-time classification.

CN115546529BActive Publication Date: 2025-07-08FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210752137.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-29
Publication Date
2025-07-08
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

The classification speed of pills in the prior art is slow, manual classification is prone to errors and cannot meet the needs of high-speed industrial production. Machine vision classification has a large delay and high power consumption on the CPU, so real-time classification cannot be achieved.

Method used

The real-time classification system for tablet residues based on FPGA is adopted, including image acquisition module, FPGA and result display module, FPGA is used for image processing and hardware acceleration, the MobileNetV2 network is deployed, and the static quantization and parallel pipeline structure is adopted to optimize storage and calculation methods, and 0 neurons in the fully connected layer are eliminated.

Benefits of technology

Effectively reduce power consumption, reduce memory consumption, improve the real-time and applicability of classifiers, and is suitable for small embedded scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546529B_ABST
    Figure CN115546529B_ABST
Patent Text Reader

Abstract

The present invention relates to a real-time classification system for defective tablets based on FPGA, which includes an image acquisition module, an FPGA, and a result display module; the FPGA includes an input image processing module, a DDR, and a hardware acceleration module; the image acquisition module includes a high-speed camera, an FMC, and a display screen; the image acquisition module acquires the surface image of the tablet, sends it into the FPGA through the FMC daughter card, and caches it to the off-chip DDR; the input image processing module reads the image data from the DDR and performs cropping and quantization processing, and stores the processed input image data in the on-chip memory; the hardware acceleration module is pre-deployed with the network structure of MobileNetV2, and the network weights are pre-trained and imported into the on-chip memory according to the read-write order required by the hardware. When it is detected that the input image is loaded, the hardware acceleration part starts forward inference acceleration and transmits the final classification result to the result display module for display. The present invention effectively improves the speed of tablet classification and greatly improves the real-time performance of the classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image classification, and particularly to a real-time classification system for defective tablets based on FPGA. Background Art

[0002] In the industrial production of tablets, defective tablets are inevitable. At present, many pharmaceutical factories still use human eyes to classify defective tablets. Due to the strong subjectivity and easy fatigue of human eyes, phenomena such as missed inspection and misinspection of tablets are inevitable, and the manual sorting speed is slow and difficult to adapt to the high-speed industrial production line. Some pharmaceutical factories use machine vision to detect and classify defective tablets. After collecting images, the image data is transmitted to the upper computer, and classification is performed through a convolutional neural network (CNN) deployed on the upper computer. The upper computer usually deploys the CNN on a traditional central processing unit (CPU) or a graphics processing unit (GPU). Deploying on the CPU will bring high latency, while deploying on the GPU will have relatively high power consumption, and it cannot meet the requirements of long-term real-time classification of defective tablets on the production line. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a real-time classification system for defective tablets based on FPGA, which can effectively improve the speed of tablet classification and greatly improve the real-time performance of the classifier.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions:

[0005] A real-time classification system for defective tablets based on FPGA includes an image acquisition module, an FPGA, and a result display module; the FPGA includes an input picture processing module, a DDR, and a hardware acceleration module; the image acquisition module includes a high-speed camera, an FMC, and a display screen; the image acquisition module acquires the surface image of the tablet, sends it into the FPGA through the FMC daughter card, and caches it in the off-chip DDR; the input image processing module reads the picture data from the DDR and performs cropping and quantization processing, and the processed input picture data is stored in the on-chip memory; the hardware acceleration module is pre-deployed with the network structure of MobileNetV2, the network weights are trained in advance and imported into the on-chip memory according to the read and write order required by the hardware. When it is detected that the input picture is loaded, the hardware acceleration part starts forward inference acceleration and transmits the final classification result to the result display module for display.

[0006] Furthermore, the image acquisition module includes a camera image acquisition part and an image display part;

[0007] The image acquisition part includes a camera for acquiring images, converting the input differential video data into parallel video data, reorganizing the parallel video data, separating the valid data and caching it to the off-chip DDR;

[0008] The image display part includes converting the video data into RGB format, converting the video stream into a parallel video signal and transmitting it to the display screen to display the acquired image effect.

[0009] Furthermore, the quantization adopts a static quantization method, replacing the 32-bit floating-point weights with INT8-type integer data, specifically as follows:

[0010]

[0011] Where q represents the INT8-type number after quantization, r is the floating-point number before quantization, s is the quantization scale, and z is the fixed-point value corresponding to the floating-point number 0;

[0012] s is obtained from formula (2), where max and min respectively refer to the maximum and minimum values of the floating-point number r and the integer number q after quantization

[0013]

[0014] z is obtained from formula (3), and round means rounding the final result

[0015]

[0016] The convolution calculation is regarded as the multiplication of two N×N-type matrices r1 and r2, and r3 is their output result, which is expressed as shown in formula (4)

[0017]

[0018] Substituting formula (1) into formula (4) and organizing it, formula (5) can be obtained

[0019]

[0020] Among them, s1, z1 are the quantization scale factor and zero point corresponding to matrix r1, and similarly, s2, z2 and s3, z3 are the quantization scale factor and zero point corresponding to matrix r2 and r3 respectively; let Find a fixed-point number M0 such that M = 2 can be achieved through shift operations -n M0, then formula (5) is completely converted into fixed-point number operations and implemented on the FPGA.

[0021] Furthermore, the MobilenetV2 includes an inverted residual bottleneck module, specifically:

[0022] First, use the 1×1 convolutional kernel of the PW layer to expand the N-dimensional network to M dimensions, then extract features through the 3×3 convolutional kernel of the DW layer, and finally reduce it to H dimensions by the 1×1 convolutional kernel of the PW layer;

[0023] The output part of the inverted residual bottleneck module is based on the direct connection structure of the residual network. When the middle DW layer does not perform downsampling, the input and output of the bottleneck module are added together. The computing parallelism of the PW is expanded in the way of combining 4 input channels and 4 output channels. Assuming the number of input channels is M, after M / 4 cycles, 4 numbers at the same position of 4 output channels are output; the parallelism of the DW layer should be the same as that of the previous layer, and the 4-input-channel expansion method is adopted.

[0024] Furthermore, the PW layer and the adjacent DW layer of the MobileNetV2 network adopt a parallel pipeline structure. Specifically: the PW layer reads the cached image and weight data of the previous layer from the on-chip memory. After the PW layer calculation array finishes the calculation, the convolutional result is stored in the cache composed of on-chip BRAM, and is stored in the arrangement mode of storing 4-channel data per address; the convolutional kernel size of the DW is 3×3, and here a sliding window structure composed of three groups of shift registers is adopted; every time the PW layer outputs a number to the feature map cache, Shirft-ram reads in a data at the same frequency. After the third layer is filled, the DW layer calculation engine starts, reads a 3×3 data window from Shirft-ram for convolutional operation, and finally writes the result back to the on-chip memory.

[0025] Furthermore, the last layer of the MobileNetV2 network is a fully connected layer. Each node is connected to all nodes of the neurons in the previous layer. Assume the output neurons of the previous layer are x1~x m , and the weights are represented as w 11 ~w nm , the bias is b n , and the output is a n , then the calculation method of the fully connected layer can be expressed as:

[0026] a n =w n1 *x1+w n2 *x2+w n3 *x3+…w nm *x m +b n (6).

[0027] Furthermore, the on-chip memory adopts an improved optimization strategy, which optimizes the storage structure of weights and caches as well as the data reading method, specifically as follows: Assume the size of the input image is k×k×n, and the weight is n×m; If it is necessary to read in the input data of 4 channels within one cycle, then the data of 4 channels are merged and stored in one BRAM, and 32-bit-wide data spliced by 4 channels are stored at each address. Similarly, 4 data of each of the 4 channels are spliced at each address in the weight memory, that is, 128-bit-wide data are stored at one address; The data reading method performs address jump reading according to the size of the input image.

[0028] The present invention has the following beneficial effects compared with the prior art:

[0029] 1. The present invention effectively reduces the power consumption of the classifier and reduces the memory consumption.

[0030] 2. The present invention adopts static quantization to compress the model data, optimizes the storage structure of weights and caches as well as the data reading, making the classifier more suitable for small embedded scenarios.

[0031] 3. The present invention adopts a parallel pipelined structure of PW and DW and eliminates 0 neurons in the fully connected layer during the forward inference process of the MobileNetV2 network to reduce the calculation amount, greatly improving the real-time performance of the classifier. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is the system framework diagram of the present invention;

[0033] Figure 2 is the program block diagram of the image acquisition module in an embodiment of the present invention;

[0034] Figure 3 is the structure diagram of the bottleneck module in an embodiment of the present invention;

[0035] Figure 4 is the parallel pipelined structure diagram of the PW and DW layers in an embodiment of the present invention;

[0036] Figure 5 is the test result diagram of the fully connected layer in an embodiment of the present invention;

[0037] Figure 6 is the calculation module diagram of the fully connected layer in an embodiment of the present invention;

[0038] Figure 7 is the data storage structure diagram in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The present invention will be further described below with reference to the drawings and embodiments.

[0040] Please refer toFigure 1-7 , the present invention provides a real-time classification system for defective tablets based on FPGA, including an image acquisition module, an FPGA, and a result display module; the FPGA includes an input picture processing module, a DDR, and a hardware acceleration module; the image acquisition module includes a high-speed camera, an FMC, and a display screen; the image acquisition module acquires the surface image of the tablet, sends it into the FPGA through the FMC daughter card, and caches it to the off-chip DDR; the input image processing module reads the picture data from the DDR and performs cropping and quantization processing, and stores the processed input picture data into the on-chip memory; the hardware acceleration module is pre-deployed with the network structure of MobileNetV2, the network weights are trained in advance and imported into the on-chip memory according to the read-write order required by the hardware. When it is detected that the input picture is loaded, the hardware acceleration part starts forward inference acceleration, and transmits the final classification result to the result display module for display.

[0041] In this embodiment, the image acquisition module uses an industrial high-speed camera to acquire the surface image of the tablets on the production line. The program block diagram is as Figure 2 shown, including the camera image acquisition part and the image display part. The image acquisition part includes: the camera acquires an image, the input differential video data is converted into parallel video data, the parallel video data is reorganized, the valid data is separated and cached to the off-chip DDR. The image display part includes: the video data is converted into RGB format, the video stream is converted into parallel video signals and transmitted to the display screen to display the acquired image effect.

[0042] In this embodiment, static quantization method is adopted for quantization, and 32-bit floating-point weights are replaced with INT8-type integer data, specifically as follows:

[0043]

[0044] where q represents the quantized INT8-type number, r is the floating-point number before quantization, s is the quantization scale, and z is the fixed-point value corresponding to the floating-point number 0;

[0045] s is obtained by formula (2), where max and min respectively refer to the maximum and minimum values of the floating-point number r and the quantized integer number q

[0046]

[0047] z is obtained by formula (3), and round means rounding the final result

[0048]

[0049] The convolution calculation is regarded as the multiplication of two N×N matrices r1 and r2, and r3 is their output result, which is expressed as shown in formula (4)

[0050]

[0051] Substituting formula (1) into formula (4) and arranging it gives formula (5).

[0052]

[0053] Where s1 and z1 are the quantization scale factor and zero point corresponding to matrix r1. Similarly, s2, z2 and s3, z3 are the quantization scale factor and zero point corresponding to matrix r2 and r3 respectively; let Find a fixed-point number M0 such that M = 2 -n M0 can be achieved through shift operations, then formula (5) is completely converted into fixed-point number operations and implemented on the FPGA.

[0054] In this embodiment, MobilenetV2 is mainly composed of inverted residual bottleneck modules (Inverted residual block);

[0055] First, use the 1×1 convolutional kernel of the PW layer to expand the N-dimensional network to M dimensions, then extract features through the 3×3 convolutional kernel of the DW layer, and finally reduce it to H dimensions by the 1×1 convolutional kernel of the PW layer. The output part of the inverted residual bottleneck module draws on the direct connection structure (short-cut) of the residual network (RseNet). When the middle DW layer does not perform downsampling (stride = 1), the input and output of the bottleneck module are added together. The specific operation is as Figure 3 shown. Here, the computational parallelism of PW is expanded in the way of combining 4 input channels and 4 output channels. Assuming the number of input channels is M, then after M / 4 cycles, 4 numbers at the same position of 4 output channels are output. In order to enable the PW layer of the previous layer and the DW of the next layer to achieve parallel pipelining, the parallelism of the DW layer should be the same as that of the previous layer, and the 4-input channel expansion method is adopted.

[0056] In this embodiment, the PW layer and the adjacent DW layer of the MobileNetV2 network adopt a parallel pipelining structure. The specific structure is as Figure 4As shown, the PW layer reads the cached pictures and weight data of the previous layer from the on-chip memory. After the calculation array in the PW layer finishes the calculation, the convolution result is stored in the cache composed of on-chip BRAMs, and is stored in the arrangement mode of storing 4-channel data per address. The convolution kernel size of the DW is 3×3, and here a sliding window structure composed of three groups of shift registers (Shirft-ram) is adopted. Every time the PW layer outputs a number to the feature map buffer (Feature_map_buffer), the Shirft-ram reads in a data at the same frequency. After the third layer is filled, the calculation engine of the DW layer starts, reads a 3×3 data window from the Shirft-ram for convolution operation, and finally writes the result back to the on-chip memory.

[0057] In this embodiment, the last layer of the MobileNetV2 network is a fully connected layer. Each node is connected to all nodes of the neurons in the previous layer. The number of input neurons is 1280, and it can classify pictures of up to 1000 categories at most. Assume that the output neurons of the previous layer are x1 to x m , and the weights are represented as w 11 ~w nm The bias is b n , and the output is a n , then the calculation method of the fully connected layer can be expressed as:

[0058] a n =w n1 *x1+w n2 *x2+w n3 *x3+…w nm *x m +b n (6)

[0059] It is observed that there are a large number of 0 neurons among the 1280 neurons in the fully connected layer. 100 random trials are carried out on the test set in the dataset, and the test results are as Figure 5 shown. For almost all 100 pictures, the proportion of 0 neurons in the fully connected layer is greater than 40%. If the 0 neurons are removed and the subsequent multiplication operations are carried out, the calculation amount of the fully connected layer can be reduced by nearly half.

[0060] Preferably, the design of the fully connected layer calculation module is as Figure 6 shown, and its implementation steps are as follows:

[0061] 1. The 1280 nerves output by the average pooling layer of the previous layer contain a large number of 0 neurons, and the neurons with 0 are removed through the non-0 monitoring unit.

[0062] 2. Store non-zero neurons in the cached BRAM and reorder the data addresses. To align the weights with the input data, the corresponding data at the same positions in the weight data also needs to be removed, and the addresses need to be reordered.

[0063] 3. The data after removing 0 neurons and the weights at the corresponding positions are read into the multiply-accumulate array simultaneously, and the accumulated result is output after adding the bias.

[0064] In this embodiment, an optimization strategy for the memory is proposed, which optimizes the storage structure of the weights and the cache and the data reading method.

[0065] Taking the PW layer as an example, the storage structure and reading method of the data are as Figure 7 shown. Assume that the size of the input picture (Map) is (k×k×n), and the weight (Weight) is (n×m). The parallelism designed in the PW layer is 16, that is, 16 groups of 1×1 pointwise convolution operations are completed in one clock cycle. It is necessary to read in the input data of 4 channels in one cycle. The simplest way to store the input picture data is to use 4 BRAMs of the same size, and each BRAM stores the data of one channel to achieve the purpose of improving the parallelism. This strategy will cause serious waste of on-chip resources. Therefore, we merge the data of 4 channels and store them in one BRAM. Each address stores 32-bit-wide data spliced from 4 channels. Similarly, each address in the weight memory (Weight-roms) spliced 4 data of 4 channels, that is, one address stores 128-bit-wide data.

[0066] The data reading method performs address jumping reading according to the size of the input picture. For example, if the size of the input picture is (k×k×n), the address jumps and reads according to addr0, addrk,... addr(n-1)k every clock cycle (CLK). The read picture data is multiplexed. In the first CLK, M11~M14 in the inter-layer cache are read in, and at the same time, 4 data at 4 positions of each of the 4 channels are read in. After multiplying W11~W14, W21~W24, W31~W34, and W41~W44 and accumulating them respectively, after the accumulation of n channels is completed, the numbers at the same positions of 4 output channels are output simultaneously, and after splicing, they are stored in one address.

[0067] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. A real-time classification system for defective tablets based on FPGA, characterized in that, It includes an image acquisition module, an FPGA, and a result display module; the FPGA includes an input picture processing module, a DDR, and a hardware acceleration module; the image acquisition module includes a high-speed camera, an FMC, and a display screen; the image acquisition module acquires the surface image of the tablet, sends it into the FPGA through the FMC daughter card, and caches it in the DDR; The input picture processing module reads the picture data from the DDR and performs cropping and quantization processing. The processed input picture data is stored in the on-chip memory; among them, static quantization method is adopted for quantization, and 32-bit floating-point weights are replaced with INT8-type integer data; the hardware acceleration module is pre-deployed with the network structure of MobileNetV2, and the network weights are trained in advance and imported into the on-chip memory according to the read-write order required by the hardware. When it is detected that the input picture is loaded, the hardware acceleration part starts forward inference acceleration, and transmits the final classification result to the result display module for display; The PW layer and the adjacent DW layer of the MobileNetV2 network adopt a parallel pipeline structure, specifically as follows: The PW layer reads the cached picture and weight data of the previous layer from the on-chip memory. After the calculation array of the PW layer finishes calculating, the convolution result is output and stored in the cache composed of on-chip BRAMs, and is stored in the arrangement mode of storing 4 channels of data per address; the convolution kernel size of the DW layer is 3×3, and here a sliding window structure composed of three groups of shift registers is adopted; for each number output by the PW layer to the feature map cache, Shirft-ram reads in a data at the same frequency. After the third layer is filled, the calculation engine of the DW layer starts, reads a 3×3 data window from Shirft-ram for convolution operation, and finally writes the result back to the on-chip memory; The on-chip memory adopts an improved optimization strategy, which optimizes the storage structure of weights and caches and the data reading method, specifically as follows: assume that the size of the input picture is k×k×n, and the weight is n×m; if it is necessary to read 4 channels of input data in one cycle, the data of 4 channels are merged and stored in a BRAM, and 32-bit wide data spliced by 4 channels is stored in each address. Similarly, each address in the weight memory spliced 4 data of each of the 4 channels, that is, 128-bit wide data is stored in one address; the data reading method performs address jump reading according to the size of the input picture.

2. The real-time defective tablet classification system based on FPGA according to claim 1, characterized in that, The image acquisition module includes a camera image acquisition part and an image display part; The camera image acquisition part includes the camera acquiring images, converting the input differential video data into parallel video data, reorganizing the parallel video data, separating the valid data and caching it on the DDR; The image display part includes converting the video data into RGB format, converting the video stream into parallel video signals and transmitting them to the display screen to display the acquired image effect.

3. The real-time classification system for defective tablets based on FPGA according to claim 1, characterized in that, The quantization adopts a static quantization method, and 32-bit floating-point weights are replaced with INT8-type integer data, specifically as follows: Among them, q represents the quantized INT8-type number, r is the floating-point number before quantization, s is the quantization scale, and z is the fixed-point value corresponding to the floating-point number 0; s is obtained from formula (2), where max and min respectively refer to the maximum and minimum values of the floating-point number r and the quantized integer number q z is obtained from formula (3), and round means rounding the final result The convolution calculation is regarded as the multiplication of two N×N matrices r1 and r2, and r3 is their output result, which is expressed as shown in formula (4) Substituting formula (1) into formula (4) and arranging it, formula (5) can be obtained where s1 and z1 are the quantization scale factor and zero point corresponding to matrix r1. Similarly, s2, z2 and s3, z3 are the quantization scale factors and zero points corresponding to matrices r2 and r3 respectively. Let find a fixed-point number M0 such that M = 2 can be achieved through shift operations -n M0, then the formula (5) is completely converted to fixed-point number operations and implemented on the FPGA.

4. The real-time classification system for defective tablets based on FPGA according to claim 1, characterized in that, The MobilenetV2 includes an inverted residual bottleneck module, specifically: First, use the 1×1 convolution kernel of the PW layer to expand the N-dimensional network to M dimensions, then extract features through the 3×3 convolution kernel of the DW layer, and finally reduce it to H dimensions by the 1×1 convolution kernel of the PW layer The output part of the inverted residual bottleneck module is based on the direct connection structure of the residual network. When the middle DW layer does not perform downsampling, the input and output of the bottleneck module are added together. The calculation parallelism of the PW is expanded in the way of combining 4 input channels and 4 output channels. Suppose the number of input channels is M, then after M / 4 cycles, 4 numbers at the same position of 4 output channels are output; the parallelism of the DW layer should be the same as that of the previous layer, and the 4 input channels are used for expansion 5. The real-time classification system for defective tablets based on FPGA according to claim 4, characterized in that The last layer of the MobileNetV2 network is a fully connected layer, where each node is connected to all nodes of the neurons in the previous layer. Let the output neurons of the previous layer be x1 to x m , and the weights are represented as w 11 to w nm , the bias is b n , and the output is a n . Then the calculation method of the fully connected layer can be expressed as: a n = w n1 * x1 + w n2 * x2 + w n3 * x3 + … w nm * x m + b n (6).

Citation Information

Patent Citations

  • FPGA parallel acceleration system based on CNN image quality enhancement algorithm

    CN110084739A

  • Method and apparatus for realizing convolutional neural network, terminal, and storage medium

    WO2019127838A1