Rice disease and pest classification method and system based on FPGA acceleration

By adopting hierarchical array-partitioning method and ping-pong buffering technology on the FPGA platform, combining multi-scale matrix blocking and three-channel parallel convolution methods, the resource utilization imbalance and computing bottleneck problems of large-scale CNN models deployed on the FPGA platform are solved, and the low power consumption, high real-time and low-cost requirements for rice leaf pest classification are achieved.

CN120219795APending Publication Date: 2025-06-27HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204406.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently deploy and infer large-scale CNN models on the FPGA platform, resulting in unbalanced resource utilization, computing bottlenecks and data transmission delays, and cannot meet the low power consumption, high real-time and low cost requirements for rice leaf pest classification.

Method used

The hierarchical array-partitioning method and ping-pong buffering technology are used to store the image data, network parameters and convolution windows output by the memory DDR into the programmable logic unit. The calculation efficiency is improved through multi-scale matrix blocking and three-channel parallel convolution methods, FPGA computing resources are dynamically managed, the execution order of computing tasks is optimized, and the pipeline and module multiplexing method of convolution operations are combined.

Benefits of technology

It improves the computing throughput of the accelerator and the utilization rate of computing resources, improves the parallelism and processing speed of model inference, realizes low power consumption, high real-time and efficient classification, reduces off-chip storage access delay and energy consumption, and provides low-cost and highly reliable pest detection solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219795A_ABST
    Figure CN120219795A_ABST
Patent Text Reader

Abstract

The invention discloses a paddy rice disease and insect pest classification method and system based on FPGA acceleration. The paddy rice disease and insect pest classification system comprises an image acquisition module, a preprocessing module, a memory DDR, a processing system, a programmable logic unit and an AXI bus. Image data output by a memory DDR and a convolution window are stored in a programmable logic unit by adopting a hierarchical array-partition method, a rice disease and insect pest classification model is calculated and decomposed into feature map sub-blocks capable of being executed in parallel, and the calculation efficiency is improved through a three-channel parallel convolution method. Memory bandwidth occupation is reduced through efficient reuse of network parameters and feature map sub-blocks; meanwhile, by optimizing the execution sequence of different calculation tasks in the rice disease and insect pest classification model and combining a pipeline of convolution operation, the calculation throughput of an accelerator and the utilization rate of calculation resources are effectively improved, and a reliable solution is provided for agricultural disease and insect pest detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of FPGA computing acceleration, and in particular relates to a rice pest and disease classification method and system based on FPGA acceleration. Background Art

[0002] With the rapid development of artificial intelligence and computer vision technology, smart agriculture has an increasingly urgent need for efficient and accurate detection and classification of pests and diseases. As a major food crop in the world, the early identification of rice foliar pests and diseases is of great significance for improving agricultural production efficiency and ensuring food security. CNN (convolutional neural network) is widely used in pest and disease classification tasks due to its excellent performance in image classification and feature extraction. However, the high computational complexity and large amount of data of the CNN model place extremely high demands on the computing power and data bandwidth of the hardware platform.

[0003] Although existing high-performance computing devices such as GPUs can meet the computing requirements of complex CNN reasoning, their high energy consumption, large size, and high cost limit their actual deployment in edge devices in agricultural scenarios. In addition, current mainstream edge devices are usually limited by limited hardware resources and data bandwidth. When performing complex CNN reasoning, data bandwidth constraints lead to frequent off-chip storage accesses, causing high storage latency and energy consumption problems, which seriously affect the real-time performance and power efficiency of the system.

[0004] At the same time, FPGA (Field Programmable Gate Array), as a low-power, highly parallel and flexible programmable hardware acceleration platform, has shown great potential in the field of edge computing. However, the limited computing resources and on-chip storage capacity of FPGA make it difficult to directly support the reasoning tasks of large-scale CNN models, such as uneven resource utilization, computing bottlenecks and data transmission delays. How to efficiently deploy CNN models on FPGA to meet the low power consumption, high real-time performance and low cost requirements of rice leaf pest classification has become an important problem that needs to be solved in the current technological development. Summary of the invention

[0005] The purpose of the present invention is to provide a rice pest and disease classification method and system based on FPGA acceleration.

[0006] In a first aspect, the present invention provides a rice pest classification method based on FPGA acceleration, which comprises the following steps:

[0007] Step 1: Obtain a labeled rice image dataset, preprocess the rice image dataset, use the preprocessed dataset to train a rice pest and disease classification model, and obtain network parameters of the rice pest and disease classification model;

[0008] Step 2: Perform fixed-point quantization on the network parameters;

[0009] Step 3: Collect the images of the rice to be measured, and perform fixed-point quantization on the images of the rice to be measured to obtain image data;

[0010] Step 4: Input the fixed-point quantized image data and network parameters into the programmable logic unit of the FPGA; use the multi-scale matrix block division method to divide the image data into different feature map sub-blocks and store them in the random access memory of the programmable logic unit; store the convolution windows of the convolution kernels in the rice pest and disease classification model in multiple linear buffers of the programmable logic unit through the three-channel parallel convolution method;

[0011] Step 5: Use the programmable logic unit to perform target detection on the images of the rice to be measured according to the feature map sub-blocks and network parameters, obtain and store the output feature maps; decode all the output feature maps and convert them into corresponding image classification labels to obtain the rice pest and disease classification results.

[0012] Preferably, in Step 4, an intermediate buffer is provided in the programmable logic unit; the intermediate buffer includes a weight buffer and a data buffer; when storing the traditional convolution layer and the depth convolution layer, the network parameters are stored in the weight buffer and the image data is stored in the data buffer; when storing the point convolution layer and the fully connected layer, the network parameters are stored in the data buffer and the image data is stored in the weight buffer.

[0013] Preferably, a plurality of convolution units for convolution calculation are provided in the programmable logic unit, and a buffer is allocated to each calculation complex provided in the calculation unit, and the data in each buffer is subjected to convolution calculation in the same stage.

[0014] Preferably, in Step 5, the calculation layer is tiled and parallelly divided according to the number of output channels, and each calculation unit is responsible for the calculation of part of the output channels and processes the part of the weight kernels related to the output channels it is responsible for to generate the corresponding output feature maps.

[0015] Preferably, the specific process of performing target detection on the images of the rice to be measured in Step 5 is as follows:

[0016] Step 5-1: According to the network structure of the rice pest and disease classification model, allocate the calculation tasks of each layer to the programmable logic unit;

[0017] Step 5-2: Use the programmable logic unit to perform convolution calculation on the feature map sub-blocks to obtain the output feature maps; after the calculation of one feature map sub-block is completed, directly store or transfer the obtained output feature maps to the next layer for calculation; repeat the above process until the calculation of all the feature map sub-blocks is completed.

[0018] Preferably, in the first step, the lightweight convolutional neural network is trained in the DIST manner.

[0019] Preferably, in the third step, before performing fixed-point quantization on the measured rice image, the measured rice image is cropped by the block cropping method, and the cropped image is normalized.

[0020] In a second aspect, the present invention provides a rice pest and disease classification system based on FPGA acceleration, which is used to execute the above-mentioned rice pest and disease classification method based on FPGA acceleration; the rice pest and disease classification system includes an image acquisition module, a preprocessing module, a memory DDR, a processing system PS, a programmable logic unit PL, and an AXI bus for data transmission; the image acquisition module is used to acquire the measured rice image; the preprocessing module is used to crop and segment the labeled rice image; the processing system PS is used to deploy the rice pest and disease classification model in the programmable logic unit PL and implement the logic control of the rice pest and disease classification model; the memory DDR is used to store the output feature map, as well as the fixed-point quantized network parameters and image data; the programmable logic unit PL includes a block random access memory BRAM and a hardware accelerator block; the block random access memory BRAM is used to input the network parameters and image data stored in the memory DDR into the hardware accelerator block, and input the output feature map obtained by the hardware accelerator block into the memory DDR; the hardware accelerator block includes a random access memory, a linear buffer, and a plurality of computing units for performing convolution calculations; the random access memory is used to store sub-blocks of the feature map; the linear buffer is used to store the convolution window of the convolution kernel in the rice pest and disease classification model.

[0021] Preferably, the block random access memory BRAM includes two intermediate buffers. In one buffer cycle, one intermediate buffer is used to store the image data and network parameters output by the memory DDR, and the other intermediate buffer is used to output the stored image data and network parameters to the hardware accelerator block; in the next buffer cycle, the functions of the two intermediate buffers are reversed; the intermediate buffer that stored data in the previous cycle is used to output the stored data to the hardware accelerator block; the intermediate buffer that output data in the previous cycle is used to store the data output by the memory DDR.

[0022] Preferably, the rice pest and disease classification system further includes a visualization module, and the visualization module reads the corresponding pre-stored image according to the label obtained by the processing system PS to complete the visualization of the rice pest and disease classification detection result.

[0023] The beneficial effects of the present invention are:

[0024] 1. The present invention adopts a hierarchical array - partitioning method and ping - pong buffering technology to store the image data, network parameters, and convolution window output from the memory DDR into the programmable logic unit PL. By decomposing the calculation of the rice pest and disease classification model into feature map sub - blocks that can be executed in parallel and using a three - channel parallel convolution method to improve the calculation efficiency, and by making efficient reuse of network parameters and feature map sub - blocks to reduce the memory bandwidth occupancy and lower the off - chip storage access latency.

[0025] 2. The present invention effectively improves the calculation throughput and utilization rate of computing resources of the accelerator by dynamically managing the FPGA computing resources, optimizing the execution order of different calculation tasks in the rice pest and disease classification model, and combining the pipelining and module reuse methods of convolution operations, enhancing the parallelism and processing speed of model inference, and ensuring that the system can achieve efficient classification under low power consumption and high real - time performance, thus providing a reliable solution for agricultural pest and disease detection.

[0026] 3. The high - speed rice disease detection system constructed by the present invention using the low - power and high - bandwidth characteristics of FPGA realizes the efficient classification and real - time processing of rice disease images, significantly improves the data transmission efficiency, reduces the calculation latency, and increases the system throughput. At the same time, the present invention introduces knowledge distillation and static fixed - point quantization technologies, greatly reducing the power consumption and storage requirements of hardware calculations, ensuring that the system has high - precision detection capabilities while being low - cost, and providing a low - cost and highly reliable solution for the precise detection and control of agricultural diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following - described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 It is the structural block diagram of the rice pest and disease classification system in the present invention.

[0029] Figure 2 It is the flow chart of optimizing the accelerator structure by adopting the ping - pong operation method in the present invention.

[0030] Figure 3 It is the schematic diagram of the method for optimizing the accelerator by adopting the multi - scale matrix partitioning method in the present invention.

[0031] Figure 4 It is the schematic diagram of the multi - channel line buffer method in the present invention.

[0032] Figure 5Schematic diagram of the method for optimizing the accelerator by adopting the DDR3 access mode of the memory in the present invention.

[0033] Figure 6 Schematic diagram of the method for optimizing the accelerator by adopting the parallel collaborative computing mode in the present invention. Detailed implementation manners

[0034] The present invention will be further described below with reference to the accompanying drawings.

[0035] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0036] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plurality" is two or more.

[0037] In the description of the present invention, it should be noted that, unless otherwise clearly defined and limited, the terms "install", "connect", "couple" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal connection of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.

[0038] Such as Figure 1As shown in the figure, a rice pest and disease classification system based on FPGA acceleration includes an image acquisition module, a preprocessing module, an external storage module, a memory DDR3 (Double Data Rate), a DDR3 control module, a processing system PS, a programmable logic unit PL, an AXI (Advanced eXtensible Interface) bus, and a visualization module. The image acquisition module uses a CMOS (Complementary Metal Oxide Semiconductor) sensor; the preprocessing module is used to crop and segment the labeled rice images; the external storage module is used to store the quantized image data and network parameters; the DDR3 control module is used to control the memory DDR3 to store and transmit the image data and network parameters; the processing system PS is used to process the rice images input by the memory DDR3 and implement the logical control of the neural network inference structure; the programmable logic unit PL includes a block random access memory BRAM and a hardware accelerator block; the block random access memory BRAM is used to input the network parameters and image data stored in the memory DDR3 into the hardware accelerator block and input the output feature map obtained by the hardware accelerator block into the memory DDR3; the hardware accelerator block is used to realize data reuse and parallel acceleration of different neural network modules; the hardware accelerator block is provided with a RAM (Random Access Memory) and a linear buffer; the random access memory is used to store the image data; the linear buffer is used to store the convolution window of the convolution kernel in the rice pest and disease classification model; the AXI bus is used for data transmission between different modules; the visualization module is used to perform output image visualization processing.

[0039] The rice pest and disease classification method adopted by the rice pest and disease classification system based on FPGA acceleration includes the following steps:

[0040] Step 1: Obtain the labeled rice image dataset, and crop and segment the rice images in the dataset through the preprocessing module; use the DIST (Knowledge Distillation from A Stronger Teacher) method to train the lightweight rice pest and disease classification model to obtain the network parameters of the trained rice pest and disease classification model, that is, the weights and biases of the rice pest and disease classification model; in this embodiment, the rice pest and disease classification model uses a convolutional neural network. Training by the DIST method can improve the network accuracy.

[0041] In some embodiments, the process of training the convolutional neural network by the DIST method is as follows:

[0042] By unifying the deep teacher network and complex training strategies, taking the change of the teacher network output as the optimization target, and increasing the weight ratio of the relationship between the student network and the teacher network. Given the traditional KD loss (Knowledge Distillation Loss) The expression is as follows:

[0043]

[0044] where KL is the KL divergence loss function (Kullback-Leibler Divergence Loss); Y t and Y s are the prediction vectors of the teacher model and the student network respectively; T is the temperature factor; x is the class index.

[0045] According to the traditional KD loss and the Pearson correlation coefficient, the inter-class relation loss and intra-class relation loss are calculated. The Pearson correlation coefficient ρ p (u, v) is expressed as:

[0046]

[0047] where u and v are two different random variables; C is the number of classes; Cov(u, v) is the covariance of the random variables u and v; Std(u) and Std(v) are the standard deviations of the random variables u and v respectively; u i is the i-th class in the random variable u; v i is the i-th class in the random variable v similarly; and are the means of the random variables u and v respectively; i = 1, 2,..., C.

[0048] By the relationship loss, the student is enabled to adaptively match the output of the teacher network, thereby greatly improving the distillation performance. The final total loss consists of three parts:

[0049]

[0050] where are the inter-class relation loss, intra-class relation loss and total loss respectively; B is the batch size; is the original classification loss between the student prediction value and the true value; α, β and γ are the three hyperparameter weights for balancing the loss; j = 1, 2,..., B.

[0051] In other embodiments, other existing methods can also be used to train the convolutional neural network.

[0052] Step 2: Fixed-point quantization

[0053] Through static fixed-point quantization, the network parameters of the trained neural network are converted into 16-bit fixed-point numbers. According to the representation structure of floating-point numbers, extract their sign, exponent, and mantissa parts, and determine the target fixed-point number format. The numerical range is expanded by multiplying by 2^n, and the floating-point decimal is mapped to the fixed-point integer range. Here, n is the fractional bit width of fixed-point quantization; multiplying by 2^n can be implemented by a hardware shift operation, such as shifting left by n bits to amplify the data and reduce resource consumption. The amplified data needs to be truncated in bit width (6-bit integer, 10-bit fraction) to retain 16-bit data: remove the redundant high bits to avoid overflow; remove the low bits to limit the precision and ensure that the data adapts to the fixed-point format. At the same time, the sign bit is reserved to correctly represent positive and negative values (in two's complement form). Finally, all computing units complete the data format conversion with a unified bit width and sign representation. The converted network parameters are stored in the external storage module according to the network layer structure, and the stored network parameters are verified in CRC (Cyclic Redundancy Check) format.

[0054] Step 3: Use the image acquisition module to capture the image of the rice to be measured, and use the processing system PS to crop the image of the rice to be measured into the format of 224×224 in a block-by-block cropping manner. Normalize the pixel values of the cropped image in the range of [-1, 1], and convert them into fixed-point format or floating-point format to adapt to subsequent calculations to obtain image data. The processed image data is stored in the external storage module through an asynchronous FIFO (First Input First Output) buffer mechanism; through the asynchronous FIFO buffer mechanism, the data writing efficiency can be improved, and the storage delay caused by hardware performance limitations can be avoided.

[0055] Step 4: Before performing target detection on the image of the rice to be measured, both the image data and network parameters in the external storage module are stored in the memory DDR3. As Figure 2As shown, the image data and network parameters stored in the memory DDR3 are input into the hardware accelerator block through the block random access memory BRAM by means of ping-pong operation; the block random access memory BRAM includes two intermediate buffers, and both of the two intermediate buffers include a weight buffer and a data buffer; in the bottleneck structure (conventional convolutional layer and depth convolutional layer), the amount of network parameter data required in each stage is much smaller than the amount of image data; while in the PW layer (point convolutional layer) and the fully connected layer, the amount of network parameter data is much larger than the amount of input data. Since the storage capacity of the designed data buffer is greater than that of the weight buffer, when storing the point convolutional layer and the fully connected layer, the weight data is stored in the feature map buffer with a larger capacity, and the input data is stored in the smaller weight buffer to optimize the data access efficiency.

[0056] In this embodiment, the specific process of storing network parameters in the block random access memory BRAM is as follows:

[0057] 4-1. Configure the initialization parameters of the memory DDR3, including clock frequency, CAS (Compare-And-Swap) latency, refresh period, etc., and start the self-check function of the memory DDR3 to calibrate the address lines and data lines to ensure the correctness of read and write operations and lock the clock signal.

[0058] 4-2. Initialize the base address and offset address registers of the memory DDR3 control module. At the same time, set the starting read and write addresses of the memory DDR3 to ensure that the reading is performed in the storage order of the network parameters.

[0059] 4-3. Issue a read command for the memory DDR3, and at the same time configure the burst read mode (BurstRead) of the memory DDR3. According to the logical address generator, gradually access the physical addresses where the network parameters are stored.

[0060] 4-4. According to the preset read logic address order, transmit the read data stream to the on-chip logic through the AXI bus interface of the memory DDR3 controller. According to the hierarchical structure of the network, such as the parameter distribution of the convolutional layer and the fully connected layer, store the read parameters into different block random access memory BRAM areas block by block. And use the address mapping table to map the physical addresses of the network parameters to the block random access memory BRAM logical addresses one by one. Perform format conversion on the read data and align it according to the bit width requirements of the in-chip computing unit.

[0061] 4-5. After each data block transmission is completed, update the starting address of the next block of data in the memory DDR3 through the address generator and mark the address of the completed area.

[0062] 4-6. Compare the checksum of the read data with the original parameters to ensure the accuracy of the read data.

[0063] In a buffer cycle, one intermediate buffer is used to store the image data and network parameters output by the memory DDR3, and the other intermediate buffer is used to transfer the stored image data and network parameters to the hardware accelerator block; in the next buffer cycle, the functions of the two intermediate buffers are reversed; the intermediate buffer that stored data in the previous cycle is used to output data to the hardware accelerator block; the intermediate buffer that output data in the previous cycle is used to store the data output by the memory DDR3. In this way, the data stream seamlessly switches between the two intermediate buffers, ensuring continuous data processing.

[0064] The image data and the convolution window of the convolution kernel are respectively stored in the random access memory and the linear buffer through the hierarchical array-partition method; as Figure 3 shown, the image data input into the random access memory is processed by using the multi-scale matrix block method, and the specific process of the multi-scale matrix block method is as follows:

[0065] According to the width col o and height row o of the output feature map, as well as the input channel N i and output channel N o of the convolution kernel, the input image data is segmented into different feature map sub-blocks, and they are stored in multiple smaller random access memories. The hardware accelerator block performs sub-block convolution calculation and processing to meet the requirements of batch data reading, and is also applicable to the multi-data reuse format; the size of the feature map sub-block is:

[0066]

[0067] B_num = B rn × B cn × N bi × N bo

[0068] where is the width of the output feature map of the l-th layer; is the width of the input feature map of the l-th layer; is the width dimension of the input convolution kernel of the l-th layer; p is the number of zero paddings of the input feature map; s is the convolution stride; is the height of the output feature map of the l-th layer; is the height of the input feature map of the l-th layer; is the height dimension of the input convolution kernel of the l-th layer; T ro is the width of the output feature map sub-block; T ri is the width of the input feature map sub-block; T cois the height of the output feature map sub - block; B rn is the number of sub - blocks in the width direction; B cn is the number of sub - blocks in the height direction; N bi is the number of sub - blocks of the convolutional kernel in the input channels; is the number of input channels of the l - th layer; T n is the value of the sub - block input channels; N bo is the number of sub - blocks of the convolutional kernel in the output channels; is the number of output channels of the l - th layer; T m is the value of the sub - block output channels; B_num is the total number of the cropped sub - blocks.

[0069] As Figure 4 shown, considering that the convolutional sizes used in the depth convolutional layers of the bottleneck structure are all 3×3, the inference is performed using a convolutional kernel of the same size with three channels. The convolutional windows of the convolutional kernel are stored in the linear buffer A and the linear buffer B respectively; this strategy utilizes the three - channel convolutional kernel, where the overlapping region of the linear buffer A is reused. Therefore, a copy of the feature map in the linear buffer A is first made, and then a one - dimensional concatenation of the three is performed with the linear buffer B, so as to be effectively stored in a buffer with a size 2 - 3 times the width T of the sub - block ro of the sub - block. By designing multiple linear buffers in the hardware accelerator block to store network parameters, memory access is minimized and latency is reduced. The convolutional window is created in each clock cycle for subsequent calculations.

[0070] Step Five: Use the hardware accelerator block to perform object detection on the measured rice image according to the image data and network parameters, and store the detection results in the memory DDR3. By dynamically managing the FPGA computing resources, the execution order of different computing tasks such as convolutional layers and fully - connected layers is optimized. Combining with the hardware pipeline mechanism, the computing throughput of the accelerator is maximized. The specific process of performing object detection on the measured rice image is as follows:

[0071] 5 - 1. Use the processing system PS to complete the logical control of the CNN inference structure. According to the network structure of the convolutional neural network through the processing system PS, the computing tasks of each layer are assigned to the programmable logic unit PL; through the interrupt or DMA start signal, the programmable logic unit PL is triggered to execute the corresponding computing tasks.

[0072] 5 - 2. When the traditional technology performs the calculation of the fully - connected layer, the data is cached in a linearly arranged data manner. Each time data is loaded, 64×128 = 8192 DMA (Direct Memory Access) transactions are required, and each burst length is only 8. This method results in extremely low bandwidth utilization of the block random access memory BRAM.

[0073] As Figure 5 shown, the present invention assigns a buffer with a length of 1024 to each of the 49 computing complexes in each PE (computing unit) in the hardware accelerator block, which is the same as the scale strategy of matrix partitioning. When calculating the ordinary convolutional layer and the depth convolutional layer, the buffer is filled one by one. To reduce the additional data routing logic for filling the buffer and maintain a long burst length when obtaining data for calculating the point convolutional layer, the weight matrix is arranged in the block random access memory BRAM. First, the entire computing unit is divided into 49×8 columns and 128 rows of computing complexes, so that one computing complex can be processed in one stage. In each computing complex, the data arrangement is as Figure 5 shown. According to Figure 5 this arrangement of data, only one DMA transaction is needed to load the entire computing complex, and the burst length can cover the size of the entire block, greatly extending the burst transfer length and significantly improving the data transfer efficiency and bandwidth utilization; in hardware implementation, this arrangement reduces the number of DMA transactions and does not require complex data routing logic when filling the buffer. The buffer within each computing module is efficiently loaded through a continuous long burst data stream, thereby maximizing the potential of the block random access memory BRAM bandwidth.

[0074] 5-3. As Figure 6 shown, the computing layer is unrolled (tiled) and parallel partitioned according to the depth of the output channels (i.e., the number of output channels), and each computing unit is responsible for calculating a part of the output channels. Each computing unit will access the input feature map and process its partial weight kernels; each computing unit only processes the partial weight kernels related to the output channels it is responsible for. Each computing unit independently calculates the output channels it is responsible for and generates the corresponding part of the output feature map. And the number of weight kernels assigned to each computing unit is proportional to the computing parallelism of the tiling or pipeline of this computing unit. This allocation method evenly distributes the computing tasks to multiple computing units, ensuring load balancing among different units, that is, all units can make full use of their computing capabilities, avoiding the problem of some units' resources being idle or some units being overloaded, and significantly improving the parallel processing ability and resource utilization efficiency.

[0075] 5-4. Store the output feature map generated by each computing unit into the block random access memory BRAM. When a sub-block of the feature map is calculated, transfer the output feature map stored in the block random access memory BRAM to the memory DDR3 or directly transfer it to the next layer of calculation, and clear the internal data of the block random access memory BRAM; repeat the above process until the calculation of all sub-blocks of the feature map is completed.

[0076] Step 6: Input all the output feature maps stored in the memory DDR3 into the processing system PS for decoding and convert them into the format corresponding to the image classification labels. Transmit the inference result to the visualization module; read the corresponding pre-stored image according to the target label to complete the visualization of the final detection result of rice leaf diseases and pests classification.

[0077] Step 7: Method evaluation

[0078] 7-1. Comparison of resource usage of multi-channel line buffer

[0079] The results of the overall circuit delay and hardware resource utilization before and after using the multi-channel line buffer are shown in Table 1. It can be seen from Table 1 that the optimization using the multi-channel line buffer can reduce the inference delay by about 41.48%. This optimization significantly increases the resource consumption, and the utilization rates of DSP (multiplier-accumulator resources) and FF (flip-flops) are increased by 40.26% and 21.69% respectively, resulting in the overall area utilization approaching saturation. The multi-channel line buffer method effectively reduces the duplicate data transmission between the memory DDR3 and the accelerator and improves the recognition speed.

[0080] Table 1 Comparison of resource usage before and after using multi-channel line buffer

[0081]

[0082] 7-2. Compare the resource utilization rates of different schemes respectively, and the results are shown in Table 2; it can be seen from Table 2 that compared with the acceleration scheme without any optimization, the utilization rates of DSP and FF units of the present invention are significantly increased by 44% and 21% respectively, and the utilization rates of FPGA resources such as DSP, BRAM, and LUT are close to saturation; compared with the DPU-P scheme, there is a significant improvement, reducing the waste of hardware resources, resulting in a 51% reduction in inference delay and an improvement in image processing efficiency.

[0083] Table 2 Comparison results of resource utilization rates of different schemes

[0084]

[0085]

[0086] 7-3. Compare the energy consumption of different schemes respectively, and the results are shown in Table 3. It can be seen from Table 3 that the DSP resources of the present invention, especially the addressing and calculation operations across different convolutional layers, have a significant increase in power consumption due to the increase in utilization rate, and the signal power consumption and logic power consumption increase by 0.26W and 0.12W respectively.

[0087] Table 3. Comparison of energy consumption before and after optimization

[0088]

[0089] The embodiments of the present invention have been described in detail with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments, including components, still fall within the protection scope of the present invention.

Claims

1. A rice pest classification method based on FPGA acceleration, characterized by: The following steps are involved: Step 1: Obtain a labeled rice image dataset, preprocess the rice image dataset, use the preprocessed dataset to train a rice pest and disease classification model, and obtain network parameters of the rice pest and disease classification model; Step 2: Fixed-point quantization of network parameters; Step 3: Collect the rice image to be tested, and obtain image data after performing fixed-point quantization on the rice image to be tested; Step 4: Input the fixed-point quantized image data and network parameters into the programmable logic unit of the FPGA; The image data is divided into different feature map sub-blocks by using a multi-scale matrix block method and stored in the random access memory of the programmable logic unit; The convolution windows of the convolution kernel in the rice pest classification model are stored in multiple linear buffers of the programmable logic unit through a three-channel parallel convolution method; Step 5: Use the programmable logic unit to perform target detection on the tested rice image according to the feature map sub-blocks and network parameters, obtain the output feature map and store it; decode all the output feature maps and convert them into corresponding image classification labels to obtain the classification results of rice pests and diseases.

2. The rice pest classification method based on FPGA acceleration according to claim 1, characterized in that: In the step 4, an intermediate buffer is provided in the programmable logic unit; the intermediate buffer includes a weight buffer and a data buffer; when storing traditional convolutional layers and deep convolutional layers, the network parameters are stored in the weight buffer, and the image data is stored in the data buffer; when storing point convolutional layers and fully connected layers, the network parameters are stored in the data buffer, and the image data is stored in the weight buffer.

3. The rice pest classification method based on FPGA acceleration according to claim 1, characterized in that: The programmable logic unit is provided with a plurality of convolution units for convolution calculations, and a buffer is allocated to each calculation complex provided in the calculation unit, and the data in each buffer is convolutionally calculated in the same stage.

4. The rice pest classification method based on FPGA acceleration according to claim 3 is characterized in that: In the step five, the computing layer is tiled and divided into blocks in parallel according to the number of output channels. Each computing unit is responsible for the calculation of part of the output channels, and processes part of the weight kernel related to the output channel it is responsible for, and generates a corresponding output feature map.

5. The rice pest classification method based on FPGA acceleration according to claim 1, characterized in that: In the step 5, the specific process of performing target detection on the detected rice image is as follows: Step 5-1. According to the network structure of the rice pest classification model, the computing tasks of each layer are assigned to the programmable logic unit; Step 5-2. Use the programmable logic unit to perform convolution calculation on the feature map sub-block to obtain the output feature map; after the calculation of a feature map sub-block is completed, the obtained output feature map is directly stored or passed to the next layer of calculation; repeat the above process until the calculation of all feature map sub-blocks is completed.

6. The rice pest classification method based on FPGA acceleration according to claim 1, characterized in that: In the step 1, the lightweight convolutional neural network is trained using the DIST method.

7. The rice pest classification method based on FPGA acceleration according to claim 1, characterized in that: In the step three, before the measured rice image is subjected to fixed-point quantization, the measured rice image is cropped using a block cropping method, and the cropped image is normalized.

8. A rice pest classification system based on FPGA acceleration, characterized by: Used to execute the rice pest classification method based on FPGA acceleration as described in claim 1; the rice pest classification system includes an image acquisition module, a preprocessing module, a memory DDR, a processing system PS, a programmable logic unit PL and an AXI bus for data transmission; the image acquisition module is used to acquire images of the tested rice; The preprocessing module is used to crop and segment the rice image after labeling processing; the processing system PS is used to deploy the rice pest and disease classification model in the programmable logic unit PL and realize the logic control of the rice pest and disease classification model; the memory DDR is used to store the output feature map and the network parameters and image data after fixed-point quantization; the programmable logic unit PL includes a block random access memory BRAM and a hardware accelerator block; the block random access memory BRAM is used to input the network parameters and image data stored in the memory DDR into the hardware accelerator block, and input the output feature map obtained by the hardware accelerator block into the memory DDR; the hardware accelerator block includes a random access memory, a linear buffer and a plurality of computing units for convolution calculation; the random access memory is used to store feature map sub-blocks; The linear buffer is used to store the convolution window of the convolution kernel in the rice pest and disease classification model.

9. The rice pest classification system based on FPGA acceleration according to claim 8, characterized in that: The block random access memory BRAM includes two intermediate buffers. In a buffer cycle, one intermediate buffer is used to store image data and network parameters output by the memory DDR, and the other intermediate buffer is used to output the stored image data and network parameters to the hardware accelerator block; in the next buffer cycle, the functions of the two intermediate buffers are reversed; the intermediate buffer that stores data in the previous cycle is used to output the stored data to the hardware accelerator block; the intermediate buffer that outputs data in the previous cycle is used to store data output by the memory DDR.

10. The rice pest classification system based on FPGA acceleration according to claim 8, characterized in that: It also includes a visualization module, which reads the corresponding pre-stored image according to the label obtained by the processing system PS to complete the visualization of the rice pest and disease classification detection results.