A recognition method for a ground-based cloud image recognition model deployed on FPGA
By deploying a foundation cloud map recognition model based on residual network on FPGA, the problem that the foundation cloud map recognition algorithm in the prior art is difficult to deploy on FPGA, high-precision recognition and portable edge computing are realized, and identification efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202410085558.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-01-22
AI Technical Summary
Existing foundation cloud map recognition algorithms are difficult to deploy on the edge of FPGA hardware, and lack portability and adaptability.
The foundation cloud map recognition model based on residual network is adopted. By building a convolutional layer, a maximum pooling layer, a residual module, an adaptive average pooling layer and a fully connected layer, it is trained and verified, the optimal model weight parameters are obtained, and quantized, and deployed on the PS and PL ends of the FPGA to achieve the generation and return of the recognition results.
The recognition accuracy of the foundation cloud map recognition model is improved, and the edge deployment on FPGA is realized, making the product portable and easy to deploy the dots, and the model acceleration is realized on the PL end of the FPGA, reducing the consumption of FPGA resources.
Smart Images

Figure CN118262140B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a recognition method for a ground-based cloud image recognition model deployed on an FPGA. Background Art
[0002] More than 60% of the earth's surface is covered by clouds. Clouds regulate short-wave radiation and long-wave radiation, which in turn has a huge impact on the earth's hydrological cycle, climate and energy balance. Therefore, accurate identification of cloud types is of great significance to climate prediction and meteorological research. The World Meteorological Organization divides clouds into 3 families, 10 genera and 29 categories based on their shape, structure, characteristics and height. At present, the identification of ground-based clouds is mainly classified according to cloud genera, that is, 10 categories of clouds. Ground-based clouds have the characteristics of many types, fast changes, similarity, and easy integration with the sky background. In addition, clouds of the same type will also show certain differences due to the influence of factors such as region, weather, and time. This brings great difficulties to the refined identification of cloud shapes in ground-based cloud images.
[0003] At present, most meteorological observation stations rely on observers to manually complete cloud classification. Due to the large number of ground-based cloud types and their complexity and variability, the accuracy of recognition is difficult to guarantee. In order to solve the drawbacks of manual recognition, in recent years, more and more researchers have begun to conduct research on automated cloud recognition. However, all of these studies are based on PCs, which require the collected images to be uploaded to a computer or server, and the algorithms are complex and do not take into account the actual application environment.
[0004] For meteorological research in remote areas, if the recognition of ground-based cloud images is achieved by using a PC or uploading images to a server, there will be problems with poor portability and high network environment requirements, making it difficult to deploy multiple points across the country. At present, ground-based cloud image recognition algorithms are all implemented using PC platforms or servers. The recognition algorithms are not lightweight models, and are written and verified using Python and Matlab. These algorithms are difficult to implement on the edge of FPGA hardware. Summary of the invention
[0005] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a recognition method for a ground-based cloud image recognition model deployed on an FPGA, thereby solving the technical problem that the existing recognition algorithms are difficult to implement edge deployment on FPGA hardware.
[0006] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0007] The present invention provides a recognition method for a ground-based cloud image recognition model deployed on an FPGA, comprising:
[0008] Construct a ground-based cloud image recognition model based on residual network;
[0009] Training and verifying the ground-based cloud image recognition model to obtain optimal model weight parameters;
[0010] Quantifying the optimal weight parameters of the model;
[0011] The ground cloud image to be identified and the model weight parameters of the quantization process are read through the PS end of the FPGA, and loaded into the storage DDR;
[0012] The ground-based cloud image recognition model is deployed through the PL end of the FPGA. The PL end reads the ground-based cloud image to be recognized and the model weight parameters of the quantization process from the PS end, generates a recognition result and returns the recognition result to the PS end.
[0013] Optionally, the ground-based cloud image recognition model includes a convolutional layer, a maximum pooling layer, a residual module Res_a, a residual module Res_b, an adaptive average pooling layer, a fully connected layer and a Softmax classifier connected in sequence;
[0014] The residual module Res_a includes a residual unit Res_a1, a residual unit Res_a2, and a ReLu unit connected to the residual unit Res_a1 and the residual unit Res_a2; the residual module Res_b includes a residual unit Res_b1, a residual unit Res_b2, a residual unit Res_b3, and a ReLu unit connected to the residual unit Res_b1, the residual unit Res_b2, and the residual unit Res_b3;
[0015] The residual unit Res_a1 and the residual unit Res_b1 have the same structure, both including a first branch and a second branch, the first branch is provided with a Conv-BN unit, the second branch is provided with a Conv-BN-ReLu unit and a Conv-BN unit in sequence, and the outputs of the first branch and the second branch are added; the residual unit Res_a2, the residual unit Res_b2 and the residual unit Res_b3 have the same structure, both including a third branch and a fourth branch, the fourth branch is provided with a Conv-BN-ReLu unit and a Conv-BN unit in sequence, and the outputs of the third branch and the fourth branch are added.
[0016] Optionally, the training and verification of the ground-based cloud image recognition model to obtain optimal model weight parameters includes:
[0017] Get a dataset of labeled ground-based cloud images;
[0018] Expanding the data set using at least one of an image blurring method, a contrast enhancement method, a mirror flipping method, and a Gaussian noise adding method;
[0019] Dividing the expanded data set into a training set and a test set according to a preset ratio;
[0020] The ground-based cloud image recognition model is trained and verified according to the training set and the test set.
[0021] Optionally, the reading of the ground-based cloud image to be identified and the quantized model weight parameters through the PS end of the FPGA and loading them into the storage DDR includes:
[0022] The storage DDR is pre-divided into a cloud map data area, a convolutional layer weight area, a BN layer weight area, a fully connected layer weight area, and a feature map data area;
[0023] Loading the ground-based cloud image to be identified directly into the cloud image data area;
[0024] Extracting the convolution layer weights corresponding to each convolution layer from the quantized model weight parameters, storing and sorting the convolution layer weights of the convolution kernels of each input channel corresponding to each output channel in order according to the convolution layer weights contained in the convolution kernel, and loading them into the convolution layer weight area after storage and sorting;
[0025] Extract the BN layer weights and BN layer biases corresponding to each BN layer from the quantized model weight parameters, and store and sort the BN layer weights of the convolution kernels of each input channel corresponding to each output channel in sequence according to the BN layer weights contained in the convolution kernels; store and sort the BN layer biases of the convolution kernels of each input channel corresponding to each output channel in sequence according to the BN layer biases contained in the convolution kernels; place the storage and sorting of the BN layer biases after the storage and sorting of the BN layer weights, and load them into the BN layer weight area;
[0026] The fully connected layer weights and fully connected layer biases corresponding to each fully connected layer are extracted from the quantized model weight parameters, and the fully connected layer weights of the convolution kernels of each input channel corresponding to each output channel are sequentially stored and sorted according to the fully connected layer weights contained in the convolution kernels; the fully connected layer biases of the convolution kernels of each input channel corresponding to each output channel are sequentially stored and sorted according to the fully connected layer biases contained in the convolution kernels; the storage and sorting of the fully connected layer biases are placed after the storage and sorting of the fully connected layer weights, and loaded into the fully connected layer weight area.
[0027] Optionally, during the recognition process, the ground-based cloud image recognition model sequentially caches the feature maps inputted by each layer in the ground-based cloud image recognition model to the feature map data area according to the forward reasoning process of the neural network.
[0028] Optionally, the PL end reads the ground-based cloud image to be identified from the PS end and the model weight parameter for quantization processing includes:
[0029] A read-write control IP core is constructed based on the AXI4 protocol, wherein the read-write control IP core includes a read-write control module and a feature map data write FIFO, a feature map data read FIFO, a convolutional layer weight read FIFO, a BN layer weight read FIFO, and a BN layer bias read FIFO connected thereto;
[0030] The model weight parameters of the quantization processing are read from the storage DDR of the PS end through the read-write control module, and are cached in the corresponding convolution layer weight read FIFO, the BN layer weight read FIFO and the BN layer bias read FIFO.
[0031] Optionally, during the recognition process, the ground-based cloud image recognition model follows the forward reasoning process of the neural network.
[0032] The feature graphs inputted from each layer in the ground-based cloud image recognition model are sequentially read from the storage DDR of the PS end through the read-write control module, and are cached in the feature graph data read FIFO;
[0033] The feature map data is written into FIFO to cache the feature maps output by each layer in the ground-based cloud map recognition model, and is loaded into the storage DDR of the PS end through the read-write control module.
[0034] Optionally, the deploying the ground-based cloud image recognition model through the PL end of the FPGA includes:
[0035] Constructing the IP core of the ground-based cloud image recognition model; the IP core includes a sliding window module, a convolution module, a pooling module, a BN layer module, an activation function module and a fully connected layer module;
[0036] The sliding window module is used to count each pixel of the input feature map according to rows and columns respectively through a row counter and a column counter. Whenever the column counter counts to the column size of the feature map, the column counter is reset to zero and the row counter is incremented by one; based on the counting result, n-1 FIFOs are used to implement matrix sliding of a sliding window size n×n for the input data;
[0037] The convolution module is used to realize the sliding of the matrix of the sliding window size m×m through the sliding window module for the convolution layer with a convolution kernel size of m×m, an input channel of a, and an output channel of b; and execute in parallel for a feature maps: multiply the pixel points corresponding to the matrix by the convolution layer weights through m×m multipliers, and execute b times in total;
[0038] The pooling module is used to, for a maximum pooling layer with a pooling kernel size of s×s, implement a matrix sliding of a sliding window size of s×s through the sliding window module; obtain the maximum value of each row of pixel points corresponding to the matrix in parallel through s maximum value modules, and obtain the maximum value of each row maximum value through 1 maximum value module; for an average pooling layer with a pooling kernel size of r×r, implement a matrix sliding of a sliding window size of r×r through the sliding window module; obtain the average value of each row of pixel points corresponding to the matrix through r average value modules in parallel, and obtain the average value of the average values of each row through 1 maximum value module;
[0039] The BN layer module is used to multiply each pixel of the input feature map by the corresponding BN layer weight through a multiplier, and add the multiplication result to the corresponding BN layer bias through an adder;
[0040] The activation function module is used to input each pixel of the feature map through the operation module. If the value of the pixel is greater than or equal to 0, the operation module outputs the value of the pixel; if the value of the pixel is less than 0, the operation module outputs 0;
[0041] The fully connected layer module is used to calculate the probabilities of p assigned categories through p parallel probability calculation modules, and to obtain the maximum value of the probabilities of the p assigned categories through a maximum value calculation module; the input of each probability calculation module is the feature map of each channel, and the input feature map is multiplied by the corresponding fully connected layer weight in sequence through a multiplier, the multiplication results of the feature maps of each channel are added through an adder, and the addition result is added to the corresponding fully connected layer bias through an adder to output a probability value.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] The present invention provides a recognition method for a ground-based cloud image recognition model deployed on an FPGA. The ground-based cloud image recognition model is constructed by constructing a residual network to improve the recognition accuracy of the model. At the same time, the ground-based cloud image recognition model is deployed on the FPGA, so that the product has the characteristics of portability and easy deployment. At the same time, the model is deployed on the PL end of the FPGA to accelerate the model. The weight and ground-based cloud image are deployed on the PS end of the FPGA to reduce the consumption of FPGA resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flow chart of a recognition method of a ground-based cloud image recognition model deployed on an FPGA provided by an embodiment of the present invention;
[0045] Figure 2 It is a structural schematic diagram of a ground-based cloud image recognition model provided by an embodiment of the present invention;
[0046] Figure 3is a schematic diagram of a confusion matrix obtained through a test set provided by an embodiment of the present invention;
[0047] Figure 4 is a schematic diagram of memory area division provided by an embodiment of the present invention;
[0048] Figure 5 is a schematic diagram of an example of the arrangement order of convolutional layer weights provided by an embodiment of the present invention;
[0049] Figure 6 Schematic diagram of the structure of the read-write control IP core provided by an embodiment of the present invention;
[0050] Figure 7 is an example schematic diagram of a characteristic graph serial input provided by an embodiment of the present invention;
[0051] Figure 8 is a schematic structural diagram of a 3×3 sliding window module provided in an embodiment of the present invention;
[0052] Fig. 9 It is a schematic diagram of the operation of a three-dimensional convolution cycle of the first convolution layer provided by an embodiment of the present invention;
[0053] Fig.10 Schematic diagram of the structure of the maximum pooling module provided by an embodiment of the present invention;
[0054] Fig.11 is a schematic diagram of the structure of an average pooling module provided in an embodiment of the present invention;
[0055] Fig.12 is a schematic diagram of the structure of a BN layer module provided in an embodiment of the present invention;
[0056] Fig.13 is an example schematic diagram of the output of the ReLu activation function provided by an embodiment of the present invention;
[0057] Fig.14 is a schematic diagram of the structure of a fully connected layer module provided by an embodiment of the present invention;
[0058] Fig.15 is a schematic diagram of the structure of a probability calculation module provided by an embodiment of the present invention;
[0059] Fig.16 It is a structural schematic diagram of a ground-based cloud image recognition system based on FPGA provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0060] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0061] Embodiment 1:
[0062] like Figure 1 As shown, the present invention provides a recognition method for a ground-based cloud image recognition model deployed on an FPGA, comprising the following steps:
[0063] Step S1, constructing a ground-based cloud image recognition model based on a residual network;
[0064] like Figure 2 As shown, specifically in this embodiment, the ground cloud image recognition model includes a convolution layer Conv-BN-ReLu, a maximum pooling layer MaxPool, a residual module Res_a, a residual module Res_b, an adaptive average pooling layer Adapt-AvgPool, a fully connected layer and a Softmax classifier Flatten-Linear-Softmax connected in sequence;
[0065] The residual module Res_a includes a residual unit Res_a1, a residual unit Res_a2, and a ReLu unit connected to the residual unit Res_a1 and the residual unit Res_a2; the residual module Res_b includes a residual unit Res_b1, a residual unit Res_b2, a residual unit Res_b3, and a ReLu unit connected to the residual unit Res_b1, the residual unit Res_b2, and the residual unit Res_b3;
[0066] The residual unit Res_a1 and the residual unit Res_b1 have the same structure, both including a first branch and a second branch, a Conv-BN unit is arranged on the first branch, a Conv-BN-ReLu unit and a Conv-BN unit are arranged in sequence on the second branch, and the outputs of the first branch and the second branch are added; the residual unit Res_a2, the residual unit Res_b2 and the residual unit Res_b3 have the same structure, both including a third branch and a fourth branch, a Conv-BN-ReLu unit and a Conv-BN unit are arranged in sequence on the fourth branch, and the outputs of the third branch and the fourth branch are added;
[0067] Among them, the convolution kernel size of the convolution layer is 7×7, the convolution kernel size of the maximum pooling layer is 3×3, the convolution kernel size of the Conv-BN-ReLu unit is 1×1, and the convolution kernel size of the Conv-BN unit is 3×3; the input of this embodiment is a 3-channel RGB image of 224×224. First, it passes through the convolution layer and the pooling kernel maximum pooling layer in sequence to extract shallow information and reduce the size of the feature map. Then, the residual modules Res_a and Res_b are used to extract deep features and output 256-channel 14×14 feature maps. Then, the feature map is downsampled to a size of 1×1 using an adaptive average pooling layer, and finally, the prediction probability of 11 types of ground-based cloud maps is output through a fully connected layer and a Softmax classifier. After all convolution operations, the feature maps are batch normalized (BN) to improve the generalization performance and training speed of the model. The residual module uses 1×1 and 3×3 convolution kernels to complete the convolution operation, and improves performance by continuously deepening the network depth. Each convolution layer is followed by a ReLu activation function that is easy to implement in hardware.
[0068] Step S2: training and verifying the ground-based cloud image recognition model to obtain the optimal model weight parameters;
[0069] The specific process is:
[0070] Step S2.1, obtaining a dataset of labeled ground-based cloud images;
[0071] Step S2.2, using at least one of an image blurring method, a contrast enhancement method, a mirror flipping method, and a Gaussian noise adding method to expand the data set; the expansion target number set in this embodiment is 12715;
[0072] Step S2.3, dividing the expanded data set into a training set and a test set according to a preset ratio; the ratio set in this embodiment is 9:1;
[0073] Step S2.4: train and verify the ground-based cloud image recognition model based on the training set and the test set.
[0074] like Figure 3 As shown in the figure, the confusion matrix of the true value and the predicted value is obtained through the test set. The confusion matrix is used to evaluate the performance of the model and reflects the accuracy of the prediction of each category. The darker the color of the matrix diagonal, the better the recognition effect. The accuracy of each category in the figure is about 96%, indicating that this model has a high recognition rate for all categories. The precision, recall rate and F1 value of each category are shown in Table 1.
[0075] Table 1: Precision, recall and F1 value
[0076] Table 1 Precision, recall, and F1 value
[0077]
[0078] The precision rate reflects the probability that the samples predicted as positive are actually positive, and the recall rate reflects the probability that the samples actually positive are predicted as positive. The F1 value is a combination of the precision rate and the recall rate, which can reflect both the precision rate and the recall rate. The ground-based cloud image recognition model proposed in the present invention has shown excellent performance in precision rate, recall rate and F1 value, especially the recognition precision rate of contrail cloud reaches 100%. In addition, in order to better evaluate the overall performance of the model, the macro average and micro average are used to calculate the overall precision and recall rate, and the macro average and micro average are averaged. The overall performance evaluation results are shown in Table 2.
[0079] Table 2 Overall performance evaluation results
[0080]
[0081] The overall score of the model ranges from 95.93% to 96.06%, with an average score of 96.02%, which proves the effectiveness and excellent generalization of the model. Compared with existing network models, it has higher accuracy and has obvious advantages in the field of ground-based cloud image recognition.
[0082] Step S3, quantizing the optimal model weight parameters;
[0083] Since PC training uses 32-bit floating point numbers, FPGA cannot directly operate on floating point numbers, so the trained weight parameters need to be quantized in advance on the PC. Using the symmetric quantization method to quantize the weight values, mapping the weight parameters of each layer to between -255 and 255, and converting the weight values to 16-bit integer data can effectively improve the computing speed and reduce resource usage.
[0084] Step S4: read the ground cloud image to be identified and the model weight parameters of the quantization processing through the PS end of the FPGA, and load them into the storage DDR;
[0085] The forward reasoning process of the convolutional neural network is performed layer by layer. The output of the previous layer is used as the input of the next layer. The calculations between layers are data dependent, and it is necessary to use storage DDR (DDR3 is used in this embodiment) to cache a large amount of data. In addition, the convolution layer, BN layer and fully connected layer all require weight input. In order to ensure that the data does not interfere with each other, the memory area is divided into 5 parts, including the cloud map data area, the convolution layer weight area, the BN layer weight area, the fully connected layer weight area and the feature map data area. Figure 4 shown.
[0086] (1) Load the ground-based cloud image to be identified directly into the cloud image data area;
[0087] (2) extracting the convolutional layer weights corresponding to each convolutional layer from the quantized model weight parameters, and storing and sorting the convolutional layer weights of the convolutional kernels of each input channel corresponding to each output channel in order according to the convolutional layer weights contained in the convolutional kernels, and loading them into the convolutional layer weight area after storage and sorting;
[0088] Taking the case where the number of input channels is 3, the number of output channels is 64, and the convolution kernel size is 7×7 as an example, the convolution layer has a total of 49×3×64=9408 weight parameters. The distribution of these weight parameters in DDR3 is as follows: Figure 5 As shown. The first row on the left of the figure, 0 to 49 × 3 × 1-1, represents the relative address of the weight parameters of the three input channels corresponding to the first output channel. There are 64 rows in total, corresponding to 64 output channels. The 49 data in each row of each part corresponds to a 7 × 7 convolution kernel. The right side of the figure shows the relative addresses corresponding to all weight parameters of the convolution layer.
[0089] (3) Extract the BN layer weights and BN layer biases corresponding to each BN layer from the quantized model weight parameters, and store and sort the BN layer weights of the convolution kernels of each input channel corresponding to each output channel in order according to the BN layer weights contained in the convolution kernels; store and sort the BN layer biases of the convolution kernels of each input channel corresponding to each output channel in order according to the BN layer biases contained in the convolution kernels; sort the storage of the BN layer biases after the storage of the BN layer weights, and load them into the BN layer weight area;
[0090] (4) Extract the fully connected layer weights and fully connected layer biases corresponding to each fully connected layer from the quantized model weight parameters, and store and sort the fully connected layer weights of the convolution kernels of each input channel corresponding to each output channel in sequence according to the fully connected layer weights contained in the convolution kernels; store and sort the fully connected layer biases of the convolution kernels of each input channel corresponding to each output channel in sequence according to the fully connected layer biases contained in the convolution kernels; and place the storage and sorting of the fully connected layer biases after the storage and sorting of the fully connected layer weights, and load them into the fully connected layer weight area.
[0091] In addition, during the recognition process, the ground-based cloud image recognition model caches the feature maps inputted by each layer in the ground-based cloud image recognition model into the feature map data area in sequence according to the forward reasoning process of the neural network.
[0092] Step S5: deploy the ground-based cloud image recognition model through the PL end of the FPGA. The PL end reads the ground-based cloud image to be recognized and the model weight parameters of the quantization processing from the PS end, generates the recognition result and returns the recognition result to the PS end.
[0093] Step S5.1, the PL end reads the ground-based cloud image to be identified from the PS end and the model weight parameters of the quantization processing include:
[0094] The read / write control IP core is built based on the AXI4 protocol. The read / write control IP core includes the read / write control module and its connected feature map data write FIFO, feature map data read FIFO, convolution layer weight read FIFO, BN layer weight read FIFO and BN layer bias read FIFO, such as Figure 6 As shown;
[0095] The model weight parameters of the quantization processing are read from the storage DDR of the PS side through the read-write control module and cached in the corresponding convolution layer weight read FIFO, BN layer weight read FIFO and BN layer bias read FIFO.
[0096] In addition, during the recognition process, the ground-based cloud image recognition model follows the forward reasoning process of the neural network.
[0097] The feature maps of each layer input in the ground-based cloud image recognition model are read from the storage DDR on the PS side in sequence through the read-write control module, and cached in the feature map data read FIFO;
[0098] The feature maps output by each layer in the ground-based cloud map recognition model are cached by writing feature map data into FIFO, and loaded into the storage DDR on the PS side through the read-write control module.
[0099] The read-write control module mainly uses a state machine to read the corresponding data from the corresponding address area in DDR3 and cache it in the corresponding FIFO. At the same time, it writes the data in the feature map data FIFO into the corresponding address area. The state machine is used to complete the arbitration of various functions, avoiding confusion between data and improving the efficiency of data transmission.
[0100] Step S5.2, deploying the ground-based cloud image recognition model through the PL end of the FPGA includes:
[0101] Construct the IP core of the ground-based cloud image recognition model; the IP core includes a sliding window module, a convolution module, a pooling module, a BN layer module, an activation function module, and a fully connected layer module;
[0102] 1) The sliding window module is used to count each pixel of the input feature map by row and column respectively through the row counter and column counter. Whenever the column counter counts to the column size of the feature map, the column counter is reset to zero and the row counter is incremented by one. Based on the counting result, n-1 FIFOs are used to implement the sliding window matrix of size n×n for the input data;
[0103] The sliding window module is an important part of the convolution layer and the pooling layer. The sliding window module is used to form a matrix of a specific size to implement convolution and pooling operations. Take the 7×7 feature map as an example. Figure 7 As shown in the figure, the feature maps are arranged in order from the first data to the 49th data (the last data) in a row-by-row manner. Assuming that the sliding window size (convolution kernel / pooling kernel size) is 3×3, since FIFO has the first-in-first-out feature, two FIFOs can be used to implement data sliding windows, as shown in Figure 8 shown.
[0104] The first row of input data is used as the third row of data of the sliding window, and the first row of input data is cached in FIFO1 in the order of input. When the second row of data of the feature map starts to be input, the data in FIFO1 is output as the second row of data of the sliding window, and the input is used as the third row of data of the sliding window. At the same time, the data output by FIFO1 is cached in FIFO2 in the order of input (the output of FIFO1 is the input of FIFO2). When the third row of data of the feature map starts to be input, the input is the third row of data of the sliding window, the output of FIFO1 is used as the second row of data of the sliding window, and the output of FIFO2 is used as the third row of data of the sliding window. At this time, the first matrix in the feature map is formed, and the sliding matrix can be realized by inputting data in sequence. When sliding windows of different step lengths need to be realized, the output can be enabled. In addition, when a 7×7 sliding window is required, 6 FIFOs can be used to achieve it.
[0105] 2) The convolution module is used to realize the sliding of the matrix of sliding window size m×m through the sliding window module for the convolution layer with a convolution kernel size of m×m, an input channel of a, and an output channel of b; for a feature map, the following is executed in parallel: the pixel points corresponding to the matrix are multiplied by the convolution layer weights through m×m multipliers, and the execution is performed b times in total;
[0106] Convolution operations account for the vast majority of the computational workload of the entire network, including convolutions with kernel sizes of 7×7, 3×3, and 1×1. 7×7 and 3×3 convolution operations are the main acceleration targets of FPGA. The convolution kernel size of the first convolution layer of the network is 7×7, the number of input channels is 3, and the number of output channels is 64. The convolution layer is split into 64 three-dimensional convolution cycles. One three-dimensional convolution cycle operation simultaneously convolves the feature maps of the three input channels with the corresponding convolution kernels. The first convolution layer can reduce the feature map size from 224×224 to 112×112. The convolution operations of the three channels can effectively improve the operation speed of the first convolution layer. Considering the problem of on-chip resource occupation of FPGA, one three-dimensional convolution operation cycle is repeated 64 times to implement all convolution operations of the first convolution layer, which optimizes resource occupation and improves performance.
[0107] A three-dimensional convolution cycle operation is as follows: Fig. 9As shown in the figure, if the 49 pixels of the specified part of the input feature map on the left are serially multiplied with the 49 weight parameters of the corresponding position of the corresponding convolution kernel in the middle part, one convolution operation of a three-dimensional convolution loop operation consumes 49 clock cycles. By utilizing the parallel capability of FPGA, 49 multipliers are deployed to perform multiplication operations simultaneously, that is, the 49 pixels of the specified part of the input feature map are multiplied with the 49 weight parameters of the corresponding position in parallel, and one convolution operation of a three-dimensional convolution loop operation only consumes 1 clock cycle. Compared with one-dimensional serial convolution, a 7×7 convolution operation saves 48 clock cycles, the speed is increased by 49 times, and the speed of a 7×7 convolution layer is increased by 147 times.
[0108] A large number of 3×3 convolutions are used in the residual module of the ground-based cloud image recognition model. The 3×3 convolution modules can be reused to reduce the FPGA resource occupation. The 3×3 convolution uses the same convolution operation method as the 7×7 convolution. Assuming that the number of input channels is N and the number of output channels is M, the convolution layer is split into M×N / 16 16-dimensional convolution cycles. A convolution operation of a 16-dimensional convolution cycle operation consumes only 1 clock cycle. Compared with serial convolution, a 3×3 convolution operation saves 8 clock cycles and increases the speed by 9 times. The speed of a 3×3 convolution layer increases by 144 times. Since the number of input channels of the residual module is 128 and 256, if the same speed optimization method of exchanging space for time as the first convolution layer is used, that is, the feature maps with 128 / 256 channels are convolved at the same time, it is completely inapplicable on the FPGA platform with limited resources. This method is only applicable to the case of a small number of channels. For the case of a large number of channels, the 16-dimensional circular convolution method is more appropriate.
[0109] The convolution layer with a convolution kernel size of 1×1 does not need to use the sliding window module to construct a matrix window. It can be realized by directly multiplying the serial input data with the weight. The number of input channels of the convolution layer of 1×1 convolution is 128 and 256. The method of using all parallel convolutions can greatly improve the speed of convolution, but the FPGA resource occupation is too large. In order to ensure the rational use of resources, the convolution layer of 1×1 convolution adopts the same method as the convolution layer of 3×3 convolution. It only needs to replace the 3×3 convolution with 1×1 convolution.
[0110] 3) The pooling module is used to, for the maximum pooling layer with a pooling kernel size of s×s, realize the sliding of the matrix with a sliding window size of s×s through the sliding window module; find the maximum value of each row of pixel points corresponding to the matrix in parallel through s maximum value finding modules, and find the maximum value of each row maximum value through 1 maximum value finding module; for the average pooling layer with a pooling kernel size of r×r, realize the sliding of the matrix with a sliding window size of r×r through the sliding window module; find the average value of each row of pixel points corresponding to the matrix in parallel through r average value finding modules, and find the average value of the average values of each row through 1 maximum value finding module;
[0111] The pooling layers in ground-based cloud image recognition mainly include maximum pooling with a pooling kernel size of 3×3 and adaptive average pooling. Since the pooling layer has a large number of input channels, if each channel is operated in parallel, it will occupy a large amount of FPGA on-chip resources. The pooling operation is implemented in a serial-parallel multiplexing manner. The pooling operation in a single feature map is operated in parallel, and the operation of all channels is performed in a serial loop, which not only ensures the speed improvement, but also saves FPGA resources to a certain extent.
[0112] The maximum pooling module is mainly composed of four modules for finding the maximum value, such as Fig.10 As shown in the figure, the maximum value module mainly realizes the maximum value of the three input data, which consumes 2 clock cycles. The input of the maximum pooling module is the 3×3 matrix formed by the sliding window module. The first row, the second row and the third row of the 3×3 matrix are simultaneously input to the three parallel maximum value modules to find the maximum value of each row of the matrix, which consumes 2 clock cycles. Then the maximum value of each row is input to the last maximum value module, and the final output is the maximum value of the 3×3 matrix, which consumes 2 clock cycles. The entire 3×3 matrix maximum value consumes 4 clock cycles in total. When the matrix maximum value is obtained in serial mode, 8 clock cycles are required. By using the parallel computing capability of FPGA, the operation of finding the maximum value of the 3×3 matrix can save half the time and increase the speed by 2 times.
[0113] The adaptive average pooling in the network is a 14×14 average pooling. Since the input feature map size is 14×14, the average pooling of the entire feature map only needs to be calculated once using the average pooling module. The average pooling module uses the same idea as the 3×3 average pooling and consists of 15 modules for averaging, such as Fig.12As shown in the figure, the averaging module averages the 14 input data, consuming 13 clock cycles. The average pooling module inputs the pixel data of 14 rows to 14 averaging modules in parallel, calculates the average value of each row, and then inputs the average value of the 14 averaging modules to the next averaging module, and finally outputs the average value of the entire feature map, which is the result of adaptive average pooling. The last averaging consumes 13 clock cycles, and the average pooling module consumes a total of 26 clock cycles. Compared with the serial average pooling, it saves 169 clock cycles and increases the speed by 7.5 times.
[0114] 4) The BN layer module is used to multiply each pixel of the input feature map with the corresponding BN layer weight through a multiplier, and add the multiplication result with the corresponding BN layer bias through an adder;
[0115] The operation principle of the BN layer is similar to the convolution operation with a convolution kernel size of 1×1, and there is no need to use a sliding window module to construct a matrix. Assuming that the feature map size of the BN layer input is 4×4, the BN layer operation diagram is as follows: Fig.12 As shown in the figure, the feature map is input serially, and the input feature map pixels are multiplied by the pre-cached weights and then added to the bias, which consumes 1 clock cycle. In order to reduce the number of times the feature map is written to DDR3, the output of the convolution layer is directly connected to the input of the BN layer, and the serial feature map data output by the convolution layer output channel directly enters the BN layer for calculation.
[0116] 5) The activation function module is used to input each pixel of the feature map through the operation module. If the value of the pixel is greater than or equal to 0, the operation module outputs the value of the pixel; if the value of the pixel is less than 0, the operation module outputs 0; the operation formula is as follows:
[0117]
[0118] Where x is the value of the pixel;
[0119] An example of the ReLu activation function output is as follows Fig.13 As shown in the figure, the ReLu activation function module is directly connected to the output of the BN layer, and the output is cached into DDR3 after the calculation is completed.
[0120] 6) The fully connected layer module is used to calculate the probabilities of p assigned categories through p parallel probability calculation modules, and to obtain the maximum value of the probabilities of the p assigned categories through a maximum calculation module; the input of each probability calculation module is the feature map of each channel, and the input feature map is multiplied by the corresponding fully connected layer weight in turn through a multiplier, and the multiplication results of the feature maps of each channel are added through an adder, and the addition result is added to the corresponding fully connected layer bias through an adder to output the probability value.
[0121] The output of the fully connected layer is the recognition result of the ground cloud map. Using 11 parallel probability modules to calculate the probability of all categories can increase the speed by 11 times. The design block diagram of the fully connected layer module is as follows: Fig.14 As shown in the figure, the output of the adaptive average pooling layer is directly connected to the fully connected layer module, and the output is simultaneously input into 11 probability calculation modules to calculate the recognition probabilities of 11 types of ground-based clouds. Then, the 11 probability values are input into the maximum probability calculation module to calculate the maximum probability value and its corresponding category, and finally the recognition result is output. Fig.15 As shown in the figure, since the input of the probability module is the feature map data arranged according to the number of channels, there is no need to use the sliding window module to construct the matrix window. The input 1×1 feature map is multiplied by the corresponding weights in turn, and then the calculation results of the 256 channel feature maps are added. Finally, the addition result is added to the bias, and the output is the obtained probability value.
[0122] Based on the recognition method of a ground-based cloud image recognition model deployed on FPGA proposed in an embodiment of the present invention, a ground-based cloud image recognition system based on FPGA is constructed. Fig.16 As shown; the PS end mainly uses the FATFS file management system to obtain the weight parameters and cloud map data in the TF card, and loads the read data into DDR3, so that the PL end can obtain data from DDR3 through ARM. After the data is written to DDR3, the enable signal of weight loading completion is transmitted to the PL end through the GPIO port. The PL end mainly implements the writing and packaging of the DDR3 read-write control IP core and the GBcNet IP core. The DDR3 read-write IP core mainly realizes data interaction with the PS end, including the reading of weights and cloud maps and the caching of feature maps. The GBcNet IP core is written in VerilogHDL language and uses FPGA to realize the recognition of ground-based cloud maps. The recognition results are printed through the serial port of the PS end, where the weight parameters only need to be loaded into DDR3 once, and each time, the cloud map data only needs to be reloaded into DDR3. The GbcNet in the figure is the ground-based cloud map recognition model proposed in this embodiment.
[0123] The system makes full use of the processing power of the ZYNQ series FPGA software and hardware collaboration, placing the loading of weight parameters and cloud map data on the PS side, reducing the pressure on FPGA resource usage. In addition, the network model GBcNet implements hardware acceleration on the PL side, giving full play to the parallel computing capabilities of FPGA.
[0124] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0125] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0126] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0128] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A recognition method for a ground-based cloud image recognition model deployed on an FPGA, characterized in that: include: Construct a ground-based cloud image recognition model based on residual network; Training and verifying the ground-based cloud image recognition model to obtain optimal model weight parameters; Quantifying the optimal weight parameters of the model; The ground cloud image to be identified and the model weight parameters of the quantization process are read through the PS end of the FPGA, and loaded into the storage DDR; The ground-based cloud image recognition model is deployed through the PL end of the FPGA, and the PL end reads the ground-based cloud image to be recognized and the model weight parameters of the quantization process from the PS end, generates a recognition result, and returns the recognition result to the PS end; The ground-based cloud image recognition model includes a convolutional layer, a maximum pooling layer, a residual module Res_a, a residual module Res_b, an adaptive average pooling layer, a fully connected layer and a Softmax classifier connected in sequence; The residual module Res_a includes a residual unit Res_a1, a residual unit Res_a2, and a ReLu unit connected to the residual unit Res_a1 and the residual unit Res_a2; the residual module Res_b includes a residual unit Res_b1, a residual unit Res_b2, a residual unit Res_b3, and a ReLu unit connected to the residual unit Res_b1, the residual unit Res_b2, and the residual unit Res_b3; The residual unit Res_a1 and the residual unit Res_b1 have the same structure, both including a first branch and a second branch, the first branch is provided with a Conv-BN unit, the second branch is provided with a Conv-BN-ReLu unit and a Conv-BN unit in sequence, and the outputs of the first branch and the second branch are added; the residual unit Res_a2, the residual unit Res_b2 and the residual unit Res_b3 have the same structure, both including a third branch and a fourth branch, the fourth branch is provided with a Conv-BN-ReLu unit and a Conv-BN unit in sequence, and the outputs of the third branch and the fourth branch are added; The method of reading the ground-based cloud image to be identified and the model weight parameters of the quantization process through the PS end of the FPGA and loading them into the storage DDR includes: The storage DDR is pre-divided into a cloud map data area, a convolutional layer weight area, a BN layer weight area, a fully connected layer weight area, and a feature map data area; Loading the ground-based cloud image to be identified directly into the cloud image data area; Extracting the convolution layer weights corresponding to each convolution layer from the quantized model weight parameters, storing and sorting the convolution layer weights of the convolution kernels of each input channel corresponding to each output channel in order according to the convolution layer weights contained in the convolution kernel, and loading them into the convolution layer weight area after storage and sorting; Extract the BN layer weights and BN layer biases corresponding to each BN layer from the quantized model weight parameters, and store and sort the BN layer weights of the convolution kernels of each input channel corresponding to each output channel in sequence according to the BN layer weights contained in the convolution kernels; store and sort the BN layer biases of the convolution kernels of each input channel corresponding to each output channel in sequence according to the BN layer biases contained in the convolution kernels; place the storage and sorting of the BN layer biases after the storage and sorting of the BN layer weights, and load them into the BN layer weight area; Extract the fully connected layer weights and fully connected layer biases corresponding to each fully connected layer from the quantized model weight parameters, and store and sort the fully connected layer weights of the convolution kernels of each input channel corresponding to each output channel in sequence according to the fully connected layer weights contained in the convolution kernels; store and sort the fully connected layer biases of the convolution kernels of each input channel corresponding to each output channel in sequence according to the fully connected layer biases contained in the convolution kernels; place the storage and sorting of the fully connected layer biases after the storage and sorting of the fully connected layer weights, and load them into the fully connected layer weight area; The PL end reads the ground-based cloud image to be identified from the PS end and the model weight parameters for quantization processing, including: A read-write control IP core is constructed based on the AXI4 protocol, wherein the read-write control IP core includes a read-write control module and a feature map data write FIFO, a feature map data read FIFO, a convolutional layer weight read FIFO, a BN layer weight read FIFO, and a BN layer bias read FIFO connected thereto; The model weight parameters of the quantization processing are read from the storage DDR of the PS end by the read-write control module, and cached in the corresponding convolution layer weight read FIFO, the BN layer weight read FIFO and the BN layer bias read FIFO; Wherein, the deploying of the ground-based cloud image recognition model through the PL end of the FPGA includes: Constructing the IP core of the ground-based cloud image recognition model; the IP core includes a sliding window module, a convolution module, a pooling module, a BN layer module, an activation function module and a fully connected layer module; The sliding window module is used to count each pixel of the input feature map according to rows and columns respectively through a row counter and a column counter. Whenever the column counter counts to the column size of the feature map, the column counter is reset to zero and the row counter is incremented by one; based on the counting result, n-1 FIFOs are used to implement matrix sliding of a sliding window size n×n for the input data; The convolution module is used to realize the sliding of the matrix of the sliding window size m×m through the sliding window module for the convolution layer with a convolution kernel size of m×m, an input channel of a, and an output channel of b; and execute in parallel for a feature maps: multiply the pixel points corresponding to the matrix by the convolution layer weights through m×m multipliers, and execute b times in total; The pooling module is used to, for a maximum pooling layer with a pooling kernel size of s×s, implement a matrix sliding of a sliding window size of s×s through the sliding window module; obtain the maximum value of each row of pixel points corresponding to the matrix in parallel through s maximum value modules, and obtain the maximum value of each row maximum value through 1 maximum value module; for an average pooling layer with a pooling kernel size of r×r, implement a matrix sliding of a sliding window size of r×r through the sliding window module; obtain the average value of each row of pixel points corresponding to the matrix through r average value modules in parallel, and obtain the average value of the average values of each row through 1 maximum value module; The BN layer module is used to multiply each pixel of the input feature map by the corresponding BN layer weight through a multiplier, and add the multiplication result to the corresponding BN layer bias through an adder; The activation function module is used to input each pixel of the feature map through the operation module. If the value of the pixel is greater than or equal to 0, the operation module outputs the value of the pixel; if the value of the pixel is less than 0, the operation module outputs 0; The fully connected layer module is used to calculate the probabilities of p allocation categories through p parallel probability calculation modules, and to obtain the maximum value of the probabilities of the p allocation categories through a maximum value calculation module; the input of each probability calculation module is the feature map of each channel, and the input feature map is multiplied by the corresponding fully connected layer weight in sequence through a multiplier, and the multiplication results of the feature maps of each channel are added through an adder, and the addition result is added to the corresponding fully connected layer bias through an adder, and the probability value is output; The output of the fully connected layer module is the recognition result of the ground-based cloud map.
2. The recognition method of the ground-based cloud image recognition model deployed on FPGA according to claim 1 is characterized in that: The training and verification of the ground-based cloud image recognition model to obtain the optimal model weight parameters includes: Get a dataset of labeled ground-based cloud images; Expanding the data set using at least one of an image blurring method, a contrast enhancement method, a mirror flipping method, and a Gaussian noise adding method; Dividing the expanded data set into a training set and a test set according to a preset ratio; The ground-based cloud image recognition model is trained and verified according to the training set and the test set.
3. The recognition method of the ground-based cloud image recognition model deployed on FPGA according to claim 1 is characterized in that: During the recognition process, the ground-based cloud image recognition model sequentially caches the feature images inputted from each layer in the ground-based cloud image recognition model to the feature image data area according to the forward reasoning process of the neural network.
4. The recognition method of the ground-based cloud image recognition model deployed on FPGA according to claim 1 is characterized in that: In the recognition process, the ground-based cloud image recognition model follows the forward reasoning process of the neural network. The feature graphs inputted from each layer in the ground-based cloud image recognition model are sequentially read from the storage DDR of the PS end through the read-write control module, and are cached in the feature graph data read FIFO; The feature map data is written into FIFO to cache the feature maps output by each layer in the ground-based cloud map recognition model, and is loaded into the storage DDR of the PS end through the read-write control module.
Citation Information
Patent Citations
Foundation cloud atlas cloud identification method based on convolutional neural network
CN110929602A
Lightweight foundation cloud segmentation method and system based on multi-scale feature fusion and alignment
CN117197462A