A lightweight rice disease prediction method based on SCL-YOLOv8n
By introducing the SCL-YOLOv8n model in rice disease detection, the network structure and module are improved, and the problems of low detection accuracy and huge model in the existing technology are solved, and efficient and real-time rice disease prediction are achieved.
Patent Information
- Application Number
- CN202510156073.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The prior art has problems such as complex features, low detection accuracy and huge models in rice disease detection, which is difficult to meet the real-time and accurate disease identification needs.
A lightweight rice disease prediction method based on SCL-YOLOv8n is proposed. By improving the GSConv module in the Slim-Neck network, building the RFCA-CSP module and using the LSCSBD detection head, the feature expression ability is enhanced, the calculation complexity is reduced, and the detection efficiency is improved.
It realizes that while maintaining high detection accuracy, it significantly reduces the number of model parameters and calculation complexity, and improves the real-time and efficiency of rice disease detection.
Smart Images

Figure CN119625543B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of rice disease prediction, and particularly relates to a lightweight rice disease prediction method based on SCL-YOLOv8n. Background Art
[0002] During the growth process of rice, diseases occur frequently. If these diseases are not identified in time and effectively controlled, they may directly affect the yield of rice and cause serious economic losses. Therefore, how to quickly and accurately identify the type of disease, determine the location and severity of the disease during the growth process of rice, and take targeted control measures has become the key to improving rice production efficiency, reducing losses, and ensuring food security.
[0003] Traditional disease detection and prediction methods mainly rely on manual visual observation and empirical judgment, which are time-consuming, laborious, and highly subjective, and are difficult to meet the actual needs of pest and disease control. Compared with traditional computer vision technologies, deep learning technologies have shown strong generalization abilities in multiple image analysis fields due to their superior feature learning capabilities and are now widely used in the identification of agricultural crop diseases. The object detection methods in deep learning can be divided into two categories: single-stage and two-stage. Single-stage methods include the YOLO series, SSD, and EfficientDet, while two-stage methods are such as Mask R-CNN and Faster R-CNN. Compared with single-stage methods, two-stage detection methods usually need to generate candidate boxes first and then perform convolutional operations, resulting in slower recognition speed and poorer real-time performance.
[0004] In contrast, the YOLO series, as single-stage methods, combine the generation of target boxes and feature extraction into one step, with faster recognition speed and stronger real-time performance. In the prior art, the YOLOv8 Rice model proposed based on YOLOv8n introduces deformable convolutions to optimize the C2f module, adopts a weighted bidirectional feature pyramid network (Bi-FPN) to improve performance and reduce computational complexity, and at the same time uses the Wise IOU loss function to improve the evaluation accuracy; the C2f-MSEC module is introduced into the backbone network to replace C2f, and a weighted bidirectional feature pyramid network is used to enhance the feature fusion ability for false smut of different sizes, and a group normalization shared convolutional lightweight detection head is designed to improve the lightweight performance; partial convolution (PConv) is introduced to replace the bottleneck layer of YOLOv8, reducing the number of parameters and improving the detection speed, reconstructing the feature pyramid BFP, and embedding Gaussian non-local attention (EGNA) to effectively reduce the aliasing effect caused by multi-layer fusion; a lightweight YOLO v8-Rice algorithm is proposed based on YOLOv8, using the ContextGuided Block to replace the C2f module, depthwise separable convolutions to replace standard convolutions, and reconstructing the lightweight shared convolutional detection head, significantly reducing the number of parameters and computational amount. Summary of the Invention
[0005] The object of the present invention is to solve the problems of complex rice disease characteristics, low detection accuracy, and large model size, and proposes a lightweight rice disease prediction method based on SCL-YOLOv8n.
[0006] To solve the above technical problems, the specific steps are as follows:
[0007] Step 1: Construct a rice disease dataset;
[0008] Step 2: Perform enhancement, annotation, and partitioning on the rice disease dataset;
[0009] Step 3: In the GSConv module of the Slim-Neck network structure, lightweight convolution is introduced in parallel with the original path before the Shuffle operation to form the GSLConv module; based on the CSP-RFA structure, after the data partitioning operation, the data is split into two branches to form the RFCA-CSP module; a lightweight shared enhanced detection head LSCSBD is adopted; the SCL-YOLOv8n model is constructed;
[0010] Step 4: Use the rice disease dataset processed in Step 2 as the input of the SCL-YOLOv8n model, and use the training set to train the SCL-YOLOv8n model;
[0011] Step 5: Use the test set to conduct rice disease category inspection and location positioning;
[0012] Step 6: Use the trained SCL-YOLOv8n model to predict rice diseases.
[0013] Further, in Step 1, the construction of the rice disease dataset is to take pictures of the target from different distances at multiple angles; the captured images cover six diseases, namely rice blast, brown spot, false smut, sheath blight, and two other diseases (the names seem incorrect in the original text, but translated as is), and pictures of normal rice.
[0014] Further, in Step 2, dataset enhancement adopts methods such as rotation, adding salt-and-pepper noise, and Gaussian noise; the LabelImg tool is used to annotate the rice disease dataset, the positions of rice disease spots are annotated with rectangular boxes, and the dataset is partitioned. The annotated dataset is randomly divided into a training set, a test set, and a validation set in a ratio of 7:2:1.
[0015] Further, in Step 3, the specific construction steps of the GSLConv module are as follows:
[0016] (1) The input feature map compresses the number of channels from C1 to C2 / 2 through a 1×1 convolution; a 5×5 depthwise separable convolution DWConv is used to extract spatial features; the output of DWConv is concatenated with the result of the 1×1 convolution in the channel dimension to obtain the feature map X2, combining local spatial features and global channel information;
[0017] (2) The concatenated feature map is processed through two paths: one path performs channel shuffling Shuffle, mixing features by adjusting the channel order and breaking channel locality; the other path first processes the feature map through the LightConv layer and then performs the Shuffle operation, with the output dimension being H×W×C_out;
[0018] (3) The outputs of the two paths are fused by addition.
[0019] Further, in step 3, the RFCA-CSP module expands the number of channels of the input feature map from C_in to twice the original through a 1x1 convolution, with the output having the number of channels C_out, and then divides it into two branches: a convolutional branch and a Transformer branch.
[0020] Further, the convolutional branch extracts local features through the Bottleneck-RFCAConv module and maintains the feature map of the input spatial dimension;
[0021] (1) During the processing, the input features are transformed through N×N grouped convolution, followed by batch normalization to stabilize the training process, and the ReLU activation function is applied to introduce non-linearity; the formula is as follows:
[0022] (1)
[0023] In the formula, g i×i represents grouped convolution of size i×i, N represents the convolution kernel size, Norm represents normalization, ReLU represents using the ReLU activation function, x represents the input features, and F represents the output features;
[0024] (2) Bottleneck-RFCAConv reduces the number of channels in the intermediate layer by introducing an expansion coefficient, thereby reducing the computational complexity, and adopts a two-layer convolutional module: the first layer reduces the number of channels through standard convolution operations, and the second layer uses the RFCAConv module to enhance the feature representation ability;
[0025] (3) After the receptive field spatial features extracted by the RFCAConv module are processed, the CA attention mechanism is used to optimize the input features; the CA attention mechanism decomposes the feature information in the height and width directions respectively, so as to obtain the feature representations along these two directions, further improving the model's ability to capture spatial information;
[0026] (4) By globally splicing the receptive fields of the feature map in the height and width directions, and applying convolution, batch normalization, and non-linear activation operations, the features are transformed;
[0027] (5) Split the intermediate features into two independent features, and perform feature transformation through convolution operations and the Sigmoid activation function, so that the output dimension is consistent with the input vector; calculate the attention weights of the input feature map in the height and width dimensions, and combine the calculated attention weights with the extracted feature map to generate the final output.
[0028] Further, the Transformer branch uses the MHSA-CGLU module to establish global dependencies through the multi-head self-attention MHSA and transform features in combination with the convolutional gating mechanism CGLU.
[0029] Further, the multi-head self-attention mechanism MHSA maps the input sequence to multiple subspaces, where each subspace corresponds to an independent attention head; its specific expression is:
[0030] (2)
[0031] In the formula, Q, K, and V represent the query, key, and value vectors respectively, d k is the dimension of the key vector, which is used to scale the dot product to prevent the dot product result from being too large, thereby avoiding the vanishing gradient. T represents the transpose matrix, and softmax represents the normalization operation;
[0032] (1) Introduce the multi-head self-attention mechanism, and multiple attention heads calculate in parallel. Each head learns different representations in its own subspace; the specific process is expressed by formula (3) and formula (4):
[0033]
[0034]
[0035] In the formula, , , , are the parameter matrices learned by the model. Concat represents the operation of concatenation, head represents the attention head, h represents the number of attention heads, and i represents the i-th;
[0036] (2) CGLU adds a 3x3 depth convolution operation before the gating linear unit GLU activation function to enhance the local modeling ability and the robustness of the model;
[0037] (3) The outputs of the two branches are concatenated along the channel dimension; and are mapped to the final output channel number C_out through a 1x1 convolution to complete feature fusion;
[0038] Further, in step 3, the operation steps of the lightweight shared enhanced detection head LSCSBD are as follows:
[0039] (1) Input feature maps of different scales into the network, and map the channel number of each layer of feature maps from C_i to the hidden channel number hidc through a 1x1 convolutional layer while keeping the spatial size unchanged;
[0040] (2) At each detection layer, apply a shared 3×3 convolution to extract spatial features; after the convolution operation, perform batch normalization BN on the output features and use the SiLU activation function for non-linear transformation; the features after SiLU processing are further subjected to a second 3×3 shared convolution operation as well as BN normalization and SiLU activation to enhance the feature expression ability;
[0041] (3) The output is a tensor containing bounding box regression and class scores, and the entire process scales the output of each detection layer through the Scale operation to ensure consistency and accuracy in multi-scale detection tasks.
[0042] Beneficial effects, the present invention has the following effects through 3 innovative points:
[0043] (1) Improve the GSConv module in the Slim-Neck network to enhance the feature expression ability, reduce the computational overhead, and at the same time improve the fusion effect of multi-scale features.
[0044] (2) Construct the RFCA-CSP module, by combining the advantages of convolution and Transformer, not only retain the high efficiency of CNN in local feature extraction, but also introduce the global dependency modeling ability of Transformer. Compared with the traditional Transformer structure, this module effectively reduces the computational complexity and achieves a good balance between capturing low-level local features and high-level global features.
[0045] (3) Adopt the LSCSBD detection head, by introducing shared convolution, separated batch normalization and deep feature learning technologies, make the model more efficient and flexible while maintaining high performance. Description of the Drawings
[0046] Figure 1It is the structural diagram of SCL-YOLOv8n;
[0047] Figure 2 It is the structural diagram of the GSLConv model;
[0048] Figure 3 It is the structural diagram of LVoVGSCSP;
[0049] Figure 4 It is the schematic diagram of the RFCA-CSP module: (a) is the RFCA-CSP structure; (b) is the original Bottleneck structure; (c) is the Bottleneck-RFCAConv structure; (d) is the MHSA-CGLU structural diagram;
[0050] Figure 5 It is the structural diagram of the LSCSBD model;
[0051] Figure 6 It is the comparison chart of different weights of the PCB dataset;
[0052] Figure 7 It is the heat map of different weights;
[0053] Figure 8 It is the overall structural diagram of the present invention. Detailed implementation manners
[0054] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0055] The following is the description of the English abbreviations appearing in the present invention: Slim-Neck: thin neck structure; GSConv module: shuffle group convolution; Shuffle operation: shuffle the data; RFCAConv: receptive field attention convolution; VoVGSCSP module: hybrid convolution and cross-stage partial network module; LightConv: lightweight convolution; CA attention: channel attention; HSA-CGLU module: multi-head self-attention-convolution gated attention; Scale operation: scaling of the bounding box; SiLU activation function: Sigmoid gated linear unit; Bottleneck-RFCAConv: bottleneck receptive field attention convolution; mAP50: mean average precision with an intersection over union of 0.5; GFLOPs: giga floating point operations per second; Params: number of parameters; Grad-CAM: gradient-weighted class activation mapping.
[0056] As Figure 8 shown, the present invention provides a rice disease detection and prediction method based on improved Y0L0v8n in a complex environment, and the process is as follows:
[0057] 1. Obtain and process the rice disease dataset:
[0058] (1) Construct a dataset. When collecting images, the shooting environment under different time periods and natural lighting conditions was fully considered, and the target was photographed from multiple angles and different distances. The photographed images cover six diseases, namely rice blast, brown spot, false smut, sheath blight, bakanae disease, and leaf scald, as well as normal rice pictures.
[0059] (2) Perform data augmentation on the collected images, using data augmentation methods such as rotation, adding salt-and-pepper noise, and Gaussian noise. Use the LabelImg tool to annotate the rice disease dataset, use rectangular boxes to annotate the positions of rice disease spots, and divide the dataset. Randomly divide the annotated dataset into a training set, a test set, and a validation set according to a ratio of 7:2:1. Finally, there are 3248 pictures in the training set, 928 pictures in the test set, and 464 pictures in the validation set.
[0060] 2. As Figure 1 shown, construct the SCL-YOLOv8n model based on YOLOv8n:
[0061] (1) In the GSConv module of the Slim-Neck network structure, lightweight convolution is introduced before the Shuffle operation and is connected in parallel with the original path to form the GSLConv module.
[0062] (2) Based on the CSP-RFA structure, after the data partitioning operation, the data is split into two branches to form the RFCA-CSP module.
[0063] (3) Adopt the LSCSBD detection head module. By introducing optimization means such as lightweight shared convolution, batch normalization BN, and dynamic anchor box calculation, improve the efficiency and stability of YOLOv8n in multi-scale object detection.
[0064] As Figure 2 - Figure 3 shown, the present invention improves the GSConv module on the basis of the Slim-Neck network.
[0065] (1) Introduce the Slim-Neck framework, which consists of two main modules: the GSConv module and the VoVGSCSP module, and the VoVGSCSP module is mainly composed of GSConv and Conv convolutions.
[0066] (2) To make up for the possible loss of key information caused by GSConv in the feature extraction process and reduce the computational amount and number of parameters of convolution in the convolution kernel operation, the present invention proposes an improved convolution module GSLConv. Combine the strategies of parallel paths and LightConv layers to optimize the traditional GSConv structure and enhance the feature extraction ability of the model. The specific process is as follows:
[0067] First, the input feature map is compressed from C1 to C2 / 2 in terms of the number of channels through a 1×1 convolution. Second, a 5×5 depthwise separable convolution (DWConv) is used to extract spatial features. The output of the DWConv is concatenated with the result of the 1×1 convolution in the channel dimension to obtain the feature map X2, combining local spatial features with global channel information. Then, the concatenated feature map is processed through two paths: one path performs channel shuffling (Shuffle) to mix features by adjusting the channel order and break channel locality; the other path first processes the feature map through a LightConv layer with an output dimension of H×W×C_out. Finally, the outputs of the two paths are fused by addition.
[0068] As Figure 4 (a) shows, the present invention proposes an RFCA-CSP module that expands the number of channels of the input feature map from C_in to twice the original through a 1x1 convolution, with an output of the number of channels C_out, and then divides it into two branches: a convolutional branch and a Transformer branch.
[0069] (1) As Figure 4 (b)-4(c) shows, the convolutional branch extracts local features through a Bottleneck-RFCAConv module and maintains the input spatial size in the feature map, specifically:
[0070] Step a: During the processing, the input features are transformed through grouped convolution, followed by batch normalization to stabilize the training process, and the ReLU activation function is applied to introduce non-linearity. The formula is as follows:
[0071] (1)
[0072] In the formula, g i×i represents grouped convolution of size i×i, N represents the convolution kernel size, Norm represents normalization, ReLU represents using the ReLU activation function, x represents the input features, and F represents the output features.
[0073] Step b: Bottleneck-RFCAConv reduces the number of channels in the intermediate layer by introducing an expansion coefficient, thereby effectively reducing the computational complexity. A two-layer convolutional module is adopted: the first layer reduces the number of channels through standard convolution operations, and the second layer uses the RFCAConv module to enhance the feature representation ability.
[0074] Step c: After the receptive field spatial features extracted by the RFCAConv module are processed, the CA attention mechanism is used to optimize the input features. The CA attention mechanism decomposes the feature information in both the height and width directions respectively, so as to obtain the feature representations along these two directions, and further improves the model's ability to capture spatial information.
[0075] Step d: By globally splicing the receptive fields of the feature map in the height and width directions, and applying operations such as convolution, batch normalization, and non-linear activation, the features are further transformed.
[0076] Step e: The intermediate features are split into two independent features, and feature transformation is performed through convolution operations and the Sigmoid activation function, so that the output dimension is consistent with the input vector. Calculate the attention weights of the input feature map in the height and width dimensions, and combine the calculated attention weights with the extracted feature map to generate the final output.
[0077] (2) As Figure 4 (d) shows, the Transformer branch uses the MHSA-CGLU module to establish global dependencies through multi-head self-attention (MHSA) and further transform the features by combining the convolutional gating mechanism (CGLU).
[0078] Step a: The multi-head self-attention mechanism (MHSA) can effectively capture complex dependencies within the sequence by mapping the input sequence to multiple subspaces, where each subspace corresponds to an independent attention head. Its specific expression is:
[0079] (2)
[0080] In the formula, Q, K, and V represent the query, key, and value vectors respectively, d k is the dimension of the key vector, which is used to scale the dot product to prevent the dot product result from being too large, thereby avoiding gradient disappearance. T represents the transpose matrix, and softmax represents the normalization operation.
[0081]
[0082]
[0083] In the formula, , , , are the parameter matrices learned by the model. Concat represents the operation of concatenation, head represents the attention head, h represents the number of attention heads, and i represents the i-th one.
[0084] Step b: CGLU adds a 3x3 depth convolution operation before the GLU (Gated Linear Unit) activation function to enhance the local modeling ability and the robustness of the model.
[0085] Step c: The outputs of the two branches are concatenated along the channel dimension and mapped to the final output channel number C_out through a 1x1 convolution to complete feature fusion.
[0086] As Figure 5 shown, the present invention proposes a lightweight shared enhanced detection head LSCSBD.
[0087] First, feature maps of different scales are input into the network. Through multiple 1x1 convolution layers, the channel number of each layer of feature map is mapped from C_i to the hidden channel number hidc while keeping the spatial size unchanged. Then, for each detection layer, a shared 3×3 convolution module is applied to extract spatial features. After the convolution operation, batch normalization (BN) is performed on the output features, and the SiLU activation function is used for non-linear transformation. The features after SiLU processing are further subjected to a second 3×3 shared convolution operation, BN normalization, and SiLU activation to enhance the feature expression ability. The final output is a tensor containing bounding box regression and class scores. The entire process adjusts the output of each detection layer through the Scale operation to ensure consistency and accuracy in multi-scale detection tasks.
[0088] 3. The improved SCL-YOLOv8n model is trained using the processed rice disease dataset. The training coefficient is set to 500 rounds, the batch size is set to 16, and 203 pictures are input for each training. After training, the trained weights are saved.
[0089] 4. The category inspection and location positioning of rice diseases are carried out using the test set. Using the optimal weight file obtained in step 3, the pictures of the test set divided in step 1 are inferred and verified.
[0090] 5. Using the optimal weight file obtained in step 3, rice diseases are predicted. The SCL-YOLOv8n model is applied to the rice disease detection system, and real-time prediction and recognition of diseases are achieved through the validation set or taking pictures.
[0091] The experimental data of this embodiment shows that:
[0092] (1) Model generalization experiment, as Figure 6 shown.
[0093] To verify the generalization ability of the improved model SCL-YOLOv8n, images of 6 different defect categories were selected from the PCB defect detection dataset open-sourced by Peking University for detection. Using the optimal weight files of the original model YOLOv8n and the improved model SCL-YOLOv8n, the data in the validation set was inferred and verified. The specific results are shown in Table 1.
[0094] Table 1 Generalization experiment
[0095] Algorithm Precision P(%) Recall R(%) mAP50(%) GFLOPs(G) Params(MB) YOLOv8n 94.3 90.7 92.0 8.2 3.01 SCL - YOLOv8n 92.1 89.1 92.4 5.2 1.93
[0096] From the experimental results in Table 1, it can be seen that compared with YOLOV8n, the improved SCL-YOLOv8n has a slight decrease in P and R on the PCB dataset. On the basis that the average precision mAP50 remains basically unchanged, the number of model parameters and the computational complexity decrease by 35.9 and 36.6 percentage points respectively. The experimental results show that the performance of this model has been significantly improved, and it has good generalization ability and robustness.
[0097] (2) Model comparison experiment, as Figure 7 shown.
[0098] To comprehensively evaluate the improvement effect of SCL-YOLOv8n, this paper compares it with other classic object detection models, covering models such as YOLOv5n, YOLOv10n, YOLOv11n, and SSD, as well as other improved algorithms in the same field. All models were experimented on the same dataset and in the same experimental environment. The specific results are shown in Table 2.
[0099] Table 2 Comparative experiments of different algorithms
[0100] Algorithm Precision P(%) Recall R(%) mAP50(%) GFLOPs(G) Params(MB) SSD 83.1 68.29 72.1 62.8 26.3 Faster - RCNN 60.9 73.3 70.8 363.4 136.7 YOLOv5n 84.9 71.4 81.6 7.2 2.51 YOLOv6 82.7 70.4 77.3 11.9 4.20 YOLOv7 - tiny 81.5 72.0 80.3 13.2 6.03 Yolov9t 79.2 75.0 81.2 7.9 2.01 YOLOv10n 83.9 76.4 82.4 8.4 2.71 YOLOv11n 84.0 77.6 83.3 6.4 2.59 RDN - YOLO 84.1 71.8 79.4 11.5 4.52 SCL - YOLOv8n 86.6 81.4 86.0 5.5 1.93
[0101] According to the experimental data in Table 2, the improved model proposed in this paper has better performance in terms of the number of parameters, computational complexity, and detection accuracy compared with the classic SSD and Faster R-CNN detection algorithms.
[0102] Compared with the YOLOv5n model, the improved model reduces the number of parameters and computational complexity by 23.1% and 23.6% respectively, while the mAP50 increases by 4.4 percentage points. Compared with YOLOv6, the SCL-YOLOv8n model has fewer parameters and computational complexity, and at the same time, the mAP50 increases significantly. For the lightweight model YOLOv7-tiny, the improved SCL-YOLOv8n reduces the number of parameters and computational complexity by 4.1MB and 7.7G respectively, and the mAP50 increases by 5.7 percentage points. Compared with YOLOv9t, although the difference in the number of parameters between the two is not significant, the improved model improves P, R, and mAP50 by 7.4, 6.4, and 4.8 percentage points respectively, while the computational complexity decreases by 2.4G. When compared with the YOLOv10n model, the improved model increases the mAP50 by 3.6 percentage points, while the computational complexity and the number of parameters are reduced by 34.5% and 28.8% respectively. Compared with the YOLOv11n model, the improved algorithm improves P, R, and mAP50 by 2.6, 3.8, and 2.7 percentage points respectively, and both the computational complexity and the number of parameters are slightly reduced.
[0103] When compared with the improved algorithm RDN-YOLO, the improved model in this paper improves P, R, and mAP50 by 2.5, 9.6, and 6.6 percentage points respectively, and the number of model parameters and computational complexity are reduced by 2.59MB and 6G respectively.
[0104] In summary, SCL-YOLOv8n performs well in terms of the number of parameters, computational complexity, and accuracy, and can be used for the detection of rice diseases.
[0105] The heatmap is used to visually show the degree of attention of the model to key features in the input image. The present invention uses Grad-CAM to generate the object detection heatmap. Figure 7 The red area in [the heatmap] indicates that the model highly focuses on a specific area, while the blue area indicates a lower degree of attention.
[0106] In the text, models with higher accuracy in the comparative experiments and the original YOLOv8n model are selected for comparison. Among them, the yellow circles represent missed detections. There are 15 missed detections in the original YOLOv8n model, and 4 missed detections in the improved SCL-YOLOv8n. The red triangles represent areas with incorrect attention. There are 2 areas with incorrect attention in the YOLOv11n model, and there is 1 area with incorrect attention in the original YOLOv8n model, the YOLOv10n model, and the improved SCL-YOLOv8n. Through the heatmap analysis, it can be proved that the improved SCL-YOLOv8n has better performance.
[0107] (3)Model ablation experiment.
[0108] To evaluate the impact of each improved module proposed by the improved algorithm in this paper on the performance of rice disease detection, corresponding ablation experiments were conducted. The experimental results are shown in Table 3.
[0109] Table 3 Ablation Experiments
[0110] Experiment Slim - Neck with GSLConv Added RFCA - CSP LSCSBD Precision P(%) Recall R(%) mAP50(%) GFLOPs(G) Params(MB) Ablation Experiment 1 × × × 81.3 75.2 81.0 8.0 3.00 Ablation Experiment 2 √ × × 84.9 77.8 83.7 6.9 2.69 Ablation Experiment 3 × √ × 81.4 76.0 82.3 7.7 2.71 Ablation Experiment 4 × × √ 82.8 76.9 80.3 6.6 2.37 Ablation Experiment 5 √ √ × 85.1 77.8 85.4 7.1 2.58 Ablation Experiment 6 √ × √ 82.7 78.9 82.7 6.0 2.24 Ablation Experiment 7 × √ √ 82.8 77.2 82.3 6.2 2.06 Ablation Experiment 8 √ √ √ 86.6 81.4 86.0 5.5 1.93
[0111] From the ablation experiment data in Table 3, it can be seen that ablation experiment 1 is the original experimental result of the unimproved YOLOv8n. Ablation experiment 2 adds the Slim-Neck structure with the improved GSLConv module on the basis of ablation experiment 1. By fusing the features of the shallow layer and the deep layer, it can capture the diverse manifestations of rice diseases more comprehensively. Compared with ablation experiment 1, on the basis of an increase of 2.7 percentage points in mAP50, the number of parameters and the computational complexity are reduced by 10.3 and 13.8 percentage points respectively. Ablation experiment 3 adds the RFCA-CSP module to YOLOv8n. Compared with ablation experiment 1, when the number of parameters and the computational complexity are both reduced, mAP50 increases by 1.3 percentage points, verifying that combining the convolutional and Transformer structures helps to combine local features with global features and extract richer disease information. Ablation experiment 4 introduces the LSCSBD detection head on the basis of the original model, and reduces the number of model parameters by applying shared convolutions. Although mAP50 drops, it only drops by 0.7 percentage points, and the number of parameters and the computational complexity drop by 21 and 17.5 percentage points respectively, verifying the effectiveness of the improvement measures.
[0112] Ablation experiments 5, 6, and 7 verify that the pairwise combination of the three modules has different improvement effects on the model performance.
[0113] Combining the above improvements on the basis of ablation experiment 1, the number of parameters of the improved SCL-YOLO v8n is only 1.93MB, and the computational complexity is 5.5G, which are reduced by 35.7 and 31.3 percentage points respectively compared with the original model, and mAP50 increases by 5.0 percentage points.
[0114] In summary, it can be proved that the improved model SCL-YOLOv8n of the present invention realizes lightweight while improving the accuracy, has better model performance and a more lightweight model framework.
[0115] The above embodiments are only to illustrate the technical concept and characteristics of the present invention, and the purpose is to enable those skilled in the relevant technical fields to understand the content of the present invention and implement it, and are not used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the present invention should be included in the protection scope of the present invention.
Claims
1. A lightweight rice disease prediction method based on SCL-YOLOv8n, characterized in that: The following steps are involved: Step 1: Construct a rice disease dataset; Step 2: Enhance, label, and divide the rice disease dataset; Step 3. Build the SCL-YOLOv8n model based on YOLOv8n: In the GSConv module of the Slim-Neck network structure, reference the lightweight convolution before the Shuffle operation and connect it in parallel with the original path to form a GSLConv module; based on the CSP-RFA structure, after the data is divided, the data is diverted to two branches to form an RFCA-CSP module; the RFCA-CSP module expands the number of channels of the input feature map from C_in to twice the original through a 1x1 convolution, and outputs the number of channels C_out, which is then divided into two branches: a convolution branch and a Transformer branch; Adopt LSCSBD detection head module, introduce lightweight shared convolution, batch normalization BN and dynamic anchor box calculation method; Step 4: Use the rice disease dataset processed in step 2 as the input of the SCL-YOLOv8n model, and use the training set to train the SCL-YOLOv8n model; Step 5: Use the test set to test the rice disease category and locate the location; Step 6: Use the trained SCL-YOLOv8n model to predict rice diseases.
2. A lightweight rice disease prediction method based on SCL-YOLOv8n according to claim 1, characterized in that: In step 1, the rice disease dataset is constructed by photographing the target from multiple angles and different distances; the photographed images cover six diseases, namely, rice blast, brown spot, heart blight, dew drop, false smoke, and sheath disease, as well as pictures of normal rice.
3. The lightweight rice disease prediction method based on SCL-YOLOv8n according to claim 1, characterized in that: In step 2, the dataset is enhanced by rotation, adding salt and pepper noise, and Gaussian noise. The rice disease dataset is annotated using the LabelImg tool, the location of rice lesions is annotated using rectangular boxes, and the dataset is divided into training, test, and validation sets at a ratio of 7:2:
1.
4. The lightweight rice disease prediction method based on SCL-YOLOv8n according to claim 1, characterized in that: In step 3, the specific construction steps of the GSLConv module are: (1) The input feature map is compressed from C1 to C2 / 2 through a 1×1 convolution; a 5×5 depthwise separable convolution DWConv is used to extract spatial features; The output of DWConv is concatenated with the 1×1 convolution result in the channel dimension to obtain the feature map X2, which combines the local spatial features with the global channel information; (2) The concatenated feature map is processed through two paths: one path performs channel shuffle, which mixes features by adjusting the channel order and breaking channel locality; the other path first processes the feature map through the LightConv layer, and then performs the shuffle operation, with an output dimension of H×W×C_out; (3) The outputs of the two paths are fused by addition.
5. The lightweight rice disease prediction method based on SCL-YOLOv8n according to claim 1, characterized in that: The convolution branch extracts local features through the Bottleneck-RFCAConv module and maintains the input spatial size feature map; (1) During the processing, N×N The input features are transformed by a group convolution, followed by batch normalization to stabilize the training process, and the ReLU activation function is applied to introduce nonlinearity; the formula is as follows: (1), where g i×i represents a grouped convolution of size i×i, N represents the convolution kernel size, Norm represents normalization, ReLU represents the use of ReLU activation function, x represents input features, and F represents output features; (2) The Bottleneck-RFCAConv module reduces the number of channels in the middle layer by introducing an expansion coefficient, thereby reducing the computational complexity. It uses a two-layer convolution module: the first layer reduces the number of channels through standard convolution operations, and the second layer uses the RFCAConv module to enhance the feature representation capability; (3) After the receptive field spatial features extracted by the RFCAConv module are processed, the input features are optimized using the CA attention mechanism; the CA attention mechanism decomposes the feature information in the height and width directions respectively to obtain feature representations along these two directions; (4) Transform the features by concatenating the feature maps globally in height and width and applying convolution, batch normalization, and nonlinear activation operations; (5) Split the intermediate features into two independent features, and perform feature transformation through convolution operation and Sigmoid activation function so that the output dimension is consistent with the input vector; calculate the attention weights of the input feature map in the height and width dimensions, combine the calculated attention weights with the extracted feature map, and generate the final output.
6. A lightweight rice disease prediction method based on SCL-YOLOv8n according to claim 1, characterized in that: The Transformer branch uses the MHSA-CGLU module to establish global dependencies through multi-head self-attention MHSA, and combines the convolutional gating mechanism CGLU to transform features.
7. A lightweight rice disease prediction method based on SCL-YOLOv8n according to claim 6, characterized in that: The multi-head self-attention mechanism MHSA maps the input sequence into multiple subspaces, each of which corresponds to an independent attention head; its specific expression is: (2), where Q, K, and V represent query, key, and value vectors, respectively, and d k is the dimension of the key vector, which is used to scale the dot product to prevent the dot product result from being too large, thereby avoiding the gradient from disappearing. T represents the transposed matrix, and softmax represents the normalization operation; (1) Introducing a multi-head self-attention mechanism, multiple attention heads are calculated in parallel, and each head learns different representations in its own subspace; The specific process is expressed by formula (3) and formula (4): , where , , , is the parameter matrix learned by the model, Concat means concatenation operation, head means attention head, h means the number of attention heads, and i means the i-th one; (2) CGLU adds a 3x3 depth convolution operation before the gated linear unit GLU activation function; (3) The outputs of the two branches are concatenated along the channel dimension and mapped to the final output channel number C_out through a 1x1 convolution to complete feature fusion.
8. The lightweight rice disease prediction method based on SCL-YOLOv8n according to claim 1, characterized in that: In step 3, the lightweight shared enhanced detection head LSCSBD operates as follows: (1) Input feature maps of different scales into the network, and map the number of channels of each feature map from C_i to the number of hidden channels hidc through a 1x1 convolutional layer, while keeping the spatial size unchanged; (2) At each detection layer, a 3×3 shared convolution is applied to extract spatial features. After the convolution operation, the output features are batch normalized (BN) and nonlinearly transformed using the SiLU activation function. The features after SiLU processing are further subjected to a second 3×3 shared convolution operation, BN normalization, and SiLU activation to enhance feature expression capabilities. (3) The output is a tensor containing bounding box regression and category scores. The entire process rescales the output of each detection layer through the Scale operation to ensure consistency and accuracy in multi-scale detection tasks.
Citation Information
Patent Citations
Rice pest detection method and device based on YOLOv5s algorithm
CN117523354A
Lightweight multi-environment tomato detection method based on improved yolov8
CN117557787A