A lightweight object detection method and system for abnormal rice detection

By improving the YOLOv11n model, combining deep convolution, SimAM and BiFPN architectures, the problem of low accuracy and efficiency in abnormal rice detection is solved, and efficient and accurate detection results are achieved, which are suitable for food quality control in large-scale production lines.

CN120047678BActive Publication Date: 2025-07-11JILIN AGRICULTURAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510525246.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-11
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing neural network technology has low accuracy and efficiency in abnormal rice detection, making it difficult to meet the needs of modern production lines for detection speed and accuracy. Especially when dealing with abnormal situations such as rice surface defects, foreign body pollution and mildew, there are problems such as insufficient sensitivity and excessive calculation volume.

Method used

Based on the YOLOv11n model, by reducing the amount of parameters and calculations, the attention mechanism and the information fusion architecture are introduced, combined with the deep convolution module, the SimAM attention mechanism and the BiFPN architecture, the model is improved to be suitable for abnormal rice detection.

Benefits of technology

It significantly improves the accuracy and efficiency of abnormal rice detection, reduces the calculation cost, and the model can operate efficiently under limited resources, is suitable for quality inspection in large-scale production lines, replaces traditional manual screening methods, improves production efficiency and ensures food safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047678B_ABST
    Figure CN120047678B_ABST
Patent Text Reader

Abstract

A lightweight object detection method and system for detecting abnormal rice. It belongs to the technical field of neural network object detection, and specifically relates to the technical field of abnormal rice detection. It solves the technical problem that the existing neural network technology has low detection accuracy and efficiency for abnormal rice. The method includes the following steps: Dataset construction: Collect pictures of different types of abnormal rice, perform category annotation, and divide the training set, validation set, and test set; Model construction: Based on the YOLOv11n model, combined with the morphological characteristics of rice, improve the YOLOv11n model to construct a model suitable for detecting abnormal rice; Model training: Use the constructed dataset to train the model suitable for detecting abnormal rice, and adjust the model parameters until the model meets the detection requirements; Use the trained model suitable for detecting abnormal rice to detect abnormal rice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural network object detection, and specifically relates to the technical field of abnormal rice detection. Background Art

[0002] With the continuous growth of the global population, food security and food quality issues have attracted increasing attention. As one of the most important foods globally, rice is widely used in the daily diets of various countries, and its quality is directly related to people's health and quality of life. Rice is susceptible to various factors during production, processing, and transportation, and problems such as foreign object contamination, mildew, and breakage may occur. These abnormal rices not only affect product quality but also pose a threat to consumer health. Therefore, how to efficiently and accurately detect and screen abnormal rices in large-scale production has become an important challenge in current food quality control.

[0003] Traditional rice quality detection methods mainly rely on manual screening. This method not only has a large workload and low efficiency but is also easily affected by human factors, making it difficult to meet the requirements of modern production lines for detection speed and accuracy. To improve the automation level and accuracy of rice detection, computer vision and deep learning technologies have gradually been introduced into the field of food quality detection. Object detection algorithms, especially deep learning-based object detection models, have achieved remarkable results in the fields of image recognition and object detection, but their application in the food industry still faces some specific challenges, including the diversity of different rice varieties, the appearance differences of abnormal rices, and the high-efficiency requirements of the model in real-time detection.

[0004] And currently, due to problems such as the inability of existing object detection models to accurately adapt to the sensitivity of abnormal situations such as rice surface defects, foreign object contamination, and mildew, and the excessive computational complexity during detection, the detection accuracy and efficiency of existing neural network technologies for abnormal rices are relatively low. Summary of the Invention

[0005] To solve the technical problem of the relatively low detection accuracy and efficiency of existing neural network technologies for abnormal rices, the present invention provides a lightweight object detection method for abnormal rice detection, and the method includes the following steps:

[0006] S1. Dataset construction: Collect pictures of different types of abnormal rices, perform category annotation, and divide them into training sets, validation sets, and test sets;

[0007] S2. Model construction: Based on the YOLOv11n model, combined with the morphological characteristics of rice, improve the YOLOv11n model to construct a model suitable for abnormal rice detection;

[0008] The improvement of the YOLOv11n model is specifically as follows:

[0009] Reduce the number of model parameters and computational complexity;

[0010] Introduce an attention mechanism;

[0011] Optimize the information fusion architecture;

[0012] S3. Model training: Use the constructed dataset to train the model applicable to abnormal rice detection, and adjust the model parameters until the model meets the detection requirements;

[0013] S4. Use the trained model applicable to abnormal rice detection to detect abnormal rice.

[0014] Furthermore, the different types of abnormal rice include four types, namely broken rice, contaminated and discolored rice, intact and edible rice, and moldy rice.

[0015] Furthermore, the specific method for reducing the number of model parameters and computational complexity is: Replace the fifth convolutional module starting from the input in the backbone network of the YOLOv11n model with a depthwise convolutional module.

[0016] Furthermore, the specific method for introducing the attention mechanism is: Add a SimAM module after the C2PSA module in the backbone network of the YOLOv11n model.

[0017] Furthermore, the specific method for optimizing the information fusion architecture is: Introduce a BiFPN architecture in the neck part of the YOLOv11n model to change the information fusion method in the neck part of the YOLOv11n model.

[0018] Furthermore, the specific method for changing the information fusion method in the neck part of the YOLOv11n model is:

[0019] S1. Connect the first C3K2 module passed through from the input in the backbone part of the YOLOv11n model to the neck part;

[0020] S2. Add a convolutional module between the modules where the backbone part and the neck part of the YOLOv11n model are connected;

[0021] S3. Replace the Concat module in the neck part with a Fusion module, and perform weighted fusion on features at different levels through learnable weights;

[0022] S4. Change the single-direction information flow method in the neck part of the YOLOv11n model to a two-way cross-scale interaction method.

[0023] Furthermore, the information flow in the neck part of the model applicable to abnormal rice detection is divided into four paths:

[0024] The modules passed in the first path are a convolution module, a Fusion module, and a C3K2 module in sequence;

[0025] In the second path, it branches into two after the convolution module; the first branch passes through the Fusion module in the first path and the C3K2 module in the first path in sequence; the second branch passes through the Fusion module, the C3K2 module, the Fusion module in the first path, and the C3K2 module in the first path in sequence;

[0026] In the third path, it branches into two after the convolution module; the first branch passes through the Fusion module, the C3K2 module, the upsampling module, the Fusion module in the second path, the C3K2 module in the second path, the Fusion module in the first path, and the C3K2 module in the first path in sequence; the second branch passes through the Fusion module, the C3K2 module, the convolution module and then inputs into the Fusion module in the fourth path;

[0027] In the fourth path, it branches into two after the convolution module; the first branch inputs into the Fusion module of the first branch in the third path after passing through the upsampling module; the second branch passes through the Fusion module and the C3K2 module in sequence;

[0028] The information output by the C3K2 module in the first path enters the head network, and the C3K2 module in the first branch also inputs the information into the Fusion module in the third path through a convolution module;

[0029] The information output by the C3K2 module of the second branch in the third path enters the head network;

[0030] The information output by the C3K2 module of the second branch in the fourth path enters the head network.

[0031] The present invention also provides a lightweight object detection system for abnormal rice detection, and the system includes the following modules:

[0032] A module for dataset construction: collecting pictures of different types of abnormal rice, performing category annotation and dividing into training set, validation set and test set;

[0033] A module for model construction: based on the YOLOv11n model, combining the morphological characteristics of rice, improving the YOLOv11n model to construct a model suitable for abnormal rice detection;

[0034] The improvement of the YOLOv11n model is specifically as follows:

[0035] Reducing the number of parameters and computational amount of the model;

[0036] Introducing an attention mechanism;

[0037] Optimize the information fusion architecture;

[0038] Module for model training: Use the constructed dataset to train the model suitable for abnormal rice detection, and adjust the model parameters until the model meets the detection requirements;

[0039] Module for detecting abnormal rice using the trained model suitable for abnormal rice detection.

[0040] The beneficial effects of the method of the present invention are:

[0041] By optimizing the YOLOv11n object detection architecture, the depth convolution module, the SimAM parameter-free attention mechanism are innovatively introduced, and the efficient BiFPN architecture is adaptively introduced in the neck, which re-changes the information processing method of the neck, enabling the model to significantly reduce the computational cost and the number of parameters while ensuring high detection accuracy. Experimental results show that the method of the present invention has achieved remarkable results in the task of abnormal rice detection. The mAP50 has increased by 5.27% compared with the original model, and its number of parameters has decreased by 36.86% (i.e., 0.952M) compared with the original model, showing high detection accuracy and low computational overhead. The lightweight design of the present invention enables the model to operate efficiently under limited device resources, with good practicability and scalability. Through the streamlined network structure and combined with innovative algorithm optimization means, the computational complexity of the model is effectively reduced, and the detection task of abnormal rice can be carried out quickly, which is particularly important for the quality detection system in large-scale production lines. In addition, the high accuracy and low computational cost of the method of the present invention enable it to be widely applied to various production environments. Especially in the automated production process, it can replace the traditional manual screening method, improve production efficiency, reduce human interference factors, and further ensure food safety. Description of the Drawings

[0042] Figure 1 Flowchart of the lightweight object detection method for abnormal rice detection in the embodiment of the present invention;

[0043] Figure 2 Schematic diagram of the rice image acquisition device in the embodiment of the present invention;

[0044] Figure 3 Schematic diagram of the labels corresponding to different types of rice in the embodiment of the present invention;

[0045] Figure 4 Structure diagram of the Rice-YOLO-AD model in the embodiment of the present invention;

[0046] Figure 5It is a simplified diagram of the YOLOv11n model structure in the embodiments of the present invention;

[0047] Figure 6 It is a simplified diagram of the structure of the Rice-YOLO-AD model in the embodiments of the present invention;

[0048] Figure 7 It is a schematic diagram of the structures of ordinary convolution and depth convolution in the embodiments of the present invention;

[0049] Figure 8 It is a schematic diagram of the structure of the SimAM attention mechanism in the embodiments of the present invention;

[0050] Figure 9 It is a schematic diagram of the structure of BiFPN in the embodiments of the present invention;

[0051] Figure 10 It is the convergence curve of the Rice-YOLO-AD model during the training process in the embodiments of the present invention;

[0052] Figure 11 It is a comparison chart of the mAP50 curves of the Rice-YOLO-AD model and other classical models in the embodiments of the present invention;

[0053] Figure 12 It is a schematic diagram of five different positions of placing SimAM into the backbone of YOLOv11 in the embodiments of the present invention;

[0054] Figure 13 It is a comparison chart of the detection results of the YOLOv11n model and the Rice-YOLO-AD model for intact edible rice in the embodiments of the present invention;

[0055] Figure 14 It is a comparison chart of the detection results of the YOLOv11n model and the Rice-YOLO-AD model for broken rice in the embodiments of the present invention;

[0056] Figure 15 It is a comparison chart of the detection results of the YOLOv11n model and the Rice-YOLO-AD model for moldy rice in the embodiments of the present invention. Detailed implementation manners

[0057] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] Example 1

[0059] This embodiment provides a lightweight target detection method for abnormal rice detection. Figure 1 As shown, the method comprises the following steps:

[0060] S1. Dataset construction: Collect different types of abnormal rice images, label them, and divide them into training sets, validation sets, and test sets;

[0061] S2. Model construction: Based on the YOLOv11n model, the YOLOv11n model is improved in combination with the morphological characteristics of rice to construct the Rice-YOLO-AD model suitable for abnormal rice detection;

[0062] The improvements to the YOLOv11n model are as follows:

[0063] Reduce the number of model parameters and computational complexity;

[0064] Introducing the attention mechanism;

[0065] Optimize information fusion architecture;

[0066] S3, model training: use the constructed data set to train Rice-YOLO-AD and adjust the model parameters until the model meets the detection requirements;

[0067] S4. Use the trained Rice-YOLO-AD to detect abnormal rice.

[0068] As an efficient target detection algorithm, YOLO (You Only Look Once) is widely used in the field of object recognition due to its advantages of fast speed and high accuracy. Nevertheless, YOLO still has certain limitations when dealing with food detection, such as the balance between processing speed and detection accuracy, especially when facing large-scale data sets. Therefore, in response to the special needs of rice detection, this embodiment proposes an improved YOLOv11n model, Rice-YOLO-AD (Rice-based You Only Look Once Anomaly Detection), which aims to solve the problem of efficiency and accuracy of rice surface anomaly detection. Compared with previous target detection algorithms, YOLOv11 introduces a more efficient architecture, including efficient attention mechanisms such as C3K2, SPPF and C2PSA. YOLOv11 aims to enhance small object detection and improve accuracy while maintaining YOLO's consistent real-time inference speed. As a lightweight model in the YOLOv11 series, YOLOv11n has a lower number of parameters and computational complexity while maintaining high accuracy, making it easier to deploy in resource-constrained environments.

[0069] Rice-YOLO-AD combines the real-time advantage of the YOLOv11 framework and is optimized for the characteristics of rice. By adding an anomaly detection module, it improves the sensitivity to anomalies such as surface defects, foreign object contamination, and mildew of rice. Compared with traditional image processing methods, Rice-YOLO-AD significantly improves the computational efficiency while ensuring high detection accuracy, meeting the requirements of industrial production.

[0070] Example 2

[0071] This example further limits Example 1 and further limits step S1.

[0072] The rice image acquisition device is as Figure 2 shown, consisting of an image capture device, a lighting device, and a black light-absorbing board. The entire shooting environment is carried out in a darkroom. The rice sample category selected in this example is Wuchang long-grain fragrant rice. The rice is randomly scattered on the black light-absorbing cloth. Without rice overlapping or adhering, the RGB image is obtained by the capture device. In this study, the capture device used is the mobile phone Honor Magic6, which is fixed on the top bracket and shot vertically. The obtained images are labeled using the Make Sense online annotation website.

[0073] In this example, the rice is divided into the following four categories.

[0074] Broken rice (category 0): Broken rice is usually caused by external forces during harvesting, processing, or transportation. Broken rice not only affects the taste but may also bring some potential food safety problems.

[0075] Contaminated and discolored rice (category 1): Discolored rice often shows discoloration due to contamination, improper storage conditions, or being invaded by microorganisms such as molds and bacteria. This type of rice may carry harmful substances and requires special attention.

[0076] Intact and edible rice (category 2): This category of rice represents good product quality and is not affected by problems such as breakage, contamination, or mildew.

[0077] Moldy rice (category 3): Moldy rice is usually caused by mildew due to humid, warm, or poorly ventilated storage conditions. Mildew not only affects the edible safety of rice but may also cause mycotoxins harmful to human health.

[0078] This detailed classification can help the model improve the recognition accuracy of each type of abnormal rice, avoid misjudging abnormal rice with similar appearances as other categories, and thus improve the precision and robustness of detection. At the same time, this classification also provides more refined data for model design, helping the model to be better optimized and debugged. Dividing rice into four categories: broken, contaminated and discolored, edible, and moldy not only meets the actual application requirements but also helps to improve the accuracy and pertinence of the object detection algorithm. This classification method enables the model to identify and process different types of abnormal rice separately in complex detection tasks, thereby achieving more accurate and efficient quality control.

[0079] As Figure 3 shown, label 0 represents broken rice, label 1 represents contaminated and discolored rice, label 2 represents intact and edible rice, and label 3 represents moldy rice.

[0080] (a) shows all normal rice, (b) shows the effect of a mixture of broken rice and normal rice, (c) shows the effect when all four types of rice are present, and (d) shows the effect when the four types of rice are mixed very densely.

[0081] The dataset has a total of 64 images, which are divided into a training set, a validation set, and a test set according to a ratio of 4:1:1. In the rice dataset, the features between samples are highly consistent and representative. In addition, the number of rice grains in each image is large and the distribution is relatively dense, which enables each sample to fully display the key features of the rice. Therefore, even though the number of images is relatively limited, the labeled data can still contain rich and effective information. These effective information can help the model effectively identify abnormal rice, thereby achieving precise detection of rice quality. By making full use of these features and information, even when the amount of data is not large, the efficiency of the model and the accuracy of detection can still be guaranteed. At the same time, it adapts to the application scenarios in the real world, avoiding the difficulty of collecting a large number of labeled samples and reducing the dependence on large-scale labeled data.

[0082] Example 3,

[0083] This example further limits Example 1 and further explains step S2.

[0084] As Figure 4 shown is the structural diagram of the Rice-YOLO-AD model in this example. Due to limited space, abbreviations are given for each module, and the corresponding full English names of each module are given here:

[0085] FFN: Feed-Forward Network, PSA: Pyramid Spatial Attention Mechanism, C2PSA: Compound Two-Path Spatial Attention, SPPF: Spatial Pyramid Pooling-Fixed, C3K: C3K Block, C3K2: C3K Block Applied Twice, SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks, BatchNorm2d: 2-Dimensional Batch Normalization Layer, SiLU: SiLU Activation Function, k: Size of the Convolutional Kernel, s: Size of Step, p: Size of Padding。

[0086] This embodiment improves the YOLOv11n as the basic model, aiming to improve its recognition and detection ability of rice. Figure 5 It is a simplified diagram of the YOLOv11n model structure. To make the model more efficient and improve the recognition accuracy, the main improvement strategies can be summarized as follows:

[0087] First, a common convolution (Conv) module in the backbone network of the model is replaced with a depthwise convolution (DWConv) module. As an efficient convolution operation method, depthwise convolution can significantly reduce the number of model parameters and computational volume compared with traditional common convolution. The main purpose of this improvement is to enhance the operation efficiency of the model when dealing with the rice recognition task. By reducing the computational complexity, DWConv enables the model to be more efficient in training and inference on the dataset, thus accelerating the response speed in the rice detection process and saving computational resources. However, although DWConv can improve the computational efficiency of the model to a certain extent, it may also lead to a decrease in the recognition accuracy of the model.

[0088] Common convolution and depthwise convolution are two common convolution methods in convolutional neural networks. They are different in structure, calculation method, and application scenarios, each with its own unique advantages and disadvantages. Their schematic diagrams are as Figure 7As shown. Ordinary convolution can flexibly adjust parameters such as the size, stride, and padding of the convolution kernel according to task requirements. By adjusting the number and parameters of the convolution kernels, the complexity and feature extraction ability of the model can be controlled. Depthwise convolution has extremely low number of parameters and computational amount. It performs independent convolution operations on each input channel, thus greatly reducing the number of parameters and computational complexity of the model. Therefore, it is more suitable for model lightweighting. Since each channel performs convolution operations independently, depthwise convolution can better preserve the spatial information of the input data.

[0089] During model training and inference, the size of the number of parameters directly determines the required computational amount. Reducing the number of parameters can speed up the training and inference speeds, improve the overall computational efficiency, enable the model to have faster inference ability for rice, and be more suitable for use in resource-constrained environments. Depthwise convolution performs independent convolution operations on each input channel, while ordinary convolution performs convolution operations on the entire input feature map. The number of convolution kernels in depthwise convolution is the same as the number of input channels, and each convolution kernel is only responsible for the convolution operation of one channel, while the number of convolution kernels in ordinary convolution is more than the number of input channels, and each convolution kernel can process data of multiple channels. Assuming that the number of input and output channels are and , is the kernel size of the convolution, the size of the picture is , and the formulas for the number of parameters and computational amount of ordinary convolution and depthwise convolution are shown as follows:

[0090] Formula for the number of parameters of ordinary convolution:

[0091] (1)

[0092] Formula for the computational amount of ordinary convolution:

[0093] (2)

[0094] Formula for the number of parameters of depthwise convolution:

[0095] (3)

[0096] Formula for the computational amount of depthwise convolution:

[0097] (4) In order to solve the problem of accuracy degradation caused by DWConv, the SimAM parameter-free attention mechanism is further introduced. SimAM is an attention mechanism that can effectively improve the performance of the model. It adjusts the weights of the feature map in a parameter-free way to enhance the representation ability of key features. The introduction of the SimAM attention mechanism can not only maintain the advantages of parameters and computation brought by DWConv, but also improve the recognition accuracy of the model by accurately capturing important features. Therefore, SimAM can make up for the negative impact of DWConv on accuracy, thereby ensuring that the recognition ability of the model is effectively improved while reducing the computational complexity.

[0098] SimAM (Similarity Attention Module) is a module based on the self-attention mechanism. Its core idea is to strengthen the model's attention to important features by calculating the similarity between samples, thereby improving the model's performance. Compared with the traditional attention mechanism, SimAM does not need to calculate complex weighted matrices or obtain the relationship between features through large-scale matrix operations. It only distributes attention by calculating the similarity between features, thereby reducing the complexity of calculation. This feature makes SimAM more efficient in calculation than some classic self-attention modules.

[0099] In the present invention, SimAM can effectively transfer important information between the layers of the network by focusing on the similarities between rice, and can better mine the potential important details in the image when processing abnormal rice detection. Since the target rice in this study is relatively small, it is relatively difficult to identify, and SimAM can emphasize and process the features more specifically, improving the sensitivity of the model to key areas and local features, thereby improving the recognition effect of rice.

[0100] Many deep learning models usually require a large amount of labeled data to improve their performance during training. SimAM can achieve good performance on smaller data sets by effectively enhancing the similarity relationship between features. This enables it to perform better than other complex models in certain data-scarce situations and has better data efficiency.

[0101] The schematic diagram of the SimAM model is as follows Figure 8 As shown in the figure, Generation represents the process of generating attention weights, and Expansion represents the process of expanding or amplifying the generated attention weights. Fusion is a fusion method that fuses feature maps or attention weights from different sources to generate richer feature representations.

[0102] The contributions of SimAM to performance improvement mainly include integrating global and local features, capturing the nuances and complex structures of input data through 3D attention weighting, making the model more accurate and efficient in handling complex tasks, the parameter-free design simplifies the model architecture, reduces computational complexity, improves performance, and is suitable for resource-constrained environments. Quantifying the uniqueness of neurons and their correlations through an energy function optimizes the attention mechanism. The simplest implementation of finding these neurons is to measure the linear separability between a target neuron and other neurons, and based on this, the following energy function is defined:

[0103] (5)

[0104] where, = and = are t and linear transformations of, t and are the target neuron and other neurons in a single channel of the input feature . i is the index in the spatial dimension, M = H×W is the number of neurons on this channel. and are the weighted sum and bias of the transformation. All values in formula (5) are scalars. When equals , and all other are , formula (5) reaches its minimum value, where and are two different values. By simplifying this formula, formula (5) is equivalent to finding the linear separability between the target neuron t and all other neurons in the same channel. For simplicity, binary labels (i.e., 1 and -1) are adopted, and a regularizer is also added to formula (5). The final energy function is as follows:

[0105] (6)

[0106] Theoretically, there are M energy functions for each channel. Solving all these equations through some iterative solvers (such as SGD) is computationally burdensome. Among them, formula (6) has a fast closed-form solution for and , which can be obtained as follows:

[0107] (7)

[0108] (8)

[0109] Among them, and are the mean and variance calculated on all neurons except t in this channel. Since the existing solutions shown in formulas (7) and (8) are obtained on a single channel, it is reasonable to assume that all pixels in a single channel follow the same distribution. Given this assumption, the mean and variance can be calculated on all neurons and reused for reporting on all neurons in this channel. It can significantly reduce the computational cost and avoid iteratively calculating µ and σ for each position. Therefore, the minimum energy can be calculated by the following formula:

[0110] (9)

[0111] and . Equation (9) shows that the lower the energy , the greater the difference between neuron t and its surrounding neurons, and the more important it is for visual processing. Therefore, the importance of each neuron can be obtained by 1 / . Using a scaling operator to refine the features, the entire optimization stage of this module is:

[0112] (10)

[0113] where E is all spanning the channel and spatial dimensions, and ⊙ represents element-wise multiplication. A sigmoid function is added to limit the overly large values in E and multiply it with the original feature map to obtain a weighted feature map. It does not affect the relative importance of each neuron because is a monochannel function.

[0114] Finally, an efficient BiFPN (Bidirectional Feature Pyramid Network) architecture is introduced at the neck of the network. BiFPN is an architecture that optimizes feature fusion through bidirectional information flow and can significantly improve the fusion effect of multi-scale features. By introducing BiFPN, the model can better transmit information in multi-scale feature maps while maintaining a low computational cost and the number of parameters, enhancing the model's sensitivity to rice.

[0115] BiFPN is an efficient network architecture for computer vision tasks, especially for effective feature fusion at different scales in images. The main purpose of BiFPN is to enhance multi-scale feature representation through an efficient feature pyramid architecture. Traditional Feature Pyramid Networks (FPN) only perform top-down feature fusion, while BiFPN supports both top-down and bottom-up feature flows by introducing bidirectional connections. This bidirectional processing can better capture feature information at different scales and improve the utilization efficiency of multi-scale information. BiFPN adopts a weighted feature fusion method. By introducing learnable weights, BiFPN can dynamically adjust the fusion method of features at different scales, thus avoiding the computationally complex fully connected structure. This method can reduce unnecessary computational volume and improve computational efficiency.

[0116] The structure of BiFPN is as Figure 9 shown. P3, P4, P5, P6, P7 represent the output layers of the backbone network. Each output layer has corresponding output features (including information such as the number of channels and the size of the features). For example, the output feature size of P3 is the input image resolution / 2^3, the output feature size of P4 is the input image resolution / 2^4, and so on. The output feature size of P7 is the input image resolution / 2^7, which are P3_in, P4_in,..., P7_in in the figure respectively. The hollow circles without color represent features, and the solid circles with color represent operators. All those connected by lines represent the weight W. The calculation formula for each operator is as follows:

[0117] (11)

[0118] (12)

[0119] (13)

[0120] (14)

[0121] (15)

[0122] (16)

[0123] (17)

[0124] (18)

[0125] Among them, Resize is usually an upsampling or downsampling operation for resolution matching. Represents a small positive number to prevent the denominator from being zero, ensuring the stability of numerical calculations and the smooth progress of training. Conv is usually a convolution operation for feature processing. BiFPN can generate more powerful feature representations through multiple cascades and efficient feature fusion. Since BiFPN can more effectively fuse multi-scale features, it can usually improve the accuracy in detection tasks, especially in scenarios where the target scale varies greatly. It can improve the detection performance of small and large objects, enhancing the overall effect of the model. At the same time, by introducing an adjustable weighting mechanism, it reduces the need for manual design and adjustment, enhancing the network's automatic learning ability.

[0126] Introducing BiFPN (Bidirectional Feature Pyramid Network) into the Neck of YOLOv11 has the following core advantages compared to the Neck architecture of the original model:

[0127] Improved feature fusion ability: In the neck part of the original model YOLOv11n, when fusing features at different levels, a simple concatenation method was used, failing to effectively distinguish the importance of different scale features, which may introduce noise interference or semantic deviation. After introducing the BiFPN architecture, different level features are weighted and fused through learnable weights, and these weights are automatically optimized during backpropagation. This method can dynamically adjust the contribution degrees of high-level semantic information and low-level detail information, significantly improving the detection accuracy for small targets (such as rice) and occluded targets.

[0128] Optimized multi-scale information flow: The neck part of the original model YOLOv11n only supports unidirectional information flow. Although it supports top-down (high level → low level) and bottom-up (low level → high level) paths, this leads to redundant calculations. After introducing the BiFPN architecture, the network can achieve bidirectional cross-scale interaction between high and low levels, forming a closed-loop feature reuse, enhancing the multi-scale expression ability, and optimizing the computational efficiency through parameter sharing.

[0129] Strong adaptability, supporting composite scaling: Due to the fixed hierarchical structure of the neck part of the original YOLOv11n, it is difficult to flexibly adapt to different hardware scenarios (such as edge devices requiring lower computational power). After introducing the BiFPN architecture, composite scaling can be achieved, uniformly adjusting the depth, width, and input resolution of the neck part, thus achieving a balance between model accuracy and computational efficiency to meet the requirements of different hardware platforms.

[0130] In summary, the YOLOv11n model has been optimized and improved by introducing DWConv, SimAM parameter-free attention mechanism, and BiFPN architecture. Through the combination of these strategies, not only the computational amount and number of parameters of the model are effectively reduced, but also its accuracy and efficiency in the rice recognition task are improved, ensuring good performance of the model in practical applications.

[0131] As shown Figure 6 in the structural simplified diagram of the Rice-YOLO-AD model, the information flow in the neck part of the model is divided into four paths:

[0132] In the first path, the modules passed through are the convolutional module, the Fusion module, and the C3K2 module in sequence;

[0133] In the second path, it branches into two after the convolutional module; the first branch passes through the Fusion module in the first path and the C3K2 module in the first path in sequence; the second branch passes through the Fusion module, the C3K2 module, the Fusion module in the first path, and the C3K2 module in the first path in sequence;

[0134] In the third path, it branches into two after the convolutional module; the first branch passes through the Fusion module, the C3K2 module, the upsampling module, the Fusion module in the second path, the C3K2 module in the second path, the Fusion module in the first path, and the C3K2 module in the first path in sequence; the second branch passes through the Fusion module, the C3K2 module, the convolutional module and then inputs into the Fusion module in the fourth path;

[0135] In the fourth path, it branches into two after the convolutional module; the first branch inputs into the Fusion module of the first branch in the third path after passing through the upsampling module; the second branch passes through the Fusion module and the C3K2 module in sequence;

[0136] The information output by the C3K2 module in the first path enters the head network, and the C3K2 module in the first branch also inputs the information into the Fusion module in the third path through a convolutional module;

[0137] The information output by the C3K2 module of the second branch in the third path enters the head network;

[0138] The information output by the C3K2 module of the second branch in the fourth path enters the head network.

[0139] Example 4

[0140] This embodiment further limits Embodiment 1 and further explains step S3. During the model training process, the input size of the dataset is set to 640 × 640, the number of training rounds is 200, the base learning rate is set to 0.01, the batch size is set to 16, and the optimizer used is SGD. The experiment is deployed on a computer equipped with an Intel(R) Xeon(R) W-2245 CPU (3.9GHz) and an NVIDIA Quadro RTX 5000 GPU (16GB), with the operating system being Windows 10, the software configuration installed as the Anaconda3-2021.11-windows version, using the PyCharm compiler, and given the Pytorch2.1.2 built-in Python3.8.19 programming language. All algorithms run in the same environment.

[0141] Embodiment 5

[0142] This embodiment further illustrates the beneficial effects achieved by the Rice-YOLO-AD model proposed in the present invention in the detection of abnormal rice through specific experimental data.

[0143] I. First, the model evaluation metrics are described.

[0144] This embodiment evaluates the performance of the model using metrics such as recall (R), precision (P), F1-score (F1), AP, and mAP. Taking a binary classification problem as an example, if the actual result is positive and the predicted result is positive, it is denoted as TP; if the actual result is negative but the predicted result is positive, it is denoted as FP; if the actual result is positive but the predicted result is negative, it is denoted as FN; if the actual result is negative and the predicted result is negative, it is denoted as TN. Precision is the ratio of the number of correctly predicted positive samples to the total number of samples predicted as positive. Recall is the ratio of the number of correctly identified positive samples to the total number of actual positive samples. The average value of the AP values for multiple classes is mAP (mean average precision), and the higher the value, the higher the average accuracy of the model's detection for each class. The F1-score is the harmonic mean of precision and recall, providing a single metric that balances the two.

[0145] GFLOPs (Giga Floating-Point Operations Per Second), that is, one billion floating-point operations per second. Weights represent weights. Postprocess per image represents postprocessing per image, that is, the steps of postprocessing a single image output by the model.

[0146] II. Ablation experiment

[0147] The specific content of the ablation experiments in this study is shown in Tables 1 and 2. When introducing depthwise convolution (DWConv) alone, the experimental results are consistent with expectations, and the overall recognition effect of the model shows a decline. However, this change leads to a 11.34% reduction in the number of model parameters, a 4.76% decrease in GFLOPs (floating point operations per second), a 10.91% reduction in weights, and a 54.72% reduction in the post-processing time per image. This result indicates that although DWConv performs well in reducing computational volume and model parameters, its negative impact on recognition performance cannot be ignored. When adding the SimAM module alone, the complexity of the model remains unchanged, but its recognition performance is significantly improved. Specifically, mAP50 increases by 4.74%, mAP50-95 increases by 5.42%, recall rate increases by 5.25%, precision increases by 2.68%, F1 score increases by 3.96%, and the post-processing time per image decreases by 56.6%. This result further verifies the effectiveness of the SimAM module in improving model performance, indicating that it can significantly enhance recognition accuracy without increasing the model complexity. When introducing both DWConv and SimAM modules, although the introduction of DWConv is a factor causing the decline in recognition accuracy, the introduction of SimAM compensates for this impact to a certain extent. Specifically, both mAP50 and mAP50-95 are improved to varying degrees, and at the same time, the number of model parameters and computational volume of the model remain decreased, indicating that SimAM can effectively balance the accuracy loss brought by DWConv, thereby improving the overall performance. When introducing the BiFPN module, the experimental results show that mAP50 is increased by 5.25%, mAP50-95 is increased by 5.42%, recall rate is increased by 4.1%, while the number of parameters is reduced by 25.55%, weights are decreased by 23.64%, and the post-processing time per image is also significantly reduced by 60.38%. These results indicate that the BiFPN module improves recognition accuracy while significantly reducing the computational burden and enhancing the efficiency of the model. When applying these three strategies (DWConv, SimAM, and BiFPN) simultaneously, the overall performance of the model is greatly improved. Specifically, mAP50 is increased by 5.27%, mAP50-95 is increased by 7.35%, recall rate is increased by 4.72%, F1 score is increased by 1.41%, while the number of model parameters of the model is reduced by 38.86%, GFLOPs are decreased by 4.76%, weights are reduced by 34.55%, and the post-processing time per image is reduced by 54.72%. This combined strategy enables the model to not only achieve a significant improvement in recognition accuracy but also greatly reduce the computational cost while maintaining light weight and high efficiency. In summary, by reasonably introducing these modules, the method described in the present invention effectively improves the recognition accuracy of the model for rice, while optimizing its computational performance, achieving a balance between high efficiency and high accuracy.

[0148] Table 1:

[0149]

[0150] Table 2:

[0151]

[0152] The convergence curve of the Rice - YOLO - AD model during training is as Figure 10 shown. Among them, train / box_loss: The bounding box regression loss on the training set.

[0153] train / cls_loss: The classification loss on the training set.

[0154] train / dfl_loss: The distribution loss on the training set.

[0155] val / box_loss: The bounding box regression loss on the validation set.

[0156] val / cls_loss: The classification loss on the validation set.

[0157] val / dfl_loss: The distribution loss on the validation set.

[0158] metrics / precision(B): Precision, which measures the proportion of samples predicted as positive classes that are actually positive classes.

[0159] metrics / recall(B): Recall, which measures the proportion of all samples that are actually positive classes and are correctly predicted as positive classes.

[0160] metrics / mAP50 (B): The value of the average precision (AP) when the IoU threshold is 0.5.

[0161] metrics / mAP50 - 95 (B): The average value of the average precision when the IoU threshold ranges from 0.5 to 0.95.

[0162] As can be observed from the figure, the recognition performance of the model is poor in the initial stage of training and it is difficult to accurately recognize the target. The emergence of this phenomenon is mainly related to the characteristics and size of the target object. At the beginning of training, the model faces the relatively small and relatively simple target of rice. Therefore, it is very difficult for the model to effectively recognize this target in the initial stage. Specifically, due to the small size of rice and the small number of samples, large errors occur in the feature extraction and target localization processes of the model, thus affecting the recognition accuracy. However, as the number of training rounds gradually increases, the recognition effect of the model shows a significant improvement. This improvement mainly benefits from the continuous optimization of the model during training and the gradual adjustment of weights, enabling the network to better adapt to the complex feature space. Through backpropagation, the model gradually reduces the prediction error and optimizes its feature extraction ability. Especially in the case of small target objects, the model can better focus on the target of rice and gradually improve its recognition accuracy. From the changes in the loss curve and accuracy curve, the model gradually tends to be stable during training. The loss curve drops rapidly in the early stage of training, indicating that the model error is continuously decreasing and the prediction results are gradually approaching the true labels. At the same time, the mAP50 curve also rises rapidly, reflecting the gradual improvement of the model's prediction accuracy and finally tending to be stable. This phenomenon shows that after multiple rounds of iteration, the model has gradually completed the process of fitting the data and finally achieved an ideal recognition effect on the training set and validation set. In addition, as the number of training rounds increases, the model can gradually recognize more detailed features and show stronger generalization ability when dealing with small targets. Generally speaking, with the deepening of training, the model finally completes the relatively accurate recognition of rice through continuous optimization and adjustment.

[0163] III. Comparative Experiments

[0164] 1. Comparison with Classical Object Detection Models

[0165] To verify the performance advantages of Rice-YOLO-AD in the task of abnormal rice detection, this embodiment conducted comparative experiments with multiple classical object detection models (including YOLOv6n, YOLOv8n, YOLOv10n, YOLOv10s, YOLOv10m, YOLOv11n, and YOLOv12n). The experimental results are shown in Table 3. The experimental results indicate that Rice-YOLO-AD achieved 95.09% in the mAP50 metric, showing excellent performance. Specifically, although the mAP50-95 of YOLOv10m is 0.63% higher than that of Rice-YOLO-AD, in other key recognition metrics, Rice-YOLO-AD outperformed YOLOv10m. More importantly, the number of model parameters of YOLOv10m is 9.39 times that of Rice-YOLO-AD, the GFLOPs is 9.82 times, and the weights is 9.31 times that of Rice-YOLO-AD, indicating a significant consumption of computing resources. Therefore, although YOLOv10m performs slightly better in some metrics, its high computational overhead significantly reduces the efficiency in practical applications. In addition, YOLOv8n also showed outstanding performance in recognition, with its mAP50 being 1.14% higher than that of the original model YOLOv11n. However, compared with Rice-YOLO-AD, although YOLOv8n has a slight advantage in accuracy, it does not show obvious advantages in other aspects. Especially in terms of controlling computing resources and the number of parameters, Rice-YOLO-AD demonstrated a more superior cost performance. Therefore, considering all metrics comprehensively, Rice-YOLO-AD not only performs excellently in recognition accuracy but also has significant advantages in optimizing the computational efficiency and the number of model parameters of the model. In summary, Rice-YOLO-AD achieved a good balance between performance and efficiency. Especially in the task of rice object detection, compared with other classical object detection models, while ensuring high recognition accuracy, it can significantly reduce the consumption of computing and storage resources, showing superior application potential.

[0166] Table 3:

[0167]

[0168] Such as Figure 11As shown in the comparison experiment with the classical object detection model, the network model Rice-YOLO-AD proposed in this application demonstrated the optimal convergence effect on the mAP50 curve. Specifically, the mAP50 curve of Rice-YOLO-AD showed a rapid and stable upward trend during the training process, indicating that the model could significantly improve the accuracy of object detection within a short training time and quickly approach the best performance. This phenomenon shows that the model effectively optimized the feature extraction and decision-making mechanisms during the learning process. Especially in the initial stage of training, through the carefully designed network architecture and optimization strategy, the detection accuracy of the model was rapidly improved, and then convergence was achieved. Compared with other classical comparison models, the performance of Rice-YOLO-AD on the mAP50 curve was particularly stable with less fluctuation, indicating that the model demonstrated strong recognition ability and excellent generalization performance in the rice detection task. In contrast, the performance of YOLOv12n in the same experiment was relatively inferior. Especially in the initial stage of training, its mAP50 curve showed a significant decline during the rapid upward stage, indicating that it faced greater challenges during the optimization process, resulting in the failure to continuously improve the recognition effect of the model during training. This further confirmed the effectiveness of the optimization strategy and model design of Rice-YOLO-AD in solving the rice detection task.

[0169] 2. Comparative Experiments with Different Backbone Networks

[0170] In the field of object detection, the backbone network is usually the core part of feature extraction, which extracts meaningful feature information from the input image. These feature information will then be passed to subsequent network layers for object localization and classification. Therefore, for object detection tasks, the choice of the backbone network directly affects the detection accuracy, computational efficiency, and model complexity. To verify the advantages of the network model Rice-YOLO-AD, taking YOLOv11n as the base model, seven classic classification backbone networks were systematically compared (including RepHGNetV2, EfficientVit, FasterNet_T0, MobileNetV4_s, UniReplkNet_S, SwinTransformer, and DynamicHGNetV2), and the comparison results are shown in Table 4. In these comparative experiments, the SwinTransformer network showed the best recognition effect. Specifically, compared with Rice-YOLO-AD, SwinTransformer was 1.61% higher in the mAP50 metric, 3.94% higher in terms of precision, and 1.17% higher in the F1 score. However, although SwinTransformer has achieved a relatively significant improvement in performance, its cost cannot be ignored. The number of parameters of SwinTransformer is 18.22 times that of Rice-YOLO-AD, the model weight is 16.67 times that of Rice-YOLO-AD, and the time consumed by calculation is 3.52 times that of Rice-YOLO-AD. This result fully reflects the trade-off between the choice of the backbone network and the model performance and computational cost. Through these comparative experiments, the advantages of Rice-YOLO-AD in computational efficiency and the number of parameters were verified.

[0171] Table 4:

[0172]

[0173] 3. Comparison with Other Attention Mechanisms

[0174] To verify the superiority of the SimAM (Attention Mechanism) adopted in the present invention, seven classical attention modules, including CPCA, AFGCAttention, CAFM, MPCA, DAttention, TripletAttention, and MLCA, were compared in this experiment. The verification results are shown in Table 5. By comparing the performance of these models in terms of average recognition accuracy, CAFM performed the most prominently among all the modules, and its mAP50 value was 0.56% higher than that of SimAM. However, this advantage came at a cost. Specifically, the number of parameters of the CAFM model increased by 13.4% compared to the original model, the GFLOPs increased by 4.76%, and the weight also increased by 12.73% accordingly. This indicates that although CAFM has improved in terms of accuracy, it has paid a relatively high price in terms of computational complexity and model size. In contrast, SimAM has maintained a low computational cost and model complexity while improving accuracy. Specifically, SimAM significantly reduced the image post-processing time by 56.6% without increasing the number of additional parameters, GFLOPs, and weight. In addition, the mAP50 value of SimAM increased by 3.89% compared to the original model. This result shows that SimAM can effectively improve the recognition accuracy of the model while maintaining low computational resource consumption, especially achieving performance improvement without significantly increasing computational complexity. Through the results of the comparative experiment, it can be clearly seen that SimAM is superior to other classical attention mechanism modules in many aspects without adding a significant computational burden, demonstrating its broad applicability and potential in the task of detecting abnormal rice.

[0175] Table 5:

[0176]

[0177] 4. Comparison of Different Positions of the Attention Mechanism

[0178] In object detection algorithms, the attention mechanism improves the performance of the model by focusing on important regions or features in the image. Incorporating the attention mechanism into different positions of the backbone network of the object detection algorithm will have different effects on the performance of the model. This is because the attention mechanism at different positions acts on different feature layers, affecting feature extraction and fusion, and ultimately affecting the effect of object detection.

[0179] Therefore, in order to verify the rationality of the placement position of the attention mechanism in the present invention, the following comparative experiment was conducted, as Figure 12As shown, SimAM was placed in five different positions in the YOLOv11 backbone to test the recognition effect. The specific results are shown in Table 6. It is worth noting that the later the SimAM attention mechanism is placed, the better the mAP50. This is mainly because the features of the front network layers are relatively rough and contain a large amount of detailed information. As the network depth increases, the abstraction level of the features improves and the information becomes more compact. At this time, placing SimAM can better focus on meaningful global information, avoid excessive feature selection or weighting in the early stage, and reduce the loss of useful information. Since the feature maps of the backend network are usually smaller and have lower computational complexity than those of the front end, placing the attention mechanism in these layers will not introduce too much computational burden, but can produce a greater improvement in accuracy. At the same time, adding an attention mechanism to the front of the network backbone may lead to unnecessary weighting of low-level features, which may have less impact on the final object recognition in object detection and may instead increase the computational overhead. In contrast, placing SimAM at the back of the network can avoid this redundant calculation and focus more on the features that have a greater impact on the final object localization and classification. This result further verifies the superiority of the implementation scheme of the attention mechanism of the present invention.

[0180] Table 6:

[0181]

[0182] IV. Influence of Image Size on the Model

[0183] In object detection algorithms, the choice of image size has an important impact on performance. Different image sizes will affect the computational efficiency, accuracy of the algorithm, and the training effect of the model. Therefore, when designing and applying object detection algorithms, it is necessary to weigh the influence of image size. Thus, to verify the superiority of the image size 640×640 used in the present invention, this study compared the influence of 7 different image sizes on the model recognition results. The specific content is shown in Table 7, where the image size of 640×640 obtained the best recognition accuracy, further reflecting the superiority of the design of the present invention.

[0184] Since the targets of rice in the image are relatively small, when the image size is smaller, the model cannot recognize small targets or targets with more details, resulting in a weakened recognition effect of rice. Because in the network Rice-YOLO-AD, a smaller image size will lead to a reduction in the resolution of the feature map, thus affecting the network's learning ability of rice details. In addition, the convolutional operations of the network will have a greater impact on the loss of details of low-resolution images, reducing the accuracy. A larger input image will retain more detailed information, especially for those small targets. Increasing the image size helps to improve the detection accuracy.

[0185] Table 7:

[0186]

[0187] V. Identification Effect Display of Different Blending Methods

[0188] As Figures 13 - 15 shown, the identification effects before and after the model improvement were tested. The top picture shows the detection results of YOLOv11n, and the bottom picture shows the detection results of the Rice-YOLO-AD model. Normal represents intact and edible rice, and the following value represents the confidence level. Broken represents broken rice, mould represents mouldy rice, and dirty represents contaminated and discolored rice. In Figure 13 and Figure 14 , the improved model significantly increased the confidence level of the overall identification effect. Especially in the identification of broken rice, the accuracy was significantly improved. It is particularly worth noting that in Figure 15 , the original model failed to correctly distinguish between broken and mouldy rice and misidentified them as normal rice. However, the Rice-YOLO-AD model can accurately identify broken and mouldy rice, showing higher identification ability. Generally speaking, the Rice-YOLO-AD model demonstrates stronger identification ability when dealing with different types of rice. Especially in the face of mixed and dense situations, the identification accuracy has been significantly improved. This indicates that the Rice-YOLO-AD model plays a key role in enhancing the detection effect and improving the accuracy.

Claims

1. A lightweight object detection method for abnormal rice detection, characterized in that, The method includes the following steps: S1. Dataset construction: Collect pictures of different types of abnormal rice, perform category annotation, and divide them into training set, validation set, and test set; S2. Model construction: Based on the YOLOv11n model, combined with the morphological characteristics of rice, improve the YOLOv11n model to construct a model suitable for detecting abnormal rice; The improvement of the YOLOv11n model is specifically as follows: Reduce the number of parameters and computational complexity of the model, specifically by replacing the fifth convolutional module starting from the input in the backbone network of the YOLOv11n model with a depth convolutional module; Introduce an attention mechanism, specifically by adding a SimAM module after the C2PSA module in the backbone network of the YOLOv11n model; Optimize the information fusion architecture, specifically by introducing a BiFPN architecture in the neck part of the YOLOv11n model to change the information fusion method in the neck part of the YOLOv11n model. The steps are as follows: S61. Connect the first C3K2 module passed from the input in the backbone part of the YOLOv11n model to the neck part; S62. Add a convolutional module between the modules connected by the backbone part and the neck part of the YOLOv11n model; S63. Replace the Concat module in the neck part with a Fusion module to perform weighted fusion of features at different levels through learnable weights; S64. Change the unidirectional information flow mode in the neck part of the YOLOv11n model to a bidirectional cross-scale interaction mode; S3. Model training: Use the constructed dataset to train the model suitable for detecting abnormal rice, and adjust the model parameters until the model meets the detection requirements; S4. Use the trained model suitable for detecting abnormal rice to detect abnormal rice.

2. The lightweight object detection method for abnormal rice detection according to claim 1, wherein, The different types of abnormal rice include four types, namely broken rice, contaminated and discolored rice, intact and edible rice, and moldy rice.

3. The lightweight object detection method for abnormal rice detection according to claim 2, characterized in that The information flow in the neck part of the model suitable for detecting abnormal rice is divided into four paths: In the first path, the modules passed through are a convolutional module, a Fusion module, and a C3K2 module in sequence; In the second path, it branches into two branches after the convolutional module; the first branch passes through the Fusion module and the C3K2 module in the first path in sequence; the second branch passes through the Fusion module, the C3K2 module, the Fusion module in the first path, and the C3K2 module in the first path in sequence; In the third path, it branches into two branches after the convolutional module; the first branch passes through the Fusion module, the C3K2 module, an upsampling module, the Fusion module in the second path, the C3K2 module in the second path, the Fusion module in the first path, and the C3K2 module in the first path in sequence; the second branch passes through the Fusion module, the C3K2 module, a convolutional module and then inputs the Fusion module in the fourth path; The fourth path branches into two after the convolutional module; the first branch is input into the Fusion module of the first branch in the third path after passing through the upsampling module; the second branch passes through the Fusion module and the C3K2 module in sequence; The information output by the C3K2 module in the first path enters the head network, and the C3K2 module in the first branch also inputs the information into the Fusion module in the third path through a convolutional module; The information output by the C3K2 module of the second branch in the third path enters the head network; The information output by the C3K2 module of the second branch in the fourth path enters the head network.

4. A lightweight object detection system for abnormal rice detection, characterized in that, The system includes the following modules: Module for dataset construction: collect abnormal rice images of different types, perform category annotation and division of training set, validation set and test set; Module for model construction: based on the YOLOv11n model, combine the morphological characteristics of rice to improve the YOLOv11n model and construct a model suitable for abnormal rice detection; The improvement of the YOLOv11n model is specifically as follows: Reduce the number of parameters and computational complexity of the model, specifically by replacing the fifth convolutional module starting from the input in the backbone network of the YOLOv11n model with a depthwise convolutional module; Introduce an attention mechanism, specifically by adding a SimAM module after the C2PSA module in the backbone network of the YOLOv11n model; Optimize the information fusion architecture, specifically by introducing a BiFPN architecture in the neck part of the YOLOv11n model to change the information fusion method in the neck part of the YOLOv11n model. The steps are as follows: S41: Connect the first C3K2 module passed through from the input in the backbone part of the YOLOv11n model to the neck part; S42: Add a convolutional module between the modules where the backbone part and the neck part of the YOLOv11n model are connected; S43: Replace the Concat module in the neck part with a Fusion module to perform weighted fusion of different-level features through learnable weights; S44: Change the unidirectional information flow mode in the neck part of the YOLOv11n model to a bidirectional cross-scale interaction mode; Module for model training: use the constructed dataset to train the model suitable for abnormal rice detection, adjust the model parameters until the model meets the detection requirements; Module for detecting abnormal rice using the trained model suitable for abnormal rice detection.

5. A computer device, comprising: A processor and a memory, characterized in that the memory is used to store executable instructions of the processor, and the processor is configured to execute the lightweight object detection method for abnormal rice detection according to any one of claims 1-3 by executing the executable instructions.

6. A computer storage medium, characterized in that, A computer program is stored in the storage medium, and when the computer program runs, it executes the lightweight object detection method for abnormal rice detection according to any one of claims 1-3.