Schisandra chinensis target detection method and system based on lightweight deep learning network

By introducing improved technology of lightweight deep learning networks into the YOLOv7 model, the problems of insufficient detection capabilities and high computing overhead in Schisandra detection are solved, and real-time detection capabilities with high accuracy and low complexity are achieved, which are suitable for deployment on edge devices.

CN120236178AActive Publication Date: 2025-07-01SHENYANG AGRI UNIV

Patent Information

Application Number
CN202510301582.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing technology has insufficient detection capabilities in the detection of small-target fruits such as Schisandra chinensis, poor ability to adapt to complex environments, high computing overhead, difficult to deploy on edge devices, and insufficient feature fusion and attention mechanism optimization.

Method used

Using the improved YOLOv7 model based on lightweight deep learning network, the BiPConv partial dimensionality reduction convolution, BiPNet module, GSConv and VoV-GSCSP module in Slim-neck, and SE attention mechanism module, the feature extraction and fusion are optimized to improve the detection capability of the model in complex backgrounds.

Benefits of technology

The accuracy and real-time performance of Northern Schisandra detection have been significantly improved. Mean Average Precision (mAP) has reached 87.1%, and Precision (P) has been increased to 89.5%. The parameter quantity and calculation complexity are reduced, making it suitable for deployment on edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236178A_ABST
    Figure CN120236178A_ABST
Patent Text Reader

Abstract

The invention discloses a fructus schisandrae target detection method and system based on a lightweight deep learning network, and relates to the technical field of agricultural intelligent detection, automatic harvesting of traditional Chinese medicinal materials and ecological environment monitoring. Comprising the steps of obtaining a schisandra chinensis plant image data set; pre-processing the acquired image data set of the schisandra chinensis plant; dividing the preprocessed schisandra chinensis plant image data set into a training set and a test set; the basic YOLOv7 model is improved, and an improved YOLOv7 model is obtained; inputting the training set into the improved YOLOv7 model to obtain a trained and improved YOLOv7 model; and inputting the test set into the trained and improved YOLOv7 model, and evaluating the detection result of the schisandra chinensis. The detection capability of the improved YOLOv7 model on the fructus schisandrae under a complex background is remarkably improved, and effective real-time detection support is provided for intelligent identification of the fructus schisandrae.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of agricultural intelligent detection, automatic harvesting of traditional Chinese medicines, and ecological environment monitoring, and particularly relates to a Schisandra chinensis target detection method and system based on a lightweight deep learning network. Background Art

[0002] Object detection technology has been widely applied in the fields of agricultural intelligent detection, automatic harvesting of traditional Chinese medicines, and ecological environment monitoring. Currently, it is mainly divided into the following two categories:

[0003] (1) Object detection technology based on traditional computer vision. Traditional computer vision methods mainly rely on manually designed features for object detection. For example: color segmentation (using the color difference between fruits and leaves for object segmentation); edge detection (Canny and Sobel algorithms are used to detect the boundaries of fruits); morphological processing (methods such as dilation and erosion are used to remove background noise); histogram analysis (classifying objects based on the RGB or HSV color space). Such methods can achieve certain object detection effects in an ideal experimental environment, but in a complex orchard environment, they are easily affected by factors such as light changes, occlusion, and background noise, resulting in unstable detection.

[0004] (2) Object detection technology based on deep learning. In recent years, deep learning methods (CNNs, Convolutional Neural Networks) have made breakthroughs in the field of object detection. Common methods include: 1. R-CNN series (two-stage object detection): Faster R-CNN and Mask R-CNN use the Region Proposal mechanism for object detection and have high accuracy. However, they have a large computational overhead and are difficult to meet the requirements of agricultural real-time detection, and are not suitable for tasks such as unmanned aerial vehicle monitoring and robot picking; 2. YOLO (You Only Look Once) series (single-stage object detection): YOLOv3 / v4 / v5 is faster than the R-CNN series, but has limited small object detection ability, and the accuracy decreases in the detection tasks of small fruits such as Schisandra chinensis; YOLOv7: Compared with the previous generation models, it has improved the accuracy and real-time performance of object detection, but the computational amount is still large, and it is difficult to be deployed on edge computing devices (such as picking robots and unmanned aerial vehicles).

[0005] The following problems exist in the detection tasks of small target fruits such as Schisandra chinensis in the prior art:

[0006] 1. The small object detection ability is insufficient. The existing YOLO models are mainly optimized for medium and large targets. In the detection of small objects (such as Schisandra chinensis), it is easy to lose objects due to insufficient feature extraction. The fruits are densely distributed and partially overlapped, and it is difficult for the existing models to distinguish multiple adjacent fruits, resulting in a low recall rate and a high miss detection rate.

[0007] 2. Poor adaptability to complex environments. Orchard environments have characteristics such as large lighting variations (shadows, backlighting, strong light), leaf occlusion, and complex backgrounds. Traditional computer vision methods are almost ineffective in these situations, while the detection accuracy of existing deep learning models is unstable under different lighting conditions. For example, when the lighting is weak or the fruit is partially occluded, the mAP (mean average precision) of YOLOv5 drops by 5% - 10%, indicating that the existing technology has limited adaptability to complex environments.

[0008] 3. High computational overhead and difficulty in deploying on edge devices. Models such as Faster R - CNN and YOLOv7 have a large number of parameters and high computational complexity, making it difficult to deploy on edge devices with limited computing power (such as picking robots and drones). For example, the detection frame rate of YOLOv7 can reach 161FPS when running on an NVIDIA RTX 4060 graphics card, but on an embedded device (such as Jetson Nano), the frame rate drops to 10 - 15FPS, unable to meet the real - time detection requirements.

[0009] 4. Insufficient feature fusion and attention mechanism optimization. When existing YOLO models process small targets, feature information is easily lost during network propagation, affecting the detection effect. Lack of an efficient attention mechanism makes it difficult for the model to accurately focus on the fruit area in complex backgrounds (such as dense leaf occlusion), resulting in a high false detection rate.

[0010] Therefore, proposing a Schisandra chinensis target detection method and system based on a lightweight deep learning network to solve the difficulties existing in the prior art is an urgent problem for those skilled in the art. Summary of the Invention

[0011] In view of this, the present invention provides a Schisandra chinensis target detection method and system based on a lightweight deep learning network. The improved YOLOv7 model significantly enhances the detection ability of Schisandra chinensis in complex backgrounds, providing effective real - time detection support for the intelligent identification of Schisandra chinensis.

[0012] To achieve the above objectives, the present invention adopts the following technical solutions:

[0013] A Schisandra chinensis target detection method based on a lightweight deep learning network, comprising the following steps:

[0014] S1. Obtain data: Obtain multi - angle images of Schisandra chinensis plants covering different growth stages, lighting conditions, and occlusion scenarios, and generate corresponding annotation labels for the images through the Make Sense tool to obtain a Schisandra chinensis plant image dataset;

[0015] S2. Data preprocessing: Preprocess the acquired Schisandra chinensis plant image dataset;

[0016] S3. Data partitioning: Partition the preprocessed Schisandra chinensis plant image dataset into a training set and a test set;

[0017] S4. Model construction: Improve the basic YOLOv7 model to obtain an improved YOLOv7 model;

[0018] S5. Model training: Input the training set into the improved YOLOv7 model, update the model weight parameters through the loss function, and after several trainings, obtain a trained improved YOLOv7 model;

[0019] S6. Detection evaluation: Input the test set into the trained improved YOLOv7 model to evaluate the Schisandra chinensis detection results.

[0020] Optionally, in S2, slice enhancement preprocessing is performed on the acquired Schisandra chinensis plant image dataset, that is, the Schisandra chinensis plant image taken from a long distance is segmented into multiple small images of 640×640 pixels.

[0021] Optionally, the specific content of improving the basic YOLOv7 model to obtain an improved YOLOv7 model in S4 is as follows:

[0022] S41. Replace some of the convolutional layers PConv in the FasterNet module of the basic YOLO7 model with BiPConv partial dimensionality reduction convolutional layers;

[0023] S42. Incorporate the construction method of the ELAN module in the feature extraction structure of the basic YOLO7 model, and propose a BiPNet module to replace the ELAN module;

[0024] S43. Introduce the GSConv and VoV-GSCSP modules in Slim-neck to replace the conventional convolutional layers and the W-ELAN module in the feature fusion structure of the basic YOLO7 model;

[0025] S44. Add an SE attention mechanism module in the feature extraction stage of the basic YOLO7 model.

[0026] Optionally, the specific content of replacing some of the convolutional layers PConv in the FasterNet module of the basic YOLO7 model with BiPConv partial dimensionality reduction convolutional layers in S41 is as follows:

[0027] After some of the convolutional layer PConv operations in the FasterNet module, use a 1x1 convolutional layer to reduce the feature channel dimension, and then perform another PConv convolutional layer to obtain the BiPConv partial dimensionality reduction convolutional layer.

[0028] Optionally, the specific content of replacing the ELAN module in the feature extraction structure of the basic YOLO7 model in S42 with the BiPNet module is as follows:

[0029] The BiPNet module adds a layer of PConv after the BiPConv layer to fully fuse the features of different channels after dimensionality reduction. Replace the ELAN module in the basic YOLO7 model with the BiPNet module and appropriately recycle the BiPNet module.

[0030] Optionally, the specific content of introducing the GSConv and VoV-GSCSP modules in Slim-neck to replace the conventional convolution and W-ELAN modules in the feature fusion structure of the basic YOLO7 model in S43 is as follows:

[0031] S431. GSConv is to solve the speed problem in the convolutional neural network CNN. The channel-dense convolutional calculation SC maximally preserves the implicit connections between each channel, while the channel-sparse convolutional DSC completely cuts off these connections. GSConv is a convolution that combines the channel-dense convolutional calculation SC and the channel-sparse convolutional DSC. The floating-point operation amount is calculated as follows:

[0032]

[0033] S432. The VoV-GSCSP module is an improvement based on GSConv. It uses the lightweight convolution method GSConv to replace the standard convolution SC. Then, on the basis of GSConv, GSbottleneck is continued to be introduced. The floating-point operation amount is calculated as follows:

[0034]

[0035] Where, H out is the height of the output feature map, W out is the width of the output feature map, C in is the number of input channels, C out is the number of output channels, K is the size of the convolution kernel, g is the number of groups of group convolution, n is the number of layers, and i is the number.

[0036] Optionally, the specific content of adding the SE attention mechanism module in the feature extraction stage of the basic YOLO7 model in S44 is as follows:

[0037] The input end of the SE attention mechanism module is connected to the output end of the BiPnet module, so as to perform weighted processing on the number of feature channels after dimensionality reduction and proportional splitting, so as to improve the accuracy of model recognition.

[0038] A Schisandra chinensis target detection system based on a lightweight deep learning network, applying a Schisandra chinensis target detection method based on a lightweight deep learning network according to any one of the above, including: a data acquisition module, a data preprocessing module, a data partitioning module, a model construction module, a model training module, and a detection and evaluation module;

[0039] The data acquisition module, connected to the input end of the data preprocessing module, is used to acquire multi-angle images of Schisandra chinensis plants covering different growth stages, lighting conditions, and occlusion scenarios, and generate corresponding annotation labels for the images through the Make Sense tool to obtain a Schisandra chinensis plant image dataset;

[0040] The data preprocessing module, connected to the input end of the data partitioning module, is used to preprocess the acquired Schisandra chinensis plant image dataset;

[0041] The data partitioning module, connected to the input end of the model construction module, is used to partition the preprocessed Schisandra chinensis plant image dataset into a training set and a test set;

[0042] The model construction module, connected to the input end of the model training module, is used to improve the basic YOLOv7 model to obtain an improved YOLOv7 model;

[0043] The model training module, connected to the input end of the detection and evaluation module, is used to input the training set into the improved YOLOv7 model, update the model weight parameters through the loss function, and after several trainings, obtain a trained improved YOLOv7 model;

[0044] The detection and evaluation module is used to input the test set into the trained improved YOLOv7 model to evaluate the Schisandra chinensis detection results.

[0045] It can be seen from the above technical solutions that, compared with the prior art, the present invention provides a Schisandra chinensis target detection method and system based on a lightweight deep learning network, which has the following beneficial effects:

[0046] (1) The present invention proposes a lighter partial dimensionality reduction convolution (BiPConv), and uses it to construct a BiPNet module to replace the ELAN feature extraction module in the basic YOLOv7 model structure;

[0047] (2) The conventional convolutions and W-ELAN modules in the neck network of the base YOLOv7 model structure are replaced by the Grouped Spatial Convolution (GSConv) and VoVNet-based Grouped Spatial Cross-Stage Partial Network (VoV-GSCSP) modules in Slim-neck;

[0048] (3) The Squeeze-and-Excitation (SE) attention mechanism is used to selectively emphasize useful features;

[0049] (4) The Mean Average Precision (mAP) of the present invention reaches 87.1%, and the Precision (P) is increased to 89.5%. Compared with the base YOLOv7 model, they are increased by 3.08% and 5.15% respectively. The Parameters (params) are reduced to 28182828, and the Floating Point Operations (FLOPs) are reduced to 74.9 FLOPs. Compared with the base YOLOv7 model, they are reduced by 22.75% and 27.42% respectively. The detection ability of the improved YOLOv7 model for Schisandra chinensis in complex backgrounds is significantly improved, providing effective real-time detection support for the intelligent recognition of Schisandra chinensis. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0051] Figure 1 It is a flowchart of a method for detecting Schisandra chinensis targets based on a lightweight deep learning network provided by the present invention;

[0052] Figure 2 It is an example diagram of an image of a Schisandra chinensis plant for basic enhancement provided by the present invention;

[0053] Figure 3 It is an example diagram of an image of a Schisandra chinensis plant for slice enhancement provided by the present invention;

[0054] Figure 4 It is a schematic diagram of the structure of the base YOLOv7 model provided by the present invention;

[0055] Figure 5Schematic diagram of the improved YOLOv7 model structure provided by the present invention;

[0056] Figure 6 Schematic diagram of the PConv convolution structure provided by the present invention;

[0057] Figure 7 Schematic diagram of the dimensionality reduction convolution structure of the BiPConv part in the improved YOLOv7 model provided by the present invention;

[0058] Figure 8 Schematic diagram of the ELAN module structure in the basic YOLOv7 model provided by the present invention;

[0059] Figure 9 Schematic diagram of the BiPNet module structure in the improved YOLOv7 model provided by the present invention;

[0060] Figure 10 Schematic diagram of the GSConv convolution structure in the improved YOLOv7 model provided by the present invention;

[0061] Figure 11 Schematic diagram of the VoV-GSCSP module structure in the improved YOLOv7 model provided by the present invention;

[0062] Figure 12 Schematic diagram of the SE attention mechanism module in the improved YOLOv7 model provided by the present invention. Detailed implementation manners

[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0064] Referring to Figure 1 as shown, the present invention discloses a Schisandra chinensis target detection method based on a lightweight deep learning network, including the following steps:

[0065] S1. Data acquisition: Obtain multi-angle images of Schisandra chinensis plants covering different growth stages, lighting conditions, and occlusion scenarios, and generate corresponding annotation labels for the images through the Make Sense tool to obtain a Schisandra chinensis plant image dataset;

[0066] S2. Data preprocessing: Preprocess the obtained Schisandra chinensis plant image dataset;

[0067] S3. Data Partition: Partition the preprocessed Schisandra chinensis plant image dataset into a training set and a test set;

[0068] S4. Model Construction: Improve the basic YOLOv7 model to obtain the improved YOLOv7 model;

[0069] S5. Model Training: Input the training set into the improved YOLOv7 model, update the model weight parameters through the loss function, and after several trainings, obtain the trained improved YOLOv7 model;

[0070] S6. Detection Evaluation: Input the test set into the trained improved YOLOv7 model to evaluate the Schisandra chinensis detection results.

[0071] Furthermore, as Figure 3 shown, in S2, slice enhancement preprocessing is performed on the obtained Schisandra chinensis plant image dataset, that is, the Schisandra chinensis plant images taken at a long distance are segmented into multiple small images of 640×640 pixels.

[0072] Specifically, as Figure 2 shown, the basic enhancement operations performed on the small images of 640×640 pixels obtained by segmenting the Schisandra chinensis plant images include rotation, flipping, cropping, color transformation (adjusting brightness, contrast, saturation), and adding Gaussian noise or salt-and-pepper noise. One to three enhancement operations will be randomly superimposed on each original image.

[0073] Furthermore, the specific content of improving the basic YOLOv7 model in S4 to obtain the improved YOLOv7 model (as Figure 5 shown) is:

[0074] S41. Replace some of the convolutional layers PConv in the FasterNet module of the basic YOLO7 model with BiPConv partial reduction convolutional layers;

[0075] S42. Incorporate the construction method of the ELAN module in the feature extraction structure of the basic YOLO7 model, and propose the BiPNet module to replace the ELAN module;

[0076] S43. Introduce the GSConv and VoV-GSCSP modules in Slim-neck to replace the conventional convolutional layers and W-ELAN module in the feature fusion structure of the basic YOLO7 model;

[0077] S44. Add the SE attention mechanism module in the feature extraction stage of the basic YOLO7 model.

[0078] Specifically, as Figure 4 shown, the basic YOLOv7 model includes:

[0079] 1. CBS Module (Conv - BN - Silu)

[0080] Function: Basic convolutional operation unit, consisting of a convolutional layer (Conv), batch normalization (BN), and Silu activation function, used for feature extraction and channel adjustment.

[0081] 2. ELAN / E - ELAN Module

[0082] Function: ELAN: Aggregates features at different levels through a multi - branch structure to optimize computational efficiency (such as reducing the number of convolutional kernels and increasing branches).

[0083] E - ELAN (Extended ELAN): An extension based on ELAN, designs a multi - path structure through gradient path analysis to enhance feature fusion ability.

[0084] 3. MP Module (MaxPool + CBS)

[0085] Function: Downsampling module, composed of a max - pooling layer (MaxPool) and a CBS module, used to compress the size of the feature map and retain key information.

[0086] 4. SPP / SPPF Module

[0087] Function: Spatial pyramid pooling, captures multi - scale features through pooling kernels of different scales (such as 5×5, 9×9, 13×13) to enhance the receptive field.

[0088] 5. Reparameterized Convolution (RepConv)

[0089] Function: During training, it adopts a multi - branch convolutional structure (such as 3×3, 1×1 convolutions and Identity connection), and merges into a single convolutional layer during inference.

[0090] 6. Auxiliary Head

[0091] Function: Introduces an additional detection head during the training phase, provides more supervision signals, and helps the main detection head learn more robust features.

[0092] 7. Feature Pyramid Network (FPN / PANet)

[0093] Function: Fuses features at different levels through lateral connections and upsampling (such as the backbone network and the neck structure) to achieve multi - scale object detection.

[0094] 8. Dynamic Label Assignment

[0095] Function: Dynamically allocate positive and negative samples according to the prediction quality during training, and optimize the matching strategy between the optimization target and the anchor box.

[0096] The overall architecture of the basic YOLOv7 model includes:

[0097] Backbone: With E-ELAN as the core, extract multi-level features through CBS and MP modules.

[0098] Neck: Adopt the FPN / PANet structure to fuse multi-scale features.

[0099] Head: The main detection head + auxiliary head (during training) outputs the prediction results, and combines RepConv to optimize the inference efficiency.

[0100] Furthermore, the specific content of replacing part of the convolutional layer PConv in the FasterNet module of the basic YOLO7 model with BiPConv (partial dimensionality reduction convolution) in S41 is as follows:

[0101] After part of the convolutional layer PConv operation in the FasterNet module, use a 1x1 convolution to reduce the feature channel dimension, and then perform another PConv convolution to obtain the BiPConv partial dimensionality reduction convolution.

[0102] Specifically, in the target detection task of Schisandra chinensis plants, since the target morphology is single and the category is fixed, most of the features in the Schisandra chinensis plant images are green leaves, and only a small area is the Schisandra chinensis fruit. Traditional convolutional operations have problems of resource waste and redundancy, especially when dealing with the background area, the computational efficiency is low; based on this situation, Partial Convolution (PConv) is introduced. PConv focuses on the effective area (Schisandra chinensis fruit) of the input features by introducing a binary mask and ignores the invalid area (background noise and occluded parts), effectively reducing computational redundancy, as Figure 6 shown.

[0103] Although using PConv alone can reduce the amount of computation to a certain extent, compared with the expected lightweight target, there is still computational redundancy, especially in images with relatively single target features such as Schisandra chinensis. The reason is that PConv only applies convolutional operations to some channels of the effective area, and the remaining channels remain unchanged, which may lead to insufficient computational efficiency, especially the redundant computation in the background area still exists; in order to further optimize the computational efficiency, BiPConv (partial dimensionality reduction convolution) is proposed as Figure 7As shown in the figure, based on PConv, a 1x1 convolution is introduced to reduce the dimension of the processed feature channels. This operation effectively reduces the computational complexity in subsequent convolution operations. Especially in the control of the number of feature channels, it can minimize redundant calculations to the greatest extent. By comparison, in terms of reducing the number of feature channels in subsequent processing, BiPConv is significantly reduced. Compared with the traditional PConv, BiPConv can significantly improve the computational efficiency while ensuring the target recognition accuracy, meeting the lightweight requirements.

[0104] Furthermore, the specific content of replacing the ELAN module in the feature extraction structure of the basic YOLO7 model in S42 by the BiPNet module is as follows:

[0105] The BiPNet module adds a layer of PConv after the BiPConv layer to fully fuse the features of different channels after dimension reduction. Replace the ELAN module in the basic YOLO7 model with the BiPNet module and appropriately recycle the BiPNet module.

[0106] Specifically, the Efficient Long-range Attention Network (ELAN) of YOLOv7 is the core unit of the backbone network, which realizes the efficient extraction and fusion of multi-scale features through multi-branch convolution operations, as Figure 8 shown. The second branch consists of 1×1 convolution and multiple layers of 3×3 convolution, which is used to extract features of different scales. By using BiPConv as the main convolution to replace multiple layers of 3×3 convolution, the computational complexity is less than that of ELAN. Through experimental analysis, it is found that by replacing 3×3 convolution at different positions with BiPConv, the accuracy and computational complexity shown by the results are different. Finally, the BiPNet module is proposed, as Figure 9 shown. The BiPNet adds a layer of PConv after the BiPConv layer to fully fuse the features of different channels after dimension reduction. In order to improve the overall optimization effect, the more lightweight BiPNet replaces the ELAN module in the initial feature extraction part and appropriately recycles this module.

[0107] Furthermore, the specific content of replacing the conventional convolution and W-ELAN module in the feature fusion structure of the basic YOLO7 model with the GSConv and VoV-GSCSP modules in Slim-neck in S43 is as follows:

[0108] S431. As Figure 10As shown, GSConv is designed to address the speed issue in the Convolutional Neural Network (CNN). The channel-dense convolutional computation SC maximally preserves the implicit connections between each channel, while the channel-sparse convolutional DSC completely cuts off these connections. GSConv is a type of convolution that combines the channel-dense convolutional computation SC and the channel-sparse convolutional DSC. The floating-point operation count is calculated as follows:

[0109]

[0110] S432. As Figure 11 shown, the VoV-GSCSP module is an improvement based on GSConv. It uses the lightweight convolutional method GSConv to replace the standard convolution SC. Then, on the basis of GSConv, GSbottleneck is further introduced. The floating-point operation count is calculated as follows:

[0111]

[0112] where, H out is the height of the output feature map, W out is the width of the output feature map, C in is the number of input channels, C out is the number of output channels, K is the size of the convolutional kernel, g is the number of groups for group convolution, n is the number of layers, and i is the number.

[0113] Furthermore, S44. The specific content of adding the SE attention mechanism module in the feature extraction stage of the basic YOLO7 model is as follows:

[0114] The input end of the SE attention mechanism module is connected to the output end of the BiPnet module, thereby performing weighted processing on the number of feature channels after dimensionality reduction and proportional splitting to improve the accuracy of model recognition.

[0115] Specifically, the Squeeze-and-Excitation (SE) module is an efficient attention mechanism that focuses on mining the global dependencies between channels, thereby significantly enhancing the feature expression ability. As Figure 12 shown, the global semantic information of each channel is extracted through global average pooling, compressing the spatial dimension into a statistical representation of the channel dimension. Subsequently, these global features pass through a two-layer fully connected network, first reducing the dimension and then increasing the dimension, generating channel attention weights through a non-linear activation function. These weights reallocate the feature importance through channel-wise weighting, dynamically enhancing the expression of key features while suppressing redundant information, thereby achieving feature recalibration and optimization. Through this concise and efficient mechanism, the SE module endows the network with stronger feature perception ability and global dependency modeling ability.

[0116] The SE attention mechanism module is placed immediately after the BiPnet module to weight the number of feature channels after dimensionality reduction and proportional segmentation, thereby improving the accuracy of model recognition.

[0117] In a specific embodiment, it includes the following:

[0118] Obtain multi-angle images of Schisandra chinensis plants covering different growth stages, light conditions, and occlusion scenarios. Use the Make Sense tool to generate corresponding annotation labels for the images. After a series of preprocessing operations such as data augmentation on the Schisandra chinensis plant images, divide the preprocessed Schisandra chinensis plant image dataset into a training set and a test set; input the training set into the improved YOLOv7 model. After several trainings, obtain the trained improved YOLOv7 model; input the test set into the trained improved YOLOv7 model. Then, the test set images are sent to the backbone network, and the backbone network part extracts features from the processed images; subsequently, the extracted features undergo feature fusion processing through the Neck module to obtain large, medium, and small-sized features; finally, the fused features are sent to the detection head, and after detection, the results are output.

[0119] The improved YOLOv7 model for Schisandra chinensis consists of four improved modules: BiPConv module, BiPNe module, SE attention mechanism module, and Slim-neck module. These modules are integrated into YOLOv7 to improve its performance on the Schisandra chinensis dataset. Ablation experiments were conducted for each module, and the corresponding results are shown in Table 1.

[0120] Table 1 Ablation experiments of the improved YOLOv7 model

[0121]

[0122] To verify the effectiveness of the proposed model, seven object detection models were selected for comparative experiments on the same dataset, including four common models (YOLOv5, YOLOv8, YOLOv9, and YOLOv10), the latest model (YOLOv11), and the baseline model YOLOv7. The comparison results are shown in Table 2. According to the data in Table 2, the SchisandraConv-YOLOv7 model is superior to the comparison models in terms of precision (P), recall (R), and mean average precision (mAP), indicating that the model proposed in the present invention can effectively extract the features of Schisandra chinensis, thereby improving the detection performance.

[0123] Table 2 Comparison test results of seven object detection models

[0124] Model P R mAP Base Yolov7 0.87 0.815 0.845 Yolov5 0.854 0.762 0.832 Yolov8 0.873 0.75 0.83 Yolov9 0.864 0.812 0.879 Yolov10 0.854 0.733 0.813 Yolov11 0.878 0.76 0.834 Our model (This invention) 0.895 0.857 0.871

[0125] Conclusion: The Mean Average Precision (mAP) of the present invention reaches 87.1%, the Precision (P) is increased to 89.5%. Compared with the basic YOLOv7 model, it is increased by 3.08% and 5.15% respectively. The Parameters (params) are reduced to 28,182,828, and the Floating Point Operations (FLOPs) are reduced to 74.9 FLOPs. Compared with the basic YOLOv7 model, they are reduced by 22.75% and 27.42% respectively. The detection ability of the improved YOLOv7 model for Schisandra chinensis under complex backgrounds is significantly improved, providing effective real-time detection support for the intelligent recognition technology of Schisandra chinensis.

[0126] Compared with Figure 1 Corresponding to the method described above, the embodiment of the present invention also provides a Schisandra chinensis target detection system based on a lightweight deep learning network, including: a data acquisition module, a data preprocessing module, a data division module, a model construction module, a model training module, and a detection and evaluation module;

[0127] The data acquisition module, connected to the input end of the data preprocessing module, is used to acquire multi-angle images of Schisandra chinensis plants covering different growth stages, lighting conditions, and occlusion scenarios, and generate corresponding annotation labels for the images through the Make Sense tool to obtain a Schisandra chinensis plant image dataset;

[0128] The data preprocessing module, connected to the input end of the data division module, is used to preprocess the acquired Schisandra chinensis plant image dataset;

[0129] The data division module, connected to the input end of the model construction module, is used to divide the preprocessed Schisandra chinensis plant image dataset into a training set and a test set;

[0130] The model construction module, connected to the input end of the model training module, is used to improve the basic YOLOv7 model to obtain an improved YOLOv7 model;

[0131] The model training module, connected to the input end of the detection and evaluation module, is used to input the training set into the improved YOLOv7 model, update the model weight parameters through the loss function, and after several trainings, obtain a trained improved YOLOv7 model;

[0132] The detection and evaluation module is used to input the test set into the trained improved YOLOv7 model to evaluate the detection results of Schisandra chinensis.

[0133] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0134] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting Schisandra chinensis targets based on a lightweight deep learning network, characterized in that: The following steps are involved: S1. Obtain data: Obtain multi-angle images of Schisandra chinensis plants under different growth stages, lighting conditions, and occlusion scenes, generate corresponding annotation labels for the images using the Make Sense tool, and obtain a Schisandra chinensis plant image dataset; S2. Data preprocessing: preprocessing the acquired Schisandra chinensis plant image dataset; S3. Data division: the preprocessed Schisandra chinensis plant image dataset is divided into a training set and a test set; S4. Model construction: Improve the basic YOLOv7 model to obtain an improved YOLOv7 model; S5. Model training: Input the training set into the improved YOLOv7 model, update the model weight parameters through the loss function, and obtain the trained improved YOLOv7 model after several trainings; S6. Detection evaluation: The test set is input into the trained improved YOLOv7 model to evaluate the detection results of Schisandra chinensis.

2. The method for detecting Schisandra chinensis target based on a lightweight deep learning network according to claim 1, characterized in that: In S2, the acquired Schisandra chinensis plant image dataset is subjected to slice enhancement preprocessing, that is, the Schisandra chinensis plant image taken at a long distance is segmented into multiple small images of 640×640 pixels.

3. The method for detecting Schisandra chinensis target based on a lightweight deep learning network according to claim 1, characterized in that: In S4, the basic YOLOv7 model is improved, and the specific contents of the improved YOLOv7 model are as follows: S41. Replace the partial convolution PConv in the FasterNet module of the basic YOLO7 model with BiPConv partial dimensionality reduction convolution; S42. A construction method of the ELAN module in the feature extraction structure of the integrated basic YOLO7 model is proposed, and an ELAN module that is replaced by the BiPNet module is proposed; S43. Introduce the GSConv and VoV-GSCSP modules in Slim-neck to replace the conventional convolution and W-ELAN modules in the feature fusion structure of the basic YOLO7 model; S44. Add the SE attention mechanism module to the feature extraction stage of the basic YOLO7 model.

4. The method for detecting Schisandra chinensis target based on a lightweight deep learning network according to claim 3, characterized in that: The specific content of replacing the partial convolution PConv in the FasterNet module of the basic YOLO7 model by BiPConv partial dimensionality reduction convolution in S41 is: After the partial convolution PConv operation in the FasterNet module, a 1x1 convolution is used to reduce the feature channel dimension, and then a PConv convolution is performed to obtain the BiPConv partial dimensionality reduction convolution.

5. The method for detecting Schisandra chinensis target based on a lightweight deep learning network according to claim 3, characterized in that: The construction method of the ELAN module in the feature extraction structure of the fusion basic YOLO7 model in S42 proposes that the BiPNet module replaces the ELAN module as follows: The BiPNet module adds a layer of PConv after the BiPConv layer to fully integrate the features of different channels after dimensionality reduction. The BiPNet module replaces the ELAN module in the basic YOLO7 model and the BiPNet module is appropriately recycled.

6. The method for detecting Schisandra chinensis target based on a lightweight deep learning network according to claim 3, characterized in that: The specific contents of introducing GSConv and VoV-GSCSP modules in Slim-neck in S43 to replace the conventional convolution and W-ELAN modules in the feature fusion structure of the basic YOLO7 model are as follows: S431.GSConv is to solve the speed problem in the convolutional neural network CNN. The channel-intensive convolution calculation SC maximizes the implicit connection between each channel, while the channel-sparse convolution DSC completely cuts off these connections. GSConv is a convolution that combines the channel-intensive convolution calculation SC and the channel-sparse convolution DSC. The floating-point operation amount is calculated as follows: The S432.VoV-GSCSP module is improved on the basis of GSConv. It uses the lightweight convolution method GSConv to replace the standard convolution SC. Then, GSBottleneck is introduced on the basis of GSConv. The floating point operation amount is calculated as follows: Among them, H out is the output feature map height, W out is the output feature map width, C in is the number of input channels, C out is the number of output channels, K is the size of the convolution kernel, g is the number of groups of group convolution, n is the number of layers, and i is the number.

7. The method for detecting Schisandra chinensis target based on a lightweight deep learning network according to claim 3, characterized in that: S44. The specific content of adding the SE attention mechanism module in the feature extraction stage of the basic YOLO7 model is: The input of the SE attention mechanism module is connected to the output of the BiPnet module, so as to perform weighted processing on the number of feature channels after dimensionality reduction and proportional segmentation to improve the accuracy of model recognition.

8. A Schisandra chinensis target detection system based on a lightweight deep learning network, characterized in that: A method for detecting Schisandra chinensis targets based on a lightweight deep learning network according to any one of claims 1 to 7 is applied, comprising: a data acquisition module, a data preprocessing module, a data partitioning module, a model building module, a model training module and a detection evaluation module; The data acquisition module is connected to the input end of the data preprocessing module to obtain multi-angle images of Schisandra chinensis plants under different growth stages, lighting conditions and occlusion scenes, and generate corresponding annotation labels for the images through the Make Sense tool to obtain the Schisandra chinensis plant image dataset; A data preprocessing module, connected to the input end of the data partitioning module, for preprocessing the acquired Schisandra chinensis plant image data set; A data partitioning module, connected to the input end of the model building module, is used to divide the preprocessed Schisandra chinensis plant image dataset into a training set and a test set; A model building module is connected to the input end of the model training module and is used to improve the basic YOLOv7 model to obtain an improved YOLOv7 model; The model training module is connected to the input end of the detection and evaluation module, and is used to input the training set into the improved YOLOv7 model, update the model weight parameters through the loss function, and obtain the trained improved YOLOv7 model after several trainings; The detection and evaluation module is used to input the test set into the trained improved YOLOv7 model to evaluate the detection results of Schisandra chinensis.

Citation Information

Patent Citations

  • Orchard apple detection method based on YOLOv7

    CN118692074A

  • Old people falling detection method based on improved YOLOv8 model

    CN118692139A

  • Method And Apparatus For Image Restoration, Storage Medium And Terminal

    US20210342977A1

Cited By

  • Greenhouse plant growth state recognition method and planting management and control system

    CN121191002A