Citrus detection method and system under complex background based on improved deep learning

By using improved deep learning technology in citrus detection methods, replacing network modules and optimizing detection heads, the existing citrus detection methods are solved in the problem of difficulty in accurately identifying and slow inference speed in complex environments, and efficient and accurate citrus detection is achieved.

CN119942530AActive Publication Date: 2025-05-06CHANGZHOU UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510014213.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing citrus detection methods rely on traditional object detection methods, are inefficient and difficult to accurately identify citrus in complex environments. Moreover, the YOLOv7-tiny algorithm has redundant calculation and parameter quantity, and the inference speed is slow, making it difficult to meet the needs of real-time decision-making in orchards.

Method used

The method of improving deep learning is adopted. By replacing the Conv module in the CBS module in the backbone network as the PConvBS module, and replacing the ELAN module with the Faster-PELAN module in the feature fusion method. The ELAN-H module is used in the neck network for multi-scale feature fusion, and the backbone network and detection head are optimized, using αShape-IoU Loss as the loss function.

Benefits of technology

It improves the detection accuracy and speed of the model, reduces the number of parameters and calculations of the model, is suitable for real-time applications, can effectively detect small and medium-sized citrus targets, and solves the problem of inaccurate positioning of small targets in fruits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942530A_ABST
    Figure CN119942530A_ABST
Patent Text Reader

Abstract

The invention relates to a citrus detection method and system under a complex background based on improved deep learning, and the method comprises the steps: collecting citrus images under the complex background in an orchard, and obtaining an initial data set; preprocessing the initial data set to obtain a first data set; labeling the first data set and performing format conversion to obtain a target data set; taking the YOLOv7-tiny network model as a deep learning reference model to construct an improved YOLOv7-tiny network model as a citrus detection network model; training and evaluating the citrus detection network model; obtaining the optimal weight of the citrus detection network model and testing the citrus detection network model; a test result is obtained through a citrus detection model, and the accuracy rate and the recall rate are calculated based on the IOU so as to judge the detection accuracy. According to the improved YOLOv7-tiny network model, small and medium-sized citrus targets can be effectively detected, the detection precision and speed of the model are improved, and the light weight of the model is considered at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method and system for detecting citrus in a complex background based on improved deep learning. Background Art

[0002] In the field of fruit recognition, citrus detection in the past mostly relied on traditional target detection methods. Traditional methods mainly rely on manual features to implement target detection, which is inefficient, time-consuming and labor-intensive. In complex environments, when faced with problems such as occlusion by branches and leaves and changes in light, it is even more helpless, making it difficult to accurately identify citrus, which greatly hinders the intelligent development of the citrus industry.

[0003] In the field of deep learning, the YOLOv7-tiny algorithm provides an innovative solution for citrus recognition, especially in complex backgrounds, such as orchard environments with different sparsity, changing lighting conditions, and different stages of fruit maturity. With its powerful feature extraction capabilities and efficient detection speed, YOLOv7-tiny can flexibly respond to changes in various environments and adapt to the detection needs of small and medium-sized targets while ensuring high detection accuracy. However, the current YOLOv7-tiny algorithm still has problems with redundant calculations and parameters, and its inference speed is slow, making it difficult to meet the needs of real-time decision-making in orchards.

[0004] Therefore, providing a new citrus detection method and system in complex backgrounds based on improved deep learning can improve the application efficiency of deep learning in the field of agricultural intelligence, promote the realization of real-time and accurate orchard management and decision-making systems, and thus bring a higher level of automation and more accurate crop management to agricultural production. Summary of the invention

[0005] Based on the above-mentioned problems existing in the prior art, the purpose of the embodiments of the present invention is to provide a citrus detection method and system in a complex background based on improved deep learning, which can effectively detect small and medium-sized citrus targets, improve the detection accuracy and speed of the model while taking into account the lightweight model.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is: a citrus detection method under complex background based on improved deep learning, comprising:

[0007] S1, collect citrus images with complex backgrounds in orchards to obtain the initial data set;

[0008] S2, preprocessing the initial data set to obtain a first data set;

[0009] S3, annotating the first data set and performing format conversion to obtain a target data set;

[0010] S4, using the YOLOv7-tiny network model as the deep learning benchmark model to build an improved YOLOv7-tiny network model as a citrus detection network model; wherein the improved YOLOv7-tiny network model includes replacing the Conv modules in all CBS modules in the backbone network with PConvBS modules; replacing the ELAN modules in the backbone network with Faster-PELAN modules by feature fusion; using the ELAN-H module in the neck network for multi-scale feature fusion; optimizing the backbone network and the detection head, including reducing the number of downsampling of the backbone network and removing the corresponding large target detection head; the detection head uses αShape-IoU Loss as the loss function;

[0011] S5, training and evaluating the citrus detection network model;

[0012] S6, obtaining the optimal weight of the citrus detection network model and testing the citrus detection network model;

[0013] S7, obtain the test results through the citrus detection model and calculate the precision and recall based on the intersection-over-union ratio (IOU) to determine the accuracy of the detection.

[0014] Furthermore, in S2, the preprocessing of the initial data set includes screening, data enhancement, and resolution formatting of the initial data set.

[0015] Furthermore, in S2, the specific steps of preprocessing the initial data set to obtain the first data set include:

[0016] Step S21, screening the initial data set, including deleting blurred images caused by shooting, light, equipment stability and other factors, and deleting citrus image data lacking key features;

[0017] Step S22, data enhancement includes color space transformation and direction inversion of the image;

[0018] Step S23, resolution formatting includes performing resolution formatting processing on the initial data set, and selecting a suitable standard resolution according to the requirements of model training and the actual situation of hardware resources.

[0019] Furthermore, in S3, the specific content of labeling the first data set includes labeling categories, labeling center points, and labeling width and height information, and format conversion of the first data set includes reorganizing and extracting the labeling information in the xml file to obtain a label file in txt format as the target data set.

[0020] Furthermore, the calculation formulas for the number of floating-point operations and memory access of the PConvBS module are as follows:

[0021]

[0022] Among them, h is the height of the feature map, w is the width of the feature map, k is the size of the convolution kernel, and c is the p is the number of output channels, where c p / c=1 / 4.

[0023] Furthermore, the Faster-PELAN module is composed of a PConvBS module, a DWConvBS module, a Concat module and a CBS module; the Faster-PELAN module has two paths, the first path will pass through the PConvBS module and the DWConvBS module for fast sampling, and the second path will be sampled through the PConvBS module; the feature map sampled by the specific layer on the first path will be spliced ​​with the feature map sampled by the second path, and finally the feature aggregation process is completed through a convolution layer; the DWConvBS module is composed of a DWConvBS module, a BatchNormalization layer and a SiLU activation function in sequence.

[0024] Furthermore, the calculation formula of the loss function αShape-IoU Loss is:

[0025] L αShape-IoU =α(1-IoU)+(1-α)distance shape +0.5×Ω shape

[0026] Among them, L αShape-IoU is the loss function, α is the weight parameter, IoU is the intersection over union ratio, and distance shape is the shape distance, Ω shape is the shape complexity related item;

[0027] The calculation formula for the intersection over union (IoU) is:

[0028]

[0029] Among them, IoU is the intersection over union ratio, B is the predicted bounding box, and B gt is the true bounding box;

[0030] IoU is used to measure the degree of overlap between the predicted box and the true box. The value of IoU is between 0 and 1. The larger the value, the higher the overlap between the predicted box and the true box.

[0031] Shape distance shapeThe calculation formula is:

[0032]

[0033] Among them, distance shape is the shape distance, x c and c is the coordinate of the center of the prediction box, and is the coordinate of the center of the real box, c is a constant, HH and WW are coefficients related to the width and height of the box, and w gt is the width of the real frame, h gt is the height of the real frame, (w gt ) scale is the new width value obtained by scaling the width of the real frame, (h gt ) scale The new height value after scaling the height of the real frame;

[0034] Shape complexity related term Ω shape The calculation formula is:

[0035]

[0036] Among them, Ω shape is the shape complexity related term, w t is the shape of the box and θ is a fixed parameter.

[0037] The citrus detection system under complex background based on improved deep learning is applied to the above-mentioned citrus detection method under complex background based on improved deep learning, and the system comprises:

[0038] The data set collection module is used to collect citrus images with complex backgrounds in the orchard to obtain the initial data set;

[0039] A data set preprocessing module, used for preprocessing the initial data set to obtain a first data set;

[0040] A data set format conversion module, used to annotate the first data set and perform format conversion to obtain a target data set;

[0041] An improved modeling module is used to construct an improved YOLOv7-tiny network model as a citrus detection network model using the YOLOv7-tiny network model as a deep learning benchmark model; wherein the improved YOLOv7-tiny network model includes replacing the Conv modules in all CBS modules in the backbone network with PConvBS modules; replacing the ELAN modules in the backbone network with Faster-PELAN modules by feature fusion; using the ELAN-H module in the neck network for multi-scale feature fusion; optimizing the backbone network and the detection head, including reducing the number of downsampling of the backbone network and removing the corresponding large target detection head; and using αShape-IoU Loss as the loss function of the detection head;

[0042] Model training and evaluation module, used to train and evaluate the citrus detection network model;

[0043] A model testing module is used to obtain the optimal weight of the citrus detection network model and test the citrus detection network model;

[0044] The detection result generation module is used to obtain the test results through the citrus detection model and calculate the precision and recall rate based on the intersection-over-union ratio (IOU) to determine the accuracy of the detection.

[0045] The embodiment of the present invention further provides a network side server, including:

[0046] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned citrus detection method in complex background based on improved deep learning.

[0047] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned citrus detection method in a complex background based on improved deep learning.

[0048] The beneficial effects of the present invention are as follows: the citrus detection method under complex background based on improved deep learning of the present invention comprises collecting citrus images under complex background in an orchard to obtain an initial data set; preprocessing the initial data set to obtain a first data set; annotating the first data set and performing format conversion to obtain a target data set; using the YOLOv7-tiny network model as a deep learning benchmark model to construct an improved YOLOv7-tiny network model as a citrus detection network model; wherein the improved YOLOv7-tiny network model comprises replacing the Conv modules in all CBS modules in the backbone network with PConvBS modules; replacing the ELAN modules in the backbone network with Faster-PELAN modules by feature fusion; using the ELAN-H module in the neck network for multi-scale feature fusion; optimizing the backbone network and the detection head, comprising reducing the number of downsampling of the backbone network and removing the corresponding large target detection head; the detection head adopts αShape-IoU Loss is used as the loss function; the citrus detection network model is trained and evaluated; the optimal weight of the citrus detection network model is obtained and the citrus detection network model is tested; the test results are obtained through the citrus detection model and the precision and recall are calculated based on the intersection-over-union (IOU) ratio to determine the accuracy of the detection. The citrus detection method under complex background based on improved deep learning of the present invention reduces redundant feature extraction and model parameter amount, optimizes memory usage and calculation amount, improves feature extraction efficiency, replaces ELAN module with Faster-PELAN module, enhances multi-scale feature aggregation capability, and realizes efficient feature sampling and aggregation; adopts ELAN-H module for multi-scale feature fusion in neck network, reduces calculation amount and parameter amount, improves reasoning speed, and is suitable for real-time application; optimizes backbone network and detection head, reduces downsampling times and removes large target detection head, reduces parameter amount, improves reasoning speed, and can effectively detect small and medium-sized citrus targets; adopts αShape-IoU Loss as loss function in detection head, solves the problem of inaccurate positioning of small fruit targets; reduces memory usage and calculation amount by converting format of model weight file; optimizes memory access, and adopts cache and data pre-fetching and other technologies to improve data processing speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0050] In the figure:

[0051] Figure 1 A flow chart of a citrus detection method under complex background based on improved deep learning provided in Example 1 of the present invention;

[0052] Figure 2A schematic diagram of the structure of the improved YOLOv7-tiny provided in the first embodiment of the present invention;

[0053] Figure 3 A schematic diagram of the structure of the PConvBS module provided in the first embodiment of the present invention;

[0054] Figure 4 A schematic diagram of the structure of a DWConvBS module provided in the first embodiment of the present invention;

[0055] Figure 5 A schematic diagram of the structure of a Faster-PELAN module provided in Embodiment 1 of the present invention;

[0056] Figure 6 A schematic diagram of the structure of an ELAN-H module provided in Embodiment 1 of the present invention;

[0057] Figure 7 A comparison diagram of algorithm detection effects under different light intensities provided in Example 1 of the present invention;

[0058] Figure 8 A comparison diagram of algorithm detection effects under different sparsities provided in Example 1 of the present invention;

[0059] Fig. 9 A comparison diagram of algorithm detection effects in an immature state provided in Example 1 of the present invention;

[0060] Fig.10 A schematic diagram of a module of a citrus detection system in a complex background based on improved deep learning provided in the second embodiment of the present invention;

[0061] Fig.11 It is a structural diagram of a network-side server provided according to a third embodiment of the present invention. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0063] First embodiment:

[0064] The first embodiment of the present invention provides a citrus detection method under a complex background based on improved deep learning, including: collecting citrus images in an orchard to obtain an initial data set; preprocessing the initial data set to obtain a first data set; annotating the first data set and performing format conversion to obtain a target data set; using the YOLOv7-tiny network model as a deep learning benchmark model to construct an improved YOLOv7-tiny network model as a citrus detection network model; wherein the improved YOLOv7-tiny network model includes replacing the Conv module in all CBS modules in the backbone network with a PConvBS module; replacing the ELAN module in the backbone network with a Faster-PELAN module by a feature fusion method; using the ELAN-H module in the neck network for multi-scale feature fusion; optimizing the backbone network and the detection head, including reducing the number of downsampling of the backbone network and removing the corresponding large target detection head; the detection head adopts αShape-IoU Loss is used as the loss function; the citrus detection network model is trained and evaluated; the optimal weight of the citrus detection network model is obtained and the citrus detection network model is tested; the test results are obtained through the citrus detection model and the precision and recall are calculated based on the intersection-over-union ratio (IOU) to judge the accuracy of the detection. The citrus detection method under complex background based on improved deep learning of the present invention can effectively detect small and medium-sized citrus targets, improve the detection accuracy and speed of the model while taking into account the lightweight model.

[0065] The following is a detailed description of the implementation details of the citrus detection method in a complex background based on improved deep learning in this embodiment. The following content is only for the convenience of understanding the implementation details, which is not necessary for the implementation of this solution. The specific process of this embodiment is as follows Figure 1 As shown, this embodiment is applied to a citrus detection system in a complex background based on improved deep learning.

[0066] S1, collect citrus images with complex background in the orchard to obtain the initial data set.

[0067] Specifically, professional camera equipment is used to photograph citrus in the orchard from different angles, distances and lighting conditions, aiming to obtain as many images as possible covering various forms and growth stages of citrus. In addition, the photography will cover different areas of the orchard, including places with different terrain heights, locations near the edge of the orchard and in the center of the orchard, etc., in order to accumulate a large amount of rich and diverse image data, ensuring that these data can truly reflect the overall picture of citrus in the actual orchard environment.

[0068] There are many complex backgrounds in the orchard. For example, some citrus grows in places with lush branches and leaves and many obstructions, and the background presents a complex situation of dense leaves interweaving; some citrus may be surrounded by weeds, forming a complex background containing weeds; and some citrus may be next to the orchard aisle, and its background includes elements such as irrigation facilities and fruit farmers' tools. Based on these different complex background situations, the collected citrus image data is classified and sorted, and images with similar complex background features are classified into the same category, and the initial data set under different complex backgrounds is constructed.

[0069] Step S2, preprocessing the initial data set to obtain a first data set.

[0070] Specifically, preprocessing the initial dataset includes screening, data enhancement, and resolution formatting of the initial dataset.

[0071] Step S21, screening the initial data set, including deleting blurred images caused by various factors such as shooting, light, equipment stability, etc., and deleting citrus image data lacking key features.

[0072] Step S22, data enhancement includes color space transformation and direction reversal of the image. By transforming the color space of the image, the color characteristics of the citrus can be highlighted. For example, in the HSV space, the orange area of ​​the ripe citrus and the green area of ​​the background leaves can be more easily separated by threshold segmentation, which facilitates fruit recognition and positioning. In addition, data in different color spaces can be used as supplementary information to allow the model to learn the characteristic performance of citrus under a variety of color description systems, enhance its adaptability to color changes, and improve robustness. The direction reversal of the image includes horizontal reversal and vertical reversal. By reversing the direction of the image, the data diversity can be further enriched, allowing the model to see more posture changes of citrus and improve its generalization ability.

[0073] Step S23, resolution formatting includes performing resolution formatting processing on the initial data set, and selecting a suitable standard resolution, such as the common 224×224 pixels or 512×512 pixels, etc., according to the requirements of model training and the actual situation of hardware resources. Through the image scaling algorithm, the original images with different resolutions are uniformly adjusted to the set standard resolution to ensure that the images in the data set are consistent in size, which is convenient for storage management, optimizes the model training process, improves overall efficiency, and ensures image details within a reasonable range to meet the model's requirements for citrus feature extraction.

[0074] Step S3: annotate the first data set and perform format conversion to obtain a target data set.

[0075] Specifically, the specific contents of labeling the first data set include labeling categories, labeling center points, and labeling width and height information. Citrus has growth stages and states, etc. When labeling categories, it is necessary to determine the category to which each citrus target belongs according to pre-set classification standards. For example, it can be divided into categories such as mandarin oranges, navel oranges, and ponkan according to common citrus varieties; or it can be divided into categories such as immature, semi-mature, and mature according to the maturity of citrus.

[0076] In order to accurately determine the position of the citrus in the image, the coordinates of the center point of each citrus target need to be marked. Usually, the upper left corner of the image is the coordinate origin (0,0), the horizontal right is the positive direction of the x-axis, and the vertical downward is the positive direction of the y-axis. Through professional annotation tools, the annotator will accurately find the coordinate value (x, y) corresponding to the approximate geometric center position of the citrus fruit. As an example, the coordinates of the center point of a citrus in an image may be (120, 80). This coordinate information is very important for the subsequent calculation of the relative position of the citrus and other objects, and the judgment of its layout in the image. It can help the model accurately lock the area where the citrus is located.

[0077] The width and height of the citrus target are also marked in pixels. The maximum horizontal length of the citrus fruit is measured along the x-axis as the width, and the maximum vertical length is measured along the y-axis as the height. For example, the width of a citrus fruit is marked as 50 pixels, and the height is marked as 40 pixels. These size information combined with the center point coordinates can fully outline the specific range of the citrus in the image, so that the model can clearly obtain which pixel areas in the image belong to the citrus target to be concerned, so as to perform accurate feature extraction and analysis.

[0078] The format conversion of the first data set includes reorganizing and extracting the annotation information in the xml file to obtain a label file in txt format as the target data set. For each citrus target, its category information is converted into a corresponding digital number according to a preset category index table (for example, tangerine corresponds to number 0, navel orange corresponds to number 1, etc.), and then the center point coordinates, width and height information are converted into numerical values ​​according to a certain proportional relationship, and the converted information is written into the txt file in a specific order, each line represents the annotation information of a citrus target, and different numerical values ​​are separated by symbols such as spaces, so that a label file in txt format is generated as the target data set.

[0079] Step S4, using the YOLOv7-tiny network model as a deep learning benchmark model to construct an improved YOLOv7-tiny network model as a citrus detection network model.

[0080] Specifically, the existing YOLOv7-tiny network model is mainly composed of a backbone network, a neck network and a detection head. Among them, the backbone network, as a feature extractor, extracts representative and distinguishing features of different levels from the input image through a series of operations such as convolution and pooling; the neck network acts as a bridge between the backbone network and the detection head. The neck network receives the feature maps of different levels extracted by the backbone network, and then integrates these features through feature fusion operations such as splicing and addition. At the same time, it will further mine and utilize contextual information, so that the features finally passed to the detection head are more comprehensive and targeted, which helps to improve the accuracy and robustness of detection; the detection head, as the last part of the network, uses the features integrated by the neck network to generate the output of the network, which is to predict the category and position of the objects in the image, and output information such as the probability of each object belonging to different categories and the coordinates of the bounding box of the object in the image, so as to realize the detection of objects in the image.

[0081] The detailed structure of the network model improved based on YOLOv7-tiny is as follows Figure 2 As shown, the specific steps are as follows:

[0082] Step S41, replacing the Conv modules in all CBS modules in the backbone network with PConvBS modules.

[0083] Specifically, Figure 3 As shown in the figure, the Conv modules in all CBS modules in the backbone network are replaced with PConvBS modules that only perform convolution on some channels to reduce the extraction of redundant features, reduce the number of model parameters, and improve feature extraction efficiency.

[0084] The PConvBS module splits the input feature map into multiple sub-feature maps and assigns an independent convolution kernel to each sub-feature map for processing, thereby performing convolution operations on only some channels. Unlike traditional convolution methods, the PConvBS module only extracts spatial features from some channels of the input feature map while keeping other channels unchanged, thereby effectively reducing the amount of calculation and the extraction of redundant features. For continuous or regular memory access, the PConvBS module optimizes memory usage and reduces unnecessary calculations by calculating the first or last continuous channel as a representative of the entire feature map. Therefore, the PConvBS module can significantly reduce the computational cost when extracting spatial features while improving the efficiency of feature extraction. Among them, the calculation formulas for the number of floating-point operations and memory access of the PConvBS module are as follows:

[0085]

[0086] Among them, h is the height of the feature map, w is the width of the feature map, k is the size of the convolution kernel, and c is thep is the number of output channels. p / c=1 / 4.

[0087] As an example, for the sake of memory continuity, the number of continuous output channels of the first or last segment is selected to represent the entire feature map. Assuming that the number of input and output channels is the same, the number of floating-point operations of the PConvBS module is (i.e., convolution kernel and input map Figure 1 The computational cost of a window is ), if c p / c=1 / 4, then the number of floating-point operations of the PConvBS module is only 1 / 16 of that of ordinary convolution, and the memory access amount of the PConvBS module is If c p / c=1 / 4, then the memory access amount will be 1 / 4 of the ordinary convolution.

[0088] Furthermore, the PConvBS module is composed of a PConvBS module, a BatchNormalization layer, and a SiLU activation function in sequence.

[0089] Step S42: replace the ELAN module in the backbone network with the Faster-PELAN module by adopting feature fusion method.

[0090] Specifically, Figure 5 As shown in the figure, the original ELAN module is replaced by the feature fusion module Faster-PELAN, and the Faster-PELAN module is embedded in the backbone feature extraction network of the model, thereby enhancing the network's feature aggregation capabilities at multiple scales to better focus on the fruit target. The Faster-PELAN module consists of multiple functional modules, and achieves more efficient feature sampling and aggregation through feature fusion of two different paths.

[0091] The Faster-PELAN module consists of the PConvBS module, the DWConvBS module, the Concat module, and the CBS module. The Faster-PELAN module has two paths. The first path will pass through the PConvBS module and the DWConvBS module for fast sampling, and the second path will pass through the PConvBS module for sampling. The feature map sampled by the specific layer on the first path will be spliced ​​with the feature map sampled by the second path, and finally a convolution layer will be used to complete the feature aggregation process.

[0092] Further, such as Figure 4As shown, the DWConvBS module is composed of a DWConvBS module, a BatchNormalization layer and a SiLU activation function in sequence.

[0093] Step S43, using the ELAN-H module in the neck network to perform multi-scale feature fusion.

[0094] Specifically, Figure 6 As shown in the figure, the ELAN-H module in the neck network reduces the amount of calculation and parameters, reduces memory consumption, and improves the reasoning speed by streamlining the feature aggregation method and adopting a more efficient convolution strategy (such as depthwise separable convolution). As a result, the improved YOLOv7-tiny network model maintains a high detection accuracy while being more suitable for real-time applications, especially in embedded devices and resource-limited environments. The ELAN-H module can achieve a better balance between efficiency and accuracy and improve target detection performance.

[0095] Step S44, optimizing the backbone network and the detection head, including reducing the number of downsampling times of the backbone network and removing the corresponding large target detection head.

[0096] Specifically, the number of downsampling times of the backbone network was reduced from five to four times, the last layer of MP module and ELAN module in the backbone network were pruned and the corresponding large target detection heads were removed, which reduced the overall parameter amount of the network model, improved the model inference speed and enabled effective detection of citrus as small and medium-sized targets.

[0097] In step S45, the detection head uses αShape-IoU Loss as the loss function.

[0098] Specifically, in the detection and recognition of citrus fruits, the existing detection heads have poor detection performance due to factors such as target overlap, occlusion, and variable object shape. The detection head uses αShape-IoU Loss as the loss function to solve the problem of inaccurate positioning of small targets in fruit targets.

[0099] The calculation formula of the loss function αShape-IoU Loss is:

[0100] L αShape-IoU =α(1-IoU)+(1-α)distance shape +0.5×Ω shape

[0101] Among them, L αShape-IoU is the loss function, α is the weight parameter, IoU is the intersection over union ratio, and distance shape is the shape distance, Ω shape is a term related to shape complexity.

[0102] Furthermore, the calculation formula of the intersection over union (IoU) is:

[0103]

[0104] Among them, IoU is the intersection over union ratio, B is the predicted bounding box, and B gt is the true bounding box.

[0105] IoU is used to measure the degree of overlap between the predicted box and the true box. The value of IoU is between 0 and 1. The larger the value, the higher the overlap between the predicted box and the true box.

[0106] Furthermore, the shape distance shape The calculation formula is:

[0107]

[0108] Among them, distance shape is the shape distance, x c and c is the coordinate of the center of the prediction box, and is the coordinate of the center of the real box, c is a constant, HH and WW are coefficients related to the width and height of the box, and w gt is the width of the real frame, h gt is the height of the real frame, (w gt ) scale is the new width value obtained by scaling the width of the real frame, (h gt ) scale The new height value after scaling the height of the real frame.

[0109] Furthermore, the shape complexity related term Ω shape The calculation formula is:

[0110]

[0111] Among them, Ω shape is the shape complexity related term, w t is the shape of the box and θ is a fixed parameter.

[0112] Step S5, training and evaluating the citrus detection network model.

[0113] Specifically, the target data set is divided into a training set, a validation set, and a test set. The training of the citrus detection network model includes: using the training set to train the improved YOLOv7-tiny citrus detection network model. During the training process, appropriate training parameters are set, such as learning rate, number of iterations, batch size, etc. The initial learning rate can be set to a smaller value, and a learning rate decay strategy is adopted to gradually reduce the learning rate as the training progresses to prevent the model from missing the optimal solution due to excessive learning rate in the early stage of training, and to continuously optimize the model parameters in the later stage of training. The batch size can be set to 16 or 32 according to hardware resources and data characteristics to ensure that the model can fully learn the characteristics of different samples in one iteration, and at the same time, the training process will not be affected by problems such as memory overflow. The number of iterations can be adjusted according to the performance of the model on the validation set, and the training is stopped when the loss function value on the validation set no longer decreases significantly or there are signs of overfitting.

[0114] Furthermore, optimizers such as stochastic gradient descent or adaptive moment estimation are used to update model parameters. In each iteration, the training sample is input into the model, and the model calculates the prediction result through forward propagation. Then, according to the difference between the prediction result and the true annotation, the loss value is calculated through the loss function αShape-IoU Loss. Then, the optimizer calculates the gradient according to the loss value through backpropagation, and updates the weight and bias parameters of the model, so that the model gradually reduces the loss value during continuous training and improves the detection accuracy.

[0115] The evaluation of the citrus detection network model includes: using the validation set to evaluate the model regularly during the training process. The main evaluation indicators include mean average precision (mAP), recall rate, precision rate, etc. Mean average precision (mAP) is an important indicator to measure the performance of the target detection model. The mean average precision comprehensively considers the detection accuracy and recall rate of different categories of citrus, and can fully reflect the detection performance of the model under various thresholds. By calculating the intersection over union (IoU) of the predicted box and the true box, and determining the correctness of the prediction according to the set IoU threshold (, the area under the precision and recall curves of different categories of citrus is then counted to obtain the mAP value. The recall rate reflects the proportion of citrus targets that the model can correctly detect to the actual citrus targets, and the precision rate indicates the proportion of targets predicted by the model as citrus that are actually citrus. By analyzing these indicators, we can understand the performance of the model in different aspects, find problems with the model in a timely manner, and adjust the training process.

[0116] Step S6, obtaining the optimal weight of the citrus detection network model and testing the citrus detection network model.

[0117] Specifically, the best performing model weights are determined during the training process in step S5. Since the previous step involves evaluating and comparing the model weights at different training stages, a comprehensive judgment is made based on indicators such as the loss function value, precision, and recall on the validation set. Once the optimal weights are determined, the trained model is comprehensively tested using the weights. The test set should cover image samples of various complex backgrounds, different lighting conditions, different citrus varieties, and growth stages to ensure that the generalization ability of the model in actual application scenarios is fully tested. During the test, the model's prediction results for each sample are recorded in detail, including the predicted citrus category, bounding box position, and corresponding confidence score, to provide sufficient data support for subsequent result analysis.

[0118] Step S7, obtaining the test results through the citrus detection model and calculating the precision and recall based on the intersection-over-union (IOU) ratio to determine the accuracy of the detection.

[0119] Specifically, for the test results, according to the standard evaluation process of target detection, the accuracy of detection is judged based on the intersection over union (IOU) of the predicted box and the label bounding box. For each predicted result, its IOU value with the true label is accurately calculated. When the IOU is greater than the pre-set threshold and the category prediction is correct, it is marked as a true positive. p ; If the intersection-over-union ratio (IOU) is greater than the threshold but the category is incorrect, it is considered a false positive example T p ; When there is no detection box on the target, it is defined as missed detection F n Among them, the calculation formulas for precision and recall are:

[0120]

[0121] Among them, P is the precision, T p For a true example, F p is a false positive example, R is the recall rate, F n For missed detection.

[0122] During the calculation process, the classification of each sample is carefully checked to ensure that the precision and recall are calculated accurately, so that the performance of the model in citrus detection under complex backgrounds can be objectively evaluated.

[0123] As an example, for the trained citrus detection network model, its weight file is converted to adapt to the operating environment of the embedded development board. During the conversion process, the accuracy conversion of the network layer weights is focused on. According to the computing power and memory resources of the embedded device, the appropriate accuracy representation method is selected, such as converting 32-bit floating point numbers to 16-bit or 8-bit fixed point numbers, so as to reduce memory usage and calculation without significantly affecting the model performance. At the same time, the memory access is optimized, and the storage structure and data reading method of the citrus detection network model are adjusted to ensure that the model weights and intermediate data can be accessed efficiently under the limited memory bandwidth of the embedded device. For example, cache strategies, data prefetching and other technologies are used to reduce memory waiting time and improve data processing speed. In addition, the multi-core processor architecture of the embedded device is fully utilized to perform parallel computing on the calculation process of the model. By reasonably dividing the computing tasks, the parallel operations are assigned to different cores for simultaneous execution, which further improves the reasoning speed of the citrus detection network model, so that the citrus detection network model can quickly respond to the user's detection request on the embedded device.

[0124] After starting the detection system on the embedded device, the camera starts to collect image data in the area being identified. The camera should have a suitable resolution and frame rate to meet the needs of real-time detection, while ensuring that the image quality can clearly present the characteristics of citrus. The collected image data is promptly transmitted to the processing unit of the embedded end through a pre-set communication interface. On the embedded end, after receiving the image data, the optimized and converted citrus detection network model is immediately called for forward reasoning. Based on the feature information of the input image, the citrus detection network model undergoes a series of convolution, pooling, activation and other operations to gradually generate prediction results for citrus, including the category, location and confidence information of the citrus.

[0125] Finally, the recognition results are processed and displayed. For images in the video stream, the detection results can be superimposed into the video stream in the form of annotation boxes and text information, and intuitively displayed to the user. At the same time, the detection results can be saved in the local storage device at any time, and classified and stored according to time, location, detection batch and other information, so that users can find and obtain historical detection data at any time. In addition, in order to realize the remote management and sharing of data, the saved results can be uploaded to the cloud server through the network communication module at the same time, and presented in the form of visual charts, reports, etc. on the client terminal, so that relevant personnel can remotely monitor the detection of citrus and provide data support for citrus planting management, picking decisions, etc.

[0126] According to the experimental results in Table 1 below, the average detection accuracy of the citrus detection network model is improved by 1.1% compared with the baseline model YOLOv7-Tiny, the total number of parameters is reduced by 21.2%, the storage space of the model is compressed to 8.4MB, and the FPS is also improved to 64.7.

[0127] Table 1 Comparison between improved model and baseline model

[0128]

[0129] It can be seen from Table 1 that the improved scheme of the present invention not only reduces the weight of the model but also improves the accuracy and detection speed of the model. Figure 7 This is a comparison chart of the detection effects of the algorithms of the present invention under different light intensities. The upper part of the figure is the algorithm effect chart of the deep learning benchmark model YOLOv7-Tiny, and the lower part is the algorithm effect chart based on the improved YOLOv7-Tiny citrus detection model. Compared with the benchmark model, the improved YOLOv7-Tiny citrus detection model has a better detection effect and can detect some targets that are missed due to the small size of the fruit. When detecting the same fruit at the same time, the confidence of the citrus detection model is generally better than that of the benchmark model. Figure 8 It can be seen that the number of repeated detections in the citrus detection method under complex background based on improved deep learning proposed by the present invention is far less than that of the original algorithm; Fig. 9 It can be seen that the citrus detection method under complex background based on improved deep learning proposed in the present invention also has good detection results on immature citrus fruits. In summary, the citrus detection method under complex background based on improved deep learning proposed in the present invention can achieve a good balance between reducing the number of model parameters and improving the model's expressiveness, which is conducive to the deployment of the model on edge devices.

[0130] The first embodiment of the present invention provides a citrus detection method under a complex background based on improved deep learning, comprising: collecting citrus images under a complex background in an orchard to obtain an initial data set; preprocessing the initial data set to obtain a first data set; annotating the first data set and performing format conversion to obtain a target data set; using the YOLOv7-tiny network model as a deep learning benchmark model to construct an improved YOLOv7-tiny network model as a citrus detection network model; training and evaluating the citrus detection network model; obtaining the optimal weight of the citrus detection network model and testing the citrus detection network model; obtaining test results through the citrus detection model and calculating the precision and recall based on the intersection-over-union (IOU) ratio to determine the accuracy of the detection. The citrus detection model of the citrus detection method under complex background based on improved deep learning of the present invention replaces the Conv module with the PConvBS module in the backbone network, reduces redundant feature extraction and model parameter amount, optimizes memory usage and calculation amount, improves feature extraction efficiency, replaces the ELAN module with the Faster-PELAN module, enhances multi-scale feature aggregation capability, and realizes efficient feature sampling and aggregation; the neck network adopts the ELAN-H module to perform multi-scale feature fusion, reduces the amount of calculation and the amount of parameters, improves the reasoning speed, and is suitable for real-time applications; the backbone network and the detection head are optimized, the number of downsampling times is reduced, and the large target detection head is removed, the parameter amount is reduced, the reasoning speed is improved, and small and medium-sized citrus targets can be effectively detected; the detection head adopts αShape-IoU Loss is used as a loss function to solve the problem of inaccurate positioning of small fruit targets; by converting the format of the model weight file, memory usage and calculation amount are reduced; memory access is optimized, and technologies such as caching and data prefetching are used to improve data processing speed; the camera collects images and transmits them to the embedded end, calling the optimization model for forward reasoning to generate citrus prediction results; the detection results are presented to users in an intuitive form, which can be saved locally and uploaded to the cloud server, and presented in a visual form to facilitate remote monitoring and provide data support for citrus management and picking decisions. Compared with the baseline model, the improved citrus detection network model has improved average detection accuracy, reduced total parameters, compressed model storage space, and improved FPS. It improves accuracy and detection speed on the basis of lightweight model, has good detection effect in complex environments, can strike a balance between reducing parameters and improving expressiveness, and is conducive to deployment on edge devices.

[0131] The step division of the above methods is only for the purpose of clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this patent.

[0132] Second implementation method:

[0133] like Fig.10 As shown, the second embodiment of the present invention provides a citrus detection system in a complex background based on improved deep learning, and the system includes: a data set acquisition module 201, a data set acquisition module 201, a data set format conversion module 203, an improved modeling module 204, a model training evaluation module 205, a model testing module 206, and a detection result generation module 207.

[0134] Specifically, a data set acquisition module 201 is used to collect citrus images with complex backgrounds in an orchard to obtain an initial data set; a data set preprocessing module 202 is used to preprocess the initial data set to obtain a first data set; a data set format conversion module 203 is used to annotate and convert the first data set to obtain a target data set; an improved modeling module 204 is used to construct an improved YOLOv7-tiny network model as a citrus detection network model using the YOLOv7-tiny network model as a deep learning benchmark model; wherein the improved YOLOv7-tiny network model includes replacing the Conv modules in all CBS modules in the backbone network with PConvBS modules; replacing the ELAN modules in the backbone network with Faster-PELAN modules by feature fusion; using the ELAN-H module in the neck network for multi-scale feature fusion; optimizing the backbone network and the detection head, including reducing the number of downsampling times of the backbone network and removing the corresponding large target detection head; the detection head uses αShape-IoU Loss is used as the loss function; a model training and evaluation module 205 is used to train and evaluate the citrus detection network model; a model testing module 206 is used to obtain the optimal weight of the citrus detection network model and test the citrus detection network model; a detection result generation module 207 is used to obtain the test result through the citrus detection model and calculate the precision and recall rate based on the intersection-over-union (IOU) ratio to determine the accuracy of the detection.

[0135] It is not difficult to find that this embodiment is a system embodiment corresponding to the first embodiment, and this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment, and in order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in the first embodiment.

[0136] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of the present invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by the present invention, but this does not mean that there are no other units in this embodiment.

[0137] A third embodiment of the present invention relates to a network side server, such as Fig.11 As shown, it includes at least one processor 302; and a memory 301 that is communicatively connected to the at least one processor 302; wherein the memory 301 stores instructions that can be executed by the at least one processor 302, and the instructions are executed by the at least one processor 302 so that the at least one processor 302 can execute the above-mentioned data processing method.

[0138] The memory 301 and the processor 302 are connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 302 and the memory 301 together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. The data processed by the processor 302 is transmitted on a wireless medium via an antenna, and further, the antenna also receives data and transmits the data to the processor 302.

[0139] The processor 302 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management and other control functions. The memory 301 can be used to store data used by the processor 302 when performing operations.

[0140] The fourth embodiment of the present invention relates to a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the citrus detection method in a complex background based on improved deep learning in the first embodiment.

[0141] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0142] The above is only an embodiment of the present invention. The common sense such as the known specific structure and characteristics in the scheme is not described in detail here. The ordinary technicians in the relevant field know all the common technical knowledge in the technical field of the invention before the application date or priority date, can obtain all the existing technologies in the field, and have the ability to apply the conventional experimental means before that date. The ordinary technicians in the relevant field can improve and implement this scheme in combination with their own abilities under the enlightenment given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the relevant field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can be made, which should also be regarded as the protection scope of the present invention, which will not affect the effect of the implementation of the present invention and the practicality of the patent. The protection scope required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

[0143] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A citrus detection method under complex background based on improved deep learning, characterized in that: include: S1, collect citrus images with complex backgrounds in orchards to obtain the initial data set; S2, preprocessing the initial data set to obtain a first data set; S3, annotating the first data set and performing format conversion to obtain a target data set; S4, using the YOLOv7-tiny network model as the deep learning benchmark model to build an improved YOLOv7-tiny network model as a citrus detection network model; wherein the improved YOLOv7-tiny network model includes replacing the Conv modules in all CBS modules in the backbone network with PConvBS modules; replacing the ELAN modules in the backbone network with Faster-PELAN modules by feature fusion; using the ELAN-H module in the neck network for multi-scale feature fusion; optimizing the backbone network and the detection head, including reducing the number of downsampling of the backbone network and removing the corresponding large target detection head; the detection head uses αShape-IoU Loss as the loss function; S5, training and evaluating the citrus detection network model; S6, obtaining the optimal weight of the citrus detection network model and testing the citrus detection network model; S7, obtain the test results through the citrus detection model and calculate the precision and recall based on the intersection-over-union ratio (IOU) to determine the accuracy of the detection.

2. The citrus detection method under complex background based on improved deep learning according to claim 1, characterized in that: In S2, the preprocessing of the initial data set includes screening, data enhancement, and resolution formatting of the initial data set.

3. The citrus detection method under complex background based on improved deep learning according to claim 2, characterized in that: In S2, the specific steps of preprocessing the initial data set to obtain the first data set include: Step S21, screening the initial data set, including deleting blurred images caused by shooting, light, equipment stability and other factors, and deleting citrus image data lacking key features; Step S22, data enhancement includes color space transformation and direction inversion of the image; Step S23, resolution formatting includes performing resolution formatting processing on the initial data set, and selecting a suitable standard resolution according to the requirements of model training and the actual situation of hardware resources.

4. The citrus detection method under complex background based on improved deep learning according to claim 1, characterized in that: In S3, the specific content of labeling the first data set includes labeling category, labeling center point, and labeling width and height information. The format conversion of the first data set includes reorganizing and extracting the label information in the xml file to obtain a label file in txt format as the target data set.

5. The citrus detection method under complex background based on improved deep learning according to claim 4, characterized in that: The calculation formulas for the number of floating-point operations and memory access of the PConvBS module are as follows: Among them, h is the height of the feature map, w is the width of the feature map, k is the size of the convolution kernel, and c is the p is the number of output channels, where c p / c=1 / 4.

6. The method for detecting citrus fruits in complex background based on improved deep learning according to claim 4, characterized in that: The Faster-PELAN module is composed of a PConvBS module, a DWConvBS module, a Concat module and a CBS module; the Faster-PELAN module has two paths, the first path will pass through the PConvBS module and the DWConvBS module for fast sampling, and the second path will pass through the PConvBS module for sampling; the feature map sampled by a specific layer on the first path will be spliced ​​with the feature map sampled by the second path, and finally the feature aggregation process will be completed through a convolution layer; the DWConvBS module is composed of a DWConvBS module, a BatchNormalization layer and a SiLU activation function in sequence.

7. The citrus detection method under complex background based on improved deep learning according to claim 1, characterized in that: The calculation formula of the loss function αShape-IoU Loss is: L αShape-IoU =α(1-IoU)+(1-α)distance shape +0.5×Ω shape Among them, L αShape-IoU is the loss function, α is the weight parameter, IoU is the intersection over union ratio, and distance shape is the shape distance, Ω shape is the shape complexity related item; The calculation formula for the intersection over union (IoU) is: Among them, IoU is the intersection over union ratio, B is the predicted bounding box, and B gt is the true bounding box; IoU is used to measure the degree of overlap between the predicted box and the true box. The value of IoU is between 0 and 1. The larger the value, the higher the overlap between the predicted box and the true box. Shape distance shape The calculation formula is: Among them, distance shape is the shape distance, x c and c is the coordinate of the center of the prediction box, and is the coordinate of the center of the real box, c is a constant, HH and WW are coefficients related to the width and height of the box, and w gt is the width of the real frame, h gt is the height of the real frame, (w gt ) scale is the new width value obtained by scaling the width of the real frame, (h gt ) scale The new height value after scaling the height of the real frame; Shape complexity related term Ω shape The calculation formula is: Among them, Ω shape is the shape complexity related term, w t is the shape of the box and θ is a fixed parameter.

8. A citrus detection system in complex background based on improved deep learning, characterized in that: Applied to the citrus detection method in complex background based on improved deep learning as described in claim 1, the system comprises: The data set collection module collects citrus images with complex backgrounds in the orchard to obtain the initial data set; A data set preprocessing module, used for preprocessing the initial data set to obtain a first data set; A data set format conversion module, used to annotate the first data set and perform format conversion to obtain a target data set; An improved modeling module is used to construct an improved YOLOv7-tiny network model as a citrus detection network model using the YOLOv7-tiny network model as a deep learning benchmark model; wherein the improved YOLOv7-tiny network model includes replacing the Conv modules in all CBS modules in the backbone network with PConvBS modules; replacing the ELAN modules in the backbone network with Faster-PELAN modules by feature fusion; using the ELAN-H module in the neck network for multi-scale feature fusion; optimizing the backbone network and the detection head, including reducing the number of downsampling of the backbone network and removing the corresponding large target detection head; and using αShape-IoU Loss as the loss function of the detection head; Model training and evaluation module, used to train and evaluate the citrus detection network model; A model testing module is used to obtain the optimal weight of the citrus detection network model and test the citrus detection network model; The detection result generation module is used to obtain the test results through the citrus detection model and calculate the precision and recall rate based on the intersection-over-union ratio (IOU) to determine the accuracy of the detection.

9. A network side server, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the citrus detection method in a complex background based on improved deep learning as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for detecting citrus in a complex background based on improved deep learning as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Lightweight target detection method based on improved Yolov7-tiny

    CN116805366A

  • Small target detection method for images acquired by unmanned aerial vehicle based on improved YOLOv8 algorithm

    CN118628939A

  • Orchard apple detection method based on YOLOv7

    CN118692074A

  • Improved YOLOv7-tiny-based citrus fruit yield estimation method

    CN118711055A