Improved YOLOv5s network-based Chinese fir seedling growth state identification and quantity statistics method and device, medium and product

By improving the training of the fir seedling detection model of the YOLOv5s network, the problem of low efficiency and accuracy of fir seedling growth status recognition and quantity statistics is solved, and more efficient and accurate identification and statistics are achieved.

CN120182818APending Publication Date: 2025-06-20ZHEJIANG FORESTRY UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510247781.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The efficiency and accuracy of fir seedling growth status identification and quantity statistics are low. Traditional methods rely on manual observation, which are inefficient and prone to large errors.

Method used

Using a method based on the improved YOLOv5s network, the target image is acquired and input into the fir seedling detection model, and the improved YOLOv5s network (constructed based on the ShuffleNetV2 model, ESCA module and YOLOv5s network) is trained to obtain the fir seedling detection model, which is used to identify and count the number of fir seedlings.

Benefits of technology

The efficiency and accuracy of fir seedling growth status identification and quantity statistics are improved, manual intervention is reduced, and data accuracy and reliability are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182818A_ABST
    Figure CN120182818A_ABST
Patent Text Reader

Abstract

The invention discloses a cunninghamia lanceolata seedling growth state recognition and quantity statistics method and device based on an improved YOLOv5s network, a medium and a product, and relates to the technical field of target detection, and the method comprises the steps: obtaining a target image; the target image is an image of a Chinese fir seedling planting area to be detected; inputting the target image into a cunninghamia lanceolata seedling detection model to obtain the target image marked with the prediction frame and a prediction value of the growth state of each cunninghamia lanceolata seedling in the cunninghamia lanceolata seedling planting area to be detected; the cunninghamia lanceolata seedling detection model is obtained by training an improved YOLOv5s network, and the improved YOLOv5s network is constructed based on a ShuffleNetV2 model, an ESCA module and the YOLOv5s network; and based on the target image marked with the prediction box, determining the number of the cedarwood seedlings in the cedarwood seedling planting area to be detected. The efficiency and precision of Chinese fir seedling growth state recognition and quantity statistics are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of object detection, and particularly to a method, device, medium and product for identifying the growth status and counting the number of Chinese fir seedlings based on an improved YOLOv5s network. Background Art

[0002] With global warming, there may be more droughts and extreme weather events in the future. Many studies have shown that drought not only affects the growth and development of trees, but also affects forest ecosystems and biogeographical processes. Therefore, cultivating high-quality and highly resistant seedlings to cope with the impact of extreme weather is an important topic in forestry breeding. Chinese fir is well-known for its valuable wood properties and high wood productivity. In recent years, with the continuous expansion of the cultivation scale of Chinese fir seedlings, it is crucial to quickly and accurately obtain the quantity and growth status of each Chinese fir seedling at different stages. Traditional methods rely on manual observation for statistical analysis. In large-scale scenarios, these methods are inefficient and prone to large errors. Summary of the Invention

[0003] The purpose of the present application is to provide a method, device, medium and product for identifying the growth status and counting the number of Chinese fir seedlings based on an improved YOLOv5s network, so as to solve the problems of low efficiency and accuracy in identifying the growth status and counting the number of Chinese fir seedlings.

[0004] To achieve the above purpose, the present application provides the following solutions:

[0005] In the first aspect, the present application provides a method for identifying the growth status and counting the number of Chinese fir seedlings based on an improved YOLOv5s network, including:

[0006] Obtain a target image; the target image is an image of the planting area of the Chinese fir seedlings to be measured;

[0007] Input the target image into a Chinese fir seedling detection model to obtain the target image with labeled prediction boxes and the predicted values of the growth status of each Chinese fir seedling in the planting area of the Chinese fir seedlings to be measured; the Chinese fir seedling detection model is obtained by training an improved YOLOv5s network, and the improved YOLOv5s network is constructed based on the ShuffleNetV2 model, the ESCA module and the YOLOv5s network;

[0008] Based on the target image with labeled prediction boxes, determine the number of Chinese fir seedlings in the planting area of the Chinese fir seedlings to be measured.

[0009] Optionally, the determination process of the Chinese fir seedling detection model includes:

[0010] Obtain a training set; the training set includes: multiple sample images of Chinese fir seedling planting areas, sample images of Chinese fir seedling planting areas with labeled ground truth boxes, and the ground truth values of the growth states of each Chinese fir seedling in each sample image of the Chinese fir seedling planting area.

[0011] Construct the improved YOLOv5s network;

[0012] Train the improved YOLOv5s network using the training set to obtain the Chinese fir seedling detection model.

[0013] Optionally, obtaining a training set includes:

[0014] Obtain an initial data set; the initial data set includes: multiple original images of sample Chinese fir seedling planting areas;

[0015] Perform augmentation processing on each original image in the initial data set to obtain an augmented data set; the augmentation processing includes: random cropping, random offset, Mosaic data augmentation, and noise processing;

[0016] Annotate each image in the augmented data set to obtain an annotated data set; the annotated data set includes: each image in the augmented data set, each image in the augmented data set with labeled ground truth boxes, and the ground truth values of the growth states of each Chinese fir seedling in each image in the augmented data set;

[0017] Divide the annotated data set according to a set ratio to obtain the training set, test set, and validation set.

[0018] Optionally, the improved YOLOv5s network includes: a feature extraction backbone network, a feature fusion neck network, and a detection head;

[0019] The feature extraction backbone network includes: a two-dimensional convolutional module, a batch normalization layer, a ReLU activation function, a max pooling layer, and 6 ShuffleNetV2 models connected in sequence;

[0020] The feature fusion neck network includes: 4 convolutional modules, 4 C3 modules, 2 upsampling modules, 4 splicing modules, and 3 ESCA modules;

[0021] The detection head includes: 3 convolutional modules.

[0022] Optionally, the ShuffleNetV2 model includes: a first ShuffleNetV2 unit or a second ShuffleNetV2 unit;

[0023] The first ShuffleNetV2 unit includes: a channel splitting module, two two-dimensional convolution modules, a depth convolution module, three batch normalization layers, two ReLU activation functions, a splicing module, and a channel shuffle layer;

[0024] The second ShuffleNetV2 unit includes: three two-dimensional convolution modules, two depth convolution modules, five batch normalization layers, three ReLU activation functions, a splicing module, and a channel shuffle layer.

[0025] Optionally, the ESCA module includes: a channel attention module, a spatial attention module, and a multiplication module;

[0026] The channel attention module includes: a two-dimensional convolution module, an average pooling module, a sigmoid activation function, and a multiplication module;

[0027] The spatial attention module includes: an average pooling module, a max pooling module, a splicing module, a two-dimensional convolution module, and a sigmoid activation function.

[0028] Optionally, training the improved YOLOv5s network using the training set to obtain the Chinese fir seedling detection model includes:

[0029] Training the improved YOLOv5s network according to the training set according to the total loss function to obtain the Chinese fir seedling detection model; the total loss function includes:

[0030] L 总 = L FocalEIoU + L BCE + L ocnf ;

[0031]

[0032] where, L 总 is the total loss value; L FocalEIoU is the localization loss value; L BCE is the growth state loss value; L conf is the confidence loss value; N is the number of sample Chinese fir seedlings in the training set; is γ times the IoU of the i-th sample Chinese fir seedling, γ is a parameter used to control the degree of outlier suppression, IoU is the intersection over union, A is the area of the ground truth box of the sample Chinese fir seedling, B is the area of the predicted box of the sample Chinese fir seedling; L EIoU,i is the L EIoU of the i-th sample Chinese fir seedling, L EIoU is the EIOU loss value of the sample Chinese fir seedling, L EIoU = L IoU + Ldis +L asp ,L IoU is the IOU loss value of the sample Chinese fir seedlings, L IoU = 1 - IoU, L dis is the distance loss of the sample Chinese fir seedlings, ρ is the Euclidean distance between the ground truth box and the predicted box of the sample Chinese fir seedlings, b pr is the x-axis coordinate of the center of the predicted box of the sample Chinese fir seedlings, b gt is the x-axis coordinate of the center of the ground truth box of the sample Chinese fir seedlings, a e is the diagonal distance of the smallest bounding box covering the ground truth box and the predicted box of the sample Chinese fir seedlings, L asp is the aspect ratio loss of the sample Chinese fir seedlings, w pr is the width of the predicted box of the sample Chinese fir seedlings, w gt is the width of the ground truth box of the sample Chinese fir seedlings, a w is the width of the smallest bounding box covering the ground truth box and the predicted box of the sample Chinese fir seedlings, h pr is the height of the predicted box of the sample Chinese fir seedlings, h gt is the height of the ground truth box of the sample Chinese fir seedlings, a h is the height of the smallest bounding box covering the ground truth box and the predicted box of the sample Chinese fir seedlings; y i is the ground truth value of the growth state of the i-th sample Chinese fir seedling; is the predicted value of the growth state of the i-th sample Chinese fir seedling; is the ground truth confidence label of the predicted box of the i-th sample Chinese fir seedling; is the confidence prediction value of the predicted box of the i-th sample Chinese fir seedling.

[0033] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the Chinese fir seedling growth state recognition and quantity statistics method based on the improved YOLOv5s network described in any one of the above.

[0034] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the Chinese fir seedling growth state recognition and quantity statistics method based on the improved YOLOv5s network described in any one of the above.

[0035] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the Chinese fir seedling growth state recognition and quantity statistics method based on the improved YOLOv5s network described in any one of the above.

[0036] According to the specific embodiments provided in this application, the following technical effects are disclosed in this application:

[0037] This application discloses a method, device, medium and product for identifying the growth status and counting the number of Chinese fir seedlings based on an improved YOLOv5s network. First, a target image is obtained; the target image is an image of the planting area of the Chinese fir seedlings to be measured. Then, the target image is input into the Chinese fir seedling detection model to obtain the target image with labeled prediction boxes and the predicted values of the growth status of each Chinese fir seedling in the planting area of the Chinese fir seedlings to be measured; the Chinese fir seedling detection model is obtained by training the improved YOLOv5s network, and the improved YOLOv5s network is constructed based on the ShuffleNetV2 model, the ESCA module and the YOLOv5s network. Finally, based on the target image with labeled prediction boxes, the number of Chinese fir seedlings in the planting area of the Chinese fir seedlings to be measured is determined. This application uses the ShuffleNetV2 model and the ESCA module to improve the YOLOv5s network to obtain the improved YOLOv5s network, and uses the Chinese fir seedling detection model obtained by training the improved YOLOv5s network to predict the prediction boxes and the growth status, and further determines the number of Chinese fir seedlings, improving the efficiency and accuracy of identifying the growth status and counting the number of Chinese fir seedlings. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0039] Figure 1 Schematic flowchart of a method for identifying the growth status and counting the number of Chinese fir seedlings based on an improved YOLOv5s network provided in an embodiment of this application;

[0040] Figure 2 Schematic diagram of the YOLOv5s network structure;

[0041] Figure 3 Schematic diagram of the improved YOLOv5s network structure;

[0042] Figure 4 Schematic diagram of the first ShuffleNetV2 unit;

[0043] Figure 5 Schematic diagram of the structure of the second ShuffleNetV2 unit;

[0044] Figure 6Schematic diagram of the ESCA module structure;

[0045] Figure 7 Schematic diagram of box-validation loss;

[0046] Figure 8 Schematic diagram of the mAP progress of the improved YOLOv5s network and the original YOLOv5s during training and validation;

[0047] Figure 9 Detection result map of SSD for Chinese fir seedling planting area 1;

[0048] Figure 10 Detection result map of SSD for Chinese fir seedling planting area 2;

[0049] Figure 11 Detection result map of Faster-R-CNN for Chinese fir seedling planting area 1;

[0050] Figure 12 Detection result map of Faster-R-CNN for Chinese fir seedling planting area 2;

[0051] Figure 13 Detection result map of YOLOv3 for Chinese fir seedling planting area 1;

[0052] Figure 14 Detection result map of YOLOv3 for Chinese fir seedling planting area 2;

[0053] Figure 15 Detection result map of YOLOv5s for Chinese fir seedling planting area 1;

[0054] Figure 16 Detection result map of YOLOv5s for Chinese fir seedling planting area 2;

[0055] Figure 17 Detection result map of YOLOv8s for Chinese fir seedling planting area 1;

[0056] Figure 18 Detection result map of YOLOv8s for Chinese fir seedling planting area 2;

[0057] Figure 19 Detection result map of YOLOv10 for Chinese fir seedling planting area 1;

[0058] Figure 20 Detection result map of YOLOv10 for Chinese fir seedling planting area 2;

[0059] Figure 21 Detection result map of the Chinese fir seedling detection model for Chinese fir seedling planting area 1;

[0060] Figure 22 It is a detection result diagram of the Chinese fir seedling detection model for the Chinese fir seedling planting area 2;

[0061] Figure 23 It is a schematic diagram of the first detection effect;

[0062] Figure 24 It is a schematic diagram of the second detection effect;

[0063] Figure 25 It is a schematic structural diagram of a computer device provided in an embodiment of the present application. Detailed implementation manners

[0064] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0065] The purpose of the present application is to provide a method, device, medium and product for identifying the growth state and counting the quantity of Chinese fir seedlings based on an improved YOLOv5s network, aiming to improve the efficiency and accuracy of identifying the growth state and counting the quantity of Chinese fir seedlings.

[0066] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0067] In an exemplary embodiment, as Figure 1 shown, a method for identifying the growth state and counting the quantity of Chinese fir seedlings based on an improved YOLOv5s network is provided, including:

[0068] Step 1: Obtain a target image; the target image is an image of the Chinese fir seedling planting area to be measured.

[0069] Step 2: Input the target image into the Chinese fir seedling detection model to obtain the target image with labeled prediction boxes and the predicted values of the growth states of each Chinese fir seedling in the Chinese fir seedling planting area to be measured.

[0070] Among them, the Chinese fir seedling detection model is obtained by training the improved YOLOv5s network, and the improved YOLOv5s network is constructed based on the ShuffleNetV2 model, the ESCA module and the YOLOv5s network.

[0071] Specifically, as Figure 2As shown in the figure, the YOLOv5s network structure mainly consists of three parts: a backbone network, a neck network, and a detection head. The backbone network includes multiple convolutional modules (Conv), multiple C3 modules, and an SPPF module. The Conv module encapsulates three functions: two-dimensional convolution (Conv2d), batch normalization (BN), and the Swish activation function. The C3 module is an improved CSPDarknet53 based on Darknet53, which can reduce the computational complexity of the model and improve the inference speed. The SPPF module is an improved spatial pyramid pooling, which aims to perform pooling operations on feature maps of different scales while keeping their sizes unchanged, thereby enhancing the expressiveness of features. The neck network combines the elements of a feature pyramid network and a path aggregation network, and uses upsampling and downsampling operations to combine feature maps from different levels. The neck network includes: multiple upsampling modules (Upsample), multiple C3 modules, and multiple concatenation modules (Concat). The detection head consists of three parts: a large (80×80) convolutional module, a medium (40×40) convolutional module, and a small (20×20) convolutional module, with an image resolution of 640×640.

[0072] As an optional implementation manner, in step 2, the determination process of the Chinese fir seedling detection model includes:

[0073] Step 21: Obtain a training set; the training set includes: multiple sample images of Chinese fir seedling planting areas, sample images of Chinese fir seedling planting areas with labeled true boxes, and the true values of the growth states of each Chinese fir seedling in each sample image of the Chinese fir seedling planting area.

[0074] As an optional implementation manner, step 21 includes:

[0075] Step 211: Obtain an initial data set; the initial data set includes: multiple original images of Chinese fir seedling planting areas.

[0076] Step 212: Perform augmentation processing on each original image in the initial data set to obtain an augmented data set; the augmentation processing includes: random cropping, random offset, Mosaic data augmentation, and noise processing.

[0077] Step 213: Label each image in the augmented data set to obtain a labeled data set; the labeled data set includes: each image in the augmented data set, each image in the augmented data set with a labeled true box, and the true values of the growth states of each Chinese fir seedling in each image in the augmented data set.

[0078] Step 214: Divide the labeled data set according to a set ratio to obtain a training set, a test set, and a validation set.

[0079] Step 22: Construct an improved YOLOv5s network.

[0080] As an alternative implementation, as Figure 3 shown, the improved YOLOv5s network includes: a feature extraction backbone network (Backbobe), a feature fusion neck network (Neck), and a detection head (YOLO Head).

[0081] The feature extraction backbone network includes: a two-dimensional convolutional module (Conv2d), a batch normalization layer, a ReLU activation function, a max pooling layer (MaxPool), and six ShuffleNetV2 models (Shuffle_Block) connected in sequence.

[0082] The feature fusion neck network includes: four convolutional modules, four C3 modules, two upsampling modules, four splicing modules, and three ESCA modules.

[0083] The detection head includes: three convolutional modules.

[0084] As an alternative implementation, the ShuffleNetV2 model includes: a first ShuffleNetV2 unit as Figure 4 shown or a second ShuffleNetV2 unit as Figure 5 shown.

[0085] The first ShuffleNetV2 unit includes: a channel splitting module, two two-dimensional convolutional modules, a depthwise convolutional module (DWConv), three batch normalization layers, two ReLU activation functions, a splicing module, and a channel shuffle layer (ChannelShuffle).

[0086] The second ShuffleNetV2 unit includes: three two-dimensional convolutional modules, two depthwise convolutional modules, five batch normalization layers, three ReLU activation functions, a splicing module, and a channel shuffle layer.

[0087] Specifically, ShuffleNetV2 is a lightweight neural network whose design follows four principles to ensure the efficiency and performance of the network: keeping the number of input channels equal to the number of output channels to minimize the memory access cost; avoiding overusing grouped convolutions which increase the memory access cost; reducing fragmented operations to improve the parallel computing efficiency; and avoiding using element-wise operations that affect the memory access cost. When the stride is 1, i.e., the first ShuffleNetV2 unit, the input feature X is first divided into two parts by a channel split module. The left branch performs an identity mapping, while the right branch performs three consecutive convolutional operations (two of the convolutional modules are non-group convolutions). Finally, a channel shuffle layer is used to ensure information interaction between the two branches. When the stride is 2, i.e., the second ShuffleNetV2 unit, the input feature X is processed through two convolutional branches, then concatenated along the channel dimension, and finally a channel shuffle layer is used to ensure information interaction between the two convolutional branches. This design follows four principles to ensure the efficiency and performance of the network.

[0088] As an alternative implementation, as Figure 6 shown, the ESCA module includes: a channel attention module, a spatial attention module, and a multiplication module.

[0089] The channel attention module includes: a two-dimensional convolutional module, an average pooling module, a sigmoid activation function, and a multiplication module.

[0090] The spatial attention module includes: an average pooling module, a max pooling module, a concatenation module, a two-dimensional convolutional module, and a sigmoid activation function.

[0091] Specifically, the ESCA module is an improved version based on the ECA attention mechanism. Considering the various limitations of the ECA attention mechanism, such as implementing local cross-channel interactions and using one-dimensional convolutional operations, it may not be able to fully capture global cross-channel dependencies. In addition, it mainly focuses on channel dependencies and ignores the importance of spatial positions, thus limiting its feature representation ability. Therefore, this application proposes an improved version, namely ESCA.

[0092] Specifically, the ESCA module includes:

[0093] (1) Channel attention operation, the formula includes:

[0094]

[0095] ChannelConv2d = σ(Mean(C, dim)).

[0096] X channel = x × ChannelConv2d.

[0097] Among them, C is the feature map after performing a two-dimensional convolution operation on x with a size of h×w×c, with a size of h×w×c, where h is the height, w is the width, and c is the number of channels; Conv2d(·) is the two-dimensional convolution operation; x is the input feature map of the ESCA module; k is the convolution kernel; Mean(·) is the mean operation; Mean(C, dim) is the output feature map after average pooling, with a size of h×w×1; dim is the dimension, dim = 1; ChannelConv2d is the output feature map after being processed by the sigmoid activation function, with a size of h×w×1; X channel is the output feature map of the ESCA module, with a size of h×w×c.

[0098] (2) Spatial attention operation, the formula includes:

[0099] Avgout = Mean(X channel , dim).

[0100] Maxout = Max(X channel , dim).

[0101] SpatialAttention = σ(Conv2d 7×7 (Concat[Avgout, Maxout])).

[0102] Among them, Avgout is the output feature map after average pooling, with a size of h×w×1; Maxout is the output feature map after max pooling, with a size of h×w×1; Conv2d 7×7 (·) is a two-dimensional convolution operation with a kernel size of 7×7; Concat[·] is the concatenation operation; Concat[Avgout, Maxout] is the feature map after concatenating Avgout and Maxout, with a size of h×w×2; Conv2d 7×7 (Concat[Avgout, Maxout] is the feature map after performing a two-dimensional convolution operation on Concat[Avgout, Maxout], with a size of h×w×1; SpatialAttention is the output feature map of the spatial attention module, with a size of h×w×c.

[0103] (3) Multiplication operation, the formula includes:

[0104] W = x × X channel × SpatialAttention.

[0105] Among them, W is the output feature map of the ESCA module.

[0106] Step 23: Train the improved YOLOv5s network using the training set to obtain a Chinese fir seedling detection model.

[0107] As an alternative implementation, step 23 includes:

[0108] Train the improved YOLOv5s network using the training set according to the total loss function to obtain a Chinese fir seedling detection model; the total loss function includes:

[0109] L 总 = L FocalEIoU + L BCE + L ocnf .

[0110]

[0111] Among them, L 总 is the total loss value; L FocalEIoU is the localization loss value; L BCE is the growth state loss value; L conf is the confidence loss value; N is the number of sample Chinese fir seedlings in the training set; is γ times the IoU of the i-th sample Chinese fir seedling, γ is a parameter used to control the degree of outlier suppression, IoU is the intersection over union, A is the area of the ground truth box of the sample Chinese fir seedling, B is the area of the predicted box of the sample Chinese fir seedling; L EIoU,i is the L EIoU of the i-th sample Chinese fir seedling, L EIoU is the EIOU loss value of the sample Chinese fir seedling, L EIoU = L IoU + L dis + L asp L IoU is the IOU loss value of the sample Chinese fir seedling, L IoU = 1 - IoU, L dis is the distance loss of the sample Chinese fir seedling, ρ is the Euclidean distance between the ground truth box and the predicted box of the sample Chinese fir seedling, b pr is the x-axis coordinate of the center of the predicted box of the sample Chinese fir seedling, b gt is the x-axis coordinate of the center of the ground truth box of the sample Chinese fir seedling, a e is the diagonal distance of the smallest bounding box covering the ground truth box and the predicted box of the sample Chinese fir seedling, L asp is the aspect ratio loss of the sample Chinese fir seedling, w pr is the width of the predicted box of the sample Chinese fir seedling, w gt is the width of the ground truth box of the sample Chinese fir seedling, a wis the width of the minimum bounding box covering the ground truth box and the predicted box of the sample Chinese fir seedlings, h pr is the height of the predicted box of the sample Chinese fir seedlings, h gt is the height of the ground truth box of the sample Chinese fir seedlings, a h is the height of the minimum bounding box covering the ground truth box and the predicted box of the sample Chinese fir seedlings; y i is the ground truth value of the growth state of the i-th sample Chinese fir seedling; is the predicted value of the growth state of the i-th sample Chinese fir seedling; is the true confidence label of the predicted box of the i-th sample Chinese fir seedling; is the confidence prediction value of the predicted box of the i-th sample Chinese fir seedling.

[0112] Step 3: Based on the target image after annotating the predicted boxes, determine the number of Chinese fir seedlings in the planting area of the Chinese fir seedlings to be measured.

[0113] Specifically, determine the number of predicted boxes in the target image after annotating the predicted boxes as the number of Chinese fir seedlings in the planting area of the Chinese fir seedlings to be measured.

[0114] Furthermore, the method of this application was also evaluated using the validation set, and the evaluation metrics include:

[0115]

[0116]

[0117] FLOPs = Ci × Kw × Kh × W × H × Co (Disregard bias).

[0118] Fps = 1000 / Forward processing time.

[0119] Among them, Precision is the precision; TP is the number of correctly detected sample Chinese fir seedlings; FP is the number of falsely detected sample Chinese fir seedlings; Recall is the recall rate; FN is the number of missed detected sample Chinese fir seedlings; F1 is the F1 score; AP is the average precision; mAP is the mean average precision; J is the number of sample Chinese fir seedlings in the validation set; AP j is the AP of the j-th sample Chinese fir seedling in the validation set; FLOPs is the number of floating-point operations per second; Ci is the input / output channel number; Kw is the width of the convolutional kernel; Kh is the height of the convolutional kernel; W is the width of the feature map; H is the height of the feature map; Co is the output channel number; Disregard bias is to ignore the bias; Fps is the number of frames per second; Forward processing time is the forward processing time.

[0120] Furthermore, the improved Chinese fir seedling detection model was used to train and test the sample dataset to obtain a weight file. The weight file was deployed using the neural network inference framework NCNN open-sourced by Tencent and the Android Studio software, and an application (APP) based on the Android mobile phone terminal was developed to identify and count the number of Chinese fir seedlings and the growth rates of different families.

[0121] The software and hardware configurations for the training and testing of the Chinese fir seedling detection model are shown in Table 1. For the training of the detection model, the SGD optimizer was used to optimize the neural network weights. The initial learning rate was 0.01, the momentum factor was 0.937, the batch size was 8, and the number of iterations (Epochs) was 250. The platforms used to develop this application included a computer running Windows 11 64-bit and a smartphone running Android 12. The experimental code was developed based on JAVA 8.0, JDK 11.0.13, SDK 31, JRE 1.8.0, NCNN, and Android Studio Dolphin 3.1.

[0122] Table 1 Software and hardware configuration table for model training and testing

[0123]

[0124] Experimental results and analysis.

[0125] Result 1: Training and validation of the model.

[0126] To more intuitively display the progress of the Chinese fir seedling detection model and the YOLOv5s network during training and validation, the loss change curves as shown in Figure 7 and Figure 8 were plotted. Figure 7 The box validation loss shown in measures the actual position of the target Chinese fir seedlings in the image. Figure 8 shows the mAP progress of the improved YOLOv5s (ImprovedYOLOv5s) network and the original YOLOv5s during training and validation. In the first 50 iterations, significant fluctuations occurred in the mAP values and loss values of both models. After approximately 50 iterations, the loss of the improved YOLOv5 was lower than the original version, while the mAP value was higher than the original version. This indicates that the improved YOLOv5 has a deeper neural network and superior performance. As the model learns, its performance continuously improves. As shown in Figure 8 the reduction in box validation loss leads to an increase in mAP. This confirms that the improved YOLOv5s is superior to YOLOv5s in terms of training performance.

[0127] Result 2: Comparison of ablation experiments.

[0128] To evaluate the impact of the proposed method on the baseline model (YOLOv5s), ablation experiments were conducted on various improved models with the same hyperparameters and configurations. The experimental results are shown in Table 2. Compared with the baseline model, the accuracy and mAP of the improved model increased by 0.5% and 1.7% respectively, while the recall decreased by 0.6%. The parameters decreased by 87.89%, the FLOPs decreased by 87.97%, and the model size decreased by 87.87%. These results confirm that the improved YOLOv5s of this application has higher detection accuracy than the original YOLOv5s, while significantly reducing the parameters, FLOPs, and model size.

[0129] Table 2 Comparison of ablation experiment results

[0130]

[0131] Result 3. Comparison between the improved YOLOv5s and other object detection algorithms.

[0132] Using the same experimental environment and dataset, the method proposed in this application was compared with six different detection models (SSD, Faster-R-CNN, YOLOv3, YOLOv5s, YOLOv8s, and YOLOv10) to analyze the performance of the Chinese fir seedling detection model. The resolution of the images is related to the network structure of the model, and the default values set by the authors were used for comparison in the experiment. The experimental results are shown in Table 3. The results show that both YOLOv5s and the improved YOLOv5s have higher accuracy than other detection models, and the mAP is above 92%. However, the improved model exceeds YOLOv5s in terms of FPS and model size, making it more suitable for deployment on mobile devices, which proves the effectiveness of the improvement in this application.

[0133] Table 3 Comparison of evaluation indicators for each algorithm

[0134]

[0135] The above models were applied to the detection of Chinese fir seedling images. The experimental results are as Figures 9 - 22 shown. Obviously, all models have problems of omission, duplication, and false detection. Specifically, SSD shows serious detection problems. In contrast, both YOLOv5s and the improved YOLOv5s model of this application show fewer false detections. These results further verify the superiority of the model in this application in detection.

[0136] The detection model of Chinese fir seedlings in this application was deployed using the open-source neural network inference framework NCNN of Tencent and the Android Studio software, and an application (APP) based on the Android mobile phone terminal was developed. Then, 30 images were randomly selected from the test set to compare the performance of the YOLOv5s network and the detection model of Chinese fir seedlings in this application on low-performance mobile devices. As Figure 23 and Figure 24 shown, the detection frame rate achieved by the detection model of Chinese fir seedlings in this application is almost three times that of the original YOLOv5s network, and a high FPS is always maintained.

[0137] In this application, a detection model of Chinese fir seedlings was used to develop an application based on the Android mobile phone terminal in combination with the NCNN framework, which is used to identify and count the number of Chinese fir seedlings and the growth rates of different families. The high-throughput evaluation of the number and growth rate of Chinese fir seedlings was realized, providing important technical support for the screening of fast-growing germplasms.

[0138] In an exemplary embodiment, a computer device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement a method for identifying the growth status and counting the number of Chinese fir seedlings based on the improved YOLOv5s network.

[0139] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements a method for identifying the growth status and counting the number of Chinese fir seedlings based on the improved YOLOv5s network.

[0140] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it implements a method for identifying the growth status and counting the number of Chinese fir seedlings based on the improved YOLOv5s network.

[0141] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 25As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes a method for identifying the growth status and counting the number of Chinese fir seedlings based on an improved YOLOv5s network.

[0142] Those skilled in the art can understand that Figure 25 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0143] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0144] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0145] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0146] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0147] In this text, specific examples are used to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. To sum up, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for identifying and counting the growth status of Chinese fir seedlings based on an improved YOLOv5s network, characterized in that: The method for identifying and counting the growth status of Chinese fir seedlings based on the improved YOLOv5s network includes: Acquire a target image; the target image is an image of the planting area of ​​the Chinese fir seedlings to be tested; The target image is input into a Chinese fir seedling detection model to obtain a target image with a marked prediction frame and a predicted value of the growth status of each Chinese fir seedling in the Chinese fir seedling planting area to be tested; the Chinese fir seedling detection model is obtained by training an improved YOLOv5s network, and the improved YOLOv5s network is constructed based on a ShuffleNetV2 model, an ESCA module and a YOLOv5s network; Based on the target image after the prediction box is marked, the number of Chinese fir seedlings in the Chinese fir seedling planting area to be tested is determined.

2. The method for identifying and counting the growth status of Chinese fir seedlings based on the improved YOLOv5s network according to claim 1, characterized in that: The process of determining the Chinese fir seedling detection model includes: Acquire a training set; the training set includes: a plurality of sample images of a planting area of ​​Chinese fir seedlings, sample images of a planting area of ​​Chinese fir seedlings marked with true frames, and true values ​​of growth states of each Chinese fir seedling in each sample image of a planting area of ​​Chinese fir seedlings; Construct the improved YOLOv5s network; The improved YOLOv5s network is trained using the training set to obtain the fir seedling detection model.

3. The method for identifying and counting the growth status of Chinese fir seedlings based on the improved YOLOv5s network according to claim 2, characterized in that: Get the training set, including: Acquire an initial data set; the initial data set includes: a plurality of original images of a sample fir seedling planting area; Performing an amplification process on each original image in the initial data set to obtain an amplified data set; the amplification process includes: random cropping, random offset, mosaic data enhancement and noise processing; Annotating each image in the augmented data set to obtain an annotated data set; the annotated data set includes: each image in the augmented data set, each image in the augmented data set annotated with a true frame, and a true value of the growth state of each Chinese fir seedling in each image in the augmented data set; The labeled data set is divided according to a set ratio to obtain the training set, test set and validation set.

4. The method for identifying and counting the growth status of Chinese fir seedlings based on the improved YOLOv5s network according to claim 2, characterized in that: The improved YOLOv5s network includes: a feature extraction backbone network, a feature fusion neck network and a detection head; The feature extraction backbone network includes: a two-dimensional convolution module, a batch normalization layer, a ReLU activation function, a maximum pooling layer and 6 ShuffleNetV2 models connected in sequence; The feature fusion neck network includes: 4 convolution modules, 4 C3 modules, 2 upsampling modules, 4 splicing modules and 3 ESCA modules; The detection head includes: 3 convolution modules.

5. The method for identifying and counting the growth status of Chinese fir seedlings based on the improved YOLOv5s network according to claim 4, characterized in that: The ShuffleNetV2 model includes: a first ShuffleNetV2 unit or a second ShuffleNetV2 unit; The first ShuffleNetV2 unit includes: a channel segmentation module, two two-dimensional convolution modules, a deep convolution module, three batch normalization layers, two ReLU activation functions, a splicing module and a channel shuffle layer; The second ShuffleNetV2 unit includes: 3 two-dimensional convolution modules, 2 deep convolution modules, 5 batch normalization layers, 3 ReLU activation functions, a splicing module and a channel shuffle layer.

6. The method for identifying and counting the growth status of Chinese fir seedlings based on the improved YOLOv5s network according to claim 4, characterized in that: The ESCA module includes: a channel attention module, a spatial attention module and a multiplication module; The channel attention module includes: a two-dimensional convolution module, an average pooling module, a sigmoid activation function and a multiplication module; The spatial attention module includes: an average pooling module, a maximum pooling module, a splicing module, a two-dimensional convolution module and a sigmoid activation function.

7. The method for identifying and counting the growth status of Chinese fir seedlings based on the improved YOLOv5s network according to claim 2, characterized in that: The improved YOLOv5s network is trained using the training set to obtain the fir seedling detection model, including: According to the total loss function, the improved YOLOv5s network is trained according to the training set to obtain the fir seedling detection model; the total loss function includes: L 总 =L FocalEIoU +L BCE +L ocnf ; Among them, L 总 is the total loss value; L FocalEIoU is the positioning loss value; L BCE is the growth state loss value; L conf is the confidence loss value; N is the number of sample Chinese fir seedlings in the training set; is γ times the IoU of the i-th sample fir seedling, γ is a parameter used to control the degree of suppression of outliers, IoU is the intersection over union ratio, A is the area of ​​the real frame of the sample Chinese fir seedling, and B is the area of ​​the predicted frame of the sample Chinese fir seedling; L EIoU,i is the L of the i-th sample Chinese fir seedling EIoU , L EIoU is the EIOU loss value of the sample Chinese fir seedlings, L EIoU =L IoU +L dis +L asp , L IoU is the IOU loss value of the sample Chinese fir seedlings, L IoU =1-IoU,L dis is the distance loss of the sample Chinese fir seedlings, ρ is the Euclidean distance between the true frame and the predicted frame of the sample Chinese fir seedling, b pr is the x-axis coordinate of the center of the prediction box of the sample fir seedling, b gt is the x-axis coordinate of the center of the real frame of the sample Chinese fir seedling, a e is the diagonal distance between the minimum bounding box of the real box and the predicted box covering the sample fir seedling, L asp is the aspect ratio loss of the sample Chinese fir seedlings, w pr is the width of the prediction box of the sample Chinese fir seedlings, w gt is the width of the real frame of the sample Chinese fir seedling, a w is the width of the minimum bounding box covering the true box and predicted box of the sample fir seedling, h pr is the height of the prediction box of the sample Chinese fir seedling, h gt is the height of the real frame of the sample Chinese fir seedling, a h is the height of the minimum bounding box covering the true box and the predicted box of the sample fir seedling; i is the true value of the growth status of the i-th sample Chinese fir seedling; is the predicted value of the growth status of the i-th sample Chinese fir seedling; is the true confidence label of the prediction box of the i-th sample Chinese fir seedling; is the confidence prediction value of the prediction box of the i-th sample Chinese fir seedling.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the method for identifying and counting the growth status of Chinese fir seedlings based on the improved YOLOv5s network as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying and counting the growth status of Chinese fir seedlings based on the improved YOLOv5s network as described in any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for identifying and counting the growth status of Chinese fir seedlings based on the improved YOLOv5s network as described in any one of claims 1 to 7 is implemented.