Multi-category banana bud suction detection method, device, equipment and medium in complex orchard environment
Through the improved YOLOv8n model and optimized loss function, the problems of inefficient and high labor costs of banana sprouting management are solved, efficient and accurate automated sprouting detection are achieved, and the efficiency and accuracy of agricultural production are improved.
Patent Information
- Application Number
- CN202510119950.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, banana bud removal management is inefficient and labor-intensive. The traditional artificial bud removal method is labor-intensive and time-consuming. Although mechanical bud removal and chemical agent bud removal method improve efficiency, it still cannot completely get rid of the dependence on labor.
The improved YOLOv8n model is used to build a better YOLOv8n model, including backbone network, neck network and detection head network. The global perception ability is improved through the C2fVMB module and VM2Block module, and the category weight and Slide Loss classification loss function are introduced to optimize the performance of the model on the category imbalanced dataset.
It effectively improves the accuracy and efficiency of banana sprout detection, reduces the model deviation caused by the imbalance in the number of samples in the category, enhances the generalization ability and computing efficiency of the model, and realizes an innovative and efficient solution for automated sprout management.
Smart Images

Figure CN120032364A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of agricultural production, and in particular to a method for detecting multiple types of banana sucking buds in a complex orchard environment, a corresponding device, an electronic device, and a computer-readable storage medium. Background Art
[0002] Banana is an important food and fruit crop in the world. The management of suckering during its cultivation is a key link in improving yield and quality. At present, the traditional manual suckering method is still widely used in banana production, but this process is labor-intensive and time-consuming. In view of the shortcomings of traditional manual suckering and the growth characteristics of suckers, a variety of new banana suckering technologies have been developed. These technologies are mainly divided into mechanical suckering and chemical suckering. Although these two methods have improved work efficiency to a certain extent, they have not been able to completely get rid of the dependence on manual labor.
[0003] In the past decade, object detection technology has experienced significant development, gradually transitioning from traditional machine learning-based methods to deep learning-driven solutions. In recent years, the YOLO series of algorithms (such as YOLOv5, YOLOv8) have received widespread attention in agricultural applications due to their excellent real-time performance and good recognition accuracy. At the same time, influenced by the rapid development of natural language processing, network structures such as Transformer and Mamba have also been applied to the field of computer vision and achieved good results.
[0004] In summary, in order to adapt to the problems of low efficiency and high labor cost in banana sucking bud management in the prior art, the applicant has made corresponding explorations to solve the problems. Summary of the invention
[0005] The purpose of the present application is to solve the above-mentioned problems and to provide a method for detecting multi-category banana sucking buds in a complex orchard environment, a corresponding device, an electronic device and a computer-readable storage medium.
[0006] In order to meet the various objectives of this application, this application adopts the following technical solutions:
[0007] A method for detecting multi-category banana sucking buds in a complex orchard environment is proposed to meet one of the purposes of this application, including:
[0008] Acquire an image frame of a banana orchard to be detected corresponding to a target object to be detected, wherein the target object to be detected includes one or more of a complete banana sucker, a regenerated banana sucker, or a banana pseudostem;
[0009] A lightweight YOLOv8n model is used as a benchmark model to construct an improved YOLOv8n model, wherein the improved YOLOv8n model includes a backbone network, a neck network and a detection head network, the backbone network includes a CBS structure, a C2fVMB module and an SPPF structure, the C2fVMB module is constructed by a C2f module and a VM2Block module, the VM2Block module includes a non-causal state space model, the neck network is an FPN-PAN network, and the detection head network includes a decoupled detection head, an Anchor-free mechanism and a TaskAlignedAssigner label matching strategy;
[0010] The banana orchard image frame to be detected is input into the improved YOLOv8n model that has been trained to a convergence state to detect the complete banana sucking buds, the regenerated banana sucking buds and / or the banana pseudostems in the banana orchard image frame to be detected, so as to complete the detection of multiple categories of banana sucking buds in the complex environment of the orchard.
[0011] Optionally, in the VM2Block module, the reconstructed state space model in the non-causal state space model is used to flatten the input feature map of (B, C, H, W) into a one-dimensional sequence of (B, L, C), where B represents the batch size, C represents the number of channels, H represents the height of the input feature map, W is the width of the input feature map, and L = H*W.
[0012] Optionally, in the VM2Block module, processing the input feature map includes:
[0013]
[0014] B, C, x:=DWConv(B, C, x),
[0015] y = NC-SSD(a, B, C, x),
[0016] Output = Linear(y⊙z);
[0017] Among them, NC-SSD(a, B, C, x) = Ch,
[0018] Among them, ⊙ represents the Hadamard product, which represents the operation of multiplying the corresponding elements one by one, Discretization is the discretization operation of continuous variables, Linear represents the linear layer, DWConv represents the depth-separable convolution, Δ represents the discretization step size parameter of the state transfer matrix, represents the continuous time input projection matrix, C represents the hidden state output projection matrix, x represents the input sequence of the state space model, z is used to adjust the nonlinear activation of the output feature, represents the scalar form of the continuous-time state transfer matrix, a and B are discrete form of .
[0019] Optionally, in the SPPF structure, a CBS structure including a 1×1 convolution is first passed to enhance information interaction between channels, and then three 5×5 maximum pooling layers are performed in series, and the feature map after the CBS structure and the feature map after each maximum pooling layer are spliced to achieve the fusion of local features and global features.
[0020] Optionally, in the C2f module, the input feature map is feature processed through a convolution, and then the feature map is divided into two parts, wherein the first part passes through N Bottleneck modules containing two convolutions to output N feature branches, and then passes through the VM2Block module to output another feature branch, and the second part is used as a feature branch alone to obtain N+1 feature branches.
[0021] Optionally, the FPN-PAN network is constructed by an FPN network and a PAN network, wherein in the FPN network, a plurality of feature maps of different scales are extracted from the backbone network, and then the high-level feature maps are gradually upsampled and fused with the feature maps of the lower layers through an upsampling operation, and each fused feature map is adjusted for the number of channels through a 1x1 convolution;
[0022] In the PAN network, the information of the low-level feature map is gradually transferred to the high-level feature map through downsampling.
[0023] Optionally, call the preset Slide Loss classification loss function to distinguish easy samples from difficult samples according to the preset threshold μ, and its expression is expressed as:
[0024]
[0025] Among them, μ represents the average of all IoU values.
[0026] A device for detecting multiple types of banana buds in a complex orchard environment is provided to meet another purpose of the present application, comprising:
[0027] An image frame acquisition module is configured to acquire an image frame of a banana garden to be detected corresponding to a target object to be detected, wherein the target object to be detected includes one or more of a complete banana sucker, a regenerated banana sucker or a banana pseudostem;
[0028] A detection model construction module is configured to use a lightweight YOLOv8n model as a benchmark model to construct an improved YOLOv8n model, wherein the improved YOLOv8n model includes a backbone network, a neck network, and a detection head network, the backbone network includes a CBS structure, a C2fVMB module, and an SPPF structure, the C2fVMB module is constructed by a C2f module and a VM2Block module, the VM2Block module includes a non-causal state space model, the neck network is an FPN-PAN network, and the detection head network includes a decoupled detection head, an Anchor-free mechanism, and a TaskAlignedAssigner label matching strategy;
[0029] The banana sucker detection module is configured to input the banana orchard image frame to be detected into the improved YOLOv8n model that has been trained to a convergence state, so as to detect the complete banana suckers, the regenerated banana suckers and / or the banana pseudostems in the banana orchard image frame to be detected, so as to complete the detection of multiple categories of banana suckers in the complex environment of the orchard.
[0030] An electronic device provided to meet another purpose of the present application includes a central processing unit and a memory, wherein the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the multi-category banana sucking bud detection method in a complex orchard environment described in the present application.
[0031] A computer-readable storage medium is provided to meet another purpose of the present application, which stores a computer program implemented according to the multi-category banana sucking bud detection method in a complex orchard environment in the form of computer-readable instructions. When the computer program is called and executed by a computer, the steps included in the corresponding method are executed.
[0032] Compared with the prior art, the present application aims at the problems of low efficiency and high labor cost in banana sucking bud management in the prior art. The present application includes but is not limited to the following beneficial effects:
[0033] First, when dealing with the problem of class imbalance, minority categories are usually neglected, affecting the accuracy and generalization ability of the model. This application introduces a class weight loss calculation strategy, which effectively reduces the model bias caused by the imbalance in the number of class samples, allowing the model to pay more attention to minority categories during training, thereby ensuring sufficient learning and recognition of these difficult categories.
[0034] Secondly, in the feature extraction stage of the backbone network, this application proposes a C2fVMB module and integrates VM2Block, which plays an important role in improving the global perception ability of the network. Compared with the traditional attention mechanism, VM2Block provides a wider receptive field and can better capture long-distance dependencies in the image. At the same time, VM2Block overcomes the high computational complexity of the traditional attention mechanism by adopting linear complexity, significantly improving computational efficiency and reasoning speed.
[0035] Third, in the training phase, in order to optimize the sample imbalance problem, this application introduces the Slide Loss classification loss function, which gives higher weights to difficult samples and drives the model to focus more on learning these difficult samples. This optimization measure not only improves the learning effect of the model, but also enhances the generalization ability of the model, ensuring its stable performance under different conditions.
[0036] Furthermore, this application solves the problems of accuracy, efficiency and adaptability of traditional banana sucker detection technology by introducing category weights, innovative module design and optimized loss function, and provides an innovative and efficient solution for automated sucker management in banana cultivation. The introduction of these technologies marks the combination of traditional agricultural technology and modern intelligent technology, which promotes the improvement of agricultural production efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0038] Figure 1 This is a flow chart of a method for detecting multiple types of banana sucking buds in a complex orchard environment in an embodiment of the present application;
[0039] Figure 2 An exemplary network structure of the improved YOLOv8n model in the embodiments of the present application;
[0040] Figure 3 This is an exemplary network structure of the C2fVMB module in the embodiment of the present application;
[0041] Figure 4 This is an exemplary network structure of the VM2Block module in the embodiment of the present application;
[0042] Figure 5 An exemplary network structure of an improved classification head network in an embodiment of the present application;
[0043] Figure 6 This is a principle block diagram of a multi-category banana sucking bud detection device in a complex orchard environment in an embodiment of the present application;
[0044] Figure 7 It is a schematic diagram of the structure of the computer device in the embodiment of the present application. DETAILED DESCRIPTION
[0045] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present application.
[0046] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0047] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0048] It will be understood by those skilled in the art that the "client", "terminal" and "terminal device" used herein include both devices with wireless signal receivers, which are devices with only wireless signal receivers without transmission capabilities, and devices with receiving and transmitting hardware, which are devices with receiving and transmitting hardware capable of two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers, tablet computers, which have single-line displays or multi-line displays or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service, personal communication system), which can combine voice, data processing, fax and / or data communication capabilities; PDA (Personal Digital Assistant, personal digital assistant), which may include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar and / or GPS (Global Positioning System, global positioning system) receiver; conventional laptop and / or palmtop computers or other devices, which have and / or include a conventional laptop and / or palmtop computer or other device with and / or including a radio frequency receiver. The "client", "terminal" and "terminal device" used herein may be portable, transportable, installed in a vehicle (air, sea and / or land), or suitable for and / or configured to run locally, and / or in a distributed form, at any other location on the earth and / or in space. The "client", "terminal" and "terminal device" used herein may also be a communication terminal, an Internet terminal, a music / video playing terminal, for example, a PDA, a MID (Mobile Internet Device) and / or a mobile phone with a music / video playing function, or a smart TV, a set-top box and other devices.
[0049] The hardware referred to by the names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann principle, such as a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit calls the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input and output devices to complete specific functions.
[0050] It should be pointed out that the concept of "server" referred to in this application can also be extended to the case of server clusters. According to the network deployment principle understood by those skilled in the art, the servers should be logically divided. In physical space, these servers can be independent of each other but can be called through interfaces, or integrated into a physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility, and should not use it to restrict the implementation of the network deployment method of this application.
[0051] Unless expressly specified, one or more technical features of the present application can be deployed on a server for implementation and accessed by a client through a remote call to obtain an online service interface provided by the server, or can be directly deployed and run on a client for access.
[0052] The neural network models referenced or may be referenced in this application, unless expressly specified, can be deployed on a remote server and remotely called on the client, or can be deployed and directly called on a client with sufficient device capabilities. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid excessive occupation of the client's hardware operating resources.
[0053] Unless explicitly specified, the various data involved in this application can be stored remotely on a server or on a local terminal device, as long as it is suitable for being called by the technical solution of this application.
[0054] Those skilled in the art should be aware that, although the various methods of the present application are described based on the same concept and thus present commonality to each other, unless otherwise specified, these methods can be independently executed. Similarly, for each embodiment disclosed in the present application, they are all proposed based on the same inventive concept, therefore, concepts with the same expression, and concepts that are appropriately changed for convenience despite different expressions, should be understood as equivalent.
[0055] Unless the mutually exclusive relationship between the embodiments to be disclosed in this application is explicitly stated, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct a new embodiment, as long as such combination does not deviate from the creative spirit of this application and can meet the needs of the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.
[0056] See also Figure 1 In one embodiment, the method for detecting multiple types of banana sucking buds in a complex orchard environment of the present application includes:
[0057] Step S10, obtaining an image frame of a banana orchard to be detected corresponding to a target object to be detected, wherein the target object to be detected includes one or more of a complete banana sucker, a regenerated banana sucker or a banana pseudostem;
[0058] The multi-category banana sucking bud detection system in the terminal device can obtain a banana garden image frame to be detected corresponding to the target object to be detected, wherein the target object to be detected includes one or any multiple of a complete banana sucking bud, a regenerated banana sucking bud or a banana pseudostem;
[0059] In some embodiments, the banana garden image frame can be taken using a Sonya 5100 digital camera with a resolution of 6000×4000, and the camera exposure mode is set to automatic exposure during shooting. 507 original images are taken in a real banana garden, and the images contain complete suckers, regenerated suckers and pseudostems. The 507 images taken are randomly cropped at a ratio of 1:1 to obtain 1202 images. These images are divided into a training set and a validation set at a ratio of 8:1, and labels are established according to the position, size and category of the banana complete suckers, regenerated suckers and pseudostems.
[0060] In a further embodiment, since the amount of banana pseudostem data in complex scenes is small, the training set is enhanced by offline data enhancement to improve the data diversity and richness in subsequent studies. The offline data enhancement strategy includes horizontal flipping, Gaussian noise, random brightness and contrast, ISO noise and random halo. The enhanced training set has 1975 images, 4879 complete suckers, 1196 regenerated suckers, 3798 pseudostems, and a total of 9873 rectangular frame targets. The validation set has 637 complete suckers, 139 regenerated suckers, and 458 pseudostems.
[0061] In a further embodiment, category weights are added: in multi-category target detection tasks, due to the imbalance of data in each category, the model tends to predict the category with more data. In order to solve this problem, the present application proposes a method: when calculating the classification loss, a specific weight is assigned to each category. These weights make the model pay more attention to those categories with lower frequency of occurrence during training, thereby improving the recognition ability of minority classes and the performance of the overall model. In this way, the performance of the model on an unbalanced data set can be effectively improved. The calculation formula for the weight of each category is shown in formula (1):
[0062]
[0063] Step S20: Use the lightweight YOLOv8n model as the baseline model to construct an improved YOLOv8n model. The improved YOLOv8n model includes a backbone network, a neck network, and a detection head network. The backbone network includes a CBS structure, a C2fVMB module, and an SPPF structure. The C2fVMB module is constructed by a C2f module and a VM2Block module. The VM2Block module includes a non-causal state space model. The neck network is an FPN-PAN network. The detection head network includes a decoupled detection head, an Anchor-free mechanism, and a TaskAlignedAssigner label matching strategy.
[0064] After obtaining the image frame of the banana orchard to be detected corresponding to the target object to be detected, use the lightweight YOLOv8n model as the baseline model to construct an improved YOLOv8n model. The improved YOLOv8n model includes a backbone network, a neck network, and a detection head network. The backbone network includes a CBS structure, a C2fVMB module, and an SPPF structure. The C2fVMB module is constructed by a C2f module and a VM2Block module. The VM2Block module includes a non-causal state space model. The neck network is an FPN-PAN network. The detection head network includes a decoupled detection head, an Anchor-free mechanism, and a TaskAlignedAssigner label matching strategy.
[0065] The YOLOv8 model is divided into five scale models: n, s, m, l, and x according to the number of parameters and computational complexity. Due to its high accuracy, it is widely used in object detection-related tasks. In the multi-class object detection of banana suckers in orchards under complex environments, the smallest YOLOv8n detection model is prone to missed detections and misclassifications, resulting in low accuracy. The larger YOLOv8m model has higher accuracy, but has the problem of large number of parameters and computational complexity, resulting in too high computational cost and being difficult to apply in practice. To address the above problems and achieve a balance between the number of parameters and computational complexity of the model and the model accuracy, the lightweight YOLOv8n is used as the baseline model, and its network structure is innovatively improved. The C2fVMB feature extraction module is proposed, and the VMS-YOLO network, that is, the improved YOLOv8n model of this application, is constructed accordingly. Specifically, this application expands on the basis of the C2f module. In addition to retaining the original Bottleneck branch, a new VM2Block branch is introduced. The essence of the VM2Block design lies in the integration of the non-causal state model (NC-SSD), which applies the state space theory to visual tasks and can accurately capture the global dependencies in the image while maintaining linear computational complexity.
[0066] Furthermore, the loss function is optimized. In the complex environment of the orchard, the number of easy samples is large, while the difficult samples are relatively sparse, and the network pays little attention to the difficult samples. This application introduces the Slide Loss classification loss function, calls the preset Slide Loss classification loss function, and distinguishes easy samples from difficult samples according to the preset threshold μ. The expression is expressed as:
[0067]
[0068] Among them, μ represents the average of all IoU values. μ is adaptively set to the average of all IoU values. For those samples close to the boundary (i.e., samples with IoU values close to the preset threshold μ), the SlideLoss classification loss function attempts to assign higher weights to these difficult samples so that the model can learn and utilize these samples more fully.
[0069] Furthermore, the improved YOLOv8n model is trained: the training set and the validation set are input into the improved YOLOv8n model (VMS-YOLO network) for training, the training hardware device is an Nvidia RTX3090 graphics card with 24G video memory, the training batch size is 32, the training rounds are 300, the optimizer uses AdamW, the learning rate is 0.001429, the mixup data enhancement parameter α is 0.4, and the other parameters are default parameters. The input of the training is the image data containing the complete banana sucker, the regenerated banana sucker and the banana pseudostem, and the output of the training is the rectangular box position coordinates, confidence and category probability of the target.
[0070] In some embodiments, in the C2f module, the input feature map is feature processed through a convolution, and then the feature map is divided into two parts, wherein the first part passes through N Bottleneck modules containing two convolutions to output N feature branches, and then passes through the VM2Block module to output another feature branch, and the second part is used as a feature branch alone to obtain N+1 feature branches.
[0071] In some embodiments, in the VM2Block module, the reconstructed state space model in the non-causal state space model is used to flatten the input feature map of (B, C, H, W) into a one-dimensional sequence of (B, L, C), where B represents the batch size, C represents the number of channels, H represents the height of the input feature map, W is the width of the input feature map, and L = H*W.
[0072] In a further embodiment, in the SPPF structure, a CBS structure including a 1×1 convolution is first passed through to enhance the information interaction between channels, and then three 5×5 maximum pooling layers are performed in series, and the feature map after the CBS structure and the feature map after each maximum pooling layer are spliced to achieve the fusion of local features and global features.
[0073] In a further embodiment, the FPN-PAN network is constructed by an FPN network and a PAN network, wherein, in the FPN network, multiple feature maps of different scales are extracted from the backbone network, and then the high-level feature maps are gradually upsampled and fused with the feature maps of the lower layers through an upsampling operation, and each fused feature map has its channel number adjusted through a 1x1 convolution; in the PAN network, the information of the low-level feature maps is gradually transferred to the high-level feature maps through downsampling.
[0074] Specifically, see Figure 2 The improved YOLOv8n model (VMS-YOLO network) of the present application includes four parts: input end, backbone network, neck network and detection head network. Among them, the backbone network of the improved YOLOv8n model (VMS-YOLO network) is mainly used for extracting multi-level feature information; the neck network is mainly used for fusing feature information of different levels; the detection head network uses a decoupled head structure, including a classification head network and a regression head network. The loss function of the regression head network adopts CIoU (Complete Intersection over Union) and Distribution Focal Loss (DFL). The classification head network uses the Slide Loss loss function with added category weights; Task Aligned Assigner positive and negative sample allocation strategy and Anchor-Free strategy are used for sample matching; the input end adopts Mosaic and Mixup data enhancement strategies. The Mixup data enhancement strategy is to perform weighted fusion of different samples to improve the generalization ability of the model and its robustness to occluded samples. The Mosaic data enhancement strategy splices four pictures by random scaling, cropping, and arrangement to form a new picture, which increases the size of the detection data set. The random scaling process also increases the number of small targets, which is beneficial to the detection of small targets.
[0075] Furthermore, the backbone network structure includes: CBS structure, C2fVMB module and SPPF structure. Among them, the CBS structure mainly has two functions: downsampling and feature extraction, and is composed of two-dimensional convolution, batch normalization and SiLU activation function. Batch normalization is performed after convolution to reduce the training instability problem caused by changes in input data distribution. Nonlinearity is introduced through the SiLU activation function, so that the network can learn and represent complex nonlinear relationships, and can effectively alleviate the gradient disappearance problem.
[0076] The C2fVMB module mainly includes C2f and VM2Block. Figure 3 and Figure 4 The exemplary network structures of the C2fVMB module and the VM2Block module are shown respectively. The input feature map of the C2f module is processed through a convolution, and then the feature map is divided into two parts: the first part passes through N Bottlenecks containing two convolutions to output N feature branches, and then passes through VM2Block to output another feature branch. The second part is used as a feature branch alone, and a total of N+1 branches are obtained.
[0077] This application adopts an n-scale model, where N is 1 and the number of branches is 3. The three branch feature maps are concatenated and then processed through a convolutional layer. The VM2Block module structure is as follows: Figure 4 As shown in the figure, its core is the non-causal state space model (NC-SSD). NC-SSD reconstructs the state space model (SSD) to allow it to process non-causal visual tasks. Before entering the module, the input shape (B, C, H, W) is deformed into (B, L, C), where L = H*W. In the VM2Block module, the processing of the input feature map includes:
[0078]
[0079] B, C, x:=DWConv(B, C, x),
[0080] y = NC-SSD(a, B, C, x),
[0081] Output = Linear(y⊙z);
[0082] Among them, NC-SSD(a, B, C, x) = Ch,
[0083] Among them, ⊙ represents the Hadamard product, which represents the operation of multiplying the corresponding elements one by one, Discretization is the discretization operation of continuous variables, Linear represents the linear layer, DWConv represents the depth-separable convolution, Δ represents the discretization step size parameter of the state transfer matrix, represents the continuous time input projection matrix, C represents the hidden state output projection matrix, x represents the input sequence of the state space model, z is used to adjust the nonlinear activation of the output feature, represents the scalar form of the continuous-time state transfer matrix, a and B are discrete form of .
[0084] In the SPPF structure, the input feature map first passes through a CBS structure containing 1×1 convolution to enhance the information interaction between channels, and then three 5×5 maximum pooling layers are serially performed. The feature map after the CBS structure and the feature map after each maximum pooling layer are spliced to achieve the fusion of local features and global features. The SPPF structure can efficiently aggregate multi-scale features while maintaining the simplicity and efficiency of calculation. Further, the neck network structure is FPN-PAN. FPN extracts multiple feature maps of different scales (such as P3, P4, P5, etc.) from the backbone network, and then gradually upsamples the high-level feature maps through upsampling operations and fuses them with the feature maps of lower layers. Each fused feature map will have the number of channels adjusted through 1x1 convolution. PAN gradually transfers the information of the low-level feature map to the high-level feature map through downsampling. Convolutional layers and pooling layers are used to achieve this bottom-up information transfer.
[0085] Furthermore, the detection head structure includes a decoupled detection head, an Anchor-free mechanism, a loss calculation, and a TaskAlignedAssigner label matching strategy. Among them, the decoupled detection head processes the bounding box regression and the target classification separately, learns them separately through different network branches, and then stacks and outputs the prediction results of three different scales. This method can effectively reduce the number of parameters and computational complexity, enhance the generalization ability and robustness of the model, and improve the prediction accuracy. For the bounding box regression branch, its output shape is (B, 4×reg_max, H, W). Among them, B is the batch size, reg_max is the number of channels used to calculate DFL, 4 represents the number of channels used to encode the bounding box position and size information, and H and W represent the height and width of the feature map used for prediction, respectively. For the target classification branch, its output shape is (B, nc, H, W), where nc represents the number of categories, and then the classification loss is calculated.
[0086] In order to deal with the problem of poor recognition ability of minority classes caused by the imbalance of the number of categories in multi-category object detection tasks, weight categories are added in the loss calculation process to alleviate the impact of the imbalance of the number of categories. Anchor-free mechanism: The number of predictions for each position is reduced from 3 to 1, and four values are directly predicted (i.e., the offset to the upper left corner of the grid and the height and width of the prediction box). In the feature map, the model marks the points close to the center of the true target box as positive samples, while the points far from the center are marked as negative samples. The positive and negative sample division and calculation process are significantly simplified, thereby simplifying the training and decoding stages. The number of parameters and GFLOPS of the detector are greatly reduced, with faster speed and better performance. The loss calculation part includes two parts: bounding box loss (BboxLoss) and classification loss (ClsLoss). Among them, the bounding box loss uses DFL and CIOU. DFL improves the accuracy of bounding box regression and the focus on difficult samples in object detection by modeling the distribution of continuous coordinates of the bounding box and introducing a focus mechanism. The CIoU loss function optimizes the bounding box regression in object detection by comprehensively considering the overlapping area, center point distance, and aspect ratio of the bounding box, thereby improving the positioning accuracy and convergence speed.
[0087] The exemplary network structure of the improved classification head network is as follows: Figure 5 As shown in the figure, in the classification loss calculation process, category weights are introduced to alleviate the problem of low detection accuracy of a few categories caused by the imbalance of the number of categories in the data set. SlideLoss is used as the regression loss function, and the difficult samples and simple samples are divided according to the mean IoU of all predicted samples. The weights are calculated using the above formula (2) to increase the model's attention to difficult samples near the mean IoU, thereby improving the model accuracy. The TaskAlignedAssigner label assignment strategy dynamically adjusts the assignment of positive and negative samples according to the IoU value and classification score of the predicted box and the true value box, which improves the model's ability to learn the features of diverse targets and enhances the target detection performance.
[0088] Step S30: input the banana orchard image frame to be detected into the improved YOLOv8n model that has been trained to a convergence state, so as to detect the complete banana sucking buds, the regenerated banana sucking buds and / or the banana pseudostems in the banana orchard image frame to be detected, so as to complete the detection of multiple categories of banana sucking buds in the complex environment of the orchard.
[0089] A lightweight YOLOv8n model is used as a benchmark model to construct an improved YOLOv8n model. When the improved YOLOv8n model is trained to a convergence state, it can be put into production and used to detect complete banana sucking buds, regenerated banana sucking buds and / or banana pseudostems in the image frame of the banana orchard to be detected. The image frame of the banana orchard to be detected can be input into the improved YOLOv8n model that has been trained to a convergence state to detect complete banana sucking buds, regenerated banana sucking buds and / or banana pseudostems in the image frame of the banana orchard to be detected, so as to complete the detection of multiple categories of banana sucking buds in a complex environment of an orchard.
[0090] It can be seen from the above embodiments that, compared with the prior art, the present application aims at the problems of low efficiency and high labor cost in banana sucking bud management in the prior art. The present application includes but is not limited to the following beneficial effects:
[0091] First, when dealing with the problem of class imbalance, minority categories are usually neglected, affecting the accuracy and generalization ability of the model. This application introduces a class weight loss calculation strategy, which effectively reduces the model bias caused by the imbalance in the number of class samples, allowing the model to pay more attention to minority categories during training, thereby ensuring sufficient learning and recognition of these difficult categories.
[0092] Secondly, in the feature extraction stage of the backbone network, this application proposes a C2fVMB module and integrates VM2Block, which plays an important role in improving the global perception ability of the network. Compared with the traditional attention mechanism, VM2Block provides a wider receptive field and can better capture long-distance dependencies in the image. At the same time, VM2Block overcomes the high computational complexity of the traditional attention mechanism by adopting linear complexity, significantly improving computational efficiency and reasoning speed.
[0093] Third, in the training phase, in order to optimize the sample imbalance problem, this application introduces the Slide Loss classification loss function, which gives higher weights to difficult samples and drives the model to focus more on learning these difficult samples. This optimization measure not only improves the learning effect of the model, but also enhances the generalization ability of the model, ensuring its stable performance under different conditions.
[0094] Furthermore, this application solves the problems of accuracy, efficiency and adaptability of traditional banana sucker detection technology by introducing category weights, innovative module design and optimized loss function, and provides an innovative and efficient solution for automated sucker management in banana cultivation. The introduction of these technologies marks the combination of traditional agricultural technology and modern intelligent technology, which promotes the improvement of agricultural production efficiency and accuracy.
[0095] See also Figure 6 , a multi-category banana sucking bud detection device in a complex environment of an orchard is provided to meet one of the purposes of the present application, including an image frame acquisition module 1100, a detection model construction module 1200 and a banana sucking bud detection module 1300. The image frame acquisition module 1100 is configured to acquire a banana orchard image frame to be detected corresponding to a target to be detected, wherein the target to be detected includes one or any multiple of a complete banana sucking bud, a regenerated banana sucking bud or a banana pseudostem; the detection model construction module 1200 is configured to use a lightweight YOLOv8n model as a reference model to construct an improved YOLOv8n model, wherein the improved YOLOv8n model includes a backbone network, a neck network and a detection head network, the backbone network includes a CBS structure, a C2fVMB module and an SPPF structure, and the C2fVMB module is composed of a C2f module and a VM2B lock module, the VM2Block module includes a non-causal state space model, the neck network is an FPN-PAN network, and the detection head network includes a decoupled detection head, an Anchor-free mechanism, and a TaskAlignedAssigner label matching strategy; the banana sucking bud detection module 1300 is configured to input the banana garden image frame to be detected into the improved YOLOv8n model that has been trained to a convergent state, so as to detect the complete banana sucking buds, the regenerated banana sucking buds and / or the banana pseudostems in the banana garden image frame to be detected, so as to complete the detection of multiple categories of banana sucking buds in the complex environment of the orchard.
[0096] Based on any embodiment of this application, please refer to Figure 7 Another embodiment of the present application further provides an electronic device, which can be implemented by a computer device, such as Figure 7 As shown, a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database may store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a method for detecting multiple categories of banana sucking buds in a complex orchard environment. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the method for detecting multiple categories of banana sucking buds in a complex orchard environment of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that Figure 7The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0097] In this embodiment, the processor is used to execute Figure 6 The memory stores the program codes and various data required to execute the above modules. The network interface is used to transmit data between user terminals or servers. The memory in this embodiment stores the program codes and data required to execute all modules in the multi-category banana sucking bud detection device under the complex environment of the orchard of this application, and the server can call the program codes and data of the server to execute the functions of all modules.
[0098] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the method for detecting multi-category banana sucking buds in a complex orchard environment as described in any embodiment of the present application.
[0099] The present application also provides a computer program product, including a computer program / instruction, which, when executed by one or more processors, implements the steps of the method for detecting multi-category banana sucking buds in a complex orchard environment as described in any embodiment of the present application.
[0100] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments of the present application can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0101] The above description is only a partial implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
[0102] In summary, this application solves the problems of accuracy, efficiency and adaptability of traditional banana sucker detection technology by introducing category weights, innovative module design and optimized loss function, and provides an innovative and efficient solution for automated sucker management in banana cultivation. The introduction of these technologies marks the combination of traditional agricultural technology and modern intelligent technology, which promotes the improvement of agricultural production efficiency and accuracy.
Claims
1. A method for detecting multi-category banana sucking buds in a complex orchard environment, characterized in that: include: Acquire an image frame of a banana orchard to be detected corresponding to a target object to be detected, wherein the target object to be detected includes one or more of a complete banana sucker, a regenerated banana sucker, or a banana pseudostem; A lightweight YOLOv8n model is used as a benchmark model to construct an improved YOLOv8n model, wherein the improved YOLOv8n model includes a backbone network, a neck network and a detection head network, the backbone network includes a CBS structure, a C2fVMB module and an SPPF structure, the C2fVMB module is constructed by a C2f module and a VM2Block module, the VM2Block module includes a non-causal state space model, the neck network is an FPN-PAN network, and the detection head network includes a decoupled detection head, an Anchor-free mechanism and a TaskAlignedAssigner label matching strategy; The banana orchard image frame to be detected is input into the improved YOLOv8n model that has been trained to a convergence state to detect the complete banana sucking buds, the regenerated banana sucking buds and / or the banana pseudostems in the banana orchard image frame to be detected, so as to complete the detection of multiple categories of banana sucking buds in the complex environment of the orchard.
2. The method for detecting multi-category banana sucking buds in a complex orchard environment according to claim 1, characterized in that: In the VM2Block module, the reconstructed state space model in the non-causal state space model is used to flatten the input feature map of (B, C, H, W) into a one-dimensional sequence of (B, L, C), where B represents the batch size, C represents the number of channels, H represents the height of the input feature map, W is the width of the input feature map, and L = H*W.
3. The method for detecting multi-category banana sucking buds in a complex orchard environment according to claim 2, characterized in that: In the VM2Block module, the processing of the input feature map includes: B,C,x:=DWConv(B,C,x), y=NC-SSD(a,B,C,x), Output = Linear(y☉z); Among them, NC-SSD(a,B,C,x)=Ch, Among them, ⊙ represents the Hadamard product, which represents the operation of multiplying the corresponding elements one by one, Discretization is the discretization operation of continuous variables, Linear represents the linear layer, DWConv represents the depth-separable convolution, Δ represents the discretization step size parameter of the state transfer matrix, represents the continuous time input projection matrix, C represents the hidden state output projection matrix, x represents the input sequence of the state space model, z is used to adjust the nonlinear activation of the output feature, represents the scalar form of the continuous-time state transfer matrix, a and B are discrete form of .
4. The method for detecting multi-category banana sucking buds in a complex orchard environment according to claim 1, characterized in that: In the SPPF structure, a CBS structure containing 1×1 convolution is first passed to enhance the information interaction between channels, and then three 5×5 maximum pooling layers are performed in series. The feature map after the CBS structure and the feature map after each maximum pooling layer are spliced to achieve the fusion of local features and global features.
5. The method for detecting multi-category banana sucking buds in a complex orchard environment according to claim 1, characterized in that: In the C2f module, the input feature map is processed through a convolution, and then the feature map is divided into two parts. The first part passes through N Bottleneck modules containing two convolutions to output N feature branches, and then passes through the VM2Block module to output another feature branch. The second part is used as a feature branch alone to obtain N+1 feature branches.
6. The method for detecting multiple types of banana sucking buds in a complex orchard environment according to claim 1, characterized in that: The FPN-PAN network is constructed by an FPN network and a PAN network, wherein in the FPN network, multiple feature maps of different scales are extracted from the backbone network, and then the high-level feature maps are gradually upsampled and fused with the feature maps of the lower layers through an upsampling operation, and each fused feature map is adjusted for the number of channels through a 1x1 convolution; In the PAN network, the information of the low-level feature map is gradually transferred to the high-level feature map through downsampling.
7. The method for detecting multiple types of banana sucking buds in a complex environment of an orchard according to any one of claims 1 to 6, characterized in that: Call the preset Slide Loss classification loss function to distinguish easy samples from difficult samples according to the preset threshold μ. Its expression is expressed as: Among them, μ represents the average of all IoU values.
8. A device for detecting multi-category banana buds in a complex orchard environment, characterized in that: include: An image frame acquisition module is configured to acquire an image frame of a banana garden to be detected corresponding to a target object to be detected, wherein the target object to be detected includes one or more of a complete banana sucker, a regenerated banana sucker or a banana pseudostem; A detection model construction module is configured to use a lightweight YOLOv8n model as a benchmark model to construct an improved YOLOv8n model, wherein the improved YOLOv8n model includes a backbone network, a neck network, and a detection head network, the backbone network includes a CBS structure, a C2fVMB module, and an SPPF structure, the C2fVMB module is constructed by a C2f module and a VM2Block module, the VM2Block module includes a non-causal state space model, the neck network is an FPN-PAN network, and the detection head network includes a decoupled detection head, an Anchor-free mechanism, and a TaskAlignedAssigner label matching strategy; The banana sucker detection module is configured to input the banana orchard image frame to be detected into the improved YOLOv8n model that has been trained to a convergence state, so as to detect the complete banana suckers, the regenerated banana suckers and / or the banana pseudostems in the banana orchard image frame to be detected, so as to complete the detection of multiple categories of banana suckers in the complex environment of the orchard.
9. An electronic device, comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.