Accurate picking-oriented multi-task learning day lily segmentation and picking parameter determination method
By constructing the ADPPL-MYOLO multi-task self-optimization segmentation model, combining dynamic loss weights and high-dimensional hyperparameter optimization algorithm, the problem of occlusion and size difference in daylily field picking is solved, and the precise positioning and efficient picking of daylily picking points are achieved.
Patent Information
- Application Number
- CN202510342976.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing crop detection and segmentation methods face problems such as occlusion, large size differences and fuzzy maturity stages in picking daylily, resulting in large positioning errors and high false detection rates, making it difficult to achieve accurate picking.
ADPPL-MYOLO multi-task self-optimization segmentation model is constructed, including IMSM intensive multi-scale feature extraction module, MBFE multi-dimensional bar feature extractor, TSDS-Head task collaborative dynamic segmentation head, combined with DLWS dynamic loss weight strategy and HCOA high-dimensional hyperparameter optimization algorithm, DCSL-Daylily picking point positioning algorithm and RSEAD-CA crop angle determination algorithm are designed to realize the precise positioning of daylily picking points.
It significantly improves the accuracy and efficiency of daylily picking, reduces human labor, and is suitable for the sustainable development of the daylily industry.
Smart Images

Figure CN120279376A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep learning and crop picking, and specifically relates to a multi-task learning method for lily flower segmentation and picking parameter determination for precise picking. Background Art
[0002] As a vegetable crop, lily flower has various nutritional components and an appropriate nutritional structure, and also has extensive medicinal values such as clearing heat and detoxifying, nourishing blood and calming the liver, and enhancing brain function. Therefore, the market value of lily flower is becoming increasingly prominent. Lily flower has now become an important local characteristic industry, and large-scale planting and efficient management play a significant role in promoting regional economic prosperity and improving farmers' living standards. However, the picking process of lily flower faces many challenges. As a labor-intensive operation, the picking of lily flower is characterized by strong seasonality and short cycle, and low picking efficiency will directly affect the subsequent processing and sales links, thereby reducing the overall economic benefits. Especially during the picking peak, the shortage of labor and high costs have become the main bottlenecks restricting the development of the lily flower industry, resulting in significant economic losses. Therefore, exploring and applying intelligent lily flower picking technology, especially achieving precise positioning of picking points, has become the key to solving the above problems.
[0003] To achieve precise segmentation of daylilies and accurate positioning of picking points is the key prerequisite for promoting precise picking operations of daylilies. In recent years, with the rapid development of computer vision and artificial intelligence technologies, crop detection methods based on deep learning have gradually become a research hotspot. Traditional image processing methods mostly rely on the extraction of features such as color, shape, and spectrum. These methods have good effects under specific conditions, but are vulnerable to the influence of light changes and background interference in complex field environments. In contrast, deep learning algorithms show strong robustness in agricultural product detection tasks through end-to-end feature learning, and their automatic feature extraction ability significantly improves the accuracy and robustness of detection. Li et al. adopted an improved YOLOv8 architecture, combined with BiFPN and SPDConv modules to detect mango fruits and stems, segmented the stem ROI through the YOLOv8n-seg model, obtained the skeleton line of the stem region based on the segmented image, and further developed a picking point positioning algorithm to determine the optimal picking point coordinates; Yan et al. proposed the MR3P-TS model, extended the network design of Mask R-CNN, determined the main part of tea buds by calculating the area of the mask connected domain, calculated the axis of the tea shoot by the minimum circumscribed rectangle, and then obtained the position coordinates of the picking point; Yu et al. used a field strawberry detection algorithm based on Mask R-CNN to not only accurately identify the mature strawberry area, but also locate the strawberry picking point by analyzing the shape and edge features of the image generated by Mask R-CNN; Liang et al. based on the YOLOv3 model, detected litchi fruits under natural night conditions, determined the region of interest of the fruit stem through the bounding box of the litchi fruit, and further used the U-Net network to finely segment the litchi fruit stalk, thus realizing the positioning of the picking area (fruit stalk); Qi et al. proposed a workflow, used YOLOv5 to detect the main stem in litchi images, performed semantic segmentation using PSPNet after extracting the ROI, and then obtained the pixel coordinates of the main stem picking point through image post-processing, and finally determined the optimal picking point position; Shuai et al. developed the YOLO-Tea model, based on the YOLOv5 network and incorporated the CARAFE upsampling operator and CBAM attention mechanism, as well as the Bottleneck Transformers module, to improve the detection accuracy of tea buds and their key points, and finally used image processing methods to locate the picking point position according to the key point information in the model inference stage; Meng et al. used the optimized YOLOX-tiny and PSP-net networks to accurately detect tea buds and segment their regions, and determined the centroid as the picking point through the segmentation results, developed a positioning algorithm to achieve tea bud positioning; Xiong et al. analyzed the color characteristics of night green grape images, adopted the method of rotating the RGB color model and processing the background, combined with the minimum circumscribed rectangle of the fruit and the Hough line detection to perform linear fitting on the fruit stem, and calculated the picking point based on the stem with the angle between the fitting line and the vertical line less than 15 degrees, achieving accurate positioning of the night green grape picking point;Rong et al. proposed an improved Swin Transformer V2 semantic segmentation model, integrated the SeMask module and the UPerNet decoder to accurately identify tomato fruits, calyces and stems, then extracted the calyx area and stem area through connected component labeling and morphological processing, and finally used image thinning, deburring algorithms and calyx position constraints to extract the picking points on the side stems; Jiang et al. proposed an improved YOLOv5-seg instance segmentation algorithm with a coordinate attention mechanism to identify and segment bitter gourds and stems, then used a thinning algorithm to thin the skeleton of the stem mask image, and finally selected the midpoint of the largest connected region as the picking point of the bitter gourd, achieving accurate identification and positioning of the picking points; Zhu et al. proposed a YOLOv5-CFD network model based on YOLOv5, which realized the identification of grapes and stems by integrating the CBAM attention mechanism, adding the fourth-layer detection and improving the Head module, and at the same time used geometric methods to quickly and accurately locate the picking points of fresh table grapes.;
[0004] In summary, the existing crop detection and segmentation methods perform excellently in dealing with scenarios where the fruits are large in volume, have significant features or grow sparsely, but face unique challenges in the field picking of daylilies. The daylily flower buds grow in spiral clusters, with multi-angle staggered occlusion and overlap between adjacent flower buds, and the immature flower buds have large size differences and blurred maturity stage boundaries, which bring great difficulties to accurate positioning. Occlusion leads to the loss of key feature information, increases the positioning error, and reduces the harvesting success rate. At the same time, the current picking point positioning methods are mostly based on the assumptions of regular fruit morphology, high color contrast or significant spatial distribution features, which do not match the characteristics of the slender tubular flower buds of daylilies, clustered growth and similar color to the stems and leaves. The picking of daylilies needs to be accurate to the connection point between the flower buds and the stems to avoid damage, but the existing research mostly focuses on the main body of the fruit, lacking in-depth exploration of the morphological characteristics of daylilies and anatomical analysis of the picking area, resulting in problems such as increased false detection rate and positioning deviation. Given that different crops need to adopt different picking strategies, for the picking point positioning of daylilies, it is necessary to conduct in-depth research closely combined with their unique growth morphology. Therefore, how to effectively reduce the influence of occlusion, size differences and blurred maturity stage on the model accuracy, and design a picking point positioning method considering the morphological characteristics of daylilies is the key problem to be solved urgently in the present invention. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a multi-task learning daylily segmentation and picking parameter determination method for precise picking, which is mainly used for the automatic picking of crops in real scenarios.
[0006] The present invention is implemented by adopting the following technical solutions:
[0007] A multi-task learning method for precision picking of daylily segmentation and picking parameter determination, which constructs an ADPPL-MYOLO multi-task self-optimizing segmentation model. The ADPPL-MYOLO multi-task self-optimizing segmentation model includes an IMSM dense multi-scale feature extraction module, an MBFE multi-dimensional bar feature extractor module, and a TSDS-Head task collaborative dynamic segmentation head module. The DLWS dynamic loss weight strategy is used to ensure the balanced performance of the model on each task. The HCOA high-dimensional hyperparameter optimization algorithm is used to improve the reliability of the model. The DCSL-Daylily and RSEAD-CA are used to determine the pixel coordinates and cropping angle information of the daylily picking points respectively.
[0008] The IMSM includes a residual structure, a DLP module, and an FCAM module. The MBFE includes a BFA module, bar convolution, and a CBS module. The TSDS-Head includes a shared convolution, average pooling, dynamic convolution, and a task decomposition module. The DCSL-Daylily includes a daylily category screening module based on the segmentation result and a multi-mask graph intersection calculation module. The RSEAD-CA includes a multi-mask graph intersection calculation module, a minimum circumscribed matrix calculation module, and a rotation angle calculation module.
[0009] Further preferably, the steps for constructing the ADPPL-MYOLO multi-task self-optimizing segmentation model are as follows:
[0010] First, construct the model backbone structure Backbone to receive the daylily image. Then, perform local feature extraction through two convolutional layers. After that, pass through the C2f module to extract the high-level semantic information of the target again. Then, enter the convolution to perform feature extraction again. After that, pass the feature information into the C2f module again to extract the high-level semantic information of the detected object again. Then, process it through the convolution module again, and pass the feature information into the IMSM module. The IMSM module, through the collaborative action of the DLP module, MGAP module, and FCAM module, integrates the complementary advantages of local features and global context information to achieve in-depth mining and efficient capture of the feature information of occluded targets. Then, pass the features through the convolution module again to the SPPF spatial pyramid pooling module to generate a fixed-size output feature map, and end the work of the backbone.
[0011] Then, two branch networks of the DPPL-MYOLO multi-task segmentation model are constructed. DPPL-MYOLO consists of two branches: one is a multi-class segmentation branch for implementing daylily segmentation, and the other is a single-class segmentation branch for implementing daylily picking area segmentation. After the feature information extracted by the backbone part is transmitted to the two branch networks through the C2f module, the IMSM module, and the SPPF spatial pyramid pooling module respectively, the two branches then extract the strip structure features of the daylily from multiple directions through the MBFE module respectively, so as to effectively capture the features of the daylily. For the multi-class segmentation branch, after the feature information processed by the MBFE is further subjected to feature fusion of shallow and deep information through the aggregation network, the TSDS-Head segmentation head integrates the task alignment method, shares the convolutional layer to extract general features, and dynamically adjusts the convolutional kernel parameters to flexibly adapt to different shapes and scales of the daylily target, realizing the efficient recognition and precise positioning of dense targets, and finally obtaining the segmentation results of different maturity levels of the daylily; for the single-class segmentation branch, the feature information processed by the MBFE is fused with the feature information extracted by the backbone part and upsampled is performed, and finally the daylily picking area segmentation result is obtained;
[0012] After that, a dynamic loss weight strategy DLWS is designed to comprehensively improve the overall segmentation effect of the model by balancing the performance of the multi-task network branches. At the same time, to further realize the adaptive optimization of the DPPL-MYOLO model and reduce the influence of artificial design subjective experience, the HCOA strategy is used to carry out the fast adaptive combination configuration of the high-dimensional hyperparameters of DPPL-MYOLO (the advanced model optimized by HCOA is named ADPPL-MYOLO), further improving the reliability of the model in the daylily segmentation task;
[0013] Finally, based on the segmentation results and combined with the characteristics of the daylily picking area, the DCSL-Daylily picking point positioning algorithm and the RSEAD-CA cutting angle determination algorithm are proposed. By calculating the center of the intersection of the double-branch mask and the direction of the minimum circumscribed matrix, the pixel coordinates and cutting angles of the daylily picking points in the dense scene are finally determined.
[0014] The technical solution provided by the present invention has the following advantages compared with the prior art:
[0015] First, aiming at the problems of serious occlusion, large scale difference, and blurred boundary in the field dense daylily picking scene, the present invention constructs an ADPPL-MYOLO multi-task self-optimizing segmentation model. By integrating multiple functional modules and strategies, the feature representation ability of dense targets is effectively improved, and the segmentation ability of the model in complex scenes is significantly improved.
[0016] Second, the present invention constructs an Intensive Multi-Scale Feature Extraction Module (IMSM). This module integrates the complementary advantages of local features and global context information through sub-modules such as the Dynamic Local Perception Module (DLP), the Multi-Head Global Attention Perception Module (MGAP), and the Feature Convolution Aggregation Module (FCAM), etc., to achieve in-depth mining and efficient capture of the feature information of occluded targets.
[0017] Third, the present invention constructs a Multi-Dimensional Bar Feature Extraction Module (MBFE). This module contains three distinctive Bar Feature Extraction (BFA) sub-modules, which extract features in different directions respectively and have different parameter settings. Through the way of cascaded connection, the MBFE module can focus on the strip structure features of daylilies in all directions and from multiple angles, achieving precise and comprehensive capture of the features of daylilies.
[0018] Fourth, the present invention constructs a Task Synergy Dynamic Segmentation Head (TSDS-Head). Through the alignment method of classification and localization tasks, it realizes the collaborative optimization between tasks. And it extracts the general features of the image through the shared convolutional layer, providing a solid foundation for subsequent tasks. On this basis, the dynamic convolutional layer can dynamically adjust the convolutional kernel parameters according to the content of the input image, thus flexibly adapting to daylily targets of different shapes and scales, promoting the effective transmission and fusion of features, and further realizing the efficient recognition and precise localization of dense targets.
[0019] Fifth, aiming at the problem that the multi-task model suppresses the performance of small tasks due to the imbalance of task complexity and data volume, the present invention proposes a Dynamic Loss Weight Strategy (DLWS) to ensure the balanced performance of the model on each task. And it combines with the HCOA High-Dimensional Hyperparameter Optimization Algorithm to establish a model adaptive optimization mechanism, improving the reliability of the model in the daylily segmentation task, providing a new direction for the performance balance and optimization of the multi-task model.
[0020] Sixth, based on the segmentation results of the ADPPL-MYOLO model proposed by the present invention, and combined with the unique characteristics of the daylily picking area, the DCSL-Daylily Daylily Picking Point Location Algorithm is further proposed. This algorithm first accurately screens out the daylily segmentation masks belonging to the mature and over-mature categories, and then precisely determines the coordinate positions of the picking points through the intersection operation of the double-branch masks. This picking point location method strictly follows the principle of only picking mature and over-mature daylilies, and resolutely does not pick daylilies that do not meet the mature standard, thus ensuring the accuracy and selectivity of picking. At the same time, this algorithm can still maintain a high degree of accuracy in the dense planting scenario, realizing the precise and efficient location of the daylily picking points.
[0021] Seventh, on the basis of successfully implementing the DCSL-Daylily daylily picking point positioning algorithm, the present invention further designs the RSEAD-CA daylily cutting angle determination algorithm. This algorithm accurately obtains the optimal picking angle by precisely calculating the angle of the shortest side of the minimum circumscribed rectangle formed by the intersection of the double-branch masks. This method ensures that the scissors of the daylily cutting device can flexibly rotate to the most suitable angle for precise cutting during the cutting task, effectively reducing the potential damage to daylilies during the cutting process.
[0022] The present invention is reasonably designed and applicable to the tasks of daylily segmentation and picking point positioning. At the same time, it helps to reduce manual labor and is of great significance to the sustainable development of the daylily industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments in accordance with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0025] Figure 1 It represents a schematic diagram of the overall framework of the present invention.
[0026] Figure 2 It represents the ADPPL-MYOLO network structure diagram.
[0027] Figure 3 It represents the IMSM network structure diagram.
[0028] Figure 4 It represents the MBFE network structure diagram.
[0029] Figure 5 It represents the TSDS-Head network structure diagram.
[0030] Figure 6 It represents the DCSL-Daylily positioning process diagram.
[0031] Figure 7 It represents the RSEAD-CA daylily cutting angle determination process diagram.
[0032] Figure 8 It represents the visualization diagram of the DLWS effect verification experiment training process.
[0033] Figure 9 It represents the experimental result diagram of comparison with other models.
[0034] Figure 10 Indicate the visualization heatmap effects of each model. Detailed implementation manners
[0035] In order to more clearly understand the above objects, features and advantages of the present invention, the solution of the present invention will be further described below.
[0036] The following will detail the specific embodiments of the present invention with reference to the accompanying drawings.
[0037] The present invention takes the multi-task self-optimizing segmentation model ADPPL-MYOLO as the core, integrates the IMSM multi-scale feature capture module, the MBFE bar feature extraction module, and the TSDS-Head task collaboration and dynamic convolution head, aiming to solve challenges such as severe occlusion, large scale differences, and blurred boundaries. The DLWS dynamic loss weight strategy is used to balance the performance of multi-task branches, and the HCOA hyperparameter optimization is utilized to significantly enhance the model's feature representation ability and automatic learning ability for dense targets. Further, the present invention designs the DCSL-Daylily daylily picking point positioning algorithm. Based on the segmentation result, this algorithm accurately determines the coordinates of the daylily picking points through the intersection of double-branch masks, and uses the RSEAD-CA cropping angle determination algorithm to accurately calculate the picking angle according to the minimum circumscribed rectangle of the intersection of double-branch masks.
[0038] A multi-task learning method for daylily segmentation and picking parameter determination for precise picking, as Figure 1 、 2 shown, constructs the ADPPL-MYOLO multi-task self-optimizing segmentation model. The ADPPL-MYOLO multi-task self-optimizing segmentation model includes the IMSM dense multi-scale feature extraction module, the MBFE multi-dimensional bar feature extractor module, and the TSDS-Head task collaboration dynamic segmentation head module; the DLWS dynamic loss weight strategy is used to ensure the balanced performance of the model in each task, the HCOA high-dimensional hyperparameter optimization algorithm is used to improve the reliability of the model, and the DCSL-Daylily and RSEAD-CA are respectively used to determine the pixel coordinates and cropping angle information of the daylily picking points. Among them, the IMSM includes a residual structure, a DLP module, an FCAM module, and an FCAM module; the MBFE includes a BFA module, bar convolution, and a CBS module; the TSDS-Head includes a shared convolution, an average pooling, a dynamic convolution, and a task decomposition module; the DCSL-Daylily includes a daylily category screening module based on the segmentation result and a multi-mask graph intersection calculation module; the RSEAD-CA includes a multi-mask graph intersection calculation module, a minimum circumscribed matrix calculation module, and a rotation angle calculation module.
[0039] First, construct the ADPPL-MYOLO multi-task self-optimizing segmentation model.
[0040] First, construct the model backbone structure Backbone, which accepts 640×640 daylily images; then, perform local feature extraction through two convolutional layers, and then through the C2f module, extract the high-level semantic information of the target again. After that, enter the convolution for feature extraction again, and then transmit the feature information to the C2f module again to extract the high-level semantic information of the detected object again. After that, process it through the convolution module again, and transmit the feature information to the IMSM module. Through the collaborative action of the DLP module, MGAP module, and FCAM module, the complementary advantages of local features and global context information are integrated to achieve in-depth mining and efficient capture of the feature information of occluded targets; then, transmit the features to the SPPF spatial pyramid pooling module through the convolution module again to generate a fixed-size output feature map, and end the work of the backbone.
[0041] Then, construct two branch networks of the multi-task segmentation model. DPPL-MYOLO consists of two branches: one is a multi-class segmentation branch for realizing daylily segmentation, and the other is a single-class segmentation branch for realizing the segmentation of the daylily picking area. After transmitting the feature information extracted by the backbone part to the two branch networks through the C2f module, the IMSM module, and the SPPF spatial pyramid pooling module respectively, the two branches respectively extract the strip structure features of the daylily from multiple directions through the MBFE module, so as to effectively capture the features of the daylily. For the multi-class segmentation branch, after the feature information processed by the MBFE is further feature-fused with the shallow and deep information through the aggregation network, the task alignment method, the shared convolutional layer is used to extract general features, and the convolutional kernel parameters are dynamically adjusted through the TSDS-Head segmentation head to flexibly adapt to different shapes and scales of daylily targets, realizing efficient recognition and precise positioning of dense targets, and finally obtaining the segmentation results of different maturity levels of daylilies; for the single-class segmentation branch, the feature information processed by the MBFE is fused with the feature information extracted by the backbone part and upsampled, and finally the segmentation result of the daylily picking area is obtained;
[0042] Finally, design the dynamic loss weight strategy DLWS to comprehensively improve the overall segmentation effect of the model by balancing the performance of the multi-task network branches. At the same time, to further realize the adaptive optimization of the DPPL-MYOLO model, reduce the influence of artificial design subjective experience, and use the HCOA strategy to carry out rapid adaptive combination configuration of the high-dimensional hyperparameters of DPPL-MYOLO (the advanced model optimized by HCOA is named ADPPL-MYOLO), further improving the reliability of the model in the daylily segmentation task.
[0043] Second, construct the IMSM. In the complex field planting scene, the intensive cultivation mode of daylilies often causes significant occlusion and overlap between fruits, which complicates the accurate segmentation task of daylilies. In this context, due to occlusion, it is difficult to comprehensively capture the key feature information of the target fruit, the visible region features become limited, and the features of the target itself are prone to concealment. All these factors together exacerbate the deviation of the positioning accuracy, thus having an adverse impact on the success rate of target segmentation. In view of this, the present invention designs a dense multi-scale feature extraction module (IMSM), and the specific structure is as shown in Figure 3 . Starting from the perspective of the synergistic effect of local features and global context information, the IMSM effectively captures the feature information of occluded targets. The IMSM consists of a dynamic local perception (DLP) mechanism, a multi-head global attention perception (MGAP) mechanism, and a feature convolution aggregation module (FCAM). Its purpose is to deeply mine and efficiently capture the feature information of occluded targets by integrating the complementary advantages of local features and global context information.
[0044] When constructing the IMSM, the feature map first enters the dynamic local perception module DLP, and the features are initially extracted through depthwise separable dynamic convolution DSDConv. Through residual connection, the input and output feature maps are fused to promote information flow and improve the model stability.
[0045] Then it enters the multi-head global attention perception module MGAP. In the process of exploring the global feature extraction of daylily images, the present invention designs the MGAP mechanism and particularly considers the impact of the intensive occlusion problem on feature extraction. This module first initializes the input features Figure X ∈R H×W×C in the channel dimension to obtain X init . Subsequently, the channel of X init is divided into four parts using the grouped convolution strategy to obtain four sub-feature maps {X1, X2, X3, X4}, and the number of channels of each sub-feature map is C / 4. These four sub-feature maps respectively enter four parallel branch networks for processing. The first branch consists of four depthwise separable dilated convolutions (3×3 convolution kernel, dilation rate of 3) in series, aiming to capture the global features. The mathematical representation is as follows:
[0046]
[0047] where represents a network composed of four depthwise separable dilated convolution layers, and each convolution layer uses a 3×3 convolution kernel and a dilation rate r = 3. This branch is particularly important for the extraction of global context information in the case of intensive occlusion. The second branch contains two dilated convolutions with the same configuration to balance the extraction of global and local features. The mathematical representation is as follows:
[0048]
[0049] Similarly, in the formula represents a network composed of two depthwise separable atrous convolutional layers. This branch can focus on both global information and local details in the presence of dense occlusions. The third branch contains only one atrous convolution, focusing on capturing local features, and its mathematical expression is:
[0050]
[0051] Similarly, in the formula represents a network that contains only one depthwise separable atrous convolutional layer. In the case of dense occlusions, this branch helps to extract local features of the unoccluded part. The fourth branch retains the original features Figure X 4 is used as a baseline (X′4 = X4), providing unprocessed original information, which helps to recover the information of the occluded part in subsequent fusion steps. The outputs of the four processed branches are concatenated along the channel dimension to obtain the concatenated features Figure X concat . Subsequently, a 1×1 convolutional layer is used to adjust the dimension of X concat so that it matches the number of channels of the feature Figure X init before grouping, and the adjusted feature Figure X adjusted is obtained. Finally, X adjusted and X init are multiplied element-wise to achieve deep fusion of features, realizing deep fusion of features. This step not only fuses the features extracted by different branches, but also helps to recover and enhance the information of the occluded part through the interaction between the original features and the adjusted features in the case of dense occlusions, and the final feature Figure X final is obtained:
[0052] X final = X adjusted ⊙ X init
[0053] where ⊙ represents the element-wise multiplication operation. Through this mechanism, when processing daylily images, the model can effectively extract and utilize global and local features even in the face of the challenge of dense occlusions. In particular, the atrous convolution design of the first and second branches helps to capture global context information and balance global and local features, thereby improving the robustness and accuracy of feature extraction in the case of dense occlusions.
[0054] After that, the features extracted by MGAP are passed to the FCAM module. The FCAM is designed to efficiently extract and aggregate image features. By using 1×1 convolution to adjust the number of channels, 3×3 convolution to capture local features, GELU activation to enhance non-linear expression, and the second 1×1 convolution to integrate global information, FCAM realizes the comprehensive extraction and efficient aggregation of image features based on DLP and MGAP. In tasks such as processing daylily images, FCAM demonstrates excellent performance, significantly improving the model's recognition ability and accuracy, and providing strong support for image analysis.
[0055] Among them, the DLP mechanism is used for local characterization of daylily images. Specifically, the present invention uses 3×3 depthwise separable dynamic convolution (DSDConv), combining the efficiency of depthwise separable convolution and the adaptability of dynamic convolution. First, the computational complexity is reduced through depthwise convolution while retaining channel independence, and then the convolution kernel parameters are dynamically adjusted to improve feature sensitivity. To enhance feature representation, the present invention introduces a residual connection to fuse the input and output feature maps, promoting information flow and improving model stability.
[0056] Third, construct MBFE. The growth directions of daylilies have significant differences. Within the field of view of a single image, their growth orientations also show diverse characteristics. This biological characteristic makes it difficult for standard networks to extract high-quality semantic features of daylilies in any direction, resulting in poor performance of the model when facing targets with low inter-class differences. To solve this problem, the present invention proposes a multi-dimensional bar feature extractor (MBFE), and the specific structure is as Figure 4 shown. MBFE extracts the strip structure features of daylilies from multiple orientations, thereby effectively capturing the features of daylilies.
[0057] The MBFE module is constructed by cascading three BFA modules with different parameters. Feature information enters three BFA modules with different parameters connected in a cascading manner. Among them, the output of the previous BFA module is used as the input of the next module and is passed sequentially. Finally, the outputs of each module are concatenated to obtain a comprehensive output result. This design realizes the feature extraction and fusion of daylily images from multiple scales by setting different convolution kernel sizes K and dilation rates d, thereby effectively capturing its overall shape and local details.
[0058] For the working process of each BFA module, first, initial feature extraction, normalization, and activation are performed through the CBS (Convolution, Batch Normalization, SiLU) module. The CBS module not only provides a basic and robust feature representation for subsequent feature processing but also captures the basic features in the image through convolution operations, batch normalization ensures the stability of features, and the SiLU activation function introduces non-linear characteristics.
[0059] F CBS = SiLU(BN(W·F in + b))
[0060] In the formula, W and b are the weight and bias of the convolutional kernel respectively. Subsequently, the BFA module uses four bar convolutional operation branches in different directions to capture the bar features in the daylily image in all directions. These branches focus on feature extraction in the horizontal, vertical, and two diagonal directions respectively. The first branch obtains the feature map F1 through a 1×k bar convolution operation (dilation rate is d), focusing on feature extraction in the horizontal direction. The second branch obtains the feature map F2 through a k×1 bar convolution operation (dilation rate is d), focusing on feature extraction in the vertical direction. The third branch first performs a horizontal transformation on the image, then through a k×1 bar convolution operation (dilation rate is d), and finally performs an inverse horizontal transformation to obtain the feature map F3, focusing on feature extraction in one diagonal direction. The fourth branch first performs a vertical transformation on the image, then through a k×1 bar convolution operation (dilation rate is d), and finally performs an inverse vertical transformation to obtain the feature map F4, focusing on feature extraction in the other diagonal direction.
[0061] F1 = Conv 1×k (F CBR ; d)
[0062] F2 = Conv k×1 (F CBR ; d)
[0063] F3 = HFlip -1 (Conv k×1 (HFlip(F CBR )); d))
[0064] F4 = VFlip -1 (Conv k×1 (VFlip(F CBR )); d))
[0065] Where HFlip represents the horizontal transformation operation, and HFlip -1 is its inverse transformation; where VFlip represents the vertical transformation operation, and VFlip -1 is its inverse transformation. After bar feature extraction in four directions, the BFA module merges the output feature maps of each branch through a feature concatenation operation to form a richer feature representation:
[0066] F concat = Concat([F1, F2, F3, F4])
[0067] The stitched feature map will be processed by the CBS module again to further enhance the feature expression ability. Each module operates in coordination to jointly achieve efficient feature extraction of daylily images from multiple scales and directions.
[0068] Fourth, construct the TSDS-Head segmentation head. When performing the daylily segmentation task in a dense scene, due to the significant difference in the target size within the immature daylily category, the small difference between the targets approaching maturity in the immature category and the targets in the mature category, and the widespread occlusion situation, the segmentation efficiency of the traditional detection head is greatly reduced in such complex scenes. Therefore, the present invention proposes a task collaborative dynamic segmentation head (TSDS-Head). As Figure 5 shown, TSDS-Head aligns by integrating classification and localization tasks, extracts general features of the image using a shared convolutional layer, and dynamically adjusts the convolutional kernel parameters according to the content of the input image using a dynamic convolutional layer, so as to flexibly adapt to daylily targets of different shapes and scales, promote the effective transmission and fusion of features, and then achieve efficient recognition and precise localization of dense targets.
[0069] TSDS-Head uses two 3×3 shared convolutional layers to perform preliminary feature extraction on the input image. These shared convolutional layers can capture the basic patterns and structures in the image and provide basic features for subsequent processing. Subsequently, through a feature fusion operation, the outputs of the two shared convolutional layers are combined to obtain the feature map F to enhance the feature expression ability. This fusion operation helps to integrate the features extracted by different convolutional layers, effectively improves the network's recognition ability of daylily targets, and provides richer feature information for subsequent classification and localization tasks. Then, the feature map F undergoes an average pooling operation to obtain F avg and then F and F avg together serve as the inputs to the Task Decomposition Component (TDC). In TDC, first, the weight W is calculated, and then weighted fusion is performed using the weight W and the feature map F to obtain the classification task feature map F cls and the localization task feature map F reg respectively. The specific process is as follows:
[0070]
[0071] F cls and F reg =σ2(BN((W⊙reshape(R))·reshape(F)))
[0072] Among them, σ1 and σ2 represent activation functions, ⊙ represents element-wise multiplication, and R is a learnable reallocation matrix. Through this mechanism, TDC can dynamically adjust the weight distribution of the feature map, thereby generating optimized feature representations for classification and localization tasks respectively. To further refine the features, F is further processed through a 3×3 convolutional layer and then through a spatial convolutional offset module (offset) to obtain the offset d and the mask m. At the same time, F obtains the classification weight map W through multi-stage feature processing cls , and then W cls and the classification task feature map F cls perform a product operation and pass through a 3×3 convolutional layer to obtain the classification task result F′ cls . Through the introduction of the classification weight map, this process further enhances the semantic information related to the daylily target in the feature map, while suppressing the interference of the background area, significantly improving the accuracy of the classification task. In the localization task branch, the offset d, the mask m, and F reg are input into the dynamic deformable convolutional layer DyDCNv2:
[0073]
[0074] where K i is the dynamically generated convolutional kernel, p i is the sampling point on the feature map, d i is the offset, m i is the mask, and N is the total number of sampling points. By adjusting the dynamic convolutional kernel, DyDCNv2 can adaptively capture the morphological and scale changes of the daylily target, especially in the case of dense distribution, significantly improving the accuracy of the localization task. After being processed by DyDCNv2, the feature map passes through a 3×3 convolutional layer and a scale adjustment operation to obtain the localization task result F′ reg , this process utilizes the flexibility of dynamic convolution and further optimizes the feature expression through scale adjustment, enabling the model to more accurately locate the boundary of the daylily target. Finally, through the feature fusion operation, the F′ cls and F′ reg are merged to obtain the output after task alignment processing. TSDS-Head significantly improves the accuracy and robustness of daylily segmentation by integrating the alignment methods of classification and localization tasks and using the dynamic convolutional layer to dynamically adjust the convolutional kernel parameters according to the input image content.
[0075] Fifth, construct the DLWS dynamic loss weight strategy. In the conventional multi-task learning framework, the calculation of the total loss usually follows a simple accumulation principle:
[0076]
[0077] where n represents the number of tasks included in the multi-task network, and loss i represents the loss of the i-th task. This calculation method implicitly assumes that all tasks contribute equally to the final model performance, so the same loss weight is assigned to each task, which is defaulted to 1. This design performs well in scenarios where the task volumes are relatively balanced.
[0078] However, the dual-branch multi-task DPPL-MYOLO network constructed in the present invention faces a unique challenge: its two branches are respectively responsible for the multi-class segmentation task of daylilies and the single-class segmentation task of their picking areas, and there is a significant imbalance between them in terms of task complexity and data volume. This imbalance causes that, in the initial stage of training, the multi-class segmentation branch with a larger task volume contributes much more to the loss function than the single-class segmentation branch, thereby suppressing the learning effect of the latter. To overcome this problem, the present invention proposes a dynamic loss weight strategy. This strategy makes a clever weight allocation for the loss terms of the multi-task segmentation task (hereinafter referred to as task A) and the single-class segmentation task (hereinafter referred to as task B). The core of this design lies in the introduction of two key variables w A and w B , which are respectively used as the weighting coefficients of the loss term loss A of task A and the loss term loss B of task B, but the calculation method is unique and aims to enhance the complementarity and balance between tasks by cross-referencing the loss terms.
[0079] Specifically, the calculation strategy of w A reflects special attention to the loss term loss B of task B.
[0080]
[0081] Among them, loss B is used as the numerator coefficient and divided by the sum of loss A and loss B . This design means that the value of w A will change with the change of loss B relative to loss A . By adding a constant term 1, it is ensured that w A is always greater than 1, so as to give a certain basic weight to the loss term loss A of task A in the loss function and dynamically adjust it according to the loss of task B on this basis. This design aims to promote the mutual learning and complementarity between multi-tasks. Especially when the loss of task B is significant, it increases the attention to task A to maintain the overall performance of multi-task learning.
[0082] Correspondingly, w BThe calculation follows the following formula, reflecting special attention to the loss term loss of task A. A Specific attention is paid.
[0083]
[0084] In this expression, loss A is used as the numerator coefficient and divided by the sum of loss A and loss B This design means that the value of w B will change with the change of loss A relative to loss B By adding the constant term 2, it is ensured that the minimum value of w B is 2, thus giving the loss term loss of task B B a higher base weight and dynamically adjusting it based on the loss of task A. This design aims to strengthen the optimization of the single-class segmentation task, especially when the loss of task A is significant, by increasing the weight of task B to ensure that the performance of a specific task is not interfered with by other tasks within the multi-task learning framework.
[0085] Finally, combining the calculated w A and w B , the present invention constructs a comprehensive loss function loss in the following form:
[0086] loss = w A ·loss A + w B ·loss B
[0087] The dynamic loss function design strategy of the present invention can utilize the loss information between tasks to dynamically adjust their respective weights, not only enhancing the complementarity between multi-task learning and single-class segmentation tasks, but also achieving efficient loss balance in a complex multi-task environment. By calculating the weights through cross-referencing the loss terms, it provides a new perspective and solution for the collaborative optimization of multi-task learning and single-class segmentation tasks, and is expected to bring new breakthroughs to the research in related fields.
[0088] Sixth, use HCOA to optimize the hyperparameters of the DPPL-MYOLO model. The optimization target of the HCOA strategy is 24 key hyperparameters in the DPPL-MYOLO model, including the initial learning rate (lr0), cyclic learning rate (lrf), learning rate momentum (momentum), weight decay coefficient (weight_decay), warm-up epochs (warmup_epochs), warm-up learning momentum (warmup_momentum), warm-up initial learning rate (warmup_bias_lr), GIoU loss coefficient (box), classification loss coefficient (cls), DFL loss coefficient (dfl), mask downsampling ratio (mask_ratio), and 13 data augmentation coefficients (translate, scale, mosaic, mixup, etc.). By adaptively adjusting the model hyperparameters during network training and giving the best hyperparameter components that maximize the fitness function value, the final ADPPL-MYOLO model is established, reducing the model design difficulty while improving the model's adaptive learning ability.
[0089] In the HCOA strategy, the fitness function value is used as the basis for evaluating the quality of individuals in the population, thereby guiding the evolution direction of the population. Therefore, the present invention defines a fitness function, which is calculated based on the performance of the ADPPL-MYOLO model in the daylily segmentation task:
[0090] fitness = 0.1·mAP@50 + 0.9·mAP@50:95
[0091] This formula is the calculation formula of the model fitness function. The obtained fitness value is used as the fitness value of HCOA for population update, and the optimal individual information is used to establish the final model. Through the adaptive learning process of the HCOA algorithm, the present invention successfully identifies the optimal hyperparameter components of the model and applies them to the final model training, as shown in Table 1.
[0092] Table 1 Results of hyperparameter optimization
[0093]
[0094]
[0095] As can be seen from Table 1, the HCOA algorithm is used to perform self-optimization of the hyperparameters of the DPPL-MYOLO model, and the optimal configuration of 24 key hyperparameters is successfully obtained. Subsequently, the initial hyperparameters and data augmentation coefficients of the model are replaced with these optimized hyperparameter components, and finally the improved ADPPL-MYOLO model is obtained. The reliability of this model in the daylily segmentation task is further improved.
[0096] Seventh, design the Daylily-DCSL daylily picking point positioning algorithm. To accurately calibrate the coordinates of the daylily picking points, the present invention proposes a Daylily Dual-Domain Centroid Synergistic Localization algorithm (Daylily-DCSL) for daylily picking point localization in dense scenarios based on ADPPL-MYOLO, which is used to solve the problem of accurate localization of daylily picking points in dense planting scenarios. The specific implementation method is as follows Figure 6 shown. This method uses the constructed dual-branch multi-task self-optimizing network architecture ADPPL-MYOLO. Under this architecture, the daylily segmentation task branch can generate the segmentation mask maps of each target (marked as maskA i , where i = 1, 2, 3,..., n, and n represents the total number of targets in the image) in the image. In parallel, the picking area segmentation task branch is responsible for drawing the segmentation mask map of the picking area (denoted as maskB). The key step in implementing the strategy is to finely process the generated segmentation mask maps. First, screen out the maskA corresponding to the categories of mature and over-mature daylilies i , and then, these screened maskA i are respectively subjected to intersection operations with the corresponding picking area masks in maskB:
[0097] Intersection i = maskA i Ι maskB
[0098] where Intersection i represents the common area between the i-th mature or over-mature daylily target and the picking area. This intersection operation aims to accurately define the overlapping part of each target object with the picking area. Finally, by accurately calculating the center points of these intersection areas, the accurate identification of the daylily picking points is achieved.
[0099]
[0100] where (x k , y k ) are the coordinates of non-zero pixel points. This strategy deeply understands and makes full use of the spatial position relationship between the main stem and the picking area, ensuring that each target corresponds to only one accurate picking point. Particularly importantly, due to the pixel-level accuracy of the semantic segmentation task, this strategy has achieved a significant improvement in accuracy compared with the object detection method.
[0101] Eighth, design the RSEAD-CA daylily cutting angle determination algorithm. After obtaining the precise coordinates of the picking points of the daylilies, this coordinate information needs to be accurately transmitted to the picking robot. Subsequently, the robot will locate to the target position based on these coordinates and start the scissors equipped on it for picking operations. However, it should be noted that the growth orientations of daylilies vary, which requires the scissors to be adjusted to an appropriate angle to ensure the smooth progress of the picking process. Based on the statistical analysis of a large number of sample data, it is found that if the picking area of the daylily is approximated as a rectangle, the scissors of the daylily cutting device should be parallel to the shortest side of this rectangle to ensure the accuracy of the cutting angle. Based on this observation result, the present invention designs a method for determining the cutting angle of the scissors of the daylily cutting device, and the implementation method is as Figure 7 shown. Therefore, accurately determining the cutting angle of the daylily has become a crucial link in the entire picking process. RSEAD-CA first uses the ADPPL-MYOLO model to generate the segmentation mask map (marked as maskA i for each target daylily in the image, where i = 1, 2, 3,..., n, and n represents the total number of targets in the image) and the segmentation mask map of the picking area (denoted as maskB). Subsequently, screen the maskA i corresponding to the categories of mature and over-mature daylilies and perform a mask intersection operation with maskB to obtain the target picking area. This intersection operation can be formally expressed as:
[0102]
[0103] In the formula, S represents the index set of the categories of mature and over-mature daylilies. Subsequently, calculate the minimum bounding rectangle of this intersection region R target , and determine the deflection angle of its shortest side in the horizontal direction as the cutting angle of the scissors of the daylily cutting device. This process can be mathematically expressed as:
[0104] θ = AngleOfShortestSide(R bounding_box (R target ))
[0105] In the formula, R bounding_box (R target ) represents calculating the minimum bounding rectangle of the intersection region R target , and the AngleOfShortestSide(·) function is used to calculate the deflection angle θ of the shortest side of this matrix relative to the horizontal direction. Through this method, the accurate determination of the cutting angle of the scissors of the daylily cutting device can be achieved, thereby improving the picking efficiency and accuracy of the daylily.
[0106] In the embodiments of the present invention, daylilies are collected, preprocessed, and pixel-level annotation is performed to distinguish daylilies from their picking areas. Among them, daylilies are subdivided into three categories: ripe, unripe, and overripe according to their growth stages for pixel-level annotation. The annotated daylily images are used to construct a dense daylily image dataset in the field according to the ratio of 8:2.
[0107] To verify the effectiveness of the proposed Dynamic Loss Weight Strategy (DLWS) of the present invention, a comparative experiment is designed in the present invention. In this experiment, the basic model and the basic model integrated with the adaptive weight (DLWS) strategy are respectively used for training under the same dataset and consistent experimental settings. To ensure fairness and accuracy, the maximum number of iterations of both models is set to 300 times to ensure that they can both reach the convergence state. During the training process, the performance changes of the two models are recorded in detail and plotted in Figure 8 where the blue curve represents the training record of the basic model, and the yellow curve shows the training record of the basic model integrated with the DLWS strategy for intuitive comparison and analysis.
[0108] As Figure 8 shown, in the initial stage of training the basic model, the learning process of the single-class segmentation branch was significantly suppressed by the multi-class segmentation branch. This phenomenon directly led to the final learning effect of the single-class segmentation branch failing to reach the expected level. To alleviate this inhibition problem between branches, the present invention integrates the adaptive weight (DLWS) strategy into the basic model. This strategy effectively avoids the potential inhibition of one branch on another by dynamically adjusting the loss weights of different task branches and balancing the influence of the two branches during the training process. The experimental results show that in the training process of the model integrated with the DLWS strategy, the training curves of both segmentation branches show good smoothness. This means that the two branches can maintain a relatively balanced learning progress during the training process, and the inhibition between them is significantly alleviated. Further analysis reveals that the model integrated with the DLWS strategy does not have a negative impact on the learning effect of the multi-class segmentation branch while improving the learning effect of the single-class segmentation branch. On the contrary, due to the more balanced weight allocation of the two branches during the training process, they can make full use of the information in the training data, thus improving the overall segmentation performance. In summary, the experimental data strongly verifies the excellent effectiveness of the proposed Dynamic Loss Weight Strategy of the present invention in dealing with the inter-branch inhibition problem caused by the imbalance of task amounts in multi-task segmentation tasks. This strategy can not only significantly improve the learning effect of the single-class segmentation branch but also maintain the learning effect of the multi-class segmentation branch, providing strong support for achieving more efficient and accurate multi-task segmentation.
[0109] To fully verify the effectiveness of each improved module in the model of the present invention for the segmentation of dense daylilies, the present invention designs ablation experiments. By means of controlling variables, each improved module and its combination are respectively added to the model. The specific experimental content and test results are shown in Table 2.
[0110] Table 2 Ablation Experiment of ADPPL-MYOLO Model
[0111]
[0112] (1) The dynamic loss weight strategy (DLWS) uses the loss information between tasks to dynamically adjust their respective weights, and enhances the complementarity and balance between tasks by cross-referencing loss terms. It effectively solves the problem that in a complex multi-task learning environment, the multi-class segmentation branch contributes significantly more to the loss function than the single-class segmentation branch due to the large amount of tasks, thus avoiding the situation where the learning effect of the single-class segmentation branch is suppressed. After adding DLWS, the single-class segmentation accuracy of the model is improved by 1.77%, and at the same time, the multi-class segmentation accuracy is not affected but slightly improved. It can be seen that using the dynamic loss weight successfully balances the learning contributions of the two segmentation branches, not only eliminating the performance imbalance caused by the task volume difference, but also promoting a positive cycle of mutual promotion between the two branches.
[0113] (2) The intensive multi-scale feature extraction module (IMSM) realizes the in-depth mining and efficient capture of the feature information of occluded targets by integrating the complementary advantages of local features and global context information. It effectively solves the problem that it is difficult to comprehensively capture the key feature information of targets due to occlusion. After adding IMSM, the segmentation accuracies of the model for the three categories of mature, immature, and over-ripe are improved by 2.03%, 1.08%, and 0.86% respectively compared with only adding DLWS, and are improved by 3.09%, 1.00%, and 0.67% respectively compared with the original model. It can be seen that combining local features with global context information has an obvious effect on capturing the information of occluded targets.
[0114] (3) The multi-dimensional bar feature extractor (MBFE) uses bar convolutions in different orientations to extract the strip structure features of daylilies from multiple directions, effectively solving the problem that it is difficult for a standard network to extract high-quality semantic features of daylilies in any direction due to the significant difference in the growth directions of daylilies. After adding MBFE to the model, the segmentation accuracy of each category is improved to a certain extent. It can be seen that capturing features from multiple directions has a positive impact on optimizing the performance of the segmentation task.
[0115] (4) The TSDS-Head proposed in this invention focuses on solving the problems of missed detection and false detection in daylily segmentation in dense scenes due to target overlap, large size differences, and mixed category features. The TSDS-Head uses the alignment method of classification and localization tasks, combines the advantages of shared convolution and dynamic convolution, effectively solves the problems caused by small differences between immature and mature targets, variable sizes, and occlusions. After adding this module, the segmentation accuracy of each category has been improved.
[0116] Through the comparative analysis of ablation experiments, it is found that after integrating all improvement strategies, the performance improvement effect of the final model is the most significant. The algorithm proposed in this invention has achieved different degrees of performance improvement for the targets of each category. Especially in dealing with the problems of poor segmentation effects caused by dense occlusion, subtle differences between immature and mature targets, and variable target sizes, it has demonstrated excellent problem-solving capabilities. Specifically, compared with the basic model, ADPPL-MYOLO has been effectively improved in four types of samples: immature, mature, over-ripe, and picking areas. This fully verifies the effectiveness and superiority of the algorithm of this invention in the dense daylily segmentation task.
[0117] To fully verify the superiority of the method of this invention, this invention selects advanced segmentation models such as YOLOv5l-seg, YOLOv5x-seg, YOLOv7-seg, YOLOv8l-seg, YOLOv8x-seg, YOLOv9s-seg, YOLOv10l-seg, YOLOv11L-seg, and YOLOv11x-seg, and conducts a comprehensive comparison from multiple dimensions such as mAP, number of parameters, computational volume, and FPS. The comparison results are shown in Table 3.
[0118] Table 3 Comparison of ADPPL-MYOLO Model with Other Advanced Models
[0119]
[0120] Combined with Table 3 and Figure 9It can be seen that in the comparative experiments on the field dense daylily dataset, the OURS model demonstrated excellent performance. The OURS model achieved 61.51% and 38.05% respectively on the two key metrics of mAP@50 and mAP@50-95, far higher than other YOLO series segmentation models, which fully proves that the OURS model has higher accuracy and robustness in the segmentation task of the field dense daylily dataset. At the same time, the precision and recall rate of the OURS model were also as high as 72.16% and 65.37% respectively, superior to other comparative models, further verifying its reliability in practical applications. The FLOPs of the OURS model is 180.5G, and the number of parameters is 50.15M, which is at a medium level compared with other YOLO series segmentation models. This indicates that the OURS model not only maintains high performance but also has good computational efficiency and storage friendliness. In addition, the frame rate of the OURS model is 44FPS. Although it is slightly lower than some YOLO series segmentation models, considering its significant advantages in mAP, precision, and recall rate, this frame rate performance is still acceptable. The experimental results show that the OURS model performed excellently in the comparative experiments on the field dense daylily dataset. It not only achieved significant advantages in key metrics but also maintained a good balance in computational complexity and the number of parameters, demonstrating broad application prospects and potential value.
[0121] It can be clearly seen from Figure 10 that other models showed obvious limitations when identifying and focusing on the entire segmentation target area. In particular, these models tend to over-focus on non-target areas outside the segmentation target, which not only introduces a large amount of useless information and redundant data but also seriously interferes with the precise segmentation of the target area, thereby weakening the overall segmentation accuracy. In contrast, the ADPPL-MYOLO model of the present invention demonstrated excellent focusing ability and was able to accurately lock and emphasize the segmentation target area. Even facing the challenges of complex and variable image backgrounds and numerous interference factors, this model can still robustly identify and focus on the characteristic area of daylilies, showing extremely high recognition accuracy and robustness. This significant advantage fully indicates that the segmentation model proposed in the present invention performs excellently in effectively capturing and enhancing key image features, thus achieving a significant improvement in segmentation accuracy.
[0122] This invention was carried out under the conditions of 13th Gen Intel(R) Core(TM) i5-13490F@2.50GHz CPU, NVIDIA GeForce RTX 4070Ti SUPER, 16GB of memory, Linux Ubuntu 20.04 operating system, and PyTorch 3.8 environment.
[0123] The present invention is applicable to a wide range of crop segmentation and harvesting tasks. The application of the present invention helps to reduce economic losses caused by manual experience and physical limitations, and accelerates the engineering process of crop automated harvesting technology.
[0124] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Although the foregoing embodiments have been described in detail, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments, and they should all be covered by the protection scope of the claims.
Claims
1. A multi-task learning method for precise picking of daylily segmentation and picking parameter determination, characterized in that: Construct an ADPPL-MYOLO multi-task self-optimizing segmentation model. The ADPPL-MYOLO multi-task self-optimizing segmentation model includes an IMSM dense multi-scale feature extraction module, an MBFE multi-dimensional bar feature extractor module, and a TSDS-Head task collaborative dynamic segmentation head module. Ensure the balanced performance of the model on each task with the DLWS dynamic loss weight strategy, improve the reliability of the model with the HCOA high-dimensional hyperparameter optimization algorithm, and determine the pixel coordinates and cropping angle information of the daylily picking points with DCSL-Daylily and RSEAD-CA respectively. The IMSM includes a residual structure, a DLP module, and an FCAM module. The MBFE includes a BFA module, bar convolution, and a CBS module. The TSDS-Head includes a shared convolution, average pooling, dynamic convolution, and a task decomposition module. The DCSL-Daylily includes a daylily category screening module based on the segmentation result and a multi-mask graph intersection calculation module. The RSEAD-CA includes a multi-mask graph intersection calculation module, a minimum bounding matrix calculation module, and a rotation angle calculation module.
2. The multi-task learning method for daylily segmentation and picking parameter determination for precise picking according to claim 1, characterized in that: Construct an ADPPL-MYOLO multi-task self-optimizing segmentation model as follows: First, construct the model backbone structure Backbone to receive the daylily image. Then, perform local feature extraction through two convolutional layers. After that, pass through the C2f module to extract the high-level semantic information of the target again. Then, enter the convolution to perform feature extraction again, and pass the feature information into the C2f module to extract the high-level semantic information of the detected object again. After that, perform convolution processing again and pass the feature information into the IMSM module. The IMSM module, through the collaborative action of the DLP module, MGAP module, and FCAM module, integrates the complementary advantages of local features and global context information to achieve in-depth mining and efficient capture of the feature information of occluded targets. Then, pass the features through the convolution module again to the SPPF spatial pyramid pooling module to generate a fixed-size output feature map, ending the work of the backbone. Then, two branch networks of the DPPL-MYOLO multi-task segmentation model are constructed; DPPL-MYOLO consists of two branches: one is a multi-class segmentation branch for implementing daylily segmentation, and the other is a single-class segmentation branch for implementing daylily picking area segmentation; after the feature information extracted by the backbone part is transmitted to the two branch networks respectively through the C2f module, IMSM module and SPPF spatial pyramid pooling module, the two branches then extract the strip structure features of the daylily from multiple directions through the MBFE module respectively, so as to effectively capture the features of the daylily; among them, for the multi-class segmentation branch, after the feature information processed by the MBFE is subjected to feature fusion of shallow and deep information through the aggregation network, it is then integrated and aligned through the TSDS-Head segmentation head, the shared convolutional layer extracts general features, and the convolutional kernel parameters are dynamically adjusted to flexibly adapt to different shapes and scales of daylily targets, realizing efficient recognition and precise positioning of dense targets, and finally obtaining the segmentation results of different maturity levels of daylily; for the single-class segmentation branch, the feature information processed by the MBFE is fused with the feature information extracted by the backbone part and upsampled is performed, and finally the daylily picking area segmentation result is obtained; Finally, a dynamic loss weight strategy DLWS is designed; and the HCOA strategy is used to carry out the fast adaptive combination configuration of the high-dimensional hyperparameters of DPPL-MYOLO. At this time, the advanced model optimized by HCOA is named the ADPPL-MYOLO multi-task self-optimizing segmentation model, which further improves the reliability of the model in the daylily segmentation task.
3. The multi-task learning method for daylily segmentation and picking parameter determination for precise picking according to claim 2, characterized in that: Construct IMSM as follows: The feature map first enters the dynamic local perception module DLP, and the features are initially extracted through the depthwise separable dynamic convolution DSDConv, and through the residual connection, the input and output feature maps are fused to promote information flow and improve the model stability; Then it enters the multi-head global attention perception module MGAP. This module first initializes the input feature map X∈R through a 1×1 convolutional layer H×W×C to obtain X init in terms of the channel dimension; subsequently, the group convolution strategy is adopted to divide the channels of X init into four parts, obtaining four sub-feature maps {X1, X2, X3, X4}, where the number of channels of each sub-feature map is C / 4. These four sub-feature maps respectively enter four parallel branch networks for processing; the first branch consists of four depthwise separable dilated convolutions in series, aiming to capture the global features, and the mathematical representation is as follows: In the formula represents a network composed of four depthwise separable atrous convolutional layers, where each convolutional layer uses a 3×3 convolutional kernel and an atrous rate r = 3; the second branch contains two atrous convolutions with the same configuration to balance the extraction of global and local features, and the mathematical representation is as follows: In the formula represents a network composed of two depthwise separable atrous convolutional layers; the third branch contains only one atrous convolution, which focuses on capturing local features, and the mathematical expression is: In the formula represents a network that only contains one depthwise separable dilated convolutional layer; the fourth branch retains the original feature map X4 as a reference (X'4 = X4), providing the original unprocessed information; the outputs of the four processed branches are concatenated along the channel dimension to obtain the concatenated feature map X concat ; Subsequently, X concat is dimensionally adjusted through a 1×1 convolutional layer to match the number of channels of the feature map X init before grouping, obtaining the adjusted feature map X adjusted ; Finally, X adjusted and X init are multiplied element-wise to achieve deep fusion of features; through the interaction between the original features and the adjusted features, the information of the occluded part is restored and enhanced, obtaining the final feature map X final : X final = X adjusted ⊙X init In the formula, ⊙ represents the element-wise multiplication operation; After that, the features extracted by MGAP are transmitted to the FCAM module, and the number of channels is adjusted through 1×1 convolution, local features are captured through 3×3 convolution, the non-linear expression is enhanced through GELU activation, and the global information is integrated through the second 1×1 convolution. FCAM realizes the comprehensive extraction and efficient aggregation of image features based on DLP and MGAP.
4. The multi-task learning method for daylily segmentation and picking parameter determination for precise picking according to claim 2, wherein: Construct MBFE as follows: The feature information enters three BFA modules with different parameters connected in a cascaded manner. Among them, the output of the previous BFA module is used as the input of the next module and is passed in turn. Finally, the outputs of each module are concatenated to obtain a comprehensive output result; For each BFA module, first the feature map is initially extracted, normalized and activated through the CBS module; the SiLU activation function in the CBS module introduces non-linear characteristics: F CBS = SiLU(BN(W·F in + b)) Where W and b are the weights and biases of the convolutional kernel respectively; Subsequently, the BFA module uses four bar convolutional operation branches in different directions to comprehensively capture the bar features in the daylily image. The four branches respectively focus on feature extraction in the horizontal, vertical, and two diagonal directions; The first branch performs a bar convolution operation of 1×k with a dilation rate of d to obtain the feature map F1, which focuses on feature extraction in the horizontal direction; The second branch performs a bar convolution operation of k×1 with a dilation rate of d to obtain the feature map F2, which focuses on feature extraction in the vertical direction; The third branch first performs an image horizontal transformation, then performs a bar convolution operation of k×1 with a dilation rate of d, and finally performs an inverse horizontal transformation to obtain the feature map F3, which focuses on feature extraction in one diagonal direction; The fourth branch first performs an image vertical transformation, then performs a bar convolution operation of k×1 with a dilation rate of d, and finally performs an inverse vertical transformation to obtain the feature map F4, which focuses on feature extraction in the other diagonal direction; F1 = Conv 1×k (F CBR ; d) F2 = Conv k×1 (F CBR ; d) F3 = HFlip -1 (Conv k×1 (HFlip(F CBR )); d)) F4 = VFlip -1 (Conv k×1 (VFlip(F CBR )); d)) where HFlip represents a horizontal transformation operation, and HFlip -1 is its inverse transformation; where VFlip represents a vertical transformation operation, and VFlip -1 is its inverse transformation; after bar feature extraction in four directions, the BFA module merges the output feature maps of each branch through a feature concatenation operation to form a richer feature representation: F concat = Concat([F1,F2,F3,F4]) The concatenated feature map is processed by the CBS module again to further enhance the feature expression ability.
5. The method for determining the daylily segmentation and picking parameters through multi-task learning for precise picking according to claim 2, wherein: Construct the TSDS-Head detection head as follows: The TSDS-Head uses two 3×3 shared convolutional layers to perform preliminary feature extraction on the input image; subsequently, through a feature fusion operation, the outputs of the two shared convolutional layers are combined to obtain the feature map F; then, the feature map F undergoes an average pooling operation to obtain F avg , and then F and F avg are jointly used as the input to the task decomposition module TDC; in TDC, first the weight W is calculated, and then the weighted fusion is performed using the weight W and the feature map F to obtain the classification task feature map F cls and the localization task feature map F reg , and the specific process is as follows: F cls , F reg = σ2(BN((W⊙reshape(R))·reshape(F))) Among them, σ1 and σ2 represent activation functions, ⊙ represents element-wise multiplication, and R is a learnable reassignment matrix; F is further processed through a 3×3 convolutional layer and then through the spatial convolutional offset module offset to obtain the offset d and the mask m; at the same time, F obtains the classification weight map W through multi-stage feature processing cls , and then W cls and the classification task feature map F cls perform a multiplication operation and pass through a 3×3 convolutional layer to obtain the classification task result F′ cls ; in the localization task branch, the offset d, the mask m, and F reg are input into the dynamic deformable convolutional layer DyDCNv2; where K i is a dynamically generated convolutional kernel, p i is a sampling point on the feature map, d i is the offset, m i is the mask, and N is the total number of sampling points; after being processed by DyDCNv2, the feature map passes through a 3×3 convolutional layer and a scale adjustment scale operation to obtain the localization task result F′ reg ; finally, through a feature fusion operation, F′ cls and F′ reg are combined to obtain the output after task alignment processing.
6. The multi-task learning method for daylily segmentation and picking parameter determination for precise picking according to claim 2, wherein: Design the DLWS dynamic loss weight strategy as follows: DLWS assigns weights to the loss terms of multitask segmentation task A and single-class segmentation task B; two key variables w A and w B are introduced, which are used as the weighting coefficients of the loss term loss A of task A and the loss term loss B of task B respectively. The calculation strategy of w A reflects special attention to the loss term loss B of task B; where loss B is used as the numerator coefficient and divided by the sum of loss A and loss B , then the value of w A will change with the change of loss B relative to loss A ; by adding a constant term 1, it is ensured that w A is always greater than 1, so as to give the loss term loss A of task A a certain basic weight in the loss function and dynamically adjust it according to the loss of task B on this basis; Accordingly, w B is calculated according to the following formula, reflecting special attention to the loss term loss A of task A; where loss A is used as the numerator coefficient and divided by the sum of loss A and loss B ; then the value of w B will change with the change of loss A relative to loss B ; by adding a constant term 2, it is ensured that the minimum value of w B is 2, so as to give the loss term loss B of task B a higher base weight in the loss function and dynamically adjust it according to the loss of task A on this basis; Finally, combining the above calculated \(w\) A and \(w\) B , a comprehensive loss function \(loss\) in the following form is constructed: loss = w A ·loss A + w B ·loss B 。 7. The multi-task learning method for daylily segmentation and picking parameter determination for precise picking according to claim 2, wherein: Utilize the HCOA hyperparameter optimization strategy as follows: The optimization target of HCOA is 24 key hyperparameters in the DPPL-MYOLO model, including the initial learning rate, cyclic learning, learning rate momentum, weight decay coefficient, warm-up learning epochs, warm-up learning momentum, warm-up initial learning rate, giou loss coefficient, classification loss coefficient, dfl loss coefficient, masks downsampling ratio, and 13 data augmentation coefficients; By adaptively adjusting the model hyperparameters during network training and giving the best hyperparameter components that maximize the fitness function value, it is used to establish the final ADPPL-MYOLO model; In the HCOA strategy, the fitness function value is used as the basis for evaluating the quality of individuals in the population, thereby guiding the evolution direction of the population. Define the fitness function, which is calculated based on the performance of the ADPPL-MYOLO model in the daylily segmentation task: fitness = 0.1·mAP@50 + 0.9·mAP@50:95 The fitness value obtained by the above formula is used as the fitness value of HCOA for population update, and the optimal individual information will be used to establish the final model.
8. The multi-task learning method for daylily segmentation and picking parameter determination for precise picking according to claim 2, characterized in that: Design the Daylily-DCSL daylily picking point location algorithm as follows: Using the constructed ADPPL-MYOLO, the daylily segmentation task branch can generate a segmentation mask map for each target in the image, denoted as maskA i , where i = 1, 2, 3,..., n, and n represents the total number of targets in the image; in parallel, the picking area segmentation task branch is responsible for drawing the segmentation mask map of the picking area, denoted as maskB; the generated segmentation mask map is finely processed: first, filter out the maskA corresponding to the categories of mature and over-mature daylilies i , then, these filtered maskA i are respectively subjected to an intersection operation with the corresponding picking area masks in maskB: Intersection i = maskA i Ι maskB Among them, Intersection i represents the common area between the i-th mature or over-mature daylily target and the picking area; this intersection operation precisely delimits the overlapping part of each target object and the picking area; finally, through the precise calculation of the center points of these intersection areas, the precise identification of the daylily picking points is achieved; where (x k , y k ) are the coordinates of non-zero pixels.
9. The multi-task learning method for daylily segmentation and picking parameter determination for precise picking according to claim 2, wherein: Design the RSEAD-CA daylily cutting angle determination algorithm as follows: First, use the constructed ADPPL-MYOLO to generate a segmentation mask map for each target of daylily in the image, denoted as maskA i , where i = 1, 2, 3,..., n, n represents the total number of targets in the image, and the segmentation mask map of the picking area, denoted as maskB; Subsequently, screen the maskA corresponding to the categories of mature and over-mature daylilies i and perform a mask intersection operation with maskB to obtain the target picking area; The formal representation of this intersection operation is: Where S represents the index set of the mature and over-mature daylily categories; Subsequently, calculate the intersection region R target to obtain its minimum circumscribed rectangle, and determine the deflection angle of its shortest side in the horizontal direction as the cutting angle of the scissors of the daylily cutting device; the mathematical expression of this process is as follows: θ = AngleOfShortestSide(R bounding_box (R target )) where R bounding_box (R target ) represents the minimum bounding rectangle for calculating the intersection region R target , and the AngleOfShortestSide(·) function is used to calculate the deflection angle θ of the shortest side of this rectangle with respect to the horizontal direction.
10. The multi-task learning method for daylily segmentation and picking parameter determination for precise picking according to claim 2, characterized in that: The model backbone structure Backbone accepts daylily images of 640×640.