Three-degree-of-freedom broccoli picking method and broccoli picking device

By introducing the YOLOv5 object detection model and adaptive spatial feature fusion technology, combined with geometric center calculation and depth camera, the efficient and low-damage picking process of broccoli picking machinery is achieved, solving the problems of low efficiency of traditional picking machinery and serious flower ball damage.

CN120036132APending Publication Date: 2025-05-27ZHEJIANG SCI-TECH UNIV

Patent Information

Application Number
CN202510092133.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing broccoli picking machinery has problems such as low picking efficiency, serious damage to the flower ball, and difficult equipment to adapt to the special growth mode of broccoli.

Method used

The three-degree of freedom broccoli picking method is adopted, combined with the YOLOv5 object detection model, adaptive spatial feature fusion (ASFF) and convolutional block attention module (CBAM), and accurate three-dimensional spatial positioning is provided through geometric center calculation and ZED2 depth camera to achieve accurate operation of the robot arm.

Benefits of technology

It significantly improves the recognition accuracy of broccoli flower balls, reduces flower ball damage, and improves picking efficiency and product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120036132A_ABST
    Figure CN120036132A_ABST
Patent Text Reader

Abstract

The invention discloses a three-degree-of-freedom broccoli picking method and a broccoli picking device, and relates to the technical field of broccoli picking, and the three-degree-of-freedom broccoli picking method comprises the following steps: S1, moving a mobile chassis to a to-be-picked place according to an instruction issued by a controller, and enabling a picking mechanical arm to enter a broccoli picking range; and S2, capturing depth and RGB image data of the surrounding environment by using a ZED2 depth camera. According to the method, YOLOv5 target detection, ASFF and CBAM modules are combined, the broccoli ball recognition precision is improved, and the problem of image feature inconsistency is solved. Through geometric center calculation and a ZED2 depth camera, accurate three-dimensional space positioning is provided, and high-precision target positioning is ensured. By adopting the three-degree-of-freedom mechanical arm design, the control algorithm and the structure are simplified, the picking efficiency is improved, the ball-flower damage is reduced, and the picking quality and efficiency are ensured, so that efficient and low-damage broccoli picking is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of broccoli picking, and specifically relates to a three-degree-of-freedom broccoli picking method and its picking device. Background Art

[0002] As an important agricultural crop, broccoli is widely cultivated and consumed due to its rich nutritional components. However, during the broccoli picking process, especially in the mature stage, the traditional manual picking method not only has a high labor intensity, but also has problems such as low picking efficiency and serious damage to the flower heads. Since the broccoli flower heads are relatively fragile, manual picking is prone to causing damage, which affects the product quality and market value. Although the mechanization level of vegetable planting in China has been improved in recent years, in the field of broccoli mechanized picking, there are still problems such as immature picking technology and difficulty for equipment to adapt to the special growth mode of broccoli. Therefore, how to improve the picking efficiency of broccoli and reduce the damage rate has become an urgent technical problem to be solved in the field of agricultural mechanization.

[0003] Existing broccoli picking machines mainly include two types: clamping and cutting two-stage manipulators and articulated robotic arms. The clamping and cutting two-stage manipulator needs to first clamp the broccoli flower head and then cut the root, which is a complex process and is prone to damaging the flower head, resulting in low efficiency. Although the articulated robotic arm can operate flexibly, it usually relies on rotating shafts and often has large coupling problems during the picking process, leading to large computational amounts and low working efficiency. On the other hand, most of the existing broccoli target detection technologies are based on feature pyramids for target detection. However, between multi-level feature maps, if an object in a certain level of the feature map is mislabeled as a positive sample, the other levels of the feature maps often regard this area as the background, resulting in inconsistencies between the feature maps, which in turn affects the gradient calculation during the training process and reduces the detection accuracy. These deficiencies in the existing technologies limit the progress of broccoli picking technology, and there is an urgent need for a new technical solution to solve these problems.

[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The object of the present invention is to provide a three-degree-of-freedom broccoli picking method and its picking device. By introducing the YOLOv5 object detection model and combining adaptive spatial feature fusion (ASFF) and convolutional block attention module (CBAM), the recognition accuracy of broccoli flower heads is significantly improved, the problem of inconsistent image features is solved, and background noise interference is reduced. Combining geometric center calculation and ZED2 depth camera provides accurate three-dimensional spatial positioning, ensuring high-precision positioning of broccoli targets and providing reliable data for the precise operation of the robotic arm. In addition, the present invention adopts a three-degree-of-freedom robotic arm design, simplifies the control algorithm and structure, improves the picking efficiency, reduces flower head damage, and ensures the picking quality and efficiency, thus realizing an efficient and low-damage broccoli picking process to solve the problems in the above-mentioned background technology.

[0006] To achieve the above object, the present invention provides the following technical solution: A three-degree-of-freedom broccoli picking method and its picking device, including the following steps:

[0007] Step S1, the mobile chassis moves to the location to be picked according to the instructions issued by the controller, so that the picking robotic arm enters the broccoli picking range;

[0008] Step S2, use the ZED2 depth camera to capture the depth and RGB image data of the surrounding environment, and preprocess the RGB image through the image processing module, extract the depth values of each pixel point in the depth image, and transmit the processed RGB image and depth information to the computer main control unit;

[0009] Step S3, the computer main control unit receives and processes the RGB image, uses the improved YOLOv5 broccoli object detection model to detect the image through adaptive spatial feature fusion (ASFF) and convolutional block attention module (CBAM), and identifies the mature broccoli flower head and determines its geometric center position;

[0010] Step S4, according to the geometric center position coordinates output by the object detection model, control the picking robotic arm to perform precise up-and-down movement to wrap the broccoli flower head and clamp it;

[0011] Step S5, after the picking robotic arm completes the clamping of the flower head, automatically complete the root cutting and put the picked broccoli flower head into the storage box on the mobile chassis;

[0012] Step S6, the mobile chassis receives a new target position instruction, and repeats steps S4 to S5 until all picking tasks are completed.

[0013] Preferably, in step S3, the computer main control unit identifies the mature broccoli flower head part through the broccoli object detection model, and calculates the position of the target center according to the rectangular box drawn by the broccoli object detection model. The specific process is as follows:

[0014] Step S31: The dataset collects broccoli images under three different weather conditions: cloudy, overcast, and sunny. A total of 4,864 images are collected. After screening, 2,348 clear images are selected for labeling. During data augmentation, image transformation is performed by vertical flipping 20% and horizontal flipping 30%. Then, the image brightness is adjusted by multiplying the pixel values by a specific value (such as 1.2 or 1.5). In addition, the Gaussian blur algorithm (σ ∈ [0 - 1.5]) is used to smooth the images, and the degree of blur is adjusted according to different situations. Finally, the images are translated (2 - 3 pixels) in the x / y axis direction and scaled to 50% - 70%, resulting in 14,088 enhanced images for training the broccoli target detection model, thereby improving the robustness of the model.

[0015] Preferably, it further includes:

[0016] Step S32: Based on the YOLOv5 model, the Neck part is improved using Adaptive Spatial Feature Fusion (ASFF), and a Convolutional Block Attention Module (CBAM) is added to address the inconsistency of image features and improve the recognition and detection effect.

[0017] Preferably, in Step S32, based on the YOLOv5 neural network, the Neck part is improved using the Adaptive Spatial Feature Fusion (ASFF) technique, and a Convolutional Block Attention Module (CBAM) is added to solve the problem of inconsistent image features. ASFF dynamically adjusts the weighting coefficients of feature maps at different layers, enabling feature maps from different scales to be fused according to their importance for target detection. The working method of ASFF is as follows:

[0018] For each level of feature map, all other levels of feature maps are adjusted to the same shape and spatially fused according to the learned weight map to ensure that feature information at different scales is effectively utilized. The specific algorithm is:

[0019]

[0020] Where, represents the output feature value of the l-th layer feature map. Specifically, it is the value of the l-th layer feature map at position i, j, which is obtained by weighted fusion of feature map information from different layers. These represent the values of different source feature maps input to the l-th layer feature map. Specifically, is the input feature map value from the l-th layer, comes from the 2nd layer, and is the feature map value from the 3rd layer. These input feature maps in the formula are weighted and fused to generate the final output feature map. is the weighting coefficient, which controls the contribution degree of each layer of feature maps in the final output and is learned during the training process. By adjusting these coefficients, the model can learn the proportion that different levels of feature maps should account for in the final prediction. i and j represent the position indices of pixels in the feature map, where i and j are the row and column indices of the feature map respectively. Each x in the formula ij refers to the feature value at a specific position, and l represents the level where the current feature map is located. This level is different stages in the network for processing input data, and different levels are responsible for extracting different levels of feature information from the data. For example, low-level feature maps may capture basic features such as edges or textures, while high-level feature maps may capture more complex object shapes or structures.

[0021] Preferably, in step S32, ASFF uses the normalized exponential function to process the weighting coefficients to ensure that the sum of the weighting coefficients is 1 and the coefficient values are between [0, 1]. The specific algorithm is as follows:

[0022]

[0023] This formula ensures that the weighting coefficients of each feature map during spatial fusion do not exceed 1, while ensuring the balance of weighting and avoiding the excessive influence of a certain layer of feature map. This optimization process significantly improves the fusion effect of image features at different scales and enhances the accuracy of broccoli target detection.

[0024] Preferably, in step S32, through adaptive spatial feature fusion, ASFF will weight each layer of feature maps according to the learned weights, and calculate the weighting coefficients of each layer through the following formula:

[0025]

[0026] This formula calculates the weighting coefficients of each feature map through the normalized exponential function and automatically adjusts these weights through the training process, thereby optimizing the performance of the ASFF module in fusing features at different scales.

[0027] Preferably, it further includes:

[0028] Step S33: Train the broccoli target detection model to generate a network file;

[0029] Step S34: Deploy the generated network file on the computer main control unit;

[0030] Step S35: Input the received preprocessed RGB image into the broccoli target detection model;

[0031] Step S36: The broccoli target detection model identifies mature and pickable broccoli based on the received image, and frames the broccoli flower head part with a rectangular box;

[0032] Step S37: After identifying the mature broccoli, calculate the position of the target center according to the rectangular box drawn by the broccoli target detection model. The diagonal coordinate points of the rectangular box are n 1 (x 1 , y 1 ) and n 2 (x 2 , y 2 ). Therefore, the geometric center point of the broccoli rectangular box, such as the system of equations:

[0033]

[0034] Thus, the coordinates of the geometric center coordinate o in the broccoli flower head of the image are obtained as (x o , y o ). Then, obtain the depth value D of the camera in the image coordinate system through the ZED2 camera, and convert the image coordinate system to the camera coordinate (x c , y c ). In this way, the spatial position of the flower head in the camera coordinate system is obtained as (x c , y c , D). Then, send the spatial position (x c , y c , D) information to the picking robotic arm through the network interface to complete the positioning task.

[0035] Preferably, in step S33, by training the YOLOv5 target detection model, a network file optimized by Adaptive Spatial Feature Fusion (ASFF) and Convolutional Block Attention Module (CBAM) is generated. During the training process, the cross-entropy loss function and the IoU (Intersection over Union) loss function are used as the optimization objectives to ensure the accuracy of the model in the target detection task. After the training is completed, the generated network file is deployed to the computer main control unit to process the image data in real time and perform target recognition;

[0036] In step S37, the spatial position (x c , y c , D) information is sent to the picking robotic arm through the network interface. The picking robotic arm adjusts its movement trajectory according to the received spatial position information, thereby completing the precise picking of broccoli. During this process, the picking robotic arm grabs the broccoli flower head through an accurate straight up and down movement path, and ensures that the picking process is fast and accurate, reducing flower head damage and improving the picking efficiency.

[0037] Preferably, in step S32, the Convolutional Block Attention Module (CBAM) processes the feature map through a two-dimensional attention mechanism. On the given intermediate feature map F, CBAM first performs the channel attention mechanism M C (F), and then performs an element-wise multiplication operation on the processed feature map and the original feature map F to obtain the weighted feature map F'. Next, the spatial attention mechanism M S (F') is applied to the feature map F', and finally the final weighted feature map F'' is obtained through an element-wise multiplication operation. This feature map distributes attention through two dimensions to strengthen the model's attention to key features:

[0038]

[0039] wherein, denotes the element-wise multiplication operation. The channel attention and spatial attention increase the response of the feature map to important regions in this way, thereby enhancing the performance of broccoli target detection.

[0040] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0041] By introducing the YOLOv5 target detection model and combining it with the Adaptive Spatial Feature Fusion (ASFF) and the Convolutional Block Attention Module (CBAM), the present invention significantly improves the recognition accuracy of broccoli flower heads. The ASFF module can dynamically adjust the weights of feature maps at different levels, enabling effective cooperation of feature maps from different scales, thereby improving the accuracy of multi-scale target detection. In addition, the CBAM module enables the model to focus on regions crucial for target detection through channel and spatial attention mechanisms, reducing interference from background noise. In this way, the present invention solves the problem of inaccurate recognition caused by inconsistent image features in traditional target detection methods, ensures accurate recognition of broccoli in complex backgrounds, and improves the efficiency and success rate of broccoli picking.

[0042] By combining geometric center calculation with depth camera information, the present invention provides precise spatial positioning for the broccoli picking robotic arm. The geometric center calculation obtains the center position of the broccoli flower head by calculating the diagonal coordinate points of the rectangular box output by the target detection model, thereby obtaining its precise position in the image coordinate system. Further, by combining the depth value obtained by the ZED2 depth camera, the image coordinate system is converted into the camera coordinate system, and finally the position coordinates of the target in three-dimensional space are obtained. This innovative method effectively solves the problem of inaccurate positioning caused by the lack of depth information in traditional technologies, ensures high-precision positioning of the broccoli target, provides reliable spatial coordinate information for the precise operation of the robotic arm, and ensures the smooth completion of the picking task.

[0043] The present invention adopts a three - degree - of - freedom robotic arm design. By reducing complex degrees of freedom and control algorithms, the structure of the robotic arm is simplified, and the picking efficiency is improved. Compared with traditional multi - degree - of - freedom robotic arms, the three - degree - of - freedom design makes the operation of the robotic arm more direct and efficient. Only simple up - and - down movements are required to complete the precise picking of broccoli. This not only reduces the control complexity of the robotic arm but also decreases the computational amount and avoids errors caused by complex movements. In addition, the simplified motion trajectory design increases the picking speed and reduces the potential risk of damage to the flower heads caused by the complexity of the robotic arm movement, thus ensuring the picking quality and efficiency of broccoli. Through this innovative design, the present invention realizes an efficient and low - damage broccoli picking process. Brief Description of the Drawings

[0044] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0045] Figure 1 It is a flowchart of a method for picking broccoli with three degrees of freedom according to the present invention.

[0046] Figure 2 It is a schematic diagram of the working process of the Adaptive Spatial Feature Fusion (ASFF) module according to the present invention. Detailed Embodiments

[0047] Now, the exemplary embodiments will be described more comprehensively with reference to the drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these exemplary embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art.

[0048] The present invention provides a method for picking broccoli with three degrees of freedom as Figure 1 shown, including the following steps:

[0049] Step S1: The mobile chassis moves to the location to be picked according to the instructions issued by the controller, so that the picking robotic arm enters the broccoli picking range.

[0050] Step S2: Use the ZED2 depth camera to capture the depth and RGB image data of the surrounding environment, and pre - process the RGB image through the image processing module to extract the depth values of each pixel point in the depth image. The processed RGB image and depth information are transmitted to the computer main control unit.

[0051] Step S3: The computer main control unit receives and processes the RGB image. Using the improved YOLOv5 broccoli target detection model, it detects the image through adaptive spatial feature fusion (ASFF) and convolutional block attention module (CBAM), identifies the mature broccoli flower heads, and determines their geometric center positions.

[0052] Step S4: According to the geometric center position coordinates output by the target detection model, control the picking robotic arm to perform precise up-and-down motion to wrap and clamp the broccoli flower heads.

[0053] Step S6: After the picking robotic arm completes clamping the flower heads, it automatically completes the cutting of the root and stem and places the picked broccoli flower heads into the storage box on the mobile chassis.

[0054] Step S6: The mobile chassis receives a new target position instruction and repeats steps S4 to S6 until all picking tasks are completed.

[0055] In step S3, the computer main control unit identifies the mature broccoli flower head part through the broccoli target detection model, and calculates the position of the target center according to the rectangular box drawn by the broccoli target detection model. The specific process is as follows:

[0056] Step S31: The data set collects broccoli images under three different weather conditions: cloudy, overcast, and sunny. A total of 4,864 images are collected. After screening, 2,348 clear images are selected for marking. During the data augmentation process, the images are transformed by vertical flipping 20% and horizontal flipping 30%. Then, the image brightness is adjusted by multiplying the pixel values by a specific value (such as 1.2 or 1.5). In addition, the Gaussian blur algorithm (σ ∈ [0 - 1.5]) is used to smooth the images, and the degree of blur is adjusted according to different situations. Finally, the images are translated (2 - 3 pixels) in the x / y axis direction and scaled to 50% - 70%, and finally 14,088 enhanced images are obtained for the training of the broccoli target detection model, thereby improving the robustness of the model.

[0057] It also includes: Step S32: Based on the YOLOv5 model, use adaptive spatial feature fusion (ASFF) to improve the Neck part and add a convolutional block attention module (CBAM) to solve the inconsistency of image features and improve the recognition and detection effect.

[0058] In step S32, based on the YOLOv5 neural network, use the adaptive spatial feature fusion (ASFF) technology to improve the Neck part and add a convolutional block attention module (CBAM) to solve the problem of inconsistent image features. The working method of ASFF is as follows:

[0059] For each level of feature maps, the feature maps of all other levels are adjusted to the same shape and spatially fused according to the learned weight maps to ensure the effective utilization of feature information at different scales. The specific algorithm is as follows:

[0060]

[0061] Among them, represents the output feature value of the feature map at the l-th layer. Specifically, it is the value of the feature map at the l-th layer at positions i and j. This feature value is obtained by weighted fusion of feature map information from different layers. These represent the values of the feature maps from different sources that are input to the feature map at the l-th layer. Specifically, is the value of the input feature map from the l-th layer, comes from the second layer, is the value of the feature map from the third layer. These input feature maps in the formula are weighted and fused to generate the final output feature map. is the weighting coefficient, which controls the contribution degree of each layer of feature map in the final output and is learned during the training process. By adjusting these coefficients, the model can learn the proportion that different levels of feature maps should occupy in the final prediction. i and j represent the position indices of the pixels in the feature map, and i and j are the row and column indices of the feature map respectively. Each x in the formula ij refers to the feature value at a specific position, and l represents the level where the current feature map is located. This level is different stages in the network for processing input data, and different levels are responsible for extracting different levels of feature information from the data. For example, low-level feature maps may capture basic features such as edges or textures, while high-level feature maps may capture more complex object shapes or structures.

[0062] In step S32, ASFF processes the weighting coefficients using the normalized exponential function to ensure that the sum of the weighting coefficients is 1 and the coefficient values are between [0, 1]. The specific algorithm is as follows:

[0063]

[0064] This formula ensures that the weighting coefficients of each feature map during spatial fusion do not exceed 1, while ensuring the balance of weighting and avoiding the excessive influence of a certain layer of feature map. This optimization process significantly improves the fusion effect of image features at different scales and enhances the accuracy of broccoli target detection.

[0065] In step S32, through adaptive spatial feature fusion, ASFF weights the feature maps of each layer according to the learned weights and calculates the weighting coefficients of each layer through the following formula:

[0066]

[0067] This formula calculates the weighting coefficients for each feature map through a normalized exponential function and automatically adjusts these weights during the training process, thereby optimizing the performance of the ASFF module in feature fusion at different scales.

[0068] It also includes:

[0069] Step S33: Train the broccoli target detection model to generate a network file;

[0070] Step S34: Deploy the generated network file on the computer main control unit;

[0071] Step S35: Input the received preprocessed RGB image into the broccoli target detection model;

[0072] Step S36: The broccoli target detection model identifies the mature and pickable broccoli according to the received image and frames the broccoli flower ball part with a rectangular box;

[0073] Step S37: After identifying the mature broccoli, calculate the position of the target center according to the rectangular box drawn by the broccoli target detection model. The diagonal coordinate points of the rectangular box are n 1 (x 1 , y 1 ) and n 2 (x 2 , y 2 ). Therefore, the geometric center point of the broccoli rectangular box, such as the system of equations:

[0074]

[0075] , from which the coordinates of the geometric center coordinate o of the broccoli flower ball in the image are obtained as (x o , y o ). Then, obtain the depth value D of the camera in the image coordinate system through the ZED2 camera, and convert the image coordinate system to the camera coordinate (x c , y c ), so as to obtain the spatial position of the flower ball in the camera coordinate system (x c , y c , D). Then, send the spatial position (x c , y c , D) information to the picking robotic arm through the network interface to complete the positioning task.

[0076] In step S33, by training the YOLOv5 object detection model, a network file optimized by Adaptive Spatial Feature Fusion (ASFF) and Convolutional Block Attention Module (CBAM) is generated. During the training process, the cross-entropy loss function and the IoU (Intersection over Union) loss function are used as optimization objectives, ensuring the accuracy of the model in object detection tasks. After training, the generated network file is deployed to the computer main control unit to process image data in real time and perform object recognition; in step S37, the spatial position (x c , y c , D) information is sent to the picking robotic arm through the network interface, and the picking robotic arm adjusts its motion trajectory according to the received spatial position information, thus completing the precise picking of broccoli. During this process, the picking robotic arm grabs the broccoli flower head through an accurate up-and-down motion path, and ensures that the picking process is fast and accurate, reducing flower head damage and improving picking efficiency.

[0077] In step S32, the Convolutional Block Attention Module (CBAM) performs two-dimensional attention mechanism processing on the feature map. On the given intermediate feature map F, CBAM first executes the channel attention mechanism M C (F), and then performs an element-wise multiplication operation on the processed feature map and the original feature map F to obtain the weighted feature map F'. Next, the spatial attention mechanism M S (F') is applied to the feature map F', and finally the final weighted feature map F'' is obtained through an element-wise multiplication operation. This feature map distributes attention through two dimensions to strengthen the model's attention to key features:

[0078]

[0079] , where represents the element-wise multiplication operation. The channel attention and spatial attention increase the response of the feature map to important regions in this way, thereby enhancing the performance of broccoli object detection.

[0080] Specific Embodiment 1: Broccoli Object Detection Based on the YOLOv5 Model and Adaptive Spatial Feature Fusion (ASFF);

[0081] In this embodiment, the YOLOv5 object detection model is adopted and improved on this basis, combined with the Adaptive Spatial Feature Fusion (ASFF) technology and the Convolutional Block Attention Module (CBAM) to improve the object recognition accuracy and robustness of the broccoli picking manipulator. During the broccoli picking process, accurate object detection is the key to ensuring efficient picking, and the introduction of ASFF and CBAM is precisely to solve the problem of inconsistent image feature processing.

[0082] First, in the data collection stage, the model is trained with a set of broccoli images under different weather conditions. Specifically, broccoli images are collected under cloudy, overcast, and sunny conditions, and a diverse image dataset is generated through data augmentation. Data augmentation techniques mainly include operations such as image flipping, brightness adjustment, Gaussian blur, translation, and scaling. The purpose of these augmentation operations is to enable the model to adapt to changes under different lighting and environmental conditions and improve its robustness in the actual picking environment.

[0083] When performing object detection, first, the augmented data is input into the YOLOv5 model. YOLOv5 itself is an efficient and accurate object detection algorithm. However, in the traditional YOLOv5 model, when processing multi-scale feature maps, there may be inconsistencies in the fusion of low-level and high-level feature maps, which can affect the final object detection accuracy. To address this issue, the Adaptive Spatial Feature Fusion (ASFF) technique is introduced. In the ASFF module, by weighting feature maps of different scales, it can dynamically select which features are most important for the current object, thereby automatically adjusting the fusion method of the feature maps. This process enables feature maps from different levels to cooperate effectively, enhancing the model's detection ability for multi-scale broccoli flower heads.

[0084] The working principle of ASFF is to add the ASFF module to the Neck part of the model, adjust the sizes of feature maps of each layer to be the same, and then weight them according to the contribution value of each layer. Specifically, ASFF weights each feature map with a normalized weighting coefficient, and the coefficient value is dynamically adjusted according to the importance of each feature map. During the training process, the ASFF module learns how to optimally weight each feature map to generate a high-quality fused feature map. In this way, the model can automatically adjust the feature map fusion strategy according to different input images, optimizing the accuracy of object detection.

[0085] At the same time, the Convolutional Block Attention Module (CBAM) is also introduced to further enhance the model's attention ability. The CBAM module performs attention mechanism processing in the channel dimension and spatial dimension respectively to ensure that the model can focus on areas crucial for object detection. Specifically, the CBAM module consists of two parts: the Channel Attention Module (CA) and the Spatial Attention Module (SA). The Channel Attention Module is used to weight the channels of the feature map to help the model determine which channel features are more important; the Spatial Attention Module weights according to the spatial dimension, enabling the model to focus on the areas decisive for classification in the image and ignore other irrelevant areas. In this way, the CBAM module can significantly improve the YOLOv5 model's recognition ability for broccoli flower heads, especially in complex backgrounds, effectively avoiding false detections and missed detections.

[0086] Through the above method, in the process of broccoli target detection in this embodiment, the recognition accuracy of the model for broccoli targets under complex environments and different weather conditions can be significantly improved. By combining the characteristics of ASFF and CBAM, the deficiencies of the traditional YOLOv5 model in processing multi-scale features are successfully solved, ensuring that the broccoli picking manipulator can efficiently and accurately identify targets in various environments.

[0087] Specific Embodiment 2: Geometric center calculation and fusion positioning with depth camera;

[0088] In this embodiment, the combination of geometric center calculation and depth camera information is introduced for precise positioning of targets during broccoli picking. During the broccoli picking process, accurate positioning is the key to ensuring efficient picking by the robotic arm. Traditional target detection methods may not be able to accurately determine the actual position of the target, especially in complex environments such as dense plant growth or cluttered backgrounds, where the positioning problem is particularly prominent. Therefore, in this embodiment, by combining the depth camera and geometric center calculation, the position of the broccoli flower head is accurately obtained, providing high-precision operation guidance for the robotic arm.

[0089] First, during the broccoli target detection process, after the model is optimized by YOLOv5 and ASFF+CBAM, it can identify the position of the broccoli flower head and draw a rectangular box. Based on the rectangular box output by the target detection model, the geometric center position of the rectangular box is further calculated. The diagonal coordinates of the rectangular box are n 1 (x 1 , y 1 ) and n 2 (x 2 , y 2 ). The coordinates of the geometric center point are calculated by the following formula:

[0090]

[0091] This geometric center point represents the precise position of the broccoli flower head in the image. However, simply calculating the position in the image coordinate system is not sufficient to provide enough spatial information for the picking robotic arm. Especially in three-dimensional space, the picking robotic arm needs to consider the depth information of the target. Therefore, next, the depth value D of the camera in the image coordinate system is obtained through the ZED2 depth camera, and it is combined with the geometric center coordinates to be converted into the spatial coordinates (x c , y c , D) in the camera coordinate system. This conversion process is achieved through camera calibration and depth data acquisition, and the depth value and coordinate information are transmitted to the computer main control unit through the network interface, thereby obtaining the precise position of the flower head in three-dimensional space.

[0092] In this way, the depth information not only helps to accurately locate the position of the broccoli in the two-dimensional image, but also enables the robotic arm to better understand the actual spatial position of the target. By combining the geometric center and the depth data, the robotic arm can adjust its motion path in real time to ensure that it can directly pick the broccoli head. The introduction of depth information significantly improves the accuracy of target positioning and solves the problem of inaccurate positioning caused by the lack of depth information in traditional methods.

[0093] The advantage of this method is that by combining the information of the geometric center and the depth camera, it ensures that the robotic arm can quickly and accurately grasp the target during the broccoli picking process, avoiding the problems of picking failure or head damage caused by positioning deviation in traditional methods. At the same time, the robotic arm no longer relies on complex visual calculations, but directly uses the depth data for spatial positioning, greatly improving the efficiency and stability of the operation.

[0094] Specific implementation method 3: combination of a three-degree-of-freedom robotic arm and precise picking;

[0095] In this implementation method, the design of the three-degree-of-freedom robotic arm is combined with the target positioning technology to form a simplified and efficient broccoli picking system. Traditional picking robotic arms usually adopt a multi-degree-of-freedom design. Although this design can provide high flexibility, in practical applications, complex degrees of freedom often bring high computational complexity and control difficulty. To simplify the picking process, a three-degree-of-freedom robotic arm is adopted, and precise picking of broccoli is achieved through precise linear motion.

[0096] First of all, the design of the three-degree-of-freedom robotic arm enables it to perform precise up-and-down motion in the vertical direction without complex rotation operations. This design greatly simplifies the control system of the robotic arm and reduces complex computational complexity. During the broccoli picking process, the computer main control unit transfers the position of the broccoli head to the robotic arm control system by receiving the target positioning information in real time. According to the geometric center coordinates and depth information of the target, the robotic arm can quickly adjust its motion trajectory, accurately wrap the broccoli head and perform the picking operation.

[0097] The advantages of this three-degree-of-freedom robotic arm are its simple structure and precise motion. The robotic arm only needs to perform simple linear motion without complex rotation or bending operations, which not only improves the picking efficiency but also avoids errors caused by complex operations. In addition, by combining the aforementioned target detection and positioning technologies, the robotic arm can ensure the accuracy of each picking and reduce the damage to the broccoli head. This design is particularly suitable for easily damaged vegetables like broccoli and can ensure that the picking process is both efficient and safe.

[0098] In summary, through the combination of a three-degree-of-freedom robotic arm and a precise target positioning system, this embodiment not only improves the picking efficiency but also ensures the stability during the picking process and the integrity of the broccoli florets. Through a streamlined motion trajectory and high-precision positioning, this system has successfully solved the control complexity and efficiency problems faced by traditional multi-degree-of-freedom robotic arms in practical applications.

[0099] The present invention significantly improves the recognition accuracy of broccoli florets by introducing the YOLOv5 object detection model and combining it with Adaptive Spatial Feature Fusion (ASFF) and Convolutional Block Attention Module (CBAM). The ASFF module can dynamically adjust the weights of feature maps at different levels, enabling effective cooperation between feature maps from different scales, thereby enhancing the accuracy of multi-scale object detection. In addition, the CBAM module, through channel and spatial attention mechanisms, enables the model to focus on regions crucial for object detection, reducing the interference of background noise. In this way, the present invention solves the problem of inaccurate recognition caused by inconsistent image features in traditional object detection methods, ensuring accurate recognition of broccoli in complex backgrounds and improving the efficiency and success rate of broccoli picking.

[0100] The present invention provides precise spatial positioning for the broccoli picking robotic arm by combining geometric center calculation and depth camera information. Geometric center calculation obtains the center position of the broccoli floret by calculating the diagonal coordinate points of the rectangular box output by the object detection model, thereby obtaining its precise position in the image coordinate system. Further, by combining the depth values obtained from the ZED2 depth camera, the image coordinate system is converted into the camera coordinate system, and finally the position coordinates of the target in three-dimensional space are obtained. This innovative method effectively solves the problem of inaccurate positioning caused by the lack of depth information in traditional technologies, ensures high-precision positioning of the broccoli target, provides reliable spatial coordinate information for the precise operation of the robotic arm, and ensures the successful completion of the picking task.

[0101] The present invention adopts a three-degree-of-freedom robotic arm design. By reducing complex degrees of freedom and control algorithms, the structure of the robotic arm is simplified and the picking efficiency is improved. Compared with traditional multi-degree-of-freedom robotic arms, the three-degree-of-freedom design makes the operation of the robotic arm more direct and efficient, and precise picking of broccoli can be completed only by simple up and down movements. This not only reduces the control complexity of the robotic arm but also reduces the computational amount and avoids errors caused by complex movements. In addition, the streamlined motion trajectory design improves the picking speed and reduces the potential risk of floret damage caused by the complexity of the robotic arm movement, thus ensuring the picking quality and efficiency of broccoli. Through this innovative design, the present invention realizes an efficient and low-damage broccoli picking process.

[0102] The above formulas are all dimensionless and only take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0103] Only some exemplary embodiments of the present invention have been described above by way of illustration. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of the claims of the present invention.

[0104] It should be noted that in this text, if there are relational terms such as first and second, they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0105] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0106] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this text can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0107] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0108] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0109] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0110] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0111] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0112] As described above, this is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

[0113] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A three-degree-of-freedom broccoli picking method, characterized in that: The following steps are involved: Step S1, the mobile chassis moves to the location to be picked according to the instruction issued by the controller, so that the picking robot arm enters the broccoli picking range; Step S2, using the ZED2 depth camera to capture the depth and RGB image data of the surrounding environment, and pre-processing the RGB image through the image processing module to extract the depth value of each pixel in the depth image, and the processed RGB image and depth information are transmitted to the computer main control unit; Step S3, the computer main control unit receives and processes the RGB image, uses the improved YOLOv5 broccoli object detection model, detects the image through adaptive spatial feature fusion and convolutional block attention module, identifies the mature broccoli head and determines its geometric center position; Step S4, according to the geometric center position coordinates output by the target detection model, controlling the picking robot arm to move up and down accurately to wrap the broccoli florets and clamp them; Step S5: After the picking robot arm completes the flower bulb gripping, it automatically completes the rhizome cutting and puts the picked broccoli flower bulb into the storage box on the mobile chassis; Step S6: The mobile chassis receives the new target position instruction and repeats steps S4 to S5 until all picking tasks are completed.

2. A three-degree-of-freedom broccoli picking method according to claim 1, characterized in that: In step S3, the computer main control unit identifies the mature broccoli head part through the broccoli target detection model, and calculates the position of the target center according to the rectangular frame drawn by the broccoli target detection model. The specific process is as follows: Step S31, the data set collects broccoli images under three different weather conditions: cloudy, overcast and sunny, including a total of 4864 images. After screening, 2348 clear images are selected for marking; in the data enhancement process, the image is transformed by flipping vertically by 20% and flipping left and right by 30%, and then the image brightness is adjusted, and the pixel value is multiplied by a fixed value; in addition, the image is smoothed using a Gaussian blur algorithm, and the degree of blur is adjusted according to different situations; finally, the image is translated in the x / y axis direction and scaled to 50% to 70%, and finally 14088 enhanced images are obtained for training the broccoli target detection model, thereby improving the robustness of the model.

3. A three-degree-of-freedom broccoli picking method according to claim 2, characterized in that: Also includes: Step S32: Based on the YOLOv5 model, the Neck part is improved using adaptive spatial feature fusion, and a convolutional block attention module is added to solve the inconsistency of image features and improve the recognition and detection effect.

4. A three-degree-of-freedom broccoli picking method according to claim 3, characterized in that: In step S32, based on the YOLOv5 neural network, the adaptive spatial feature fusion technology is used to improve the Neck part, and a convolutional block attention module is added to solve the inconsistency problem of image features; ASFF dynamically adjusts the weighting coefficients of feature maps of different layers so that feature maps from different scales can be fused according to their importance to target detection; ASFF works as follows: For each level of feature maps, all other levels of feature maps are adjusted to the same shape and spatially fused according to the learned weight map to ensure that feature information of different scales is effectively utilized. The specific algorithm is: in, Represents the output feature value of the feature map of the lth layer, that is, the value of the feature map of the lth layer at position i, j, obtained by weighted fusion of feature map information from different layers. Represents the values ​​of feature maps from different sources input to the feature map of the lth layer, is the input feature map value from layer l, From the 2nd floor, Feature map values ​​from layer 3, It is a weighting coefficient that controls the contribution of each layer of feature maps in the final output. It is learned during the training process. i and j represent the position index of the pixel in the feature map. i and j are the row and column indexes of the feature map, respectively. l represents the level of the current feature map. The level is used in the network to process the input data at different stages. Different levels are responsible for extracting feature information at different levels in the data.

5. A three-degree-of-freedom broccoli picking method according to claim 4, characterized in that: In step S32, ASFF processes the weighting coefficients using a normalized exponential function to ensure that the sum of the weighting coefficients is 1 and the coefficient value is between [0, 1]. The specific algorithm is as follows: The formula ensures that the weighted coefficient of each feature map does not exceed 1 during spatial fusion, while ensuring the balance of weighting to avoid excessive influence of feature maps of a certain layer.

6. A three-degree-of-freedom broccoli picking method according to claim 3, characterized in that: In step S32, through adaptive spatial feature fusion, ASFF will weight the feature map of each layer according to the learned weights, and calculate the weight coefficient of each layer by the following formula: The formula calculates the weighted coefficient of each feature map through a normalized exponential function, and automatically adjusts the weight through the training process to optimize the performance of the ASFF module in the fusion of features at different scales.

7. The three-degree-of-freedom broccoli picking method according to claim 2, characterized in that: Also includes: Step S33: training the broccoli target detection model to generate a network file; Step S34: deploying the generated network file on the computer main control unit; Step S35: passing the received pre-processed RGB image into the broccoli object detection model; Step S36: the broccoli object detection model identifies the broccoli that is ripe and ready for picking based on the received image, and frames the broccoli head with a rectangular frame; Step S37: After the mature broccoli is identified, the position of the target center is calculated based on the rectangular box drawn by the broccoli target detection model. The diagonal coordinates of the rectangular box are n1 (x1, y1) and n2 (x2, y2). Therefore, the geometric center point of the broccoli rectangular box is as follows: The coordinates of the geometric center o in the broccoli head in the image are obtained as (x o ,y o ), and then use the ZED2 camera to obtain the depth value D of the camera in the image coordinate system, and convert the image coordinate system into the camera coordinate system (x c ,y c ), and thus obtain the spatial position of the flower ball in the camera coordinate system (x c ,y c , D), and then the spatial position (x c ,y c , D) The information is sent to the picking robot arm through the network interface to complete the positioning task.

8. A three-degree-of-freedom broccoli picking method according to claim 7, characterized in that: In step S33, by training the YOLOv5 target detection model, a network file optimized by adaptive spatial feature fusion and convolutional block attention module is generated. During the training process, the cross entropy loss function and the IoU loss function are used as optimization targets to ensure the accuracy of the model in the target detection task. After the training is completed, the generated network file is deployed to the computer main control unit to process image data in real time and perform target recognition; In step S37, the spatial position (x c ,y c , D) information is sent to the picking robot arm through the network interface. The picking robot arm adjusts the motion trajectory according to the received spatial position information to complete the precise picking of broccoli. During this process, the picking robot arm grabs the broccoli head through a precise straight up and down motion path, and ensures that the picking process is fast and accurate, reduces the damage to the head, and improves the picking efficiency.

9. The three-degree-of-freedom broccoli picking method according to claim 3, characterized in that: In step S32, the convolutional block attention module performs a two-dimensional attention mechanism on the feature map. On a given intermediate feature map F, CBAM first performs a channel attention mechanism M C (F), then the processed feature map is element-wise multiplied with the original feature map F to obtain the weighted feature map F′; then, the spatial attention mechanism M S (F′) is applied to the feature map F′, and finally the final weighted feature map F″ is obtained by element-wise multiplication. The feature map distributes attention in two dimensions to strengthen the model's attention to key features. The specific expression is: in, Represents an element-wise multiplication operation. Channel attention and spatial attention increase the response of the feature map to important areas in this way, thereby enhancing the performance of broccoli object detection.

10. A three-degree-of-freedom broccoli picking device, used to implement the three-degree-of-freedom broccoli picking method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Vehicle-mounted thermal image detection based on improved YOLOv5

    CN115546765A

  • Intelligent identification harvesting method suitable for broccoli

    CN117016200A

  • Broccoli selective harvesting robot and control method thereof

    CN117044496A

  • Bionic perception broccoli selective harvesting claw integrating clamping and cutting

    CN117958021A

  • Remote control type lotus seedpod picking machine

    CN119278771A

Cited By

  • Trollius chinensis recognition and harvesting method, system and equipment and medium

    CN121236615A

  • A method, system, device and medium for identifying and harvesting trollius chinensis

    CN121236615B