Target detection method and system based on improved YOLOv11

By improving the YOLOv11 object detection algorithm and combining RGB-D information, robotic arm kinematics and three-dimensional coordinate conversion technology, the problem of inaccurate target recognition and low grasping accuracy in traditional robotic arms in dynamic environments is solved, and efficient and accurate target recognition and grasping in complex environments is achieved.

CN119992048APending Publication Date: 2025-05-13NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510058417.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional robotic arm grasping systems have inaccurate target recognition, low grab accuracy and slow response speed in dynamic environments, making it difficult to achieve efficient and accurate target recognition and grabbing in complex contexts.

Method used

The target detection method based on improved YOLOv11 is adopted, combined with RGB-D information, robotic arm kinematics and three-dimensional coordinate conversion technology, to realize the target recognition of the robotic arm in any position and make corresponding decisions, improving the grasping accuracy and efficiency.

Benefits of technology

Achieve high-accuracy target recognition and efficient grab operation in complex environments, improving the grab accuracy and efficiency of the robotic arm and reducing the time delay between target detection and grabbing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992048A_ABST
    Figure CN119992048A_ABST
Patent Text Reader

Abstract

The invention relates to a target detection method and system based on improved YOLOv11, and aims to improve the grabbing precision and operation efficiency of a mechanical arm in a complex environment. The method combines an augmented reality technology, a target detection algorithm, a depth sensing technology and an advanced motion control system, and specifically comprises the following steps that an RGB-D camera is used for collecting data; establishing a three-dimensional space model, and performing accurate alignment processing on the depth image and the color image; based on a mercuric chloride Atlas 200I DK A2 development board, in combination with an improved YOLOv11 model, efficient detection and identification of a target are realized; a target is converted from a pixel coordinate system to a camera coordinate system through a coordinate conversion method, and then the target is further mapped to the three-dimensional space position of a mechanical arm base coordinate system; finally, grabbing operation is completed through the elephant mechanical arm MyCobot 280M5. By combining RGB-D information, a deep learning detection algorithm and a three-dimensional coordinate conversion method, efficient and accurate target grabbing of the mechanical arm under the complex environment is achieved through the system, and the system has the advantages of being high in precision and real-time performance, easy to maintain, easy to integrate and deploy and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, automation and robotics, and is specifically a method for realizing a target detection method and a robotic arm grasping system based on an improved YOLOv11. The present invention combines augmented reality technology, target detection algorithm, depth sensing technology and advanced motion control system, aiming to improve the grasping accuracy and efficiency of the robotic arm in a complex environment. Background Art

[0002] With the rapid development of automation technology, the application of robotic arms in multiple fields has gradually become popular, covering multiple industries such as industrial manufacturing, logistics, medical care, agriculture and home services. The performance of robotic arms in specific tasks, especially the level of grasping ability, directly affects production efficiency and work quality. Therefore, improving the ability of robotic arms to quickly and accurately detect target objects and make intelligent decisions has become one of the focuses of current research and development.

[0003] Traditional robotic gripping systems usually rely on static cameras and basic image processing technology, but they show many limitations in dynamic environments and complex gripping tasks. Static cameras cannot provide real-time depth information, which makes it easy for errors to occur in the positioning and identification of target objects. In addition, due to insufficient processing power, traditional image processing algorithms often face problems such as slow response speed and low recognition accuracy in the analysis of real-time video streams. These technical shortcomings directly affect the accuracy and efficiency of gripping operations and limit the application scope of robotic arms in complex operating scenarios.

[0004] In recent years, the development of RGB-D camera technology has significantly improved the three-dimensional space perception ability. By acquiring depth information and color images at the same time, RGB-D cameras can accurately reconstruct three-dimensional scenes and provide more comprehensive information for environmental perception. This brings new room for improvement in the decision-making ability of the robot arm in the target grasping task. However, how to efficiently process and utilize these rich data to optimize the motion control of the robot arm is still a key technical problem. The challenge lies not only in the efficient coordination and analysis of depth maps and color maps, but also in how to quickly and accurately convert the processed information into the motion instructions of the robot arm to meet the real-time needs in dynamic environments. In terms of target detection, although some existing deep learning algorithms have shown good performance in specific environments, how to balance detection accuracy and processing speed in real-time grasping scenarios still puts higher requirements on system design. At the same time, the detection algorithm for moving or dynamic targets needs to be further optimized to improve the robustness in complex backgrounds.

[0005] Developing a robotic gripping system that integrates advanced target detection algorithms, deep image processing, and precise motion control has extremely important practical application value for achieving efficient automatic gripping operations in various complex environments. This type of gripping system based on a domestic platform can not only effectively improve the automation level of industrial production lines, but also promote the further development of artificial intelligence and robotics technology to meet the ever-changing market needs. Summary of the invention

[0006] In order to solve the above problems, the present invention provides a target detection method and system based on improved YOLOv11. The method uses RGB-D information, robot kinematics, improved YOLOv11 and three-dimensional coordinate conversion method to enable the robot to recognize targets in any posture and make corresponding decisions to realize the robot to grasp the target object, solving the problems of inaccurate target recognition, low grasping accuracy and slow response speed of traditional robot in dynamic environment.

[0007] In order to achieve the above object, the present invention is achieved through the following technical solutions:

[0008] The present invention is a target detection method and system based on improved YOLOv11, and the implementation method of the system comprises the following steps:

[0009] Step 1: Combine the robotic arm and RGB-D camera to build a 3D space model and achieve accurate alignment of depth image and color image.

[0010] Step 2: Integrate the Ascend Atlas 200IDK A2 development board and the improved YOLOv11 model to perform target detection and recognition on the object to be grasped.

[0011] Step 3: Using the information of the object to be grasped obtained in the previous steps, perform three-dimensional space coordinate transformation to achieve the transformation from the pixel coordinate system to the camera coordinate system, and from the camera coordinate system to the robot base coordinate system.

[0012] Step 4: Through the motion control of the Elephant Robot Arm MyCobot 280 M5, the robot arm performs the grabbing operation on the target coordinate position.

[0013] A further improvement of the present invention is that in step 1, the method of combining the robotic arm and the RGB-D camera to construct a three-dimensional space model and realize accurate alignment of the depth image and the color image includes the following:

[0014] 1) Use Obi’s camera tools to extract camera internal parameters.

[0015] 2) Efficient image alignment is achieved through correlation algorithms, making the depth and color information accurately matched.

[0016] A further improvement of the present invention is that in step 2, the Ascend Atlas 200I DK A2 development board and the improved YOLOv11 model are integrated, and the method for target detection and recognition of the object to be grasped includes the following:

[0017] 1) Preprocess the objects to be grasped and annotate them to create a dataset.

[0018] 2) The improved YOLOv11 model is selected as the target detection algorithm, and the MindSpore framework is used for training and optimization to generate a model that can efficiently detect target objects.

[0019] 3) Convert the trained model to generate an om model that can be adapted to the Atlas 200I DK A2 development board.

[0020] 4) In the data processing stage, the CANN computing architecture is used to achieve rapid response to dynamic environments and perform real-time video stream processing.

[0021] A further improvement of the present invention is that in step 3, the method for realizing the conversion from the pixel coordinate system to the camera coordinate system, and the conversion from the camera coordinate system to the robot arm base coordinate system includes the following:

[0022] 1) In the RGB-D vision system, the intrinsic parameter matrix and depth value of the depth camera are used to complete the conversion from the pixel coordinate system to the camera coordinate system.

[0023] 2) Through hand-eye calibration, the camera coordinate system is converted to the robot arm base coordinate system.

[0024] The beneficial effects of the present invention are as follows: the present invention is based on an improved target detection method and system of YOLOv11, adopts RGB-D information, an improved YOLOv11 target detection algorithm in deep learning, and three-dimensional coordinate conversion technology, so that the robot arm can perform efficient and accurate target recognition and grasping in a complex environment. Specifically, the present invention has the following advantages:

[0025] Accurate recognition and grasping: By combining the depth information provided by the RGB-D camera with the improved YOLOv11 target detection algorithm, the present invention can achieve high-accuracy target recognition in complex backgrounds, thereby ensuring that the grasping operation of the robot arm achieves the ideal effect and improving the overall operation efficiency.

[0026] High real-time performance: In the data processing stage, the CANN computing architecture is combined with real-time video stream processing to greatly reduce the time delay between target detection and capture, achieve fast response and efficient operation, and effectively improve the overall work efficiency of the system.

[0027] Lightweight and efficient computing: This paper integrates the universal reverse bottleneck (UIB) module and the bidirectional feature pyramid network (BiFPN) in the YOLOv11 algorithm, realizing the organic combination of model lightweight and efficient computing. This optimized design not only reduces the system's computing resource requirements, but also ensures high-precision target detection effects.

[0028] Accurate mapping between multiple coordinate systems: The present invention ensures high accuracy of coordinate transformation through multi-level transformation of pixel coordinate system, camera coordinate system and robot base coordinate system, combined with hand-eye calibration technology and rotation matrix modeling. This multi-coordinate system mapping technology can dynamically adjust the spatial position of the target object and improve the grasping robustness of the robot arm.

[0029] Easy to integrate and deploy: Using domestic hardware platforms (such as Ascend Atlas 200I DK A2 development board and Elephant Robot Arm MyCobot 280 M5), it has good compatibility and high cost performance, which is convenient for integration with existing automation systems. The overall system design is modular, the process is simplified, and the hardware deployment and algorithm update are convenient, which reduces the complexity of implementation and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a schematic diagram of the workflow of each module.

[0031] Figure 2 This is the improved YOLOv11 structure diagram used in this invention.

[0032] Figure 3 Schematic diagram of the coordinate system transformation. DETAILED DESCRIPTION

[0033] The hardware structure of the present invention mainly includes the following three modules: MyCobot 280 M5, Huawei Ascend development board Atlas 200I DK A2, and Obbec RGB-D camera Gemini 2. Figure 1 as shown.

[0034] The MyCobot 280 M5 robot arm mainly provides joint control and coordinate control.

[0035] Huawei Ascend development board Atlas 200I DK A2 is mainly responsible for data processing and algorithm operation.

[0036] The Orbbec RGB-D camera Gemini 2 is mainly responsible for obtaining the RGB-D information of the object to be grasped.

[0037] The deployment process of the robotic arm grasping system implementation solution proposed in the present invention mainly includes the following steps:

[0038] Step 1: Using the depth camera of Orbbec, the camera's intrinsic parameter matrix, including focal length (fx, fy) and optical center coordinates (cx, cy) are successfully obtained through its built-in OrbbecViewer tool. These parameters define the mapping relationship between the pixel points in the image and the actual space points, ensuring the accuracy of the subsequent conversion process. In order to obtain more accurate three-dimensional spatial information and ensure that the depth information can accurately correspond to the pixel points in the color image, the relevant algorithm provided by Orbbec is used to achieve efficient alignment of the depth image and the color image. This step provides the necessary conditions for subsequent three-dimensional reconstruction and calculation of the three-dimensional coordinates of the pixel points.

[0039] Step 2: Preprocessing the object to be grasped is the key to ensuring the target detection effect. First, the depth image and color image of the target object are obtained through the Orbbec RGB-D camera, and the target object in the color image is annotated using the annotation tool, including the bounding box and category of the object, to build the data set required for training. This data set will provide an important foundation for the subsequent training of the target detection algorithm.

[0040] Step 3-1: The present invention proposes an innovative improvement method based on the YOLOv11 target detection algorithm. Based on the original algorithm, this method optimizes its core module to achieve a coordinated improvement in detection accuracy and model lightweight. Specifically, the present invention incorporates a Universal Inverted Bottleneck (UIB) module designed based on MobileNetV4 into the C3k2 module of the YOLOv11 algorithm, and performs secondary innovative optimization on it. The UIB module used as a flexible architecture unit integrates the Inverted Bottleneck (IB), ConvNext architecture, Feed Forward Network (FFN) and novel additional depth variants (Extra Depthwise, ExtraDW). This improvement significantly improves the algorithm's balance between model accuracy and computational efficiency. In addition, in the neck network part of the model, the present invention introduces an improved bidirectional feature pyramid network (BiFPN). As an efficient multi-scale feature fusion architecture, BiFPN optimizes and upgrades the shortcomings of the traditional feature pyramid network (FPN). By establishing a bidirectional connection mechanism between the top-down path and the bottom-up path, BiFPN maximizes the information flow between features of different scales. This bidirectional information flow mode effectively enhances the interactive ability of the feature map and improves the efficiency of feature fusion, thereby significantly improving the overall accuracy and effect of detection. The specific improved structure diagram is as follows Figure 2 as shown.

[0041] Step 3-2: Select the improved YOLOv11 model as the target detection algorithm, and use the MindSpore framework to train and optimize the model. Input the created labeled data set into the model for model training. By setting appropriate hyperparameters and loss functions, the model can learn the characteristics of the target object. During the training process, data enhancement techniques (such as translation, rotation, scaling, etc.) are used to increase the diversity of training data and improve the generalization ability of the model. By controlling the learning rate and adopting an adaptive learning mechanism, ensure that the model can converge quickly and achieve a high accuracy rate. After the training, use MindSpore to convert the trained improved YOLOv11 model into the om format suitable for Huawei Ascend development board Atlas 200IDK A2, so as to facilitate subsequent real-time reasoning on edge devices.

[0042] Step 4: In the real-time reasoning process of target detection, the CANN computing architecture is used to enable Atlas 200I DKA2 to efficiently handle complex target detection tasks. By optimizing the real-time performance of the algorithm in static and dynamic environments, the target detection model can respond quickly to environmental changes. In this step, the input real-time video stream is processed by the improved YOLOv11 model, and the system can instantly identify and locate the target object in the image and extract its specific location information. This information will be used for subsequent three-dimensional space coordinate conversion, tracking the position change of the target object, and providing real-time feedback for the robot arm's grasping operation.

[0043] Step 5: In the field of three-dimensional space coordinate transformation, it is usually necessary to realize the transformation from pixel coordinate system to camera coordinate system, and then from camera coordinate system to robot base coordinate system, such as Figure 3 As shown, there are three steps:

[0044] Step 5-1: In the three-dimensional space coordinate conversion, in order to use the perception information of the visual system for the control of the robot arm, it is necessary to realize the conversion from the pixel coordinate system to the camera coordinate system, and then from the camera coordinate system to the robot arm base coordinate system. First, the conversion from the pixel coordinate system to the camera coordinate system must be completed. In the RGB-D visual system, the pixel coordinate system represents the two-dimensional coordinates (u, v) on the image plane, while the camera coordinate system is used to describe the three-dimensional spatial position (x C ,y C , z C ). This conversion depends on the camera's intrinsic matrix and depth information Z C , the calculation method is described below. The camera intrinsic parameter matrix K usually has the following form:

[0045]

[0046] Among them, f xand f y is the focal length, c x and c y are the pixel coordinates of the principal point of the image.

[0047] Step 5-2: After completing the conversion from pixel coordinates to camera coordinates, the next step is to convert the points in the camera coordinate system to the robot base coordinate system. This process relies on the description of the relative position and posture relationship between the camera and the robot, which is usually obtained through the hand-eye calibration method.

[0048] The goal of hand-eye calibration is to calculate the transformation matrix T from the camera coordinate system to the robot base coordinate system. Base-Camera , this matrix can be decomposed into the rotation matrix R and the translation vector t:

[0049]

[0050] Among them, R is a 3×3 rotation matrix, which describes the rotation of the camera coordinate system relative to the robot base coordinate system, and t is a 3×1 translation vector, which describes the translation position of the camera coordinate system.

[0051] The coordinate transformation formula is:

[0052]

[0053] Through this matrix, the points in the camera coordinate system can be transformed and mapped to the robot base coordinate system, providing accurate geometric relationship support for subsequent operations.

[0054] Step 5-3: When integrating the vision system of the robot with the operating system, in order to realize the conversion from the vision coordinates to the robot base coordinate system and control the robot, it is also necessary to determine the rotation and translation of the end of the robot, and establish a mapping relationship from the camera coordinate system to the robot base coordinate system. Through the interface provided by the robot, the position and posture [x, y, z, r] of the end effector in the robot base coordinate system can be obtained. x , r y , r z ], where [x, y, z] represents the coordinates of the end in the robot base coordinate system, and [r x , r y , r z ] represents the rotational posture of the end, which can be converted into a rotation matrix R. The specific conversion method is:

[0055] The matrix R for rotation around the X axis x (r x ):

[0056]

[0057] The matrix R for rotation around the Y axisy (r y ):

[0058]

[0059] The matrix R of rotation around the Z axis z (r z ):

[0060]

[0061] Calculate the combined rotation matrix. For the Euler angle rotation in the Roll-Pitch-Yaw order, the final rotation matrix R can be obtained by multiplying the above matrices in order:

[0062] R=R z (r z )·R y (r y )·R x (r x )

[0063] Combining the above methods and steps, a complete three-dimensional space coordinate system conversion algorithm was written in Python language to achieve accurate conversion from pixel coordinate system to camera coordinate system and then to robot base coordinate system. This algorithm combines camera internal and external parameter calibration, hand-eye calibration, and mathematical modeling and calculation of rotation matrix and translation vector to achieve efficient mapping between multiple coordinate systems.

[0064] Step 6: Use the API officially provided by the Elephant Robotic Arm for precise control. Specifically, the gripper of the robot arm can automatically locate and smoothly reach the target position according to the converted coordinates. In this way, the robot arm can efficiently complete complex grasping operations.

Claims

1. A target detection method and system based on improved YOLOv11, characterized in that ,The implementation method of the system includes the following steps: Step 1: Combine the robotic arm and RGB-D camera to build a 3D space model and achieve accurate alignment of depth image and color image. Step 2: Based on the Ascend Atlas 200I DK A2 development board and the improved YOLOv11 model, perform target detection and recognition on the object to be grasped. Step 3: Using the information of the object to be grasped obtained in the previous steps, perform three-dimensional space coordinate transformation to achieve the transformation from the pixel coordinate system to the camera coordinate system, and from the camera coordinate system to the robot base coordinate system. Step 4: Through the motion control of the Elephant Robot Arm MyCobot 280 M5, the robot arm performs the grabbing operation on the target coordinate position.

2. According to claim 1, a target detection method and system based on improved YOLOv11, characterized in that: In step 1, the method of combining the robotic arm and the RGB-D camera to construct a three-dimensional space model and achieve accurate alignment of the depth image and the color image includes the following: 1) Use Obi’s camera tools to extract camera internal parameters. 2) Efficient image alignment is achieved through correlation algorithms, allowing accurate matching of depth and color information.

3. According to claim 1, a target detection method and system based on improved YOLOv11, characterized in that: In step 2, the Ascend Atlas 200I DK A2 development board and the improved YOLOv11 model are integrated to perform target detection and recognition on the object to be grasped, including the following methods: 1) Preprocess the objects to be grasped and annotate them to create a dataset. 2) The improved YOLOv11 model is selected as the target detection algorithm, and the MindSpore framework is used for training and optimization to generate a model that can efficiently detect target objects. 3) Convert the trained model to generate an om model that can be adapted to the Atlas 200I DK A2 development board. 4) Use the CANN computing architecture to achieve rapid response to dynamic environments and perform real-time video stream processing.

4. According to claim 1, a target detection method and system based on improved YOLOv11, characterized in that: In step 3, the method for realizing the conversion from the pixel coordinate system to the camera coordinate system, and the conversion from the camera coordinate system to the robot base coordinate system includes the following: 1) In the RGB-D vision system, the intrinsic parameter matrix and depth value of the depth camera are used to complete the conversion from the pixel coordinate system to the camera coordinate system. 2) Through hand-eye calibration, complete the conversion from the camera coordinate system to the robot arm base coordinate system.

Citation Information

Cited By

  • Multi-target capturing method and system based on deep learning

    CN120182383A

  • Signature handwriting authenticity digital identification method based on multi-modal features

    CN120932312A

  • Ginger seed taking system and method based on depth camera and mechanical arm

    CN121264243A

  • High-efficiency automatic distance control clamping system for euphausia superba freezing tray

    CN121989207A