Deep Learning-Based Adaptive Method and System for Robot Grasping
By constructing digital 3D models and using edge computing based on deep learning, robots can identify and respond to abnormal states in complex environments, enabling adaptive grasping of multiple objects. This solves the problem of existing technologies being unable to adapt to complex environments and improves the adaptability and accuracy of grasping.
Patent Information
- Application Number
- CN202510617907.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Existing technologies can only accurately capture stacked object scenes under specific rules and cannot adapt to more complex environments, especially stacked scenes with more than two objects.
By using a deep learning-based approach and an RGB-D camera to construct a digital 3D model, combined with a central control system and control modules, abnormal states are identified and emergency grasping options are triggered, enabling the robot to adaptively grasp multiple objects. Furthermore, more control modules and the robot are connected via external interfaces to perform edge computing to reduce data transmission latency.
It enables the robot to adaptively grasp at least two objects, improving the grasping adaptability and accuracy in complex environments, reducing data transmission latency, and enhancing the robot's response speed and grasping accuracy in abnormal situations.
Smart Images

Figure CN120496054B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot grasping technology, specifically a deep learning-based adaptive method and system for robot grasping. Background Technology
[0002] Robotic grasping technology is a core capability for realizing intelligent manufacturing, warehousing and logistics, and home service automation. Adaptive grasping methods, through a real-time perception-decision-execution closed loop, enable robots to autonomously adjust to environmental changes, and have become a research hotspot in recent years.
[0003] Chinese invention patent (CN119407773A) discloses a "robot visual grasping and detection method, system, and medium for object stacking scenarios," specifically disclosing: extracting image features from the RGB images of objects in the work scene; determining object detection boxes and grasping rectangles based on the image features, where the object detection boxes represent the object category and position in the work scene, and the grasping rectangles represent the grasping pose of the objects in the work scene; matching the object detection boxes and grasping rectangles; reflecting the matched object detection boxes and grasping rectangles into the image features, and then performing operation relationship detection on the objects in the work scene to obtain an object operation relationship tree; the robot grasps the objects in the work scene based on the object operation relationship tree, object detection boxes, and grasping rectangles. This enables accurate grasping by the robot in object stacking scenarios.
[0004] The above technical solutions achieve accurate grasping in object stacking scenarios. However, they can only accurately grasp stacked objects, and the object grasping is performed under specific rules. They cannot adapt to more complex environments, such as accurately grasping stacks of at most two objects. Therefore, there is an urgent need for a deep learning-based adaptive robot grasping method and system to solve the above problems. Summary of the Invention
[0005] The purpose of this invention is to provide a deep learning-based adaptive method and system for robot grasping, in order to solve the problems raised in the prior art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a deep learning-based adaptive method for robot grasping.
[0007] A digital 3D model of the object grasped by the robot is acquired and constructed using an RGB-D camera.
[0008] The central control system uses a deep learning model to identify abnormal states of objects in a digital 3D model and determine whether an emergency grabbing condition has been triggered.
[0009] The control module receives instructions from the central control system, activates the emergency grasp option, and controls the robot to adaptively grasp objects in abnormal states.
[0010] The overall control system is also equipped with an external interface, which is used to connect other control modules and robots when the object placement is abnormal, and cooperate with the original control module and robot to perform adaptive object grasping.
[0011] According to the above technical solution, the training of the deep learning model is carried out by collecting a dataset of images of objects without abnormalities and a dataset of images of real abnormal objects, defining a loss function, training the model with the dataset, and completing the test to finally obtain the deep learning model.
[0012] The deep learning model identifies abnormal states and uses annotation tools to define the stacked objects, forming anomaly identification boxes.
[0013] The annotation tool is also used to mark any single object, with the marking point located at the physical center of the object, forming a calibration point;
[0014] For objects within the anomaly identification box, their 3D data is extracted, and image processing technology is used to extract the outlines of the objects within the anomaly identification box, resulting in several sets of closed outlines. The similarity of these closed outlines is compared one by one with the outline of a single object to determine the outline of the topmost object in the stacked objects. The topmost object is then specially marked using a labeling tool.
[0015] According to the above technical solution, the deep learning model sends a digital 3D model containing anomaly calibration boxes, calibration points, and special markers to the control module. The control module modifies the robot's original grasping scheme and sends the modified grasping scheme to the grasping robot. The grasping robot executes the grasping command to complete the adaptive grasping of objects in abnormal states.
[0016] According to the above technical solution, for an abnormal state where the object being grasped has an abnormal bounding box, the grasping robot is used to grasp the object with special markings. Then, the RGB-D camera installed on the grasping robot is used to collect three-dimensional data, and the collected three-dimensional data is transmitted back to the control module. The control module continues to perform image processing on the objects within the abnormal bounding box based on the three-dimensional data, and uses a labeling tool to specially mark the topmost object within the abnormal bounding box until there is no stacking phenomenon among the objects within the abnormal bounding box.
[0017] According to the above technical solution, the deep learning model also includes a data analysis unit, which is used to analyze other abnormal situations that occur during the object transportation process. The data analysis unit is connected to the object transportation control system and is used to receive relevant information data about the object transportation.
[0018] According to the above technical solution, the data analysis unit defines the first unit area in the digital three-dimensional model, counts the number of calibration points in the first unit area, and calculates the object density within the first unit area.
[0019] When the data analysis unit receives an increase or decrease in the speed of the object being transported, the data analysis unit continues to define the second unit area of the digital 3D model, count the number of calibration points in the second unit area, and analyze the object density within the second unit area.
[0020] When the speed of object transport increases, the first unit area is larger than the second unit area; when the speed of object transport decreases, the first unit area is smaller than the second unit area.
[0021] When the density of objects within the first or second unit area is greater than or equal to a set threshold, other abnormal situations are determined to have occurred.
[0022] According to the above technical solution, when other abnormal situations occur, additional control modules and grasping robots are connected through the external interface of the central control system. The central control system sends data information to different control modules respectively, and at least two grasping robots grasp the objects conveyed on the conveyor belt.
[0023] According to the above technical solution, at least two grasping robots are arranged in a grasping order. After the first grasping robot completes the grasping of the object, it collects three-dimensional data of the object being transported through an RGB-D camera and transmits the collected three-dimensional data to the control module of the next grasping robot. The next grasping robot plans a grasping scheme for the object being transported based on the latest collected three-dimensional data, and collects three-dimensional data again when the next grasping robot completes the grasping.
[0024] If this grasping robot is the last grasping robot in the sequence, the acquired 3D data is transmitted to the control module of the first grasping robot in the sequence. This allows the control module to complete the cycle of the entire grasping scheme planning and achieve adaptive grasping of the object.
[0025] A deep learning-based adaptive robot grasping system, wherein the RGB-D camera used for establishing digital 3D models is located at the front end of the object conveying direction and directly above the object in the spatial plane, and the grasping robot is located in the middle of the conveying direction and behind the RGB-D camera.
[0026] The control module includes a capture planning unit and an instruction sending unit;
[0027] The grasping planning unit is used to replan the grasping scheme of the grasping robot based on the abnormal state of the object to be grasped when the emergency grasping option is enabled, and transmit the grasping scheme to the instruction sending unit. The instruction sending unit sends the revised grasping scheme to the grasping robot, and the grasping robot executes the instruction to adaptively grasp the object in the abnormal state.
[0028] According to the above technical solution, the grasping robot is also equipped with an RGB-D camera. After grasping an object, the grasping robot collects three-dimensional data through the RGB-D camera and transmits the collected three-dimensional data to the grasping planning unit for planning subsequent grasping schemes on the control module side.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] This invention establishes a digital 3D model, identifies and analyzes it through a central control system, and transmits the results to a control module. This control module enables the robot to adaptively grasp at least two stacked objects. Furthermore, the invention includes an external interface to connect to more control modules and the robot, allowing for adaptive grasping even in more complex environments. This improves the robot's adaptability to adaptive grasping and, for more complex environments, enables data analysis and edge computing on the robot's end, reducing latency in data transmission and ensuring the accuracy of object grasping. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the adaptive grasping control process of the robot according to the present invention;
[0032] Figure 2 This is a further schematic diagram of the adaptive grasping control of the robot according to the present invention;
[0033] Figure 3 This is a schematic diagram illustrating the relationship between the overall control system and the control modules of this invention;
[0034] Figure 4 This is a further flowchart illustrating the relationship between the overall monitoring system, control module, and robot of the present invention. Detailed Implementation
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0036] Example: Figure 1-Figure 2 As shown, the robot in this embodiment adaptively grasps objects transported on a conveyor belt, specifically including the following steps:
[0037] S1. Construct a digital twin model of the object grasped by the robot;
[0038] Specifically, the robot acquires 3D data of the object it grasps using an RGB-D camera and constructs a digital 3D model of the object.
[0039] Furthermore, the RGB-D camera is located at the front end of the conveyor belt in the conveying direction and directly above the conveyor belt in the spatial plane. It is used to collect three-dimensional data of objects passing through the RGB-D camera and construct digital three-dimensional models. The robot is located in the middle of the conveyor belt in the conveying direction and behind the RGB-D camera. It is used to adaptively grasp objects on the conveyor belt.
[0040] After acquiring 3D data, the RGB-D camera completes the construction of a digital 3D model through a process of point cloud registration, surface reconstruction, texture mapping, and model optimization.
[0041] Before the robot grasps an object, by establishing a digital 3D model of the object to be grasped, on the one hand, the robot can plan the grasping scheme in advance for adaptive grasping, and on the other hand, abnormal states of object placement can be identified so that the robot can take timely countermeasures and achieve adaptive grasping.
[0042] S2, The central control system processes and analyzes data from the digital 3D model;
[0043] Specifically, the central control system uses a deep learning model to identify abnormal states in the constructed digital 3D model and determine whether emergency capture conditions are triggered.
[0044] The deep learning model is mainly used to identify and analyze the state of objects in digital 3D models, such as objects stacking on the surface of a conveyor belt.
[0045] For training deep learning models, we collect datasets of images without abnormal objects and datasets of images of real abnormal objects, define a loss function, train the model using the datasets, and complete the testing to finally obtain the deep learning model.
[0046] The deep learning model identifies abnormal states and uses annotation tools to define the objects in the abnormal states, forming anomaly labeling boxes.
[0047] The annotation tool is also used to mark any single object, with the marking point located at the physical center of the object, forming a calibration point;
[0048] An abnormal state refers to the phenomenon of objects being stacked. The purpose of framing the stacked objects is to facilitate the robot's adaptive grasping of the stacked objects. The purpose of marking any single object is to analyze and judge whether there are other abnormal situations in the transport of objects on the conveyor belt, so that the robot can make timely responses and perform adaptive grasping.
[0049] For objects within the anomaly calibration box, their 3D data is extracted, and image processing technology is used to extract the outlines of the objects within the anomaly calibration box to obtain several sets of closed outlines. The similarity of these sets of closed outlines is compared one by one with the outlines of the objects transported by the conveyor belt to determine the outline of the topmost object in the stacked objects, and it is specially marked by the annotation tool.
[0050] When an object being grasped has an abnormal bounding box, the grasping robot grasps the object with the special mark. Then, the RGB-D camera installed on the grasping robot collects three-dimensional data and transmits the collected three-dimensional data back to the control module. The control module continues to process the images of the objects within the abnormal bounding box based on the three-dimensional data, and uses a labeling tool to specially mark the topmost object within the abnormal bounding box until there is no stacking of objects within the abnormal bounding box.
[0051] The deep learning model also includes an additional data analysis unit; the data analysis unit is used to analyze other abnormal situations that occur on the conveyor belt, such as: increased object density per unit area on the surface of the conveyor belt or increased conveying speed of the conveyor belt for the object, etc.
[0052] The design of the data analysis unit enables the optimal grasping scheme to be adjusted for abnormal object transport states in different scenarios, allowing the grasping robot to adaptively grasp objects in different scenarios and improving the robot's adaptability.
[0053] Specifically, in this embodiment, the data analysis unit defines a unit area in the digital 3D model, where the unit area is S, and counts the number of calibration points in the unit area to determine the number of calibration points as N. The data analysis unit then calculates the density P of the object within the unit area according to the formula P=N / S.
[0054] When the density P of objects per unit area is greater than or equal to a set threshold, it is determined that the density of objects per unit area on the surface of the conveyor belt has increased and other abnormal situations have occurred.
[0055] The data analysis unit is connected to the conveyor belt control system and is used to receive the conveyor belt's transmission speed data V. When the conveyor belt's transmission speed V increases, the data analysis unit continues to define the unit area of the digital 3D model and analyze the object density P within the unit area. ’ However, the defined unit area is S. ’ =a*S, where a is the proportionality coefficient;
[0056] Because the transmission speed of the conveyor belt increases, in order to ensure that the data analysis results are not affected while the density threshold remains unchanged, it is necessary to scale the unit area defined in the digital 3D model proportionally to ensure the uniformity of density calculation and avoid increasing the system's computational load by setting multiple thresholds or modifying the threshold.
[0057] The proportionality coefficient 'a' is affected by the change in the transmission speed of the conveyor belt. When the change in transmission speed is greater than 0, a < 1; when the change in transmission speed is less than 0, a > 1.
[0058] S3. The control module receives instructions from the central control system, enables the emergency grasping option, and controls the robot to adaptively grasp objects in abnormal areas.
[0059] Specifically, such as Figure 2 As shown, the control module includes a capture planning unit and an instruction sending unit;
[0060] The grasping planning unit is used to replan the robot's grasping scheme based on the abnormal state of the object to be grasped when the emergency grasping option is enabled, and transmit the grasping scheme to the instruction sending unit. The instruction sending unit sends the modified grasping scheme to the robot, and the robot executes the instruction to adaptively grasp the object in the abnormal state.
[0061] Furthermore, the robot's gripper is also equipped with an RGB-D camera, which is used to collect three-dimensional data of the object before and after grasping it.
[0062] In this embodiment, when the abnormal situation is that objects are stacked, the grasping planning unit first grasps the objects with special marks within the abnormality marking box when planning the grasping scheme. When grasping the specially marked objects, the RGB-D camera is used to collect three-dimensional data again, that is, point cloud data. Then, the robot uses image processing technology to extract the contour lines of the remaining objects within the abnormality marking box, and obtains several sets of closed contour lines. The similarity of these sets of closed contour lines is compared one by one with the contour lines of the objects transported by the conveyor belt to determine the contour of the topmost object among the remaining stacked objects. Then, the robot grasps the topmost object. The robot repeats the above actions until there are no more stacked objects within the abnormality marking box.
[0063] In this embodiment, the abnormal situation is when the density of objects per unit area on the surface of the conveyor belt increases or the conveyor belt's conveying speed for objects increases;
[0064] When the density of objects per unit area is less than a set threshold, the grasping planning unit increases the robot's grasping speed when changing the grasping scheme, thereby ensuring that the grasping function can be performed normally in the event of abnormal situations.
[0065] like Figures 3-4 As shown, the central control system is also equipped with an external interface. When the density of objects in a unit area is greater than or equal to a set threshold, an additional control module and robot are connected through the external interface. At this time, the central control system sends data information to different control modules respectively, and at least two robots grab the objects conveyed on the conveyor belt to solve the problem of the density of objects in a unit area on the conveyor belt being greater than or equal to the set threshold.
[0066] Two or more robots are arranged along the side of a conveyor belt in a grasping sequence. After the preceding robot completes the grasping of an object, it uses an RGB-D camera to collect 3D data (point cloud data) of the object on the conveyor belt and transmits the collected 3D data to the grasping planning unit of the control module of the next robot. The grasping planning unit of the next robot plans a grasping scheme for the object on the conveyor belt based on the latest collected 3D data. When the next robot completes the grasping, it collects 3D data again. If this robot is the last one in the sequence, it transmits the collected 3D data to the grasping planning unit of the first robot in the sequence. This completes the cycle of the entire grasping scheme planning at the control module. Edge computing reduces the amount of data transmitted, making the robot respond faster and grasp more accurately when grasping objects.
[0067] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A deep learning-based adaptive method for robot grasping, characterized by: A digital 3D model of the object grasped by the robot is acquired and constructed using an RGB-D camera. The central control system uses a deep learning model to identify abnormal states of objects in a digital 3D model and determine whether an emergency grabbing condition has been triggered. The control module receives instructions from the central control system, activates the emergency grasp option, and controls the robot to adaptively grasp objects in abnormal states. The overall control system is also equipped with an external interface, which is used to connect other control modules and robots when the object placement is abnormal, and cooperate with the original control module and robot to perform adaptive object grasping. For training deep learning models, we collect datasets of images without abnormal objects and datasets of images of real abnormal objects, define a loss function, train the model using the datasets, and complete the testing to finally obtain the deep learning model. The deep learning model identifies abnormal states and uses annotation tools to define the stacked objects, forming anomaly identification boxes. The annotation tool is also used to mark any single object, with the marking point located at the physical center of the object, forming a calibration point; For objects within the anomaly identification box, their 3D data is extracted, and image processing technology is used to extract the outline of the objects within the anomaly identification box to obtain several sets of closed outlines. The similarity of these sets of closed outlines is compared one by one with the outline of a single object to determine the outline of the topmost object in the stacked objects. The topmost object is then specially marked using a labeling tool. The deep learning model also includes a data analysis unit, which is used to analyze other abnormal situations that occur during the object transportation process. The data analysis unit is connected to the object transportation control system and is used to receive relevant information data about the object transportation. The data analysis unit defines the first unit area in the digital 3D model, counts the number of calibration points in the first unit area, and calculates the object density within the first unit area. When the data analysis unit receives an increase or decrease in the speed of the object being transported, the data analysis unit continues to define the second unit area of the digital 3D model, count the number of calibration points in the second unit area, and analyze the object density within the second unit area. When the speed of object transport increases, the first unit area is larger than the second unit area; when the speed of object transport decreases, the first unit area is smaller than the second unit area. When the density of objects within the first or second unit area is greater than or equal to a set threshold, other abnormal situations are determined to have occurred.
2. The deep learning-based adaptive robot grasping method according to claim 1, characterized in that: The deep learning model sends a digital 3D model containing anomaly bounding boxes, calibration points, and special markers to the control module. The control module then modifies the robot's original grasping scheme and sends the modified grasping scheme to the grasping robot. The grasping robot executes the grasping command to complete the adaptive grasping of objects in abnormal states.
3. The deep learning-based adaptive robot grasping method according to claim 2, characterized in that: When an object being grasped has an abnormal bounding box, the grasping robot grasps the object with the special mark. Then, the RGB-D camera installed on the grasping robot collects three-dimensional data and transmits the collected three-dimensional data back to the control module. The control module continues to process the images of the objects within the abnormal bounding box based on the three-dimensional data, and uses a labeling tool to specially mark the topmost object within the abnormal bounding box until there is no stacking of objects within the abnormal bounding box.
4. The deep learning-based adaptive robot grasping method according to claim 1, characterized in that: When other abnormal situations occur, additional control modules and grasping robots are connected through the external interface of the central control system. The central control system sends data information to different control modules, and at least two grasping robots grasp the objects conveyed on the conveyor belt.
5. The deep learning-based adaptive robot grasping method according to claim 4, characterized in that: At least two grasping robots are arranged in a grasping order. After the first grasping robot completes the grasping of the object, it uses an RGB-D camera to collect three-dimensional data of the object being transported and transmits the collected three-dimensional data to the control module of the next grasping robot. The next grasping robot plans a grasping scheme for the object being transported based on the latest collected three-dimensional data, and collects three-dimensional data again when the next grasping robot completes the grasping. If this grasping robot is the last grasping robot in the sequence, the acquired 3D data is transmitted to the control module of the first grasping robot in the sequence. This allows the control module to complete the cycle of the entire grasping scheme planning and achieve adaptive grasping of the object.
6. A deep learning-based adaptive robot grasping system for implementing the deep learning-based adaptive robot grasping method of claim 2, characterized in that: The RGB-D camera used for creating the digital 3D model is located at the front end of the object transport direction and directly above the object's spatial plane, while the grasping robot is located in the middle of the transport direction and behind the RGB-D camera. The control module includes a capture planning unit and an instruction sending unit; The grasping planning unit is used to replan the grasping scheme of the grasping robot based on the abnormal state of the object to be grasped when the emergency grasping option is enabled, and transmit the grasping scheme to the instruction sending unit. The instruction sending unit sends the revised grasping scheme to the grasping robot, and the grasping robot executes the instruction to adaptively grasp the object in the abnormal state.
7. The deep learning-based adaptive robot grasping system according to claim 6, characterized in that: The grasping robot is also equipped with an RGB-D camera. After grasping an object, the grasping robot collects three-dimensional data through the RGB-D camera and transmits the collected three-dimensional data to the grasping planning unit, which is used to plan the subsequent grasping scheme at the control module.
Citation Information
Patent Citations
Object stacking scene-oriented robot vision grabbing detection method and system and medium
CN119407773A
Garbage classification device and classification method
CN113334368A
Method for identifying stacked special-shaped objects and sorting work station thereof
CN114463634A
Stacked material identification method and device for scattered materials and storage medium
CN118762366A