A robot arm object recognition and grasping method and system based on multi-sensor fusion
Patent Information
- Application Number
- CN202610904572.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-25
AI Technical Summary
然而,这种方法存在固有局限性:首先,视觉系统易受光照、遮挡、物体表面反光等因素干扰,导致识别失败或数据不准;其次,视觉系统仅能提供物体的外部几何信息,对于抓取过程中所需的力度、物体的确切尺寸(尤其在视觉盲区)无法提供反馈,可能导致抓取不稳定或抓取失败
[0038]1、高准确性:通过视觉与触觉两种不同原理的传感器数据进行交叉验证,从根本上避免了单一传感器失效带来的错误,实现了自我校验,保障了识别与抓取的极高准确率。
Smart Images

Figure CN122807866A_ABST
Abstract
Description
Technical Field
[0001] This invention discloses a method and system for object recognition and grasping by a robotic arm based on multi-sensor fusion, belonging to the field of robotics technology. Background Technology
[0002] Robotic arms play a crucial role in modern industrial automation, especially in scenarios such as logistics sorting and warehousing, where accurate object identification and stable grasping are core components.
[0003] Currently, the mainstream solutions mainly rely on machine vision. 2D or 3D cameras acquire the position and appearance information of objects to guide a robotic arm in grasping. However, this method has inherent limitations: First, vision systems are susceptible to interference from factors such as lighting, occlusion, and surface reflections, leading to recognition failures or inaccurate data; second, vision systems can only provide external geometric information about objects and cannot provide feedback on the force required during grasping or the exact size of the object (especially in blind spots), potentially causing unstable or failed grasping.
[0004] Another approach is to introduce force control or haptic sensors to adjust the gripping force based on the feedback. However, these approaches typically only use force sensors to prevent pinching or slipping, without fully utilizing their measurement data to inversely calculate the object's physical dimensions or performing in-depth cross-validation with visual information.
[0005] Therefore, existing technologies lack a method or system that can deeply integrate visual guidance and tactile perception, and utilize multi-source data for cross-verification, thereby ensuring high accuracy in object recognition and grasping. Summary of the Invention
[0006] The purpose of this invention is to solve the problems in the prior art and provide a method and system for object recognition and grasping of a robot arm based on multi-sensor fusion.
[0007] This invention achieves the above objective through the following technical solution: a method for object recognition and grasping by a robotic arm based on multi-sensor fusion, comprising the following steps:
[0008] Visual guidance and initial grasping: Based on the initial three-dimensional spatial data of the target object, control the robot arm to move to the target area and drive the gripping mechanism to perform gripping action on the object with initial parameters;
[0009] Tactile perception and data acquisition: During the execution of the clamping action, the stress value change data fed back by the stress sensor is monitored in real time, and the stroke data of the drive motor that drives the clamping mechanism is recorded simultaneously.
[0010] First dimension calculation: Based on the stress value change data and the stroke data of the drive motor, calculate the first contour dimension data of the target object on the plane perpendicular to the direction of the clamping force;
[0011] Visual precision recognition: After the clamping mechanism holds the object, the robotic arm is controlled to move the object to the preset recognition station. The object image is acquired through the machine vision system, and the second contour dimension data and category information of the object are parsed based on the pre-trained AI recognition model.
[0012] Data fusion and verification: The first contour dimension data and the second contour dimension data are compared and fused. If the data consistency is verified, the object recognition and grasping are determined to be successful.
[0013] Preferably, in the visual guidance and initial grasping step, the initial three-dimensional spatial data is acquired through a front-end 3D vision sensor or directly issued by the upper-level control system.
[0014] Preferably, the first dimension calculation step specifically includes:
[0015] A stress-stroke relationship curve is established, which records the change of stress value with motor stroke during the clamping process;
[0016] Identify the characteristic inflection point in the relationship curve, which corresponds to the instant when the gripping mechanism contacts the object surface and the moment when a stable gripping force is reached;
[0017] Based on the motor stroke data corresponding to the feature inflection point, and combined with the kinematic model of the clamping mechanism, the first contour dimension data of the target object is calculated.
[0018] Preferably, the first contour dimension data includes the object's width, diameter, or cross-sectional contour information.
[0019] Preferably, in the data fusion and verification step, the condition for the data consistency to pass the verification is: the error between the first contour dimension data and the second contour dimension data is less than a preset threshold.
[0020] Preferably, when data consistency verification fails, one or a combination of the following steps are performed:
[0021] Trigger an abnormal alarm to notify the system administrator for intervention;
[0022] Control the robotic arm to place the object into the reset area;
[0023] Initiate the secondary capture and recognition process.
[0024] A robotic arm system for implementing object recognition and grasping methods includes:
[0025] Robotic arm body;
[0026] A gripping mechanism is installed at the end of the robotic arm body;
[0027] A drive motor is used to drive the clamping mechanism to perform opening and closing actions;
[0028] A stress sensor is installed on the clamping mechanism to detect stress changes during the clamping process;
[0029] A machine vision system, including at least one camera, for acquiring images of objects;
[0030] The control system is communicatively connected to the drive motor, the stress sensor, and the machine vision system, and is configured to execute the steps of the robotic arm object recognition and grasping method.
[0031] Preferably, the stress sensor is a distributed stress sensor array, which is attached to the contact surface of the clamping mechanism to sense the pressure distribution on the surface of the object.
[0032] Preferably, the control system includes:
[0033] The motion control module is used to control the movement of the robotic arm and gripping mechanism;
[0034] The data acquisition module is used to receive data from the stress sensor and the motor;
[0035] A visual processing module is used to run the pre-trained AI recognition model;
[0036] The data fusion judgment module is used to perform the data fusion and verification steps.
[0037] Compared with the prior art, the beneficial effects of the present invention are:
[0038] 1. High accuracy: By cross-validating data from sensors based on two different principles, namely vision and touch, errors caused by the failure of a single sensor are fundamentally avoided, achieving self-verification and ensuring extremely high accuracy in recognition and grasping.
[0039] 2. Strong robustness: Tactile perception makes up for the shortcomings of vision in terms of changes in lighting and partial occlusion; visual recognition makes up for the limitations of tactile perception in acquiring information such as object category and color. The two complement each other perfectly, enhancing the system's adaptability in complex environments.
[0040] 3. Maximizing data value: Innovatively elevating stress sensor data from simple "force control" to the level of "dimensional calculation", thus uncovering the deeper value of the data.
[0041] 4. Wide range of applications: This method is particularly suitable for scenarios with extremely high requirements for grasping accuracy, such as handling precision electronic components, sorting high-value goods, and handling mixed materials. Attached Figure Description
[0042] Figure 1 This is a flowchart of a robot arm object recognition and grasping method based on multi-sensor fusion according to the present invention;
[0043] Figure 2 This is a flowchart illustrating the application of the method of the present invention in a logistics parcel sorting robot system;
[0044] Figure 3 This is a schematic diagram of the gripping mechanism and sensor in the robot arm system of the present invention;
[0045] Figure 4 This is a stress-stroke relationship curve in this invention. Detailed Implementation
[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] like Figure 1 As shown, a method for object recognition and grasping using a robotic arm based on multi-sensor fusion is described. This method includes the following steps:
[0048] Visual guidance and initial grasping: First, acquire the initial 3D data of the known object, such as from the upstream system or preliminary scan. Then, based on the initial 3D spatial data of the target object, control the robot arm to move to the target position and drive the gripping mechanism to open the gripper to the preset width and begin to close to grasp the object.
[0049] Tactile sensing and data acquisition: During the clamping process, a high-precision stress sensor starts to work, monitoring the pressure changes on the clamping surface in real time. At the same time, the system accurately records the rotation stroke of the servo motor that drives the clamping opening and closing. These two sets of data are synchronized in time.
[0050] The first dimension of calculation: By analyzing the "stress-stroke" curve, the system can accurately determine when the gripper contacts the object's surface. Contact indicates a sudden increase in gripper stress from zero, and the system indicates that the gripper is properly engaged with the object when the clamping force stabilizes. Based on the motor stroke difference between these two key points, and combined with the gripper's kinematic model (e.g., the proportional relationship of a linkage mechanism or gear rack), the system can accurately calculate the object's actual dimensions in the clamping direction, such as width or diameter. This data comes from physical contact and is unaffected by visual errors.
[0051] Visual precision recognition: After the robotic arm stably grasps the object, it moves it to a recognition station with stable lighting conditions and a single background. A fixed high-resolution camera takes a picture of the object, and the image is recognized by a pre-trained deep learning AI model, such as a convolutional neural network (CNN). The image not only determines the object category, but also calculates the two-dimensional or three-dimensional contour data of the object through image analysis technology.
[0052] Data fusion and verification: The system compares the "first contour dimension data" calculated by touch with the "second contour dimension data" recognized by vision. If the two are consistent within the preset tolerance range, for example, the width error is less than 2mm, the grasping and recognition are considered successful and the robot can proceed with the subsequent placement operation. If they are inconsistent, the abnormal handling process is triggered, such as alarm or re-placement and re-inspection, so as to ensure 100% operation accuracy.
[0053] Reference Figure 2 This method is implemented in a logistics parcel sorting robot system.
[0054] S1: Visual Guidance and Initial Grasping
[0055] The system has initial length, width, and height data of a rectangular cardboard box, for example, obtained from a warehouse management system. The robotic arm moves above the package, the grippers open to a position slightly larger than the width of the package, and then begin to close.
[0056] S2: Tactile Perception and Data Acquisition
[0057] During the closing process, the stress sensor array installed on the inside of the gripper begins to transmit data, and the system synchronously records the encoder readings of the servo motor, such as... Figure 4 As shown, when the left gripper contacts the package, the stress value on the left side reaches its first inflection point, which is designated as point A, the initial contact point; when the right gripper contacts the package, the stress value on the right side reaches its inflection point, and the overall stress value begins to rise steadily; when the clamping force reaches the preset value, it is recorded as the stable gripping point, designated as point B.
[0058] S3: First Dimension Calculation
[0059] The system calculates the total stroke S of the motor between point A and point B. Based on the transmission ratio of the gripper, for example, the gripper closes by 1mm for every 1000 pulses of motor rotation, the actual closing distance L of the gripper can be calculated. Then the actual width W1 of the object is equal to the initial opening width of the gripper - L. The calculated W1 value is the "first contour dimension data".
[0060] S4: Visual Recognition
[0061] The robotic arm holds the package and moves it to a recognition area with a solid-color background. A fixed camera captures an image of the package. A pre-trained AI model, such as a convolutional neural network, identifies the package as a "cardboard box" and calculates the pixel width of the package in the image using an edge detection algorithm. Then, based on the camera calibration parameters, it converts the width into the actual physical width W2, which is the "second contour dimension data". At the same time, it outputs the category information "cardboard box".
[0062] S5: Data Fusion and Validation
[0063] The system compares W1 and W2. If |W - W'| < 5mm, the verification is successful, and the system determines that the object recognition and grasping are successful. The robotic arm places the package into the corresponding sorting basket. If the error exceeds 5mm, the system determines that there may be slippage or visual misrecognition during the grasping, and will trigger the abnormal handling process, such as placing the package on the manual re-inspection station and issuing an alarm.
[0064] This method begins with visually guided initial grasping, then collects data through tactile perception and calculates the first contour dimension, while simultaneously obtaining the second contour dimension through visual precision recognition. Finally, it ensures the absolute accuracy of the operation through data fusion and verification, forming an intelligent closed loop with self-verification capabilities.
[0065] exist Figure 3 The invention demonstrates the key execution components of its system. The drive motor drives the left and right gripper fingers to move in opposite directions or in the same direction through a transmission mechanism. The key feature is that a stress sensor array is integrated on the inner contact surface of the gripper fingers. This array is used to detect changes in pressure distribution on the contact surface with high precision and high resolution when gripping an object, providing raw data for subsequent calculation of the object's contour dimensions.
[0066] exist Figure 4 The paper demonstrates the relationship between stress sensor readings and motor stroke during the clamping process.
[0067] Point A is the initial contact point: the stress value increases significantly from zero or the background noise level, indicating that the gripper has made initial contact with the object surface.
[0068] Point B is the stable gripping point: a force platform where the stress value reaches the preset value, enabling stable gripping of the object without slipping or being damaged.
[0069] Stroke difference ΔS: The difference between the stroke S_A at the contact point and the stroke S_B at the gripping point, i.e., ΔS = S_B - S_A. This ΔS is directly related to the actual distance the gripper closes to accommodate the object. By combining the kinematic model of the gripper, the actual size of the object can be accurately calculated, such as width W = initial opening width - K·ΔS, where K is the transmission coefficient.
[0070] This embodiment verifies that the method of the present invention can effectively utilize multi-sensor fusion technology to ensure a high degree of accuracy in object recognition and grasping.
[0071] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0072] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for object recognition and grasping using a robotic arm based on multi-sensor fusion, characterized in that, Includes the following steps: Visual guidance and initial grasping: Based on the initial three-dimensional spatial data of the target object, control the robot arm to move to the target area and drive the gripping mechanism to perform gripping action on the object with initial parameters; Tactile perception and data acquisition: During the execution of the clamping action, the stress value change data fed back by the stress sensor is monitored in real time, and the stroke data of the drive motor that drives the clamping mechanism is recorded simultaneously. First dimension calculation: Based on the stress value change data and the stroke data of the drive motor, calculate the first contour dimension data of the target object on the plane perpendicular to the direction of the clamping force; Visual precision recognition: After the clamping mechanism holds the object, the robotic arm is controlled to move the object to the preset recognition station. The object image is acquired through the machine vision system, and the second contour dimension data and category information of the object are parsed based on the pre-trained AI recognition model. Data fusion and verification: The first contour dimension data and the second contour dimension data are compared and fused. If the data consistency is verified, the object recognition and grasping are determined to be successful.
2. The method for object recognition and grasping by a robotic arm based on multi-sensor fusion according to claim 1, characterized in that, In the visual guidance and initial grasping step, the initial three-dimensional spatial data is acquired through a front-end 3D vision sensor or directly issued by the upper-level control system.
3. The method for object recognition and grasping by a robotic arm based on multi-sensor fusion according to claim 1, characterized in that, The first dimension calculation steps specifically include: A stress-stroke relationship curve is established, which records the change of stress value with motor stroke during the clamping process; Identify the characteristic inflection point in the relationship curve, which corresponds to the instant when the gripping mechanism contacts the object surface and the moment when a stable gripping force is reached; Based on the motor stroke data corresponding to the feature inflection point, and combined with the kinematic model of the clamping mechanism, the first contour dimension data of the target object is calculated.
4. The method for object recognition and grasping by a robotic arm based on multi-sensor fusion according to claim 1, characterized in that, The first contour dimension data includes the object's width, diameter, or cross-sectional contour information.
5. The method for object recognition and grasping by a robotic arm based on multi-sensor fusion according to claim 1, characterized in that, In the data fusion and verification step, the condition for the data consistency to pass the verification is that the error between the first contour dimension data and the second contour dimension data is less than a preset threshold.
6. The method for object recognition and grasping by a robotic arm based on multi-sensor fusion according to claim 5, characterized in that, When data consistency verification fails, perform one or a combination of the following steps: Trigger an abnormal alarm to notify the system administrator for intervention; Control the robotic arm to place the object into the reset area; Initiate the secondary capture and recognition process.
7. A robotic arm system for implementing the method of any one of claims 1-6, characterized in that, include: Robotic arm body; A gripping mechanism is installed at the end of the robotic arm body; A drive motor is used to drive the clamping mechanism to perform opening and closing actions; A stress sensor is installed on the clamping mechanism to detect stress changes during the clamping process; A machine vision system, including at least one camera, for acquiring images of objects; A control system, communicatively connected to the drive motor, the stress sensor, and the machine vision system, is configured to perform the steps of the method according to any one of claims 1-6.
8. The robotic arm system according to claim 7, characterized in that, The stress sensor is a distributed stress sensor array, which is attached to the contact surface of the clamping mechanism to sense the pressure distribution on the object surface.
9. The robotic arm system according to claim 7, characterized in that, The control system includes: The motion control module is used to control the movement of the robotic arm and gripping mechanism; The data acquisition module is used to receive data from the stress sensor and the motor; A visual processing module is used to run the pre-trained AI recognition model; The data fusion judgment module is used to perform the data fusion and verification steps.