Robot follow-up grabbing method based on proximity sensing and related equipment thereof

By deploying a photoelectric sensor array and a dynamic feature encoding convolutional neural network on the inside of the robot's end gripper, combined with the ROS architecture, the perception blind spot and response delay problems of traditional robot grasping systems are solved, high-precision follow-up grasping operations are achieved, and the robot's operating reliability in complex environments is improved.

CN120620179APending Publication Date: 2025-09-12QINGDAO UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510720681.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional robotic grasping systems rely on depth cameras and tactile sensors, which leads to perception blind spots and response delays, making it difficult to achieve real-time, high-precision posture perception and reliable control in dynamic grasping tasks.

Method used

By deploying a photoelectric sensor array on the inside of the robot's end gripper, proximity perception signals are acquired in real time. Combined with the dynamic feature encoding convolutional neural network and ROS distributed architecture, the 5D pose parameters in the gripper coordinate system are output, and joint angle control instructions are generated through inverse kinematics solution to achieve high-precision follow-up grasping.

Benefits of technology

It makes up for the blind spots of visual sensors, improves the accuracy and robustness of dynamic object pose estimation, ensures the robot's motion planning and execution efficiency in complex environments, and improves safety and operational reliability in unstructured scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120620179A_ABST
    Figure CN120620179A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent robots, and discloses a robot follow-up grabbing method based on proximity sensing and related equipment thereof. The method comprises the following steps: acquiring an approaching sensing signal between a target object and a clamping jaw in real time through a photoelectric sensor array deployed on the inner side of the clamping jaw at the tail end of the robot; performing filtering and noise reduction processing on the proximity sensing signal to obtain a preprocessed proximity sensing signal; the preprocessed proximity sensing signals are processed through the trained dynamic feature coding convolutional neural network, and 5D pose parameters of the target object under the clamping jaw coordinate system are output; the 5D pose parameters are converted into a target pose under a robot base coordinate system through coordinate conversion nodes of an ROS distributed architecture, and a joint angle control instruction is generated through inverse kinematics solution; and a joint angle control instruction is dynamically adjusted based on the real-time pose deviation, and the robot is driven to complete follow-up grabbing operation. Based on the method, high-precision follow-up grabbing operation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent robot technology, and in particular to a robot follow-up grasping method based on proximity perception and related equipment. Background Art

[0002] With the rapid development of industrial automation and intelligent robotics, robotic grasping plays a vital role in scenarios such as smart manufacturing, warehousing and logistics, and human-robot collaboration. However, in complex, unstructured environments, traditional robotic grasping systems generally rely on the collaborative work of depth camera imaging and tactile sensors. The vision system plans the grasping path, while the tactile sensor verifies the contact state. While this staged perception architecture improves grasping accuracy to a certain extent, it is still inherently limited by the visual sensor's field of view and the tactile sensor's response latency.

[0003] Specifically, depth cameras struggle to cover obscured areas of an object's surface or back, leading to blind spots. Tactile sensors, on the other hand, are unable to promptly report pose deviations due to response lag when the robot's end-effector approaches the target. This synergistic effect of "visual blind spots" and "tactile lag" can easily lead to failures in the dynamic grasping process, leading to issues such as grasping strategy failure and increased collision risk. Furthermore, existing deep learning-based pose estimation models rely on large amounts of precisely labeled data, while labeling the six-degree-of-freedom pose of dynamic targets is costly and difficult, further limiting the algorithm's generalization and deployment efficiency in real-world scenarios.

[0004] Therefore, existing technologies rely on visual and tactile sensors, resulting in perception blind spots and response delays, making it difficult to achieve real-time, high-precision posture perception and reliable control in dynamic grasping tasks. Summary of the Invention

[0005] The embodiments of the present invention provide a robot follow-up grasping method based on proximity perception and related equipment, which at least solve the problem in related technologies that reliance on visual and tactile sensors leads to perception blind spots and response delays, making it difficult to achieve real-time high-precision posture perception and reliable control in dynamic grasping tasks.

[0006] According to a first aspect of an embodiment of the present invention, a robot follow-up grasping method based on proximity perception is provided, comprising:

[0007] The photoelectric sensor array deployed inside the end gripper of the robot acquires the proximity sensing signal between the target object and the gripper in real time. The proximity sensing signal is used to represent the two-dimensional position information and distance information of the target object.

[0008] Sampling, filtering, noise reduction, and normalization processing are performed on the proximity sensing signal to obtain a preprocessed proximity sensing signal;

[0009] Processing the pre-processed proximity sensing signal through a trained dynamic feature encoding convolutional neural network to output 5D pose parameters of the target object in a gripper coordinate system, wherein the 5D pose parameters include plane position coordinates and quaternion posture;

[0010] The 5D pose parameters are converted into the target pose in the robot base coordinate system through the coordinate conversion node of the ROS distributed architecture, and the joint angle control instructions are generated through inverse kinematics solution;

[0011] The joint angle control instruction is dynamically adjusted based on the real-time posture deviation to drive the robot to complete the follow-up grasping operation. The real-time posture deviation includes the deviation between the target posture and the actual posture of the end effector.

[0012] According to a second aspect of an embodiment of the present invention, a robot follow-up grasping method based on proximity perception is provided, comprising:

[0013] An acquisition module is configured to acquire, in real time, a proximity sensing signal between a target object and the gripper through a photoelectric sensor array disposed inside the gripper at the end of the robot. The proximity sensing signal is used to represent two-dimensional position information and distance information of the target object.

[0014] a processing module, configured to perform sampling, filtering, noise reduction, and normalization processing on the proximity sensing signal to obtain a preprocessed proximity sensing signal;

[0015] The processing module is further configured to process the pre-processed proximity sensing signal through a trained dynamic feature encoding convolutional neural network, and output 5D pose parameters of the target object in a gripper coordinate system, wherein the 5D pose parameters include plane position coordinates and quaternion pose;

[0016] A conversion module is used to convert the 5D pose parameters into the target pose in the robot base coordinate system through the coordinate conversion node of the ROS distributed architecture, and generate joint angle control instructions through inverse kinematics solution;

[0017] An adjustment module is used to dynamically adjust the joint angle control instruction based on the real-time posture deviation to drive the robot to complete the follow-up grasping operation. The real-time posture deviation includes the deviation between the target posture and the actual posture of the end effector.

[0018] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, comprising: a processor, and a memory storing a program, wherein the program comprises instructions that, when executed by the processor, cause the processor to execute the method according to the first aspect.

[0019] According to a fourth aspect of an embodiment of the present invention, a non-transitory machine-readable medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method according to the first aspect.

[0020] Beneficial effects of the embodiments of the present invention:

[0021] A robot follow-up grasping method based on proximity perception provided by an embodiment of the present invention obtains the proximity perception signal between the target object and the gripper in real time through a photoelectric sensor array deployed on the inner side of the robot's end gripper. The signal can accurately characterize the two-dimensional position information and distance information of the object, thereby making up for the perception blind spot problem of traditional visual sensors in occluded areas. By sampling, filtering, denoising and normalizing the proximity perception signal, and combining the dynamic feature encoding convolutional neural network to perform deep feature extraction and pose regression on the processed signal, the 5D pose parameters (including plane position coordinates and quaternion pose) in the gripper coordinate system can be output efficiently and accurately, thereby improving the accuracy and robustness of the dynamic object pose estimation. Furthermore, the pose parameters in the gripper coordinate system are mapped to the robot base coordinate system through the coordinate conversion node of the ROS distributed architecture, and the joint angle control instructions are generated in combination with the inverse kinematics solution, thereby ensuring the robot's motion planning and execution efficiency in complex dynamic environments. Ultimately, the closed-loop feedback based on real-time posture deviation dynamically adjusts the joint angle control instructions, allowing the robot to quickly respond to changes in the position and posture of the target object and achieve high-precision follow-up grasping operations. This effectively solves the grasping failure problem caused by visual blind spots and tactile lag in traditional methods, and significantly improves the safety and operational reliability of the robot in unstructured scenarios.

[0022] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below so that other features, objects, and advantages of the invention are more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be derived from these drawings without inventive effort.

[0024] Figure 1 A flowchart of a robot follow-up grasping method based on proximity perception provided by an embodiment of the present invention.

[0025] Figure 2 A schematic diagram of a proximity sensor circuit based on a photoelectric sensor provided in an embodiment of the present invention.

[0026] Figure 3 A schematic diagram of a coordinate system definition for a micro-photoelectric sensor element provided by an embodiment of the present invention.

[0027] Figure 4 A schematic diagram of tilt detection using a proximity sensor provided by an embodiment of the present invention.

[0028] Figure 5 A schematic diagram of the layout of a photoelectric sensor array in a proximity sensor provided by an embodiment of the present invention.

[0029] Figure 6 A schematic diagram of the installation of a three-finger flexible gripper with an integrated proximity sensor provided by an embodiment of the present invention.

[0030] Figure 7 A flow chart of a training method for a dynamic feature encoding convolutional neural network provided by an embodiment of the present invention.

[0031] Figure 8 A schematic diagram of a tracker and a preset object coordinate system provided by an embodiment of the present invention.

[0032] Figure 9 A schematic diagram of various coordinate systems and their relationships within a robot grasping system provided by an embodiment of the present invention.

[0033] Figure 10 A schematic diagram of the orientation of a hexagonal prism gripping surface provided in an embodiment of the present invention.

[0034] Figure 11 A schematic diagram of the architecture of a dynamic feature encoding convolutional neural network provided by an embodiment of the present invention.

[0035] Figure 12 A schematic diagram of the processing flow of a dynamic feature encoding module provided in an embodiment of the present invention.

[0036] Figure 13 A schematic diagram of the architecture of a time series feature extraction module provided in an embodiment of the present invention.

[0037] Figure 14 A schematic diagram of the architecture of a spatial feature extraction module provided by an embodiment of the present invention.

[0038] Figure 15 A schematic diagram of the software architecture of a ROS-based robot follow-up grasping system provided in an embodiment of the present invention.

[0039] Figure 16 The pitch angle and position output x provided by the embodiment of the present invention c Schematic diagram of the relationship.

[0040] Figure 17The roll angle and position output y provided by the embodiment of the present invention c Schematic diagram of the relationship.

[0041] Figure 18 A schematic diagram of distance measurement performance experimental results provided by an embodiment of the present invention.

[0042] Figure 19 A schematic diagram of the voltage rise time when an LED is turned on, provided by an embodiment of the present invention.

[0043] Figure 20 A schematic diagram of voltage drop time when an LED is turned off provided by an embodiment of the present invention.

[0044] Figure 21 A schematic diagram of a loss curve for posture prediction provided in an embodiment of the present invention.

[0045] Figure 22 A schematic diagram of a one-way translation follow-up grasping process provided by an embodiment of the present invention.

[0046] Figure 23 A schematic diagram of a compound motion follow-up grasping process provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The following describes embodiments of the present invention in more detail with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0048] In complex, unstructured environments, traditional robotic grasping systems rely on the collaborative operation of depth cameras and tactile sensors. However, due to limited visual field and delayed tactile response, they are prone to blind spots and coordination failures, leading to grasping failures. Furthermore, the difficulty of labeling dynamic objects in deep learning-based pose estimation models hinders the generalization and deployment of algorithms, making it difficult for existing technologies to meet the real-time, high-precision pose perception and reliable control requirements of dynamic grasping tasks.

[0049] In order to solve the above problems, an embodiment of the present invention provides a robot follow-up grasping method based on proximity perception and related equipment, which collects proximity signals through a photoelectric sensor array to make up for the visual blind spot, and outputs 5D posture parameters in the gripper coordinate system through signal processing and convolutional neural network; then uses the ROS architecture to convert coordinates, combines inverse kinematics to generate control instructions, and dynamically adjusts through closed-loop feedback to achieve high-precision follow-up grasping, solve the grasping failure problem of traditional methods, and improve operation reliability.

[0050] Figure 1 The flowchart of a robot follow-up grasping method based on proximity perception provided by an embodiment of the present invention. Figure 1 As shown, the method includes the following steps.

[0051] In step S101, a photoelectric sensor array is deployed inside the end gripper of the robot to obtain a proximity sensing signal between the target object and the gripper in real time. The proximity sensing signal is used to represent the two-dimensional position information and distance information of the target object.

[0052] Step S102: filtering and denoising the proximity sensing signal to obtain a pre-processed proximity sensing signal.

[0053] In step S103, the pre-processed proximity sensing signal is processed by the trained dynamic feature encoding convolutional neural network to output the 5D pose parameters of the target object in the gripper coordinate system. The 5D pose parameters include plane position coordinates and quaternion pose.

[0054] In step S104, the 5D pose parameters are converted into the target pose in the robot base coordinate system through the coordinate conversion node of the ROS distributed architecture, and the joint angle control instructions are generated through inverse kinematics solution.

[0055] Step S105 , dynamically adjusting the joint angle control instructions based on the real-time posture deviation to drive the robot to complete the follow-up grasping operation. The real-time posture deviation includes the deviation between the target posture and the actual posture of the end effector.

[0056] First, a photoelectric sensor array deployed inside the robot's end gripper acquires a proximity sensing signal between the target object and the gripper in real time. In this embodiment, the proximity sensing signal is used to represent the two-dimensional position and distance information of the target object.

[0057] In this embodiment, a photoelectric sensor array is deployed inside the robot's end gripper to detect the relative position of the target object and the gripper in real time. Based on the principle of non-contact detection, this photoelectric sensor array emits infrared light and receives light signals reflected from the object's surface, generating a proximity sensing signal representing the target object's two-dimensional position and distance information.

[0058] Deploying the photoelectric sensor array on the inside of the gripper can avoid the blind spot problem of traditional visual sensors caused by occlusion of the manipulator. At the same time, combined with the fast response characteristics of the photoelectric sensor (response time less than 1ms), it can achieve accurate positioning during the dynamic process of the robot approaching the target object, providing high-precision, low-latency raw data for subsequent pose estimation.

[0059] After obtaining the proximity sensing signal between the target object and the gripper, the proximity sensing signal can be sampled, filtered, denoised and normalized to obtain a preprocessed proximity sensing signal.

[0060] In this embodiment, the unprocessed raw proximity sensing signal may contain environmental interference (such as stray light, electromagnetic noise) and sensor noise. Therefore, the raw proximity sensing signal can be subjected to filtering, noise reduction, and normalization.

[0061] Preprocessing the proximity sensing signal more accurately reflects the actual position and posture changes of the target object, providing high-quality input data for subsequent neural network processing. This step balances noise suppression with signal fidelity, avoiding the loss of dynamic information caused by excessive filtering.

[0062] After obtaining the pre-processed proximity sensing signal, it can be input into the trained dynamic feature encoding convolutional neural network. The trained dynamic feature encoding convolutional neural network processes the pre-processed proximity sensing signal and outputs the 5D pose parameters of the target object in the gripper coordinate system. The 5D pose parameters include the plane position coordinates and the quaternion posture.

[0063] In this embodiment, the pre-trained Dynamic Feature Encoding Convolutional Neural Network (DFENet) can receive the pre-processed proximity sensing signal and output the 5D pose parameters of the target object in the gripper coordinate system through the following process:

[0064] DFENet first performs a first-order difference operation on the input preprocessed proximity perception signal to extract the speed characteristics of the target object's motion, enhance the model's ability to capture dynamic changes, and process the dynamic features through a one-dimensional convolution layer. The original features and the dynamic features after convolution processing are fused to generate fused features; the fused features are then input into the spatiotemporal feature extraction module, and temporal features are extracted through convolution, CBAM attention mechanism, multi-scale hole convolution and maximum pooling. The spatial features are then divided through a grouped convolution strategy, local patterns are learned, and spatiotemporal features are extracted by combining CBAM attention mechanism, convolution and adaptive pooling; finally, the high-level spatiotemporal features are mapped to 5D pose parameters through a double fully connected layer.

[0065] In this embodiment, the training of the neural network can be based on a large-scale data set covering different object postures and motion scenarios to ensure the generalization ability of the model in dynamic environments.

[0066] Subsequently, the 5D pose parameters can be converted into the target pose in the robot base coordinate system through the coordinate conversion node of the ROS distributed architecture, and the joint angle control instructions can be generated through inverse kinematics solution.

[0067] In this embodiment, the system converts the pose parameters from the gripper coordinate system to the robot base coordinate system through the Robot Operating System (ROS) distributed architecture.

[0068] Specifically, a homogeneous transformation matrix (containing a rotation matrix and a translation vector) is first used to map the gripper's pose to the robot's base coordinate system, ensuring that the robot can plan the grasping path. The target pose is then converted into the angular configuration of the robot's joints.

[0069] Finally, the joint angle control instructions are dynamically adjusted based on the real-time posture deviation to drive the robot to complete the follow-up grasping operation. The real-time posture deviation includes the deviation between the target posture and the actual posture of the end effector.

[0070] In this embodiment, the system dynamically adjusts the grasping action through a closed-loop feedback mechanism. Specifically, during the deviation calculation phase, the target pose (calculated from the output signal of the photoelectric sensor array) is compared with the actual pose of the end effector in real time to generate a pose deviation signal. Then, based on the deviation signal, the joint angle correction is calculated and the correction instruction is sent to each joint servo mechanism. During the execution process, the robot continuously senses and corrects errors, ultimately achieving precise follow-up grasping.

[0071] In an embodiment of the present invention, a near-field perception layer is constructed by integrating highly sensitive photoelectric sensors inside the gripper, overcoming the limitations of visual perception with high precision and fast response. With the help of a dynamic feature encoding convolutional neural network (DFENet), motion features and spatiotemporal feature extraction are integrated to solve the problem of dynamic object pose estimation. Based on the ROS distributed architecture, efficient collaboration of perception, computation, and execution is achieved, and closed-loop feedback is used to dynamically correct the trajectory. The method provided by the embodiment of the present invention, with its innovations in hardware, algorithms, and system architecture, can propel robotic follow-up grasping from passive response to active perception and decision-making, enhancing its autonomy and safety in complex scenarios.

[0072] In an alternative embodiment, the photosensor array consists of three rows and two columns of microphotoelectric sensors. Each microphotoelectric sensor detects the intensity of reflected light from a target object based on the infrared reflection principle and outputs a proximity sensing signal through a resistor network. The photosensor array has an operating range of 0.6 to 40 mm, a response time of less than 1 ms, and calculates two-dimensional position and distance information based on the center position and total current of the photocurrent distribution.

[0073] In this embodiment, the photoelectric sensor array achieves high-precision real-time perception of the two-dimensional position information and distance information of the target object through a compact layout design of three rows and two columns, a signal integration mechanism of a resistor network, and a mathematical calculation method of photocurrent distribution.

[0074] Specifically, the sensor array consists of six miniature photoelectric sensors arranged in a rectangular pattern of three rows and two columns. This structure not only ensures coverage of the detection area, but also reduces the system complexity and size by reducing the number of sensors, thus solving the installation blind spot problem caused by the complex wiring and rigid structure of traditional proximity sensors. Each miniature photoelectric sensor operates based on the principle of infrared reflection and may contain an infrared light-emitting diode and a phototransistor. It achieves perception by emitting infrared light and receiving the intensity of the reflected light from the surface of the target object. When the target object approaches, the reflected light intensity changes with distance and shows a specific pattern. The sensor detects this change to infer the object's position and distance.

[0075] To integrate the signals from multiple sensors into a usable proximity sensing signal, a resistor network design can be used. This network, constructed using resistors in layers A and B, converts the output voltage signals from the six photosensors in the middle layer into voltage values ​​for four electrodes: E1, E2, E3, and E4. This structure not only simplifies external wiring requirements (requiring only eight wires), but also enables efficient processing of photocurrent distribution through weighted calculations within the resistor network.

[0076] Specifically, the photocurrent centroid position (x c ), and the voltage difference between the electrodes E2 and E4 in the B layer is used to calculate the center of mass position in the y-axis direction (y c ), these two coordinate values ​​together constitute the two-dimensional position information of the target object. At the same time, the total current (I all ) can be used to deduce the distance between the object and the sensor, where the total current is proportional to the intensity of the reflected light, which in turn is inversely proportional to the distance to the object. In practical applications, the operating range of the sensor array can be set to 0.6 to 40 mm. This range covers the dynamic adjustment phase before the robot gripper contacts the target object, and can effectively detect the object's tilt angle (determined by the center of mass offset) and small displacement (calculated by the distance change). Its response time of less than 1 ms ensures the system's real-time feedback capability for dynamic objects, avoiding the problem of grasping failure caused by response delays in traditional tactile sensors.

[0077] Figure 2 A schematic diagram of a proximity sensor circuit based on a photoelectric sensor provided in an embodiment of the present invention.

[0078] In this embodiment, the proximity sensing system based on photoelectric sensors uses a resistor network integrated with micro photoelectric sensors to construct a three-row and two-column photoelectric sensor array. The photoelectric sensor works based on the infrared reflection principle and consists of a light source (such as an infrared light emitting diode), a light receiver (such as a phototransistor) and a signal processing unit. Figure 2 As shown in FIG, the proximity sensor structure adopts a three-layer design, with layer A and layer B connected by a photoelectric sensor layer in the middle.

[0079] A-layer structure: The collector terminals of the phototransistors are connected in series in the y-axis direction and then connected in the x-axis direction through a resistor r. The two ends are connected to a positive voltage (+V0) through the resistor r and an external resistor (R0), respectively.

[0080] Layer B structure: The emitter terminals of the phototransistors are connected in series in the x-axis direction and connected in the y-axis direction through a resistor r. The two ends are connected to a negative voltage (-V0) through the resistor r and an external resistor (R0), respectively.

[0081] Based on the above structure, the sensor can effectively detect objects through only four voltage outputs (electrodes E1, E3, E2, and E4).

[0082] When working, the light emitted by the infrared light emitting diode illuminates the surface of the object, and is reflected and attenuated according to the distance from the object. The phototransistor detects the reflected infrared light and converts it into photocurrent. The target information is determined by calculating the distribution of photocurrent:

[0083] Two-dimensional position: Calculated using the coordinates of the center point of the photocurrent distribution. Specifically, the center of mass of the proximity sensor's photocurrent distribution along the x-axis is calculated using the output voltages of electrodes E1 and E3 at both ends of layer A. The center of mass of the photocurrent distribution along the y-axis is calculated using the output voltages of electrodes E2 and E4 at both ends of layer B. The center of the photocurrent distribution is then determined.

[0084] Distance: Calculated by the total amount of photocurrent distribution.

[0085] Figure 3 A schematic diagram of a coordinate system definition for a micro-photoelectric sensor element provided by an embodiment of the present invention.

[0086] like Figure 3 As shown, the microphotoelectric sensor elements are arranged in an m×n grid structure, with each element identified by an index (i, j), where i = 1, 2, 3…, m, and j = 1, 2, 3…, n. This coordinate system is a Cartesian o-xy coordinate system, where i corresponds to the x-axis and j corresponds to the y-axis. In this grid layout, the (x, y) coordinates of the lower left corner element are defined as (-1, -1), and the (x, y) coordinates of the upper right corner element are (1, 1).

[0087] When the photosensor elements are arranged at uniform intervals, the coordinates (x, y) of the element (i, j) can be calculated using the following formula (1):

[0088]

[0089] The above formula (1) clarifies the specific coordinate calculation method of each element (i, j) in the Cartesian coordinate system o-xy in the evenly spaced m×n grid, providing a coordinate basis for subsequent position detection and analysis based on the sensor array.

[0090] At the signal processing level, the proximity sensor calculates the analog output through a resistor network, and its output is the voltage V on the four electrodes E1, E2, E3 and E4. out1 、V out2 、V out3 、V out4 Based on these voltages, the total current and the center position of the photocurrent distribution can be measured.

[0091] The center position of the photocurrent distribution in the x direction c It can be calculated based on the following formula (2):

[0092]

[0093] The center position of the photocurrent distribution in the y direction c It can be calculated based on the following formula (3):

[0094]

[0095] Total current I all It can be calculated based on the following formula (4):

[0096]

[0097] The first-order moment of the photocurrent distribution in the x direction I x It can be calculated based on the following formula (5):

[0098]

[0099] The first-order moment of the photocurrent distribution in the y direction I y It can be calculated based on the following formula (6):

[0100]

[0101] Wherein, V0 is the power supply voltage of the sensor, r is the internal resistance, R0 is the external resistance, m is the number of micro-photoelectric sensor elements in the x direction, and n is the number of micro-photoelectric sensor elements in the y direction. These values ​​are all constants.

[0102] In the embodiments of the present invention, hardware-level calculations using a resistor network reduce reliance on subsequent software processing, thereby improving the real-time performance of the overall system. Furthermore, the miniaturization of the sensor array (for example, using the Omron EE-SY1200 sensor) enables its flexible installation inside the robot gripper, avoiding the blind spots caused by occlusion in traditional depth cameras. Its low power consumption and small size also meet the demand for compact actuators in industrial scenarios.

[0103] In this embodiment, when there is reflected light on the surface of the object and the reflectivity is evenly distributed, the output of the proximity sensing system is It will change with the distance from the object surface. Since the intensity of infrared reflected light decreases with the increase of the distance to the target object, the output of the sensing system Approximately inversely proportional to distance.

[0104] x calculated from the proximity sensor system output c and y c The value is based on the center of the sensor array and its value range is [-1,1]. When the object to be measured is small relative to the size of the sensor array, these two values ​​will change with the change of the center position of the object; when the surface of the object to be measured is large, x c and y c It depends on the inclination of the sensor array relative to the surface of the object being measured.

[0105] Figure 4 A schematic diagram of tilt detection using a proximity sensor system provided by an embodiment of the present invention.

[0106] like Figure 4 As shown in (a), if the sensor array is parallel to the object surface (0°), the photocurrents of each phototransistor in the system are equal, and the center of the photocurrent distribution is the origin. At this time, y c is zero.

[0107] like Figure 4 As shown in (b), when the sensor array is not parallel to the object surface, the photocurrent distribution is tilted and the center position moves to the side closer to the object, so the y c Detects tilt in the roll angle direction.

[0108] Similarly, if Figure 4 As shown in (d), it can be obtained by x c Detects the tilt in the pitch angle direction.

[0109] In practical applications, proximity sensors can use the Omron ee-sy1200 miniature photoelectric sensor, which integrates six sensor elements on a flexible printed circuit (FPC) board, forming a three-row, two-column rectangular array. Based on the infrared reflection principle, this sensor consists of a light-emitting diode (LED) and a phototransistor pair. It senses the position of the target object by emitting infrared light and receiving reflected signals from its surface.

[0110] To accommodate the limited space inside the robot finger, the sensor array adopts a compact design to ensure that the detection area inside the gripper is covered while maintaining a lightweight structure. Figure 5 shown.

[0111] Each photoelectric sensor in the proximity sensor is connected via a resistor network. Layers A and B are connected to the collector and emitter of the phototransistor, respectively. Both ends are connected to a power supply via external resistors, forming four voltage output channels (E1-E4). When a target object approaches, the reflected light intensity changes, causing the currents of each photoelectric sensor to differ. The two-dimensional position coordinates are calculated using the calculation method described in the previous embodiment, and distance information is derived from the total current.

[0112] The installation diagram of the three-finger flexible manipulator with integrated proximity sensor is as follows Figure 6 As shown, the sensor module is fixed to the inside of a three-finger flexible gripper via an FPC board, ensuring real-time monitoring of the target object's tilt angle and dynamic displacement during the robot's grasping process. Experiments have shown that the proximity sensor provides stable distance measurement within a range of 0.6 to 40 mm, with a response time of less than 1 ms.

[0113] In an optional embodiment, the proximity sensing signal is sampled, filtered, denoised, and normalized to generate a preprocessed proximity sensing signal. This includes sampling and filtering the proximity sensing signal at a sampling frequency of 60 Hz to eliminate environmental interference and sensor noise, followed by normalization. The normalized proximity sensing signal is divided into time series windows (step size 1) of eight consecutive time steps. The eight-step time series signals from three distributed photoelectric sensor arrays (each with four voltage signal outputs) inside the gripper are recombined into a 12-channel spatiotemporal tensor to generate the preprocessed proximity sensing signal. After filtering, the input data is normalized, linearly mapping the voltage values ​​of each channel to the interval [0, 1]. The core purpose of this step is to eliminate differences in signal amplitude between sensors or under different operating conditions, such as output fluctuations caused by changes in light source intensity or sensor aging, while accelerating the convergence of the neural network. Through this processing, the dynamic range of the input data is unified, avoiding the problem of model training being dominated by the excessively large numerical range of some channels, while also improving the model's robustness to posture changes of different target objects.

[0114] Finally, the system divides the normalized signal into time series windows (step size 1) of eight consecutive time steps. The eight-step time series signals from three distributed photoelectric sensor arrays (each with four voltage signal outputs) inside the gripper are recombined into 12 channels of spatiotemporal input data. The 12 channels are derived from the hardware design of the proximity sensing unit. The system utilizes three photoelectric sensor arrays (each with four voltage signal outputs), resulting in a total of 3 × 4 = 12 signal channels. The time series windowing logic arranges the eight time-step signals sequentially to form a time series. The 12-channel time series signals are then converted into a two-dimensional spatiotemporal tensor (number of channels × number of time steps). This recombining approach not only preserves the signal's temporal characteristics but also, through multi-channel parallel input, enhances the neural network's ability to perceive spatially distributed features, such as the differential responses of different sensor locations to the target object's distance and position.

[0115] This embodiment, through a phased signal processing flow, not only addresses the vulnerability of proximity sensing signals to environmental interference but also provides structured, standardized input data for the subsequent pose estimation network model. Temporal windowing constructs a spatiotemporal feature representation suitable for deep learning models, while normalization optimizes the stability and efficiency of model training, laying a reliable data foundation for robotic follow-up grasping tasks.

[0116] In an optional embodiment, the dynamic feature coding convolutional neural network includes a dynamic feature coding module, a spatiotemporal feature extraction module and a posture regression module.

[0117] The Dynamic Feature Encoding Convolutional Neural Network (DFENet) in this embodiment is a deep learning architecture designed for the task of estimating the pose of dynamic objects within a robot hand. Its core is to achieve high-precision prediction of 5D pose parameters (planar position x, y and quaternion attitude) through the collaborative processing of proximity perception signals by multiple modules. The network consists of three components: a dynamic feature encoding module, a spatiotemporal feature extraction module, and a pose regression module. The overall structure not only captures temporal dynamic information but also strengthens the local and global representation of spatial features, effectively improving pose estimation accuracy in dynamic scenarios.

[0118] In this embodiment, the preprocessed proximity sensing signal is processed by a pre-trained dynamic feature encoding convolutional neural network to output the 5D pose parameters of the target object in the gripper coordinate system, including the following steps.

[0119] First, the preprocessed proximity sensing signal is input into the dynamic feature encoding module, and the first-order difference of the preprocessed proximity sensing signal is performed to extract the dynamic features. The dynamic features are then processed through a one-dimensional convolution layer, and the original features and the dynamic features after convolution processing are fused to generate fused features.

[0120] The dynamic feature encoding module, serving as the network's input processing unit, first extracts dynamic features from the preprocessed proximity sensing signal. The input preprocessed signal is partitioned using a sliding window to form time series data with the shape [B, K, T, 1] (B is the batch size, K is the number of sensor features, and T is the time step). This is converted to the shape [B, K, T] by removing the last dimension. The module then performs first-order differences on the time series of each sensor feature to extract object motion information. These dynamic features reflect the real-time motion trends of the target object within the gripper coordinate system. The first-order differenced dynamic features are processed in a one-dimensional convolutional layer (Conv1d) and fused with the original features. The convolution kernel size is 3, and padding is used to maintain feature dimensionality. This design not only preserves the spatial distribution of the original signal but also concatenates the dynamic and static features along the channel dimension to generate fused features. This fused feature simultaneously represents both the static position information and dynamic characteristics of the object, providing a richer input for subsequent spatiotemporal feature extraction.

[0121] Then, the fused features are input into the spatiotemporal feature extraction module, and the temporal features are extracted through convolution, CBAM attention mechanism, multi-scale dilated convolution and maximum pooling. The spatial features are then divided through the grouped convolution strategy to learn local patterns, and the spatiotemporal features are extracted by combining the CBAM attention mechanism, convolution and adaptive pooling.

[0122] In the spatiotemporal feature extraction module, the network extracts temporal features by combining multi-scale dilated convolutions with the Convolutional Block Attention Module (CBAM) attention mechanism. This multi-scale dilated convolution design enables the network to capture motion patterns at different time scales. By adjusting the dilation rate of the dilated holes, the network can expand its receptive field, thereby enhancing its ability to model temporal dependencies without increasing the number of parameters. Furthermore, the CBAM attention mechanism is introduced to dynamically adjust the importance of different regions in the feature map. CBAM consists of a channel attention module and a spatial attention module. This dual attention mechanism enables the network to focus on key features relevant to pose estimation, suppress irrelevant noise, and improve the robustness of feature representation. Furthermore, the spatial feature extraction module employs a grouped convolution strategy, partitioning the 64-dimensional features into 16 response units, each containing four channels. Local patterns are learned using independent convolution kernels, while the CBAM mechanism is used to dynamically evaluate the contribution of each unit. This grouping design not only reduces computational complexity but also enhances the network's ability to capture local features. Finally, adaptive pooling is used to further extract a joint spatiotemporal representation.

[0123] Finally, the extracted spatiotemporal features are input into the pose regression module, and the 5D pose parameters in the gripper coordinate system are output through a double fully connected layer.

[0124] In this embodiment, the pose regression module maps the extracted spatiotemporal features to 5D pose parameters in the gripper coordinate system through a dual fully connected layer. The input of this module is the spatiotemporal features after spatiotemporal feature extraction, and after processing by the fully connected layer, the planar position coordinates (x, y) and quaternion pose (dw, dx, dy, dz) of the target object are output. To prevent overfitting, Dropout and regularization techniques are introduced in the regression layer to improve the generalization ability of the model by randomly discarding some neurons and constraining the parameter range. The structural design of the dual fully connected layer enables the network to gradually abstract features, from low-level local patterns to high-level global representations, and ultimately achieve accurate prediction of 5D pose parameters. In addition, the accuracy of the output results can be further ensured by jointly optimizing the position loss function and the attitude loss function during network training. The position loss function combines the square error and the error term with adjustment parameters to distinguish small errors from large errors, while the attitude loss function calculates the rotation angle error through the quaternion dot product to ensure the stability of the attitude estimation.

[0125] This embodiment extracts the motion trend of an object through a dynamic feature encoding module, combines multi-scale dilated convolution with the CBAM attention mechanism to enhance the representation of temporal features, and utilizes a grouped convolution strategy combined with the CBAM mechanism and adaptive pooling to optimize the extraction of spatial features. Ultimately, high-precision 5D pose prediction is achieved through a dual fully connected layer. Based on this method, it can solve the problems of pose estimation lag and error accumulation in dynamic scenes caused by traditional methods. It also improves the computational efficiency and robustness of the model through a modular structure and efficient algorithm, thereby achieving high-precision positioning and rapid response capabilities in robot follow-up grasping tasks.

[0126] The following describes in detail the training method of the dynamic feature encoding convolutional neural network provided by an embodiment of the present invention. Figure 7 A flow chart of a training method for a dynamic feature coding convolutional neural network provided by an embodiment of the present invention. Figure 7 As shown, the method includes the following steps.

[0127] In step S701, the preprocessed proximity sensing signal of the preset object used for model training is synchronously aligned with the real posture data corresponding to the preset object in the gripper coordinate system to construct a training sample set containing temporal features and spatial features. The sampling frequency of the proximity sensing signal is 60 Hz. The training sample set is generated by sliding window division with a window length of 8 time steps. The real posture data is collected using an optical tracking system.

[0128] In step S702, the preprocessed proximity sensing signal is input into the dynamic feature encoding module, a first-order difference is performed on the preprocessed proximity sensing signal, dynamic features are extracted, and the dynamic features are processed through a one-dimensional convolution layer. The original features and the dynamic features after convolution processing are fused to generate fused features.

[0129] In step S703, convolution, CBAM attention mechanism, multi-scale dilated convolution and maximum pooling are used to extract temporal features from the fused features. Then, spatial features are divided through the grouped convolution strategy to learn local patterns. The spatiotemporal features are extracted by combining the CBAM attention mechanism, convolution and adaptive pooling.

[0130] Step S704: Based on the spatiotemporal features, the predicted pose parameters are output through a dual fully connected layer, and the position loss and attitude loss between the predicted pose parameters and the actual pose data are calculated.

[0131] In step S705, the Adam optimization algorithm is used to update the parameters of the dynamic feature encoding convolutional neural network. The initial learning rate of the Adam optimization algorithm is set to 0.01, and the learning rate is adjusted through a staged attenuation strategy. The number of training iterations is 160 rounds until the position loss and posture loss converge to the preset threshold.

[0132] In this embodiment, the actual pose data of the preset object in the gripper coordinate system must first be synchronously aligned with the corresponding preprocessed proximity sensing signal. This process ensures the temporal consistency of the input data and avoids feature-label mismatches caused by time deviations. The proximity sensing signal sampling frequency is 60Hz, and the training sample set is constructed using a sliding window partitioning method with a window length of 8 time steps (step length 1). This design not only ensures the continuity of the time series data, but also balances the computational complexity and the ability to capture dynamic features through the 8-time step window length.

[0133] In the feature extraction stage, multi-scale dilated convolution is introduced to enhance the model's perception of temporal features. Dilated convolution expands the receptive field by inserting holes (i.e., interval sampling) in the convolution kernel, enabling the model to capture longer temporal dependencies without increasing the number of parameters. For example, multi-scale dilated convolution may use convolution kernels with different expansion rates to extract local details and global trend features respectively, thereby adapting to the different motion patterns that may exist in the object during dynamic grasping. At the same time, the CBAM attention mechanism is integrated into the temporal feature extraction module, and the weights of key features are dynamically adjusted through the joint optimization of channel attention and spatial attention. This design not only improves the model's ability to capture dynamic features, but also effectively suppresses noise interference, such as ambient light interference or sensor noise that may exist in proximity perception signals.

[0134] In terms of spatial feature processing, the grouped convolution strategy is used to divide spatial features and learn local patterns. Grouped convolution divides the input feature channels into multiple groups, performs convolution operations on each group, and then concatenates the results. This reduces computational complexity and enhances the model's ability to express local features. For example, assuming the input feature has 64 channels, grouped convolution may divide it into 16 groups, each with 4 channels. After each group undergoes independent convolution operations, a richer spatial representation is generated through cross-group feature fusion. This strategy, combined with the CBAM attention mechanism, can dynamically assign weights to different units, thereby improving the recognition accuracy of local patterns.

[0135] Dual fully connected layers are used to map high-level features to 5-DOF pose parameters (planar position x, y and quaternion pose). For the joint regression task of 2D position and 3D pose, the total loss function of this embodiment is composed of a position loss function and a pose loss function to comprehensively measure the difference between the model-predicted pose parameters and the actual pose data.

[0136] Position loss function L position By combining the square error and the error term with adjustment parameters, the error between the predicted position and the true position is quantified in sections. Its specific form is shown in the following formula (7):

[0137]

[0138] Where n is the number of samples, and are the predicted position vector and the true position vector of the i-th training sample, respectively. ||·|| represents the Euclidean distance, and δ is a tuning parameter used to distinguish small errors from large errors. In this embodiment, δ is set to 5.

[0139] For the pose loss function L orientation , which quantifies the error by calculating the rotation angle difference between the predicted pose and the true pose. The specific formula is:

[0140]

[0141] in, and are the predicted pose and true pose (expressed in quaternion) of the i-th training sample, respectively. |·| represents the absolute value operation. 2arccos() maps the quaternion dot product result to the rotation angle range [0,2π], and calculates the average rotation angle of all samples as the pose loss value.

[0142] Finally, the total loss function L total The sum of position loss and attitude loss is obtained:

[0143] L total =L position +L orientation (9)

[0144] Finally, the parameters are updated using the Adam optimization algorithm, with an initial learning rate set to 0.01 and adjusted using a phased decay strategy. Training iterations are repeated for 160 epochs. The Adam optimizer combines the advantages of momentum and adaptive learning rates, enabling rapid convergence and avoiding local optima. The phased decay strategy might use a learning rate of 0.01 for the first 40 epochs, then reduce it to 0.001, and then adjust it to 0.0001 after 30 epochs. This strategy ensures rapid initial learning while allowing for fine-tuning of parameters later to prevent overfitting. Training continues until the loss function converges to a preset threshold. For example, in the experiment, the loss stabilized after 160 epochs, indicating that the model has fully learned the patterns in the data. Through the synergistic effect of these technical features, the dynamic feature encoding convolutional neural network can effectively process the dynamic characteristics of proximity perception signals and achieve high-precision pose estimation.

[0145] In an optional embodiment, the process of producing a training sample set involves multiple steps, including pose acquisition, system calibration, coordinate conversion, and final data set construction. In this embodiment, the optical tracking system for pose acquisition can include a base station, a tracker, a computer, and data processing software. The base station emits a global infrared laser through an LED array and is equipped with two orthogonally placed infrared laser transmitter rotors, which only serve as laser emission sources; the tracker has multiple built-in photosensors for receiving base station signals, and the data processing software is responsible for integrating and parsing the data obtained by the tracker. When the base station LED flashes, the base station records the timestamp, and the photosensor on the tracker calculates the angle of the sensor relative to the two rotors of the base station by measuring the arrival timestamp of the laser emitted by the two rotors. Since the position of the photosensor on the tracker has been pre-set, the overall pose information of the tracker can be derived through the multi-sensor pose data.

[0146] The tracker's pose is expressed in millimeters, and its attitude is described as a quaternion (in radians), with the origin of its coordinate system located at the center of the base. The quaternion is converted to a matrix and the translation information is incorporated to obtain the pose transformation matrix of the tracker coordinate system T relative to the base station coordinate system V. This matrix form is shown in the following formula (10):

[0147]

[0148] in, Represents the pose transformation matrix of the tracker coordinate system T relative to the base station coordinate system V.

[0149] Figure 8 A schematic diagram of a tracker and a preset object coordinate system provided by an embodiment of the present invention.

[0150] like Figure 8 As shown, establish the tracker coordinate system O T -X T Y T Z T With the grasped preset object coordinate system O t -X t Y t Z t , the matrix formula (11) for converting the preset object coordinate system to the tracker coordinate system is as follows:

[0151]

[0152] Where -l represents the translation offset between the tracker coordinate system and the preset object coordinate system.

[0153] Therefore, the position of the grasped preset object in the base station coordinate system is

[0154] Figure 9 A schematic diagram of various coordinate systems and their relationships within a robot grasping system provided by an embodiment of the present invention.

[0155] like Figure 9 As shown in the figure, the calibration process of the robot and the pose acquisition system needs to process multiple coordinate system relationships, including the tracker coordinate system T, the base station coordinate system V, the gripper coordinate system G, the robot base coordinate system B and the end effector coordinate system E. Establish the transformation matrix from the end effector coordinate system E to the gripper coordinate system G And the transformation matrix from the base station coordinate system V to the robot base coordinate system B The transformation matrix from the tracker coordinate system T to the end effector coordinate system E is as follows:

[0156]

[0157] in It is the matrix from the robot base coordinate system B to the end effector coordinate system E, which can be obtained through the robot's internal system; is the matrix from the tracker coordinate system T to the base station coordinate system V, provided by the pose acquisition system; i is the control point index.

[0158] During the calibration process, since the tracker is rigidly connected to the robot end, their relative pose parameters are constant, and the constraint equation (13) is established:

[0159]

[0160] The rotation and translation parts are separated and solved. The rotation part uses the Sylvester equation to construct an overdetermined equation group through vectorization and Kronecker product transformation, and the least squares method is used to solve the rotation matrix; the translation part is solved by sorting out the known equations to construct a linear equation group. Finally, the system calibration parameters are obtained Then derive the position of the grasped preset object in the robot base coordinate system And the pose in the gripper coordinate system

[0161] During the dataset construction phase, the z-axis coordinate of the preset object relative to the gripper coordinate system is fixed to zero (z=0). Each proximity sensor node outputs four channel signals. The data of the three sensor nodes are input into the neural network model and finally mapped to 5D pose parameters (plane position x, y and quaternion posture dw, dx, dy, dz) in the gripper coordinate system. This embodiment uses an equilateral hexagonal prism as the standard grasping target, such as Figure 10As shown in the figure, a dual constraint strategy is adopted during data acquisition: the initial grasping orientation of the preset object is fixed to eliminate the multi-solution mapping problem of "sensor data-absolute pose" caused by rotational symmetry, and at the same time, the preset object is allowed to rotate slightly within the range of ±20° around the z-axis to simulate the alignment error in actual grasping.

[0162] Using a data acquisition system (such as the AD7616), a 60Hz sampling frequency and a ±5V input voltage range were set for 10 minutes, and the raw signals were stored in a .txt format. During preprocessing, the voltage signal was normalized to the [0, 1] interval. Samples were partitioned using a sliding window of eight consecutive time steps (step size 1). Feature data and pose labels were synchronized using timestamps, ensuring that timestamp deviation was within ±2ms. This generated a dataset of 36,000 valid samples, divided into training, validation, and test sets in a 7:2:1 ratio.

[0163] The architecture of the dynamic feature encoding convolutional neural network provided by the embodiment of the present invention is described below with reference to the accompanying drawings.

[0164] Figure 11 A schematic diagram of the architecture of a dynamic feature encoding convolutional neural network provided by an embodiment of the present invention.

[0165] like Figure 11 As shown, the input layer of the dynamic feature encoding convolutional neural network has a data dimension of ([B, 3, 4, 8]), where B represents the batch size, corresponding to the three distributed proximity sensors (each sensor has four voltage signal outputs, for a total of eight time-step signals). After the input preprocessing stage, these signals are reorganized into a 12-channel spatiotemporal tensor with a dimension of ([B, 12, 8, 1]). This then enters the dynamic feature encoding module (DFE), which introduces a one-dimensional temporal convolution to calculate differential velocity features between adjacent time steps, fusing the original signal with the velocity information, resulting in an output dimension of ([B, 24, 8, 1]). This is then processed by the temporal feature extraction module, reducing the dimension to ([B, 64, 4, 1]). The spatial feature extraction module then extracts spatiotemporal features, with an output dimension of ([B, 256, 1, 1]). Finally, a fully connected layer integrates and processes these features to complete the task of dynamic object pose estimation within the robot hand. This entire architecture achieves effective pose estimation of dynamic objects through the hierarchical extraction and fusion of spatiotemporal features.

[0166] Figure 12 A schematic diagram of the processing flow of a dynamic feature encoding module provided in an embodiment of the present invention.

[0167] like Figure 12As shown in the figure, the input feature shape is ([B, K, T, 1]) (B is the batch size, K is the number of sensor features, and T is the time step). First, dimensionality reduction is performed by removing the last dimension, resulting in a data dimension of ([B, K, T]). Next, first-order differences are performed on the time series readings of each sensor feature to extract dynamic features reflecting the object's motion speed. A one-dimensional convolutional layer (Conv1d) is then used to process the dynamic information with a kernel size of 3. Padding is performed to maintain feature dimensionality. After processing, the original features are concatenated with the dynamic features along the channel dimension, resulting in a tensor of dimension ([B, 2K, T]). Finally, dimensionality increase is performed, resulting in an output feature shape of ([B, 2K, T, 1]). This process enhances the model's ability to capture object motion trends.

[0168] Figure 13 A schematic diagram of the architecture of a time series feature extraction module provided in an embodiment of the present invention.

[0169] like Figure 13 As shown in the figure, the DFE output first passes through the convolution layer, and then the CBAM attention mechanism is introduced to enhance the key motion features. Then, the multi-scale temporal receptive field is constructed through the void convolution, and finally the maximum pooling layer is used to achieve effective extraction of temporal information.

[0170] Figure 14 A schematic diagram of the architecture of a spatial feature extraction module provided by an embodiment of the present invention.

[0171] like Figure 14 As shown in the figure, the output of the temporal feature extraction module enters the grouped convolution layer, which divides the 64-dimensional features into 16 response units (4 channels per group). The local patterns of specific sensors or sensors are learned through independent convolution kernels, and the contribution of each unit is dynamically evaluated by the CBAM attention mechanism. Finally, the adaptive pooling layer extracts a robust spatial-temporal joint representation.

[0172] The regression head uses a dual fully connected layer with dropout and regularization to map high-level features into 5-DOF poses for accurate pose estimation.

[0173] In an optional embodiment, the 5D pose parameters are converted into the target pose in the robot base coordinate system through the coordinate conversion node of the ROS distributed architecture, and the joint angle control instructions are generated by inverse kinematics solution, including:

[0174] Based on the homogeneous transformation matrix between the gripper coordinate system and the robot base coordinate system, the 5D pose parameters in the gripper coordinate system are mapped to the target pose in the robot base coordinate system. The inverse kinematics solution is performed on the target pose to generate the angle configuration of each joint of the six-axis collaborative robot.

[0175] In this embodiment, we first rely on the homogeneous transformation matrix between the gripper coordinate system and the robot base coordinate system, which is established through the system calibration process. The 5D pose parameters (including plane position coordinates and quaternion posture) in the gripper coordinate system need to be converted into the target pose in the robot base coordinate system through this matrix. This process involves translation and rotation transformations in three-dimensional space. Its mathematical basis is the theory of homogeneous coordinate transformation, that is, the point coordinates are converted from one reference system to another through a 4×4 homogeneous transformation matrix. For example, if the object pose in the gripper coordinate system is the position and posture relative to the gripper, it needs to be converted into the absolute position and posture in the robot base coordinate system through matrix multiplication, thereby providing a global reference for subsequent motion control.

[0176] The target pose is then solved by inverse kinematics to generate the angular configurations for each joint of the six-axis collaborative robot. The essence of the inverse kinematics problem is to infer the kinematic angles of each joint based on the target pose of the end effector. In this solution, the algorithm needs to calculate the joint angles of a six-axis collaborative robot (such as the Innfos Gluon robot) so that its end effector reaches the target pose.

[0177] Through the ROS communication mechanism, the generated joint angle instructions are sent to each joint servo system through the underlying drive unit, coordinating the rotation of each axis according to the predetermined angle difference, and finally achieving a smooth transition of the end effector from the current posture to the target posture.

[0178] The adoption of ROS distributed architecture enables each node (such as coordinate conversion nodes and motion control nodes) to collaborate efficiently. This solution transmits data through topic communication.

[0179] The following describes in detail the ROS-based software architecture provided by the embodiments of the present invention with reference to the accompanying drawings. Figure 15 A schematic diagram of a ROS-based software architecture provided in an embodiment of the present invention.

[0180] The system of this embodiment is built on the robot operating system ROS in the Linux environment, which realizes the deep integration of heterogeneous hardware and software components through a distributed architecture. ROS adopts standardized communication interfaces and publish-subscribe mode, and ensures efficient collaboration of various functional units through lightweight message passing protocols. Specifically, Figure 15 As shown in the figure, the system is divided into four core nodes: proximity sensor node, prediction model calculation node, coordinate transformation node and robot motion control node, all of which work together around the ROS Master.

[0181] The proximity sensor node is responsible for collecting the voltage signal output by the photoelectric sensor array in real time and filtering and denoising the raw data. The processed signal is transmitted to the prediction model computing node via a ROS topic at a frequency of 60Hz.

[0182] The prediction model computation node receives the voltage signal stream from the proximity sensor node and first performs data preprocessing to adapt the input specifications of the Dynamic Feature Encoding Convolutional Neural Network (DFENet). The model regresses the 5D pose parameters of the target object in the gripper coordinate system, and the prediction results are transmitted to the coordinate conversion node via a ROS topic.

[0183] The coordinate transformation node is a key middleware module in the system. It utilizes the ROS built-in tf (Transform Library) toolkit to accurately transform between multiple coordinate systems. This node maintains a complete transformation chain between the robot base coordinate system (B), the end-effector coordinate system (E), and the gripper coordinate system (G). It converts the pose parameters in the gripper coordinate system into the global pose in the robot base coordinate system. This process takes into account the robot arm's structural parameters and calibration error compensation to ensure the positioning accuracy of the end-effector's motion trajectory.

[0184] After receiving the desired pose output from the coordinate transformation node, the robot motion control node first analyzes the Cartesian displacement vector and quaternion attitude matrix to perform inverse kinematics and generate the joint angle configuration for the six-axis collaborative robot. The underlying drive unit then synchronously sends commands to the servo systems of each joint, coordinating the rotation of each axis by a predetermined angle difference, ultimately achieving a smooth transition of the end effector from the current pose to the target pose. A closed-loop feedback mechanism continuously monitors actual pose deviations and dynamically adjusts joint angle corrections to ensure robustness and stability in the grasping operation.

[0185] The entire ROS architecture ensures data consistency between nodes through standardized communication protocols (such as TCPROS / Rosbridge), while leveraging ROS visualization tools (such as rviz) for real-time monitoring of system status. This design not only simplifies the development and debugging of complex systems but also provides flexible interface support for future expansion, such as the introduction of force control or vision modules.

[0186] In an optional embodiment, joint angle control instructions are dynamically adjusted based on real-time posture deviations to drive the robot to perform a follow-up grasping operation, including: obtaining the actual posture of the end effector through closed-loop feedback, and calculating the target posture using proximity sensing signals obtained by the photoelectric sensor array. Based on the deviation between the target posture and the actual posture, joint angle corrections are calculated. The joint angle corrections are synchronously sent to the servo mechanisms of each joint of the robot, driving the robot end effector to perform the grasping operation according to the corrected joint angles.

[0187] In this embodiment, the real-time posture deviation detection is used to accurately correct the robot joint angles, thereby driving the end effector to complete a high-precision grasping operation.

[0188] First, the target pose information of the end effector is obtained. The calculation of this target pose relies on the proximity sensing signal collected by the photoelectric sensor array. The photoelectric sensor array is deployed on the inside of the gripper and can monitor the relative position and posture changes between the target object and the gripper in real time. The proximity sensing signal it outputs is filtered and denoised, and then combined with an algorithm model (such as a dynamic feature encoding convolutional neural network) for feature extraction and pose estimation, ultimately generating the target pose data in the gripper coordinate system. The key to this process lies in the high response speed (less than 1ms) and high positioning accuracy (0.6-40mm working range) of the photoelectric sensor array, which enables the system to quickly capture the pose changes of dynamic targets and provide reliable data support for subsequent control.

[0189] After obtaining the actual pose, the system compares the target pose with the actual pose and calculates the deviation between the two. This deviation calculation includes not only the coordinate difference of the planar position (x, y direction) but also the difference in quaternion posture (dw, dx, dy, dz), thus fully reflecting the pose error between the end effector and the target.

[0190] The generated correction amount is further transmitted to the servo mechanism of each joint of the robot. The servo mechanism adjusts the joint angle according to the received instructions and drives the end effector to perform the grasping action according to the corrected posture. This process must ensure that the servo mechanism of each joint receives and executes instructions within the same time step to avoid grasping failure or collision risks due to delays or asynchrony. The entire control process achieves efficient communication through the ROS distributed architecture. Each node (such as sensor data acquisition node, prediction model node, coordinate conversion node, motion control node) exchanges data through a topic (Topic) or service (Service) mechanism to ensure the performance of the system. Ultimately, this technical solution can effectively solve the response delay and positioning deviation problems of traditional grasping systems in dynamic target tracking through closed-loop feedback, and realize high-precision and high-robust follow-up grasping of target objects by the robot end effector, which is especially suitable for complex grasping tasks in unstructured environments.

[0191] The following, combined with the accompanying drawings and experimental data, describes in detail the specific implementation method of the technical solution of this patent to realize the feasibility of the robot follow-up grasping method based on proximity perception and verify its technical effect.

[0192] In the tilt detection characteristic experiment, a rectangular whiteboard is used as the object to be measured, and the position output data of the sensor at different pitch angles and roll angles are recorded. The experimental results are as follows: Figure 16 and Figure 17 When the pitch angle and roll angle are both at zero, the position output value is close to zero regardless of how the distance between the target object and the sensor changes, indicating that the sensor can adjust the fingertip direction to make the output x c and yc The value is reset to zero.

[0193] In the distance measurement performance experiment, the object to be measured is placed parallel to the front of the sensor and moved step by step. The total current change within the range of 0.6mm to 60mm is measured. The results are as follows: Figure 18 As shown in the figure, the sensor exhibits stable distance measurement capability in the range of 0.6mm to 40mm, which can accurately reflect the relative distance between the object and the sensor.

[0194] In view of the transient response characteristics, the sensor is fixed at a distance of 2mm from the target object, and the voltage change process is measured by controlling the on / off state of the LED. The results are as follows: Figure 19 and Figure 20 As shown in the figure: when the LED is turned on, the time required for the voltage to rise to 90% of the steady-state value is 44μs; when the LED is turned off, the time required for the voltage to drop to 10% of the steady-state value is 48μs. The overall response time is less than 1ms, verifying the sensor's fast response capability in dynamic environments.

[0195] During the model training and verification phase, the experimental environment used a server equipped with an Intel Xeon E5-2630 v4 processor and an NVIDIA TITAN X graphics card (8GB of video memory), running a 64-bit Ubuntu 18.04 operating system, programming language Python 3.6, and building a network model based on the PyTorch 1.10.0 framework.

[0196] During the training process, the batch size is set to 32, and the Adam optimizer is used to update the weights. The initial learning rate is 0.01, which is adjusted to 0.001 after 40 rounds of training, and further reduced to 0.0001 after 30 rounds. The total number of training rounds is 160 rounds. During the training process, the loss value of the pose prediction shows a rapid convergence trend on both the training set and the validation set, as shown in Figure 21 This shows that the model has good regression ability.

[0197] To evaluate the contribution of the dynamic feature encoding module (DFE) and CBAM attention mechanism to model performance, four sets of ablation experiments were designed:

[0198] (1) DFENet (complete model, integrating original signal and velocity differential features);

[0199] (2) DFENet-NoDFE-NoCBAM (DFE and CBAM are removed as the baseline model);

[0200] (3) DFENet-NoDFE (only DFE is removed and the original signal is input);

[0201] (4) DFENet-NoCBAM (keep DFE but replace CBAM with ordinary convolutional layers).

[0202] The model performance is quantitatively evaluated by the two indicators of average position error and average attitude error. The calculation formulas for the average position error and the average attitude error are shown in the following formulas (14) and (15), respectively:

[0203]

[0204] The results are shown in Table 1. The average position error of the baseline model DFENet-NoDFE-NoCBAM is 2.79mm and the attitude error is 3.23°, while the errors of the complete model DFENet are reduced to 1.91mm and 2.07° respectively, verifying the key role of the DFE and CBAM modules in improving the accuracy of pose estimation.

[0205] Table 1 Ablation experiment results

[0206] Model Name Average position error / (mm) Average attitude error / (°) DFENet-NoDFE-NoCBAM 2.79 3.23 DFENet-NoCBAM 2.39 2.63 DFENet-NoDFE 2.45 2.75 DFENet 1.91 2.07

[0207] For the actual grasping verification, the experimental system was constructed using an Innfos Gluon six-axis collaborative robot and a Tunan Intelligent Control three-finger flexible gripper. The Innfos Gluon robot boasts a lightweight design, weighing 6.2kg, a 1.2kg end-load, a maximum range of motion of 456mm, and a repeatability of 5mm. The three-finger flexible gripper is controlled via the CAN bus and is compatible with workpieces up to 50mm in diameter. A proximity sensor integrated within the gripper senses the target object's position in real time, enabling closed-loop control in conjunction with the ROS communication framework. The experiment involved two scenarios: unidirectional translation-based grasping and compound motion-based grasping.

[0208] like Figure 22 and Figure 23 In the former, the tester continuously translates the target object in a single direction, while in the latter, the tester simultaneously adjusts the object's position and attitude. The results show that, after filtering, noise reduction, and feature extraction, the model network can simultaneously resolve the object's planar position and quaternion attitude. The collaborative work between ROS nodes ensures stable follow-up grasping of dynamic objects, validating the applicability of this method in complex, unstructured environments.

[0209] Based on the above method, an embodiment of the present invention further provides a robot follow-up grasping device based on proximity perception, comprising:

[0210] The acquisition module is used to obtain the proximity sensing signal between the target object and the gripper in real time through the photoelectric sensor array deployed on the inner side of the robot's end gripper. The proximity sensing signal is used to characterize the two-dimensional position information and distance information of the target object.

[0211] The processing module is used to filter and reduce noise on the proximity sensing signal to obtain a preprocessed proximity sensing signal.

[0212] The processing module is also used to process the pre-processed proximity perception signal through the trained dynamic feature encoding convolutional neural network, and output the 5D pose parameters of the target object in the gripper coordinate system. The 5D pose parameters include plane position coordinates and quaternion posture.

[0213] The conversion module is used to convert the 5D pose parameters into the target pose in the robot base coordinate system through the coordinate conversion node of the ROS distributed architecture, and generate joint angle control instructions through inverse kinematics solution.

[0214] The adjustment module is used to dynamically adjust the joint angle control instructions based on the real-time posture deviation to drive the robot to complete the follow-up grasping operation. The real-time posture deviation includes the deviation between the target posture and the actual posture of the end effector.

[0215] An embodiment of the present invention further provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, wherein the computer program, when executed by the at least one processor, causes the electronic device to perform the method of an embodiment of the present invention.

[0216] An embodiment of the present invention further provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform the method of the embodiment of the present invention.

[0217] An embodiment of the present invention further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to enable the computer to perform the method of the embodiment of the present invention.

[0218] The computer programs for implementing the methods of the embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer programs are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0219] In the context of an embodiment of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0220] It should be noted that the term "including" and its variations used in the embodiments of the present invention are open-ended, i.e., "including but not limited to." The term "based on" means "based at least in part on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; and the term "some embodiments" means "at least some embodiments." The modifications of "one" and "a plurality of" mentioned in the embodiments of the present invention are illustrative and non-restrictive. Those skilled in the art should understand that, unless the context clearly indicates otherwise, they should be understood as "one or more."

[0221] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0222] The various steps described in the method implementation scheme provided in the embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method implementation scheme may include additional steps and / or omit the steps shown. The scope of protection of the present invention is not limited in this respect.

[0223] The term "embodiment" in this specification refers to specific features, structures, or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. The various embodiments in this specification are described in a related manner, and the same or similar parts between the various embodiments are referenced to each other. In particular, for the embodiments of the device, equipment, and system, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts are referred to the partial description of the method embodiment.

[0224] The above-described embodiments merely illustrate several implementations of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of the present invention. Therefore, the scope of the present invention shall be determined by the appended claims.

Claims

1. A robot follow-up grasping method based on proximity perception, characterized in that: include: The photoelectric sensor array deployed inside the end gripper of the robot acquires the proximity sensing signal between the target object and the gripper in real time. The proximity sensing signal is used to represent the two-dimensional position information and distance information of the target object. Performing filtering and noise reduction processing on the proximity sensing signal to obtain a preprocessed proximity sensing signal; Processing the pre-processed proximity sensing signal through a trained dynamic feature encoding convolutional neural network to output 5D pose parameters of the target object in a gripper coordinate system, wherein the 5D pose parameters include plane position coordinates and quaternion posture; The 5D pose parameters are converted into the target pose in the robot base coordinate system through the coordinate conversion node of the ROS distributed architecture, and the joint angle control instructions are generated through inverse kinematics solution; The joint angle control instruction is dynamically adjusted based on the real-time posture deviation to drive the robot to complete the follow-up grasping operation. The real-time posture deviation includes the deviation between the target posture and the actual posture of the end effector.

2. The method according to claim 1, characterized in that The photoelectric sensor array is constructed by three rows and two columns of micro-photoelectric sensors through a resistor network. Each of the micro-photoelectric sensors detects the intensity of reflected light from the target object based on the infrared reflection principle and converts it into a photocurrent. The photoelectric sensor array outputs the proximity sensing signal through the resistor network. The working range of the photoelectric sensor array is 0.6 to 40 mm, the response time is less than 1 ms, and the two-dimensional position information and the distance information are calculated by the center position and total current of the photocurrent distribution.

3. The method according to claim 1, characterized in that Sampling, filtering, denoising and normalizing the proximity sensing signal to obtain a preprocessed proximity sensing signal includes: The proximity sensing signal is sampled and filtered for noise reduction at a sampling frequency of 60 Hz to eliminate environmental interference and sensor noise, and then normalized; The normalized proximity sensing signal is divided into a timing window according to 8 consecutive time steps, and the 8-step timing signals of the three photoelectric sensor arrays distributedly deployed on the inner side of the gripper are recombined into a 12-channel spatiotemporal tensor to generate the preprocessed proximity sensing signal.

4. The method according to claim 1, wherein The dynamic feature coding convolutional neural network includes a dynamic feature coding module, a spatiotemporal feature extraction module and a posture regression module; The pre-processed proximity sensing signal is processed by the trained dynamic feature encoding convolutional neural network to output the 5D pose parameters of the target object in the gripper coordinate system, including: Inputting the preprocessed proximity sensing signal into the dynamic feature encoding module, performing first-order difference on the preprocessed proximity sensing signal to extract dynamic features, processing the dynamic features through a one-dimensional convolution layer, and fusing the original features with the dynamic features after convolution to generate a fused feature; The fused features are input into the spatiotemporal feature extraction module, and temporal features are extracted through convolution, CBAM attention mechanism, multi-scale dilated convolution and maximum pooling. Then, spatial features are divided through grouped convolution strategy, local patterns are learned, and spatiotemporal features are extracted by combining CBAM attention mechanism, convolution and adaptive pooling; The spatiotemporal features are input into the pose regression module, and the 5D pose parameters in the gripper coordinate system are output through a double fully connected layer.

5. The method according to claim 4, wherein the training method of the dynamic feature encoding convolutional neural network comprises: The pre-processed proximity sensing signal of the preset object used for model training is synchronously aligned with the real pose data corresponding to the preset object in the gripper coordinate system to construct a training sample set containing temporal and spatial features. The proximity sensing signal sampling frequency is 60 Hz, and the training sample set is generated by sliding window partitioning with a window length of 8 time steps. The real pose data is collected using an optical tracking system; Inputting the preprocessed proximity sensing signal into a dynamic feature encoding module, performing first-order difference on the preprocessed proximity sensing signal to extract dynamic features, processing the dynamic features through a one-dimensional convolution layer, and fusing the original features with the dynamic features after convolution to generate a fused feature; Convolution, CBAM attention mechanism, multi-scale dilated convolution and maximum pooling are used to extract temporal features from the fused features. Then, spatial features are divided through a grouped convolution strategy to learn local patterns. The spatiotemporal features are extracted by combining CBAM attention mechanism, convolution and adaptive pooling. Based on the spatiotemporal features, outputting predicted pose parameters through a dual fully connected layer, and calculating the position loss and attitude loss between the predicted pose parameters and the true pose data; The Adam optimization algorithm is used to update the parameters of the dynamic feature encoding convolutional neural network. The initial learning rate of the Adam optimization algorithm is set to 0.01, and the learning rate is adjusted by a staged attenuation strategy. The number of training iterations is 160 rounds until the position loss and attitude loss converge to the preset threshold.

6. The method according to claim 1, characterized in that The coordinate conversion node of the ROS distributed architecture converts the 5D pose parameters into the target pose in the robot base coordinate system, and generates joint angle control instructions through inverse kinematics solution, including: Based on the homogeneous transformation matrix between the gripper coordinate system and the robot base coordinate system, the 5D pose parameters in the gripper coordinate system are mapped to the target pose in the robot base coordinate system; The target posture is solved by inverse kinematics to generate the angle configuration of each joint of the six-axis collaborative robot.

7. The method according to claim 1, characterized in that The method of dynamically adjusting the joint angle control instruction based on the real-time posture deviation to drive the robot to complete the follow-up grasping operation includes: The actual posture of the end effector is obtained through closed-loop feedback, and the target posture is calculated by the proximity sensing signal obtained by the photoelectric sensor array; Calculating a joint angle correction based on a deviation between the target posture and the actual posture; The joint angle correction amount is synchronously sent to the servo mechanism of each joint of the robot, and the robot end effector is driven to perform a grasping operation according to the corrected joint angle.

8. A robot follow-up grasping device based on proximity perception, characterized in that: include: An acquisition module is used to acquire a proximity sensing signal between a target object and the gripper in real time through a photoelectric sensor array deployed inside the gripper at the end of the robot. The proximity sensing signal is used to represent the two-dimensional position information and distance information of the target object; a processing module, configured to perform sampling, filtering, noise reduction, and normalization processing on the proximity sensing signal to obtain a preprocessed proximity sensing signal; The processing module is further configured to process the pre-processed proximity sensing signal through a trained dynamic feature encoding convolutional neural network, and output 5D pose parameters of the target object in a gripper coordinate system, wherein the 5D pose parameters include plane position coordinates and quaternion pose; A conversion module is used to convert the 5D pose parameters into the target pose in the robot base coordinate system through the coordinate conversion node of the ROS distributed architecture, and generate joint angle control instructions through inverse kinematics solution; An adjustment module is used to dynamically adjust the joint angle control instruction based on the real-time posture deviation to drive the robot to complete the follow-up grasping operation. The real-time posture deviation includes the deviation between the target posture and the actual posture of the end effector.

9. An electronic device comprising: A processor, and a memory storing a program, wherein the program comprises instructions which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory machine-readable medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent robot with body, material transfer method, storage medium and program product

    CN121374644A