Robot hand-eye positioning method and device based on 3D camera and medium
Through real-time scanning and dynamic adjustment of prediction models of 3D cameras, the problem of failed dynamic target capture in traditional methods is solved, and efficient and accurate capture of robots in complex environments is achieved.
Patent Information
- Application Number
- CN202510872110.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Traditional static mapping methods cannot adapt to object position and posture changes in real time in dynamic target scenarios where high-speed motion or frequent pose changes, resulting in robot grabbing failure or delay, affecting efficiency and accuracy.
The target object is scanned in real time through a 3D camera, and the prediction model is used to predict the motion state and match it with the point cloud data. The model is adjusted to generate a dynamic crawling path, and combined with the historical point cloud data to process the occlusion situation, realizing hand-eye matrix mapping and path planning.
It improves the accuracy and efficiency of the robot's grasping of high-speed moving targets, enhances the robustness and stability of the system, and ensures accurate operation in complex environments.
Smart Images

Figure CN120422248A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robotics technology, and in particular to a 3D camera-based robot hand-eye positioning method, device, and medium. Background Art
[0002] In robot grasping or interaction scenarios, when the target object is in motion, even if the hand-eye matrix calibration has been completed, to obtain the real-time position and posture of the object, it is still necessary to use the data collected by the camera and accurately map it to the robot coordinate system. However, the traditional static mapping method has obvious drawbacks. When the object moves at high speed, there is a time difference between the data collected by the camera and the execution of the robot action. This time synchronization error will continue to accumulate. Because the traditional method cannot adapt to the rapid changes in the position and posture of the object in real time and dynamically, when faced with dynamic targets that move at high speed or change posture frequently, it is very easy to fail to grasp or generate significant delays in the grasping process when guiding the robot operation based on the static mapping results, which seriously affects the efficiency and accuracy of the robot operation and cannot meet the needs of efficient and precise operation of dynamic targets. Summary of the Invention
[0003] In order to solve the above problems, the present application proposes a robot hand-eye positioning method based on a 3D camera, comprising: scanning a scene in which a target object is located in real time through a 3D camera to determine point cloud data of the target object, wherein the point cloud data includes three-dimensional coordinate information of each point on the surface of the target object; predicting the motion state of the target object according to a predetermined prediction model to obtain a predicted trajectory, matching the predicted trajectory with the real-time scanned point cloud data to determine a prediction difference, wherein the prediction difference includes a distance difference and a shape difference, and adjusting the prediction model according to the prediction difference; determining whether the target object is obscured, and if the target object is obscured, determining historical point cloud data of the target object, and adjusting the prediction model according to the historical point cloud data; re-predicting the target object through the adjusted prediction model to obtain a prediction result, determining a preset hand-eye matrix, mapping the prediction result with the hand-eye matrix to dynamically generate a path for the target object, and controlling the robot's end effector to reach a corresponding predicted position according to the path to achieve grasping at the predicted position.
[0004] In one example, the motion state of the target object is predicted according to a predetermined prediction model to obtain a predicted trajectory, specifically including: determining an initial state according to the point cloud data of the target object at the current moment, the initial state including position and speed; using the historical point cloud data and related motion features of the target object as input to train an LSTM network, thereby outputting the position information of the target object at the future moment, and then obtaining a predicted trajectory.
[0005] In one example, the method further includes: installing the 3D camera on the end effector of the robot to construct a hand-eye system; placing a calibration object in the workspace of the robot, and controlling the robot to drive the 3D camera to shoot the calibration object in multiple postures to obtain multiple sets of point cloud data of the calibration object and corresponding end effector posture data.
[0006] In one example, the method further includes: determining a hand-eye transformation matrix according to a preset hand-eye calibration equation, calibrating the robot hand-eye according to the hand-eye transformation matrix, thereby obtaining the 3D camera coordinate system, and determining the conversion relationship between the 3D camera coordinate system and the end effector coordinate system.
[0007] In one example, the method further includes: collecting point cloud data of the target object through the 3D camera, and preprocessing the collected point cloud data, wherein the preprocessing includes denoising, filtering, and downsampling; identifying the preprocessed point cloud data to determine the category of the target object; and aligning the point cloud data with a pre-established target object pose model to obtain the pose information of the target object in the coordinate system of the 3D camera.
[0008] In one example, the method further includes: calibrating the 3D camera according to a pre-set calibration plate to obtain a hand-eye matrix of the 3D camera, the hand-eye matrix including an intrinsic parameter matrix and an extrinsic parameter matrix, the intrinsic parameter matrix including the focal length and principal point coordinates of the 3D camera, and the extrinsic parameter matrix including the conversion relationship between the 3D camera coordinate system and the end effector coordinate system; and optimizing the intrinsic parameter matrix and the extrinsic parameter matrix using the least squares method to improve the calibration accuracy of the 3D camera.
[0009] In one example, the prediction model is adjusted according to the historical point cloud data, specifically including: analyzing the geometric information of the historical point cloud data, reconstructing the prediction model according to the geometric information and the motion law of the target object, and determining the minimum reconstruction ratio of the prediction model according to the category and grasping angle of the target object.
[0010] In one example, dynamically generating the path of the target object specifically includes: generating a dynamic grasping path according to a dynamic mapping relationship, and determining the predicted motion trajectory of the target object according to the dynamic grasping path, so that the end effector of the robot reaches the corresponding predicted position in advance to achieve accurate grasping operation.
[0011] On the other hand, the present application also proposes a robot hand-eye positioning device based on a 3D camera, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the robot hand-eye positioning device based on a 3D camera can perform the following: real-time scanning of the scene in which the target object is located by a 3D camera to determine point cloud data of the target object, wherein the point cloud data includes three-dimensional coordinate information of each point on the surface of the target object; predicting the motion state of the target object according to a predetermined prediction model to obtain a predicted trajectory, and comparing the predicted trajectory with the target object. The point cloud data scanned in real time are matched to determine the prediction difference, which includes distance difference and shape difference, and the prediction model is adjusted according to the prediction difference; whether the target object is blocked is determined, and if the target object is blocked, the historical point cloud data of the target object is determined, and the prediction model is adjusted according to the historical point cloud data; the target object is re-predicted by the adjusted prediction model to obtain a prediction result, a preset hand-eye matrix is determined, and the prediction result is mapped with the hand-eye matrix to dynamically generate the path of the target object, and the end effector of the robot is controlled to reach the corresponding predicted position according to the path to achieve grasping at the predicted position.
[0012] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to: scan a scene in which a target object is located in real time using a 3D camera to determine point cloud data of the target object, wherein the point cloud data includes three-dimensional coordinate information of each point on the surface of the target object; predict the motion state of the target object according to a predetermined prediction model to obtain a predicted trajectory, match the predicted trajectory with the real-time scanned point cloud data to determine a prediction difference, wherein the prediction difference includes a distance difference and a shape difference, and adjust the prediction model according to the prediction difference; determine whether the target object is obscured, and if the target object is obscured, determine the historical point cloud data of the target object, and adjust the prediction model according to the historical point cloud data; re-predict the target object using the adjusted prediction model to obtain a prediction result, determine a preset hand-eye matrix, map the prediction result with the hand-eye matrix to dynamically generate a path for the target object, and control the end effector of the robot to reach the corresponding predicted position according to the path to achieve grasping at the predicted position.
[0013] This application utilizes a prediction model to predict the target object's motion state. By matching it with real-time point cloud data to identify discrepancies and adjusting the model, the system can promptly adapt to changes in the object's motion, improving prediction accuracy. This is particularly advantageous in complex motion scenarios. When the target object is occluded, the prediction model is adjusted using historical point cloud data to avoid positioning failures due to occlusion, enhancing system robustness and ensuring stable operation in complex environments. A hand-eye system is constructed. Using multiple sets of calibration object point cloud data and end-effector pose data, combined with hand-eye calibration equations, the hand-eye transformation matrix is determined. This defines the transformation relationship between the 3D camera coordinate system and the end-effector coordinate system, laying the foundation for subsequent precise control. The target object's category is identified and its pose information is registered, providing an accurate basis for subsequent path planning and grasping operations. A 3D camera hand-eye matrix is obtained using a calibration plate. The least squares method is used to optimize the intrinsic and extrinsic parameter matrices, improving camera calibration accuracy and positioning accuracy. A dynamic grasping path is generated based on the dynamic mapping relationship. Combined with the predicted motion trajectory, the robot's end-effector reaches the predicted position in advance, achieving accurate grasping and improving grasping efficiency and success rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0015] Figure 1 This is a flow chart of a robot hand-eye positioning method based on a 3D camera in an embodiment of the present application;
[0016] Figure 2 This is a schematic diagram of a robot hand-eye positioning device based on a 3D camera in an embodiment of the present application. DETAILED DESCRIPTION
[0017] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0018] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0019] like Figure 1 As shown, in order to solve the above problems, an embodiment of the present application provides a robot hand-eye positioning method based on a 3D camera, the method comprising:
[0020] S101 . Scanning a scene where a target object is located in real time using a 3D camera to determine point cloud data of the target object, where the point cloud data includes three-dimensional coordinate information of each point on the surface of the target object.
[0021] In robotic operations, advanced 3D cameras are used to perform real-time scanning of the complex environment surrounding the target object. 3D cameras can capture scene information from all directions and angles. During the scanning process, they utilize specific optical principles and techniques to acquire rich geometric data from the target object's surface, point by point. Through precise calculations and analysis, the target object's point cloud data is ultimately determined. This point cloud data contains the three-dimensional coordinate information of every point on the target object's surface, accurate to the micron level. Each coordinate point is interconnected and mutually supportive, collectively outlining the target object's complete three-dimensional form. With this precise point cloud data, the robot can clearly "see" the target object's appearance and position, providing solid and reliable data support for subsequent precise operations such as grasping and handling.
[0022] S102. Predicting the motion state of the target object according to a predetermined prediction model to obtain a predicted trajectory, matching the predicted trajectory with the point cloud data scanned in real time to determine a prediction difference, wherein the prediction difference includes a distance difference and a shape difference, and adjusting the prediction model according to the prediction difference.
[0023] In the robot's intelligent perception and decision-making system, a prediction model, pre-built and optimized through extensive historical data training, accurately predicts the target object's motion state. This prediction model integrates various kinematic principles, dynamic characteristics, and potential influencing factors of the target object's environment. By inputting various state parameters of the target object at the current moment, such as velocity vector, acceleration components, and force conditions, it uses recursive neural network architectures within deep learning algorithms, such as LSTM, to iteratively calculate the target object's motion trend over a period of time. This generates a continuous and smooth predicted trajectory, consisting of a series of three-dimensional spatial coordinate points arranged in chronological order, that accurately depicts the target object's expected motion path.
[0024] At the same time, a 3D camera is used to perform high-frequency, high-precision real-time scanning of the scene where the target object is located to obtain point cloud data of the target object at each moment. These point cloud data contain the three-dimensional coordinate information of a large number of discrete points on the surface of the target object, which can truly reflect the actual shape and position of the target object in real space. Subsequently, the predicted trajectory is strictly matched with the point cloud data obtained by real-time scanning. The matching process uses advanced point cloud registration algorithms, such as the iterative closest point (ICP) algorithm, to accurately determine the predicted difference by calculating the distance error and shape similarity index between each point on the predicted trajectory and the corresponding area point in the real-time point cloud data. The predicted difference specifically includes distance difference, that is, the Euclidean distance deviation between the predicted trajectory point and the actual point cloud data point; and shape difference, which is measured by calculating indicators such as the Hausdorff distance between the geometric shape enclosed by the predicted trajectory and the object shape presented by the actual point cloud data.
[0025] Based on the predicted discrepancies, the prediction model is dynamically optimized using an adaptive adjustment algorithm. Adjustment strategies include, but are not limited to, modifying the model's motion parameter weights, updating the neural network's connection weights, and optimizing the model's input feature combinations. This improves the accuracy and robustness of the prediction model's predictions of the target object's motion state, ensuring the robot can accurately track and manipulate the target object in complex and changing environments.
[0026] S103 , determining whether the target object is blocked; if the target object is blocked, determining historical point cloud data of the target object, and adjusting the prediction model according to the historical point cloud data.
[0027] In dynamic scenes of robot operations, in order to ensure accurate positioning and operation of target objects, it is necessary to determine in real time whether the target object is in an occluded state. By comprehensively analyzing the point cloud data features collected by the 3D camera in real time and the depth information distribution of the scene, an occlusion detection algorithm based on deep learning and geometric feature fusion is used to determine whether the target object is occluded. The algorithm first pre-processes the point cloud data at the current moment, extracts the local geometric features of the target object area, such as surface curvature, normal direction, etc., and combines the gradient change information of the depth image to construct a three-dimensional feature descriptor of the target object. The descriptor is compared with a pre-established standard feature library of target objects in an unoccluded state. If the matching degree between the current feature descriptor and the standard feature library is lower than the preset threshold, and there is an obvious depth mutation or loss in the target object area in the depth image, the target object is determined to be occluded.
[0028] Once the target object is determined to be occluded, the system immediately retrieves the target object's historical point cloud data from the storage module. This historical point cloud data contains the target object's 3D geometric information at different times and in different postures, recording its motion patterns and morphological changes. Based on this historical point cloud data, the prediction model is adjusted using a combination of time series analysis and machine learning. Specifically, time series modeling is performed on the historical point cloud data to extract dynamic features such as the target object's velocity and acceleration. A clustering algorithm is then used to classify the target object's different motion patterns. Based on the classification results, the dynamic parameters in the prediction model are retrained to optimize the model's ability to predict the target object's motion trends. Furthermore, the prediction model's structure is adaptively adjusted based on the target object's geometric characteristics and motion patterns. For example, the number of hidden layer nodes in the neural network is increased or decreased, and hyperparameters such as the learning rate are adjusted to improve the model's robustness and accuracy in the presence of occlusion, ensuring that the robot can continue to accurately track and operate the target object based on the adjusted prediction model.
[0029] S104. Re-predict the target object using the adjusted prediction model to obtain a prediction result, determine a pre-set hand-eye matrix, map the prediction result with the hand-eye matrix to dynamically generate a path for the target object, and control the robot's end effector to reach a corresponding predicted position according to the path to achieve grasping at the predicted position.
[0030] After completing the targeted adjustments to the prediction model, the adjusted prediction model was put into operation to accurately grasp the future motion of the target object. This model is built based on the deep integration of a deep learning framework and kinematic principles. By inputting the point cloud features, motion state parameters, and environmental constraint information of the target object at the current moment, a multi-layer neural network is used to extract high-dimensional features and perform nonlinear mapping. Within the model, a long short-term memory network (LSTM) unit is used to capture the temporal dependencies of the target object's motion. Combined with a fully connected layer, it iteratively predicts the position, posture, and other state variables of the target object at future moments. The final output is a prediction result containing detailed information such as the target object's three-dimensional spatial coordinates and rotation angles at multiple time steps in the future.
[0031] At the same time, the pre-set hand-eye matrix is read from the system storage module. Determination of the hand-eye matrix involves a complex calibration process. This involves controlling the robot to drive the 3D camera to capture the calibration object in multiple poses, collecting a large amount of point cloud data from the calibration object and the corresponding robot end-effector pose data. This matrix is then solved using a hand-eye calibration algorithm, such as the Tsai two-step method or a method based on nonlinear optimization. This hand-eye matrix precisely describes the spatial transformation relationship between the 3D camera coordinate system and the robot end-effector coordinate system and is a key parameter for achieving precise robot operation.
[0032] The prediction results are mapped to the hand-eye matrix, and the predicted pose information of the target object in the 3D camera coordinate system is converted to the robot end-effector coordinate system through matrix multiplication. Based on this converted pose information, a path planning algorithm, such as the rapidly expanding random tree algorithm (RRT*) or a model predictive control-based path planning method, is used to dynamically generate a grasping path for the target object. This path planning process fully considers the robot's kinematic constraints, dynamic characteristics, and obstacle information in the workspace to ensure that the generated path is smooth, feasible, and efficient.
[0033] Based on the generated path, the robot control system sends control commands to each joint of the robot. Using real-time feedback control algorithms, such as PID control or model-based control algorithms, the robot's joint angles are precisely adjusted, driving the robot's end effector along the planned path, ultimately reaching the predicted position. At the predicted position, the robot's end effector adjusts its gripping posture and force according to the pre-set grasping strategy to achieve a stable grasp of the target object, ensuring the accuracy and reliability of the entire grasping process.
[0034] In one embodiment, a 3D camera is installed on the robot end effector to construct a hand-eye system; a calibration object is placed in the robot workspace, and the robot is controlled to drive the 3D camera to shoot the calibration object in different postures, obtaining multiple sets of calibration object point cloud data in the 3D camera coordinate system and the corresponding robot end effector posture data.
[0035] Based on the hand-eye calibration equation AX=XB, where A is the transformation matrix of the robot end effector between different poses, B is the transformation matrix of the 3D camera observing the calibration object at different poses, and X is the hand-eye transformation matrix to be solved, the hand-eye transformation matrix X is solved using singular value decomposition (SVD) or other numerical methods to implement robot hand-eye calibration and obtain the precise transformation relationship between the 3D camera coordinate system and the robot end effector coordinate system.
[0036] In one embodiment, a robot is controlled to move a 3D camera to the area where the target object is located, and the 3D camera collects point cloud data of the target object. The collected point cloud data of the target object is preprocessed, including denoising, filtering, downsampling, and other operations to improve the quality of the point cloud data and the efficiency of subsequent processing. The preprocessed point cloud data of the target object is identified using a feature extraction or deep learning method to determine the category of the target object. At the same time, a point cloud registration algorithm (such as the iterative closest point algorithm (ICP)) is used to align the point cloud data of the target object with the pre-established point cloud data of the target object model to obtain the precise position information of the target object in the 3D camera coordinate system.
[0037] In one embodiment, a 3D camera is calibrated using a calibration plate to obtain the camera's intrinsic and extrinsic parameter matrices. The intrinsic parameter matrix contains information about the camera's internal parameters, such as focal length and principal point coordinates, while the extrinsic parameter matrix describes the transformation between the 3D camera coordinate system and the world coordinate system. Using multiple sets of imaging data from the calibration plate in different positions and postures within the 3D camera, the intrinsic and extrinsic parameter matrices are solved and optimized using the least squares method or other optimization algorithm to improve calibration accuracy.
[0038] In one embodiment, when the target object is occluded, a method based on geometric features, motion patterns or deep learning models is used to reconstruct the local object motion model based on the historical point cloud data of the target object; the method based on geometric features analyzes the shape, size and other geometric information of the historical point cloud data and reconstructs the model in combination with the motion law of the object; the method based on motion patterns infers the possible motion of the object during the occlusion period based on the motion trajectory and speed change law of the object at the historical moment; the method based on deep learning models uses a pre-trained neural network, inputs historical point cloud data, and outputs a reconstructed object model.
[0039] The minimum reconstruction ratio of the motion model is determined based on the actual object category and grasping angle. Different categories of objects have different motion characteristics and structural features, and the grasping angle will also affect the complete requirements of the object model. Taking these factors into consideration, the minimum reconstruction ratio is determined to reduce the amount of calculation while ensuring the effectiveness of the model.
[0040] In one embodiment, a dynamic mapping relationship between the hand-eye matrix and the motion trajectory is established based on the reconstructed object model and the predicted motion trajectory. This dynamic mapping relationship takes into account the object's position and posture changes at different times, as well as the calibration parameters of the hand-eye system, ensuring that the robot can adjust its operations based on the object's real-time motion. Based on this established dynamic mapping relationship, a dynamic grasping path is generated. This dynamic grasping path, based on the predicted motion trajectory of the target object, enables the end-arm to reach the predicted position in advance to achieve accurate grasping.
[0041] like Figure 2 As shown, the embodiment of the present application also provides a robot hand-eye positioning device based on a 3D camera, including:
[0042] at least one processor; and,
[0043] a memory communicatively connected to the at least one processor; wherein,
[0044] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the robot hand-eye positioning device based on a 3D camera to perform:
[0045] Scanning a scene where a target object is located in real time using a 3D camera to determine point cloud data of the target object, wherein the point cloud data includes three-dimensional coordinate information of each point on the surface of the target object;
[0046] Predicting the motion state of the target object according to a predetermined prediction model to obtain a predicted trajectory, matching the predicted trajectory with the point cloud data scanned in real time to determine a prediction difference, the prediction difference including a distance difference and a shape difference, and adjusting the prediction model according to the prediction difference;
[0047] determining whether the target object is occluded, and if the target object is occluded, determining historical point cloud data of the target object, and adjusting the prediction model according to the historical point cloud data;
[0048] The target object is re-predicted using the adjusted prediction model to obtain a prediction result, a pre-set hand-eye matrix is determined, the prediction result is mapped to the hand-eye matrix to dynamically generate the path of the target object, and the robot's end effector is controlled to reach the corresponding predicted position according to the path to achieve grasping at the predicted position.
[0049] The embodiment of the present application further provides a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to:
[0050] Scanning a scene where a target object is located in real time using a 3D camera to determine point cloud data of the target object, wherein the point cloud data includes three-dimensional coordinate information of each point on the surface of the target object;
[0051] Predicting the motion state of the target object according to a predetermined prediction model to obtain a predicted trajectory, matching the predicted trajectory with the point cloud data scanned in real time to determine a prediction difference, the prediction difference including a distance difference and a shape difference, and adjusting the prediction model according to the prediction difference;
[0052] determining whether the target object is occluded, and if the target object is occluded, determining historical point cloud data of the target object, and adjusting the prediction model according to the historical point cloud data;
[0053] The target object is re-predicted using the adjusted prediction model to obtain a prediction result, a pre-set hand-eye matrix is determined, the prediction result is mapped to the hand-eye matrix to dynamically generate the path of the target object, and the robot's end effector is controlled to reach the corresponding predicted position according to the path to achieve grasping at the predicted position.
[0054] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device and medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0055] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0056] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0057] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0058] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0060] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0061] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0062] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0063] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0064] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A robot hand-eye positioning method based on a 3D camera, characterized in that: include: Scanning a scene where a target object is located in real time using a 3D camera to determine point cloud data of the target object, wherein the point cloud data includes three-dimensional coordinate information of each point on the surface of the target object; Predicting the motion state of the target object according to a predetermined prediction model to obtain a predicted trajectory, matching the predicted trajectory with the point cloud data scanned in real time to determine a prediction difference, the prediction difference including a distance difference and a shape difference, and adjusting the prediction model according to the prediction difference; determining whether the target object is occluded, and if the target object is occluded, determining historical point cloud data of the target object, and adjusting the prediction model according to the historical point cloud data; The target object is re-predicted using the adjusted prediction model to obtain a prediction result, a pre-set hand-eye matrix is determined, the prediction result is mapped to the hand-eye matrix to dynamically generate the path of the target object, and the robot's end effector is controlled to reach the corresponding predicted position according to the path to achieve grasping at the predicted position.
2. The method according to claim 1, characterized in that Predicting the motion state of the target object according to a predetermined prediction model to obtain a predicted trajectory specifically includes: Determine an initial state based on the point cloud data of the target object at the current moment, wherein the initial state includes position and speed; The historical point cloud data and related motion features of the target object are used as input to train the LSTM network, thereby outputting the position information of the target object at the future moment and obtaining the predicted trajectory.
3. The method according to claim 1, characterized in that The method further comprises: Mounting the 3D camera on the end effector of a robot to construct a hand-eye system; A calibration object is placed in the workspace of the robot, and the robot is controlled to drive the 3D camera to shoot the calibration object in multiple postures to obtain multiple sets of point cloud data of the calibration object and corresponding end effector posture data.
4. The method according to claim 1, wherein The method further comprises: The hand-eye transformation matrix is determined according to a preset hand-eye calibration equation, so as to perform hand-eye calibration on the robot according to the hand-eye transformation matrix, thereby obtaining the 3D camera coordinate system and determining the conversion relationship between the 3D camera coordinate system and the end effector coordinate system.
5. The method according to claim 1, wherein The method further comprises: Collecting point cloud data of the target object by the 3D camera, and preprocessing the collected point cloud data, wherein the preprocessing includes denoising, filtering, and downsampling; Identifying the preprocessed point cloud data to determine the category of the target object; The point cloud data is aligned with a pre-established target object pose model to obtain the pose information of the target object in the coordinate system of the 3D camera.
6. The method according to claim 1, characterized in that The method further comprises: Calibrate the 3D camera according to a pre-set calibration plate to obtain a hand-eye matrix of the 3D camera, wherein the hand-eye matrix includes an intrinsic parameter matrix and an extrinsic parameter matrix, wherein the intrinsic parameter matrix includes the focal length and principal point coordinates of the 3D camera, and the extrinsic parameter matrix includes the conversion relationship between the 3D camera coordinate system and the end effector coordinate system; The least squares method is used to optimize the intrinsic parameter matrix and the extrinsic parameter matrix to improve the calibration accuracy of the 3D camera.
7. The method according to claim 1, characterized in that Adjusting the prediction model according to the historical point cloud data specifically includes: The geometric information of the historical point cloud data is analyzed, the prediction model is reconstructed according to the geometric information and the motion law of the target object, and the minimum reconstruction ratio of the prediction model is determined according to the category and grasping angle of the target object.
8. The method according to claim 6, characterized in that Dynamically generating the path of the target object specifically includes: A dynamic grasping path is generated according to the dynamic mapping relationship, and a predicted motion trajectory of the target object is determined according to the dynamic grasping path, so that the end effector of the robot reaches the corresponding predicted position in advance to achieve an accurate grasping operation.
9. A robot hand-eye positioning device based on a 3D camera, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the robot hand-eye positioning device based on a 3D camera to perform: Scanning a scene where a target object is located in real time using a 3D camera to determine point cloud data of the target object, wherein the point cloud data includes three-dimensional coordinate information of each point on the surface of the target object; Predicting the motion state of the target object according to a predetermined prediction model to obtain a predicted trajectory, matching the predicted trajectory with the point cloud data scanned in real time to determine a prediction difference, the prediction difference including a distance difference and a shape difference, and adjusting the prediction model according to the prediction difference; determining whether the target object is occluded, and if the target object is occluded, determining historical point cloud data of the target object, and adjusting the prediction model according to the historical point cloud data; The target object is re-predicted using the adjusted prediction model to obtain a prediction result, a pre-set hand-eye matrix is determined, the prediction result is mapped to the hand-eye matrix to dynamically generate the path of the target object, and the robot's end effector is controlled to reach the corresponding predicted position according to the path to achieve grasping at the predicted position.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Scanning a scene where a target object is located in real time using a 3D camera to determine point cloud data of the target object, wherein the point cloud data includes three-dimensional coordinate information of each point on the surface of the target object; Predicting the motion state of the target object according to a predetermined prediction model to obtain a predicted trajectory, matching the predicted trajectory with the point cloud data scanned in real time to determine a prediction difference, the prediction difference including a distance difference and a shape difference, and adjusting the prediction model according to the prediction difference; determining whether the target object is occluded, and if the target object is occluded, determining historical point cloud data of the target object, and adjusting the prediction model according to the historical point cloud data; The target object is re-predicted using the adjusted prediction model to obtain a prediction result, a pre-set hand-eye matrix is determined, the prediction result is mapped to the hand-eye matrix to dynamically generate the path of the target object, and the robot's end effector is controlled to reach the corresponding predicted position according to the path to achieve grasping at the predicted position.
Citation Information
Patent Citations
3D (three-dimensional) quick positioning and grabbing method for industrial mechanical arm
CN118334089A
Target positioning and tracking method and system based on cooperation of PTZ camera and radar
CN118549924A
Interactive prediction learning and training method and device and computer equipment
CN119567280A
Method and system for generating a 3D reconstruction of a human
US20210256776A1
Reactive interactions for robotic applications and other automated systems
US20230294277A1
Cited By
Vision-based automatic graph card alignment method and system for mechanical arm
CN121505043A