Multi-industrial robot 2D and 3D key point estimation method based on RGB-D camera
Through the RGB-D camera-based multi-industrial robot 2D and 3D key point estimation method, the problem of difficult online calibration of the base coordinate system of multiple industrial robots is solved, low-cost, real-time key point measurement is achieved, and the system stability and human-machine collaboration efficiency are improved.
Patent Information
- Application Number
- CN202511236874.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-01
AI Technical Summary
In the existing technology, online calibration of the base coordinate system of industrial robots is difficult, especially in scenarios with multiple industrial robots. Traditional passive estimation methods are costly and affected by the environment, while active estimation methods require physical connection between sensors and robots, making them difficult to adapt to complex scenarios and human-machine collaboration.
A 2D and 3D keypoint estimation method for multiple industrial robots based on RGB-D cameras is adopted. By establishing hardware configuration, generating simulation datasets and network training, real-time keypoint measurement of multiple industrial robots is achieved. YOLOv8, HRNet and IRDepthNet networks are used for keypoint detection and estimation.
It realizes low-cost, real-time 2D and 3D coordinate measurement of key points of multiple industrial robots, reduces equipment costs, and improves system stability and human-machine collaboration efficiency.
Smart Images

Figure CN120740611A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of industrial robot measurement, and specifically relates to a method for estimating 2D and 3D key points of multiple industrial robots based on an RGB-D camera. Background Art
[0002] In high-end equipment industries like aerospace and energy, the machining and assembly of large, complex components, such as large wind turbine blades and space station capsules, is evolving from traditional machine tool operations to collaborative operations involving industrial robots and humans, in order to improve precision and efficiency. Real-time pose estimation of humans and robots during this collaborative process provides a data foundation for human-robot safety and is a crucial component in ensuring the smooth operation of the overall collaborative operation.
[0003] Current approaches to joint pose estimation for industrial robots used in human-machine safety monitoring can be categorized as active and passive. Passive estimation involves establishing a communication link between a computer and the robot, with the industrial robot controller cyclically sending its joint angles to the computer, thereby achieving perception of the robot's actual state. However, the robot's base coordinate system still requires additional calibration, which not only requires costly, wide-field-of-view measurement equipment, but also, in complex field environments, the calibrated measurement field of view can be affected by multiple factors, including the equipment itself. Furthermore, in large component processing scenarios, multiple mobile industrial robots are typically deployed, with AGVs mounted underneath to process different parts of the component. Due to the varying positions of the AGVs and AGV positioning errors, the base coordinate system of the industrial robot being loaded by the AGVs also changes, making online calibration of the base coordinate system extremely difficult. Therefore, passive estimation cannot be used for online measurement of the robot's pose. Finally, this method depends entirely on the industrial robot controller's support for transmitting joint angle information, requiring varying degrees of development and adaptation for different brands of industrial robots.
[0004] Active estimation refers to actively estimating the posture of industrial robots through sensors such as cameras and lidar. This method relies entirely on the sensor's own perception information and has no physical connection with the industrial robot, which reduces the degree of system coupling. While having stronger scene adaptability, it can monitor the robot's safety from an objective perspective and improve system stability. In addition, the robot's active posture estimation can be combined with human posture estimation, human-computer interaction and other technologies to complete the simultaneous estimation of human-computer status and interaction information. It is an important technical approach to achieve human-computer safety monitoring and improve the efficiency of human-computer collaboration. Summary of the Invention
[0005] In response to the problems existing in the above-mentioned prior art, the present invention proposes a method for estimating 2D and 3D key points of multiple industrial robots based on RGB-D cameras, so as to achieve accurate real-time estimation of 2D and 3D key points of multiple industrial robots and objectively and stably monitor the robot posture.
[0006] In order to achieve the above technical objectives, the present invention provides the following technical solutions: A method for estimating 2D and 3D key points of multiple industrial robots based on RGB-D cameras, which specifically includes: S1. Establish the hardware configuration of the industrial robot measurement system, which includes an RGB-D camera and multiple industrial robots. S2. Establish a simulation data set generation system, select key points of the industrial robot, and generate a simulation data set of key point coordinates based on the virtual environment; S3. Build a coordinate detection network based on RGB-D camera and train the coordinate detection network using simulation data sets; S4. Use the trained coordinate detection network to achieve real-time observation of the 2D and 3D coordinates of key points of multiple industrial robots.
[0007] Furthermore, in step S1, multiple industrial robots are set to be within the measurement field of view of the RGB-D camera, and the flanges of the multiple industrial robots can be connected to various actuator mechanisms that do not block the camera's measurement field of view.
[0008] Furthermore, step S2 specifically includes: S21. There are 9 key points for each industrial robot. S22. Setting generation targets of the simulation data set generation system to include: two-dimensional coordinates of multiple key points of the industrial robot in the image, three-dimensional coordinates of multiple key points of the industrial robot in the camera coordinate system, a color image captured by an RGB-D camera, a depth image captured by an RGB-D camera, and a bounding box of the industrial robot in the image; S23. Build a simulation dataset generation system based on the Unreal game engine. The system consists of a simulation environment, multiple virtual RGB-D cameras, and multiple virtual industrial robots. The virtual cameras capture the virtual industrial robots in real time from different angles. S24. The operating logic of the designed dataset generation system is as follows: first, the virtual robot performs random movements by generating random numbers in the feasible joint space, and detects its own collisions and external collisions to form the robot's achievable posture in the real environment; then, multiple virtual RGB-D cameras sequentially shoot and calculate the 3D point coordinates of the 9 key points of the pre-selected industrial robot relative to the camera, and project the 3D points into the 2D space to obtain the image coordinates. Finally, the data are combined into the same label file, where the depth map is encoded in exr format, the color map is in png format, and the label file is in txt format.
[0009] Furthermore, step S3 specifically includes: S31, the coordinate detection network consists of the YOLOv8 network, the HRNet network and the industrial robot key point depth estimation network IRDepthNet; S32. Use the simulation data set to train the YOLOv8 network, HRNet network, and IRDepthNet network separately, and use the data enhancement method to add Gaussian noise, Gaussian blur the image, and randomly adjust the contrast, saturation, and brightness of the image to simulate the image effects in different real environments; S33. Establish a top-down detection mechanism, specifically: the coordinate detection network inputs the color image captured by the RGB-D camera, first passes it through the YOLOv8 network, and obtains the 2D bounding box of the industrial robot in the color image and the cropped image within the 2D bounding box; the cropped image within the 2D bounding box passes through the HRNet network to detect the key points of the industrial robot and obtain a heat map of 9 key points; the heat map Figure 1 On the one hand, the 2D coordinates of the key points of the industrial robot are obtained through maximum response processing, and on the other hand, the 2D coordinates of the key points of the industrial robot are obtained through depth map and color image processing. Figure 1 The 3D coordinates of the key points of the industrial robot are obtained by sending them into the IRDepthNet network.
[0010] Furthermore, during the training process, the HRNet network adjusts the number of output channels to 18 channels, and each channel outputs a 2D coordinate value of a key point.
[0011] Furthermore, the IRDepthNet network is: The IRDepthNet network consists of a splicing layer, a ResNet50 network and several fully connected layers; the input depth map and color image are first cropped and padded, and the length, width and number of channels of the image are adjusted, and then spliced with the key point heat map in the splicing layer; then it passes through the ResNet50 network and several fully connected layers of different scales, and through continuous convolution operations, finally obtains the xyz three-axis coordinate values of the 3D coordinates of the 9 key points.
[0012] Furthermore, the loss function of the IRDepthNet network is It is designed as the sum of the absolute MSE error and the relative MSE error, specifically: ; ; ; ; ; in, is the absolute MSE error, is the relative MSE error, Respectively i The x-axis, y-axis, and z-axis coordinates of the predicted key points; Respectively i The x-axis, y-axis, and z-axis coordinates of the theoretical key points; For the i The prediction key point and The distance between the predicted key points, For the i The key theoretical point and The distance between the theoretical key points.
[0013] Based on the above technical solution, the present invention has at least the following beneficial effects: The method proposed in this invention can complete the real-time key point measurement of multiple industrial robots using only a low-cost RGB-D camera, reducing equipment costs while achieving real-time measurement output of the 2D and 3D coordinates of the key points of multiple industrial robots, thus proposing an effective solution for the measurement of the key points of multiple industrial robots. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Schematic diagram of the layout of the RGB-D camera and industrial robot in this application; 1 in the figure is the RGB-D camera, and 2 is the industrial robot; Figure 2 A schematic diagram of the key points of the industrial robot selected in this application; Figure 3 Generate a system operation flow chart for the simulation data set in this application; Figure 4 This is the overall flow chart of the 2D and 3D key point estimation method for multiple industrial robots based on RGB-D cameras proposed in this application; Figure 5 The IRDepthNet network structure diagram designed for this invention; Figure 6This is the robot bounding box detection effect diagram on the simulation dataset; Figure 7 This is the effect diagram of the robot key point 2D coordinate detection on the simulation dataset; Figure 8 This is the estimated effect diagram of the 3D coordinates of the robot's key points. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solutions and advantages of the present invention more clear, the following Figure 1-8 It should be understood that the specific embodiments described herein are only used to illustrate the present invention and are not intended to limit the present invention.
[0016] Although the steps in the present invention are arranged with numbers, they are not intended to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a step requires other steps as a basis, the relative order of the steps can be adjusted. It is understood that the term "and / or" used herein refers to and covers any and all possible combinations of one or more of the associated listed items.
[0017] like Figure 4 As shown, the present invention proposes a method for estimating 2D and 3D key points of multiple industrial robots based on RGB-D cameras, which specifically includes the following steps: S1. Establish the hardware configuration of the industrial robot's measurement system, such as Figure 1 As shown, the measurement system hardware configuration includes an RGB-D camera and multiple industrial robots. In this embodiment, the multiple industrial robots are set within the measurement field of view of the RGB-D camera, and the flanges of the multiple industrial robots can be connected to various actuator mechanisms that do not block the camera's measurement field of view.
[0018] S2. Establish a simulation data set generation system, select key points of the industrial robot, and generate a simulation data set of key point coordinates based on the virtual environment; As a preferred embodiment, step S2 specifically includes: S21, such as Figure 2 As shown, there are 9 key points selected for each industrial robot; S22. Setting generation targets of the simulation data set generation system to include: two-dimensional coordinates of multiple key points of the industrial robot in the image, three-dimensional coordinates of multiple key points of the industrial robot in the camera coordinate system, a color image captured by an RGB-D camera, a depth image captured by an RGB-D camera, and a bounding box of the industrial robot in the image; S23. Build a simulation dataset generation system based on the Unreal game engine. The system consists of a simulation environment, multiple virtual RGB-D cameras, and multiple virtual industrial robots. The virtual cameras capture the virtual industrial robots in real time from different angles. S24. The operating logic of the designed dataset generation system is as follows: first, the virtual robot performs random movements by generating random numbers in the feasible joint space, and detects its own collisions and external collisions to form the robot's achievable posture in the real environment; then, multiple virtual RGB-D cameras sequentially shoot and calculate the 3D point coordinates of the 9 key points of the pre-selected industrial robot relative to the camera, and project the 3D points into the 2D space to obtain the image coordinates. Finally, the data are combined into the same label file, where the depth map is encoded in exr format, the color map is in png format, and the label file is in txt format.
[0019] S3. Build a coordinate detection network based on RGB-D camera and train the coordinate detection network using simulation data sets; As a preferred embodiment, step S3 specifically includes: S31, the coordinate detection network consists of the YOLOv8 network, the HRNet network and the industrial robot key point depth estimation network IRDepthNet; S32. Use the simulation data set to train the YOLOv8 network, HRNet network, and IRDepthNet network separately, and use the data enhancement method to add Gaussian noise, Gaussian blur the image, and randomly adjust the contrast, saturation, and brightness of the image to simulate the image effects in different real environments; S33. Establish a top-down detection mechanism, specifically: the coordinate detection network inputs the color image captured by the RGB-D camera, first passes it through the YOLOv8 network, and obtains the 2D bounding box of the industrial robot in the color image and the cropped image within the 2D bounding box; the cropped image within the 2D bounding box passes through the HRNet network to detect the key points of the industrial robot and obtain a heat map of 9 key points; the heat map Figure 1 On the one hand, the 2D coordinates of the key points of the industrial robot are obtained through maximum response processing, and on the other hand, the 2D coordinates of the key points of the industrial robot are obtained through depth map and color image processing. Figure 1 The 3D coordinates of the key points of the industrial robot are obtained by sending them into the IRDepthNet network.
[0020] In addition, in this embodiment, during the training process, the HRNet network adjusts the number of output channels to 18 channels, and each channel outputs a 2D coordinate value of a key point.
[0021] The specific structure of the IRDepthNet network is as follows Figure 5As shown in the figure, it includes a splicing layer, a ResNet50 network and several fully connected layers; the input depth map and color map are first cropped and filled, and the length, width and number of channels of the image are adjusted, and then spliced with the key point heat map in the splicing layer; then it passes through the ResNet50 network and several fully connected layers of different scales, and through continuous convolution operations, the xyz three-axis coordinate values of the 3D coordinates of the 9 key points are finally obtained.
[0022] At the same time, this application also designs a loss function for the training of the IRDepthNet network , specifically designed as the sum of the absolute MSE error and the relative MSE error, the formula is expressed as: ; ; ; ; ; in, is the absolute MSE error, is the relative MSE error, Respectively i The x-axis, y-axis, and z-axis coordinates of the predicted key points; Respectively i The x-axis, y-axis, and z-axis coordinates of the theoretical key points; For the i The prediction key point and The distance between the predicted key points, For the i The key theoretical point and The distance between the theoretical key points.
[0023] S4. Use the trained coordinate detection network to achieve real-time observation of the 2D and 3D coordinates of key points of multiple industrial robots.
[0024] So far, the entire process of the method proposed in the present invention has been introduced.
[0025] This embodiment also provides the specific detection effect of the method proposed by the present invention, such as Figure 6-8 As shown; in, Figure 6 The 2D bounding box detection effect of an industrial robot in a color image in a simulation environment is demonstrated. The red rectangular box in the image represents the industrial robot. It can be seen that the industrial robot can be accurately identified in any scene. Figure 7The results of estimating the 2D coordinates of key points of an industrial robot in a color image in a simulation environment are shown. Adjacent key points in the image are connected by straight lines, and it can be seen that the detected key point lines accurately outline the skeleton of the industrial robot, reflecting the accurate estimation of 2D key points. Figure 8 The figure shows the estimated effect of the 3D coordinates of the key points, where blue is the predicted value and red is the theoretical value. It can be seen that the predicted value of the 3D coordinates of the key points is also highly consistent with the theoretical value.
[0026] In summary, the method proposed in the present invention can complete the real-time key point measurement of multiple industrial robots using only low-cost RGB-D cameras, reducing equipment costs while achieving real-time measurement output of the 2D and 3D coordinates of the key points of multiple industrial robots, thus proposing an effective solution for the measurement of key points of multiple industrial robots.
[0027] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
[0028] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. A method for estimating 2D and 3D key points of multiple industrial robots based on RGB-D cameras, characterized in that: The specific steps include: S1. Establish the hardware configuration of the industrial robot measurement system, which includes an RGB-D camera and multiple industrial robots. S2. Establish a simulation data set generation system, select key points of the industrial robot, and generate a simulation data set of key point coordinates based on the virtual environment; S3. Build a coordinate detection network based on an RGB-D camera and train it using a simulation dataset. This includes: S31, the coordinate detection network consists of the YOLOv8 network, the HRNet network and the industrial robot key point depth estimation network IRDepthNet; S32. Use the simulation data set to train the YOLOv8 network, HRNet network, and IRDepthNet network separately, and use the data enhancement method to add Gaussian noise, Gaussian blur the image, and randomly adjust the contrast, saturation, and brightness of the image to simulate the image effects in different real environments; S33. Establish a top-down detection mechanism, specifically: the coordinate detection network inputs the color image captured by the RGB-D camera, first passes it through the YOLOv8 network, and obtains the 2D bounding box of the industrial robot in the color image and the cropped image within the 2D bounding box; the cropped image within the 2D bounding box is passed through the HRNet network to detect the key points of the industrial robot, and obtain a heat map of 9 key points; the heat map is processed by maximum response to obtain the 2D coordinates of the key points of the industrial robot, and is sent to the IRDepthNet network together with the depth map and color map to obtain the 3D coordinates of the key points of the industrial robot; S4. Use the trained coordinate detection network to achieve real-time observation of the 2D and 3D coordinates of key points of multiple industrial robots.
2. The method for estimating 2D and 3D key points of multiple industrial robots based on RGB-D cameras according to claim 1, characterized in that: In step S1, multiple industrial robots are set to be within the measurement field of view of the RGB-D camera, and the flanges of the multiple industrial robots can be connected to various actuator mechanisms that do not block the camera's measurement field of view.
3. The method for estimating 2D and 3D key points of multiple industrial robots based on RGB-D cameras according to claim 1, characterized in that: Step S2 specifically includes: S21. There are 9 key points for each industrial robot. S22. Setting generation targets of the simulation data set generation system to include: two-dimensional coordinates of multiple key points of the industrial robot in the image, three-dimensional coordinates of multiple key points of the industrial robot in the camera coordinate system, a color image captured by an RGB-D camera, a depth image captured by an RGB-D camera, and a bounding box of the industrial robot in the image; S23. Build a simulation dataset generation system based on the Unreal game engine. The system consists of a simulation environment, multiple virtual RGB-D cameras, and multiple virtual industrial robots. The virtual cameras capture the virtual industrial robots in real time from different angles. S24. The operating logic of the designed dataset generation system is as follows: first, the virtual robot performs random movements by generating random numbers in the feasible joint space, and detects its own collisions and external collisions to form the robot's achievable posture in the real environment; then, multiple virtual RGB-D cameras sequentially shoot and calculate the 3D point coordinates of the 9 key points of the pre-selected industrial robot relative to the camera, and project the 3D points into the 2D space to obtain the image coordinates. Finally, the data are combined into the same label file, where the depth map is encoded in exr format, the color map is in png format, and the label file is in txt format.
4. The method for estimating 2D and 3D key points of multiple industrial robots based on RGB-D cameras according to claim 1, characterized in that: During the training process, the HRNet network adjusts the number of output channels to 18 channels, and each channel outputs a 2D coordinate value of a key point.
5. The method for estimating 2D and 3D key points of multiple industrial robots based on RGB-D cameras according to claim 1, characterized in that: The IRDepthNet network is: The IRDepthNet network consists of a splicing layer, a ResNet50 network and several fully connected layers; the input depth map and color image are first cropped and padded, and the length, width and number of channels of the image are adjusted, and then spliced with the key point heat map in the splicing layer; then it passes through the ResNet50 network and several fully connected layers of different scales, and through continuous convolution operations, finally obtains the xyz three-axis coordinate values of the 3D coordinates of the 9 key points.
6. The method for estimating 2D and 3D key points of multiple industrial robots based on RGB-D cameras according to claim 5, characterized in that: Loss function of IRDepthNet network It is designed as the sum of the absolute MSE error and the relative MSE error, specifically: ; ; ; ; ; in, is the absolute MSE error, is the relative MSE error, Respectively i The x-axis, y-axis, and z-axis coordinates of the predicted key points; Respectively i The x-axis, y-axis, and z-axis coordinates of the theoretical key points; For the i The prediction key point and The distance between the predicted key points, For the i The key theoretical point and The distance between the theoretical key points.
Citation Information
Patent Citations
Melon and fruit vegetable size measurement method based on depth camera and key points
CN113932712A
Visual SLAM method for indoor dynamic scene based on SOLOv2
CN118089730A
Deburring method combining 2D image and 3D point cloud
CN120495371A
Heavy-load industrial robot space safety man-machine cooperation control method
CN120552011A
Robotic system with 3D box location functionality
US20150224648A1