Target board pose tracking device and method based on deep learning

CN120655708APending Publication Date: 2025-09-16CHINA NUCLEAR POWER OPERATION TECH CORP +2

Patent Information

Application Number
CN202510682195.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

[0004]本发明提供一种基于深度学习的靶标板位姿跟踪装置及跟踪方法,用于解决现有技术中智能抓取机械臂精度不够的问题

Benefits of technology

[0037] 1. The present invention proposes a target plate pose tracking device based on deep learning. Traditional target pose estimation methods are easily affected by illumination, motion, occlusion, perspective change, distortion, etc., resulting in target detection loss and reduced pose estimation accuracy. The software system of this device is based on a tracking method based on deep learning. Taking into account the imaging quality deterioration mechanism, it constructs imaging quality deterioration data through data simulation, and first trains a pre-training model using a simulated data set. Then, part of the data is collected in an artificial simulation environment as real data for training the target detection model. Based on the continuous frame depth detection features and the continuity of motion, combined with the deepsort framework and the kalman filtering method, the accuracy and robustness of real-time target tracking are improved. Finally, accurate pose results are obtained by combining standard template data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655708A_ABST
    Figure CN120655708A_ABST
Patent Text Reader

Abstract

The invention relates to the field of robot automation control, in particular to a target board pose tracking device and method based on deep learning, the device comprises a sensor, a target board, an image processor and a displayer, the target board is arranged at the tail end of a mechanical arm, and the sensor is used for collecting image data of the target board and the mechanical arm; the sensor is connected with an image processor through a data line, a software system is integrated in the image processor, the image processor is connected with a displayer, and the displayer displays the posture of the mechanical arm. A software system of the device is based on a deep learning tracking method, an imaging quality deterioration mechanism is considered, imaging quality deterioration data is constructed in a data simulation mode, a simulation data set is adopted to train a pre-training model, then part of data is collected in an artificial simulation environment to serve as real data to be used for training a target detection model, and a target detection result is obtained. And on the basis of continuous frame depth detection features and motion continuity, the accuracy and robustness of target real-time tracking are improved, and finally, an accurate pose result is obtained in combination with standard template data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot automation control, and in particular to a target plate posture tracking device and tracking method based on deep learning. Background Art

[0002] Autonomous grasping technology with robotic arms is a research hotspot and a key challenge in the field of robotics. Currently, this technology is widely used in intelligent logistics sorting, smart warehousing, smart homes, and other fields. Obtaining high-precision positioning of the grasped target in complex environments is crucial for successful autonomous grasping tasks, placing higher demands on the robotic arm's intelligence.

[0003] Existing technologies have used deep learning methods to train robotic arms, achieving intelligent grasping through target recognition and grasping point positioning. However, current research requires high-precision robotic arms for target recognition and detection. However, due to factors such as their inherent configuration, manufacturing accuracy, calibration accuracy, and load, existing robotic arms often have significant issues with their final feedback pose accuracy and control precision. This is particularly true for robotic arms with unconventional configurations performing feature-based operations. The final accuracy of the robotic arms often struggles to meet the high demands of specialized tasks. Summary of the Invention

[0004] The present invention provides a target plate posture tracking device and tracking method based on deep learning, which are used to solve the problem of insufficient precision of intelligent grasping robotic arms in the prior art.

[0005] The technical solutions of the present invention are as follows:

[0006] The present invention proposes a target plate posture tracking device based on deep learning. The device includes a sensor, a target plate, an image processor and a display. The target plate is arranged at the end of the robotic arm. The sensor is used to collect image data of the target plate and the robotic arm. The sensor is connected to the image processor via a data cable. The image processor is connected to the display, and the display displays the posture of the robotic arm.

[0007] In some embodiments, the sensor is a visual sensor, and the sensor is a monocular camera visual sensor, a binocular stereo camera visual sensor, or a structured light camera visual sensor.

[0008] In some embodiments, the sensor is an underwater visual sensor, and the sensor is equipped with a matching light source.

[0009] In some embodiments, the target plate is a hexagonal snowflake-shaped structure, and six different Apriltag targets are respectively set at the six corners of the snowflake-shaped structure of the target plate, and the six Apriltag targets are in the same plane.

[0010] In some embodiments, the image processor is an industrial computer, and the industrial computer integrates a software system. The software system part mainly includes an image processing module, a manual interaction module and a guidance control module. The image processing module mainly includes a target pose estimation unit, a blocking plate air lock hole target pose estimation unit and a target recognition and pose estimation unit; the manual interaction module transmits the image measured by the sensor back to the display for display, and marks and displays the working target and the end of the robotic arm in the image; the guidance control module obtains real-time robotic arm control pose parameters according to the robotic arm pose and target pose output by the image processing module and the difference between the robotic arm pose and the target pose. The guidance control module provides the robotic arm control pose parameters to the robotic arm controller to guide the end control of the working robotic arm.

[0011] In some embodiments, the target pose estimation unit obtains the target pose of the end of the robot arm based on deep learning to obtain the pose of the end of the robot arm, the air-lock hole target pose estimation unit obtains the air-lock pose based on deep learning, and the target recognition and pose estimation module realizes the recognition and pose determination of the working target of the robot arm based on binocular stereo vision.

[0012] In some embodiments, the target pose estimation unit includes a target tracking function, a target pose estimation function and a target plate pose solution function. The target tracking function uses the yoloV8 network to perform target position detection and the DeepSort tracking module to perform target position tracking. The yoloV8 network uses a simulated data set to first train a pre-training model, and then collects part of the data as real data in an artificial simulation environment to fine-tune the pre-training model; the DeepSort tracking module includes kalman prediction, deep feature matching and Hungarian algorithm; the target pose estimation unit uses the yoloV8 network output vector as the deep feature of the target pattern, uses the cosine function to calculate the similarity of the deep feature vector, combines the state information of the kalman prediction, and performs matching association of the detection targets of the previous and next consecutive frames according to the Hungarian algorithm to achieve correct target detection output.

[0013] In some embodiments, after the target plate tracking and identification is completed, the target pose estimation function of the target pose estimation unit detects and extracts the target point of the Apriltag target for the monocular camera vision sensor, obtains the pixel coordinates of the target point, matches the corner points of the Apriltag target with the corner points of the standard target template, obtains the rotation and translation transformation strategy from the Apriltag target image to the standard target template image, obtains the correspondence between the pixel coordinates of the target point in the Apriltag target and the world coordinates of the standard target template, and uses the PNP algorithm to calculate the pose of the target plate center in the camera coordinate system. The PNP algorithm equation is as shown in formula (1):

[0014]

[0015] Where [u,v] is the pixel coordinate of the Apriltag target, K is the camera intrinsic parameter matrix, [R|t] is the required pose matrix, and [x,y,z] is the world coordinate of the standard target template;

[0016] For binocular stereo camera vision sensors, the pose of the same Apriltag target is calculated in the field of view of two cameras at the same time. The three-dimensional coordinates of the four vertices of the target plate are solved using the coordinate system conversion matrix calibrated between the cameras. The three-dimensional coordinate calculation equations are as follows:

[0017]

[0018] Among them, (u1, v1) and (u2, v2) are the pixel coordinates of the corresponding points in camera 1 and camera 2 of the binocular stereo camera, K1 and K2 are the intrinsic parameter matrices of camera 1 and camera 2 respectively, [R1|t1] and [R2|t2] are the extrinsic parameter matrices of camera 1 and camera 2, (x w ,y w , z w ) are three-dimensional coordinates;

[0019] For a structured light camera visual sensor, the three-dimensional coordinates of the four vertices are obtained through the correspondence between the image and the point cloud. After the three-dimensional coordinates of the target point in the camera coordinate system are calculated, the three-dimensional coordinates of the target point in the local coordinate system with the target as the target are obtained based on the relative relationship between the target points. The pose of the target is calculated by solving the equation, as shown in formula (4):

[0020]

[0021] Among them, (x w ,y w , z w ) is the three-dimensional coordinate of the camera coordinate system, (x o ,y o , z0) is the three-dimensional coordinate of the target point in the local coordinate system with the target, [R c |t c ] is the pose of the target to be solved.

[0022] In some embodiments, after the Apriltag target pose estimation is completed, the target plate pose solving function of the target pose estimation unit determines the pose of the target plate according to the target plate calibration matrix; for multiple Apriltag target pose results, the pose result of the target plate is obtained by averaging, as shown in formula (5):

[0023]

[0024] Among them, C Tag is the pose result of the target plate, Target collection.

[0025] The present invention proposes a target plate posture tracking device based on deep learning, the method comprising:

[0026] Step 1: Assemble the sensor, image processor, and display, ensure that the sensor data cable is connected to the image processing industrial computer, connect the device power supply, and calibrate the sensor.

[0027] Step 2: Use a structured light 3D camera to collect target plate calibration data and calibrate the conversion relationship between the coordinate system of each target surface of the target plate and the overall coordinate system of the target plate;

[0028] Step 3: Install the target plate to the target to be measured, and calibrate the conversion relationship between the tool coordinate system to be measured and the target plate coordinate system.

[0029] Step 4: Place the sensor in the measurement scene to ensure that the field of view covers the range of movement of the target to be measured, run the software system of the industrial computer, collect the target plate image when the target to be measured moves in real time, output the target plate pose in real time, and convert it into the pose of the target coordinate system according to the conversion relationship.

[0030] In some embodiments, step 2 specifically includes:

[0031] Step 2.1: Place the target plate facing the structured light 3D camera at a distance of 400mm, 500mm, 600mm, and 700mm from the camera.

[0032] Step 2.2: The camera structured light 3D camera extracts each Apriltag target of the three target plates and calculates the pose of each Apriltag target;

[0033] Step 2.3: For each target plate, extract the four corner points of one of the Apriltag targets on the target plate and obtain the coordinates of the target plate center point. Calculate the translation vector of the Apriltag target relative to the target plate center and obtain the transformation matrix of the Apriltag target relative to the target plate center.

[0034] Step 2.4: Calculate the transformation matrices of the remaining five Apriltag target patterns relative to the Apriltag target, and convert the transformation matrix between the Apriltag target and the target plate center into the transformation matrix between the remaining five Apriltag targets and the target plate center to complete the calibration of each target plate;

[0035] Step 2.5: The calibration results at different positions output a set of rotation angle components and translation components obtained by decomposition of the pose matrix. The final calibration result is obtained by averaging the results of multiple groups of angle components and translation components.

[0036] The implementation of the present invention has the following beneficial effects:

[0037] 1. The present invention proposes a target plate pose tracking device based on deep learning. Traditional target pose estimation methods are easily affected by illumination, motion, occlusion, perspective change, distortion, etc., resulting in target detection loss and reduced pose estimation accuracy. The software system of this device is based on a tracking method based on deep learning. Taking into account the imaging quality deterioration mechanism, it constructs imaging quality deterioration data through data simulation, and first trains a pre-training model using a simulated data set. Then, part of the data is collected in an artificial simulation environment as real data for training the target detection model. Based on the continuous frame depth detection features and the continuity of motion, combined with the deepsort framework and the kalman filtering method, the accuracy and robustness of real-time target tracking are improved. Finally, accurate pose results are obtained by combining standard template data.

[0038] 2. This invention proposes a deep learning-based target plate pose tracking device. This device utilizes a hexagonal snowflake-shaped target plate composed of multiple targets. This overcomes the tracking failures often associated with traditional single-target detection, which are susceptible to loss, occlusion, or detection failure. This ensures the success rate and robustness of target pose tracking. Furthermore, the multi-target detection results are leveraged to improve the overall detection accuracy of the target plate.

[0039] 3. This invention proposes a target plate pose tracking device based on deep learning. Traditional target tracking methods can only track the pose of the end-arm, but cannot take into account the observation of other targets in the scene. When tracking the pose of a single target at the end-arm, it is easily lost due to occlusion, changes in viewing angle, etc. The software system of this device can simultaneously track the pose of the end-arm while observing other targets in the environment. It can effectively overcome the occlusion caused by the movement of the robot arm, the tracking loss caused by changes in viewing angle, and the robustness issues. It is used for precise measurement and real-time tracking of the six-degree-of-freedom pose of the end-arm, with an accuracy better than 1mm. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A schematic diagram of a target plate posture tracking device based on deep learning proposed in an embodiment of the present invention;

[0041] Figure 2 A schematic diagram of a target pattern of a target plate pose tracking device based on deep learning proposed in an embodiment of the present invention;

[0042] Figure 3A schematic diagram of a target plate of a target plate posture tracking device based on deep learning proposed in an embodiment of the present invention;

[0043] Figure 4 This is a system flow chart of a target pose estimation unit of a target plate pose tracking device based on deep learning proposed in an embodiment of the present invention;

[0044] Figure 5 A target plate pose graph of a target plate pose tracking device based on deep learning proposed in an embodiment of the present invention;

[0045] Figure 6 A flow chart of a target pose tracking algorithm for a target plate pose tracking device based on deep learning proposed in an embodiment of the present invention;

[0046] Figure 7 A flow chart of the Kalman filter algorithm for a target plate posture tracking device based on deep learning proposed in an embodiment of the present invention;

[0047] Figure 8 This is a flow chart of a target point extraction and pose estimation algorithm for a target plate pose tracking device based on deep learning proposed in an embodiment of the present invention;

[0048] Figure 9 This is an example diagram of the target detection effect of a target plate posture tracking device based on deep learning proposed in an embodiment of the present invention;

[0049] Description of the accompanying drawings: 1. Sensor; 2. Industrial computer; 3. Display; 4. Robotic arm; 5. Target board. DETAILED DESCRIPTION

[0050] The technical solution of the present invention is clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0051] like Figures 1 to 9 As shown, the present invention proposes a target plate posture tracking device based on deep learning, which includes a sensor 1, a target plate 5, an image processor and a display 3. The target plate 5 is arranged at the end of the robotic arm 4. The sensor 1 is a visual sensor 1, which is used to collect image data of the target plate 5. The sensor 1 is connected to the image processor through a data cable, and the image processor is connected to the display 3. The display 3 displays the posture of the robotic arm 4.

[0052] Sensor 1 can be a monocular camera, a binocular stereo camera, or a structured light camera. This patent solution can be used for underwater robotic arm end-of-arm position tracking. Therefore, sensor 1 can be extended to underwater scenes and can be an underwater monocular, binocular, or structured light camera. Sensor 1 also includes a matching light source.

[0053] like Figure 2 and Figure 3 As shown, the target plate 5 utilizes a hexagonal snowflake structure, with six different Apriltag targets positioned at each corner. This ensures that at least one target is detected during robotic arm 4 operation, allowing for a relatively accurate determination of the target plate 5's pose. The six different Apriltag targets are coplanar, facilitating accurate calibration of the target plate 5's pose and the conversion matrix between each target's pose. The size of the target plate 5 can be adjusted based on factors such as the measurement distance and camera focal length to ensure that the target can be properly identified and located in the camera image.

[0054] Target plate 5 is primarily used to mark and locate the end-point of robotic arm 4. Its design principles prioritize visual saliency, clear distinction from the background, high tracking efficiency, high positioning accuracy, 3D pose calculation capabilities, high integration, and ease of installation. The design principles for target plate 5 are as follows: Saliency: distinguishable from the background image, making it easily detectable; Positioning Accuracy: high positioning accuracy and easy calibration; Redundancy: capable of detection and positioning from multiple viewing angles or even in the presence of partial occlusion; End-point installation requirements: minimal size and light weight to minimize interference with robotic arm 4 operations; Corrosion, wear, and deformation resistance: surface resistance to wear and deformation during use; and Ease of Installation and Removal: easy replacement.

[0055] The image processor is the industrial computer 2. The image processor is also the industrial computer 2. The industrial computer 2 receives the image, depth, point cloud and other data collected by the sensor 1, processes the data, identifies the target plate 5, and solves the end position of the robotic arm 4.

[0056] like Figure 4As shown, a software system is integrated within the industrial computer 2. The software system primarily comprises an image processing module, a human interaction module, and a guidance control module. The image processing module primarily includes a target pose estimation unit, a blocking plate air-lock hole target pose estimation unit, and a target recognition and pose estimation module. The target pose estimation unit uses deep learning to determine the target pose at the end of the robotic arm 4, thereby determining the end pose of the robotic arm 4. The air-lock hole target pose estimation unit also uses deep learning to determine the air-lock pose. The air-lock is a pneumatically controlled locking device located at the end of the robotic arm 4 and primarily used by the robotic arm 4 to grasp the blocking plate. The target recognition and pose estimation module uses binocular stereo vision to identify the working target and determine the pose of the robotic arm 4. The human interaction module transmits and displays images measured by the sensor 1 and annotates the working target and the end of the robotic arm 4 in the image, facilitating operator observation and emergency stop control. The guidance control module mainly obtains the real-time control posture parameters of the robot arm 4 based on the posture and target posture of the robot arm 4 output by the image processing module and the difference between the posture of the robot arm 4 and the working target posture of the robot arm 4, and provides it to the robot arm 4 controller to guide the end control of the working robot arm 4.

[0057] The target pose estimation unit includes target tracking, target pose estimation, and target plate pose solution. The target tracking function of the target pose estimation unit mainly uses the yoloV8 (You Only Look Once V8) network for target position detection and the DeepSort tracking module for target position tracking.

[0058] The existing yoloV8 pre-trained model does not have a data training model for the target being used, so it is necessary to collect a data set and retrain the detection model for the target being used. Deep learning network model training mainly involves obtaining a training set. Since the actual application scenario is inside a nuclear power plant and the radiation dose is high, it is difficult to enter with a handheld device for shooting. In addition, the recognition scenario used is relatively simple and the recognition target does not change much. Therefore, artificial simulation is used to generate data to obtain more simulated data. The pre-trained model is first trained using the simulated data set, and then some data is collected in the artificial simulation environment as real data to fine-tune the pre-trained model.

[0059] The YoloV8 network's simulated dataset uses industrial scene images featuring steel and other equipment, resembling on-site operational scenarios, as background images. Standard target template images are subjected to geometric transformations such as rotation, translation, shearing, perspective, and scaling, with grayscale variations and noise added. These images are then randomly superimposed onto the background image. The background image is also randomly affine transformed, its color shading altered, noise added, and motion blur added accordingly. A mask box serves as a pixel-level annotation tool for the target template image. The mask box of the target template image undergoes the same geometric processing as the target image to obtain the ground truth box of the target block. Using the simulated dataset as the training set, an initial pre-trained model for target plate 5 detection is first trained. A series of images containing and excluding the target scene are then captured in a simulated field block scenario. These images serve as the actual scene dataset, and the pre-trained model is fine-tuned to obtain the final target detection model.

[0060] The DeepSort tracking module further tracks the target detection results of the yoloV8 network. After using the yoloV8 network to detect each frame of the video image to obtain the target area, since the yoloV8 network detection results cannot absolutely guarantee the accuracy of the detection results, it is necessary to combine the detection results of the previous frames to judge the current detection results. And sometimes due to problems such as changes in lighting and occlusion, the target in some frame images cannot be successfully detected, so it is also necessary to combine the detection results of the previous moment for optimization. The DeepSort target tracking module is introduced in this solution. The DeepSort framework mainly includes Kalman prediction, deep feature matching and Hungarian algorithm.

[0061] The Kalman prediction algorithm process is as follows Figure 7 As shown, Figure 7 middle is the state vector of the target at time k-1, P k-1 is the covariance matrix between the state quantities at time k-1. k-1 represents the control amount of the system at time k, z k represents the measured value at time k, A, B, and H are state transformation matrices, which are adjustment coefficients in the state transformation process, and are constants here. Matrix B represents the gain of the optional control input variable u, and matrix H represents the state variable x. k For the measured variable z k The gain. Q is the process excitation noise covariance matrix, R is the observation noise covariance matrix, Q and R are constants, K k is the Kalman gain.

[0062] The target tracking function of the target pose estimation unit uses the output vectors of the BackBone main frame of the YoloV8 network as the deep features of the target pattern. These deep feature vectors contain information such as the shape, texture, and color of the target in the image, which can be used to distinguish different targets. The cosine function is used to calculate the similarity of the feature vectors. By calculating the cosine similarity of the feature vectors of the targets in the previous and next frames, it can be determined whether they are the same target. At the same time, combined with the state information predicted by Kalman, the Hungarian algorithm is used to match and associate the detection targets of the previous and next consecutive frames. In particular, for partially occluded low-confidence candidate targets, the correct target detection is output based on the degree of similarity between the previous and next frames.

[0063] The target pose estimation function is for the pose estimation of the monocular camera visual sensor 1. After the target plate 5 is tracked and identified, the target point in the target plate 5 pattern needs to be detected and extracted to obtain the pixel coordinates of the target point. Then, the target point in the target pattern needs to be matched with the standard target template to obtain the correspondence between the pixel point in the image block and the target template point, and then the world coordinate of the target point is obtained. The world coordinate is a rectangular coordinate system established with the center of the standard target template as the circle point, the length and width of the target template as the X and Y axes respectively, and the plane where the target template is located as the Z axis. The world coordinate of the target point can be obtained in advance based on the actual size of the target template and the design of the point. After obtaining the correspondence between the world coordinate and the pixel coordinate of the target point, the PNP (Perspective-n-Point) algorithm is used to calculate the pose of the center of the target plate 5 in the camera coordinate system.

[0064] like Figure 9 As shown, the characteristic points of the Apriltag target are the four corners of the pattern border. When extracting target points, the corresponding standard target template can be found based on the tracked Apriltag target pattern ID. Based on the tracked image border, the local image can be extracted from the overall image. Corner points in the local image are then extracted and matched with the characteristic points of the standard target template image using Harris corner detection and Fast corner detection. A rotation and translation transformation strategy is then derived from the Apriltag target local image to the standard target template image. The corners of the Apriltag target are then located based on the corners of the standard target template image. Sub-pixel corner location is then used to precisely locate the Apriltag target corners.

[0065] After obtaining the alignment between the Apriltag target and the standard target template, the one-to-one correspondence between the pixel coordinates of the target points in the Apriltag target and the world coordinates of the standard target template can be obtained based on the pre-designed world coordinates of the standard target template points. The relationship is then brought into the PNP equation to solve the corresponding target plate 5 pose.

[0066] For the pose estimation of the monocular camera visual sensor 1, the PNP equation used is as follows:

[0067]

[0068] Where [u,v] is the pixel coordinate of the Apriltag target, K is the camera intrinsic parameter matrix, [R|t] is the desired pose matrix (extrinsic parameter matrix), and [x,y,z] is the world coordinate of the standard target template.

[0069] The target pose estimation function is used for the pose estimation of the binocular stereo camera vision sensor 1 or the structured light camera vision sensor 1. When the binocular stereo camera vision sensor 1 is used, when the pose of the same target is calculated in both camera fields of view at the same time, the three-dimensional coordinates of the four vertices of the target plate 5 can be solved according to the reprojection principle, through the reprojection equation, and using the coordinate system conversion matrix calibrated between the cameras, as shown in formulas (2) and (3):

[0070]

[0071] Among them, (u1, v1) and (u2, v2) are the pixel coordinates of the corresponding points in camera 1 and camera 2 of the binocular stereo camera, K1 and K2 are the intrinsic parameter matrices of camera 1 and camera 2 respectively, [R1|t1] and [R2|t2] are the extrinsic parameter matrices of camera 1 and camera 2, (x w ,y w , z w ) is the three-dimensional coordinate. Usually the coordinate system of camera 1 is used as the reference system, then [R1|t1] is the unit matrix, which is the external parameter matrix of camera 2 during camera calibration (i.e. the rotation and translation matrix relative to camera 1). By combining the two equations [R2|t2], we can solve (x w ,y w , z w ) three-dimensional coordinates.

[0072] When using a structured light camera visual sensor 1, the three-dimensional coordinates of the four vertices can be directly obtained through the correspondence between the image and the point cloud. After solving the three-dimensional coordinates of the target point in the camera coordinate system, the three-dimensional coordinates of the target point in the local coordinate system with the target as the target can be obtained based on the relative relationship between the target points (x o ,y o , z0), the target’s position and posture can be calculated by solving the equation [R c |t c ], as shown in formula (4):

[0073]

[0074] Among them, (x w ,yw , z w ) is the three-dimensional coordinate of the camera coordinate system, (x o ,y o , z0) is the three-dimensional coordinate of the target point in the local coordinate system with the target, [R c |t c ] is the pose of the target to be solved.

[0075] After the pose estimation of each target is completed, the pose solution function of target plate 5 is based on the calibration matrix of target plate 5. The calibration matrix of target plate 5 is a matrix that describes the relative position and direction between multiple targets. Through one Apriltag target, the position of other Apriltag targets can be determined through the calibration matrix of target plate 5, thereby determining the pose of target plate 5. At the same time, since the pose calculation of target plate 5 only requires one target pose result, and the six target patterns on target plate 5 often detect more than one target in actual applications, the pose results of multiple targets are further used to reduce the error and uncertainty caused by a single target, making the final target plate 5 estimation result more accurate and reliable. Assume that the target set captured and recognized by the camera is The target number is composed of any one or more of 0-5. According to the calibration matrix of target plate 5, the pose result of target plate 5 is obtained by averaging. Tag , as shown in formula (5):

[0076]

[0077] Among them, C Tag is the pose result of target plate 5, is the target set, C tagi is the pose of the i-th target, T tagi_T is the coordinate system transformation matrix of the pre-calibrated target i relative to the center of the target plate 5, For collection The number of elements. Usually T tagi_T *C tagi The pose matrix is ​​converted into rotation and translation vectors, averaged, and then converted back into a pose matrix.

[0078] The present invention proposes a target plate posture tracking method based on deep learning, which includes:

[0079] Step 1: Install the device and calibrate sensor 1;

[0080] Step 1.1: Install the device. Select one of the three visual sensors 1: a monocular camera, a binocular stereo camera, or a structured light camera. For underwater scenes, select the corresponding underwater sensor 1 part. Adjust the light source to ensure that the sensor 1 is well illuminated. Assemble the sensor 1, image processor, and display 3. Ensure that the data cable of the sensor 1 is connected to the image processing industrial computer 2, and connect the device to a power source.

[0081] Step 1.2: Calibrate sensor 1. Before using sensor 1, the visual sensor 1 is calibrated using a calibration plate with a special geometric pattern. This calibrates the camera's internal imaging geometry, facilitating the calculation of the target's pose based on the tracked target feature points. For the monocular visual sensor 1, a checkerboard calibration method is used to obtain the camera's intrinsic parameter matrix and distortion coefficients to complete the calibration. For the binocular stereo visual sensor 1, a checkerboard calibration plate is also used for calibration. First, the single camera calibration of the two cameras is calibrated to obtain the intrinsic parameters and distortion coefficients of the two cameras. Then, based on the intersection of the binocular cameras, a collinearity equation is constructed, and the extrinsic parameter matrix between the cameras is solved to obtain the rotation and translation parameters between the cameras. For the structured light camera visual sensor 1, the coordinate system transformation relationship between the calibrated camera and laser is used to obtain the correspondence between the point cloud and the image, and the correspondence between the point cloud and the image pixels. When using an underwater camera to track the 5-pose of an underwater target, the above parameters need to be calibrated in an underwater environment. During calibration, the calibrated observation distance must be consistent with the observation distance range in actual application to ensure the accuracy of the calibration results in real scenes.

[0082] Step 2: Calibrate the target plate 5. Use a high-precision structured light 3D camera to collect calibration data of the target plate 5 and calibrate the conversion relationship between the coordinate system of each target surface of the target plate 5 and the overall coordinate system of the target plate 5. The position and posture of the target plate 5 are obtained by detecting the position and posture of the Apriltag targets outside the target plate 5. Only the position and posture of one target needs to be detected. After the coordinate system conversion, it can be converted into the position and posture of the center coordinate system of the target plate 5. Establish a local coordinate system at the center position of each of the six Apriltag targets. Assume that the local coordinate systems of the six targets from Apriltag target 1 to Apriltag target 6 are C tag0 、C tag1 、C tag2 、C tag3 、C tag4 、C tag5 , the target plate 5 center coordinate system is C Tag , the central coordinate system is defined as the origin at the center of the upper surface of the target plate 5, and the coordinate axis of Tag is consistent with the coordinate axis of tag0, such as Figure 5 As shown, where C Tag The X and Y directions are defined with C tag0The X and Y coordinates of the target are consistent and in the same plane. Therefore, it is necessary to calibrate the coordinate system of each target and C Tag The coordinate system transformation matrix is ​​shown in formula (6):

[0083]

[0084] Among them, C tag0 、C tag1 、C tag2 、C tag3 、C tag4 、C tag5 The local coordinate systems C are established for the center positions of the six Apriltag targets. Tag is the target plate 5 center coordinate system, T tag0_Tag is the transformation matrix from the local coordinate system of the tag0 target to the center coordinate system of the tag, T tag1_Tag is the transformation matrix from the local coordinate system of the tag1 target to the center coordinate system of the tag, T tag2_Tag is the transformation matrix from the local coordinate system of the tag2 target to the center coordinate system of the tag, T tag3_Tag is the transformation matrix from the local coordinate system of the tag3 target to the Tag center coordinate system, T tag4_Tag is the transformation matrix T from the local coordinate system of the tag4 target to the Tag center coordinate system tag5_Tag is the transformation matrix from the local coordinate system of target tag 0 to the center coordinate system of tag 5. Target plate 5 is calibrated using a structured light 3D camera, ensuring an accuracy range of 0.17-0.34mm in the X and Y directions and 0.047-0.14mm in the Z direction. A color image, depth map, and point cloud data of the captured scene are simultaneously obtained, with one-to-one correspondence between the three data points.

[0085] Step 2.1: Place the target plate 5 at 400mm, 500mm, 600mm, and 700mm from the structured light 3D camera, facing the structured light 3D camera to ensure that all target patterns are within the field of view and acquire images;

[0086] Step 2.2: The camera structured light 3D camera extracts each Apriltag target of the three target plates 5 and calculates the pose of each Apriltag target;

[0087] Step 2.3: For each target plate 5, extract the four corner points of one Apriltag target on the target plate 5 and obtain the coordinates of the center point of the target plate 5, calculate the translation vector of the Apriltag target relative to the center of the target plate 5, and obtain the transformation matrix of the Apriltag target relative to the center of the target plate 5;

[0088] Step 2.4: Calculate the transformation matrices of the remaining five Apriltag target patterns relative to the Apriltag target, and convert the transformation matrix between the Apriltag target and the center of the target plate 5 into the transformation matrix between the remaining five Apriltag targets and the center of the target plate 5 to complete the calibration of each target plate 5;

[0089] Step 2.5: The calibration results at different positions output a set of rotation angle components and translation components obtained by decomposition of the pose matrix (each set of calibration results contains the rotation and translation components of 6 target patterns relative to the center of the target plate 5). The results of multiple sets of angle components and translation components are averaged to obtain the final calibration result, that is, the rotation and translation components of the 6 target patterns relative to the center of the target plate 5.

[0090] Testing has shown that the calibration accuracy of the pose conversion matrix of the six target patterns relative to the center of the target plate 5 is guaranteed to be 0.5mm. This means that the difference between the measured poses of the six targets, converted to the pose of the center of the target plate 5 using the calibration matrix, and the actual pose of the center of the target plate 5 is less than 0.5mm, making it suitable for pose positioning in simulated scenarios. The pose tracking of a single target achieves a 97% detection success rate during the motion of the robotic arm 4, assuming no occlusion and minimal view angle changes, and a pose estimation accuracy better than 0.7mm.

[0091] Step 3: Install the target plate 5 to the target to be measured. Take the end of the robot arm 4 as an example. Install the target plate 5 to the end of the robot arm 4 and calibrate the conversion relationship between the tool coordinate system to be measured and the target plate 5 coordinate system. This can be measured based on the installation position relationship. If the installation is coaxial, only the height measurement is needed to obtain the conversion relationship.

[0092] Step 4: Place sensor 1 in the measurement scene, ensuring that its field of view covers the range of the target's motion. Start running the target tracking and pose estimation modules in the industrial computer 2's software system. This system captures images of the target plate 5 as the target moves, outputs the pose of the target plate 5 in real time, and converts it into the pose of the target's coordinate system based on the transformation relationship.

[0093] The above embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A target plate posture tracking device based on deep learning, characterized in that: The device comprises a sensor (1), a target plate (5), an image processor (2) and a display (3); the target plate (5) is arranged at the end of a robotic arm; the sensor (1) is used to collect image data of the target plate (5) and the robotic arm; the sensor (1) is connected to the image processor (2) via a data line; the image processor (2) is connected to the display (3); and the display (3) displays the posture of the robotic arm (4).

2. A target plate posture tracking device based on deep learning according to claim 1, characterized in that: The sensor (1) is a visual sensor, and the sensor (1) is a monocular camera visual sensor, a binocular stereo camera visual sensor, or a structured light camera visual sensor.

3. A target plate posture tracking device based on deep learning according to claim 2, characterized in that: The sensor (1) is an underwater visual sensor, and the sensor (1) is equipped with a matching light source.

4. A target plate posture tracking device based on deep learning according to claim 3, characterized in that: The target plate (5) is a hexagonal snowflake-shaped structure, and six different Apriltag targets are respectively arranged at the six corners of the snowflake-shaped structure of the target plate (5), and the six Apriltag targets are located in the same plane.

5. The target plate posture tracking device based on deep learning according to claim 4, characterized in that: The image processor (2) is an industrial computer, and the industrial computer integrates a software system. The software system mainly includes an image processing module, a manual interaction module and a guidance control module. The image processing module mainly includes a target posture estimation unit, an air lock hole target posture estimation unit and a target recognition and posture estimation unit. The manual interaction module transmits the image measured by the sensor (1) back to the display (3) for display, and annotates and displays the operation target and the end of the robot arm (4) in the image. The guidance control module obtains real-time robot arm (4) control posture parameters based on the robot arm (4) posture and target posture output by the image processing module and the difference between the robot arm (4) posture and the target posture. The guidance control module provides the robot arm control posture parameters to the robot arm (4) controller to guide the end control of the operation robot arm (4).

6. A target plate posture tracking device based on deep learning according to claim 5, characterized in that: The target pose estimation unit obtains the target pose of the end of the manipulator (4) based on deep learning, thereby obtaining the end pose of the manipulator (4); the air-lock hole target pose estimation unit obtains the air-lock pose based on deep learning; and the target recognition and pose estimation module realizes the recognition and pose determination of the working target of the manipulator (4) based on binocular stereo vision.

7. A target plate posture tracking device based on deep learning according to claim 6, characterized in that: The target pose estimation unit includes a target tracking function, a target pose estimation function and a target plate (5) pose solution function. The target tracking function uses a yoloV8 network to perform target position detection and a DeepSort tracking module to perform target position tracking. The yoloV8 network uses a simulated data set to first train a pre-training model, and then collects part of the data as real data in an artificial simulation environment to fine-tune the pre-training model; the DeepSort tracking module includes Kalman prediction, deep feature matching and Hungarian algorithm; the target pose estimation unit uses the output vector of the yoloV8 network as the deep feature of the target pattern, uses the cosine function to calculate the similarity of the deep feature vector, combines the state information of the Kalman prediction, and performs matching association of the detection targets of the previous and next consecutive frames according to the Hungarian algorithm to achieve correct target detection output.

8. The target plate posture tracking device based on deep learning according to claim 7, characterized in that: After the target plate (5) is tracked and identified, the target pose estimation function of the target pose estimation unit detects and extracts the target point of the Apriltag target for the monocular camera visual sensor (1), obtains the pixel coordinates of the target point, matches the corner points of the Apriltag target with the corner points of the standard target template, obtains the rotation and translation transformation strategy from the Apriltag target image to the standard target template image, obtains the corresponding relationship between the pixel coordinates of the target point in the Apriltag target and the world coordinates of the standard target template, and uses the PNP algorithm to calculate the pose of the center of the target plate (5) in the camera coordinate system. The PNP algorithm equation is as shown in formula (1): Where [u,v] is the pixel coordinate of the Apriltag target, K is the camera intrinsic parameter matrix, [R|t] is the required pose matrix, and [x,y,z] is the world coordinate of the standard target template; For the binocular stereo camera vision sensor (1), the pose of the same Apriltag target is calculated in the field of view of two cameras at the same time. The three-dimensional coordinates of the four vertices of the target plate (5) are solved using the coordinate system conversion matrix calibrated between the cameras. The three-dimensional coordinate calculation equations are as follows: Among them, (u1, v1) and (u2, v2) are the pixel coordinates of the corresponding points in camera 1 and camera 2 of the binocular stereo camera, K1 and K2 are the intrinsic parameter matrices of camera 1 and camera 2 respectively, [R1|t1] and [R2|t2] are the extrinsic parameter matrices of camera 1 and camera 2, (x w ,y w , z w ) are three-dimensional coordinates; For the structured light camera vision sensor (1), the three-dimensional coordinates of the four vertices are obtained through the correspondence between the image and the point cloud; after the three-dimensional coordinates of the target point in the camera coordinate system are calculated, the three-dimensional coordinates of the target point in the local coordinate system with the target as the target are obtained based on the relative relationship between the target points, and the target pose is calculated by solving the equation, as shown in formula (4): Among them, (x w ,y w , z w ) is the three-dimensional coordinate of the camera coordinate system, (x o ,y o , z0) is the three-dimensional coordinate of the target point in the local coordinate system with the target, [R c |t c ] is the pose of the target to be solved.

9. The target plate posture tracking device based on deep learning according to claim 8, characterized in that: The target plate (5) pose solving function of the target pose estimation unit determines the pose of the target plate (5) according to the target plate (5) calibration matrix after the Apriltag target pose estimation is completed; for multiple Apriltag target pose results, the pose result of the target plate (5) is obtained by averaging as C Tag , as shown in formula (5): Among them, C Tag is the pose result of the target plate (5), is the target set, C tagi is the pose of the i-th target, T tagi_T is the coordinate system transformation matrix of the pre-calibrated target i relative to the center of the target plate (5), For collection The number of elements.

10. A target plate posture tracking device based on deep learning according to any one of claims 1 to 9, characterized in that: The method comprises: Step 1: Assemble the sensor (1), image processor (2) and display (3), ensure that the data line of the sensor (1) is connected to the image processing industrial computer, connect the power supply of the device, and calibrate the sensor (1). Step 2: using a structured light 3D camera to collect calibration data of the target plate (5), and calibrate the conversion relationship between the coordinate system of each target surface of the target plate (5) and the overall coordinate system of the target plate (5); Step 3: Install the target plate (5) to the target to be measured, and calibrate the conversion relationship between the tool coordinate system to be measured and the target plate (5) coordinate system. Step 4: Arrange the sensor (1) in the measurement scene to ensure that the field of view covers the range of movement of the target to be measured, run the software system of the industrial computer, collect the image of the target plate (5) when the target to be measured moves in real time, output the position and posture of the target plate (5) in real time, and convert it into the position and posture of the target to be measured based on the conversion relationship.

11. A target plate posture tracking device based on deep learning according to claim 10, characterized in that: The second step specifically includes: Step 2.1: Place the target plate (5) facing the structured light 3D camera at a distance of 400 mm, 500 mm, 600 mm, and 700 mm from the structured light 3D camera respectively; Step 2.2: The camera structured light 3D camera extracts each Apriltag target of the three target plates (5) and calculates the pose of each Apriltag target; Step 2.3: For each target plate (5), extract the four corner points of one Apriltag target on the target plate (5) and obtain the coordinates of the center point of the target plate (5), calculate the translation vector of the Apriltag target relative to the center of the target plate (5), and obtain the transformation matrix of the Apriltag target relative to the center of the target plate (5); Step 2.4: Calculate the conversion matrices of the remaining five Apriltag target patterns relative to the Apriltag target, and convert the conversion matrix between the Apriltag target and the center of the target plate (5) into the conversion matrix between the remaining five Apriltag targets and the center of the target plate (5), thereby completing the calibration of each target plate (5); Step 2.5: The calibration results at different positions output a set of rotation angle components and translation components obtained by decomposition of the pose matrix. The final calibration result is obtained by averaging the results of multiple groups of angle components and translation components.

Citation Information

Patent Citations

  • Mechanical arm dynamic grabbing method based on visual servo

    CN117428785A

  • Device, system and method for visual guidance of block board robot

    CN119839856A

Cited By

  • Intelligent remote sensing target cooperative control system and method based on unmanned aerial vehicle

    CN120848357A

  • A collaborative control system and method for intelligent remote sensing targets based on unmanned aerial vehicles (UAVs)

    CN120848357B