A method, system, device and medium for estimating target object posture

By establishing a keyframe offline data set and performing online multi-frame matching, the robustness and real-time problems of 6-Dof pose estimation in an open environment are solved, and fast and real-time multi-objective pose estimation is achieved, avoiding the cumbersome process of data set annotation.

CN119600088BActive Publication Date: 2025-05-16HUNAN INSTITUTE OF ENGINEERING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411636615.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-05-16
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

The prior art is difficult to achieve robust 6-Dof pose estimation in an open environment, especially in the case of occlusion and lighting changes, and the depth measurement accuracy is limited and cannot be operated in real time. The deep learning method requires a large number of manual labeling data sets, and the scalability is limited.

Method used

By obtaining the position pose, effective point cloud and target effective texture of the target object in all keyframes, establishing a keyframe offline data set, and real-time target pose estimation by online multi-frame matching, a large number of dataset annotation and feature matching processes are avoided.

Benefits of technology

It realizes fast and real-time multi-objective 6-Dof pose estimation in an open environment, avoids the tedious process of data set annotation, breaks through the limitations of identification types, and significantly improves the matching speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600088B_ABST
    Figure CN119600088B_ABST
Patent Text Reader

Abstract

The invention discloses a method, system, device and medium for estimating the pose of a target object, and relates to the technical field of pose estimation. The method comprises: establishing a data set offline and an online real-time multi-target pose estimation process; in the offline process, continuous RGBD streams surrounding a single object are collected, the relative poses of previous and next frames are calculated frame by frame, and the three-dimensional surface of the target object is reconstructed to obtain a key frame offline data set; in the online process, a visible target set in the data stream is first detected, each target is segmented by VOS to obtain real-time effective texture and point cloud, offline key frame images of the target to be detected are arranged to achieve a single match of all feature points and find the closest reference frame at the same time, and the relative pose of the current frame target and the reference frame is calculated by using RANSAC; the method avoids the expensive data annotation process, realizes the rapid integrated deployment of the detection of new target objects that have not been seen, improves the matching speed, and can realize online real-time multi-target pose estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of posture estimation, and in particular to a method, system, device and medium for estimating the posture of a target object. Background Art

[0002] Accurate and stable 6-Dof pose estimation technology is indispensable for the rapidly developing embodied intelligent robot applications, autonomous driving perception, and assembly and unloading operations in industrial production environments. However, the actual operating environment of the intelligent body is relatively open and susceptible to external interference, which has a great impact on the adaptability of the 6-Dof information estimation algorithm.

[0003] From the perspective of the operation process, the current robust six-dimensional position and posture information extraction algorithm generally adopts the method of first obtaining the CAD three-dimensional model of the target to be inspected and then using the PnP method to solve the pose, or using the SLAM method to establish a reference key frame to store the prior pose information in advance. In an open environment, traditional feature point matching or PnP methods are difficult to cope with occlusion and lighting texture changes. Moreover, the feature point extraction and matching of existing algorithms are generally limited to a single object, and multiple key frames need to be matched, resulting in the inability to operate in real time. From the perspective of data sources, due to the cost and the frequency of multimodal data updates, the perception part of the intelligent body's target pose estimation generally uses a low-cost RGBD camera, whose depth measurement accuracy is very limited, usually at a relative accuracy of 1% to 2%, and it is difficult to guarantee the accuracy of directly using point cloud registration (point cloud extraction of key points to establish FPFH feature description).

[0004] The current deep learning-based approach generally first determines the effective point cloud position mask of the target object to be inspected on the texture map through a deep segmentation network, and then extracts the point cloud feature representation through representative networks such as PointNet++, and inputs the obtained fixed-dimensional point cloud feature vector into the MLP (multi-layer perceptron) to regress the target's 6-Dof information. However, whether it is RGB-D multi-channel information fusion or pure point cloud input, this deep learning approach requires manual annotation of a large number of pose data sets, and each category requires the collection of a three-dimensional point cloud training data set. It can only classify and identify objects that already exist in the data set, and its scalability is very limited. Summary of the invention

[0005] In view of the shortcomings of the prior art, such as difficulty in data set annotation, weak adaptability of feature matching, and failure to meet the requirements of multi-target real-time performance, the present invention proposes a method, system, device and medium for target object pose estimation. By acquiring the position and posture, valid point cloud, and target effective texture of the target object in all key frames, a key frame offline data set of the target object is established, and online multi-frame matching is performed on the key frame offline data set to estimate the real-time target object pose, thereby solving the problems existing in the prior art.

[0006] A method for estimating a target object's position and posture, comprising the following steps:

[0007] Establish an offline key frame dataset of the target to be detected;

[0008] Real-time target object pose estimation through online multi-frame matching through key frame offline dataset;

[0009] The step of establishing an offline key frame dataset of multiple targets to be detected specifically includes the following steps:

[0010] Collect multiple texture images and depth images surrounding a single target object, extract the target mask from the multiple texture images, extract the effective depth and effective texture of the target object according to the target mask; and convert the effective depth into an effective point cloud;

[0011] Detect the coarse matching feature points in the effective texture of two consecutive frames of images, and determine the three-dimensional matching feature point pairs with effective depth according to the pixel positions of the coarse matching feature points on the original image and the effective point cloud; randomly sample the three-dimensional matching feature point pairs to obtain effective matching three-dimensional feature point pairs, and decompose the effective matching three-dimensional feature point pairs to obtain the relative pose T of the two consecutive frames of images. ij , and take the first frame posture T0 as the initial frame, and determine the posture T of each subsequent frame relative to the initial frame i ; Calculate the pose T of the current frame i The rotation matrix R in i With each key frame rotation matrix R saved k The distance between the rotational postures; when R i With all R k If the distance value is greater than the threshold, it is a new key frame, otherwise continue to process the next frame until all images are processed;

[0012] According to the pose, valid point cloud and valid texture of the target object in all key frames, an offline key frame dataset of the target object is established.

[0013] Furthermore, the step of extracting target masks from the plurality of texture images specifically comprises the following steps:

[0014] Manually mark the mask where the target object is located in the first texture image;

[0015] The transformerVOS target segmentation network based on position memory is used to extract the target mask in the subsequent continuous texture map, and then the mask is eroded.

[0016] Furthermore, the effective depth is converted into a valid point cloud using camera intrinsic parameters.

[0017] Furthermore, the method also includes filtering out point cloud noise according to its spatial point cloud and point cloud normal vector by calculating the point cloud normal vector, and updating the mask information of the texture image using the point cloud from which the noise is filtered out.

[0018] Furthermore, the detecting of coarse matching feature points in effective textures of two consecutive frames of images specifically comprises the following steps:

[0019] By scaling the effective texture of the target object to a pixel size of 512*512;

[0020] The Efficient LOFTR model is used to detect the coarse matching feature points in the effective texture after scaling of the previous and next frames.

[0021] Furthermore, the RANSAC algorithm is used to randomly sample the three-dimensional matching feature point pairs to obtain effective matching three-dimensional feature point pairs.

[0022] Furthermore, online multi-frame matching real-time target pose estimation is performed through the key frame offline data set, which specifically includes the following steps:

[0023] The valid texture images of all key frames of the target object are stitched into a reference frame, and the SupperPoint feature points and descriptors are extracted from the image;

[0024] Drive the camera to obtain real-time RGBD data stream, and perform YOLO target detection on the texture image to determine the target and its bounding box position information appearing in the camera's field of view;

[0025] The number of threads is determined according to the number of targets in the camera's field of view, and each thread processes one target. In one thread, the precise manually annotated mask of the target object in the initial frame is retrieved, and the rough mask of the target object in the current real-time frame is extracted using the VOS model in combination with the bounding box position information. The rough mask is filtered to extract the effective texture and depth point cloud of the target object.

[0026] Scale the current target effective texture to a pixel size of 512*512, extract the SupperPoint feature points and descriptors of the scaled effective texture; match the current target feature points and descriptors with the SupperPoint feature points and descriptors of the entire reference frame using a lightweight LightGlue model to obtain a coarse matching feature point pair;

[0027] The RANSAC algorithm is used to randomly sample the coarse matching feature point pairs to obtain effective matching three-dimensional feature point pairs;

[0028] The relative pose Tt between the current frame and the selected key frame is obtained by using SVD decomposition for the effective matching 3D feature point pairs;

[0029] The relative posture T t Offline pose T with keyframe k Combined with the first frame pose T0 as the reference, the final 6Dof is: T = T t *T k *T0.

[0030] The present invention also includes a target object position and posture estimation system, comprising:

[0031] An offline data set building module is used to build an offline data set of key frames of the target to be detected;

[0032] The target object pose estimation module is used to perform real-time target pose estimation through online multi-frame matching using key frame offline data sets;

[0033] The offline data set establishment module specifically includes:

[0034] An acquisition unit is used to acquire multiple texture images and depth images surrounding a single target object, extract a target mask from the multiple texture images, extract an effective depth and effective texture of the target object according to the target mask, and convert the effective depth into an effective point cloud;

[0035] The pose calculation unit is used to detect the coarse matching feature points in the effective texture of two consecutive frames of images, determine the three-dimensional matching feature point pairs with effective depth according to the pixel positions of the coarse matching feature points on the original image and the effective point cloud; randomly sample the three-dimensional matching feature point pairs to obtain effective matching three-dimensional feature point pairs, and decompose the effective matching three-dimensional feature point pairs to obtain the relative pose T of the two consecutive frames of images. ij , and take the first frame posture T0 as the initial frame, and determine the posture T of each subsequent frame relative to the initial frame i ; Calculate the pose T of the current frame i The rotation matrix R in i With each key frame rotation matrix R saved k The distance between the rotational postures; when R i With all R k If the distance value is greater than the threshold, it is a new key frame, otherwise continue to process the next frame until all images are processed;

[0036] The data set establishment unit is used to establish a key frame offline data set of the target object according to the position, valid point cloud and target valid texture of the target object in all key frames obtained.

[0037] The present invention also includes a computer device for estimating the position and posture of a target object, comprising: a memory, a processor, and a computer program stored in the memory, wherein the processor implements the steps of the method for estimating the position and posture of a target object when executing the computer program.

[0038] The present invention also includes a readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the target object pose estimation method.

[0039] The present invention provides a method, system, device and medium for estimating the position and posture of a target object, which has the following features:

[0040] Beneficial effects:

[0041] The present invention establishes a key frame offline data set of the target to be detected, determines three-dimensional matching feature points with valid depth values ​​at the pixel positions of the original image and the visible valid point cloud according to the coarse matching feature points in the effective texture of two consecutive frames of images, and is used to remove matching points whose depth values ​​may be invalid at the corresponding positions of the feature points on the texture image; randomly samples the three-dimensional matching feature point pairs, further removes the mismatched feature point pairs, and obtains valid matching three-dimensional feature point pairs; combines the relative posture between the current frame and the selected key frame with the offline posture of the key frame, takes the posture of the first frame as a reference, estimates the posture of the target object, and establishes a key frame offline data set, thereby avoiding the process of labeling and learning a large number of data sets, breaking through the recognition type limitation, greatly improving the matching speed, and facilitating the subsequent online multi-frame matching of the real-time target object posture estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a schematic diagram of the target object pose estimation process in an embodiment of the present invention;

[0043] Figure 2 Schematic diagram of a 3D reconstruction model of three objects, a toy car, a socket and a milk bottle, in an embodiment of the present invention;

[0044] Figure 3 A schematic diagram of real-time online matching in an embodiment of the present invention;

[0045] Figure 4 A multi-target segmentation map in an embodiment of the present invention;

[0046] Figure 5 Schematic diagram of real-time online matching of multiple targets in an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0048] The present invention proposes a fast multi-frame matching real-time multi-target 6-Dof information estimation technology based on an offline database of target key frames in an open environment, which is divided into an offline process and an online process:

[0049] The offline process is used to establish the key frame offline database; it specifically includes the following steps:

[0050] S1. First, use a depth camera to collect clear texture images and depth images surrounding a single object.

[0051] S2. Manually mark the mask of the target object in the first texture image and store it locally to obtain a key frame. Use the transformerVOS target segmentation network based on position memory to extract the target mask in the subsequent continuous texture image, and erode the mask to eliminate the uncertain part of the segmentation edge.

[0052] S3. Extract the effective depth of the target location according to the target mask in the continuous texture map, convert the effective depth into a structured point cloud using the camera internal parameters, calculate the point cloud normal, filter out the point cloud noise according to the spatial point cloud and its normal information, use the point cloud with the noise filtered out to update the texture map mask information, and store the visible structured effective point cloud and mask of the target object.

[0053] S4. Use the mask to extract the effective texture of the target object, scale the effective texture to a pixel size of 512*512, and use the scaled effective texture to calculate the relative pose between the previous and next frames.

[0054] S4.1. Use the Efficient LOFTR model to directly detect the coarse matching feature points in the effective texture after scaling the previous and next frames.

[0055] S4.2. Determine the three-dimensional matching feature points with valid depth values ​​according to the pixel positions of the coarse matching feature points in the original image and the visible valid point cloud (the corresponding positions of the feature points on the texture image may have invalid depth values, so the matching point pairs are eliminated).

[0056] S4.3. Perform a random sampling consistency RANSAC operation on the three-dimensional matching feature point pairs to further remove the mismatched feature point pairs and obtain valid matching three-dimensional feature point pairs.

[0057] S4.4. Effectively match the 3D feature point pairs using SVD decomposition to obtain the relative pose T of the previous and next frames ij , taking the first frame pose T0 as the reference, calculate the pose T of each subsequent frame relative to the initial frame i .

[0058] S5, key frame extraction: calculate the pose T of the current frame i The rotation matrix R in iWith each key frame rotation matrix R saved k The distance d between the rotational postures. i With all R k If the distance value is greater than the threshold, it is a new key frame and is saved. Otherwise, the next frame is processed until all images are processed. In practice, this value is taken as 0.2 to limit the number of key frames while ensuring the matching effect.

[0059]

[0060] Where: R i is the pose T of the current frame i The rotation matrix in R k is the rotation matrix of the saved keyframe.

[0061] S6. Local 3D reconstruction based on Nerf: The position and posture of the target object, the valid point cloud, and the target effective texture in all key frames are input into the neural radiation field Nerf for training. The TSDF (truncated signed distance function) is used to eliminate the uncertainty of the 3D surface of the target object obtained from the key frame information at various angles. Since it is a continuous image of the surround view, a high-quality 3D reconstruction model of the target object can be obtained.

[0062] S7. Use this 3D reconstructed model and all the existing pose information of the key frames to back-project to the camera imaging plane to obtain the overall reprojection error, and use Bundle Adjustment (BA, bundle adjustment method) to correct the existing pose information of all the key frames at one time; store the updated key frame pose information, target object effective texture, structured effective point cloud information and initial frame mask locally.

[0063] S8. For multiple targets, the above-mentioned offline operation process is used to establish corresponding key frame pose information, target object effective texture, structured effective point cloud information and initial frame mask, and store and establish a key frame database of multiple targets to be inspected.

[0064] The online process is used for online multi-frame matching and real-time multi-target pose estimation; specifically, it includes the following steps:

[0065] S1. Initialization of online detection algorithm: load the YOLO target detection model weights, VOS target segmentation model weights, SupperPoint feature point extraction model weights, LightGlue feature point matching model weights, load the key frame database of multiple targets to be detected into memory, stitch all key frame valid texture images of a single target into a reference frame, extract SupperPoint feature points and descriptors from the image and save them.

[0066] S2: Drive the camera to obtain real-time RGBD data stream and perform YOLO target detection on the texture map to determine the visible targets and their bounding box position information that appear in the camera's field of view.

[0067] S3. Create a corresponding number of threads according to the number of visible targets in the camera's field of view, and process one target in each thread. In one thread, retrieve the precise manually annotated mask of the target's initial frame, combine it with the bounding box position information, use the VOS model to extract the rough mask of the target in the current real-time frame, and extract the target's effective texture and depth point cloud through filtering.

[0068] S4. Scale the effective texture of the current target extracted in the previous step to a size of 512*512, extract the SupperPoint feature points and descriptors; match the feature points and descriptors with the SupperPoint feature points and descriptors of the entire reference frame of the current target using a lightweight LightGlue model to obtain a coarse matching feature point pair.

[0069] S5. Check the validity of the coarse matching feature point pair: use the feature point in the original image pixel position as an index to find out whether there is a corresponding visible valid depth. If the depth value of the corresponding position of the feature point on the texture image is invalid, the matching point pair is eliminated.

[0070] S6, performing a random sampling consistency RANSAC operation on the matching feature point pairs with three-dimensional coordinates xyz values ​​in the camera coordinate system obtained in the previous step to remove mismatched feature point pairs, and obtain valid matching three-dimensional feature point pairs;

[0071] S7. Use SVD decomposition to obtain the relative pose T for the effective matching 3D feature point pairs. t , which is the relative pose between the current frame and the selected key frame, combined with the offline pose T of the key frame k , taking the first frame pose T0 as the reference, the final 6Dof is calculated as: T = T t *T k *T0;

[0072] The present invention performs an erosion operation on the mask to eliminate the uncertain part of the segmentation edge; filters out point cloud noise according to the spatial point cloud and its normal vector information, and uses the point cloud from which the noise is filtered out to update the texture map mask information; determines three-dimensional matching feature points with valid depth values ​​according to the pixel positions of the coarse matching feature points in the original image and the visible valid point cloud, and removes matching points whose depth values ​​may be invalid at the corresponding positions of the feature points on the texture image; inputs the position and posture of the target object in all key frames, the valid point cloud, and the target valid texture into the neural radiation field Nerf for training, and uses the truncated signed distance function TSDF to eliminate the uncertainty of the three-dimensional surface of the target object obtained from the key frame information at various angles, thereby obtaining a high-quality three-dimensional reconstruction model of the target object; combines the relative posture between the current frame and the selected key frame with the offline posture of the key frame, and takes the posture of the first frame as a reference to estimate the posture of the target object; the method avoids the process of labeling and learning a large number of data sets, breaks through the limitation of recognition types, and greatly improves the matching speed.

[0073] In actual operation, since there is a fixed offset between the actual grasping posture and the target model space coordinates, T0 can be set to the actual offset of the target grasp, which is often obtained by manual measurement (the rotation part of T0 is set to the unit matrix, and the deviation between the grasping posture after recognition and the ideal grasping posture is manually measured).

[0074] Experimental proof:

[0075] (1) Offline process: key frame offline database is established.

[0076] For the toy car, socket and milk bottle, three targets are respectively surrounded by depth cameras to obtain clear texture maps and depth maps of each object. The key frames with the highest degree of distinction are stored in the reference frame offline data set, and then the 3D reconstruction model of the target is obtained based on the local 3D reconstruction of Nerf, such as Figure 2 As shown, (a) is the 3D reconstruction model of a toy car; (b) is the 3D reconstruction model of a socket; and (c) is the 3D reconstruction model of a milk bottle.

[0077] (2) Online process: online multi-frame matching for real-time multi-target pose estimation.

[0078] like Figure 3 The milk bottle on the far left is a real-time effective texture map of the milk bottle obtained through Yolo target detection. The full image on the right is composed of the reference frames of nine milk bottles. After the lightweight LightGlue model operation and RANSAC operation, the effective matching 3D feature point pairs are obtained. Figure 4 is a segmentation map of multiple targets (toy car, socket and milk bottle). The real-time online matching of multiple targets is as follows: Figure 5 shown.

[0079] The invention has been proven to be feasible through experiments and simulations. The method avoids the process of labeling and learning a large number of data sets, breaks through the limitation of recognition types, greatly improves the matching speed, and the online process can complete the posture detection of multiple targets within 80ms.

[0080] Based on the same inventive concept, the present invention also proposes a target object pose estimation system, comprising:

[0081] An offline data set building module is used to build an offline data set of key frames of the target to be detected;

[0082] The target object pose estimation module is used to perform real-time target pose estimation through online multi-frame matching using key frame offline data sets;

[0083] The offline data set establishment module specifically includes:

[0084] An acquisition unit is used to acquire multiple texture images and depth images surrounding a single target object, extract a target mask from the multiple texture images, extract an effective depth and effective texture of the target object according to the target mask, and convert the effective depth into an effective point cloud;

[0085] The pose calculation unit is used to detect the coarse matching feature points in the effective texture of two consecutive frames of images, determine the three-dimensional matching feature point pairs with effective depth according to the pixel positions of the coarse matching feature points on the original image and the effective point cloud; randomly sample the three-dimensional matching feature point pairs to obtain effective matching three-dimensional feature point pairs, and decompose the effective matching three-dimensional feature point pairs to obtain the relative pose T of the two consecutive frames of images. ij , and take the first frame posture T0 as the initial frame, and determine the posture T of each subsequent frame relative to the initial frame i ; Calculate the pose T of the current frame i The rotation matrix R in i With each key frame rotation matrix R saved k The distance between the rotational postures; when R i With all R k If the distance value is greater than the threshold, it is a new key frame, otherwise continue to process the next frame until all images are processed;

[0086] The data set establishment unit is used to establish a key frame offline data set of the target object according to the position, valid point cloud and target valid texture of the target object in all key frames obtained.

[0087] Based on the same inventive concept, the present invention also proposes a computer device for estimating the position and posture of a target object, comprising: a memory, a processor, and a computer program stored in the memory, and when the processor executes the computer program, the steps of the method for estimating the position and posture of a target object are implemented.

[0088] Based on the same inventive concept, the present invention also proposes a readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the target object pose estimation method.

[0089] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for estimating a target object's position and posture, characterized in that: The following steps are involved: Establish an offline key frame dataset of the target to be detected; Real-time target object pose estimation through online multi-frame matching using key frame offline dataset; The step of establishing an offline key frame dataset of multiple targets to be detected specifically includes the following steps: Collect multiple texture images and depth images surrounding a single target object, extract the target mask from the multiple texture images, extract the effective depth and effective texture of the target object according to the target mask; and convert the effective depth into an effective point cloud; Detect the coarse matching feature points in the effective texture of two consecutive frames of images, and determine the three-dimensional matching feature point pairs with effective depth according to the pixel positions of the coarse matching feature points on the original image and the effective point cloud; randomly sample the three-dimensional matching feature point pairs to obtain effective matching three-dimensional feature point pairs, and decompose the effective matching three-dimensional feature point pairs to obtain the relative pose T of the two consecutive frames of images. ij , and take the first frame posture T0 as the initial frame, and determine the posture T of each subsequent frame relative to the initial frame i ; Calculate the pose T of the current frame i The rotation matrix R in i The saved keyframe rotation matrix R k The distance between the rotational postures; when R i With all R k If the distance value is greater than the threshold, it is a new key frame, otherwise continue to process the next frame until all images are processed; According to the pose, valid point cloud and valid texture of the target object in all key frames, an offline key frame dataset of the target object is established.

2. A method for estimating a target object's position and posture according to claim 1, characterized in that: The step of extracting target masks from multiple texture images specifically includes the following steps: Manually mark the mask where the target object is located in the first texture image; The transformerVOS object segmentation network based on position memory is used to extract the object mask in the subsequent continuous texture map, and the mask is eroded.

3. A method for estimating a target object's position and posture according to claim 1, characterized in that: The effective depth is converted into a valid point cloud using camera intrinsic parameters.

4. A method for estimating a target object's position and posture according to claim 1, characterized in that: It also includes converting the effective depth into a valid point cloud, calculating the point cloud normal vector, filtering out the point cloud noise according to its spatial point cloud and the point cloud normal vector, and updating the mask information of the texture image using the point cloud from which the noise is filtered out.

5. A method for estimating a target object's position and posture according to claim 1, characterized in that: The method of detecting coarse matching feature points in effective textures of two consecutive frames of images specifically comprises the following steps: By scaling the effective texture of the target object to a pixel size of 512*512; The Efficient LOFTR model is used to detect the coarse matching feature points in the effective texture after scaling of the previous and next frames.

6. A method for estimating a target object's position and posture according to claim 1, characterized in that: The RANSAC algorithm is used to randomly sample three-dimensional matching feature point pairs to obtain effective matching three-dimensional feature point pairs.

7. A method for estimating a target object's position and posture according to claim 1, characterized in that: The real-time target pose estimation is performed by online multi-frame matching through the key frame offline data set, which specifically includes the following steps: The valid texture images of all key frames of the target object are stitched into a reference frame, and the SupperPoint feature points and descriptors are extracted from the image; Drive the camera to obtain real-time RGBD data stream, and perform YOLO target detection on the texture image to determine the target and its bounding box position information appearing in the camera's field of view; The number of threads is determined according to the number of targets in the camera's field of view, and each thread processes one target. In one thread, the precise manually annotated mask of the target object in the initial frame is retrieved, and the rough mask of the target object in the current real-time frame is extracted using the VOS model in combination with the bounding box position information. The rough mask is filtered to extract the effective texture and depth point cloud of the target object. Scale the current target effective texture to a pixel size of 512*512, extract the SupperPoint feature points and descriptors of the scaled effective texture; match the current target feature points and descriptors with the SupperPoint feature points and descriptors of the entire reference frame using a lightweight LightGlue model to obtain a coarse matching feature point pair; The RANSAC algorithm is used to randomly sample the coarse matching feature point pairs to obtain effective matching 3D feature point pairs. The SVD decomposition is used to obtain the relative pose T between the current frame and the selected key frame. t ; The relative posture T t Offline pose T with keyframe k Combined with the first frame pose T0 as the reference, the final 6Dof is: T = T t *T k *T0.

8. A target object pose estimation system, characterized in that: include: An offline data set building module is used to build an offline data set of key frames of the target to be detected; The target object pose estimation module is used to perform real-time target pose estimation through online multi-frame matching using key frame offline data sets; The offline data set establishment module specifically includes: An acquisition unit is used to acquire multiple texture images and depth images surrounding a single target object, extract a target mask from the multiple texture images, extract an effective depth and effective texture of the target object according to the target mask, and convert the effective depth into an effective point cloud; The pose calculation unit is used to detect the coarse matching feature points in the effective texture of two consecutive frames of images, determine the three-dimensional matching feature point pairs with effective depth according to the pixel positions of the coarse matching feature points on the original image and the effective point cloud; randomly sample the three-dimensional matching feature point pairs to obtain effective matching three-dimensional feature point pairs, and decompose the effective matching three-dimensional feature point pairs to obtain the relative pose T of the two consecutive frames of images. ij , and take the first frame posture T0 as the initial frame, and determine the posture T of each subsequent frame relative to the initial frame i ; Calculate the pose T of the current frame i The rotation matrix R in i With each key frame rotation matrix R saved k The distance between the rotational postures; when R i With all R k If the distance value is greater than the threshold, it is a new key frame, otherwise continue to process the next frame until all images are processed; The data set establishment unit is used to establish a key frame offline data set of the target object according to the position, valid point cloud and target valid texture of the target object in all key frames obtained.

9. A computer device for estimating the position and posture of a target object, characterized in that: include: A memory, a processor, and a computer program stored in the memory, wherein the processor implements the steps of the target object pose estimation method according to any one of claims 1 to 7 when executing the computer program.

10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the target object pose estimation method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target feature based monocular vision rapid relative pose estimation system and method

    CN108225319A

  • Drainage pipeline three-dimensional reconstruction method, device and equipment based on neural radiation field

    CN118781295A