A Method and System for Identifying the Pose of a Local High-Precision Object in a Large-Scene
By combining deep learning and three-dimensional point cloud data, the real-time and computing resource problems of pose estimation in large scenarios are solved, and fast and accurate object pose recognition is achieved, which is suitable for industrial production environments.
Patent Information
- Application Number
- CN202510460484.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-14
AI Technical Summary
In large scenarios, it is difficult for the prior art to improve the real-timeness of pose estimation while maintaining high-precision recognition capabilities and reduce dependence on computing resources, resulting in increased computing costs and insufficient real-timeness.
Combining the recognition results of deep learning and three-dimensional point cloud data, through image segmentation, depth image generation, three-dimensional model reconstruction and pose estimation, the computing power requirements are reduced and local high-precision object pose pose recognition is achieved.
It realizes fast and accurate object recognition, with a recognition speed of 20FPS, meeting the needs of high-speed processing of real-time images and reducing the dependence of computing resources.
Smart Images

Figure CN119991816B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial automation, and particularly relates to a method and system for local high-precision object pose recognition in a large scene. Background Art
[0002] In industrial manufacturing, the accurate recognition of object poses is of great significance for improving production efficiency and ensuring product quality. At present, many industrial applications adopt end-to-end deep learning pose estimation networks to directly predict the spatial position and orientation of objects from raw images or depth data. These methods learn the appearance features of objects by training models and can accurately recognize object poses in various complex scenarios, providing a relatively direct solution.
[0003] When it comes to pose recognition in a large scene, due to the sharp increase in data volume, the pressure on computing resources and network bandwidth of using existing models also rises, resulting in increased costs and challenges to real-time guarantee. At the same time, the actual observed objects are usually small, while a large amount of computing resources are wasted on processing irrelevant objects. Therefore, how to improve the real-time performance of pose estimation and reduce the dependence on computing resources while maintaining high-precision recognition ability has become a key problem faced in current industrial production. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for local high-precision object pose recognition in a large scene, which combines the two-dimensional position information output by the recognition result of deep learning with three-dimensional point cloud data for more efficient object pose recognition, and reduces the computing power requirements without reducing the recognition ability as much as possible, so as to solve the problems raised in the above background art.
[0005] To solve the above technical problems, the present invention provides the following technical solution: A method for local high-precision object pose recognition in a large scene, comprising the following steps:
[0006] Step S1: Use an industrial camera to collect gear images and perform preprocessing;
[0007] Step S2: Input the preprocessed image into an image segmentation model for rough recognition of object positions to obtain object recognition results;
[0008] Step S3: Extract local images according to the object recognition results, calculate the depth value of each pixel by the triangulation method, and generate a depth image with the same size as the original image;
[0009] Preferably, generating a depth image with the same size as the original image specifically includes:
[0010] Step S301: Feature extraction and matching: Extract local images based on the object recognition result, and respectively extract key point features from the local images. The features include, but are not limited to: corner points, edges, and calibration points. Compare the two sets of features to obtain matching points of the corresponding relationship existing between the two views;
[0011] Step S302: Determine pixel pairs: Form pixel pairs between the left and right images according to the matching points. For each pair of matching points (f Li , f Rj ), calculate the similarity where f Li is the pixel point coordinate vector of the left image, and f Rj is the pixel point coordinate vector of the right image. Filter the matching points through the threshold of the similarity, and retain the optimal matching point pairs;
[0012] Step S303: Triangulation to calculate depth information: Calculate the depth values of each pixel according to the matching point pairs, and organize them into a matrix form with the same size as the original image. Each element represents the distance from the object surface at the corresponding position to the camera;
[0013] Step S304: Smoothing processing: The smoothing processing includes bilateral edge filtering and median filtering algorithms.
[0014] Step S4: Reconstruct a 3D model based on the depth image;
[0015] Preferably, reconstructing the 3D model based on the depth image specifically includes the following running steps:
[0016] Step S401: Calculate the exact coordinates of each pixel point in the 3D space through the local depth image, generate point cloud data, and input the point cloud data into the point cloud data processing model to output filtered point cloud data;
[0017] Step S402: Identify noise points with large differences from surrounding points by analyzing the point cloud data, and remove the noise points using statistical filtering and radius filtering;
[0018] Step S403: Combine the point cloud data obtained by the plane vision recognition method with the feature points of the 3D geometric data of the object to construct a 3D model.
[0019] Step S5: Perform pose estimation and correction on the 3D model and the original object model, and perform reverse verification through the pose data.
[0020] Preferably, performing pose estimation and correction and performing reverse verification specifically include:
[0021] Object pose estimation includes the following steps;
[0022] Step S501: Coarse alignment is achieved by feature point matching;
[0023] Step S502: Fitting the coordinate mapping function: Using the least squares method to fit the coordinate mapping function based on the matched feature point pairs, that is, the transformation relationship from the pose coordinate system of the original object model to the pose coordinate system of the reconstructed object;
[0024] Step S503: Fusing historical trajectories to estimate pose: Incorporating historical trajectories into pose estimation and filtering the pose using time as a weight. During the movement of the object, the pose data at each time step is recorded, and filtering is performed using a time-related function;
[0025] Step S504: Reverse verification: Applying the estimated pose to the object model and checking the geometric fitting under the known model. Among them, for the object models to be identified, they are all entered, and reverse verification can be performed through the obtained pose data, and the recognition result can be corrected within the credible range; Reverse verification is to check the mapped position to determine whether the pose conforms to the actual observation in the scene. If the estimated pose is not within the credible range, it can be corrected.
[0026] Step S505: When the error meets the requirements, it is considered that the pose estimation is completed. When the error does not meet the requirements, optimize or adjust the initial pose, and output the 6D pose information of the gear.
[0027] According to the above technical solution, a local high-precision object pose recognition system in a large scene is provided, including: an image acquisition module, an object recognition module, a depth image generation module, a 3D modeling module, and a pose estimation module;
[0028] Preferably, the image acquisition module is used to collect images containing the recognition object from different angles by using an industrial camera; the object recognition module is used to input the collected images into an image segmentation model for recognition and output the approximate position of the recognition object; the depth image generation module is used to generate a depth image with depth information of each pixel in a small range according to the object recognition result through binocular imaging technology; the 3D modeling module is used to generate a 3D model of the object according to the depth image, and combine the point cloud data obtained by the planar vision recognition method with the feature points of the three-dimensional geometric data of the object to construct a 3D model; the pose estimation module is used to estimate the pose of the 3D model.
[0029] Compared with the prior art, the beneficial effects achieved by the present invention are: The present invention realizes the fast and accurate recognition of objects through local high-precision point cloud recognition based on image segmentation. And in a standard hardware environment, the recognition speed of the method of the present invention is 20FPS, meeting the requirements of high-speed processing of real-time images, and having stronger advantages in industrial production environments that require efficient recognition of object poses. Brief Description of the Drawings
[0030] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.
[0031] In the accompanying drawings:
[0032] Figure 1 It is a schematic application diagram of a high-precision gear pose recognition system in a large scene provided by an embodiment of the present invention;
[0033] Figure 2 It is a module composition diagram of a local high-precision object pose recognition system in a large scene provided by an embodiment of the present invention;
[0034] Figure 3 It is a schematic flow diagram of the operation of an object recognition module in a local high-precision object pose recognition system in a large scene provided by an embodiment of the present invention;
[0035] Figure 4 It is a general flow chart of a method for local high-precision object pose recognition in a large scene provided by an embodiment of the present invention;
[0036] Figure 5 It is a schematic flow diagram of local depth image processing in a method for local high-precision object pose recognition in a large scene provided by an embodiment of the present invention;
[0037] Figure 6 It is a schematic flow diagram of 3D modeling in a method for local high-precision object pose recognition in a large scene provided by an embodiment of the present invention;
[0038] Figure 7 It is a schematic flow diagram of object pose estimation in a method for local high-precision object pose recognition in a large scene provided by an embodiment of the present invention. Detailed Embodiments
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0040] The embodiments of the present invention are combined with Figures 1 to 7 as shown, and the following technical solutions are provided: A local high-precision object pose recognition system in a large scene, taking a gear as an example for the recognition object. Exemplarily, in combination with Figure 1As shown, the target gear 2 is the object to be recognized, and the conveyor belt 4 transports the gear from the outside to the working area of the robotic arm; the binocular camera 3 has a high resolution and can obtain a large amount of scene information; after obtaining the accurate gear pose information, the robotic arm 1 is responsible for grasping the gear and further operations.
[0041] Exemplarily, in combination with Figure 2 As shown, the system includes an image acquisition module, an object recognition module, a depth image generation module, a 3D modeling module, and a pose estimation module.
[0042] Image acquisition module: This module collects multi-view images of the object by configuring different angles and parameters of the industrial camera; uses a high-resolution camera to ensure clear details, and automatically adjusts parameters such as exposure and focal length to adapt to different lighting conditions.
[0043] Object recognition module: This module preprocesses the image output of the image acquisition module and uses it as the input of the image segmentation model, and the recognition result is output in the XML file format; optionally, for the image segmentation model used for rough recognition of the object position in the object recognition module, traditional image segmentation methods or image segmentation methods based on deep learning can be used.
[0044] Further, in combination with Figure 3 As shown, the operation process steps of the object recognition module include:
[0045] Step S201: Data collection and annotation, by collecting original images or data samples, adding necessary labels including gear types or other identifiers for training.
[0046] Step S202: Perform data augmentation processing, applying techniques such as rotation, flipping, scaling, or color adjustment to expand the training dataset and improve the robustness of the model.
[0047] Step S203: Set the model hyperparameters, define the hyperparameters of the model including learning rate, batch size, and number of layers to ensure training optimization.
[0048] Step S204: Model training and validation, use the prepared dataset to train the model and verify its performance on a separate validation set to tune the parameters.
[0049] Step S205: Model performance evaluation, evaluate the accuracy, precision, recall, and F1 score of the model on the test set to measure the effectiveness of the model.
[0050] Step S206: Model format conversion, convert the model to the deployment format TensorRT suitable for real-time applications.
[0051] Step S207: Model deployment and optimization, deploy the model to the target environment and optimize it to ensure efficient operation.
[0052] Step S208: Real-time object recognition and detection, using the deployed model for real-time detection to identify objects in a real-time environment or video stream.
[0053] Depth image generation module: This module reads the XML file containing the object recognition results, selects local images according to the recognition results to generate depth maps, and can only process the same recognized object at the same time; optionally, the depth image generation module measures the degree of noise by calculating the standard deviation of the depth image pixel values, and uses a smoothing algorithm to denoise the depth maps that do not meet the standards, and uses the denoised depth maps as the output of the depth image generation module. In the same system, multiple depth image generation modules are allowed to work in parallel to achieve parallel recognition of multiple objects.
[0054] 3D modeling module: This module takes the depth image data and the recognition object information in the XML file as inputs and outputs the 3D model of the recognized object. This module has a one-to-one correspondence with the depth image generation module; optionally, the 3D modeling module can combine the point cloud data obtained by the planar vision recognition method with the feature points of the three-dimensional geometric data of the object to construct a 3D model, and directly uses the point cloud data to construct a 3D model for objects lacking three-dimensional geometric data.
[0055] Pose estimation module: This module takes the 3D model as an input, fits the coordinate mapping function according to the spatial positions of the feature points, maps the pose coordinate system of the original object model to the reconstructed object, and calculates the 6D pose (position and orientation) of the object.
[0056] In this embodiment, in combination with Figure 4 as shown, a method for local high-precision object pose recognition in a large scene is provided, including the following steps:
[0057] Step S1: An industrial camera captures a graphic and performs preprocessing;
[0058] Exemplarily, the industrial camera captures a high-resolution image through precise optical imaging and uses a Gaussian filter to remove high-frequency noise in the image and smooth the image, that is, the industrial camera captures images of gears from different angles. To improve the recognition accuracy, the captured images are preprocessed to reduce noise interference and highlight object features, where the preprocessing part includes denoising, grayscale conversion, or enhancing image contrast, etc.
[0059] Step S2: Identify the object and estimate its position;
[0060] Exemplarily, the image obtained in step S1 is segmented into a grid of S×S (S is a hyperparameter) and input into the target detection model for recognition to obtain the prediction results of each grid, and the results include the category, position, and confidence of the recognized object.
[0061] Step S3: Generate and optimize the local depth image;
[0062] Exemplarily, extract the local image according to the recognition result in step S2, calculate the depth value of each pixel by the triangulation method, and generate a depth map with the same size as the original image; In a specific implementation, this step includes as Figure 5 the process shown, Figure 5 is a schematic flow diagram of local depth image processing in a local high-precision object pose recognition method provided by an embodiment of the present invention for a large scene.
[0063] Step S301: Feature extraction and matching;
[0064] Exemplarily, extract key point features from the local images output in step S2 respectively, and these features include but are not limited to: corner points, edges, calibration points. Compare these two sets of features and find the key points with the corresponding relationship (i.e., the same physical point from the world coordinate) existing between the two views;
[0065] Step S302: Determine pixel pairs;
[0066] Exemplarily, the matching points obtained through the above steps form pixel pairs between the left and right images. For each pair of matching points (f Li , f Rj ), calculate the similarity Filter the matching points through the threshold of the similarity, and retain the optimal pair of matching points; where f Li is the pixel point coordinate vector of the left image, and f Rj is the pixel point coordinate vector of the right image.
[0067] Step S303: Triangulation to calculate depth information;
[0068] Exemplarily, calculate the depth value of each pixel according to the pair of matching points output in step S302, and organize it into a matrix form with the same size as the original image, and each element represents the distance from the object surface at the corresponding position to the camera.
[0069] Step S304: Smoothing processing;
[0070] Exemplarily, apply algorithms such as double-edge filtering and median filtering to the depth map generated in step S303 to smooth the result and reduce the noise in the local depth image.
[0071] Step S4: Reconstruct a three-dimensional model according to the image;
[0072] Exemplarily, the feature points of the partial depth image obtained in step S3 are extracted and input into the model to obtain the corresponding relationship between the feature points in the depth image and the three-dimensional geometric data of the object, determine the position corresponding relationship of each feature point, and construct a 3D model according to this function;
[0073] Further, in this embodiment, in combination with Figure 6 as shown, the specific steps of 3D modeling include the following;
[0074] Step S401: Calculate the exact coordinates of each pixel point in the world coordinate system through the partial depth image to generate point cloud data; the data is saved in the.obj format, storing additional information such as three-dimensional coordinates, colors, and normals;
[0075] Among them, the exact coordinates (X W , Y W , Z W ) of the pixel point (i, j) in the world coordinate system are:
[0076]
[0077] Z C = D(i, j)
[0078] Among them, R is the rotation matrix from the camera coordinate system to the world coordinate system, T is the translation vector from the origin of the camera coordinate system to the origin of the world coordinate system, (X C , Y C , Z C ) are the coordinates of the pixel point (i, j) in the camera coordinate system, D(i, j) is the depth value of the pixel point (i, j), (u, v) are the projection coordinates of the camera optical axis on the image plane, and fx is the horizontal focal length of the camera.
[0079] Exemplarily, the point cloud model adjusts the resolution and density of the point cloud according to actual needs, so that the model can ensure a certain efficiency while meeting the refinement requirements. At the same time, the generated point cloud model is further optimized, such as simplification, smoothing, etc., to improve the usability and rendering efficiency of the model.
[0080] Step S402: Identify the noise points that are significantly different from the surrounding points by analyzing the point cloud data, and apply statistical filtering and radius filtering to remove these noise points;
[0081] Exemplarily, the point cloud data in step S401 is input into the point cloud data processing model, and the filtered point cloud data is output.
[0082] Step S403: Combine the point cloud data obtained by the planar vision recognition method with the feature points of the three-dimensional geometric data of the object to construct a 3D model;
[0083] Exemplarily, the point cloud data output from step S402 and the original object model are used as inputs to output the three-dimensional reconstructed model of the recognized object. The combination of the point cloud data and the three-dimensional geometric data of the object is achieved by performing function fitting on the feature points of both sides, and the finally obtained fitting function is used as the connection between the two different-dimensional data.
[0084] Step S5: Pose estimation and correction and verification;
[0085] Exemplarily, the model obtained in step S4 and the original object model are used for pose estimation and correction, and the 6D pose information of the recognized object within the confidence range is output. The basis for confidence judgment is as follows: is the error in translation and rotation. The confidence range is defined as ∈. If then it is said that the pose is within the confidence range, where is the mapped pose obtained from the estimated data, is the actual observed large-scene pose.
[0086] Further combination Figure 7 As shown, object pose estimation includes the following steps;
[0087] Step S501: Coarse alignment is achieved by using feature point matching;
[0088] Exemplarily, by selecting a suitable 3D feature descriptor SHOT to describe the local geometric information around each feature point, the stability of the matching is improved. At the same time, appropriate multi-scale feature extraction is performed according to the object size to obtain a robust matching effect.
[0089] Step S502: Fit the coordinate mapping function;
[0090] Exemplarily, the least squares method is used to fit the coordinate mapping function based on the matched feature point pairs, that is, the transformation relationship from the pose coordinate system of the original object model to the pose coordinate system of the reconstructed object.
[0091] Step S503: Fuse the historical trajectory to estimate the pose;
[0092] Exemplarily, the historical trajectory is incorporated into the pose estimation, and time is used as the weight to filter the pose. During the movement of the object, the pose data at each time step is recorded, and filtering processing is performed using a time-related function;
[0093] Assume that the pose at the current moment is P t , and the historical trajectory is where h represents the historical trajectory time, then the pose weighted average can be expressed as:
[0094]
[0095] where ω h = exp(-μ(t - t h )) is the weight coefficient, μ is the decay coefficient, t h is the historical moment, t is the current moment, N is the number of historical trajectories, is the pose corresponding to the historical moment t h corresponding thereto.
[0096] Step S504: Reverse verification;
[0097] Exemplarily, apply the estimated pose to the object model and check the geometric fitting under the known model. Among them, for the object models to be recognized, all are entered, and the reverse verification can be performed through the obtained pose data, and the recognition result can be corrected within the credible range; the reverse verification is to check the mapped position to determine whether the pose conforms to the actual observation in the scene. If the estimated pose is not within the credible range, it can be corrected.
[0098] Step S505: If the error meets the requirements, it is considered that the pose estimation is completed; if the error does not meet the requirements, further optimization or adjustment of the initial pose can be performed, and the 6D pose information of the gear is output. In this embodiment, an xml file is selected to save the relevant information;
[0099] When the error of the pose exceeds the threshold, it can be adjusted by the following correction formula:
[0100]
[0101] where ΔT is the correction amount calculated by Kalman filtering.
[0102] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0103] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for identifying the pose of a local high-precision object in a large-scale scene, characterized in that: It includes the following steps: Step S1: Use an industrial camera to collect gear images and perform preprocessing; Step S2: Input the preprocessed image into an image segmentation model for rough recognition of the object position to obtain an object recognition result; Step S3: Extract a local image according to the object recognition result, calculate the depth value of each pixel by the triangulation method, and generate a depth image with the same size as the original image; Step S4: Reconstruct a 3D model according to the depth image; Step S5: Perform pose estimation and correction on the 3D model and the original object model, and perform reverse verification through pose data; The performing of pose estimation and correction includes the following steps; Step S501: Use feature point matching to achieve rough alignment; Step S502: Fit a coordinate mapping function: Use the least squares method to fit a coordinate mapping function based on the matched feature point pairs, that is, the transformation relationship from the pose coordinate system of the original object model to the pose coordinate system of the reconstructed object; Step S503: Fuse the historical trajectory to estimate the pose: Incorporate the historical trajectory into the pose estimation, and use time as a weight to filter the pose. During the movement of the object, record the pose data at each time step and perform filtering processing using a time-related function; Step S504: Reverse verification: Apply the estimated pose to the object model and check the geometric fitting under the known model. Among them, for the object models to be recognized, they are all entered, and reverse verification can be performed through the obtained pose data, and the recognition result is corrected within the credible range; Reverse verification is to check the mapped position to determine whether the pose conforms to the actual observation in the scene, and correct it when the estimated pose is not within the credible range; Step S505: When the error meets the requirements, it is considered that the pose estimation is completed. When the error does not meet the requirements, optimize or adjust the initial pose, and output the 6D pose information of the gear.
2. The local high-precision object pose recognition method in a large scene according to claim 1, characterized in that: The step S3 further includes: Step S301: Feature extraction and matching: Extract a local image according to the object recognition result, and respectively extract key point features from the local image. The key point features include but are not limited to: corner points, edges, calibration points, and obtain matching points of the corresponding relationship existing between the two views by comparing the two groups of features; Step S302: Determine pixel pairs: Form pixel pairs between the left and right images based on the matching points. For each pair of matching points (f Li , f Rj ), calculate the similarity where f Li is the pixel point coordinate vector of the left image, and f Rj is the pixel point coordinate vector of the right image. Filter the matching points through the similarity threshold and retain the optimal matching point pairs; Step S303: Triangulation to calculate depth information: Calculate the depth value of each pixel according to the matching point pairs, organize it into a matrix form with the same size as the original image, and each element represents the distance from the object surface at the corresponding position to the camera; Step S304: Smoothing processing: The smoothing processing includes using a bilateral edge filter and a median filter algorithm.
3. The method for identifying the pose of a local high-precision object in a large scene according to claim 2, wherein: The reconstructing the 3D model according to the depth image specifically includes the following running steps: Step S401: Calculate the exact coordinates of each pixel point in the 3D space through the local depth image to generate point cloud data, input the point cloud data into a point cloud data processing model, and output filtered point cloud data; Step S402: Identify noise points with large differences from surrounding points by analyzing the point cloud data, and remove the noise points using statistical filtering and radius filtering; Step S403: Combine the point cloud data obtained by the planar vision recognition method with the feature points of the three-dimensional geometric data of the object to construct a 3D model.
4. A method for identifying the pose of a local high-precision object in a large scene according to claim 3, characterized in that: The inputting of the point cloud data into the point cloud data processing model includes: adjusting the resolution and density of the point cloud according to actual requirements through the point cloud model, and performing optimizations including simplification and smoothing on the generated point cloud model; Among them, the exact coordinates (X W , Y W , Z W ) of the pixel point (i, j) in the world coordinate system are: Z C = D(i, j) where R is the rotation matrix from the camera coordinate system to the world coordinate system, T is the translation vector from the origin of the camera coordinate system to the origin of the world coordinate system, (X C , Y C , Z C ) are the coordinates of the pixel point (i, j) in the camera coordinate system, D(i, j) is the depth value of the pixel point (i, j), (u, v) are the projection coordinates of the camera optical axis on the image plane, and fx is the horizontal focal length of the camera.
5. A method for local high-precision object pose recognition in a large scene according to claim 4, characterized in that: The constructing of the 3D model takes the point cloud data and the original object model as inputs and outputs the three-dimensionally reconstructed model of the recognized object. The combination of the point cloud data and the three-dimensional geometric data of the object is achieved by performing function fitting on the feature points of both sides, and using the finally obtained fitting function as the connection between the two different-dimensional data.
6. The method and system for local high-precision object pose recognition in a large scene according to claim 5, characterized in that: The specific steps of step S5 include: performing pose estimation and correction on the 3D model and the original object model, and outputting the 6D pose information of the recognized object within the confidence range. The basis for confidence judgment is as follows: is the error in translation and rotation. The confidence range is defined as ∈. If then it is said that the pose is within the confidence range, where is the mapped pose obtained according to the estimated data, is the large-scale scene pose actually observed.
7. A method for local high-precision object pose recognition in a large scene according to claim 1, characterized in that: The pose estimated by fusing the historical trajectory specifically includes: when the pose at the current moment is P t , the historical trajectory is where h represents the historical trajectory time, then the weighted average of the pose can be expressed as: where ω h = exp(-μ(t - t h )) is the weight coefficient, μ is the decay coefficient, t h is the historical moment, t is the current moment, N is the number of historical trajectories, is the pose corresponding to the historical moment t h .
8. A method and system for local high-precision object pose recognition in a large-scale scene according to claim 7, characterized in that: The optimizing or adjusting of the initial pose when the error is not satisfied includes: when the pose error exceeds the threshold, adjusting through the following correction formula: where ΔT is the correction amount calculated by Kalman filtering.
9. A local high-precision object pose recognition system in a large scene, which executes the method for recognizing the local high-precision object pose in a large scene described in claim 1, and is characterized in that: It includes an image acquisition module, an object recognition module, a local depth image generation module, a 3D modeling module, and a pose estimation and correction module; the image acquisition module is used to collect images containing the recognized object from different angles using an industrial camera; the object recognition module is used to input the collected images into an image segmentation model for recognition and output the position of the recognized object; the local depth image generation module is used to generate a depth image containing depth information of each pixel in a small range through binocular imaging technology according to the object recognition result; the 3D modeling module is used to generate a 3D model of the object based on the depth image, and combine the point cloud data obtained by the planar vision recognition method with the feature points of the three-dimensional geometric data of the object to construct a 3D model; the pose estimation and correction module is used to perform pose estimation on the 3D model.
Citation Information
Patent Citations
Irregular object pose estimation method and device based on depth camera
CN113450408A