Local high-precision object pose recognition method and system in large scene
By combining deep learning and three-dimensional point cloud data, image segmentation and triangulation methods are used to identify objects in large scenarios, the problem of insufficient computing resource dependence and real-time in the existing technology is solved, and high-precision and high-speed recognition are achieved.
Patent Information
- Application Number
- CN202510460484.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
In large scenarios, it is difficult for the prior art to improve the real-timeness of pose estimation while maintaining high-precision recognition capabilities and reduce dependence on computing resources.
By combining the two-dimensional position information output from the recognition results of deep learning and three-dimensional point cloud data, an image segmentation model is used to coarsely identify the object position, local images are extracted and depth images are generated through triangulation method, the three-dimensional model is reconstructed, and posture correction and reverse verification are performed through the posture estimation module.
It realizes high-precision object pose recognition in standard hardware environments, with a recognition speed of 20FPS, meeting the needs of high-speed processing of real-time images, and has stronger industrial production environment advantages.
Smart Images

Figure CN119991816A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial automation technology, and in particular to a method and system for local high-precision object posture recognition in a large scene. Background Art
[0002] In industrial manufacturing, accurate recognition of object pose is of great significance for improving production efficiency and ensuring product quality. Currently, many industrial applications use end-to-end deep learning pose estimation networks to predict the spatial position and pose of objects directly from raw images or depth data. These methods learn the appearance features of objects by training models, and can accurately identify the pose of objects in various complex scenarios, providing a more direct solution.
[0003] When it comes to pose recognition in large scenes, due to the sharp increase in data volume, the pressure on computing resources and network bandwidth of existing models also increases, resulting in increased costs and challenges to real-time assurance. At the same time, the objects actually observed are usually small, and a large amount of computing resources are wasted on processing irrelevant objects. Therefore, how to maintain high-precision recognition capabilities while improving the real-time performance of pose estimation and reducing dependence on computing resources has become a key issue facing current industrial production. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for local high-precision object pose recognition in a large scene, by combining the two-dimensional position information output by the recognition results of deep learning with the three-dimensional point cloud data to perform more efficient object pose recognition, while reducing the computing power requirements without reducing the recognition ability as much as possible, so as to solve the problems raised in the above background technology.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for local high-precision object posture recognition in a large scene, comprising the following steps:
[0006] Step S1: using an industrial camera to collect gear images and perform preprocessing;
[0007] Step S2: inputting the preprocessed image into an image segmentation model to perform rough recognition of the object position to obtain an object recognition result;
[0008] Step S3: extracting a local image according to the object recognition result, calculating the depth value of each pixel by a triangulation method, and generating a depth image of the same size as the original image;
[0009] Preferably, generating a depth image of the same size as the original image specifically includes:
[0010] Step S301: feature extraction and matching: extracting a local image according to the object recognition result, extracting key point features from the local image respectively, wherein the features include but are not limited to corner points, edges, and calibration points, and comparing two sets of features to obtain matching points of the corresponding relationship between the two views;
[0011] Step S302: Determine pixel pairs: According to the matching points, pixel pairs between the left and right images are formed. For each pair of matching points , calculate the similarity ,in, is the pixel coordinate vector of the left image, is the pixel coordinate vector of the right image. The matching points are filtered by the similarity threshold and the best matching point pair is retained.
[0012] Step S303: Calculate depth information by triangulation: Calculate the depth value of each pixel according to the matching point pair, organize it into a matrix form with the same size as the original image, and each element represents the distance from the object surface at the corresponding position to the camera;
[0013] Step S304: Smoothing processing: The smoothing processing includes double edge filtering and median filtering algorithms.
[0014] Step S4: reconstructing a three-dimensional model according to the depth image;
[0015] Preferably, reconstructing a three-dimensional model according to the depth image specifically includes the following operating steps:
[0016] Step S401: Calculate the precise coordinates of each pixel in the three-dimensional space through the local depth image, generate point cloud data, input the point cloud data into a point cloud data processing model, and output filtered point cloud data;
[0017] Step S402: identifying noise points that are significantly different from surrounding points by analyzing the point cloud data, and removing the noise points by using statistical filtering and radius filtering;
[0018] Step S403: Combining the point cloud data obtained by the plane vision recognition method with the feature points of the three-dimensional geometric data of the object to construct a 3D model.
[0019] Step S5: performing pose estimation and correction on the three-dimensional model and the original object model, and performing reverse verification through the pose data.
[0020] Preferably, performing pose estimation and correction and performing reverse verification specifically include:
[0021] Object pose estimation includes the following steps:
[0022] Step S501: achieving rough alignment by feature point matching;
[0023] Step S502: Fitting the coordinate mapping function: using the least square method to fit the coordinate mapping function based on the matched feature point pairs, that is, the transformation relationship from the pose coordinate system of the original object model to the pose coordinate system of the reconstructed object;
[0024] Step S503: Estimating posture by integrating historical trajectories: Incorporating historical trajectories into posture estimation, and filtering posture using time as a weight. During the object's motion, the posture data of each time step is recorded, and filtering is performed using a time-related function;
[0025] Step S504: Reverse verification: Apply the estimated pose to the object model to check the geometric fit under the known model, wherein all object models that need to be identified are entered, and reverse verification can be performed through the obtained pose data to correct the recognition result within a credible range; reverse verification is to check the mapped position to determine whether the pose is consistent with the actual observation in the scene. If the estimated pose is not within the credible range, it can be corrected.
[0026] Step S505: When the error meets the requirement, the posture estimation is considered to be completed. When the error does not meet the requirement, the initial posture is optimized or adjusted, and the 6D posture information of the gear is output.
[0027] According to the above technical solution, a local high-precision object posture recognition system in a large scene is provided, including: an image acquisition module, an object recognition module, a depth image generation module, a 3D modeling module, and a posture estimation module;
[0028] Preferably, the image acquisition module is used to use an industrial camera to acquire images containing the identified object from different angles; the object recognition module is used to input the acquired image into an image segmentation model for recognition, and output the approximate position of the identified object; the local depth image generation module is used to generate a depth image containing the depth information of each pixel in a small range through binocular imaging technology based on the object recognition result; the 3D modeling module is used to generate a 3D model of the object based on the depth image, and to combine the point cloud data acquired by the plane vision recognition method with the feature points of the three-dimensional geometric data of the object to construct a 3D model; the pose estimation and correction module is used to perform pose estimation on the 3D model.
[0029] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: the present invention realizes rapid and accurate recognition of objects through local high-precision point cloud recognition based on image segmentation, and under a standard hardware environment, the recognition speed of the method of the present invention is 20FPS, which meets the requirements of high-speed processing of real-time images and has stronger advantages in industrial production environments that require efficient recognition of object posture. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0031] Figure 1 A schematic diagram of the application of a high-precision gear position recognition system in a large scenario provided by an embodiment of the present invention;
[0032] Figure 2 A module composition diagram of a local high-precision object posture recognition system in a large scene provided by an embodiment of the present invention;
[0033] Figure 3 A schematic diagram of the operation flow of an object recognition module in a local high-precision object posture recognition system in a large scene provided by an embodiment of the present invention;
[0034] Figure 4 A general flow chart of a method for local high-precision object pose recognition in a large scene provided by an embodiment of the present invention;
[0035] Figure 5 A schematic diagram of the process of local depth image processing in a method for local high-precision object pose recognition in a large scene provided by an embodiment of the present invention;
[0036] Figure 6 A schematic diagram of the process of 3D modeling in a method for local high-precision object pose recognition in a large scene provided by an embodiment of the present invention;
[0037] Figure 7 A schematic diagram of the process of object pose estimation in a method for local high-precision object pose recognition in a large scene provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0039] Embodiments of the present invention are combined Figures 1 to 7 As shown, the following technical solution is provided: a local high-precision object posture recognition system in a large scene, the recognition object is taken as an example of a gear, and is combined with Figure 1 As shown, the target gear 2 is the object to be identified, and the conveyor belt 4 transports the gear from the outside to the working area of the robotic arm; the binocular camera 3 has a high resolution and obtains a large amount of scene information; after obtaining the precise gear posture information, the robotic arm 1 is responsible for grasping the gear and further work.
[0040] Exemplary, combined Figure 2 As shown, the system includes an image acquisition module, an object recognition module, a depth image generation module, a 3D modeling module, and a pose estimation module.
[0041] Image acquisition module: This module collects multi-view images of the object by configuring different angles and parameters of industrial cameras; uses high-resolution cameras to ensure clear details, and automatically adjusts parameters such as exposure and focus to adapt to different lighting conditions;
[0042] Object recognition module: This module pre-processes the image output of the image acquisition module as the input of the image segmentation model, and outputs the recognition result in XML file format; optionally, the image segmentation model used for coarse recognition of object position in the object recognition module can adopt traditional image segmentation methods or image segmentation methods based on deep learning.
[0043] Further, combined with Figure 3 As shown, the operation process steps of the object recognition module include:
[0044] Step S201: data collection and annotation, by collecting original images or data samples and adding necessary labels including gear types or other identifiers for training;
[0045] Step S202: performing data enhancement processing, applying techniques such as rotation, flipping, scaling or color adjustment to expand the training data set and improve the robustness of the model;
[0046] Step S203: setting model hyperparameters, defining model hyperparameters including learning rate, batch size, and number of layers to ensure training optimization;
[0047] Step S204: Model training and validation, using the prepared data set to train the model and verifying its performance on a separate validation set to tune the parameters;
[0048] Step S205: Model performance evaluation: evaluating the accuracy, precision, recall and F1 score of the model on the test set to measure the effectiveness of the model;
[0049] Step S206: Model format conversion, converting the model into TensorRT, a deployment format suitable for real-time applications;
[0050] Step S207: Model deployment and optimization: deploy the model to the target environment and optimize it to ensure efficient operation.
[0051] Step S208: Real-time object recognition and detection, using the deployed model for real-time detection to identify objects in a real-time environment or video stream.
[0052] Local depth image generation module: This module reads the XML file containing the object recognition results, selects the local image to generate the depth map according to the recognition results, and the module can only process the same recognized object at the same time; optionally, the local depth image generation module measures the degree of noise by calculating the standard deviation of the depth map pixel values, and uses a smoothing algorithm to denoise the depth map that does not meet the standard, and uses the denoised depth map as the output of the local depth image generation module. In the same system, multiple local depth image generation modules are allowed to work in parallel to achieve parallel recognition of multiple objects.
[0053] 3D modeling module: This module takes depth image data and the identification object information in the XML file as input, and outputs a 3D model of the identified object. This module has a one-to-one correspondence with the local depth image generation module. Optionally, the 3D modeling module can combine the point cloud data obtained by the plane vision recognition method with the feature points of the object's three-dimensional geometric data to construct a 3D model. For objects that lack three-dimensional geometric data, the point cloud data is directly used to construct a 3D model.
[0054] Pose estimation and correction module: This module takes the 3D model as input, fits the coordinate mapping function according to the spatial position of the feature points, maps the pose coordinate system of the original object model to the reconstructed object, and calculates the 6D pose (position and attitude) of the object.
[0055] In this embodiment, combined Figure 4 As shown, a method for local high-precision object pose recognition in a large scene is provided, comprising the following steps:
[0056] Step S1: The industrial camera collects graphics and performs preprocessing;
[0057] Exemplarily, an industrial camera captures high-resolution images through precise optical imaging, and uses a Gaussian filter to remove high-frequency noise in the image and smooth the image. That is, an industrial camera is used to capture images of gears from different angles. In order to improve recognition accuracy, the captured images are preprocessed to reduce noise interference and highlight object features, wherein the preprocessing part includes denoising, graying or enhancing image contrast.
[0058] Step S2: Identify the object and estimate its location;
[0059] Exemplarily, the image obtained in step S1 is divided into S×S grids (S is a hyperparameter) and input into the target detection model for recognition, and a prediction result of each grid is obtained, which includes the category, position and confidence of the recognized object.
[0060] Step S3: generating and optimizing a local depth image;
[0061] Exemplarily, a local image is extracted according to the recognition result in step S2, a depth value of each pixel is calculated by a triangulation method, and a depth map of the same size as the original image is generated; in a specific implementation, this step includes the following steps: Figure 5 The process shown, Figure 5 A schematic diagram of the process of local depth image processing in a method for local high-precision object pose recognition in a large scene provided by an embodiment of the present invention.
[0062] Step S301: feature extraction and matching;
[0063] Exemplarily, key point features are extracted from the local images output in step S2, including but not limited to corner points, edges, and calibration points. The two sets of features are compared, and key points of corresponding relationships (i.e., the same physical point from the world coordinates) between the two views are found;
[0064] Step S302: determining pixel pairs;
[0065] Exemplarily, the matching points obtained through the above steps form pixel pairs between the left and right images. For each pair of matching points , calculate similarity ,Filter the matching points by similarity threshold and retain the best matching point pairs; is the pixel coordinate vector of the left image, is the pixel coordinate vector of the right image.
[0066] Step S303: Calculate depth information by triangulation;
[0067] Exemplarily, the depth value of each pixel is calculated based on the matching point pairs outputted in step S302 and organized into a matrix form with the same size as the original image, where each element represents the distance from the object surface at the corresponding position to the camera.
[0068] Step S304: smoothing processing;
[0069] Exemplarily, a double edge filter, a median filter or other algorithm is applied to the depth map generated in step S303 to smooth the result and reduce noise in the local depth image.
[0070] Step S4: reconstructing a three-dimensional model according to the image;
[0071] Exemplarily, the local depth image feature points obtained in step S3 are extracted and input into the model to obtain the correspondence between the depth image and the feature points in the three-dimensional geometric data of the object, determine the position correspondence between each feature point, and construct the 3D model according to the function;
[0072] Furthermore, in this embodiment, Figure 6As shown, 3D modeling specifically includes the following steps:
[0073] Step S401: Calculate the precise coordinates of each pixel in the world coordinate system through the local depth image to generate point cloud data; the data is saved in .obj format, storing the three-dimensional coordinates and additional information such as color and normal;
[0074] Among them, the pixel Exact coordinates in the world coordinate system for:
[0075]
[0076]
[0077]
[0078]
[0079] in, is the rotation matrix from the camera coordinate system to the world coordinate system, is the translation vector from the origin of the camera coordinate system to the origin of the world coordinate system, Pixel The coordinates in the camera coordinate system, Pixel The depth value of is the projection coordinate of the camera optical axis on the image plane, is the horizontal focal length of the camera.
[0080] Exemplarily, the point cloud model adjusts the resolution and density of the point cloud according to actual needs, so that the model can ensure a certain efficiency while meeting the refinement requirements. At the same time, the generated point cloud model is further optimized, such as simplification and smoothing, to improve the usability and rendering efficiency of the model.
[0081] Step S402: identifying noise points that are significantly different from surrounding points by analyzing the point cloud data, and applying statistical filtering and radius filtering to remove these noise points;
[0082] Exemplarily, the point cloud data of step S401 is input into a point cloud data processing model, and filtered point cloud data is output.
[0083] Step S403: combining the point cloud data obtained by the plane vision recognition method with the feature points of the three-dimensional geometric data of the object to construct a 3D model;
[0084] Exemplarily, the point cloud data outputted from step S402 and the original object model are used as input to output a three-dimensional reconstructed model of the identified object, wherein the point cloud data and the three-dimensional geometric data of the object are combined by fitting a function to the feature points of both parties, with the final fitting function serving as the connection between the two data of different dimensions.
[0085] Step S5: pose estimation and correction and verification;
[0086] Exemplarily, the model obtained in step S4 is subjected to pose estimation and correction with the original object model, and the 6D pose information of the identified object within the credibility range is output. The credibility judgment is based on the following: For the errors in translation and rotation, the credible range is defined as ,like Then the pose is said to be within the credibility range, where , is the mapping pose obtained from the estimated data, is the actual observed large scene pose.
[0087] Further integration Figure 7 As shown, the object pose estimation includes the following steps:
[0088] Step S501: achieving rough alignment by feature point matching;
[0089] Exemplarily, a suitable 3D feature descriptor SHOT is selected to describe the local geometric information around each feature point, thereby improving the stability of the matching, and appropriate multi-scale feature extraction is performed according to the object size to obtain a robust matching effect.
[0090] Step S502: fitting a coordinate mapping function;
[0091] Exemplarily, the least square method is used to fit the coordinate mapping function based on the matched feature point pairs, that is, the transformation relationship from the pose coordinate system of the original object model to the pose coordinate system of the reconstructed object.
[0092] Step S503: Estimating the pose by fusing the historical trajectory;
[0093] Exemplarily, the historical trajectory is incorporated into the pose estimation, and the pose is filtered using time as a weight. During the movement of the object, the pose data of each time step is recorded, and a function related to time is used for filtering;
[0094] Assume that the current position is , the historical trajectory is Where h represents the historical trajectory moment, the weighted average of the posture can be expressed as:
[0095]
[0096] in is the weight coefficient, is the attenuation coefficient, It's a historical moment. is the present moment, is the number of historical trajectories, It's a historical moment The corresponding posture.
[0097] Step S504: reverse verification;
[0098] Exemplarily, the estimated pose is applied to the object model to check the geometric fit under the known model. Among them, all object models that need to be identified are entered, and reverse verification can be performed through the obtained pose data to correct the recognition results within the credible range; reverse verification is to check the mapped position to determine whether the pose is consistent with the actual observation in the scene. If the estimated pose is not within the credible range, it can be corrected.
[0099] Step S505: If the error meets the requirement, the posture estimation is considered to be completed; if the error does not meet the requirement, further optimization or adjustment of the initial posture can be performed, and the 6D posture information of the gear is output. In this embodiment, an xml file is selected to save the relevant information;
[0100] When the posture error exceeds the threshold, it can be adjusted using the following correction formula:
[0101]
[0102] in, is the correction calculated by Kalman filtering.
[0103] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0104] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for local high-precision object pose recognition in a large scene, characterized by: The following steps are involved: Step S1: using an industrial camera to collect gear images and perform preprocessing; Step S2: inputting the preprocessed image into an image segmentation model to perform rough recognition of the object position to obtain an object recognition result; Step S3: extracting a local image according to the object recognition result, calculating the depth value of each pixel by a triangulation method, and generating a depth image of the same size as the original image; Step S4: reconstructing a three-dimensional model according to the depth image; Step S5: performing pose estimation and correction on the three-dimensional model and the original object model, and performing reverse verification through the pose data.
2. According to the method for local high-precision object pose recognition in a large scene according to claim 1, it is characterized by: The step S3 further comprises: Step S301: feature extraction and matching: extracting a local image according to the object recognition result, extracting key point features from the local image respectively, wherein the features include but are not limited to corner points, edges, and calibration points, and comparing two sets of features to obtain matching points of the corresponding relationship between the two views; Step S302: Determine pixel pairs: According to the matching points, pixel pairs between the left and right images are formed. For each pair of matching points , calculate the similarity ,in, is the pixel coordinate vector of the left image, is the pixel coordinate vector of the right image. The matching points are filtered by the similarity threshold and the best matching point pair is retained. Step S303: Calculate depth information by triangulation: Calculate the depth value of each pixel according to the matching point pair, organize it into a matrix form with the same size as the original image, and each element represents the distance from the object surface at the corresponding position to the camera; Step S304: Smoothing processing: The smoothing processing includes adopting double edge filtering and median filtering algorithms.
3. The method for local high-precision object pose recognition in a large scene according to claim 2, characterized in that: The reconstruction of the three-dimensional model according to the depth image specifically includes the following operation steps: Step S401: Calculate the precise coordinates of each pixel in the three-dimensional space through the local depth image, generate point cloud data, input the point cloud data into a point cloud data processing model, and output filtered point cloud data; Step S402: identifying noise points that are significantly different from surrounding points by analyzing the point cloud data, and removing the noise points by using statistical filtering and radius filtering; Step S403: Combining the point cloud data obtained by the plane vision recognition method with the feature points of the three-dimensional geometric data of the object to construct a 3D model.
4. The method for local high-precision object pose recognition in a large scene according to claim 3, characterized in that: Inputting the point cloud data into the point cloud data processing model includes: adjusting the resolution and density of the point cloud according to actual needs through the point cloud model, and optimizing the generated point cloud model including simplification and smoothing; Among them, the pixel Exact coordinates in the world coordinate system for: ; ; ; ; in, is the rotation matrix from the camera coordinate system to the world coordinate system, is the translation vector from the origin of the camera coordinate system to the origin of the world coordinate system, Pixel The coordinates in the camera coordinate system, Pixel The depth value of is the projection coordinate of the camera optical axis on the image plane, is the horizontal focal length of the camera.
5. The method for local high-precision object pose recognition in a large scene according to claim 4, characterized in that: The 3D model construction is to take the point cloud data and the original object model as input, and output a three-dimensional reconstructed model of the identified object, wherein the point cloud data is combined with the three-dimensional geometric data of the object by fitting a function to the feature points of both parties, and the final fitting function is used as the connection between the two different dimensional data.
6. The method for local high-precision object pose recognition in a large scene according to claim 5, characterized in that: The step S5 specifically includes: performing pose estimation and correction on the 3D model and the original object model, and outputting 6D pose information of the identified object within a credibility range. The credibility judgment is based on the following: For the error in translation and rotation, the credible range is defined as ,like Then the pose is said to be within the credibility range, where , is the mapping pose obtained from the estimated data, is the actual observed large scene pose.
7. The method for local high-precision object pose recognition in a large scene according to claim 6, characterized in that: The object pose estimation comprises the following steps: Step S501: achieving rough alignment by feature point matching; Step S502: Fitting the coordinate mapping function: using the least square method to fit the coordinate mapping function based on the matched feature point pairs, that is, the transformation relationship from the pose coordinate system of the original object model to the pose coordinate system of the reconstructed object; Step S503: Estimating posture by integrating historical trajectories: Incorporating historical trajectories into posture estimation, and filtering posture using time as a weight. During the movement of the object, the posture data of each time step is recorded, and filtering is performed using a function related to time; Step S504: reverse verification: applying the estimated pose to the object model to check the geometric fit under the known model, wherein all object models to be identified are entered, and reverse verification can be performed through the obtained pose data, and the recognition result is corrected within the credible range; reverse verification is to check the mapped position to determine whether the pose is consistent with the actual observation in the scene, and to correct the estimated pose when it is not within the credible range; Step S505: When the error meets the requirement, the posture estimation is considered to be completed. When the error does not meet the requirement, the initial posture is optimized or adjusted, and the 6D posture information of the gear is output.
8. The method for local high-precision object pose recognition in a large scene according to claim 7, characterized in that: The fusion history trajectory estimation posture specifically includes: when the posture at the current moment is , the historical trajectory is Where h represents the historical trajectory moment, the weighted average of the posture can be expressed as: ; in, is the weight coefficient, is the attenuation coefficient, It's a historical moment. is the present moment, is the number of historical trajectories, It's a historical moment The corresponding posture.
9. The method for local high-precision object pose recognition in a large scene according to claim 8, characterized in that: When the error is not satisfied, optimizing or adjusting the initial posture includes: when the error of the posture exceeds the threshold, adjusting by the following correction formula: ; in, is the correction calculated by Kalman filtering.
10. A local high-precision object pose recognition system in a large scene, characterized by: include: Image acquisition module, object recognition module, depth image generation module, 3D modeling module, posture estimation module; the image acquisition module is used to use an industrial camera to collect images containing recognized objects from different angles; the object recognition module is used to input the collected images into an image segmentation model for recognition, and output the approximate position of the recognized object; the local depth image generation module is used to generate a depth image containing the depth information of each pixel in a small range according to the object recognition result through binocular imaging technology; the 3D modeling module is used to generate a 3D model of the object according to the depth image, and combine the point cloud data obtained by the plane vision recognition method with the feature points of the three-dimensional geometric data of the object to construct a 3D model; the posture estimation and correction module is used to perform posture estimation on the 3D model.
Citation Information
Patent Citations
Irregular object pose estimation method and device based on depth camera
CN113450408A
Workpiece pose estimation method and device oriented to disordered sorting scene
CN115359119A
Stacked workpiece pose estimation method based on combination of two-dimensional image and three-dimensional point cloud
CN116309847A
Local refinement mapping system and method based on SLAM and semantic segmentation
CN116772820A
Calibration block and hand eye calibration method for line laser sensor
JP2022039903A