Visual SLAMMOT system and method fusing object plane features
By integrating the extraction and association of object plane features in the visual SLAMMOT system, the problem of insufficient position estimation accuracy in dynamic environments is solved, and higher precision target tracking and environmental mapping are achieved, improving the performance of robots and autonomous driving.
Patent Information
- Application Number
- CN202510326155.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-19
AI Technical Summary
When dealing with dynamic environments, existing visual SLAMMOT methods are difficult to effectively utilize plane features, resulting in insufficient accuracy of camera position estimation and dynamic object position estimation, affecting the performance of tasks such as target tracking and obstacle avoidance.
The visual SLAMMOT system and method that integrates the plane features of an object are extracted and correlated with the plane features of a dynamic object and incorporated them into the optimization framework as constraints, thereby improving the accuracy of target pose estimation and reconstruction.
It significantly improves the position estimation and reconstruction accuracy of the visual SLAMMOT system in dynamic environments, enhances the robustness and applicability of the system, especially in complex environments, improves the robot positioning and navigation capabilities and the perception capabilities of autonomous driving.
Smart Images

Figure CN120339376A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a visual SLAM MOT system and method integrating object plane features. Background Art
[0002] Visual Simultaneous Localization and Mapping (SLAM) is a technology that enables a platform equipped with a visual sensor to simultaneously perform self-localization and map construction in an unknown environment. With the rapid development of fields such as robotics and autonomous driving, visual SLAM, as an important technology in the field of computer vision, has received extensive attention in recent years and has become one of the active research directions.
[0003] Existing visual SLAM methods usually assume that the scene is static, that is, the changes between adjacent frames are only caused by camera movement. However, the actual application environment often contains a large number of dynamic objects, which may lead to errors in camera pose estimation. To solve this problem, some methods use the Random Sample Consensus (RANSAC) strategy to detect and remove outliers, while others reduce errors by detecting movable objects and removing dynamic feature points related to them. However, information such as the pose and motion trend of dynamic objects is crucial for tasks such as target tracking and obstacle avoidance. Therefore, recent research has begun to explore the combination of visual SLAM and Multiple Object Tracking (MOT), that is, the visual SLAM MOT technology.
[0004] Currently, most visual SLAM MOT methods are mainly based on point features. However, the real-world environment usually contains rich and stable plane features. Existing research has shown that these features can be used to improve the accuracy of localization and mapping and have been applied in some visual SLAM methods. For example, some methods use static plane features such as walls or floors to optimize camera pose estimation. However, the application of plane features in the pose estimation of dynamic objects is still in the exploratory stage and has not yet formed a mature method system.
[0005] Therefore, how to effectively utilize plane features in the visual SLAM MOT task to improve the pose estimation and reconstruction accuracy of dynamic objects remains an urgent technical problem to be solved. Summary of the Invention
[0006] The present invention provides a visual SLAMMOT system and method for integrating the plane features of objects, so as to solve the defects existing in the prior art. By accurately extracting and associating the plane features of dynamic objects and integrating them into the optimization framework as constraints, the accuracy of target pose estimation and reconstruction is improved, and the performance of the visual SLAMMOT system in dynamic environments is enhanced.
[0007] In a first aspect, the present invention provides a visual SLAMMOT system for integrating plane features of an object, comprising: RGB-D cameras, Isaac Sim simulation platform and computing equipment; The RGB-D camera collects RGB images and depth information of the target scene; The Isaac Sim simulation platform synthesizes the RGB image and depth information to generate an RGB-D image; The computing device receives the RGB-D image, performs camera pose estimation and static mapping, object pose estimation and dynamic reconstruction, and generates a target scene environment map.
[0008] In a second aspect, the present invention further provides a visual SLAMMOT method for fusing object planar features, comprising: Obtain an RGB-D image of the target scene, perform instance segmentation and normal estimation on the RGB-D image, and obtain instance segmentation results and normal estimation results; Extract static point features of the input image, filter outlier point features, estimate the initial camera pose by tracking static point features and minimizing reprojection errors, and build a local camera map to optimize the camera pose and static map point positions; The instance segmentation results are used to associate data of objects in different frames, track key points of associated objects and estimate their initial poses. The plane features of dynamic objects are extracted based on the normal estimation results. Through an association strategy based on matching features and plane parameters, a local optimization framework for objects that integrate plane constraints is constructed to optimize object poses and dynamic feature parameters. Combining the static processing results and the dynamic processing results, the output is a target environment map including camera pose, static map points, object pose, dynamic map points and dynamic plane.
[0009] According to a visual SLAMMOT method for fusing object plane features provided by the present invention, static point features of an input image are extracted, abnormal point features are filtered, and initial camera pose estimation is performed by tracking the static point features and minimizing the reprojection error, including: Extract Shi-Tomasi feature points from the input image, remove the feature points on the dynamic objects, and retain the feature points in the static background; The Lucas-Kanade optical flow method is used to track the extracted feature points frame by frame, and the corresponding relationship of key points is established between consecutive frames; After completing the key point tracking, a reprojection error model between the static map points and the feature points in the current frame is established through the matching relationship with the existing static map points :
[0010] Among them, is the pixel coordinate of the feature point corresponding to the static map point in the th frame, is the camera projection model, is the camera pose of the th frame, is the static map point coordinates in the world coordinate system; By minimizing the first objective function, the initial camera pose estimate of the th frame is obtained :
[0011] Among them, is the set of static map points corresponding to the feature points in the th frame, is the constant covariance matrix.
[0012] According to a visual SLAM MOT method for fusing object plane features provided by the present invention, a camera local map is constructed to optimize the camera pose and the position of static map points, including: Construct a sliding window optimization framework to locally optimize the initial camera pose and the position of static map points. Set the frames in the sliding window to include a fixed part of frames and a part of frames to be optimized. Among them, the fixed part of frames constrains the part of frames to be optimized, and the part of frames to be optimized is adjusted in the local optimization; Construct a second objective function in the sliding window optimization to optimize the camera pose and the position of static map points within the sliding window:
[0013] Among them, and respectively represent the set of frame indices and the set of static map point indices in the part of the sliding window to be optimized, is the set of all frame indices within the sliding window.
[0014] A visual SLAM MOT method integrating object plane features provided by the present invention uses the instance segmentation results to perform data association on objects in different frames, and performs key point tracking and initial object pose estimation on the associated objects, including: Calculate the intersection over union of the object instance segmentation mask of the current frame and the previous frame of the object, construct an association cost matrix, and use the KM algorithm to solve the association cost matrix to determine the optimal matching pair; Uniformly sample key points according to the size of the object, track the key points between adjacent frames by the Lucas-Kanade optical flow method, and verify the key points so that the key points are located within the matching object mask; Based on the correspondence between the object map points and the key points obtained from the object data association, construct an object map point reprojection error term :
[0015] where is the map point corresponding to the pixel coordinates of the key points in the frame, is the projection matrix of the camera, is the transformation matrix from the world coordinate system to the camera coordinate system of the frame, is the transformation matrix from the object coordinate system to the world coordinate system of the object at the frame, is the map point of the object in its object coordinate system;
[0016] where is the set of map point indices of the object associated with the key points in the frame, is a constant covariance matrix.
[0017] A visual SLAM MOT method integrating object plane features provided by the present invention extracts the plane features of dynamic objects based on the normal estimation results, including: Randomly select a key point from the key point set on the object surface as a seed point, and use the normal vector of the seed point as a benchmark, and gradually expand the area by comparing the similarity of the normal vectors. The normal vector of the current point to be expanded is , expand the seed point into a planar region if the following similarity conditions are met:
[0018] where, is the normal vector similarity threshold; During the process of region expansion, the planar normal vector is continuously updated according to the characteristics of the newly added region to adapt to the expanded planar feature region:
[0019] where, is the updated planar normal vector, is the planar normal vector before update, is the total number of planar points; Continue the process of region expansion until no more key points meet the similarity conditions; When the planar key point set and the corresponding planar normal vector are extracted, the final planar parameters are calculated as follows:
[0020] where, is the planar normal vector in the camera coordinate system, is the coordinate of the point cloud center corresponding to the planar key point set in the camera coordinate system, is the planar distance parameter in the camera coordinate system, obtaining the representation of the planar parameters in the camera coordinate system ; When the plane is first observed, initialize the plane by transforming it from the camera coordinate system to the object coordinate system:
[0021] where, and respectively represent the transformation matrices between the camera and the world coordinate system and between the object and the world coordinate system.
[0022] According to a visual SLAM MOT method for fusing object plane features provided by the present invention, by means of an association strategy based on matching features and plane parameters, construct a local optimization framework for the object that fuses plane constraints to optimize the object pose and dynamic feature parameters, including: After completing the extraction and association of object plane features, construct an error term from point to plane and an error term from plane to plane :
[0023]
[0024] Among them, is the object plane parameters in the object coordinate system, is the object map point coordinates in its object coordinate system, is the object plane parameters in the object coordinate system, is the transformation matrix from the world coordinate system to the frame camera coordinate system, is the object at the frame transformation matrix from the object coordinate system to the world coordinate system, is the object plane at the frame observation; Based on the smoothness of the object motion and the characteristic that the position does not change significantly, construct the object constant velocity motion model error term :
[0025] Among them, represents the mapping from the Lie group to the Lie algebra, represents the object at the frame to the frame motion transformation, is the object at the frame transformation matrix from the object coordinate system to the world coordinate system, is the object at the frame transformation matrix from the object coordinate system to the world coordinate system; Perform local BA optimization on the object pose, object map points, and object planes in the sliding window, and construct the overall objective function of object BA optimization:
[0026] Among them, are respectively the index sets of the camera frames, objects, object map points, and object planes of the part to be optimized in the sliding window, are respectively the covariance matrices corresponding to the object map point reprojection error term, object constant velocity motion model error term, point-to-plane error term, and plane-to-plane error term.
[0027] Thirdly, the present invention also provides a visual SLAM MOT device integrating object plane features, including: An acquisition module, configured to acquire RGB-D images of a target scene, perform instance segmentation and normal estimation on the RGB-D images respectively, and obtain an instance segmentation result and a normal estimation result; A static processing module, configured to extract static point features of an input image, filter abnormal point features, estimate an initial camera pose by tracking the static point features and minimizing a reprojection error, and construct a local camera map to optimize the camera pose and the positions of static map points; A dynamic processing module, configured to perform data association on objects in different frames by using the instance segmentation result, perform key point tracking and initial object pose estimation on the associated objects, extract planar features of dynamic objects based on the normal estimation result, and construct a local object optimization framework integrating planar constraints through an association strategy based on matching features and planar parameters to optimize the object pose and dynamic feature parameters; An integration module, configured to integrate the static processing result and the dynamic processing result, and output a target environment map including a camera pose, static map points, object poses, dynamic map points, and dynamic planes.
[0028] In a fourth aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the visual SLAMMOT method for fusing object planar features as described in any one of the above is implemented.
[0029] In a fifth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the visual SLAMMOT method for fusing object planar features as described in any one of the above is implemented.
[0030] The visual SLAMMOT system and method for fusing object planar features provided by the present invention add object planar features on the basis of point features commonly used in existing visual SLAMMOT systems, propose a method for extracting and matching object planar features based on normal vectors, determine a planar region through the consistency of adjacent point normal vectors, and perform planar matching by combining the point feature matching result and the similarity of planar parameters, so as to achieve stable and accurate extraction and association of multi-plane features in a multi-frame and multi-object scene; in addition, a method for optimizing the object pose by fusing object planar constraints is also proposed, and object map point reprojection, object constant velocity motion models, object point-to-plane and plane-to-plane constraints are incorporated into a factor graph optimization framework to improve the accuracy of object surface reconstruction and pose estimation; the solution of the present invention is simple to implement and highly practical, effectively solves the limitations of the prior art, improves the robustness and applicability of the system, and has important market application value. Description of the Drawings
[0031] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0032] Figure 1 It is one of the schematic flowcharts of the visual SLAMMOT method that fuses the planar features of objects provided by the present invention; Figure 2 It is the second schematic flowchart of the visual SLAMMOT method that fuses the planar features of objects provided by the present invention; Figure 3 It is the schematic diagram of the object plane constraint provided by the present invention; Figure 4 It is the object BA optimization factor graph provided by the present invention; Figure 5 It is the schematic diagram of the image and trajectory of the synthetic dataset provided by the present invention; Figure 6 It is the visualization result of the object plane extraction and association provided by the present invention; Figure 7 It is the schematic structural diagram of the visual SLAMMOT device that fuses the planar features of objects provided by the present invention; Figure 8 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0033] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0034] The embodiments of the present invention provide a visual SLAMMOT system and method that fuse the planar features of objects. The visual SLAMMOT system and method provided by the embodiments of the present invention can be integrated into an electronic device, which can be a terminal, a server, or other devices. Among them, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a micro-processing box, or other devices, etc.
[0035] A Visual SLAMMOT system and method integrating object plane features provided by an embodiment of the present invention can also be integrated on a special system related to the scenario of the solution. The present invention provides a Visual SLAMMOT system integrating object plane features, including: An RGB-D camera, the Isaac Sim simulation platform, and a computing device; The RGB-D camera collects RGB images and depth information of the target scene; The Isaac Sim simulation platform synthesizes the RGB image and the depth information to generate an RGB-D image; The computing device receives the RGB-D image, performs camera pose estimation, static mapping, object pose estimation, and dynamic reconstruction to generate an environmental map of the target scene.
[0036] Specifically, the hardware devices and their uses adopted in the embodiments of the present invention include: an RGB depth (RGB-Deepth, RGB-D) camera. The RGB-D sensor therein is mainly used to collect RGB images and depth information of the scene, providing visual and spatial data input for the Visual SLAMMOT algorithm. The RGB image can be used for feature extraction, instance segmentation, etc., and the depth information can be directly used to generate a three-dimensional point cloud, supporting scene mapping and object reconstruction.
[0037] In this embodiment, the device used in the experiment is an Intel RealSense D435 depth camera, and the NVIDIA Isaac Sim simulation platform is also used to generate high-quality synthetic images and depth data. A depth camera is a device that can capture the depth information of objects in a scene (i.e., the distance between the object and the camera). Different from traditional two-dimensional cameras, in addition to capturing the color and brightness of the image, the depth camera can also generate depth data about the distance of each pixel in the scene from the camera, that is, a depth map. The NVIDIA Isaac Sim simulation platform is a cloud-based robot simulation platform designed to accelerate the development and testing of robots through a high-fidelity simulation environment. Isaac Sim utilizes NVIDIA's GPU acceleration technology to provide real-time physical simulation and rendering, helping developers quickly iterate and optimize robot designs in a virtual environment.
[0038] The computing device is used to support the operation of the algorithm. Since Visual SLAMMOT involves computationally intensive tasks such as instance segmentation, normal estimation, and motion optimization, a high-performance computing device is required for deep learning inference and real-time data processing. In this embodiment, an Intel i7 processor and an NVIDIA RTX 3050 GPU are used in the experiment, which can meet the real-time processing requirements of medium-scale scenes.
[0039] It should be noted that the above hardware configuration is the implementation solution of the embodiment of the present invention, which can effectively support the operation and optimization of the visual SLAM MOT algorithm, and can be adjusted according to specific scenario requirements in practical applications.
[0040] Specifically, in the field of autonomous robots, the present invention can be applied to scenarios such as warehousing logistics and service robots. By utilizing the planar features of dynamic objects, the positioning and navigation capabilities of robots in complex environments can be improved, ensuring that they can stably track moving goods or other dynamic targets. In addition, this method can be used for dynamic target grasping tasks in industrial robots. By accurately tracking the plane of the target object, the grasping success rate and operation accuracy can be improved. In the field of autonomous driving, the present invention can be used to enhance the vehicle's perception ability of the dynamic environment and improve the stable tracking performance of dynamic targets such as vehicles. By fusing the planar information on the vehicle body, the accuracy of target trajectory prediction can be effectively improved, and the decision-making of the autonomous driving system can be optimized. In augmented reality (AR) and virtual reality (VR) applications, the present invention can be used to enhance the interactive experience and motion scene modeling. In AR / VR interaction devices, by utilizing the planar features of dynamic objects, the stable tracking ability of moving objects can be improved, and more accurate virtual-real combination can be achieved.
[0041] Figure 1 is one of the schematic flowcharts of the visual SLAM MOT method for fusing object planar features provided by the embodiment of the present invention. As Figure 1 shown, it includes: Step 100: Obtain the RGB-D image of the target scene, perform instance segmentation and normal estimation on the RGB-D image respectively, and obtain the instance segmentation result and the normal estimation result; Step 200: Extract the static point features of the input image, filter the abnormal point features, estimate the initial pose of the camera by tracking the static point features and minimizing the reprojection error, and construct a local camera map to optimize the camera pose and the position of the static map points; Step 300: Use the instance segmentation result to perform data association on the objects in different frames, perform key point tracking and initial pose estimation of the associated objects, extract the planar features of the dynamic objects based on the normal estimation result, and construct an object local optimization framework with fused plane constraints through an association strategy based on matching features and plane parameters to optimize the object pose and dynamic feature parameters; Step 400: Integrate the static processing results and the dynamic processing results, and output a target environment map including the camera pose, static map points, object pose, dynamic map points, and dynamic planes.
[0042] Specifically, as Figure 2As shown, the embodiment of the present invention takes RGB-D images as input, and respectively obtains instance segmentation and surface normal images for subsequent object tracking processing. The system mainly consists of two core modules: the camera pose estimation and static mapping module, and the object pose estimation and dynamic reconstruction module. In the camera pose estimation and static mapping module, first, point features are extracted from the input image, and potential abnormal points caused by dynamic objects are filtered to avoid introducing errors in camera pose estimation. Subsequently, initial pose estimation is performed through feature point tracking and minimization of reprojection error, and finally, local map optimization is used to further improve the accuracy of the camera pose. In the object pose estimation and dynamic reconstruction module, instance segmentation masks are used to associate objects in different frames, and key point tracking and initial pose estimation are performed on the successfully associated objects. On this basis, plane features of dynamic objects are extracted based on the normal image, and a stable tracking of the object plane features is achieved through an association strategy based on matching features and plane parameters. Subsequently, an optimization framework integrating plane constraints is constructed to further optimize the initial pose of the object and improve the pose estimation and reconstruction accuracy of dynamic targets. Finally, the system of the present invention can generate a high-precision environmental map including camera poses, static map points, object poses, dynamic map points, and dynamic planes, providing more stable and reliable mapping and positioning capabilities for applications such as robot autonomous navigation and augmented reality.
[0043] In one embodiment, as Figure 2 shown, the input of the embodiment of the present invention is RGB-D images, and the YOLACT algorithm and the Metric3D algorithm are respectively used to obtain the instance segmentation result and the normal estimation result. The subsequent implementation process is mainly divided into two modules: the camera pose estimation and static mapping module, and the object pose estimation and dynamic reconstruction module.
[0044] YOLACT (You Only Look At Coefficients) is a real-time instance segmentation algorithm that decomposes the instance segmentation task into two parallel subtasks: generating a set of prototype masks and predicting the mask coefficients of each instance, and then linearly combining the prototypes and the mask coefficients to generate the segmented instances. The Metric3D algorithm is a geometric-based model algorithm for zero-shot monocular depth and surface normal estimation, used to perform normal estimation of objects.
[0045] First, in the camera pose estimation and static mapping module, the present invention extracts Shi-Tomasi feature points from the input image. To improve the accuracy of camera pose estimation, the feature points on dynamic objects are removed, and only the feature points in the static background are retained. Then, the Lucas-Kanade (LK) optical flow method is used to track these feature points frame by frame, thereby establishing the correspondence of key points between consecutive frames. After completing the key point tracking, the present invention establishes a reprojection error model between the static map points and the feature points in the current frame through the matching relationship with the existing static map points. The specific expression is as follows:
[0046] where, is the pixel coordinate of the feature point corresponding to the static map point in the th frame, is the camera projection model, is the camera pose of the th frame, is the static map point coordinate in the world coordinate system.
[0047] By minimizing the following objective function, the camera pose of the th frame can be estimated:
[0048] where, is the set of static map points corresponding to the feature points in the th frame, is a constant covariance matrix.
[0049] To reduce the pose accumulation error and ensure the local consistency of static map points, the present invention constructs a sliding window optimization framework to locally optimize the camera pose and the position of static map points. The number of frames in the sliding window is fixed and divided into two parts: a fixed part and a part to be optimized. The fixed part contains earlier frames that only provide constraints; the part to be optimized contains newer frames that are mainly adjusted in local optimization. In the present invention, the number of frames in the sliding window is 20 frames, with 10 frames in each of the fixed part and the part to be optimized.
[0050] In the sliding window optimization, the present invention constructs the following objective function to optimize the camera pose and the position of static map points within the sliding window:
[0051] where, and respectively represent the set of frame indices and the set of static map point indices in the part to be optimized in the sliding window, It is the set of all frame indices within the sliding window.
[0052] Through the above embodiments, the present invention can accurately eliminate the influence of dynamic objects in a dynamic scene, retain the valid information in the static background, thereby achieving high-precision estimation of the camera pose, and ensuring the local consistency and optimization effect of static map points through sliding window optimization.
[0053] After completing the camera pose estimation and static reconstruction, data association of dynamic objects is performed to accurately match dynamic objects between adjacent frames to support the object pose estimation and 3D reconstruction in a dynamic scene. Specifically, first, the intersection over union (IoU) of the object instance segmentation masks of the current frame and the previous frame is calculated to construct an association cost matrix, and the KM algorithm is used to solve the cost matrix to determine the optimal matching pairs. To ensure the accuracy of object pose estimation and reconstruction, key points are uniformly sampled according to the size of the object, and the key points are tracked between adjacent frames by the LK optical flow method, and the key points are verified to ensure that they are located within the matching object mask, further improving the reliability of the association result.
[0054] When estimating the pose of a dynamic object, based on the correspondence between the object map points and the key points obtained from object data association, an object map point reprojection error term can be constructed:
[0055] where is the map point corresponding to the pixel coordinates of the key point in the th frame, is the projection matrix of the camera, is the transformation matrix from the world coordinate system to the camera coordinate system of the th frame, is the transformation matrix from the object coordinate system to the world coordinate system of the object at the th frame, is the coordinate of the map point of the object
[0056] Next, the initial pose of the object is estimated by minimizing the objective function:
[0057] where is the set of map point indices of the object associated with the key points in the th frame, is a constant covariance matrix.
[0058] After the initial pose estimation of the object is completed, a region growing algorithm based on normal vectors is used to extract plane features. The specific implementation steps are as follows: First, a key point is randomly selected from the key point set on the object surface as the seed point, and based on the normal vector of this point, the region is gradually expanded by comparing the similarity of normal vectors. Specifically, assuming the normal vector of the seed point is , and the normal vector of the current point to be expanded is , if the following similarity conditions are met, this point is expanded into a plane region:
[0059] where is the normal vector similarity threshold, which is set to 0.9 in the present invention.
[0060] During the process of region expansion, the plane normal vector is continuously updated according to the characteristics of the newly added region to accurately adapt to the expanded plane feature region:
[0061] where is the updated plane normal vector, is the plane normal vector before update, is the total number of plane points.
[0062] The process of region expansion continues until no more key points meet the similarity conditions. Through the above process, multiple plane features can be extracted from the target object.
[0063] When the plane key point set and the corresponding plane normal vector are extracted, the final plane parameters can be calculated in the following way:
[0064] where is the plane normal vector in the camera coordinate system, is the coordinate of the point cloud center corresponding to the plane key point set in the camera coordinate system, is the plane distance parameter in the camera coordinate system. Thus, the representation of the plane parameters in the camera coordinate system is obtained.
[0065] When the plane is first observed, it can be initialized by converting it from the camera coordinate system to the object coordinate system. This conversion is expressed as:
[0066] where and respectively represent the transformation matrices between the camera and the world coordinate system and between the object and the world coordinate system.
[0067] The present invention uses the KM algorithm to associate the planar features of an object. The cost function in this algorithm is constructed based on the number of matching key point pairs between the newly extracted plane and the existing object planes. After obtaining the preliminary association, the present invention further evaluates the similarity between the plane parameters to ensure the correctness of the matching and exclude the obviously incorrect associations. In practical applications, some newly extracted planes may not be able to establish an effective association with the existing planes. In such cases, the present invention uses the aforementioned transformation formula to initialize a new object plane and incorporate it into the subsequent plane feature matching and association process.
[0068] After the extraction and association of the object plane features are completed, the extracted object planes are used to construct plane constraints, specifically including point-to-plane constraints and plane-to-plane constraints, as Figure 3 shown. Considering the structural characteristics of the object surface, ideally, the distance between a point located on a plane and the plane should be zero, that is, satisfying . However, in practical applications, this ideal condition is often difficult to fully meet. To solve this problem, a point-to-plane error term is proposed to improve the accuracy of object surface reconstruction and pose estimation:
[0069] where, is the parameter of the plane of the object in the object coordinate system, is the map point of the object in its object coordinate system.
[0070] In addition, since the plane feature is extracted from multiple points, even if individual points are affected by noise, the plane feature can still maintain high stability, which makes the plane feature itself have strong anti-interference ability. Therefore, the inter-frame plane association relationship provides reliable information for object pose estimation. To make full use of this advantage, the present invention proposes a plane-to-plane error term, and the expression is as follows:
[0071] where, is the parameter of the plane of the object in the object coordinate system, is the transformation matrix from the world coordinate system to the camera coordinate system of the th frame, is the transformation matrix from the object coordinate system to the world coordinate system of the object at the th frame, is the plane of the object Observation at the frame.
[0072] The above error terms fully incorporate the structural characteristics of the object surface and utilize the stability of planar features to provide additional geometric constraints for object pose estimation. By minimizing the corresponding error function, not only can the accuracy of surface reconstruction be improved, but also the object pose estimation can be effectively optimized.
[0073] After completing the initial object pose estimation, the present invention further optimizes the object pose in the local map. This method separates the BA (Bundle Adjustment) optimization process of the camera and the object to reduce the mutual interference between the initial rough poses, thereby improving the computational efficiency of the overall optimization. During the optimization process, in addition to the reprojection error term of the object map points, an error term of the object constant velocity motion model, as well as point-to-plane and plane-to-plane error terms, are introduced. These error terms are uniformly constructed into the factor graph of the object BA, as Figure 4 shown.
[0074] In view of the smoothness of the object motion and the characteristic that there is no obvious mutation in the position, an error term of the object constant velocity motion model is constructed:
[0075] where represents the mapping from the Lie group to the Lie algebra, represents the object at the frame to the frame, is the transformation matrix from the object coordinate system to the world coordinate system of the object at the frame, is the transformation matrix from the object coordinate system to the world coordinate system of the object at the frame.
[0076] Finally, the present invention performs local BA optimization on the object pose, object map points, and object planes in the sliding window to improve the object pose estimation and reconstruction accuracy. The overall objective function of the object BA optimization is defined as follows:
[0077] where are the index sets of the camera frames, objects, object map points, and object planes of the part to be optimized in the sliding window, respectively, are the covariance matrices corresponding to the reprojection error term of the object map points, the error term of the object constant velocity motion model, the point-to-plane error term, and the plane-to-plane error term, respectively.
[0078] The system of the present invention can generate a high-precision environmental map including camera poses, static map points, object poses, dynamic map points, and dynamic planes, providing more stable and reliable mapping and positioning capabilities for applications such as robot autonomous navigation and augmented reality.
[0079] The visual SLAM MOT method and system integrating object plane features proposed by the present invention can effectively improve the accuracy of object pose estimation and reconstruction, which is mainly reflected in the following aspects.
[0080] First, to verify the effectiveness of the plane constraint method proposed in the present invention, the present invention uses the Isaac Sim platform to construct a warehouse environment and generates five experimental sequences: cube_3dof_cam_static, cube_3dof_cam_moving, cube_6dof_cam_static, cube_6dof_cam_moving, and box_cam_moving. The corresponding RGB images, depth images, and instance segmentation masks are generated through the renderer. The generated images and trajectory schematic diagrams are as Figure 5 shown.
[0081] In the experiment, the present invention extracts and associates the plane features of the object, and the results are as Figure 6 shown. The experimental results show that the extracted plane features can accurately capture the geometric structure characteristics of the object. In addition, by assigning a unique color to the object plane in each frame, it can be observed that in the case of the object performing six-degree-of-freedom motion, the association of the plane features remains stable and accurate between multiple frames.
[0082] For each experimental sequence, the present invention designs four groups of experimental modes: no plane constraint, only using point-to-plane constraint, only using plane-to-plane constraint, and using both plane constraints simultaneously. Since object local BA optimization does not affect the camera pose, the experiment focuses on comparing the object pose accuracy under different modes. The results are shown in Table 1. The experimental data show that even in the absence of plane constraints, the availability of point features can still ensure the basic functions of the system. However, after adding plane constraints to object local BA, the evaluation indexes of most experimental sequences have been improved. Further analysis shows that compared with the point-to-plane constraint, the plane-to-plane constraint has a more significant improvement in accuracy. The reason is that the plane-to-plane constraint aligns the object surface by adjusting the object pose, directly providing geometric measurement information; while the point-to-plane constraint indirectly adjusts the object pose by optimizing the position of object map points. In addition, plane constraints reduce the relative rotation error ( The effect in terms of
[0083] Table 1 Object pose estimation results under different object plane constraint modes on the synthetic dataset
[0084] In addition, Table 2 shows the comparison results with existing visual SLAMMOT methods in terms of camera and object pose estimation. The experimental results show that, compared with existing methods, especially the latest method SDPL-SLAM that uses line features, the present invention performs equivalently in terms of camera pose estimation, but shows significant advantages in terms of object pose estimation. Specifically, the relative translation error ( ) and relative rotation error ( ) of the object are reduced by 60.0% and 62.6% respectively. This significant improvement is attributed to two key factors. First, the present invention optimizes the position of object map points by using the plane characteristics of the object surface, thereby enhancing the accuracy of object pose estimation, especially in the case where there is noise in depth measurement, and this optimization effect is more obvious. Second, as a surface structure feature, the object plane feature is relatively stable and insensitive to noise, and its inter-frame correlation further provides additional useful information, thus improving the accuracy of object pose estimation.
[0085] Table 2 Comparison results with existing visual SLAMMOT methods in terms of camera and object pose estimation
[0086] Next, the visual SLAMMOT device integrating object plane features provided by the present invention will be described. The visual SLAMMOT device integrating object plane features described below can be mutually referred to the visual SLAMMOT method integrating object plane features described above.
[0087] Figure 7 is a schematic structural diagram of the visual SLAMMOT device integrating object plane features provided by an embodiment of the present invention. As shown in Figure 7 , it includes: an acquisition module 71, a static processing module 72, a dynamic processing module 73, and a comprehensive module 74, where: The acquisition module 71 is used to acquire the RGB-D image of the target scene, perform instance segmentation and normal estimation on the RGB-D image respectively, and obtain the instance segmentation result and the normal estimation result; the static processing module 72 is used to extract the static point features of the input image, filter out the abnormal point features, estimate the initial pose of the camera by tracking the static point features and minimizing the reprojection error, and construct a local camera map to optimize the camera pose and the position of the static map points; the dynamic processing module 73 is used to perform data association on the objects in different frames by using the instance segmentation result, track the key points and estimate the initial pose of the associated objects, extract the plane features of the dynamic objects based on the normal estimation result, and construct an object local optimization framework integrating plane constraints through an association strategy based on matching features and plane parameters to optimize the object pose and the dynamic feature parameters; the comprehensive module 74 is used to integrate the static processing result and the dynamic processing result, and output a target environment map including the camera pose, the static map points, the object pose, the dynamic map points and the dynamic planes.
[0088] Figure 8 An example of the physical structure diagram of an electronic device is shown as Figure 8 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the visual SLAM MOT method integrating object plane features, and this method includes: acquiring the RGB-D image of the target scene, performing instance segmentation and normal estimation on the RGB-D image respectively, and obtaining the instance segmentation result and the normal estimation result; extracting the static point features of the input image, filtering out the abnormal point features, estimating the initial pose of the camera by tracking the static point features and minimizing the reprojection error, and constructing a local camera map to optimize the camera pose and the position of the static map points; performing data association on the objects in different frames by using the instance segmentation result, tracking the key points and estimating the initial pose of the associated objects, extracting the plane features of the dynamic objects based on the normal estimation result, and constructing an object local optimization framework integrating plane constraints through an association strategy based on matching features and plane parameters to optimize the object pose and the dynamic feature parameters; integrating the static processing result and the dynamic processing result, and outputting a target environment map including the camera pose, the static map points, the object pose, the dynamic map points and the dynamic planes.
[0089] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, external hard drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0090] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the visual SLAM MOT method for fusing object plane features provided by the above-mentioned various methods. The method includes: acquiring an RGB-D image of a target scene, respectively performing instance segmentation and normal estimation on the RGB-D image to obtain an instance segmentation result and a normal estimation result; extracting static point features of the input image, filtering abnormal point features, estimating the initial pose of the camera by tracking the static point features and minimizing the reprojection error, and constructing a local camera map to optimize the camera pose and the positions of static map points; using the instance segmentation result to perform data association on objects in different frames, tracking key points of the associated objects and estimating the initial pose of the objects, extracting the plane features of dynamic objects based on the normal estimation result, and constructing an object local optimization framework that integrates plane constraints through an association strategy based on matching features and plane parameters to optimize the object pose and dynamic feature parameters; integrating the static processing result and the dynamic processing result, and outputting a target environment map including the camera pose, static map points, object poses, dynamic map points, and dynamic planes.
[0091] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.
[0092] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A visual SLAM MOT system integrating object plane features, characterized in that Including: RGB-D camera, IsaacSim simulation platform and computing device; The RGB-D camera acquires the RGB image and depth information of the target scene; The Isaac Sim simulation platform synthesizes the RGB image and depth information to generate an RGB-D image; The computing device receives the RGB-D image, performs camera pose estimation, static mapping, object pose estimation and dynamic reconstruction, and generates a target scene environment map.
2. A visual SLAM MOT method that fuses object plane features, which executes the visual SLAM MOT system for fusing object plane features described in claim 1, characterized in that Including: Obtain the RGB-D image of the target scene, perform instance segmentation and normal estimation on the RGB-D image respectively to obtain the instance segmentation result and normal estimation result; Extract the static point features of the input image, filter out the abnormal point features, estimate the initial camera pose by tracking the static point features and minimizing the reprojection error, and construct a local camera map to optimize the camera pose and the positions of the static map points; Use the instance segmentation result to perform data association on the objects in different frames, perform key point tracking and initial object pose estimation on the associated objects, extract the plane features of the dynamic objects based on the normal estimation result, and construct an object local optimization framework that integrates plane constraints to optimize the object pose and dynamic feature parameters; Integrate the static processing results and dynamic processing results, and output a target environment map including camera pose, static map points, object pose, dynamic map points and dynamic planes.
3. The visual SLAM MOT method for fusing object plane features according to claim 1, characterized in that Extract the static point features of the input image, filter out the abnormal point features, and estimate the initial camera pose by tracking the static point features and minimizing the reprojection error, including: Extract Shi-Tomasi feature points from the input image, remove the feature points on the dynamic objects, and retain the feature points in the static background; Use the Lucas-Kanade optical flow method to track the extracted feature points frame by frame and establish key point correspondence relationships between consecutive frames; After the key point tracking is completed, a reprojection error model between the static map points and the feature points in the current frame is established through the matching relationship with the existing static map points. : Among them, is the pixel coordinate of the feature point of the corresponding frame, is the camera projection model, is the camera pose of the frame, and is the coordinate of the static map point in the world coordinate system; By minimizing the first objective function, the initial camera pose estimation for the first frame is obtained : Among them, is the set of static map points corresponding to the feature points in the frame, and is the constant covariance matrix.
4. The visual SLAM MOT method for fusing planar features of an object according to claim 3, wherein Construct a local camera map to optimize the camera pose and the positions of the static map points, including: Construct a sliding window optimization framework to locally optimize the initial camera pose and the positions of the static map points. Set the frames in the sliding window to include a fixed part of frames and a part of frames to be optimized. Among them, the fixed part of frames constrains the part of frames to be optimized, and the part of frames to be optimized is adjusted in the local optimization; Construct a second objective function in the sliding window optimization to optimize the camera pose and the positions of the static map points within the sliding window: Among them, and respectively represent the set of frame indices and the set of static map point indices in the part of the sliding window to be optimized, is the set of all frame indices within the sliding window.
5. The visual SLAM MOT method for fusing object plane features according to claim 1, characterized in that Use the instance segmentation result to perform data association on the objects in different frames, perform key point tracking and initial object pose estimation on the associated objects, including: Calculate the intersection over union of the object instance segmentation masks in the current frame and the previous frame of the object, construct an association cost matrix, solve the association cost matrix using the KM algorithm, and determine the optimal matching pairs; Uniformly sample key points according to the size of the object, track the key points between adjacent frames by the Lucas-Kanade optical flow method, and verify the key points to make the key points located within the matching object mask; Based on the correspondence between the object map points and key points obtained by associating object data, construct the object map point reprojection error term : Among them, is the object corresponding to the map point of the pixel coordinates of the key points in the frame, is the projection matrix of the camera, is the transformation matrix from the world coordinate system to the camera coordinate system of the frame, is the transformation matrix from the object coordinate system to the world coordinate system of the object at the frame, is the map point of the object in its object coordinate system; Minimize the third objective function to obtain the initial pose estimate of the object : Among them, is the object associated with the key point in the th frame, and is the set of map point indices, and is the constant covariance matrix.
6. The visual SLAM MOT method for fusing object plane features according to claim 5, wherein Extract the plane features of the dynamic objects based on the normal estimation result, including: A key point is randomly selected from the key point set on the surface of the object as the seed point, and the normal vector of the seed point is As a benchmark, the area is gradually expanded by comparing the similarity of normal vectors. The normal vector of the current point to be expanded is , if the following similarity conditions are met, the seed point is expanded into a plane area: Among them, is the normal vector similarity threshold; During the process of region expansion, the plane normal vector is continuously updated according to the characteristics of the newly added region to adapt to the expanded plane feature region: Among them, is the updated plane normal vector, is the plane normal vector before update, is the total number of plane points; Continue the process of region expansion until no more key points meet the similarity conditions. When the planar key point set and the corresponding planar normal vector are extracted, the final planar parameters are calculated as follows: Among them, is the plane normal vector in the camera coordinate system, is the coordinate of the point cloud center corresponding to the plane key point set in the camera coordinate system, is the plane distance parameter in the camera coordinate system, and the representation of the plane parameter in the camera coordinate system is obtained ; When the plane is first observed, the plane is initialized by transforming it from the camera coordinate system to the object coordinate system: Among them, and respectively represent the transformation matrices between the camera and the world coordinate system, and between the object and the world coordinate system.
7. The visual SLAM MOT method for fusing object plane features according to claim 6, wherein Based on the association strategy between the matching features and the planar parameters, a local optimization framework of the object that incorporates planar constraints is constructed to optimize the object pose and dynamic feature parameters, including: After completing the extraction and association of the planar features of the object, construct the point-to-plane error term and the plane-to-plane error term : Among them, is the object plane parameters in the object coordinate system, is the object map point coordinates in its object coordinate system, is the object plane parameters in the object coordinate system, is the transformation matrix from the world coordinate system to the frame camera coordinate system, is the object at the frame transformation matrix from the object coordinate system to the world coordinate system, is the object plane at the frame observation; Based on the smoothness of object motion and the characteristic of no obvious sudden change in position, construct the error term of the constant-speed motion model of the object : Among them, represents the mapping from the Lie group to the Lie algebra, represents the object at the th frame to the th frame of the motion transformation, is the object at the th frame of the object coordinate system to the world coordinate system conversion matrix, is the object at the th frame of the object coordinate system to the world coordinate system conversion matrix; Performing local BA optimization on the object pose, object map points, and object planes in the sliding window to construct the overall objective function for object BA optimization: Among them, are respectively the index sets of the camera frame, object, object map point, and object plane in the part to be optimized in the sliding window. are respectively the covariance matrices corresponding to the object map point reprojection error term, object constant velocity motion model error term, point-to-plane error term, and plane-to-plane error term.
8. A visual SLAM MOT device integrating object plane features, characterized in that, Including: An acquisition module for acquiring the RGB-D image of the target scene, performing instance segmentation and normal estimation on the RGB-D image respectively to obtain the instance segmentation result and the normal estimation result; A static processing module for extracting the static point features of the input image, filtering out abnormal point features, estimating the initial camera pose by tracking the static point features and minimizing the reprojection error, and constructing a local map of the camera to optimize the camera pose and the positions of the static map points; A dynamic processing module for associating the objects in different frames using the instance segmentation result, tracking the key points of the associated objects and estimating the initial object pose, extracting the planar features of the dynamic objects based on the normal estimation result, and constructing a local optimization framework of the object that incorporates planar constraints to optimize the object pose and dynamic feature parameters based on the association strategy between the matching features and the planar parameters; A comprehensive module for integrating the static processing result and the dynamic processing result and outputting the target environment map including the camera pose, static map points, object pose, dynamic map points, and dynamic planes.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the visual SLAMMOT method for fusing object planar features as described in any one of claims 2 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the visual SLAMMOT method for fusing object planar features as described in any one of claims 2 to 7.
Citation Information
Patent Citations
Visual SLAM method based on multi-feature fusion
CN110060277A
Laser SLAM (Simultaneous Localization and Mapping) method fusing ground constraint and closed-loop constraint
CN118392171A
Cited By
Image target intelligent tracking and positioning method and system under Beidou position constraint
CN121559569A