Object-level environment reconstruction method based on hyper-quadric surface

Through modular design and hyperquadratic surface modeling based on ROS2 framework, combined with lightweight edge computing and dynamic data association, the real-time semantic modeling and high-precision positioning problems of robot autonomous navigation in complex scenarios are solved, and lightweight semantic perception and efficient pose estimation are realized.

CN120279095APending Publication Date: 2025-07-08NANJING UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510400480.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to realize real-time environment perception, semantic modeling and high-precision positioning of robot autonomous navigation in complex scenarios, especially in unstructured scenarios, object-level semantic information extraction and data association are poor, and edge-end deployment lacks optimization methods.

Method used

The modular design based on the ROS2 framework is adopted, combined with a monocular camera, YOLOv8 and NanoSAM for instance segmentation, and hyperquadratic surface modeling and optimization, dynamic data association and loopback detection are performed through improved KM algorithm to achieve lightweight edge computing.

Benefits of technology

Real-time high-precision object-level environment reconstruction of robots in complex scenarios, improve semantic perception ability and pose estimation accuracy, and ensure real-time processing capabilities at the edge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279095A_ABST
    Figure CN120279095A_ABST
Patent Text Reader

Abstract

The invention discloses an object-level environment reconstruction method based on a hyper-quadric surface, and the method specifically comprises the steps: constructing four core assemblies, namely a visual data collection module, a semantic preprocessing module, an object modeling and optimization module, and a data association and loopback detection module, based on an ROS2 frame and employing a modular design; the visual data acquisition module adopts a monocular camera to capture RGB image flow in real time and publish an original image, the semantic preprocessing module carries out semantic information extraction based on decoupling type instance segmentation, the object modeling and optimization module carries out super-quadratic surface initialization and outer point elimination, and then carries out super-quadratic surface optimization, and the object modeling and optimization module carries out super-quadratic surface optimization. Obtaining the precisely fitted geometric representation of the object; and the data association and loopback detection module performs dynamic data association and loopback detection enhancement. According to the method, the object modeling error is reduced, the pose estimation precision is improved, the environment perception capability is enhanced, the real-time processing capability can be maintained, and high-level semantic environment cognition support is provided for autonomous navigation of the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and robot autonomous navigation, and particularly to an object-level environment reconstruction method based on superquadrics. Background Art

[0002] Traditional visual SLAM technology takes geometric point clouds as the core. Although it can achieve environment reconstruction and positioning, it lacks the ability to express object semantics and structured information. In recent years, semantic SLAM extracts object-level semantic information by introducing deep learning models, but it still faces the following technical challenges:

[0003] 1. Limitations of geometric representation: Existing parametric models (such as cuboids, quadrics) are difficult to adapt to complex object shapes. Cuboids rely on vanishing point detection and are prone to failure in unstructured scenes; dual quadrics can only fit symmetric ellipsoids and cannot represent irregular objects (such as desks, chairs, mechanical parts).

[0004] 2. Insufficient real-time performance: Semantic segmentation models based on deep learning have high computational complexity and are difficult to achieve real-time inference on edge devices, resulting in difficulties in synchronizing semantic information extraction and the SLAM process.

[0005] 3. Poor robustness of data association: Existing object-level data association methods mostly rely on a single feature (such as appearance or geometry), and do not unify and model multi-source information such as category, shape, and spatial distribution, resulting in a high false matching rate.

[0006] 4. Lack of edge deployment: Superquadric modeling is mostly used for offline simulation, lacking optimization methods for monocular sparse point clouds and adaptation schemes for edge computing platforms (such as Jetson).

[0007] Among existing representative methods, CubeSLAM realizes object-level SLAM through cuboid modeling, but it depends on vanishing point detection and is difficult to handle unstructured scenes; QuadricSLAM represents objects using quadrics, but it is only applicable to symmetric ellipsoid-like objects; deep learning models can extract semantic information, but they are not designed for lightweight on mobile devices. In addition, existing data association methods (such as EAO-SLAM) do not fully integrate the mathematical characteristics of superquadrics, resulting in insufficient matching constraints. Therefore, there is an urgent need for an object-level environment reconstruction method that takes into account geometric representation ability, real-time performance, and edge deployment to achieve real-time environment perception, semantic modeling, and high-precision positioning of mobile robots in complex scenes. Summary of the Invention

[0008] The purpose of the present invention is to provide a robot autonomous navigation object-level environment reconstruction method with low object modeling error, high pose estimation accuracy, strong environment perception ability, strong real-time processing ability, and low latency.

[0009] The technical solution for achieving the object of the present invention is as follows: An object-level environment reconstruction method based on superquadrics, comprising the following steps:

[0010] Step 1, based on the ROS2 framework, using modular design, construct four core components: a visual data acquisition module, a semantic preprocessing module, an object modeling and optimization module, and a data association and loop detection module;

[0011] Step 2, the visual data acquisition module uses a monocular camera to capture RGB image streams in real time and publish the original images;

[0012] Step 3, the semantic preprocessing module performs semantic information extraction based on decoupled instance segmentation;

[0013] Step 4, the object modeling and optimization module performs superquadric initialization and outlier rejection, and then performs superquadric optimization to obtain an accurately fitted object geometric representation;

[0014] Step 5, the data association and loop detection module performs dynamic data association and loop detection enhancement.

[0015] Further, the visual data acquisition module described in step 2 uses a monocular camera to capture RGB image streams in real time and publish the original images, specifically as follows:

[0016] Step 2.1, the visual data acquisition module uses a monocular camera to capture RGB image streams in real time;

[0017] Step 2.2, publish the original images through a topic, and attach the timestamp and camera intrinsic information for downstream modules to subscribe.

[0018] Further, the semantic preprocessing module described in step 3 performs semantic information extraction based on decoupled instance segmentation, specifically as follows:

[0019] Step 3.1, the semantic preprocessing module constructs a lightweight instance segmentation framework through a separation of object detection and mask generation modules, uses YOLOv8 for object detection and outputs bounding boxes, and uses NanoSAM for pixel-level segmentation of the local area covered by the bounding boxes to generate object masks;

[0020] Step 3.2, in the model conversion stage, convert the segmentation model from the PyTorch format to the ONNX intermediate representation, use FP16 mixed-precision quantization and construct a TensorRT engine to achieve real-time inference on Jetson edge devices;

[0021] Step 3.3, after quantizing and accelerating the segmentation model through TensorRT, deploy it to edge computing devices.

[0022] Further, the object modeling and optimization module in step 4 performs superquadric surface initialization and outlier removal, and then performs superquadric surface optimization to obtain an accurately fitted object geometric representation, which is specifically as follows:

[0023] Step 4.1: The object modeling and optimization module performs superquadric surface parameter initialization based on the sparse point cloud of the object obtained by monocular SLAM, and calculates the initial size parameters based on the point cloud;

[0024] Step 4.2: Use the reprojection error of map points to remove the points outside the object mask, use the 3σ criterion to remove the outliers beyond the threshold range, perform DBSCAN clustering on the remaining point cloud, and retain the point set within the largest connected domain;

[0025] Step 4.3: Based on the object point cloud after outlier removal, construct a least squares optimization model using the differentiable mathematical characteristics of the superquadric surface, and obtain an accurately fitted object geometric representation by iteratively updating the shape parameters ε1, ε2, the center coordinate t, and the rotation matrix R.

[0026] Further, constructing the least squares optimization model using the differentiable mathematical characteristics of the superquadric surface in step 4.3 is specifically as follows:

[0027] Construct a non-linear least squares problem based on the inner and outer functions of the superquadric surface:

[0028]

[0029] where F is the improved inner and outer functions of the superquadric surface, ε1 is the shape parameter, p i represents the coordinates of the i-th 3D point, and the Levenberg-Marquardt algorithm is used for iterative optimization.

[0030] Further, the data association and loop detection module in step 5 performs dynamic data association and loop detection enhancement, which is specifically as follows:

[0031] Step 5.1: The data association and loop detection module models the matching of the currently detected object with the existing objects in the map as a bipartite graph optimal matching problem;

[0032] Step 5.2: Use the improved KM algorithm to synthesize the scores of four types of features, namely the total number of common map points, the object category confidence, the intersection over union of bounding boxes, and the inner and outer function values of the superquadric surface, to obtain a matching score function;

[0033] Step 5.3: Based on the successfully matched objects, construct an object-level co-visibility relationship graph, introduce object category consistency constraints and superquadric surface pose geometry verification in the traditional visual bag-of-words loop detection, generate loop candidates and perform global optimization.

[0034] Furthermore, the matching score function of the improved KM algorithm described in step 5.2 is defined as:

[0035]

[0036] where is the map point score, is the class confidence score, is the IoU score, is the superquadric score, k m , k c , k i , k s are learnable weight coefficients and satisfy k m + k c + k i + k s = 1.

[0037] Furthermore, the scores of the four types of features, namely the total number of map points, object class confidence, bounding box intersection over union, and superquadric inner and outer function values, described in step 5.2 are as follows:

[0038] (1) The total number of map points score is defined as:

[0039]

[0040] where represents the map points that are the same between the newly created object o i in the current key frame and the existing object o j in the map, N i represents all the map points observed on the object o j in the current key frame;

[0041] (2) The object class confidence score is defined as:

[0042]

[0043] where represents the class confidence;

[0044] (3) The bounding box intersection over union score, i.e., the IOU score, is defined as:

[0045]

[0046] where represents the bounding box of the object in the current key frame, represents the bounding box of the object in the previous key frame, represents the volume of the object in the current key frame, represents the volume of the object in the previous key frame;

[0047] (4) The score of the inner and outer functions of the superquadric surface is defined as:

[0048]

[0049] where is the number of object map points newly created in the current key frame among the existing objects in the map, and N i represents all the map points on the observed object o j in the current key frame.

[0050] Furthermore, the weight coefficients k m , k c , k i , k s are determined through ablation experiments, and the specific values are k m = 0.3, k c = 0.2, k i = 0.2, k s = 0.3.

[0051] Furthermore, for the object-level co-visibility relationship graph constructed based on the successfully matched objects in step 5.3, object category consistency constraints and superquadric surface pose geometric verification are introduced in the traditional visual bag-of-words loop detection to generate loop candidates and perform global optimization, as follows:

[0052] When the traditional bag-of-words model detects a loop candidate, extract the pairs of successfully matched objects between the candidate frames. When the number of consistent object pairs exceeds the threshold N obj , it is confirmed that the loop is established, where N obj = 4.

[0053] Compared with the prior art, the significant advantages of the present invention are as follows: (1) Lightweight semantic perception framework: An instance segmentation scheme that decouples detection and segmentation is adopted. YOLOv8 is used for object detection, and NanoSAM is used to generate masks, reducing the dependence on labeled data. The model is quantized to FP16 and graph-optimized through TensorRT to achieve real-time inference on Jetson edge devices; (2) Superquadric multi-stage modeling: For monocular sparse point clouds, an initialization method based on principal component analysis is adopted, combined with a three-modal outlier rejection strategy of reprojection error, 3σ filtering, and DBSCAN clustering, improving the point cloud quality, reducing the object modeling error, and enhancing the pose estimation accuracy; (3) Dynamic data association and loop closure enhancement: Object matching is modeled as a bipartite graph optimal matching problem, and a composite scoring function that fuses map points, categories, IoU, and superquadric features is defined. The KM algorithm is improved to achieve robust association, and object-level geometric verification is introduced in traditional bag-of-words loop closure detection, enhancing the success rate of loop closure detection in complex scenarios; (4) Edge system integration: Based on the ROS2 framework, modular design is implemented, and the algorithms are deployed to the Jetson AGX Orin platform to ensure the real-time performance of the entire process; (5) A complete technical chain from semantic perception to object modeling is constructed, providing high-level semantic environment perception support for robot autonomous navigation. Description of the Drawings

[0054] Figure 1 It is a schematic flowchart of a method for object-level environment reconstruction based on superquadrics according to the present invention.

[0055] Figure 2 It is a schematic communication flowchart of ROS2 nodes in the present invention. Detailed Embodiments

[0056] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0057] Combined with Figure 1 , a method for object-level environment reconstruction based on superquadrics according to the present invention includes the following steps:

[0058] Step 1: Based on the ROS2 framework, adopt modular design to construct four core components: a visual data acquisition module, a semantic preprocessing module, an object modeling and optimization module, and a data association and loop closure detection module;

[0059] Step 2: The visual data acquisition module uses a monocular camera to capture RGB image streams in real time and publishes the original images;

[0060] Step 3: The semantic preprocessing module performs semantic information extraction based on decoupled instance segmentation;

[0061] Step 4: The object modeling and optimization module performs superquadric surface initialization and outlier removal, and then performs superquadric surface optimization to obtain an accurately fitted object geometric representation;

[0062] Step 5: The data association and loop detection module performs dynamic data association and loop detection enhancement.

[0063] As a specific example, in Step 1, based on the ROS2 framework, a modular design is adopted to construct four core components: the visual data acquisition module, the semantic preprocessing module, the object modeling and optimization module, and the data association and loop detection module. Each module conducts data interaction through the distributed communication mechanism of ROS2 to ensure real-time performance and scalability; The ROS2 node communication is as Figure 2 shown.

[0064] As a specific example, in Step 2, the visual data acquisition module uses a monocular camera to capture the RGB image stream in real time and publishes the original image, specifically as follows:

[0065] Step 2.1: The visual data acquisition module uses a monocular camera to capture the RGB image stream in real time;

[0066] Step 2.2: Publish the original image through a topic, and attach the timestamp and camera intrinsic parameter information for downstream modules to subscribe.

[0067] As a specific example, in Step 3, the semantic preprocessing module performs semantic information extraction based on decoupled instance segmentation, specifically as follows:

[0068] Step 3.1: The semantic preprocessing module constructs a lightweight instance segmentation framework through a target detection and mask generation separation module, uses YOLOv8 for target detection and outputs the bounding box, and uses NanoSAM to perform pixel-level segmentation on the local area covered by the bounding box to generate the object mask;

[0069] Step 3.2: In the model conversion stage, convert the segmentation model from the PyTorch format to the ONNX intermediate representation, adopt FP16 mixed-precision quantization and construct a TensorRT engine to achieve real-time inference on the Jetson edge device;

[0070] Step 3.3: After quantizing and accelerating the segmentation model through TensorRT, deploy it to the edge computing device.

[0071] As a specific example, in Step 4, the object modeling and optimization module performs superquadric surface initialization and outlier removal, and then performs superquadric surface optimization to obtain an accurately fitted object geometric representation, specifically as follows:

[0072] Step 4.1: Based on the sparse point cloud of the object obtained by monocular SLAM, the object modeling and optimization module performs initialization of the superquadric surface parameters and calculates the initial size parameters based on the point cloud.

[0073] Step 4.2: Use the reprojection error of the map points to remove the points outside the object mask, use the 3σ criterion to remove the outlier points beyond the threshold range, and perform DBSCAN clustering on the remaining point cloud to retain the point set within the largest connected domain.

[0074] Step 4.3: Based on the object point cloud after removing the outlier points, use the differentiable mathematical characteristics of the superquadric surface to construct a least squares optimization model, and obtain the accurately fitted geometric representation of the object by iteratively updating the shape parameters ε1, ε2, the center coordinate t, and the rotation matrix R, as follows:

[0075] Construct a non-linear least squares problem based on the inner and outer functions of the superquadric surface:

[0076]

[0077] where F is the improved inner and outer functions of the superquadric surface, ε1 is the shape parameter, p i represents the coordinates of the i-th 3D point, and the Levenberg-Marquardt algorithm is used for iterative optimization.

[0078] As a specific example, in Step 5, the data association and loop detection module enhances the dynamic data association and loop detection, as follows:

[0079] Step 5.1: The data association and loop detection module models the matching between the detected objects in the current frame and the existing objects in the map as a bipartite graph optimal matching problem.

[0080] Step 5.2: Use the improved KM algorithm to comprehensively obtain the matching score function based on the scores of four types of features, namely the total number of map points, the object category confidence, the intersection over union (IoU) of the bounding boxes, and the values of the inner and outer functions of the superquadric surface, as follows:

[0081] The matching score function of the improved KM algorithm is defined as:

[0082]

[0083] where is the map point score, is the category confidence score, is the IoU score, is the superquadric surface score, k m ,k c ,k i ,k s are learnable weight coefficients and satisfy k m +kc +k i +k s = 1;

[0084] (1) The number of map points score is defined as:

[0085]

[0086] where represents the newly created object o in the current key frame i and the existing object o in the map j with the same map points, N i represents all the map points on the observed object o j in the current key frame;

[0087] (2) The object category confidence score is defined as:

[0088]

[0089] where represents the category confidence;

[0090] (3) The bounding box intersection over union score, IOU score, is defined as:

[0091]

[0092] where represents the bounding box of the object in the current key frame, represents the bounding box of the object in the previous key frame, represents the volume of the object in the current key frame, represents the volume of the object in the previous key frame;

[0093] (4) The score of the superquadric inner and outer function values is defined as:

[0094]

[0095] where is the number of map points of the newly created object in the current key frame among the existing objects in the map, N i represents all the map points on the observed object o j in the current key frame;

[0096] The weight coefficient k m , k c , k i , k s is determined by ablation experiments, and the specific values are k m = 0.3, k c = 0.2, k i = 0.2, k s= 0.3.

[0097] Step 5.3: Based on the successfully matched objects, construct an object-level co-visibility relationship graph, introduce object category consistency constraints and superquadric pose geometric verification in traditional visual bag-of-words loop detection, generate loop candidates and perform global optimization, specifically as follows:

[0098] When the traditional bag-of-words model detects a loop candidate, extract the pairs of successfully matched objects between the candidate frames. When the number of consistent object pairs exceeds the threshold N obj it is confirmed that the loop is established, where N obj = 4.

[0099] Embodiment 1

[0100] Combined with Figure 1 , a method for object-level environment reconstruction based on superquadrics provided in this embodiment includes the following steps:

[0101] Step 1: Based on the ROS2 framework, adopt a modular design to construct four core components: a visual data acquisition module, a semantic preprocessing module, an object modeling and optimization module, and a data association and loop detection module. Each module conducts data interaction through the distributed communication mechanism of ROS2 to ensure real-time performance and scalability; ROS2 node communication is as Figure 2 shown.

[0102] Step 2: The visual data acquisition module uses a monocular camera to capture RGB image streams in real time and publishes the original images, specifically as follows:

[0103] Step 2.1: Use an Intel RealSense D435i monocular camera to capture RGB image streams in real time through the ` / camera` node of ROS2;

[0104] Step 2.2: The original images are published through the topic ` / camera / color / image_raw` and are appended with timestamp and camera intrinsic information for downstream modules to subscribe.

[0105] Step 3: The semantic preprocessing module performs semantic information extraction based on decoupled instance segmentation, specifically as follows:

[0106] Step 3.1: The semantic preprocessing module constructs a lightweight instance segmentation framework through a target detection and mask generation separation module. Use YOLOv8 for target detection and output bounding boxes, and use NanoSAM to perform pixel-level segmentation on the local area covered by the bounding boxes to generate object masks, specifically as follows:

[0107] Model deployment: Integrate the YOLOv8 target detection and NanoSAM instance segmentation models in the ` / yolov8_node` node;

[0108] Object Detection: YOLOv8 performs multi-scale feature extraction on the input image, outputs object bounding boxes and class confidence scores, and the confidence score threshold ≥ 0.5;

[0109] Mask Generation: Input the bounding box coordinates into NanoSAM to generate pixel-level object masks for precise segmentation of the target area.

[0110] Step 3.2, In the model conversion stage, convert the segmentation model from PyTorch format to ONNX intermediate representation, adopt FP16 mixed-precision quantization, and build a TensorRT engine to achieve real-time inference on Jetson edge devices;

[0111] Step 3.3, After quantifying and accelerating the segmentation model through TensorRT, deploy it to the edge computing device, specifically as follows:

[0112] Model Acceleration: Optimize the model through the TensorRT engine:

[0113] (1) Format Conversion: Convert the PyTorch model to ONNX intermediate representation;

[0114] (2) Quantization Compression: Adopt FP16 mixed-precision quantization to reduce the model size and computational volume;

[0115] (3) Engine Construction: Generate an optimized TensorRT engine for the GPU architecture of Jetson AGX Orin;

[0116] Data Publishing: The processing results are published to the ` / detection_result` topic in a custom message format, including object labels, masks, and bounding boxes.

[0117] Step 4, The object modeling and optimization module performs superquadric initialization and outlier rejection, and then performs superquadric optimization to obtain an accurately fitted object geometric representation, specifically as follows:

[0118] Step 4.1, The object modeling and optimization module performs superquadric parameter initialization based on the sparse point cloud of the object obtained by monocular SLAM, and calculates the initial size parameters based on the point cloud, specifically as follows:

[0119] (1) Point Cloud Acquisition: Generate sparse map points based on ORB-SLAM3, and extract the point cloud belonging to the target object in combination with the instance segmentation mask;

[0120] (2) Parameter Initialization:

[0121] Set the initial shape parameters (ε1, ε2) according to the object semantic category.

[0122] Step 4.2: Use the reprojection error of map points to remove the points outside the object mask, and use the 3σ criterion to remove the outlier points beyond the threshold range. Perform DBSCAN clustering on the remaining point cloud and retain the point set within the largest connected domain, as follows:

[0123] Reprojection filtering: Reproject the point cloud onto the image plane and remove the mis-matched points outside the mask;

[0124] 3σ statistical filtering: Calculate the mean and standard deviation of the distance between the depth of the point cloud and the centroid, and remove the outlier points that deviate more than 3σ;

[0125] DBSCAN clustering: Set the neighborhood radius to 0.1 m and the minimum number of points to 5, and retain the point set within the largest connected domain.

[0126] Step 4.3: Based on the object point cloud after removing the outlier points, use the differentiable mathematical characteristics of the superquadric surface to construct a least squares optimization model. By iteratively updating the shape parameters ε1, ε2, the center coordinate t, and the rotation matrix R, obtain the accurate geometric representation of the object, as follows:

[0127] (1) Optimization modeling: Construct a non-linear least squares problem, and the objective function is:

[0128]

[0129] where F is the improved inner and outer functions of the superquadric surface, ε1 is the shape parameter, and p i represents the coordinates of the i-th 3D point;

[0130] (2) Optimization algorithm: Use the Levenberg-Marquardt algorithm to iteratively solve, and introduce the Huber kernel function to suppress the influence of depth outliers, where δ = 0.5;

[0131] (3) Parameter constraint: According to the semantic category, limit the optimization range of the shape parameters (such as for cube-like objects, ε1 ∈ [0.05, 0.3]) to avoid falling into local optima.

[0132] Step 5: The data association and loop detection module performs dynamic data association and loop detection enhancement, as follows:

[0133] Step 5.1: The data association and loop detection module models the matching between the detected objects in the current frame and the existing objects in the map as a bipartite graph optimal matching problem;

[0134] Step 5.2: Adopt the improved KM algorithm, and synthesize the scores of four types of features, namely the total number of common map points, the confidence of the object category, the intersection over union of the bounding boxes, and the values of the inner and outer functions of the superquadric surface, to obtain the matching score function, as follows:

[0135] The matching score function of the improved KM algorithm is defined as:

[0136]

[0137] where is the map point score, is the category confidence score, is the IoU score, is the superquadric score, k m , k c , k i , k s are learnable weight coefficients and satisfy k m + k c + k i + k s = 1;

[0138] (1) The score definition of the total number of map points is:

[0139]

[0140] where represents the same map points of the object o newly created in the current key frame i and the object o existing in the map j , N i represents all the map points on the object o observed in the current key frame j ;

[0141] (2) The score definition of the object category confidence is:

[0142]

[0143] where represents the category confidence;

[0144] (3) The score definition of the bounding box intersection over union (IOU score) is:

[0145]

[0146] where represents the bounding box of the object in the current key frame, represents the bounding box of the object in the previous key frame, represents the volume of the object in the current key frame, represents the volume of the object in the previous key frame;

[0147] (4) The score definition of the superquadric inner and outer function values is:

[0148]

[0149] where The number of object map points newly created for the current key frame among the existing objects in the map, N i Indicates all the map points on the object o observed in the current key frame j ;

[0150] The weight coefficient k m , k c , k i , k s Determined by ablation experiments, and the specific values are k m = 0.3, k c = 0.2, k i = 0.2, k s = 0.3.

[0151] Step 5.3: Based on the successfully matched objects, construct an object-level co-visibility relationship graph, introduce object category consistency constraints and superquadric pose geometric verification in the traditional visual bag-of-words loop detection, generate loop candidates and perform global optimization, specifically as follows:

[0152] (1) Candidate frame screening: On the basis of the traditional bag-of-words model, add object-level constraints, and there should be ≥ 4 groups of homogeneous object pairs between candidate frames;

[0153] (2) Global optimization: Add object pose constraints to the pose graph through the g2o library to optimize the camera trajectory and ground Figure 1 consistency.

[0154] The edge-side deployment and real-time guarantee of an object-level environment reconstruction method based on superquadrics provided in this embodiment are specifically as follows:

[0155] (1) ROS2 node communication:

[0156] Node splitting: Divided into three independent nodes, ` / camera`, ` / yolov8_node`, and ` / slam_node`, and communicate asynchronously through topics.

[0157] Message definition: Customize the `DetectionResult` message type, which includes object labels, masks, bounding boxes, and confidence levels.

[0158] (2) Performance verification: Achieve 30fps full-process processing on Jetson AGX Orin, and the latency of key modules is as follows:

[0159] Semantic preprocessing: ≤ 15ms / frame

[0160] Superquadric optimization: ≤ 8ms / object

[0161] Data association: ≤ 5ms / key frame

[0162] Example 2

[0163] The scenario of this example is the autonomous navigation of an indoor service robot. In a dynamic office environment, the service robot needs to continuously perceive objects such as desks, chairs, and electronic devices in real time, and construct a semantic map to support tasks such as autonomous navigation and item delivery.

[0164] Input: Real-time RGB image stream, including complex lighting changes and dynamic pedestrian interference;

[0165] Processing flow:

[0166] (1) Semantic information extraction: YOLOv8 quickly detects targets such as office desks, chairs, and display screens. NanoSAM generates a high-precision mask based on the detection box to distinguish the boundaries of overlapping objects;

[0167] (2) Object modeling: For regular objects such as filing cabinets, use superquadric surfaces to parameterize and represent their geometric shapes. Determine the initial orientation through principal component analysis, and combine an outlier rejection strategy to filter out the noisy point clouds generated by personnel movement;

[0168] (3) Dynamic scene adaptation: In scenarios such as corridor turning and multi-room switching, continuously update the poses of objects in the map based on object-level data association, and optimize the robot's motion trajectory in combination with semantic constraints;

[0169] (4) Loop closure enhancement: When the robot returns to the office area, use the matching object poses, such as printers and conference tables at fixed positions, to enhance loop closure detection and correct the cumulative positioning errors during long-term operation;

[0170] Output: A three-dimensional semantic map containing object semantic labels, geometric models, and passable areas, supporting robot obstacle avoidance, target object retrieval, and multi-room path planning.

[0171] Example 3

[0172] The scenario of this example is the intelligent inspection and maintenance of industrial equipment. In industrial scenarios such as petrochemical and power facilities, high-precision three-dimensional modeling of equipment such as pipelines and valves is required to support tasks such as corrosion detection and component life assessment.

[0173] Input: RGB image stream in a complex industrial environment, including low-texture metal surfaces and dense equipment occlusion;

[0174] Processing flow:

[0175] (1) Complex geometric representation: For equipment such as cylindrical pipelines and special-shaped flanges, use superquadric surfaces to adaptively fit their geometric features, optimize the model parameters through multi-view observations, and accurately represent the surface curvature and connection structure;

[0176] (2) Low-texture environment processing: On metal surfaces with no significant texture, utilize object semantic information, such as pipe diameter and valve type, to construct data association constraints to make up for the deficiencies of traditional feature matching;

[0177] (3) Abnormality detection support: Based on the geometric comparison between the superquadric surface model and historical data, automatically identify abnormal states such as pipe deformation and bolt loss, and generate a visual inspection report;

[0178] (4) Cross-period data association: In periodic inspections, associate device models of different periods through object-level matching, track the changing trend of corrosion degree, and provide a decision-making basis for preventive maintenance;

[0179] Output: A device-level high-precision 3D model library that supports digital twin, defect location, and automatic generation of maintenance work orders.

[0180] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An object-level environment reconstruction method based on superquadrics, characterized in that It includes the following steps: Step 1: Based on the ROS2 framework, adopt modular design to construct four core components: a visual data acquisition module, a semantic preprocessing module, an object modeling and optimization module, and a data association and loop detection module; Step 2: The visual data acquisition module uses a monocular camera to capture the RGB image stream in real time and publishes the original image; Step 3: The semantic preprocessing module performs semantic information extraction based on decoupled instance segmentation; Step 4: The object modeling and optimization module performs superquadric surface initialization and outlier rejection, and then performs superquadric surface optimization to obtain an accurately fitted object geometric representation; Step 5: The data association and loop detection module performs dynamic data association and loop detection enhancement.

2. The object-level environment reconstruction method based on superquadrics according to claim 1, wherein The visual data acquisition module described in Step 2 uses a monocular camera to capture the RGB image stream in real time and publishes the original image, specifically as follows: Step 2.1: The visual data acquisition module uses a monocular camera to capture the RGB image stream in real time; Step 2.2: Publish the original image through a topic, and attach the timestamp and camera intrinsic information for downstream modules to subscribe.

3. The object-level environment reconstruction method based on superquadrics according to claim 1, wherein, The semantic preprocessing module described in Step 3 performs semantic information extraction based on decoupled instance segmentation, specifically as follows: Step 3.1: The semantic preprocessing module constructs a lightweight instance segmentation framework through a target detection and mask generation separation module, uses YOLOv8 for target detection and outputs bounding boxes, and uses NanoSAM to perform pixel-level segmentation on the local area covered by the bounding boxes to generate object masks; Step 3.2: In the model conversion stage, convert the segmentation model from the PyTorch format to the ONNX intermediate representation, adopt FP16 mixed-precision quantization and construct a TensorRT engine to achieve real-time inference on the Jetson edge device; Step 3.3: After quantifying and accelerating the segmentation model through TensorRT, deploy it to the edge computing device.

4. The object-level environment reconstruction method based on superquadrics according to claim 1, characterized in that, The object modeling and optimization module described in Step 4 performs superquadric surface initialization and outlier rejection, and then performs superquadric surface optimization to obtain an accurately fitted object geometric representation, specifically as follows: Step 4.1: The object modeling and optimization module performs superquadric surface parameter initialization based on the sparse point cloud of the object obtained by monocular SLAM, and calculates the initial size parameters based on the point cloud; Step 4.2: Use the map point reprojection error to reject the points outside the object mask, adopt the 3σ criterion to reject the outliers beyond the threshold range, and perform DBSCAN clustering on the remaining point cloud to retain the point set within the largest connected domain; Step 4.3: Based on the object point cloud after outlier rejection, use the differentiable mathematical characteristics of the superquadric surface to construct a least squares optimization model, and obtain an accurately fitted object geometric representation by iteratively updating the shape parameters ε1, ε2, the center coordinate t, and the rotation matrix R.

5. The object-level environment reconstruction method based on superquadrics according to claim 4, wherein The construction of the least squares optimization model by using the differentiable mathematical characteristics of the superquadric surface described in Step 4.3 is specifically as follows: Construct a non-linear least squares problem based on the inner and outer functions of the superquadric surface: where F is the inner and outer function of the improved superquadric surface, ε1 is the shape parameter, and p i represents the coordinates of the i-th 3D point, and the Levenberg-Marquardt algorithm is used for iterative optimization.

6. The object-level environment reconstruction method based on superquadrics according to claim 1, wherein The data association and loop detection module described in Step 5 performs dynamic data association and loop detection enhancement, specifically as follows: Step 5.1: The data association and loop detection module models the matching between the detected objects in the current frame and the existing objects in the map as an optimal bipartite matching problem; Step 5.2: An improved KM algorithm is adopted to obtain a matching score function by synthesizing the scores of four types of features, namely the total number of map points, object category confidence, bounding box intersection over union (IOU), and the inner and outer function values of the superquadric surface; Step 5.3: Based on the successfully matched objects, an object-level co-visibility relationship graph is constructed. Object category consistency constraints and superquadric surface pose geometric verification are introduced into the traditional visual bag-of-words loop detection to generate loop candidates and perform global optimization.

7. The object-level environment reconstruction method based on superquadrics according to claim 6, characterized in that The matching score function of the improved KM algorithm described in Step 5.2 is defined as: where is the map point score, is the class confidence score, is the IoU score, is the superquadric score, k m , k c , k i , k s are learnable weight coefficients and satisfy k m + k c + k i + k s = 1.

8. The object-level environment reconstruction method based on superquadrics according to claim 6, characterized in that The scores of the four types of features, namely the total number of map points, object category confidence, bounding box intersection over union (IOU), and the inner and outer function values of the superquadric surface, described in Step 5.2 are specifically as follows: (1) The score of the total number of map points is defined as: Among them represents the newly created object o in the current key frame i and the existing object o in the map j at the same map point, N i represents all the map points observed on the object o in the current key frame j ; (2) The score of object category confidence is defined as: Among them indicates the class confidence; (3) The bounding box intersection over union (IOU) score is defined as: Among them represents the bounding box of the object in the current key frame, represents the bounding box of the object in the previous key frame, represents the volume of the object in the current key frame, represents the volume of the object in the previous key frame; (4) The score of the inner and outer function values of the superquadric surface is defined as: Among them is the number of object map points newly created for the current key frame among the existing objects in the map, N i represents all the map points on the object o observed in the current key frame j in the above.

9. The object-level environment reconstruction method based on superquadrics according to claim 8, wherein, The weight coefficient k m , k c , k i , k s is determined by ablation experiments, and the specific values are k m = 0.3, k c = 0.2, k i = 0.2, k s = 0.

3.

10. The object-level environment reconstruction method based on superquadrics according to claim 9, wherein Based on the successfully matched objects described in Step 5.3, an object-level co-visibility relationship graph is constructed. Object category consistency constraints and superquadric surface pose geometric verification are introduced into the traditional visual bag-of-words loop detection to generate loop candidates and perform global optimization, specifically as follows: When the traditional bag-of-words model detects a loop candidate, extract the object pairs that are successfully matched between the candidate frames. When the number of consistent object pairs exceeds the threshold N obj the loop is confirmed to be established, where N obj = 4.

Citation Information

Cited By

  • Multi-robot collaborative semantic SLAM and dynamic exploration method

    CN120765918A