Dynamic Voronoi skeleton constrained multi-modal door body detection and position correction method and dynamic Voronoi skeleton constrained multi-modal door body detection and position correction system
Through the multimodal detection method of dynamic Voronoi skeleton constraints, combined with RGBD cameras and multi-line lidar, the dynamic Voronoi skeleton model is constructed using improved YOLO network and point cloud data, and the door body detection error of single-modal sensors in dynamic environment is solved, achieving high-precision door body positioning and position correction.
Patent Information
- Application Number
- CN202511080216.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-08-04
AI Technical Summary
In the prior art, single-modal sensors have a high error detection rate in door body detection in light change, occlusion or complex geometric scenarios, and do not fully integrate semantic-geometric information, and lack methods to correct the position of the door body, resulting in insufficient detection accuracy in dynamic environments.
The multimodal detection method of dynamic Voronoi skeleton constraints is adopted, and data is obtained through RGBD cameras and multi-line lidar, and the gate body is coarsely positioned in combination with the improved YOLO deep learning network, point cloud data is fused for geometric verification, a dynamic Voronoi skeleton model is constructed, and the gate body position is corrected through the curvature constraint projection algorithm.
It significantly improves the robustness and navigation reliability of door body positioning in complex dynamic environments, solves the adaptability problem of traditional methods in dynamic environments, and realizes high-precision door body detection and position correction.
Smart Images

Figure CN120580294A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mobile robot environment perception and navigation, and in particular relates to a multimodal door detection and position correction method and system with dynamic Voronoi skeleton constraints. Background Art
[0002] In recent years, with the continuous development of computer technology, robotics, and electronic information technology, mobile robots have been widely used in various fields such as logistics, industry, agriculture, and services. Mobile robots such as indoor service robots, logistics robots, and security robots often need to identify and locate doors in their environment during autonomous exploration and navigation to better complete their tasks. However, current door detection and positioning methods still have the following technical shortcomings: Current methods often use single-modal detection, resulting in unreliable results. Traditional methods rely on a single sensor (such as lidar or RGB cameras), resulting in high false positive rates in scenes with varying lighting, occlusion, or complex geometry.
[0003] Current methods do not fully integrate semantic and geometric information. Existing methods do not deeply couple visual semantic detection (door frame recognition) with point cloud geometric verification (door open and closed status, precise location).
[0004] Current methods lack a method to correct the door position to further improve detection accuracy. This patent proposes using the environment's Voronoi skeleton to further correct the door position. At the same time, to address the problem of insufficient static skeleton constraints and inability to adapt to dynamic environmental changes (such as moving obstacles and temporarily stacked items), which leads to deviations in door position mapping, this patent proposes constructing a dynamic Voronoi skeleton environment representation model and designing a curvature-constrained projection algorithm to accurately map the detection results to the skeleton centerline, thereby correcting the detected door position and improving adaptability to dynamic environments. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a multimodal door detection and position correction method and system with dynamic Voronoi skeleton constraints, which significantly improves the robustness of door positioning and navigation reliability in complex dynamic environments.
[0006] On one hand, to achieve the above-mentioned purpose, the present invention provides a multimodal door detection and position correction method based on dynamic Voronoi skeleton constraints, comprising: Acquire sensor data through RGBD cameras and multi-line lidar; The improved YOLO deep learning network is used to process the RGB image to achieve rough door positioning and output the initial 3D coordinates of the door frame. Combine point cloud data fusion for 3D geometry verification, including ground point cloud elimination, door frame plane fitting, and point cloud density analysis to determine the door open or closed status; Build a dynamic Voronoi skeleton model and modify the environment topology in real time through an incremental distance field update algorithm; The detection results are mapped to the skeleton centerline using a curvature constrained projection algorithm; Output the coordinates of the door center on the two-dimensional grid map.
[0007] Optionally, the process of acquiring sensor data includes: Perform time synchronization and build a weighted least squares model to align the timestamps of the RGBD camera and the lidar. The model is: ; in, is the optimal synchronization time offset; For time; is the sensor weight coefficient of the RGBD camera; is the sensor weight coefficient of the lidar; 、 are the timestamp deviations of the RGBD camera and the lidar respectively; Perform spatial calibration, solve the homogeneous transformation matrix through the calibration plate, and unify the sensor coordinate system. The matrix is: ; in, is the homogeneous transformation matrix from the laser radar to the camera; is the 3D rotation matrix; is the translation vector; More accurate transformation matrices calculated for optimization methods; 、 are the corresponding point coordinates of the lidar and camera respectively; is the transformation matrix.
[0008] Optionally, the process of roughly positioning the door includes: The YOLO network is used to identify the door frame bounding box in the RGB image, and then processed by confidence filtering and non-maximum suppression; Combined with the depth map or lidar point cloud data, the two-dimensional pixel coordinates are converted into three-dimensional space coordinates. The conversion formula is: ; in, is the X coordinate in the camera coordinate system; is the pixel coordinate; is the column coordinate of the principal point of the image; is the focal length in the X-axis direction; is the depth value; is the Y coordinate in the camera coordinate system; is the focal length in the Y-axis direction; is the row coordinate of the principal point of the image.
[0009] Optionally, the three-dimensional geometry verification process includes: Perform ground point cloud removal, use the RANSAC algorithm to fit the ground plane and filter out interference points; Perform door frame plane fitting and use a robust fitting algorithm to construct a door frame plane model. The door frame plane model is: ; in, is the unit normal vector of the door frame plane; is the offset from the plane to the origin; A three-dimensional point in space ; is a point on the door frame plane; is the German-McClure robust kernel parameter; The door opening and closing status is analyzed based on the point cloud density, and the point cloud density inside the door frame is calculated to determine the status. The determination formula is: ; in, represents the set of door states; The door is in the closed state; The door is in the open state; is the point cloud density inside the door frame; is the point cloud density; for Other situations besides.
[0010] Optionally, the process of constructing the dynamic Voronoi skeleton model includes: Create a two-dimensional raster map based on lidar data; The environment topology skeleton network is dynamically generated by an incremental distance field update algorithm, and the skeleton points of the environment topology skeleton network meet the following conditions: ; in, is any point on the grid map; is the Voronoi skeleton point set; is the coordinate set of all obstacles in the environment; is any point on the grid map; is any point on the grid map; is the resolution of the raster map.
[0011] Optionally, the process of the curvature constrained projection algorithm includes: Project the initial portal coordinates to the Voronoi skeleton; Within the skeleton search radius, the optimal projection point of the low curvature area is found by combining the curvature optimization objective function. The objective function is: ; in, is the door center coordinate position after correction; is any point in the Voronoi skeleton point set; is the Voronoi skeleton point set; is the door center coordinate of the initial detection; is the weight coefficient; is the curvature.
[0012] Optionally, the process of outputting the door center coordinates includes mapping the corrected coordinates to a two-dimensional grid map and marking them.
[0013] On the other hand, to achieve the above-mentioned purpose, the present invention also provides a multimodal door detection and position correction system with dynamic Voronoi skeleton constraints, comprising: Multi-sensor spatiotemporal alignment module, used to synchronize the spatiotemporal data of RGBD cameras and multi-line lidar; The door rough detection module is used to identify the door frame in the RGB image through the improved YOLO network and output the initial 3D coordinates; Point cloud refinement processing module, used to fuse point cloud data to perform ground removal, door frame plane fitting, and open / close status determination; Dynamic Voronoi skeleton generation module, used to construct the environment topology skeleton in real time through incremental distance field update algorithm; Position correction module, used to map the detection results to the skeleton centerline through the curvature constrained projection algorithm; The coordinate output module is used to output the coordinates of the door center on the two-dimensional grid map.
[0014] Technical effects of the present invention: (1) Multimodal sensor fusion and spatiotemporal alignment technology: Through the weighted time synchronization model and joint calibration algorithm, the spatiotemporal data of RGBD camera and lidar are accurately aligned, solving the problems of timing misalignment and coordinate deviation in cross-modal data fusion, and improving the reliability of detection and consistency of environmental perception.
[0015] (2) Dynamic Voronoi skeleton generation and real-time update mechanism: Using an incremental distance field update method, the skeleton topology structure is dynamically adjusted in combination with the obstacle movement speed to ensure that the environmental representation reflects the latest status (such as moving obstacles and temporary piles) in real time, overcoming the defect that traditional static skeletons cannot adapt to dynamic scenes.
[0016] (3) Curvature-constrained topological projection optimization algorithm: A curvature penalty term is introduced into the skeleton mapping to avoid projection ambiguity in high-curvature areas such as corners and bifurcation points, ensuring that the center position of the door is always mapped to the safe center line of the passage area.
[0017] (4) Door opening and closing status verification mechanism: Through the sparsity of the point cloud inside the door frame (low density when closed), the door opening and closing status can be effectively judged. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings: Figure 1 Schematic diagram of the process of a multimodal door detection and position correction method with dynamic Voronoi skeleton constraints according to an embodiment of the present invention; Figure 2 Schematic diagram of the structure of a multimodal door detection and position correction system with dynamic Voronoi skeleton constraints according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the door recognition process of the robot during autonomous exploration according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a door recognition by a robot when exploring an actual environment according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0020] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0021] like Figure 1 As shown, this embodiment provides a multimodal door detection and position correction method based on dynamic Voronoi skeleton constraints, including: Acquire sensor data through RGBD cameras and multi-line lidar; The improved YOLO deep learning network is used to process the RGB image to achieve rough door positioning and output the initial 3D coordinates of the door frame. Combine point cloud data fusion for 3D geometry verification, including ground point cloud elimination, door frame plane fitting, and point cloud density analysis to determine the door open or closed status; Build a dynamic Voronoi skeleton model and modify the environment topology in real time through an incremental distance field update algorithm; The detection results are mapped to the skeleton centerline using a curvature constrained projection algorithm; Output the coordinates of the door center on the two-dimensional grid map.
[0022] Figure 1 The patent flowchart shows that it starts from the multi-sensor data synchronization module, unifies the spatiotemporal reference of the RGBD camera and lidar through timestamp alignment and spatial calibration, and inputs the synchronized RGB image into the door body coarse detection module. The door frame bounding box and confidence level are output using the YOLO algorithm. At the same time, the point cloud data enters the refined processing module for ground removal, plane modeling and density analysis, extracts the three-dimensional parameters of the door frame and determines the open and closed status; the dynamic Voronoi skeleton generation module builds an environmental topology network based on the real-time point cloud, and adapts to obstacle changes through incremental distance field updates. The position correction module combines the initial detection coordinates with the skeleton topology for curvature optimization projection, avoids high curvature areas, and finally outputs the coordinates of the door center on the two-dimensional grid map, forming a complete closed-loop process from data acquisition, feature extraction, environmental modeling to error suppression.
[0023] Furthermore, RGBD cameras and LiDARs have different hardware characteristics. The former captures color images and depth information at high frequency, while the latter scans the environment's geometry in the form of a point cloud. First, a timestamp-weighted calibration is performed to ensure that both sensors collect data at the same moment. Then, a spatial coordinate transformation is performed to align the camera's image coordinate system and the LiDAR's 3D point cloud coordinate system to the robot's coordinate system. This process eliminates data bias caused by hardware differences, providing accurate and consistent input for subsequent processing.
[0024] In the time synchronization phase, a weighted least squares synchronization model was established to achieve spatiotemporal alignment of multi-sensor data. This model eliminates timing deviations between the RGBD camera and the LiDAR due to hardware clock discrepancies. The data acquisition times of the two sensors were aligned through weighted least squares optimization.
[0025] ; in, is the optimal synchronization time offset (unit: seconds), Value , Value Is the sensor weight coefficient, the weight coefficient can be set according to the sensor sampling rate, , It is the timestamp deviation between the RGBD camera and the lidar.
[0026] In the spatial calibration phase, the coordinate transformation matrix between sensors is solved through joint calibration with a calibration plate, thereby solving the problem of unifying the sensor coordinate system and solving the spatial alignment problem of cross-modal data (images and point clouds).
[0027] ; in, is the homogeneous transformation matrix from the laser radar to the camera; is the 3D rotation matrix; is the translation vector; More accurate transformation matrices calculated for optimization methods; 、 are the corresponding point coordinates of the lidar and camera respectively; is the transformation matrix.
[0028] Furthermore, the YOLO-based door detection and point cloud mapping positioning technology mainly includes the following steps: First, the RGB image is detected in real time through a pre-trained YOLO model (such as YOLOv5) to identify the 2D bounding box of the door frame. The model input is The output includes the center pixel coordinates, width, height and confidence of the bounding box. The detection results need to be filtered by confidence (threshold 0.7) and non-maximum suppression (NMS, loU threshold 0.45) to eliminate false detections and overlapping frames. Subsequently, combined with the depth map of the RGBD camera or the lidar point cloud data, the two-dimensional pixel coordinates are converted into three-dimensional space coordinates. Specifically, the depth value of the corresponding point in the depth map is used to calculate the depth of the camera through the intrinsic parameters (focal length and main point ) Calculate the three-dimensional position of the door frame center using the formula: , in, is the X coordinate in the camera coordinate system; is the pixel coordinate; is the column coordinate of the principal point of the image; is the focal length in the X-axis direction; is the depth value; is the Y coordinate in the camera coordinate system; is the focal length in the Y-axis direction; is the row coordinate of the principal point of the image.
[0029] For lidar data, the coordinates in the camera coordinate system need to be converted to the radar coordinate system through a pre-calibrated extrinsic matrix.
[0030] Furthermore, the LiDAR point cloud is processed in multiple stages to verify and optimize the rough detection results. First, the intelligent filtering algorithm is used to remove ground reflection points (such as floors and carpets) to avoid interference. Among them, the ground point cloud is removed using ; in, is the ground plane normal vector (fitted by RANSAC), is the ground plane equation offset, Represents a collection of raw point clouds.
[0031] Next, we extract point cloud clusters near the door frame location provided by the coarse detection. A robust fitting algorithm is used to construct a door frame plane model, calculating the precise 3D dimensions and normal orientation of the door. Robust RANSAC is used to fit the door frame plane, suppressing the influence of outliers and accurately fitting the door frame plane.
[0032] ; in, represents the door frame plane unit normal vector, is the offset from the plane to the origin, represents the German-McClure robust kernel parameter, A three-dimensional point in space : is a point on the door frame plane.
[0033] Finally, by analyzing the density and distribution of the point cloud inside the door frame, we can determine whether the door is open or closed. If the door is closed, the point cloud will be blocked by the door panel and concentrated. If the door is open, the point cloud can penetrate the door frame area, and the density is significantly reduced. The door state is determined by the point cloud density and the direction of the normal vector.
[0034] ; in, Indicates the point cloud density inside the door frame, is the number of point clouds in the valid area. Indicates the width and height of the door frame (unit: meter), where the height and height of the door frame are related to the point cloud density , represents the set of door states; The door is in the closed state; The door is in the open state; for Other circumstances other than It can be obtained through pre-measurement or testing to facilitate accurate identification of the door and judgment of the door's open and closed state.
[0035] Furthermore, a real-time topological map of the environment is constructed to provide geometric constraints for position correction. Based on a two-dimensional grid map created from lidar data, a dynamic wavefront algorithm is used to locally update the grid affected by obstacle changes. A single-pixel-wide Voronoi diagram is generated by combining conditional comparison and skeletonization methods. Using integer square distance optimization, an incremental update algorithm dynamically generates a Voronoi skeleton, a network structure that reflects the centerline of the environment's traffic area. Skeleton points are located near the equidistant boundary between two obstacles. Skeleton points are determined as follows: ; in, is any point on the grid map; is the Voronoi skeleton point set; is the coordinate set of all obstacles in the environment; is any point on the grid map; is any point on the grid map; is the resolution of the raster map. Considering that points on the Voronoi skeleton are at the same distance from nearby obstacles, the center of a door is usually at the same distance from the door frames on both sides and is also located at or very close to the environment's Voronoi skeleton. Therefore, we propose to use the Voronoi skeleton to correct the detected door position.
[0036] Furthermore, multi-source information is integrated to achieve high-precision positioning. First, the door frame coordinates, processed from the point cloud, are projected onto the Voronoi skeleton. A curvature optimization algorithm is used to avoid bifurcations or corners in the skeleton, ensuring that the projected point is located on the straight part of the Voronoi skeleton. Finally, the optimized door center point is mapped onto a two-dimensional grid map.
[0037] During the robot's movement, the camera outputs image data at a high frequency, and the robot observes the same door multiple times. The door detection module continuously returns the door position, and the door position in the map coordinate system is recorded as , ,in Indicates the number of times the door has been observed. At the same time, due to the large uncertainty of the robot's positioning information during the mapping process, there is often a deviation between the positions of the door returned multiple times. Therefore, the door position output by the door detection module is stored and its center of mass is found using the following formula. The center of mass position is used as the initial value of the door. In order to further correct the position of the door, the generalized Voronoi diagram is used. Since the generalized Voronoi diagram stores points that are equidistant from the obstacle, the center line of the door is generally included in it. In order to improve the accuracy of the correction and also increase the search speed, the edge branches on the generalized Voronoi diagram that are directly in contact with the obstacle are removed. Taking into account the curvature Reflects the local bending degree of the skeleton. :The frame is straight, usually the center of the door also appears here, which is the ideal projection area; when : The skeleton needs to avoid sharp turns or bifurcations. Therefore, it is proposed to search for the nearest point along the skeleton and combine it with curvature constraint optimization. First, direct nearest point search: only use Euclidean distance as an indicator to find the nearest point on the Voronoi skeleton with the initial coordinates. The nearest point, based on the distance, introduces the local geometric features (curvature) of the skeleton as a penalty term, and the optimization goal is: ; in, is the weight coefficient, is the door center coordinate of the initial detection, is the corrected door center coordinate position, is the curvature; is any point in the Voronoi skeleton point set; is the Voronoi skeleton point set; About weight coefficient choice, , can be adjusted Controlling curvature affects strength; degenerates to the direct closest point; The door center is projected onto the low-curvature skeleton segment to avoid topological ambiguity.
[0038] The curvature is calculated as follows: ; in, and are the first-order and second-order derivatives of the skeleton curve respectively; the curvature calculation is optimized using discretized skeleton points Approximate derivatives are obtained by differencing at adjacent points: ; ; in, is the skeleton point spacing.
[0039] ; in, The set of skeleton points for finding the center of the door on the Voronoi skeleton. is the door center coordinate of the initial detection, The corrected door center coordinates. Search within a radius of r = 1.0m to avoid global traversal and finally output the corrected door center position And marked on the two-dimensional raster map.
[0040] like Figure 2 As shown, this embodiment provides a multimodal door detection and position correction system with dynamic Voronoi skeleton constraints, including: Multi-sensor spatiotemporal alignment module, used to synchronize the spatiotemporal data of RGBD cameras and multi-line lidar; The door rough detection module is used to identify the door frame in the RGB image through the improved YOLO network and output the initial 3D coordinates; Point cloud refinement processing module, used to fuse point cloud data to perform ground removal, door frame plane fitting, and open / close status determination; Dynamic Voronoi skeleton generation module, used to construct the environment topology skeleton in real time through incremental distance field update algorithm; Position correction module, used to map the detection results to the skeleton centerline through the curvature constrained projection algorithm; The coordinate output module is used to output the coordinates of the door center on the two-dimensional grid map.
[0041] In this embodiment, each module forms a closed-loop processing flow, significantly improving the robustness and navigation reliability of door positioning in complex dynamic environments. In this embodiment, each module has a clear division of labor and is closely linked to form a complete closed loop from environmental perception to precise positioning.
[0042] like Figure 3 As shown in the figure, the method proposed in this example is used to identify doors during the autonomous exploration process of a robot equipped with a laser radar and an RGBD camera. Figure 3 On the left, you can see that there are 9 doors in the simulation environment. During the robot's autonomous exploration, it continuously recognizes the doors and uses Figure 3 The purple square on the right represents the projection of the center position of the door on the grid map. It can be seen that the robot can finally fully identify the nine doors and correctly mark their locations.
[0043] Figure 4 is the recognition of the door when the robot explores the actual environment, Figure 4 The left side shows the door in the actual environment, and the right side shows the marking of the door position during autonomous exploration. It can be seen that the robot can accurately mark the actual door position in the grid map, which also verifies the effectiveness of the method proposed in this embodiment.
[0044] The technical advantages of the present invention are: (1) Multimodal sensor fusion and spatiotemporal alignment technology: Through the weighted time synchronization model and joint calibration algorithm, the spatiotemporal data of RGBD camera and lidar are accurately aligned, solving the problems of timing misalignment and coordinate deviation in cross-modal data fusion, and improving the reliability of detection and consistency of environmental perception.
[0045] (2) Dynamic Voronoi skeleton generation and real-time update mechanism: Using an incremental distance field update method, the skeleton topology structure is dynamically adjusted in combination with the obstacle movement speed to ensure that the environmental representation reflects the latest status (such as moving obstacles and temporary piles) in real time, overcoming the defect that traditional static skeletons cannot adapt to dynamic scenes.
[0046] (3) Curvature-constrained topological projection optimization algorithm: A curvature penalty term is introduced into the skeleton mapping to avoid projection ambiguity in high-curvature areas such as corners and bifurcation points, ensuring that the center position of the door is always mapped to the safe center line of the passage area.
[0047] (4) Door opening and closing status verification mechanism: Through the sparsity of the point cloud inside the door frame (low density when closed), the door opening and closing status can be effectively judged.
[0048] Through the above innovations, the present invention overcomes the bottlenecks of traditional solutions in terms of dynamic environment adaptability, multi-sensor collaboration, and complex door status determination, providing mobile robots with high-precision and robust door positioning capabilities. It can be widely used in indoor autonomous navigation scenarios such as warehousing and logistics, and medical services. The module decoupling design: each module can be optimized independently (such as replacing YOLO with other detectors) to improve system flexibility; dynamic adaptability: the Voronoi skeleton is updated in real time to support moving obstacles and temporary environmental changes; precision-efficiency balance: a hierarchical strategy of coarse detection (YOLO) and fine processing (point cloud) takes into account both real-time and accuracy.
[0049] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A multimodal door detection and position correction method based on dynamic Voronoi skeleton constraints, characterized by: include: Acquire sensor data through RGBD cameras and multi-line lidar; The improved YOLO deep learning network is used to process the RGB image to achieve rough door positioning and output the initial 3D coordinates of the door frame. Combine point cloud data fusion for 3D geometry verification, including ground point cloud elimination, door frame plane fitting, and point cloud density analysis to determine the door open or closed status; Build a dynamic Voronoi skeleton model and modify the environment topology in real time through an incremental distance field update algorithm; The detection results are mapped to the skeleton centerline using a curvature constrained projection algorithm; Output the coordinates of the door center on the two-dimensional grid map.
2. The multimodal door detection and position correction method based on dynamic Voronoi skeleton constraints as claimed in claim 1, characterized in that: The process of acquiring sensor data includes: Perform time synchronization and build a weighted least squares model to align the timestamps of the RGBD camera and the lidar. The model is: ; in, is the optimal synchronization time offset; For time; is the sensor weight coefficient of the RGBD camera; is the sensor weight coefficient of the lidar; 、 are the timestamp deviations of the RGBD camera and the lidar respectively; Perform spatial calibration, solve the homogeneous transformation matrix through the calibration plate, and unify the sensor coordinate system. The matrix is: ; in, is the homogeneous transformation matrix from the laser radar to the camera; is the 3D rotation matrix; is the translation vector; More accurate transformation matrices calculated for optimization methods; 、 are the corresponding point coordinates of the lidar and camera respectively; is the transformation matrix.
3. The multimodal door detection and position correction method based on dynamic Voronoi skeleton constraints as claimed in claim 1, characterized in that: The process of door body rough positioning includes: The YOLO network is used to identify the door frame bounding box in the RGB image, and then processed by confidence filtering and non-maximum suppression; Combined with the depth map or lidar point cloud data, the two-dimensional pixel coordinates are converted into three-dimensional space coordinates. The conversion formula is: ; in, is the X coordinate in the camera coordinate system; is the pixel coordinate; is the column coordinate of the principal point of the image; is the focal length in the X-axis direction; is the depth value; is the Y coordinate in the camera coordinate system; is the focal length in the Y-axis direction; is the row coordinate of the principal point of the image.
4. The multimodal door detection and position correction method based on dynamic Voronoi skeleton constraints as claimed in claim 1, characterized in that: The process of 3D geometry verification includes: Perform ground point cloud removal, use the RANSAC algorithm to fit the ground plane and filter out interference points; Perform door frame plane fitting and use a robust fitting algorithm to construct a door frame plane model. The door frame plane model is: ; in, is the unit normal vector of the door frame plane; is the offset from the plane to the origin; A three-dimensional point in space ; is a point on the door frame plane; is the German-McClure robust kernel parameter; The door opening and closing status is analyzed based on the point cloud density, and the point cloud density inside the door frame is calculated to determine the status. The determination formula is: ; in, represents the set of door states; The door is in the closed state; The door is in the open state; is the point cloud density inside the door frame; is the point cloud density; for Other situations besides.
5. The multimodal door detection and position correction method based on dynamic Voronoi skeleton constraints as claimed in claim 1, characterized in that: The process of constructing the dynamic Voronoi skeleton model includes: Create a two-dimensional raster map based on lidar data; The environment topology skeleton network is dynamically generated by an incremental distance field update algorithm, and the skeleton points of the environment topology skeleton network meet the following conditions: ; in, is any point on the grid map; is the Voronoi skeleton point set; is the coordinate set of all obstacles in the environment; is any point on the grid map; is any point on the grid map; is the resolution of the raster map.
6. The multimodal door detection and position correction method based on dynamic Voronoi skeleton constraints as claimed in claim 1, characterized in that: The process of the curvature constrained projection algorithm includes: Project the initial portal coordinates to the Voronoi skeleton; Within the skeleton search radius, the optimal projection point of the low curvature area is found by combining the curvature optimization objective function. The objective function is: ; in, is the door center coordinate position after correction; is any point in the Voronoi skeleton point set; is the Voronoi skeleton point set; is the door center coordinate of the initial detection; is the weight coefficient; is the curvature.
7. The multimodal door detection and position correction method based on dynamic Voronoi skeleton constraints as claimed in claim 1, characterized in that: The process of outputting the door center coordinates includes mapping the corrected coordinates to a two-dimensional grid map and marking them.
8. The system of the multimodal door detection and position correction method based on dynamic Voronoi skeleton constraints according to any one of claims 1 to 7, characterized in that: include: Multi-sensor spatiotemporal alignment module, used to synchronize the spatiotemporal data of RGBD cameras and multi-line lidar; The door rough detection module is used to identify the door frame in the RGB image through the improved YOLO network and output the initial 3D coordinates; Point cloud refinement processing module, used to fuse point cloud data to perform ground removal, door frame plane fitting, and open / close status determination; Dynamic Voronoi skeleton generation module, used to construct the environment topology skeleton in real time through incremental distance field update algorithm; Position correction module, used to map the detection results to the skeleton centerline through the curvature constrained projection algorithm; The coordinate output module is used to output the coordinates of the door center on the two-dimensional grid map.
Citation Information
Patent Citations
Point cloud simplification method for maintaining geometrical characteristics based on absolute Gaussian curvature estimation
CN112184869A
Dynamic environment positioning and mapping method for mobile service robot based on deep clustering
CN116758260A
Depth camera and laser radar fused three-dimensional dense point cloud mapping method and system
CN119048600A
News scene three-dimensional reconstruction and visualization method based on multi-source remote sensing data
CN119904592A
Human body skeleton point positioning identification method and system based on OpenPose
CN120411545A
Cited By
Method and device for monitoring opening and closing states of three-column horizontal rotary disconnecting switch
CN121305468A