A method for monitoring the distance of key targets using mobile lidar and multi-camera collaboration
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-24
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]针对现有监控感知方法在三维表达不足、跨摄像头空间统一困难及图像与点云融合能力有限等问题,本发明提供一种移动激光雷达与多相机协同的关键目标距离监管方法,实现目标场地三维重建、动态监管与风险评估
[0014] (1) This invention combines the LiDAR scanning of the cruise robot with the images of the monitoring camera to complete the 3D point cloud mapping, external parameter estimation and spatial unification of the scene, thus solving the problems of insufficient spatial perception and difficulty in unifying camera coordinates in traditional monitoring systems.
Smart Images

Figure CN122546239A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for monitoring the distance to key targets in collaboration with mobile lidar and multiple cameras, belonging to the fields of computer vision, multi-sensor and 3D perception fusion technology. Background Technology
[0002] Surveillance cameras are widely deployed in scenarios such as power line inspection, park security, and industrial supervision. However, traditional monitoring systems mainly rely on two-dimensional image information, which has problems such as insufficient spatial perception capabilities, lack of target depth information, and inaccurate target positioning under occlusion. They are difficult to meet the needs of refined supervision and risk assessment in complex scenarios. At the same time, LiDAR can acquire high-precision three-dimensional geometric structure information, but its point cloud data is usually sparse and lacks rich texture and color information. When used alone, it still has limitations in target recognition, anomaly detection, and scene representation.
[0003] In existing technologies, one approach uses fixed LiDAR and cameras for joint perception, but this typically requires additional dedicated equipment, resulting in high system deployment costs and insufficient utilization of existing surveillance networks. Another approach, while utilizing visual images for scene analysis, lacks a reliable 3D spatial benchmark, making it difficult to achieve unified modeling and accurate spatial correlation across cameras. This leads to difficulties in accurately calculating the true spatial distances between key targets and between key targets and hazardous boundaries. Particularly in large-scale target areas, how to rapidly construct a 3D point cloud benchmark model of the scene using mobile LiDAR, and further achieve spatial unification, collaborative fusion, enhanced reconstruction, and safe distance monitoring with existing surveillance cameras, has become a pressing technical problem that needs to be solved.
[0004] Therefore, this paper proposes a method for monitoring the distance of key targets in collaboration between mobile lidar and multiple cameras. This method is of great significance for improving the three-dimensional representation of scenes, the spatial positioning accuracy of key targets, the dynamic calculation capability of safe distances, and the level of risk warning. Summary of the Invention
[0005] To address the shortcomings of existing monitoring and perception methods, such as insufficient 3D representation, difficulty in unifying images across camera spaces, and limited image-point cloud fusion capabilities, this invention provides a method for monitoring the distance to key targets using mobile lidar and multi-camera collaboration, enabling 3D reconstruction, dynamic monitoring, and risk assessment of target sites.
[0006] The above objectives are achieved through the following technical solutions:
[0007] The present invention provides a method for monitoring the distance to key targets using a mobile lidar system in collaboration with multiple cameras, comprising the following steps:
[0008] Step S1: Continuously scan the target site using a mobile LiDAR to acquire 3D point cloud data of the target site, and construct a 3D point cloud map of the target site through point cloud registration and pose estimation; Simultaneously collect image data from multiple surveillance cameras deployed in the target site, and solve the extrinsic parameters of each surveillance camera relative to the 3D point cloud map based on the cross-modal structural consistency constraints between the 3D point cloud map of the target site and the surveillance camera images, thereby achieving spatial unification between the coordinate systems of multiple surveillance cameras and the mobile LiDAR.
[0009] Step S2: Based on the spatial unification completed in Step S1, establish a multi-monitoring camera network model. According to the intrinsic parameters, extrinsic parameters, installation position and field of view parameters of each monitoring camera, calculate the field of view range, visible area, field of view overlap relationship and spatial coverage relationship of each monitoring camera in the 3D point cloud map. Combine the acquisition time, observation area and target motion information of the monitoring cameras to construct a spatiotemporal correlation model of the monitoring area.
[0010] Step S3: Based on the spatiotemporal correlation model constructed in Step S2, construct a camera selection optimization strategy. Using target area coverage, field of view overlap, observation distance, imaging clarity, and detection reliability as constraints or optimization indicators, select the optimal subset of cameras from multiple surveillance cameras that meet the requirements of collaborative perception of key targets.
[0011] Step S4: Based on the intrinsic and extrinsic parameters of each surveillance camera in the optimal camera subset selected in step S3, project the 3D points in the 3D point cloud map onto the corresponding surveillance camera image plane, establish the correspondence between the 3D points and image pixels, and extract the color information, texture information or semantic information of the image pixels according to the correspondence, assign them to the corresponding 3D points, and form color point cloud data with image attributes.
[0012] Step S5: Based on the color point cloud data, the correspondence between 3D points and image pixels, point cloud geometric constraints, and multi-view image consistency constraints obtained in Step S4, the 3D scene model is jointly optimized to generate a dense 3D scene model of the target site; in the dense 3D scene model, key targets and their corresponding safety supervision objects are identified, the 3D spatial distance between the key targets and the safety supervision objects is calculated, and the presence of safety distance anomalies is determined according to a preset safety distance threshold, thereby realizing safety distance supervision of key targets.
[0013] Compared with the prior art, the beneficial effects of the present invention are:
[0014] (1) This invention combines the LiDAR scanning of the cruise robot with the images of the monitoring camera to complete the 3D point cloud mapping, external parameter estimation and spatial unification of the scene, thus solving the problems of insufficient spatial perception and difficulty in unifying camera coordinates in traditional monitoring systems.
[0015] (2) By combining spatiotemporal correlation modeling and camera selection optimization, this invention achieves monitoring area coverage analysis and optimal camera subset selection, solving the problems of excessive redundant camera calls and low collaborative perception efficiency in existing methods.
[0016] (3) This invention establishes a point cloud-pixel correspondence and constructs a dense three-dimensional scene model by projecting and fusing monitoring images with lidar point clouds. It calculates the spatial distance between key targets and between key targets and dangerous boundary objects in a unified three-dimensional coordinate system, thereby realizing dynamic monitoring and early warning of the safety distance of key targets. Attached Figure Description
[0017] Figure 1 This is an overall flowchart of a key target distance monitoring method proposed in this invention, which combines mobile lidar with multi-camera collaboration. Figure 2 This is a schematic diagram illustrating the principle of extrinsic parameter estimation and spatial unification of surveillance cameras and lidar proposed in this invention. Figure 3 This is a schematic diagram illustrating the spatiotemporal correlation modeling of the monitoring area and the optimization of camera selection in this invention; Figure 4 This is a schematic diagram of the point cloud image fusion and enhancement reconstruction process of the present invention; Figure 5 This is a schematic diagram of the simulation environment described in an embodiment of the present invention. Detailed Implementation
[0018] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] The present invention adopts the following technical solution, combined with Figure 1 To complete the detection and evaluation, this invention proposes a key target distance monitoring method using mobile lidar and multi-camera collaboration, comprising the following steps:
[0020] Step S1: Continuously scan the target site using a mobile LiDAR to acquire 3D point cloud data of the target site, and construct a 3D point cloud map of the target site through point cloud registration and pose estimation; Simultaneously collect image data from multiple surveillance cameras deployed in the target site, and solve the extrinsic parameters of each surveillance camera relative to the 3D point cloud map based on the cross-modal structural consistency constraints between the 3D point cloud map of the target site and the surveillance camera images, thereby achieving spatial unification between the coordinate systems of multiple surveillance cameras and the mobile LiDAR.
[0021] Based on the spatial unification completed in step S1, a multi-monitoring camera network model is established. According to the intrinsic parameters, extrinsic parameters, installation position and field of view parameters of each monitoring camera, the field of view range, visible area, field of view overlap relationship and spatial coverage relationship of each monitoring camera in the 3D point cloud map are calculated. Combined with the acquisition time, observation area and target motion information of the monitoring cameras, a spatiotemporal correlation model of the monitoring area is constructed.
[0022] Step S3: Based on the spatiotemporal correlation model constructed in Step S2, construct a camera selection optimization strategy. Using target area coverage, field of view overlap, observation distance, imaging clarity, and detection reliability as constraints or optimization indicators, select the optimal subset of cameras from multiple surveillance cameras that meet the requirements of collaborative perception of key targets.
[0023] Step S4: Based on the intrinsic and extrinsic parameters of each surveillance camera in the optimal camera subset selected in step S3, project the 3D points in the 3D point cloud map onto the corresponding surveillance camera image plane, establish the correspondence between the 3D points and image pixels, and extract the color information, texture information or semantic information of the image pixels according to the correspondence, assign them to the corresponding 3D points, and form color point cloud data with image attributes.
[0024] Step S5: Based on the color point cloud data, the correspondence between 3D points and image pixels, point cloud geometric constraints, and multi-view image consistency constraints obtained in Step S4, the 3D scene model is jointly optimized to generate a dense 3D scene model of the target site; in the dense 3D scene model, key targets and their corresponding safety supervision objects are identified, the 3D spatial distance between the key targets and the safety supervision objects is calculated, and the presence of safety distance anomalies is determined according to a preset safety distance threshold, thereby realizing safety distance supervision of key targets.
[0025] Combination Figure 2 Preferably, step S1 includes:
[0026] Step S1-1: Control the mobile platform equipped with LiDAR to perform mobile scanning in the target site, continuously collect LiDAR point cloud data of the site, and use the LiDAR SLAM method to process the multi-frame point cloud data to obtain the pose information of LiDAR at each sampling time and construct a three-dimensional point cloud map of the target site.
[0027] Step S1-2: Collect image sequences continuously acquired by the surveillance camera in the target site, perform visual structure restoration processing on the image sequences, establish matching relationships between multiple frames based on the key feature points extracted from the images, restore the camera pose sequence and visual 3D map points, thereby obtaining camera pose change information and scene visual 3D structure information.
[0028] Step S1-3: Using the LiDAR pose change information obtained in step S1-1 and the camera pose change information obtained in step S1-2, establish hand-eye calibration constraints between the LiDAR and the camera, and solve for the initial values of the extrinsic parameters between the LiDAR and the camera using the HECalib method.
[0029] Step S1-4: Based on the initial values of the extrinsic parameters obtained in step S1-3, the LiDAR point cloud is projected onto the image plane, and a point cloud-image reprojection error model is constructed. The extrinsic parameters between the camera and the LiDAR are globally optimized using a nonlinear minimization squares optimization method to achieve spatial unification of multiple cameras in the three-dimensional point cloud coordinate system.
[0030] Preferably, in step S1-3, firstly, based on the LiDAR pose change information obtained in step S1-1 and the camera pose change information obtained in step S1-2, a hand-eye calibration constraint relationship between the camera and the LiDAR is established: Let the camera at the... Frame and the The relative pose between frames is The relative pose of the lidar between corresponding frames is The extrinsic parameter transformation from the lidar coordinate system to the camera coordinate system is as follows: Then the camera pose transformation matrix With lidar pose transformation matrix The following hand-eye calibration constraints must be satisfied: (1) in, Indicates the camera from the first Frame to the Pose transformation matrix between frames; Indicates that the lidar starts from the first Frame to the Pose transformation matrix between frames; This represents the extrinsic transformation matrix from the lidar coordinate system to the camera coordinate system;
[0031] Expanding equation (1) according to the rotation and translation parts, we obtain the corresponding rotation constraints and translation constraints, where the rotation part satisfies The translation part satisfies , For the camera in frame With frames Rotation matrix between; For lidar in frames With frames Rotation matrix between; The rotational extrinsic parameter between the lidar and the camera; This represents the translation vector between the lidar coordinate system and the camera coordinate system; Indicates the camera in frame With frames Translation vectors between them; Indicates that the lidar is in the frame With frames Translation vectors between them; Represents the identity matrix; The scale factor is used to convert the relative scale shift recovered by monocular visual SLAM to the actual scale consistent with that of LiDAR.
[0032] When the sensor's motion trajectory lacks sufficient rotational excitation, the rotational constraint degenerates into an identity relation and the sparse matrix of the translation term tends to zero, leading to calibration degradation. To improve the stability of the extrinsic parameter solution, a regularization constraint term is introduced during the optimization process: (2) in, This represents the extrinsic parameter optimization error function; Indicates the regularization weight coefficient; This represents the initial estimate of the translation parameters;
[0033] By minimizing the regularized extrinsic parameter optimization error function constructed by equation (2), the rotational extrinsic parameters between the lidar and the camera are obtained. and the translation vector between the lidar coordinate system and the camera coordinate system This provides initial parameters for subsequent cross-modal structural consistency optimization processes.
[0034] Preferably, in steps S1-4, after obtaining the initial values of the extrinsic parameters between the lidar and the camera, the lidar point cloud is transformed to the camera coordinate system and further projected onto the image plane.
[0035] Let the three-dimensional points in the lidar point cloud be... The extrinsic transformation matrix obtained in steps S1-3 Transform the LiDAR points to the camera coordinate system to obtain ,in, Indicates the first The first frame of the lidar point cloud Three-dimensional points; This indicates the coordinates of the 3D point in the camera coordinate system;
[0036] The 3D map points were then recovered using visual SLAM. Construct the visual BA reprojection error, which is expressed as: ,in, Represents the camera projection function; Represents the pixel coordinates of key feature points in an image;
[0037] In adjacent image frames, the reprojection error of the aforementioned 3D points after camera pose transformation is expressed as: ,in, This represents the rotation matrix between camera frames; This represents the translation vector between camera frames; Indicates the first The location of feature points in a frame image;
[0038] Based on this, LiDAR point cloud data is introduced to construct cross-modal reprojection error: (3) in, Indicates the first The pixel coordinates of the corresponding feature point in the frame image; and Let represent the rotation matrix and translation vector from the lidar coordinate system to the camera coordinate system, respectively.
[0039] In adjacent image frames, we can further obtain: (4) in, Indicates the camera from the first Frame to the The frame rotation matrix; Camera from the first Frame to the The translation vector of the frame; The scale factor is used to convert the relative scale shift recovered by monocular visual SLAM to the actual scale consistent with that of LiDAR.
[0040] The cross-modal reprojection error is obtained from equations (3) and (4), and effective matching points are selected accordingly. After obtaining effective matching points, point-to-plane error is further introduced to enhance the consistency between the local geometry of the point cloud and the visual 3D structure. For the 3D map points recovered by visual SLAM... In local point cloud collection Search for the nearest neighbor point in the spatial location. ,Right now satisfy Based on the neighborhood point and its surrounding point cloud, a local plane is fitted. Let the normal vector of the local plane be... The point-to-plane error is expressed as: (5) in, Indicates the first A collection of local point clouds in a frame; Represents the normal vector of the local plane; Represents the Euclidean distance between the 3D map points recovered by visual SLAM and the nearest neighbor points in the local point cloud set; This represents the perpendicular distance from the visual 3D point to the fitted local plane;
[0041] To ensure the stability of planar structural constraints, a planar validity screening is performed on the local point cloud set. Specifically, assume that the local point cloud set contains... The residual threshold for the plane fitting is ... The minimum radius threshold for neighborhood distribution is When the sum of the residuals from each neighborhood point in the local point cloud set to the fitting plane is less than And the maximum spatial distance between a neighboring point and the central neighboring point is greater than If the local point cloud set is considered to be able to form a stable local plane, then the local point cloud set is discarded and not included in the subsequent point-to-plane error calculation.
[0042] Finally, a global optimization objective function for cross-modal structural consistency is constructed: (6) in, This represents the point-to-plane error term; This represents the cross-modal reprojection error term; This represents the corresponding weight coefficient. This represents the global optimization objective function used to constrain the consistency between the point cloud geometry and the visual reprojection;
[0043] By minimizing the global optimization objective function represented by equation (6), the external parameters are... , and scale factor Iterative optimization is performed to obtain the optimal extrinsic parameters between the camera and the LiDAR, thereby achieving spatial unification of multiple cameras in the three-dimensional point cloud coordinate system.
[0044] Combination Figure 3 Preferably, step S2 includes:
[0045] Step S2-1: Obtain the rotational extrinsic parameters between the laser radar and the camera based on step S1. Translation vector between the lidar coordinate system and the camera coordinate system The coordinate systems of each camera are unified to the 3D point cloud reference coordinate system of the LiDAR, thereby obtaining the relative positional relationships of the cameras in a unified 3D space. Based on this spatial unification result, the target site is divided into multiple monitoring areas. Furthermore, it is divided into several sub-regions, and a spatial correlation model is established by statistically analyzing the movement relationship of the target between different sub-regions;
[0046] Step S2-2: For the target from the known monitoring area Let the spatial movement relationship after leaving a certain sub-region be given. For the target from the monitored area subregion The set of all directly accessible sub-regions Indicates the target is from the sub-region Move to sub-region The number of times in history This indicates that the target has left the sub-region. And based on the historical number of times the user has left the monitoring network, the spatial correlation probability is defined as: (7) in, Indicates the target is from the sub-region Move to sub-region Spatial correlation probability; Indicates from sub-region Any sub-region that can be directly reached;
[0047] The aforementioned spatial correlation probabilities must satisfy a normalization constraint, that is, for a target from a sub-region All possible destinations from the starting point, and their spatial correlation probabilities to reach each reachable sub-region. Probability of leaving the monitoring network The sum is 1, where Indicates from sub-region A set of directly accessible sub-regions;
[0048] Step S2-3: After obtaining the spatial correlation, to further describe the temporal characteristics of the target's movement between different sub-regions, let... Indicates the target is from the sub-region Move to sub-region Total number of historical times Indicates the target within the time interval Inner subregion Reaching sub-region The historical frequency of a given event is then defined as follows: .in, Indicates the target within the time interval Inner subregion Reaching sub-region Time-related probability; Indicates the time interval The historical number of times a move has been completed within the specified timeframe; Indicates from sub-region Reaching sub-region Total number of historical times;
[0049] Based on the above spatial and temporal correlation probabilities, the spatiotemporal movement pattern of the target in the monitoring area is obtained, providing prior constraints for subsequent visible area calculation and camera scheduling.
[0050] Step S2-4: After obtaining the spatial correlation probability and temporal correlation probability between regions, further combine the camera's detection capability for each sub-region to calculate the region detection reliability, assuming the first... Each camera pair in the sub-area The detection accuracy rate is When there is a total One camera was activated to monitor the sub-area. At that time, this sub-region The overall detection reliability is defined as: (8) in, Subregion Overall testing reliability; This indicates that it is activated for monitoring this sub-region. The number of cameras; Indicates the first One camera for this sub-area Detection accuracy;
[0051] The overall detection reliability of each sub-region can be obtained through reliability parameters. This reliability parameter will serve as an important input for modeling the success probability of cross-regional target tracking in step S3, and will be used to evaluate the detection capabilities of different camera combinations in target tracking tasks.
[0052] Combination Figure 3 Preferably, step S3 includes:
[0053] Step S3-1: Obtain the spatial correlation probability in step S2 Time-related probability and regional detection reliability Afterwards, set For edge nodes Whether selected for tracking targets binary decision variables, For edge nodes The corresponding number below Whether a sub-region is selected for tracking the target Given binary decision variables, then the objective... The probability of successful tracking is defined as: (9) in, Indicate target The probability of successful tracking; Indicates the size of the edge node set; Indicate whether to select the first Tracking target at each edge node ; Indicates the first The first edge node associated with the Sub-regions; Indicates the first The number of sub-regions associated with each edge node; Indicates whether the sub-region is selected for the target. Tracking; This indicates the starting sub-region from which the target is currently leaving; Indicates that the target leaves the sub-region. Move to sub-region Spatial correlation probability; Indicates from sub-region To sub-region In the time interval Time-related probabilities within; Subregion The reliability of the detection;
[0054] The success probability of target tracking will be used as an evaluation metric for subsequent scheduling optimization and will be used to construct reliable row constraints for target tracking.
[0055] Step S3-2: To ensure continuous tracking of the target during its movement across monitoring areas, constraints need to be set on the target tracking success probability. When the target leaves the current monitoring area, it may enter an adjacent reachable area or leave the monitoring network directly. If the next monitoring area fails to complete the detection in time, the risk of the target leaving the network will increase. Based on this, the target tracking success probability, the target leaving the network probability, and the target importance level are all included in the constraints, requiring the scheduling results to ensure that the target tracking success probability meets the preset reliability requirements.
[0056] Based on the above constraints, we further obtain linear inequality constraints: (10) in, This indicates the probability that the target will leave the monitoring network; Indicate target Importance level; Indicate target Equation (10) gives the linear constraint condition for the target tracking success rate, which is used to ensure that the target tracking success rate always meets the preset requirements during the scheduling optimization process.
[0057] Step S3-3: After obtaining the constraints on regional detection reliability and target tracking success rate, establish an edge node scheduling optimization model to minimize the overall system computational overhead. The optimization objective is expressed as: (11) in, Indicates the total computational cost; Indicates the size of the target set; Indicates the first Tracking target at each edge node The amount of data that needs to be processed at that time; Indicate whether to select the first Tracking target at each edge node ; Indicates whether the sub-region is selected for the target. Tracking;
[0058] After completing edge node scheduling optimization and obtaining a set of candidate cameras that meet the tracking success rate constraint, in order to further reduce the number of activated cameras in the system and ensure that all sub-regions within the field of view are effectively covered, a camera selection optimization model is established within a single field of view.
[0059] set up This refers to the set of cameras within the field of view. This represents the set of sub-regions within the field of view, where Indicates the first Sub-regions Indicates camera The first one that can cover the field of view Sub-regions ,otherwise , Indicates camera If a camera is selected for the next round of tracking, the camera selection model is written as: (12) in, Indicates the number of cameras within the field of view; For binary decision variables, when the first... When a camera is selected for the next round of tracking ;otherwise, ;
[0060] While minimizing the number of cameras, further consideration is given to the detection accuracy of the cameras to improve overall detection performance. The optimization objective is further expressed as: (13) in, Indicates the first The detection accuracy of each camera within the field of view; (14) in, Indicates the first The first camera focuses on the field of view of the first... Sub-regions Coverage relationship;
[0061] Step S3-4: After completing the construction of the camera selection optimization model, a greedy strategy-based camera selection method is used to solve the camera selection problem. First, all cameras within the field of view are sorted in descending order according to the number of sub-regions they cover. If multiple cameras cover the same number of sub-regions, the selection is further based on the detection reliability of the camera within that field of view. The cameras are sorted by size; then, cameras are selected sequentially from the sorted list, and the set of currently covered sub-regions is updated; the algorithm terminates when all sub-regions are covered by at least one camera and the detection reliability constraint is satisfied, thus obtaining the final set of cameras. .
[0062] Combination Figure 4 Preferably, step S4 includes:
[0063] Step S4-1: Based on the set of cameras selected in step S3 Geometrically unfold the panoramic images captured by the corresponding cameras to establish a mapping relationship between the virtual image plane and the panoramic spherical coordinate system:
[0064] Let the latitude and longitude of the center of the virtual image plane in spherical coordinates be respectively... and Then the rotation matrix from the spherical coordinate system to the virtual image coordinate system can be expressed as: ,in This represents the rotation matrix about the x-axis. Let represent the rotation matrix about the z-axis; for any pixel in the virtual image plane, its coordinates are represented as . ,in For pixel coordinates, The virtual focal length is used; through rotation transformation, the pixels in the virtual image plane are mapped to the spherical coordinate system to obtain their three-dimensional coordinate representation. , and record as Further based on , and Calculate the corresponding latitude angles based on the spatial relationship between them. With longitude angle This establishes a mapping relationship between virtual image planar pixels and panoramic spherical pixels, enabling the projection and cropping of panoramic images onto planar images;
[0065] Step S4-2: Extract SIFT feature points from the planar image obtained in step S4-1, and calculate the feature density of each region in the image by kernel density estimation;
[0066] Let the set of image feature points be: ,in Indicates the first The two-dimensional coordinates of each image feature point; then the pixel position The characteristic density function is: (15) in, The feature density at the pixel location; Represents the Gaussian kernel function; For bandwidth parameters; Represents the position of a two-dimensional pixel in the image plane; Indicates the first Two-dimensional coordinates of image feature points;
[0067] After obtaining the feature density distribution of the image, according to the feature density function The value of determines the region with a high density of local feature points and treats it as a region with rich features.
[0068] Step S4-3: Based on the camera intrinsic parameter matrix and the rotational extrinsic parameters between the lidar and the camera Translation vector between the lidar coordinate system and the camera coordinate system Projecting the lidar point cloud onto the image plane:
[0069] Let the i-th point in the point cloud to be projected in the lidar coordinate system be... Three-dimensional points are Its coordinates are represented as Based on the extrinsic parameter relationship between the lidar and camera coordinate systems, this 3D point can be transformed to the camera coordinate system, where its coordinates are denoted as: The coordinate transformation relationship is as follows: (16) in, Let be the rotation matrix from the lidar coordinate system to the camera coordinate system. This is the translation vector from the lidar coordinate system to the camera coordinate system;
[0070] Further based on the camera intrinsic parameter matrix The projection relationship between the 3D points in the camera coordinate system and the image plane is as follows: (17) in, The projection scale factor; This is the camera intrinsic parameter matrix; Representing a three-dimensional point Depth components in the camera coordinate system;
[0071] When the projection point When a point falls into the feature-rich region determined in step S4-2, the corresponding 3D point is used as a candidate point. To further improve the matching reliability between the candidate points and image information, depth consistency constraints are used to filter the candidate points, that is, to determine the depth component of the 3D point in the camera coordinate system. With pixel position Depth value at Is the difference between them less than the depth threshold? ;
[0072] The final feature-enhanced point cloud is obtained: (18) in, The output is a feature-enhanced point cloud set, which serves as the input data for point cloud reprojection, color information extraction, and 3D scene modeling in step S5.
[0073] Combination Figure 4 Preferably, step S5 includes:
[0074] Step S5-1: The feature-enhanced point cloud set obtained in step S4 Based on this, according to the camera intrinsic parameter matrix and external parameters ,in This indicates the camera's intrinsic parameters. This represents the rotation matrix from the lidar coordinate system to the camera coordinate system. This represents the corresponding translation vector, and is combined with the set of cameras selected in step S3. Using the point cloud projection relationship established in step S4-3, the 3D points in the feature-enhanced point cloud are projected again onto the corresponding image plane of the selected camera to obtain the 3D points. Its projected pixel position The correspondence between them is established, forming a set of correspondences between feature-enhanced point clouds and image pixels. This correspondence is used to subsequently read the color information corresponding to the 3D points from the image, and provides the basis for the construction of the color point cloud in step S5-2;
[0075] Step S5-2: Based on the correspondence between the 3D points and image pixels established in Step S5-1, read the color value of the corresponding projected pixel position from the selected camera image and use it as the color vector of the 3D point. This process forms a colored point cloud; after obtaining the colored point cloud data, it is used as the initial input for a 3D Gaussian model; to reduce training complexity and preserve local geometric structure, the colored point cloud is voxelized and aggregated; for the first... Individual unit voxel space center of mass From all three-dimensional points within this voxel The coordinate mean is determined, and the average color of the voxels is determined. The color vector of all three-dimensional points within this voxel The mean is determined; thus, the initial point set used for 3D Gaussian modeling is obtained. Each initialization point is determined by the centroid of the voxel space. and its corresponding average voxel color composition;
[0076] Step S5-3: Based on the initialization point set obtained in step S5-2 A three-dimensional Gaussian field model is constructed and trained and optimized using the camera images selected in step S3; for the initial point set No. There are initialization points, and their corresponding Gaussian primitives are denoted as . The Gaussian primitive is located at the center. Covariance matrix Color attributes and opacity parameter Together they form; among them, Indicates the first The central position of a Gaussian primitive; This represents the covariance matrix of the corresponding Gaussian primitive, used to describe its spatial distribution range; Indicates the first The color attribute of a Gaussian primitive is initialized by the average color of the voxels in the corresponding initialization point. The opacity parameter represents the Gaussian primitive;
[0077] Render the current Gaussian model from the perspectives of each camera, and construct a multi-view reconstruction error function: (19) in, Indicates the first Real images from the perspective of the camera used in the training. This represents the rendered image of the current Gaussian model from this viewpoint. The number of viewpoints involved in training and optimization is represented; the Gaussian model parameters are iteratively optimized by minimizing the error function shown in equation (19);
[0078] Furthermore, to reduce the parameter scale and improve training efficiency during model training, the order of the spherical harmonic function expansion of the Gaussian primitive color attribute is restricted, and only low-order spherical harmonic components are retained for color modeling, thereby reducing memory consumption and improving 3D reconstruction efficiency; finally, a 3D Gaussian model that can simultaneously express scene geometry and texture information is obtained.
[0079] Step S5-4: Based on the three-dimensional Gaussian scene model constructed in step S5-3, obtain the spatial position of the key targets in a unified three-dimensional coordinate system, calculate the spatial distance between key targets and between key targets and dangerous boundary objects, and make a judgment based on a preset safety threshold to realize the supervision of the safety distance of key targets.
[0080] Assuming a unified three-dimensional coordinate system, at time... No. The three-dimensional positions of the key targets are Their coordinate components are respectively , , ;
[0081] For any two key objectives and At any given moment Spatial distance is defined as: (20) in Indicate key objectives With key objectives At any moment The spatial distance between two key targets, according to the definition of Euclidean distance, can be determined by the distance between them. , , The difference in the three coordinate directions is calculated.
[0082] Let the set of boundary points of the dangerous boundary object in a unified three-dimensional coordinate system be... The key objective The minimum spatial distance to a dangerous boundary object is defined as: (twenty one) in, Represents any boundary point on a dangerous boundary object. This represents the set of boundary points of a dangerous boundary object in a unified three-dimensional coordinate system. Indicate key objectives At any moment Minimum spatial distance to the dangerous boundary object;
[0083] Let the safe distance threshold between key targets be... The safe distance threshold between critical targets and hazardous boundary objects is When key objectives With key objectives At any moment spatial distance Less than When a safe distance risk exists between the two, it is determined that there is one; when a critical target Minimum spatial distance to dangerous boundary objects Less than At that time, it was determined that there was a risk of the critical target crossing the boundary with the dangerous boundary object;
[0084] If any of the above risk assessment conditions are met, the corresponding safety warning information is output; otherwise, the current state is determined to be safe. Dynamic monitoring of the safety distance to key targets is achieved through continuous updates to the spatial location and distance calculations.
[0085] To verify the feasibility of the key target distance monitoring method based on the collaboration of mobile lidar and multiple monitoring cameras described in this invention, in a specific embodiment, a simulation environment was constructed based on ROS Noetic and Gazebo11. For example... Figure 5 As shown, the simulation scenario is a rectangular monitoring area of 20m×16m×5m. Within the scenario, there are 8 fixed obstacles, a mobile patrol robot equipped with a 3D LiDAR, three fixed surveillance cameras, a key monitored object, and a moving target. The maximum speed of the mobile patrol robot is set to 0.6m / s, and the cruising speed is set to 0.4m / s. The LiDAR scanning frequency is 10Hz, the maximum detection distance is 30m, and the distance noise standard deviation is 0.03m. Point cloud data is published via the " / lidar_points" topic. The three surveillance cameras are installed at different locations within the monitored area, with installation heights ranging from 3.6m to 4.2m. The image resolution is 1920×1080, the frame rate is 30fps, the horizontal field of view is 75°, and the vertical field of view is 50°. Image data is published via the " / camera_i / image_raw" topic. The moving target measures 0.5m × 0.5m × 1.7m, with an initial position of (3.20, 2.00, 1.50)m and a speed ranging from 0.3m / s to 0.5m / s. The center point of the key monitored object is (10.00, 8.00, 0.00)m, the safety distance threshold is set to 3.0m, and the simulation time is 60s. Simulation time and sensor parameters are shown in Table 1.
[0086] Table 1. Simulation environment and sensor parameters: Simulation platform ROS version ROS Noetic Simulation platform Gazebo version Gazebo11 Scene size Size of regulatory area 20m×16m Scene obstacles Fixed number of obstacles 8 mobile platform Robot's maximum speed 0.6m / s mobile platform Robot cruise speed 0.4m / s LiDAR Scan frequency 10Hz LiDAR Maximum detection range 30m LiDAR Distance noise standard deviation 0.03m Surveillance cameras Number of cameras 3 units Surveillance cameras Image resolution 1920×1080 Surveillance cameras Frame rate 30fps Surveillance cameras Horizontal field of view 75° Surveillance cameras Vertical field of view 50° Target movement Moving target speed 0.3–0.5 m / s Safety supervision safe distance threshold 3.0m Simulation duration single simulation time 60s
[0087] During the simulation, the system establishes TF transformation relationships between the world coordinate system "world", the robot's body coordinate system "base_link", the LiDAR coordinate system "lidar_link", and the coordinate systems of each monitoring camera "camera_i_link", and publishes coordinate transformation information via " / tf" and " / tf_static". Subsequently, the LiDAR point cloud and multi-camera observation results are unified into the world coordinate system. The system determines whether the target is within the field of view of the corresponding camera based on the projection results of the target in each camera coordinate system, and performs 3D localization by combining the point cloud data. The fused localization results are updated at a frequency of 10Hz and published via the " / target_pose_world" topic.
[0088] Finally, the system uses the center point of the key monitored object as the distance benchmark to calculate the spatial distance between the moving target and the key monitored object in real time, and performs a safety assessment every 0.1 seconds. When the estimated distance is less than 3.0m, an early warning message is issued through the " / distance_alarm" topic; when the estimated distance is greater than or equal to 3.0m, the system maintains normal monitoring status. This simulation environment can be used to verify the effectiveness of multi-sensor coordinate unification, multi-camera field-of-view judgment, target 3D localization, and key target distance monitoring.
[0089] The simulation results are shown in Table 2. During the process of the moving target approaching the key monitored object along a preset trajectory, this invention can continuously output the estimated three-dimensional coordinates of the target in the world coordinate system and calculate the spatial distance between it and the key monitored object in real time. For example, at a simulation time of 10 seconds, the estimated target distance is 4.68 m, which is greater than the safe distance threshold of 3.0 m, and the system does not trigger an alarm; at a simulation time of 20 seconds, the estimated target distance decreases to 2.09 m, which is less than the safe distance threshold, and the system triggers an alarm; at a simulation time of 40 seconds, the estimated target distance increases to 3.57 m, the system cancels the alarm and returns to normal monitoring status.
[0090] Table 2. Simulation results of target localization and distance monitoring: 0 (3.20,2.00,1.50) (3.17,2.00,1.55) 9.19 9.22 0.02 1、2 no 5 (5.00,3.50,1.50) (4.94,3.51,1.48) 6.89 6.93 0.03 1、2 no 10 (6.80,4.80,1.50) (6.86,4.84,1.42) 4.77 4.68 0.09 1、2 no 15 (8.20,6.20,1.50) (8.14,6.17,1.50) 2.95 3.01 0.06 1、2 no 20 (9.00,7.00,1.50) (8.96,6.98,1.50) 2.06 2.09 0.03 1、2 yes 25 (10.40,7.50,1.50) (10.38,7.44,1.51) 1.63 1.65 0.02 2、3 yes 30 (11.20,8.30,1.50) (11.21,8.30,1.45) 1.94 1.91 0.03 2、3 yes 35 (12.50,8.60,1.50) (12.48,8.59,1.54) 2.98 2.98 0.01 2、3 yes 40 (13.00,9.40,1.50) (12.94,9.38,1.48) 3.63 3.57 0.07 1、3 no 45 (11.50,10.00,1.50) (11.45,10.06,1.52) 2.92 2.94 0.03 1、3 yes
[0091] Meanwhile, the visible camera results at different time points show that when the target moves within the monitored area, the three surveillance cameras can form complementary observations based on the target's location. From 0s to 20s, the target is primarily observed by surveillance cameras 1 and 2; from 25s to 35s, it is primarily observed by surveillance cameras 2 and 3; and from 40s to 45s, it is primarily observed by surveillance cameras 1 and 3. Therefore, even when the target leaves the field of view of a camera or is partially obstructed, the system can still maintain target positioning and distance monitoring using the perspectives of other cameras and LiDAR point cloud information.
[0092] In this embodiment, the error between the fused estimated distance and the actual distance remains within a small range, with a maximum distance error of 0.09m and a minimum of 0.01m. This result demonstrates that, through the collaborative perception of mobile lidar and multiple surveillance cameras, and the fusion calculation of multi-sensor observations in a unified world coordinate system, this invention can achieve continuous positioning of key targets, spatial distance calculation, and safety threshold early warning, thereby improving the accuracy and stability of target distance monitoring in complex surveillance scenarios.
Claims
1. A method for monitoring the distance to key targets using a mobile lidar system in collaboration with multiple cameras, characterized in that, Includes the following steps: Step S1: Use a mobile lidar to continuously scan the target site to obtain three-dimensional point cloud data of the target site, and construct a three-dimensional point cloud map of the target site through point cloud registration and pose estimation. Simultaneously collect image data from multiple surveillance cameras deployed in the target site. Based on the cross-modal structural consistency constraints between the 3D point cloud map of the target site and the surveillance camera images, solve the extrinsic parameters of each surveillance camera relative to the 3D point cloud map to achieve spatial unification between the coordinate system of multiple surveillance cameras and the mobile lidar. Step S2: Based on the spatial unification completed in Step S1, establish a multi-monitoring camera network model. According to the intrinsic parameters, extrinsic parameters, installation position and field of view parameters of each monitoring camera, calculate the field of view range, visible area, field of view overlap relationship and spatial coverage relationship of each monitoring camera in the 3D point cloud map. Combine the acquisition time, observation area and target motion information of the monitoring cameras to construct a spatiotemporal correlation model of the monitoring area. Step S3: Based on the spatiotemporal correlation model constructed in Step S2, construct a camera selection optimization strategy. Using target area coverage, field of view overlap, observation distance, imaging clarity, and detection reliability as constraints or optimization indicators, select the optimal subset of cameras from multiple surveillance cameras that meet the requirements of collaborative perception of key targets. Step S4: Based on the intrinsic and extrinsic parameters of each surveillance camera in the optimal camera subset selected in step S3, project the 3D points in the 3D point cloud map onto the corresponding surveillance camera image plane, establish the correspondence between the 3D points and image pixels, and extract the color information, texture information or semantic information of the image pixels according to the correspondence, assign them to the corresponding 3D points, and form color point cloud data with image attributes. Step S5: Based on the color point cloud data, the correspondence between 3D points and image pixels, point cloud geometric constraints, and multi-view image consistency constraints obtained in Step S4, the 3D scene model is jointly optimized to generate a dense 3D scene model of the target site; in the dense 3D scene model, key targets and their corresponding safety supervision objects are identified, the 3D spatial distance between the key targets and the safety supervision objects is calculated, and the presence of safety distance anomalies is determined according to a preset safety distance threshold, thereby realizing safety distance supervision of key targets.
2. The method for monitoring the distance to key targets using mobile lidar and multi-camera collaboration according to claim 1, characterized in that, Step S1 includes: Step S1-1: Control the mobile platform equipped with LiDAR to perform mobile scanning in the target site, continuously collect LiDAR point cloud data of the site, and use the LiDAR SLAM method to process the multi-frame point cloud data to obtain the pose information of LiDAR at each sampling time and construct a three-dimensional point cloud map of the target site. Step S1-2: Collect image sequences continuously acquired by the surveillance camera in the target site, perform visual structure restoration processing on the image sequences, establish matching relationships between multiple frames based on the key feature points extracted from the images, restore the camera pose sequence and visual 3D map points, thereby obtaining camera pose change information and scene visual 3D structure information. Step S1-3: Using the LiDAR pose change information obtained in step S1-1 and the camera pose change information obtained in step S1-2, establish hand-eye calibration constraints between the LiDAR and the camera, and solve for the initial values of the extrinsic parameters between the LiDAR and the camera using the HECalib method. Step S1-4: Based on the initial values of the extrinsic parameters obtained in step S1-3, the LiDAR point cloud is projected onto the image plane, and a point cloud-image reprojection error model is constructed. The extrinsic parameters between the camera and the LiDAR are globally optimized using a nonlinear minimization squares optimization method to achieve spatial unification of multiple cameras in the three-dimensional point cloud coordinate system.
3. The method for monitoring the distance to key targets using mobile lidar and multi-camera collaboration according to claim 2, characterized in that, The specific methods for steps S1-3 are as follows: First, based on the LiDAR pose change information obtained in step S1-1 and the camera pose change information obtained in step S1-2, a hand-eye calibration constraint relationship between the camera and the LiDAR is established: Let the camera be in the... Frame and the The relative pose between frames is The relative pose of the lidar between corresponding frames is The extrinsic parameter transformation from the lidar coordinate system to the camera coordinate system is as follows: Then the camera pose transformation matrix With lidar pose transformation matrix The following hand-eye calibration constraints must be satisfied: (1) in, Indicates the camera from the first Frame to the Pose transformation matrix between frames; Indicates that the lidar starts from the first Frame to the Pose transformation matrix between frames; This represents the extrinsic transformation matrix from the lidar coordinate system to the camera coordinate system; Expanding equation (1) according to the rotation and translation parts, we obtain the corresponding rotation constraints and translation constraints, where the rotation part satisfies The translation part satisfies , For the camera in frame With frames Rotation matrix between; For lidar in frames With frames Rotation matrix between; The rotational extrinsic parameter between the lidar and the camera; This represents the translation vector between the lidar coordinate system and the camera coordinate system; Indicates the camera in frame With frames Translation vectors between them; Indicates that the lidar is in the frame With frames Translation vectors between them; Represents the identity matrix; The scale factor is used to convert the relative scale shift recovered by monocular visual SLAM to the actual scale consistent with that of LiDAR. When the sensor's motion trajectory lacks sufficient rotational excitation, the rotational constraint degenerates into an identity relation and the sparse matrix of the translation term tends to zero, leading to calibration degradation. To improve the stability of the extrinsic parameter solution, a regularization constraint term is introduced during the optimization process: (2) in, This represents the extrinsic parameter optimization error function; Indicates the regularization weight coefficient; This represents the initial estimate of the translation parameters; By minimizing the regularized extrinsic parameter optimization error function constructed by equation (2), the rotational extrinsic parameters between the lidar and the camera are obtained. and the translation vector between the lidar coordinate system and the camera coordinate system This provides initial parameters for subsequent cross-modal structural consistency optimization processes.
4. The method for monitoring the distance to key targets using mobile lidar and multi-camera collaboration according to claim 3, characterized in that, Steps S1-4 include: After obtaining the initial values of the extrinsic parameters between the LiDAR and the camera, the LiDAR point cloud is transformed into the camera coordinate system and then projected onto the image plane.
5. Let the three-dimensional points in the lidar point cloud be... The extrinsic transformation matrix obtained in steps S1-3 Transform the LiDAR points to the camera coordinate system to obtain ,in, Indicates the first The first frame of the lidar point cloud Three-dimensional points; This indicates the coordinates of the 3D point in the camera coordinate system; The 3D map points were then recovered using visual SLAM. Construct the visual BA reprojection error, which is expressed as: ,in, Represents the camera projection function; Represents the pixel coordinates of key feature points in an image; In adjacent image frames, the reprojection error of the aforementioned 3D points after camera pose transformation is expressed as: ,in, This represents the rotation matrix between camera frames; This represents the translation vector between camera frames; Indicates the first The location of feature points in a frame image; Based on this, LiDAR point cloud data is introduced to construct cross-modal reprojection error: (3) in, Indicates the first The pixel coordinates of the corresponding feature point in the frame image; and Let represent the rotation matrix and translation vector from the lidar coordinate system to the camera coordinate system, respectively. In adjacent image frames, we can further obtain: (4) in, Indicates the camera from the first Frame to the The frame rotation matrix; Camera from the first Frame to the The translation vector of the frame; The scale factor is used to convert the relative scale shift recovered by monocular visual SLAM to the actual scale consistent with that of LiDAR. The cross-modal reprojection error is obtained from equations (3) and (4), and effective matching points are selected accordingly. After obtaining effective matching points, point-to-plane error is further introduced to enhance the consistency between the local geometry of the point cloud and the visual 3D structure. For the 3D map points recovered by visual SLAM... In local point cloud collection Search for the nearest neighbor point in the spatial location. ,Right now satisfy Based on the neighborhood point and its surrounding point cloud, a local plane is fitted. Let the normal vector of the local plane be... The point-to-plane error is expressed as: (5) in, Indicates the first A collection of local point clouds in a frame; Represents the normal vector of the local plane; Represents the Euclidean distance between the 3D map points recovered by visual SLAM and the nearest neighbor points in the local point cloud set; This represents the perpendicular distance from the visual 3D point to the fitted local plane; To ensure the stability of planar structural constraints, a planar validity screening is performed on the local point cloud set. Specifically, assume that the local point cloud set contains... The residual threshold for the plane fitting is ... The minimum radius threshold for neighborhood distribution is When the sum of the residuals from each neighborhood point in the local point cloud set to the fitting plane is less than And the maximum spatial distance between a neighboring point and the central neighboring point is greater than If the local point cloud set is considered to be able to form a stable local plane, then the local point cloud set is discarded and not included in the subsequent point-to-plane error calculation. Finally, a global optimization objective function for cross-modal structural consistency is constructed: (6) in, This represents the point-to-plane error term; This represents the cross-modal reprojection error term; This represents the corresponding weight coefficient. This represents the global optimization objective function used to constrain the consistency between the point cloud geometry and the visual reprojection; By minimizing the global optimization objective function represented by equation (6), the external parameters are... , and scale factor Iterative optimization is performed to obtain the optimal extrinsic parameters between the camera and the LiDAR, thereby achieving spatial unification of multiple cameras in the three-dimensional point cloud coordinate system.
6. The method for monitoring the distance to key targets using mobile lidar and multi-camera collaboration according to claim 1, characterized in that, Step S2 includes: Step S2-1: Obtain the rotational extrinsic parameters between the laser radar and the camera based on step S1. Translation vector between the lidar coordinate system and the camera coordinate system The coordinate systems of each camera are unified to the 3D point cloud reference coordinate system of the LiDAR, thereby obtaining the relative positional relationships of the cameras in a unified 3D space. Based on this spatial unification result, the target site is divided into multiple monitoring areas. Furthermore, it is divided into several sub-regions, and a spatial correlation model is established by statistically analyzing the movement relationship of the target between different sub-regions; Step S2-2: For the target from the known monitoring area Let the spatial movement relationship after leaving a certain sub-region be given. For the target from the monitored area subregion The set of all directly accessible sub-regions Indicates the target is from the sub-region Move to sub-region The number of times in history This indicates that the target has left the sub-region. And based on the historical number of times the user has left the monitoring network, the spatial correlation probability is defined as: (7) in, Indicates the target is from the sub-region Move to sub-region Spatial correlation probability; Indicates from sub-region Any sub-region that can be directly reached; The aforementioned spatial correlation probabilities must satisfy a normalization constraint, that is, for a target from a sub-region All possible destinations from the starting point, and their spatial correlation probabilities to reach each reachable sub-region. Probability of leaving the monitoring network The sum is 1, where Indicates from sub-region A set of directly accessible sub-regions; Step S2-3: After obtaining the spatial correlation, to further describe the temporal characteristics of the target's movement between different sub-regions, let... Indicates the target is from the sub-region Move to sub-region Total number of historical times Indicates the target within the time interval Inner subregion Reaching sub-region The historical frequency of a given event is then defined as follows: .in, Indicates the target within the time interval Inner subregion Reaching sub-region Time-related probability; Indicates the time interval The historical number of times a move has been completed within the specified timeframe; Indicates from sub-region Reaching sub-region Total number of historical times; Based on the above spatial and temporal correlation probabilities, the spatiotemporal movement pattern of the target in the monitoring area is obtained, providing prior constraints for subsequent visible area calculation and camera scheduling. Step S2-4: After obtaining the spatial correlation probability and temporal correlation probability between regions, further combine the camera's detection capability for each sub-region to calculate the region detection reliability, assuming the first... Each camera pair in the sub-area The detection accuracy rate is When there is a total One camera was activated to monitor the sub-area. At that time, this sub-region The overall detection reliability is defined as: (8) in, Subregion Overall testing reliability; This indicates that it is activated for monitoring this sub-region. The number of cameras; Indicates the first One camera for this sub-area Detection accuracy; The overall detection reliability of each sub-region can be obtained through reliability parameters. This reliability parameter will serve as an important input for modeling the success probability of cross-regional target tracking in step S3, and will be used to evaluate the detection capabilities of different camera combinations in target tracking tasks.
7. A method for monitoring the distance to key targets using mobile lidar and multi-camera collaboration as described in claim 5, characterized in that, Step S3 includes: Step S3-1: Obtain the spatial correlation probability in step S2 Time-related probability and regional detection reliability Afterwards, set For edge nodes Whether selected for tracking targets binary decision variables, For edge nodes The corresponding number below Whether a sub-region is selected for tracking the target Given binary decision variables, then the objective... The probability of successful tracking is defined as: (9) in, Indicate target The probability of successful tracking; Indicates the size of the edge node set; Indicate whether to select the first Tracking target at each edge node ; Indicates the first The first edge node associated with the Sub-regions; Indicates the first The number of sub-regions associated with each edge node; Indicates whether the sub-region is selected for the target. Tracking; This indicates the starting sub-region from which the target is currently leaving; Indicates that the target leaves the sub-region. Move to sub-region Spatial correlation probability; Indicates from sub-region To sub-region In the time interval Time-related probabilities within; Subregion The reliability of the detection; The success probability of target tracking will be used as an evaluation metric for subsequent scheduling optimization and will be used to construct reliable row constraints for target tracking. Step S3-2: To ensure continuous tracking of the target during its movement across monitoring areas, constraints need to be set on the target tracking success probability. When the target leaves the current monitoring area, it may enter an adjacent reachable area or leave the monitoring network directly. If the next monitoring area fails to complete the detection in time, the risk of the target leaving the network will increase. Based on this, the target tracking success probability, the target leaving the network probability, and the target importance level are all included in the constraints, requiring the scheduling results to ensure that the target tracking success probability meets the preset reliability requirements. Based on the above constraints, we further obtain linear inequality constraints: (10) in: This indicates the probability that the target will leave the monitoring network; Indicate target Importance level; Indicate target The probability of successful tracking; Equation (10) gives the linear constraint on the target tracking success rate, which is used to ensure that the target tracking success probability always meets the preset requirements during the scheduling optimization process; Step S3-3: After obtaining the constraints on regional detection reliability and target tracking success rate, establish an edge node scheduling optimization model to minimize the overall system computational overhead. The optimization objective is expressed as: (11) in, Indicates the total computational cost; Indicates the size of the target set; Indicates the first Tracking target at each edge node The amount of data that needs to be processed at that time; Indicate whether to select the first Tracking target at each edge node ; Indicates whether the sub-region is selected for the target. Tracking; After completing edge node scheduling optimization and obtaining a set of candidate cameras that meet the tracking success rate constraint, in order to further reduce the number of activated cameras in the system and ensure that all sub-regions within the field of view are effectively covered, a camera selection optimization model is established within a single field of view. set up This refers to the set of cameras within the field of view. This represents the set of sub-regions within the field of view, where Indicates the first Sub-regions Indicates camera The first one that can cover the field of view Sub-regions ,otherwise , Indicates camera If a camera is selected for the next round of tracking, the camera selection model is written as: (12) in, Indicates the number of cameras within the field of view; For binary decision variables, when the first... When a camera is selected for the next round of tracking ;otherwise, ; While minimizing the number of cameras, further consideration is given to the detection accuracy of the cameras to improve overall detection performance. The optimization objective is further expressed as: (13) in, Indicates the first The detection accuracy of each camera within the field of view; (14) in, Indicates the first The first camera focuses on the field of view of the first... Sub-regions Coverage relationship; Step S3-4: After completing the construction of the camera selection optimization model, a greedy strategy-based camera selection method is used to solve the camera selection problem. First, all cameras within the field of view are sorted in descending order according to the number of sub-regions they cover. If multiple cameras cover the same number of sub-regions, the selection is further based on the detection reliability of the camera within that field of view. The cameras are sorted by size; then, cameras are selected sequentially from the sorted list, and the set of currently covered sub-regions is updated; the algorithm terminates when all sub-regions are covered by at least one camera and the detection reliability constraint is satisfied, thus obtaining the final set of cameras. .
8. A method for monitoring the distance to key targets using mobile lidar and multi-camera collaboration as described in claim 6, characterized in that, Step S4 includes: Step S4-1: Based on the set of cameras selected in step S3 Geometrically unfold the panoramic images captured by the corresponding cameras to establish a mapping relationship between the virtual image plane and the panoramic spherical coordinate system: Let the latitude and longitude of the center of the virtual image plane in spherical coordinates be respectively... and Then the rotation matrix from the spherical coordinate system to the virtual image coordinate system can be expressed as: ,in This represents the rotation matrix about the x-axis. Let represent the rotation matrix about the z-axis; for any pixel in the virtual image plane, its coordinates are represented as . ,in For pixel coordinates, The virtual focal length is used; through rotation transformation, the pixels in the virtual image plane are mapped to the spherical coordinate system to obtain their three-dimensional coordinate representation. , and record as Further based on , and Calculate the corresponding latitude angles based on the spatial relationship between them. With longitude angle This establishes a mapping relationship between virtual image planar pixels and panoramic spherical pixels, enabling the projection and cropping of panoramic images onto planar images; Step S4-2: Extract SIFT feature points from the planar image obtained in step S4-1, and calculate the feature density of each region in the image by kernel density estimation; Let the set of image feature points be: ,in Indicates the first The two-dimensional coordinates of each image feature point; then the pixel position The characteristic density function is: (15) in, The feature density at the pixel location; Represents the Gaussian kernel function; For bandwidth parameters; Represents the position of a two-dimensional pixel in the image plane; Indicates the first Two-dimensional coordinates of image feature points; After obtaining the feature density distribution of the image, according to the feature density function The value of determines the region with a high density of local feature points and treats it as a region with rich features. Step S4-3: Based on the camera intrinsic parameter matrix and the rotational extrinsic parameters between the lidar and the camera Translation vector between the lidar coordinate system and the camera coordinate system Projecting the lidar point cloud onto the image plane: Let the i-th point in the point cloud to be projected in the lidar coordinate system be... Three-dimensional points are Its coordinates are represented as Based on the extrinsic parameter relationship between the lidar and camera coordinate systems, this 3D point can be transformed to the camera coordinate system, where its coordinates are denoted as: The coordinate transformation relationship is as follows: (16) in, Let be the rotation matrix from the lidar coordinate system to the camera coordinate system. This is the translation vector from the lidar coordinate system to the camera coordinate system; Further based on the camera intrinsic parameter matrix The projection relationship between the 3D points in the camera coordinate system and the image plane is as follows: (17) in, The projection scale factor; This is the camera intrinsic parameter matrix; Representing a three-dimensional point Depth components in the camera coordinate system; When the projection point When a point falls into the feature-rich region determined in step S4-2, the corresponding 3D point is used as a candidate point. To further improve the matching reliability between the candidate points and image information, depth consistency constraints are used to filter the candidate points, that is, to determine the depth component of the 3D point in the camera coordinate system. With pixel position Depth value at Is the difference between them less than the depth threshold? ; The final feature-enhanced point cloud is obtained: (18) in, The output is a feature-enhanced point cloud set, which serves as the input data for point cloud reprojection, color information extraction, and 3D scene modeling in step S5.
9. A method for monitoring the distance to key targets using mobile lidar and multi-camera collaboration as described in claim 1, characterized in that, Step S5 includes: Step S5-1: The feature-enhanced point cloud set obtained in step S4 Based on this, according to the camera intrinsic parameter matrix and external parameters ,in This indicates the camera's intrinsic parameters. This represents the rotation matrix from the lidar coordinate system to the camera coordinate system. This represents the corresponding translation vector, and is combined with the set of cameras selected in step S3. Using the point cloud projection relationship established in step S4-3, the 3D points in the feature-enhanced point cloud are projected again onto the corresponding image plane of the selected camera to obtain the 3D points. Its projected pixel position The correspondence between them is established, forming a set of correspondences between feature-enhanced point clouds and image pixels. This correspondence is used to subsequently read the color information corresponding to the 3D points from the image, and provides the basis for the construction of the color point cloud in step S5-2; Step S5-2: Based on the correspondence between the 3D points and image pixels established in Step S5-1, read the color value of the corresponding projected pixel position from the selected camera image and use it as the color vector of the 3D point. This process forms a colored point cloud; after obtaining the colored point cloud data, it is used as the initial input for a 3D Gaussian model; to reduce training complexity and preserve local geometric structure, the colored point cloud is voxelized and aggregated; for the first... Individual unit voxel space center of mass From all three-dimensional points within this voxel The coordinate mean is determined, and the average color of the voxels is determined. The color vector of all three-dimensional points within this voxel The mean is determined; thus, the initial point set used for 3D Gaussian modeling is obtained. Each initialization point is determined by the centroid of the voxel space. and its corresponding average voxel color composition; Step S5-3: Based on the initialization point set obtained in step S5-2 A three-dimensional Gaussian field model is constructed and trained and optimized using the camera images selected in step S3; for the initial point set No. There are initialization points, and their corresponding Gaussian primitives are denoted as . The Gaussian primitive is located at the center. Covariance matrix Color attributes and opacity parameter Together they form; among them, Indicates the first The central position of a Gaussian primitive; This represents the covariance matrix of the corresponding Gaussian primitive, used to describe its spatial distribution range; Indicates the first The color attribute of a Gaussian primitive is initialized by the average color of the voxels in the corresponding initialization point. The opacity parameter represents the Gaussian primitive; Render the current Gaussian model from the perspectives of each camera, and construct a multi-view reconstruction error function: (19) in, Indicates the first Real images from the perspective of the camera used in the training. This represents the rendered image of the current Gaussian model from this viewpoint. Indicates the number of viewpoints involved in training and optimization; The Gaussian model parameters are iteratively optimized by minimizing the error function shown in equation (19); Furthermore, to reduce the parameter scale and improve training efficiency during model training, the order of the spherical harmonic function expansion of the Gaussian primitive color attribute is restricted, and only low-order spherical harmonic components are retained for color modeling, thereby reducing memory consumption and improving 3D reconstruction efficiency; finally, a 3D Gaussian model that can simultaneously express scene geometry and texture information is obtained. Step S5-4: Based on the three-dimensional Gaussian scene model constructed in step S5-3, obtain the spatial position of the key targets in a unified three-dimensional coordinate system, calculate the spatial distance between key targets and between key targets and dangerous boundary objects, and make a judgment based on a preset safety threshold to realize the supervision of the safety distance of key targets. Assuming a unified three-dimensional coordinate system, at time... No. The three-dimensional positions of the key targets are Their coordinate components are respectively , , ; For any two key objectives and At any given moment Spatial distance is defined as: (20) in Indicate key objectives With key objectives At any moment The spatial distance between two key targets, according to the definition of Euclidean distance, can be determined by the distance between them. , , The difference in the three coordinate directions is calculated. Let the set of boundary points of the dangerous boundary object in a unified three-dimensional coordinate system be... The key objective The minimum spatial distance to a dangerous boundary object is defined as: (21) in, Represents any boundary point on a dangerous boundary object. This represents the set of boundary points of a dangerous boundary object in a unified three-dimensional coordinate system. Indicate key objectives At any moment Minimum spatial distance to the dangerous boundary object; Let the safe distance threshold between key targets be... The safe distance threshold between critical targets and hazardous boundary objects is When key objectives With key objectives At any moment spatial distance Less than When a safe distance risk exists between the two, it is determined that there is one; when a critical target Minimum spatial distance to dangerous boundary objects Less than At that time, it was determined that there was a risk of the critical target crossing the boundary with the dangerous boundary object; If any of the above risk assessment conditions are met, the corresponding safety warning information is output; otherwise, the current state is determined to be safe. Dynamic monitoring of the safety distance to key targets is achieved through continuous updates to the spatial location and distance calculations.