Building construction safety assessment system based on computer vision and unmanned aerial vehicle inspection

CN122598041APending Publication Date: 2026-08-18SHENZHEN JINDING SAFETY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610726527.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0003]上述现有技术存在单视角视觉盲区与施工场景动态演变导致危险区域评估失真的技术问题

Benefits of technology

1.本发明通过将无人机采集的多视角连续图像序列生成局部三维点云,并将二维图像特征投影至局部三维点云所在三维空间构建跨模态时空映射关系,利用三维空间信息消除了单视角视觉盲区对危险区域与人员位置关系的遮挡影响;通过引入施工进度时序维度,将前序时刻的危险区域边界作为先验知识,结合当前时刻局部三维点云中建筑结构的增量变化特征,利用时序递归网络对先验知识进行修正以动态更新危险区域的空间边界,使危险边界跟随施工物理结构的演变而自适应调整;基于跨模态时空映射关系确定人员目标的三维坐标,计算人员目标的三维坐标与动态更新的危险区域的空间边界之间的空间距离并输出安全评估结果,克服了二维静态评估的偏差,获得了与实际物理空间状态相符的安全评估结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598041A_ABST
    Figure CN122598041A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer vision, and particularly relates to a building construction safety evaluation system based on computer vision and unmanned aerial vehicle inspection. A local three-dimensional point cloud is generated by collecting a multi-view continuous image sequence of a construction site, and a cross-modal space-time mapping relationship is constructed by projecting two-dimensional image features to the local three-dimensional point cloud. A construction progress time sequence dimension is introduced, the boundary of a dangerous area at a previous time is taken as prior knowledge, and the spatial boundary of the dangerous area is dynamically updated by a time sequence recursive network in combination with the incremental change features of the building structure in the local three-dimensional point cloud at the current time. Based on the cross-modal space-time mapping relationship, the three-dimensional coordinates of a personnel target are determined, the spatial distance between the three-dimensional coordinates of the personnel target and the dynamically updated spatial boundary of the dangerous area is calculated, and a safety evaluation result is output. The present application eliminates the influence of single-view visual blind area obstruction, overcomes the defect that a static dangerous boundary cannot adapt to the evolution of a construction structure, and improves the accuracy of dangerous area boundary delineation and the reliability of personnel over-border evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, specifically to a construction safety assessment system based on computer vision and drone inspection. Background Technology

[0002] In existing technologies, construction safety assessment systems based on drones typically use single-view two-dimensional images for hazardous area identification and personnel boundary crossing detection. Specifically, the drone flies along a pre-set fixed route, its onboard camera acquiring single-view two-dimensional images of the construction site. Image processing algorithms are then used to extract the hazardous area boundaries and personnel pixel positions from the two-dimensional images. The planar distance between the bounding box of the personnel pixels and the hazardous area pixel boundary is calculated within the two-dimensional image plane. Based on this planar distance, it is determined whether personnel have crossed the boundary, and the safety assessment result is output.

[0003] The aforementioned existing technologies suffer from technical problems such as blind spots in single-view vision and distortion in hazard assessment due to the dynamic evolution of construction scenarios. Because construction scenarios involve numerous building structures that obstruct the view, single-view 2D images cannot capture spatial information behind these obstructions, leading to overlap or omissions in the spatial relationship between hazard areas and personnel in the 2D projection. Simultaneously, the construction process evolves dynamically over time, and changes in building structures cause spatial shifts in the actual physical boundaries of hazard areas. Static hazard boundary settings based on single-frame 2D images cannot adaptively adjust boundary positions to keep pace with construction progress, resulting in discrepancies between safety assessment results and the actual physical spatial state. Summary of the Invention

[0004] The purpose of this invention is to provide a construction safety assessment system based on computer vision and drone inspection, which can effectively solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A construction safety assessment system based on computer vision and drone inspection includes: The UAV multi-waypoint temporal image acquisition subsystem performs the operation of acquiring a series of continuous images of the construction site from multiple perspectives along a preset trajectory. The cross-modal spatiotemporal mapping construction subsystem performs the operation of processing the multi-view continuous image sequence using the structure of motion recovery algorithm to generate local three-dimensional point clouds, and the operation of projecting the extracted two-dimensional image features onto the three-dimensional space where the local three-dimensional point cloud is located to construct the cross-modal spatiotemporal mapping relationship between the two-dimensional image and the three-dimensional point cloud; The dynamic hazard boundary update subsystem performs the following operations: introducing the construction progress time sequence dimension and using the hazard area boundary of the previous time as prior knowledge; and combining the incremental change characteristics of the building structure in the local three-dimensional point cloud at the current time to modify the prior knowledge through a time-series recursive network to dynamically update the spatial boundary of the hazard area. The safety assessment output subsystem performs the operation of determining the three-dimensional coordinates of the personnel target based on the cross-modal spatiotemporal mapping relationship, and the operation of calculating the spatial distance between the three-dimensional coordinates of the personnel target and the spatial boundary of the dynamically updated danger zone and outputting the safety assessment result based on the spatial distance.

[0006] Preferably, in the operation performed by the UAV multi-waypoint temporal image acquisition subsystem, based on the spatial occlusion relationship analysis results in the local three-dimensional point cloud at the previous moment, the volume and distribution location of the occlusion blind spots not visually covered in the construction site are calculated. According to the distribution location of the occlusion blind spot volume in three-dimensional space and the size of the occlusion blind spot volume, the three-dimensional spatial coordinates of the UAV waypoint and the camera optical acquisition pitch angle at the next moment are dynamically planned to generate a non-fixed preset trajectory. The multi-view continuous image sequence is acquired along the non-fixed preset trajectory to realize a mechanism for adaptively adjusting the visual acquisition perspective based on the spatial occlusion state.

[0007] Preferably, in the operation of projecting the extracted two-dimensional image features onto the three-dimensional space where the local three-dimensional point cloud is located, the cross-modal spatiotemporal mapping construction subsystem, for the projection rays formed by the same two-dimensional image features projected from multiple different viewpoints, introduces epipolar geometric constraints to calculate the spatial intersection error distribution between the projection rays, uses multi-view feature photometric consistency verification and spatial epipolar distance measurement to iteratively eliminate the spatial intersection error, determines the unique three-dimensional spatial coordinates to eliminate depth ambiguity, and establishes the cross-modal spatiotemporal mapping relationship between the two-dimensional image features and the local three-dimensional point cloud based on the unique three-dimensional spatial coordinates.

[0008] Preferably, in the operation of the dynamic hazard boundary update subsystem combining the incremental change features of the building structure in the local three-dimensional point cloud at the current moment, the spatiotemporal difference operation after spatial coordinate registration is performed on the local three-dimensional point cloud at the previous moment and the current moment to obtain the point cloud spatial change set. Curvature anisotropic diffusion filtering is applied to the point cloud spatial change set to filter out scattered point cloud noise generated by dynamic interference with non-rigid deformation characteristics, and retain and extract spatial point cloud clusters with rigid geometric continuous topological characteristics as the incremental change features of the building structure.

[0009] Preferably, in the operation of the dynamic hazard boundary update subsystem to modify the prior knowledge through a temporal recursive network to dynamically update the spatial boundary of the hazard area, the boundary topology node sequence in the prior knowledge is input into the time-gated channel of the temporal recursive network, the incremental change features of the building structure are input into the spatial feature transformation channel of the temporal recursive network, and the offset vector field is output by fusing the product of the time-gated channel and the spatial feature transformation channel. The offset vector field is then superimposed on the boundary coordinates of the prior knowledge to obtain the spatial boundary of the dynamically updated hazard area.

[0010] Preferably, in the operation of determining the three-dimensional coordinates of the personnel target based on the cross-modal spatiotemporal mapping relationship, the safety assessment output subsystem locates the two-dimensional pixel bounding box of the personnel target based on the cross-modal spatiotemporal mapping relationship, extracts the two-dimensional coordinates of key points of the human skeleton within the two-dimensional pixel bounding box, and uses the cross-modal spatiotemporal mapping relationship to back-project the two-dimensional coordinates of the key points of the human skeleton onto three-dimensional space to form a spatial ray beam. Combining the multi-view spatial ray intersection geometric constraints and the prior knowledge of the proportion of human physical scale, the three-dimensional coordinates of the personnel target are jointly solved.

[0011] Preferably, in the operation performed by the UAV multi-waypoint time-series image acquisition subsystem, when dynamically planning the three-dimensional spatial coordinates of the UAV waypoint and the camera optical acquisition pitch angle at the next moment based on the distribution position of the occlusion blind zone volume in three-dimensional space, the system acquires the three-axis attitude angle change rate of the UAV in real time, establishes a negative correlation coupling constraint model between the three-axis attitude angle change rate of the UAV and the image acquisition exposure time, and adaptively adjusts the exposure parameters of the acquired multi-view continuous image sequence based on the negative correlation coupling constraint model to suppress motion blur during dynamic waypoint switching.

[0012] Preferably, in the operation of iteratively eliminating the spatial intersection error using multi-view feature photometric consistency verification and spatial epipolar distance measurement, weak texture regions in the two-dimensional image corresponding to the local three-dimensional point cloud are detected, local phase consistency features are extracted in the weak texture regions as illumination-invariant feature descriptors, and spatial rotation normalization is performed on the illumination-invariant feature descriptors in combination with the normal vector direction of the corresponding spatial point in the local three-dimensional point cloud. The multi-view feature photometric consistency verification is then performed using the normalized illumination-invariant feature descriptors.

[0013] Preferably, after retaining and extracting spatial point cloud clusters with rigid geometric continuous topological properties as incremental change features of the building structure, a statistical outlier removal operation based on spatial proximity is performed on the incremental change features of the building structure. The incremental change features of the building structure after removing outliers are transformed into a voxelized spatial mesh representation. Laplacian smoothing iteration processing that preserves the weights of geometric local features is applied to the edge voxels of the voxelized spatial mesh to obtain rigid building incremental features with continuous spatial boundaries and no floating isolated noise.

[0014] Preferably, in the operation of jointly solving the three-dimensional coordinates of the personnel target by combining multi-view spatial ray intersection geometric constraints and human physical scale proportion prior, a skeletal joint spatial angle constraint model based on human kinematic constraints is constructed. The three-dimensional spatial coordinates of the skeletal key points obtained by the preliminary solution of the multi-view spatial ray intersection geometric constraints are used as observation data items. The skeletal joint spatial angle constraint model and the human physical scale proportion prior are used together as spatial topological regularization constraint items. The final three-dimensional coordinates of the personnel that conform to the laws of physical motion are solved by a nonlinear least squares optimization algorithm.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention generates a local 3D point cloud from a series of continuous multi-view images collected by a UAV, and projects 2D image features onto the 3D space of the local 3D point cloud to construct a cross-modal spatiotemporal mapping relationship. This utilizes 3D spatial information to eliminate the occlusion effect of single-view blind spots on the relationship between hazardous areas and personnel positions. By introducing a construction progress time-series dimension, the hazardous area boundary from previous moments is used as prior knowledge. Combined with the incremental change characteristics of the building structure in the local 3D point cloud at the current moment, a temporal recursive network is used to correct the prior knowledge to dynamically update the spatial boundary of the hazardous area, allowing the hazardous boundary to adaptively adjust with the evolution of the construction physical structure. Based on the cross-modal spatiotemporal mapping relationship, the 3D coordinates of personnel targets are determined, the spatial distance between the 3D coordinates of personnel targets and the dynamically updated spatial boundary of the hazardous area is calculated, and a safety assessment result is output. This overcomes the bias of 2D static assessment and obtains a safety assessment result that conforms to the actual physical space state.

[0016] 2. Based on the mechanism of adaptive adjustment of visual acquisition perspective under spatial occlusion, the three-dimensional spatial coordinates of the UAV waypoint and the camera's optical acquisition pitch angle at the next moment are dynamically planned by analyzing the blind zone of point cloud occlusion, avoiding the problem of repeated blind zone acquisition caused by fixed waypoints; the spatial intersection error is iteratively eliminated by using epipolar geometric constraints and multi-view feature photometric consistency verification, eliminating the depth ambiguity when mapping two-dimensional features to three-dimensional space; curvature anisotropic diffusion filtering is applied to the spatiotemporal differential point cloud to filter out scattered noise generated by non-rigid dynamic interference objects and retain the incremental change characteristics of rigid building structures; a negative correlation coupling constraint model is established between the three-axis attitude angle change rate of the fuselage and the image acquisition exposure time to suppress motion blur during dynamic waypoint switching; local phase consistency features are extracted in weak texture regions as illumination-invariant feature descriptors and spatial rotation normalization is performed to improve the feature matching robustness under weak texture and variable illumination conditions; the three-dimensional coordinates of the personnel target are jointly solved by the skeletal joint spatial angle constraint model and the prior of human physical scale ratio as regularization constraints, and the spatial positioning results of personnel that conform to the laws of physical motion are obtained. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the overall workflow of the construction safety assessment system based on computer vision and drone inspection of the present invention. Figure 2 This is a flowchart illustrating the workflow of the cross-modal spatiotemporal mapping construction subsystem of the present invention. Figure 3 This is a flowchart of the UAV waypoint dynamic planning based on spatial occlusion status according to the present invention. Figure 4 This is a flowchart of the projection error elimination method based on epipolar geometric constraints according to the present invention. Figure 5 This is a flowchart of the building structure incremental change feature extraction process of the present invention; Figure 6 This is a flowchart of the multi-constraint joint solution of the three-dimensional coordinates of the personnel target in this invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Please refer to Figure 1This embodiment provides a construction safety assessment system based on computer vision and UAV inspection, comprising four interconnected subsystems. Each subsystem executes corresponding operations according to a preset data flow logic. The UAV multi-waypoint temporal image acquisition subsystem acquires a series of continuous images of the construction site from multiple perspectives along a preset trajectory. The preset trajectory is generated based on an initial 3D model of the construction site, covering multiple different perspectives of the main construction areas. Each image frame carries corresponding timestamp information and UAV pose data, including the UAV's 3D spatial coordinates in the world coordinate system and the camera's intrinsic and extrinsic parameters. The acquisition frequency of the multi-view continuous image sequence is adjusted according to the construction progress. When the construction progress is fast, the acquisition frequency is appropriately increased to capture rapid changes in the building structure; when the construction progress is slow, the acquisition frequency is appropriately decreased to reduce the amount of data processing.

[0020] refer to Figure 2 The cross-modal spatiotemporal mapping construction subsystem utilizes the Structure in Motion (SIM) algorithm to process multi-view continuous image sequences and generate local 3D point clouds. Specifically, the SIM algorithm's processing includes four stages: feature extraction and matching, camera pose estimation, triangulation to generate point clouds, and bundle adjustment optimization. In the feature extraction and matching stage, scale-invariant feature transform (SMT) feature points are extracted from each image. Each feature point contains a 128-dimensional descriptor, which characterizes the local image texture information surrounding the feature point. Feature points from adjacent images are matched using a fast nearest neighbor search library to obtain an initial set of matching pairs. Subsequently, the initial set of matching pairs is processed using a random sampling consensus algorithm to remove erroneous matches, resulting in an accurate set of matching pairs. In the camera pose estimation stage, the eight-point method is used to estimate the fundamental matrix between adjacent images based on the accurate set of matching pairs, and then the camera's rotation matrix and translation vector are decomposed to determine the pose of each camera in the world coordinate system. In the triangulation to generate point clouds stage, for each pair of matched feature points, the 3D spatial coordinates corresponding to the feature point are calculated using the projection matrices of the two cameras, generating an initial local 3D point cloud. In the bundle adjustment optimization stage, a reprojection error objective function is constructed, and a nonlinear optimization algorithm is used to simultaneously optimize the intrinsic and extrinsic parameters of all cameras and the coordinates of all 3D point clouds, minimizing the reprojection error and improving the accuracy of the local 3D point clouds.

[0021] The camera projection matrix describes the projection transformation from a point in three-dimensional space to a two-dimensional image plane, and its mathematical expression is: Where P is a 3×4 camera projection matrix, and K is a 3×3 camera intrinsic parameter matrix, in the form of: and These are the focal lengths of the camera along the x and y axes, respectively. and Let be the coordinates of the camera principal point on the image plane; R is a 3×3 rotation matrix that satisfies the orthogonality constraint. And determinant , describes the rotation relationship between the camera coordinate system and the world coordinate system; t is a 3×1 translation vector, describing the position of the origin of the camera coordinate system in the world coordinate system.

[0022] The projection relationship between a point in three-dimensional space and a feature point in a two-dimensional image is expressed by the following formula: in, These are the homogeneous pixel coordinates of feature points in a two-dimensional image. The coordinates of a point in three-dimensional space are homogeneous coordinates in the world coordinate system.

[0023] The cross-modal spatiotemporal mapping construction subsystem projects extracted 2D image features onto the 3D space containing local 3D point clouds, constructing a cross-modal spatiotemporal mapping relationship between the 2D image and the 3D point cloud. Specifically, for each 2D image feature point, its corresponding 3D space point is found through the above projection relationship, establishing a corresponding mapping relationship. The cross-modal spatiotemporal mapping relationship is stored in the form of a mapping table, which contains multiple entries, each corresponding to a 2D image feature point and a 3D space point. Each entry includes fields such as image frame index, feature point unique identifier, feature point descriptor, feature point pixel coordinates, corresponding 3D point cloud unique identifier, 3D point coordinates, timestamp, and matching confidence. Through this mapping table, bidirectional fast mapping between 2D image features and 3D space points can be achieved; that is, the corresponding 3D space coordinates can be quickly retrieved based on the pixel coordinates of the 2D image feature point, and the projection position of the 3D space point in each image frame can be quickly retrieved based on the coordinates of the 3D space point.

[0024] The dynamic hazard boundary update subsystem introduces a construction progress time-series dimension, using the hazard zone boundary from previous time steps as prior knowledge. This prior knowledge is then combined with the incremental changes in the building structure within the current local 3D point cloud. A temporal recursive network is used to correct this prior knowledge, dynamically updating the spatial boundary of the hazard zone. Specifically, the initial hazard zone boundary is pre-defined based on construction drawings and building construction safety specifications, consisting of a closed 3D surface composed of a series of 3D topological nodes. Each topological node contains attributes such as 3D spatial coordinates, boundary type, safety threshold, and hazard threshold. During construction, the local 3D point cloud generated at each time step is spatially registered with the local 3D point cloud from previous time steps to obtain the incremental changes in the building structure relative to previous time steps. The sequence of hazard zone boundary topological nodes from previous time steps is input into the temporal recursive network as prior knowledge, along with the incremental changes in the building structure at the current time step. The temporal recursive network outputs an offset vector field of the hazard zone boundary. This offset vector field is then superimposed onto the hazard zone boundary coordinates from previous time steps to obtain the dynamically updated spatial boundary of the hazard zone at the current time step.

[0025] The safety assessment output subsystem determines the 3D coordinates of personnel targets based on cross-modal spatiotemporal mapping, calculates the spatial distance between the 3D coordinates of the personnel targets and the dynamically updated spatial boundary of the danger zone, and outputs the safety assessment result based on the spatial distance. Specifically, firstly, a target detection algorithm is used to process a multi-view continuous image sequence to detect personnel targets in the images and generate 2D pixel bounding boxes for the personnel targets. Then, based on the cross-modal spatiotemporal mapping, the center coordinates of the 2D pixel bounding boxes are back-projected into 3D space to obtain the preliminary 3D coordinates of the personnel targets. Subsequently, the shortest spatial distance from the preliminary 3D coordinates of the personnel targets to the dynamically updated spatial boundary of the danger zone is calculated. If the shortest spatial distance is greater than a preset safety threshold, the personnel target is determined to be in a safe state; if the shortest spatial distance is less than or equal to the safety threshold but greater than a preset danger threshold, the personnel target is determined to be in a warning state; if the shortest spatial distance is less than or equal to the danger threshold, the personnel target is determined to be in a dangerous state. The safety assessment results include unique identifiers of personnel targets, 3D coordinates, shortest distance to hazardous areas, safety status, unique identifiers of hazardous areas, boundary coordinates, and update time. These are visualized and overlaid on a 3D model of the construction site, while corresponding alerts are issued based on the safety status. Table 1 shows a statistical comparison of the errors between 2D image feature types and 3D point cloud projection mapping.

[0026] Table 1. Comparison of Statistical Data on Two-Dimensional Image Feature Types and Three-Dimensional Point Cloud Projection Mapping Errors Table 1 shows the statistical results of 3D point cloud projection mapping errors corresponding to different types of 2D image features under the same image resolution and local 3D point cloud density. The average projection error is the average of the reprojection errors of all matched feature points, the matching success rate is the proportion of correctly matched feature points to the total number of extracted feature points, and the processing time is the average time required for feature extraction and matching in a single frame. As can be seen from the table, SIFT features have the smallest average projection error and the highest matching success rate, but the processing time is relatively long; ORB features have the shortest processing time, but the average projection error is relatively large and the matching success rate is relatively low; SURF features have performance between SIFT and ORB. In practical applications, the appropriate feature type can be selected based on different requirements for accuracy and real-time performance. When high accuracy is required, SIFT features should be chosen; when high real-time performance is required, ORB features should be chosen; and when both accuracy and real-time performance need to be considered, SURF features should be chosen.

[0027] This embodiment generates local 3D point clouds through a multi-view continuous image sequence, constructing a cross-modal spatiotemporal mapping relationship between 2D images and 3D point clouds. This converts information from 2D images into 3D spatial information, eliminating the occlusion effect of single-view blind spots on the relationship between hazardous areas and personnel positions. By introducing a construction progress time-series dimension, the hazardous area boundaries from previous moments are used as prior knowledge. Combined with the incremental changes in the building structure at the current moment, the spatial boundaries of the hazardous areas are dynamically updated, allowing the hazardous boundaries to adaptively adjust with the evolution of the construction physical structure. Safety assessment based on 3D spatial distance overcomes the bias of 2D static assessments, obtaining safety assessment results consistent with the actual physical space state.

[0028] In another embodiment, based on the technical solution of the above basic embodiment, the UAV multi-waypoint temporal image acquisition subsystem further performs the operation of adaptively adjusting the visual acquisition perspective based on the spatial occlusion state, and the cross-modal spatiotemporal mapping construction subsystem further performs the projection error elimination operation based on epipolar geometric constraints and multi-view feature photometric consistency verification.

[0029] refer to Figure 3The UAV multi-waypoint temporal image acquisition subsystem, based on the spatial occlusion relationship analysis results of the local 3D point cloud at the previous moment, calculates the volume and distribution location of the occlusion blind spots not visually covered in the construction site. According to the distribution location and size of the occlusion blind spot volume in 3D space, it dynamically plans the 3D spatial coordinates of the UAV waypoints and the camera's optical acquisition pitch angle for the next moment, generating a non-fixed preset trajectory, and acquiring a multi-view continuous image sequence along the non-fixed preset trajectory. Specifically, the local 3D point cloud generated at the previous moment is first voxelized, dividing the 3D space into a uniform cubic voxel grid. The side length of each voxel is determined according to the required spatial resolution. For each voxel, it is determined whether it is covered by at least one camera viewpoint. If covered, it is marked as a visible voxel; if not covered, it is marked as an invisible voxel, i.e., an occlusion blind spot. The total volume of the occlusion blind spot is the sum of the volumes of all invisible voxels, calculated using the following formula: in, To cover the total volume of the blind spot, N is the total number of voxels in the 3D spatial mesh. Let i be the volume of the i-th voxel. For binary labeled variables, when the i-th voxel is an invisible voxel... The value is 1 if it is not 1, otherwise the value is 0.

[0030] After obtaining the volume and distribution of the occlusion blind spots, a greedy algorithm is used to dynamically plan the UAV waypoints for the next time step. First, all occlusion blind spots are sorted in descending order of volume, prioritizing the largest ones. For each blind spot, its geometric center coordinates are calculated, and the UAV waypoint position is determined, ensuring the camera's field of view completely covers the blind spot while maintaining a safe distance between the UAV and existing structures. The camera's optical acquisition pitch angle is adjusted according to the height of the blind spot, ensuring the geometric center of the blind spot is at the center of the camera's field of view. After planning a waypoint, the remaining volume of the blind spot is recalculated. If the remaining volume is greater than a preset threshold, the next waypoint is planned, continuing until the remaining volume is less than the preset threshold. This method generates a non-fixed preset trajectory, enabling an adaptive adjustment of the visual acquisition angle based on spatial occlusion conditions, avoiding the problem of repeated blind spot acquisition caused by fixed waypoints.

[0031] The UAV multi-waypoint temporal image acquisition subsystem dynamically plans the 3D spatial coordinates of the UAV waypoints and the camera's optical acquisition pitch angle for the next moment. It then acquires the real-time rate of change of the UAV's three-axis attitude angles and establishes a negative correlation coupling constraint model between the rate of change of the three-axis attitude angles and the image acquisition exposure time. Based on this model, it adaptively adjusts the exposure parameters for acquiring continuous multi-view image sequences to suppress motion blur during dynamic waypoint switching. Specifically, the UAV's three-axis attitude angles include roll, pitch, and yaw, corresponding to the rotation angles of the fuselage around the x, y, and z axes, respectively. The rate of change of the three-axis attitude angles is acquired in real-time by the UAV's inertial measurement unit and denoted as... , and The unit is radians per second. The magnitude of the rate of change of attitude angle is calculated using the following formula: Wherein, ω is the magnitude of the rate of change of the three-axis attitude angle of the fuselage, which characterizes the intensity of the overall motion of the UAV fuselage.

[0032] A negative correlation coupling constraint model is established between exposure time and the magnitude of the rate of change of attitude angle, and its expression is: Where T is the image acquisition and exposure time at the current moment. The preset maximum exposure time is denoted by k, which is a scaling factor determined based on the camera's imaging characteristics and the drone's motion characteristics. When the drone's attitude angle change rate increases, the intensity of the drone's motion increases, and the exposure time decreases accordingly, thereby reducing the camera's displacement within the exposure time and suppressing motion blur. Conversely, when the drone's attitude angle change rate decreases, the intensity of the drone's motion decreases, and the exposure time increases accordingly, thereby improving the image's signal-to-noise ratio and image quality.

[0033] refer to Figure 4The cross-modal spatiotemporal mapping construction subsystem, in the operation of projecting extracted 2D image features onto the 3D space of a local 3D point cloud, introduces epipolar geometric constraints to calculate the spatial intersection error distribution between projection rays formed by projecting the same 2D image feature from multiple different viewpoints. It iteratively eliminates spatial intersection errors using multi-view feature photometric consistency verification and spatial epipolar distance measurement, determining a unique 3D spatial coordinate to eliminate depth ambiguity. Based on this unique 3D spatial coordinate, a cross-modal spatiotemporal mapping relationship between the 2D image feature and the local 3D point cloud is established. Specifically, for the same 2D image feature point, there are corresponding matching points in m images from different viewpoints. Each matching point corresponds to a spatial projection ray originating from the camera projection center. Ideally, these m projection rays should converge at the same point, i.e., the 3D spatial point corresponding to the feature point. However, due to feature matching errors and camera pose estimation errors, these projection rays usually do not strictly converge at the same point, but rather form a spatial error region.

[0034] Epipolar geometric constraints describe the projection geometry between two cameras with different viewpoints. For two cameras C1 and C2, their projection centers are O1 and O2, respectively. The projections of a 3D point X onto the image planes of the two cameras are respectively... and ,but Must be located On the corresponding epipolar line in the C2 image plane, and vice versa. The epipolar geometric constraint can be represented by the fundamental matrix F, whose mathematical expression is: in, and Let X be the homogeneous pixel coordinates of a point X in 3D space on the image planes of the two cameras, and F be the 3×3 fundamental matrix with a rank of 2, containing the relative pose information between the two cameras.

[0035] The spatial intersection error between projected rays is calculated using epipolar geometric constraints. For any two projected rays, the shortest distance between them is calculated; this shortest distance is the intersection error between the two rays. For m projected rays, the shortest distance between all pairs of rays is calculated, and their average value is taken as the spatial intersection error of the feature point. If the spatial intersection error is greater than a preset threshold, the feature point is considered to have depth ambiguity and requires further processing.

[0036] Multi-view feature photometric consistency verification and spatial epipolar distance measurement are used to iteratively eliminate spatial intersection errors. The principle of multi-view feature photometric consistency verification is that the projections of the same 3D spatial point under different views should have similar photometric characteristics. For the corresponding pixels of the same feature point under different views, the difference in their grayscale values ​​is calculated. If the difference is greater than a preset threshold, the matching pair is considered erroneous and is eliminated. The principle of spatial epipolar distance measurement is that the matching point should be located on the corresponding epipolar line. The distance from the matching point to the corresponding epipolar line is calculated. If the distance is greater than a preset threshold, the matching pair is considered erroneous and is eliminated. This process is iteratively executed until the spatial intersection error of all remaining projection rays is less than the preset threshold. Then, the centroid of the spatial intersection point of these rays is taken as the unique 3D spatial coordinates corresponding to the feature point.

[0037] The cross-modal spatiotemporal mapping construction subsystem, in its iterative elimination of spatial intersection errors using multi-view feature photometric consistency verification and spatial epipolar distance measurement, detects weakly textured regions in the 2D image corresponding to a local 3D point cloud. Within these weakly textured regions, local phase consistency features are extracted as illumination-invariant feature descriptors. Combined with the normal vector direction of the corresponding spatial points in the local 3D point cloud, spatial rotation normalization is performed on the illumination-invariant feature descriptors. Finally, the normalized illumination-invariant feature descriptors are used to perform multi-view feature photometric consistency verification. Specifically, weakly textured region detection is achieved by calculating the gray-level variance of local image regions. The image is divided into equally sized local windows, and the gray-level variance within each window is calculated. If the variance is less than a preset threshold, the region corresponding to that window is determined to be a weakly textured region.

[0038] Local phase consistency features are extracted within weakly textured regions. Based on the image's phase information, these features are robust to changes in illumination. The formula for calculating local phase consistency features is as follows: in, For pixels Local phase consistency value at that location, 's' is the direction index of the Gabor filter, and 's' is the scale index of the Gabor filter. For weighting coefficients of different directions and scales, For Gabor filters in direction and the response magnitude at scale s, For the corresponding response phase, The average phase across all directions and scales. It is a very small constant used to prevent the denominator from being zero.

[0039] Obtain the normal vector n of the 3D space point corresponding to the feature point. The normal vector is calculated using the coordinates of this point and its neighboring points. Rotate the normal vector n until it coincides with the Z-axis of the world coordinate system to obtain the rotation matrix. Then, the local phase-consistency feature descriptors are subjected to the same rotation transformation to obtain normalized feature descriptors. Spatial rotation normalization eliminates the influence of rotation from different viewpoints, improving the robustness of feature matching. Using the normalized illumination-invariant feature descriptors to perform multi-view feature photometric consistency verification effectively improves feature matching accuracy in weakly textured regions and under varying illumination conditions. Table 2 shows a comparison of UAV waypoint planning results under different occlusion blind zone distribution patterns.

[0040] Table 2 Comparison of UAV waypoint planning results under different blind spot distribution patterns Table 2 shows the UAV waypoint planning results for different blind zone distribution patterns under the same initial waypoint settings and total blind zone volume. A centralized distribution pattern refers to blind zones concentrated in a small area; a distributed distribution pattern refers to blind zones scattered across multiple different areas; and a hybrid distribution pattern refers to the simultaneous presence of both centralized and distributed blind zones. The blind zone coverage rate is the complement of the ratio of the remaining blind zone volume after waypoint planning to the initial blind zone volume. The table shows that the centralized distribution pattern requires the fewest waypoints, has the shortest total flight distance, the shortest data collection time, and the highest blind zone coverage rate; the distributed distribution pattern requires the most waypoints, has the longest total flight distance, the longest data collection time, and the lowest blind zone coverage rate; the hybrid distribution pattern's indicators fall between those of the centralized and distributed patterns. These results indicate that the distribution pattern of blind zones significantly impacts the efficiency and effectiveness of waypoint planning. Blind zones in a centralized distribution are easier to cover, while blind zones in a distributed distribution require more waypoints and longer data collection times.

[0041] This embodiment improves the visual coverage of the construction site and reduces the impact of blind spots on safety assessment by adaptively adjusting the waypoints and acquisition perspectives of the UAV based on spatial occlusion. By establishing a negatively correlated coupling constraint model between the fuselage attitude angle change rate and exposure time, adaptive adjustment of exposure parameters is achieved, effectively suppressing motion blur during dynamic waypoint switching and improving image acquisition quality. By introducing epipolar geometric constraints to calculate spatial intersection errors and using multi-view feature photometric consistency verification and spatial epipolar distance measurement for iterative elimination, depth ambiguity in mapping two-dimensional features to three-dimensional space is eliminated, improving the accuracy of cross-modal spatiotemporal mapping. By extracting local phase consistency features as illumination-invariant feature descriptors in weakly textured regions and performing spatial rotation normalization, the robustness of feature matching under weak texture and varying illumination conditions is improved, further enhancing the reliability of cross-modal spatiotemporal mapping.

[0042] In another embodiment, based on the technical solution of the above embodiments, the dynamic hazard boundary update subsystem further performs the operation of extracting incremental change features of building structure based on curvature anisotropic diffusion filtering and the operation of hazard boundary correction based on dual-channel temporal recursive network, and the safety assessment output subsystem further performs the operation of determining the three-dimensional coordinates of personnel based on the joint solution of human skeleton key points and multiple constraints.

[0043] refer to Figure 5 The dynamic hazard boundary update subsystem, in its operation of combining the incremental change characteristics of the building structure in the local 3D point cloud at the current moment, performs spatiotemporal difference operations on the local 3D point clouds at previous and current moments after spatial coordinate registration, obtains a set of spatial changes in the point cloud, and applies curvature anisotropic diffusion filtering to this set of spatial changes to filter out scattered point cloud noise generated by dynamic interference with non-rigid deformation characteristics, retaining and extracting spatial point cloud clusters with rigid geometric continuous topological characteristics as the incremental change characteristics of the building structure. Specifically, firstly, the iterative nearest point algorithm is used to register the local 3D point cloud at the current moment with the local 3D point cloud at the previous moment in terms of spatial coordinates, obtaining the transformation matrix of the current moment point cloud relative to the previous moment point cloud, transforming the current moment point cloud into the coordinate system of the previous moment, and achieving spatial alignment of the point clouds at the two moments.

[0044] The spatiotemporal difference operation is performed. For each point in the point cloud at the current time step, its k nearest neighbors in the point cloud at previous time steps are found, and the average distance from the point to these k nearest neighbors is calculated. If the average distance is greater than a preset difference threshold, the point is considered a spatially changed point and is added to the point cloud spatial change set. The point cloud spatial change set contains all points whose spatial location has changed between two time steps. These points include those caused by changes in building structure, as well as those caused by dynamic disturbances (such as people, vehicles, and construction machinery).

[0045] Curvature anisotropic diffusion filtering is applied to the spatial variation set of the point cloud. Curvature anisotropic diffusion filtering is a point cloud smoothing algorithm based on local geometric features, which can smooth noise while preserving the edge and detail features of the point cloud. The core idea of ​​this algorithm is to adjust the diffusion coefficient according to the local curvature of the point cloud. In edge regions with greater curvature, the diffusion coefficient is smaller, thus preserving edge features; in flat regions with less curvature, the diffusion coefficient is larger, thus smoothing noise. The iterative formula for curvature anisotropic diffusion filtering is: in, Let be the 3D coordinates of the i-th point at the t-th iteration, and λ be the iteration step size, used to control the smoothness of each iteration. Let i be the set of neighborhood points of the i-th point. Let be the diffusion weight between the i-th point and the j-th point.

[0046] Diffusion weight It is determined by the difference in Gaussian curvature between two points, and its calculation formula is: in, and Let be the Gaussian curvatures of the i-th and j-th points, respectively. The standard deviation of the Gaussian kernel is used to control the influence of curvature difference on diffusion weight. When the Gaussian curvature difference between two points is small, the diffusion weight is large, and the diffusion effect between the two points is strong; when the Gaussian curvature difference between two points is large, the diffusion weight is small, and the diffusion effect between the two points is weak.

[0047] Point clouds generated by dynamic disturbances with non-rigid deformation characteristics are scattered, have large curvature variations, and are topologically discontinuous. After curvature anisotropic diffusion filtering, these points are gradually smoothed and filtered out. Point clouds generated by building structures with rigid geometric continuity and topological characteristics, however, have continuous curvature distributions and topological structures, and are preserved after filtering. In this way, changes in building structures and dynamic disturbances can be effectively distinguished, and accurate incremental change features of building structures can be extracted.

[0048] The dynamic hazard boundary update subsystem, after retaining and extracting spatial point cloud clusters with rigid geometric continuity as incremental change features of the building structure, performs a statistical outlier removal operation based on spatial proximity on these incremental change features. The outlier-removed incremental change features are then transformed into a voxelized spatial mesh representation. Laplace smoothing iteration, which preserves the weights of geometrically local features, is applied to the edge voxels of the voxelized spatial mesh to obtain rigid building incremental features with continuous spatial boundaries and no floating or isolated noise. Specifically, the statistical outlier removal operation based on spatial proximity determines whether a point is an outlier by calculating the average distance from each point to its k nearest neighbors. For each point, the average distance to its k nearest neighbors is calculated, and then the mean and standard deviation of the average distances of all points are calculated. If the average distance of a point is greater than the mean plus twice the standard deviation, the point is determined to be an outlier and removed.

[0049] The point cloud after outlier removal is transformed into a voxelized spatial mesh representation. The 3D space is divided into a uniform cubic voxel mesh, with the side length of each voxel determined according to the required spatial resolution. For each voxel, if it contains at least one point, it is marked as an occupied voxel; otherwise, it is marked as an empty voxel. This voxelized spatial mesh representation transforms discrete point cloud data into continuous spatial mesh data, facilitating subsequent boundary extraction and smoothing processes.

[0050] Laplacian smoothing is applied iteratively to the edge voxels of the voxelized spatial mesh, preserving weights that maintain local geometric features. An edge voxel is an occupied voxel with at least one empty neighbor voxel. Laplacian smoothing achieves smoothing by shifting the center coordinate of each edge voxel towards the average of the center coordinates of its neighboring voxels. To preserve local geometric features, local feature weights are introduced: smaller smoothing weights are assigned to edge voxels with greater curvature, and larger smoothing weights are assigned to edge voxels with less curvature. The iterative formula for Laplacian smoothing is: in, Let be the center coordinates of the i-th edge voxel during the t-th iteration. The global smoothing coefficient. Let be the local feature weight of the i-th edge voxel. Let i be the set of neighboring voxels of the i-th edge voxel. The number of neighborhood voxels. Local feature weights. It is inversely proportional to the curvature of the edge voxels; the greater the curvature, the higher the curvature. The smaller the value, the less smooth the surface; the smaller the curvature, The larger the value, the smoother the surface. In this way, the local geometric features of the building structure can be preserved while smoothing the voxel mesh boundaries, resulting in rigid building incremental features with continuous spatial boundaries and no floating or isolated noise.

[0051] The dynamic hazard boundary update subsystem, in its operation of dynamically updating the spatial boundary of the hazard area by modifying prior knowledge through a temporal recursive network, inputs the boundary topology node sequence from the prior knowledge into the temporal gating channel of the temporal recursive network, and inputs the incremental change features of the building structure into the spatial feature transformation channel of the temporal recursive network. Through the product and fusion of the temporal gating channel and the spatial feature transformation channel, an offset vector field is output. This offset vector field is then superimposed onto the boundary coordinates of the prior knowledge to obtain the dynamically updated spatial boundary of the hazard area. Specifically, the temporal recursive network adopts a dual-channel structure, including a temporal gating channel and a spatial feature transformation channel. The temporal gating channel consists of gated recurrent units, used to process the hazard area boundary topology node sequence from previous time steps, capturing the temporal evolution pattern of the hazard boundary. The spatial feature transformation channel consists of a convolutional neural network, used to process the incremental change features of the building structure at the current time step, extracting feature information of spatial structural changes.

[0052] The sequence of topological nodes representing the hazardous area boundary at previous time steps is input into the time-gated channel. The gated recurrent unit processes this sequence and outputs a boundary feature sequence containing temporal evolution information. The incremental rigid building features at the current time step are input into the spatial feature transformation channel. A convolutional neural network performs multi-layer convolution and pooling processing on these features, outputting a spatial feature map containing information on spatial structural changes. The boundary feature sequence output from the time-gated channel and the spatial feature map output from the spatial feature transformation channel are fused element-wise to obtain the offset vector of each boundary topological node, forming an offset vector field. The formula for calculating the offset vector field is: in, Let be the boundary offset vector field at time t. This is the sequence of topological nodes representing the boundary of the danger zone at time t-1. For the time-gated channel gated loop unit network, Let be the incremental characteristics of the rigid building at time t. A convolutional neural network for transforming spatial features through channels. This is an element-wise multiplication operation.

[0053] The offset vector field is superimposed onto the boundary coordinates of the danger zone from the previous time step to obtain the spatial boundary of the dynamically updated danger zone at the current time step. The calculation formula is as follows: in, This represents the dynamically updated sequence of topological nodes representing the hazardous area boundary at time t. This dual-channel product fusion method allows for the simultaneous utilization of both the temporal evolution information of the hazardous boundary and the spatial variation information of the building structure, enabling accurate dynamic updates of the hazardous area boundary.

[0054] refer to Figure 6 In the operation of determining the 3D coordinates of personnel targets based on cross-modal spatiotemporal mapping, the safety assessment output subsystem locates the 2D pixel bounding box of the personnel target based on the cross-modal spatiotemporal mapping. Within the 2D pixel bounding box, it extracts the 2D coordinates of key human skeletal points. Using the cross-modal spatiotemporal mapping, it back-projects these 2D coordinates onto 3D space to form a spatial ray beam. Combining multi-view spatial ray intersection geometric constraints and prior knowledge of human physical scale proportions, it jointly solves for the 3D coordinates of the personnel target. Specifically, firstly, a target detection algorithm is used to process a multi-view continuous image sequence to detect personnel targets in the images and generate 2D pixel bounding boxes for the personnel targets. Then, a human pose estimation algorithm is used to extract the 2D coordinates of 17 key human skeletal points within the 2D pixel bounding boxes, including the head, neck, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, and left and right ankles.

[0055] For the two-dimensional coordinates of each skeletal keypoint, a cross-modal spatiotemporal mapping relationship is used to back-project them into three-dimensional space, forming a spatial ray originating from the camera projection center. For the same human target, corresponding skeletal keypoints from multiple viewpoints will form multiple spatial rays, and the ideal intersection of these rays is the three-dimensional coordinate of the keypoint. However, due to feature extraction and mapping errors, these rays usually do not strictly intersect at the same point, requiring a joint solution combining the geometric constraints of multi-view spatial ray intersection and the prior knowledge of human physical scale proportions.

[0056] The safety assessment output subsystem, in its operation of jointly solving for the 3D coordinates of a person's target using multi-view spatial ray intersection geometric constraints and human physical scale proportion priors, constructs a skeletal joint spatial angle constraint model based on human kinematic constraints. The 3D spatial coordinates of key skeletal points obtained from the initial solution of the multi-view spatial ray intersection geometric constraints are used as observation data. The skeletal joint spatial angle constraint model and the human physical scale proportion priors are used together as spatial topological regularization constraints. A nonlinear least squares optimization algorithm is then used to jointly solve for the final 3D coordinates of the person, conforming to the laws of physical motion. Specifically, the human kinematic constraint model, based on the anatomical structure of the human skeleton, defines the allowable range of motion for each joint. For example, the elbow joint's flexion angle ranges from 0 to 145 degrees, the shoulder joint's flexion angle ranges from 0 to 180 degrees, and the hip joint's flexion angle ranges from 0 to 120 degrees. For each joint, its current angle is calculated; if the angle exceeds the allowable range, a corresponding penalty is applied.

[0057] The prior knowledge of human physical scale proportions is based on the average body proportions of adults, defining the length ratios between various skeletal segments. For example, the ratio of an adult's height to arm length is approximately 1:1, the ratio of thigh length to calf length is approximately 1:0.8, and the ratio of upper arm length to forearm length is approximately 1:0.75. These ratios are used as prior constraints to ensure that the solved human skeletal dimensions conform to physical laws.

[0058] The objective function for nonlinear least squares optimization is constructed as follows: Where E is the objective function value, and M is the number of skeletal keypoints. Let π be the two-dimensional observation coordinates of the i-th keypoint, and π be the camera projection function. Let i be the three-dimensional coordinates of the i-th key point. and is the regularization coefficient, used to balance the weights of the observed data terms and the regularization constraint terms; K is the number of joints. Let be the current angle of the j-th joint. Let the center angle of the j-th joint be . Let L represent the allowable range of motion of the j-th joint, and L be the number of bone segments. This represents the current length of the l-th bone segment. Let be the prior length of the l-th bone segment.

[0059] The first term of the objective function is the observation data term, which measures the error between the 3D coordinates projected onto the 2D image plane and the actual observed coordinates. The second term is the joint angle constraint term, which penalizes joint angles that exceed the allowed range of motion. The third term is the bone length constraint term, which penalizes bone segment lengths that deviate from the prior proportions. By minimizing this objective function, the final 3D coordinates of the person can be obtained, which simultaneously satisfy multi-view geometric constraints, human kinematic constraints, and the prior physical scale proportions.

[0060] After obtaining the final three-dimensional coordinates of the personnel target, the shortest spatial distance between these coordinates and the dynamically updated spatial boundary of the hazardous area is calculated. The spatial boundary of the hazardous area is a closed three-dimensional surface, and the shortest spatial distance is the shortest Euclidean distance from the center coordinates of the personnel target to this three-dimensional surface. Based on the comparison between this shortest spatial distance and the preset safety threshold and hazard threshold, the corresponding safety assessment result is output. Table 3 shows a comparison of the hazardous boundary update error under different temporal recursive network input combinations.

[0061] Table 3 Comparison of Dangerous Boundary Update Errors under Different Temporal Recursive Network Input Combinations Table 3 shows the hazard boundary update results for different combinations of temporal recursive network inputs under the same construction progress and local 3D point cloud density. The average boundary offset error is the average distance between the updated hazard boundary and the actual hazard boundary, and the boundary topology preservation rate is the proportion by which the updated hazard boundary retains the original topology. The table shows that when only the preceding boundary is used as input, the average boundary offset error is large, and the boundary topology preservation rate is low, failing to adapt to changes in the building structure. When only incremental features are used as input, the average boundary offset error is even larger, and the boundary topology preservation rate is even lower, easily leading to boundary topology confusion. When both preceding boundaries and incremental features are used as input, the average boundary offset error is the smallest, and the boundary topology preservation rate is the highest, accurately updating hazard boundaries to follow changes in the building structure while maintaining the boundary topology. These results demonstrate that the dual-channel temporal recursive network, by fusing the temporal information of the preceding boundary and the spatial information of the incremental features, can significantly improve the accuracy and reliability of hazard boundary updates.

[0062] This embodiment effectively filters out scattered point cloud noise generated by non-rigid dynamic interference by applying curvature anisotropic diffusion filtering to the spatiotemporal difference point cloud, accurately extracting the incremental change features of building structures with rigid geometric continuity topology. Through statistical outlier removal, voxelization, and Laplacian smoothing that preserves local geometric features, the quality of the incremental change features of the building structures is further optimized, resulting in rigid building incremental features with continuous spatial boundaries and no suspended isolated noise. Dynamic updates of hazardous boundaries are achieved through a dual-channel temporal recursive network, enabling the hazardous boundaries to accurately adapt to the construction progress while maintaining the topological structure of the boundaries. By combining multi-view spatial ray intersection geometric constraints, human kinematic constraints, and prior knowledge of human physical scale proportions, a nonlinear least squares optimization algorithm is used to jointly solve for the three-dimensional coordinates of personnel targets, obtaining high-precision personnel spatial positioning results that conform to the laws of physical motion, further improving the accuracy of safety assessment.

Claims

1. A building construction safety assessment system based on computer vision and unmanned aerial vehicle inspection, characterized in that, include: The UAV multi-waypoint temporal image acquisition subsystem performs the operation of acquiring a series of continuous images of the construction site from multiple perspectives along a preset trajectory. The cross-modal spatiotemporal mapping construction subsystem performs the operation of processing the multi-view continuous image sequence using the structure of motion recovery algorithm to generate local three-dimensional point clouds, and the operation of projecting the extracted two-dimensional image features onto the three-dimensional space where the local three-dimensional point cloud is located to construct the cross-modal spatiotemporal mapping relationship between the two-dimensional image and the three-dimensional point cloud; The dynamic hazard boundary update subsystem performs the following operations: introducing the construction progress time sequence dimension and using the hazard area boundary of the previous time as prior knowledge; and combining the incremental change characteristics of the building structure in the local three-dimensional point cloud at the current time to modify the prior knowledge through a time-series recursive network to dynamically update the spatial boundary of the hazard area. The safety assessment output subsystem performs the operation of determining the three-dimensional coordinates of the personnel target based on the cross-modal spatiotemporal mapping relationship, and the operation of calculating the spatial distance between the three-dimensional coordinates of the personnel target and the spatial boundary of the dynamically updated danger zone and outputting the safety assessment result based on the spatial distance.

2. The construction safety assessment system based on computer vision and UAV inspection according to claim 1, characterized in that, In the operation performed by the UAV multi-waypoint time-series image acquisition subsystem, based on the spatial occlusion relationship analysis results in the local three-dimensional point cloud at the previous moment, the volume and distribution location of the occlusion blind spots not visually covered in the construction site are calculated. According to the distribution location of the occlusion blind spot volume in three-dimensional space and the size of the occlusion blind spot volume, the three-dimensional spatial coordinates of the UAV waypoint and the camera optical acquisition pitch angle at the next moment are dynamically planned to generate a non-fixed preset trajectory, and the multi-view continuous image sequence is acquired along the non-fixed preset trajectory.

3. The construction safety assessment system based on computer vision and UAV inspection according to claim 1, characterized in that, The cross-modal spatiotemporal mapping construction subsystem executes the operation of projecting the extracted two-dimensional image features onto the three-dimensional space where the local three-dimensional point cloud is located. For the projection rays formed by projecting the same two-dimensional image feature from multiple different viewpoints, epipolar geometric constraints are introduced to calculate the spatial intersection error distribution between the projection rays. The spatial intersection error is iteratively eliminated using multi-view feature photometric consistency verification and spatial epipolar distance measurement to determine a unique three-dimensional spatial coordinate that eliminates depth ambiguity. Based on the unique three-dimensional spatial coordinate, the cross-modal spatiotemporal mapping relationship between the two-dimensional image feature and the local three-dimensional point cloud is established.

4. The construction safety assessment system based on computer vision and UAV inspection according to claim 1, characterized in that, In the operation of the dynamic hazard boundary update subsystem, which combines the incremental change features of the building structure in the local three-dimensional point cloud at the current moment, a spatiotemporal difference operation is performed on the local three-dimensional point cloud after spatial coordinate registration between the previous moment and the current moment to obtain a set of spatial changes in the point cloud. Curvature anisotropic diffusion filtering is applied to the set of spatial changes in the point cloud to filter out scattered point cloud noise generated by dynamic interference with non-rigid deformation characteristics, and to retain and extract spatial point cloud clusters with rigid geometric continuous topological characteristics as the incremental change features of the building structure.

5. The construction safety assessment system based on computer vision and UAV inspection according to claim 1, characterized in that, In the operation of the dynamic hazard boundary update subsystem to modify the prior knowledge through a temporal recursive network to dynamically update the spatial boundary of the hazard area, the boundary topology node sequence in the prior knowledge is input into the time-gated channel of the temporal recursive network, and the incremental change features of the building structure are input into the spatial feature transformation channel of the temporal recursive network. The offset vector field is output by fusing the product of the time-gated channel and the spatial feature transformation channel. The offset vector field is superimposed on the boundary coordinates of the prior knowledge to obtain the spatial boundary of the dynamically updated hazard area.

6. The construction safety assessment system based on computer vision and UAV inspection according to claim 1, characterized in that, In the operation of determining the three-dimensional coordinates of the personnel target based on the cross-modal spatiotemporal mapping relationship, the safety assessment output subsystem locates the two-dimensional pixel bounding box of the personnel target based on the cross-modal spatiotemporal mapping relationship, extracts the two-dimensional coordinates of key points of the human skeleton within the two-dimensional pixel bounding box, and uses the cross-modal spatiotemporal mapping relationship to back-project the two-dimensional coordinates of the key points of the human skeleton onto three-dimensional space to form a spatial ray beam. Combining the multi-view spatial ray intersection geometric constraints and the prior knowledge of the proportion of human physical scale, the three-dimensional coordinates of the personnel target are jointly solved.

7. The construction safety assessment system based on computer vision and UAV inspection according to claim 2, characterized in that, In the operation performed by the UAV multi-waypoint time-series image acquisition subsystem, when dynamically planning the three-dimensional spatial coordinates of the UAV waypoint and the camera optical acquisition pitch angle at the next moment based on the distribution position of the occlusion blind zone volume in three-dimensional space, the system acquires the three-axis attitude angle change rate of the UAV in real time, establishes a negative correlation coupling constraint model between the three-axis attitude angle change rate of the UAV and the image acquisition exposure time, and adaptively adjusts the exposure parameters of the acquired multi-view continuous image sequence based on the negative correlation coupling constraint model to suppress motion blur during dynamic waypoint switching.

8. The construction safety assessment system based on computer vision and UAV inspection according to claim 3, characterized in that, In the operation of iteratively eliminating the spatial intersection error using multi-view feature photometric consistency verification and spatial epipolar distance measurement, weak texture regions in the two-dimensional image corresponding to the local three-dimensional point cloud are detected. Local phase consistency features are extracted in the weak texture regions as illumination-invariant feature descriptors. Combined with the normal vector direction of the corresponding spatial point in the local three-dimensional point cloud, a spatial rotation normalization operation is performed on the illumination-invariant feature descriptors. The multi-view feature photometric consistency verification is performed using the normalized illumination-invariant feature descriptors.

9. The construction safety assessment system based on computer vision and UAV inspection according to claim 4, characterized in that, After preserving and extracting spatial point cloud clusters with rigid geometric continuous topological properties as incremental change features of the building structure, a statistical outlier removal operation based on spatial proximity is performed on the incremental change features of the building structure. The incremental change features of the building structure after removing outliers are transformed into a voxelized spatial mesh representation. Laplacian smoothing iteration processing that preserves the weights of geometric local features is applied to the edge voxels of the voxelized spatial mesh to obtain rigid building incremental features with continuous spatial boundaries and no floating isolated noise.

10. The construction safety assessment system based on computer vision and UAV inspection according to claim 6, characterized in that, In the operation of jointly solving the three-dimensional coordinates of the human target by combining multi-view spatial ray intersection geometric constraints and human physical scale proportion prior, a skeletal joint spatial angle constraint model based on human kinematic constraints is constructed. The three-dimensional spatial coordinates of the skeletal key points obtained by the preliminary solution of the multi-view spatial ray intersection geometric constraints are used as observation data items. The skeletal joint spatial angle constraint model and the human physical scale proportion prior are used together as spatial topological regularization constraint items. The final three-dimensional coordinates of the human target that conform to the laws of physical motion are solved by a nonlinear least squares optimization algorithm.