Building progress evaluation method based on multi-unmanned-aerial-vehicle visual collaboration
By constructing a ternary line-of-sight volume and optimizing the perspective relay logic, the problems of dynamic occlusion and fragmented collaborative observation in UAV building inspection were solved, improving the accuracy and automation of construction progress assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-04-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing drone-based building inspection technology faces challenges in construction sites, including dynamic occlusion avoidance failures and fragmented multi-drone collaborative observation, leading to a decrease in the accuracy of component progress assessment.
By constructing a ternary line-of-sight volume of components, drones, and time, and combining dynamic object trajectory sets to calculate occlusion risk, a candidate set of component view responsibility is generated. The drone view succession logic is optimized through wavefront extension search to generate a component view responsibility timetable. Finally, multi-view observation data is aggregated for local 3D reconstruction.
Explicit modeling of dynamic occlusion and optimization of UAV perspective relay logic improve the automation and accuracy of building component progress assessment and solve the problems of obstructed line of sight and fragmented observation data at the construction site.
Smart Images

Figure CN121810091A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of building digitization, and particularly relates to a building progress evaluation method based on multi-unmanned aerial vehicle (UAV) visual cooperation. BACKGROUND
[0002] With the deep integration of building industrialization and digitization, construction progress monitoring has become a core link for ensuring project delivery on schedule, controlling cost and quality management. The use of unmanned aerial vehicle (UAV) clusters equipped with visual sensors for automatic inspection, with the advantages of high mobility, all-around view and non-contact data acquisition, is gradually replacing traditional time-consuming and inefficient manual site inspection, and has become a key technical means for realizing smart construction sites and building life cycle management.
[0003] In the prior art, the unmanned aerial vehicle (UAV) building inspection is mainly in the single-machine operation mode or the multi-machine simple cooperation mode based on static area division. The path planning mainly relies on the pre-set GPS waypoints or the coverage algorithm based on static maps, and focuses on two-dimensional plane full coverage shooting or fixed route reciprocating flight. In the data processing link, a large amount of images are imported into modeling software for overall reconstruction after flight, and then the reconstructed model is compared with the design model by manual operation.
[0004] However, in the face of complex and variable construction sites, the prior art mainly faces two major problems of dynamic occlusion avoidance failure and multi-UAV cooperative observation fragmentation. Specifically, the traditional static path planning cannot perceive the real-time motion of dynamic obstacles such as tower cranes and transport vehicles, resulting in frequent occlusion of the component line of sight under the preset route, and the collected image data contains a large amount of invalid waste pieces. In addition, the simple region division-based cooperation strategy lacks a time sequence relay mechanism for a single component, and the view angle switching between multiple unmanned aerial vehicles is random and disordered, resulting in fragmented and discontinuous observation data in space and time for a specific component, which cannot meet the view angle coverage requirements for high-quality local three-dimensional reconstruction, and ultimately causes a significant decrease in the accuracy of component-level progress evaluation. SUMMARY
[0005] The application aims to provide a building progress evaluation method based on multi-UAV visual cooperation to solve the above problems in the prior art.
[0006] The technical scheme is a building progress evaluation method based on multi-UAV visual cooperation, which comprises:
[0007] A component-UAV-time ternary line-of-sight body is constructed using pre-defined BIM component data, and a dynamic object trajectory set is extracted from a pre-stored aligned multi-UAV image sequence to calculate the occlusion risk and generate a component view responsibility candidate set;
[0008] A component responsibility grid is constructed, the component perspective responsibility candidate set is mapped into feasible units in the component responsibility grid, a wavefront expansion search is performed, and a component perspective responsibility schedule is generated by calculating a cumulative cost of a minimum handover cost;
[0009] According to the component perspective responsibility schedule, multi-perspective observation data is aggregated from the aligned multi-UAV image sequence, local three-dimensional reconstruction and state determination are performed, and a component short-time state sequence is generated.
[0010] Advantageously, the application solves the problems of obstructed view and fragmented observation data in the construction site by explicitly modeling dynamic occlusion and optimizing UAV perspective handover logic, and improves the automation degree and accuracy of building component progress evaluation. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 A step flowchart of a multi-UAV vision cooperative building progress evaluation method provided by an embodiment of the application.
[0012] Figure 2 A step flowchart of mapping a component perspective responsibility candidate set into feasible units in a component responsibility grid provided by an embodiment of the application.
[0013] Figure 3 A step flowchart of generating a component perspective responsibility schedule provided by an embodiment of the application.
[0014] Figure 4 A step flowchart of correcting a component perspective responsibility schedule provided by an embodiment of the application. DETAILED DESCRIPTION
[0015] In order to enable persons skilled in the art to better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should fall within the protection scope of the present application.
[0016] It should be noted that the terms include and have and any variations thereof are intended to cover inclusive rather than exclusive inclusion, for example, a process, method, system, product, or apparatus that comprises a list of steps or units is not necessarily limited to those steps or units that are clearly listed, but can include other steps or units that are not clearly listed or inherent to such processes, methods, products, or apparatuses.
[0017] As shown in FIG. Figure 1 A multi-UAV vision cooperative building progress evaluation method, comprising the following steps:
[0018] A component-UAV-time triadic view volume is constructed using predefined BIM component data, and a set of dynamic object trajectories extracted from pre-stored aligned multi-UAV image sequences are combined to calculate occlusion risk, and generate a candidate set of component visual responsibility.
[0019] In other words, BIM component data containing geometric information and attribute information in a building information model (BIM) is obtained, and a component-UAV-time triadic view volume is constructed based on the BIM component data. A set of dynamic object trajectories is extracted from pre-stored aligned multi-UAV image sequences, and the set of dynamic object trajectories and the component-UAV-time triadic view volume are combined to calculate occlusion risk, and generate a candidate set of component visual responsibility.
[0020] In the present embodiment, the input data is derived from a basic information acquisition system of a construction site. The predefined building information model (BIM) component data specifically contains geometric information and attribute information in a building information model, such as a three-dimensional geometric grid of each building component, spatial position coordinates, a unique identification code of the component, and construction attributes defined in the design stage. The aligned multi-UAV image sequence refers to UAV acquisition data after clock synchronization and coordinate system registration of multiple machines, which contains sequential images of multiple UAVs under the same time reference and camera pose parameters corresponding to each frame of image.
[0021] Specifically, high-quality observation opportunities are selected through geometric calculation and spatio-temporal analysis. According to the position of the target component in the BIM data, combined with the intrinsic and extrinsic parameters of the UAV camera, a conical visible region pointing from the camera optical center to the surface of the component is constructed in three-dimensional space, i.e. a triadic view volume. This view volume is a spatial volume that changes with time, and its shape is determined by the field of view angle and effective observation distance of the camera. At the same time, the system extracts the motion trajectories of dynamic objects such as tower cranes, engineering vehicles, and construction personnel from the image sequence to form a set of dynamic object trajectories. By calculating the overlap of the spatial bounding boxes of these dynamic objects and the triadic view volume in the spatio-temporal dimension, the possibility of occlusion of the line of sight, i.e. the occlusion risk, is quantitatively evaluated. Those observation windows with an observation quality score higher than a preset threshold and an occlusion risk lower than a safety limit value are defined as a candidate set of component visual responsibility, serving as the basis for subsequent task allocation.
[0022] In some embodiments, the extraction of the set of dynamic object trajectories can use a deep learning-based object detection algorithm combined with multi-view geometric constraints for three-dimensional positioning and tracking. The calculation of the occlusion risk can use a voxelization method to discretize the continuous view volume space into a voxel grid and count the proportion of voxels occupied by dynamic objects.
[0023] A responsibility grid is constructed, and the candidate set of responsibility is mapped to the feasible cells in the grid. A wavefront expansion search is performed on the grid, and a responsibility schedule is generated by minimizing the cumulative cost of handoff.
[0024] In this embodiment, the system converts the complex multi-UAV path planning problem into an optimal path search problem in discrete space. The horizontal axis of the responsibility grid represents discrete time steps, and the vertical axis represents different UAV numbers. Each cell in the grid represents the observation responsibility state of a specific UAV for a specific component at a specific time step. The candidate set of responsibility is read, and the time-UAV combinations that satisfy the observation condition are marked as feasible cells in the grid, i.e., the cell is allowed to be selected.
[0025] Specifically, the system starts from the initial time step and proceeds backward, calculating the minimum cumulative cost to reach each feasible cell. The cost function not only considers the cost of maintaining observation by a single UAV but also introduces a handoff cost. The handoff cost refers to the additional cost generated when the observation responsibility is switched from one UAV to another, usually involving a sudden change in perspective, energy consumption for UAV maneuvering, and potential risks during the handover process. By minimizing this cumulative cost, the path with the lowest cost can be searched on the time axis, which is the responsibility schedule of the component. The responsibility schedule explicitly specifies which UAV should be responsible for observing the component in each time segment during the entire construction period and the specific handover time.
[0026] In one possible example, if switching from UAV A to UAV B at time point t can avoid an upcoming occlusion, and the spatial positions of the two UAVs are close, then the handoff cost generated by this switching path is low, and this path will be preferred. Conversely, if the switching leads to a dramatic change in perspective affecting the reconstruction quality, then the cost is high, and the current UAV's responsibility will be preferred.
[0027] According to the responsibility schedule of the component, multi-perspective observation data is aggregated from the aligned multi-UAV image sequence, local 3D reconstruction and state determination are performed, and a short-time state sequence of the component is generated.
[0028] In this embodiment, the component view responsibility timetable acts as a data index. Based on the responsibility allocation information in this table, the system accurately retrieves and extracts valid image frames belonging to specific components from a massive array of aligned multi-UAV image sequences. These image frames are not randomly selected but are high-quality data optimized by planning algorithms, ensuring continuous and unobstructed views. In practice, these time-spanning and multi-UAV image frames are aggregated into multi-view observation data. Using multi-view stereo vision technology, this aggregated data is processed to generate a local dense point cloud or 3D mesh model of the target component. This model is then geometrically compared with the original predefined BIM component data. By calculating the shape differences, positional deviations, or feature point matching degrees between the two, the system can automatically determine the current construction status of the component. The short-time state sequence of the component is the time-series record of the determination results, and its state values can include installed, not installed, partially installed, or with temporary support, etc.
[0029] For example, if the reconstructed point cloud and the BIM model overlap by more than 90% at the design location, and the surface texture matches the characteristics of poured concrete, it is considered completed. If there are objects at the location but their shapes do not match, it may be considered as scaffolding obstruction or temporary material storage. The serialized state output reflects the actual physical progress of the building construction.
[0030] In one possible implementation, a three-dimensional line-of-sight volume of component-UAV-time is constructed using predefined BIM component data, including:
[0031] Extract the 3D geometric center and target area of the component from the predefined BIM component data, and obtain the pose parameters of the UAV camera from the aligned multi-UAV image sequence.
[0032] Specifically, the system parses predefined BIM component data to extract the geometric features of each component to be monitored. The target area is defined as the outer surface of the component, key connection nodes, or inspection surfaces specified in the design drawings. To simplify subsequent calculations, the geometric center or bounding box center of the component can be used as the focal point of the view. Simultaneously, the system reads camera pose parameters from the aligned multi-UAV image sequence. These parameters include the camera's 3D position coordinates in the world coordinate system and rotation matrices or quaternions describing the camera's attitude. Furthermore, intrinsic camera parameters such as focal length and principal point coordinates are also necessary for constructing the view frustum.
[0033] Construct a cone-shaped field of view pointing from the optical center of the UAV camera to the target area, and map the cone-shaped field of view to a unified three-dimensional voxel space to generate a ternary line of view volume indexed by the component, the UAV, and time.
[0034] Specifically, this embodiment constructs a solid model describing the observation geometry. For any given time step t, UAV u, and component c, the system constructs a spatial cone or frustum with the optical center of the UAV camera as the vertex and the target area of the component as the base. This represents the spatial area that the camera must penetrate to observe the component in its current pose, i.e., the effective field of view. To perform efficient Boolean operations and spatial analysis, the system can employ voxelization technology. The three-dimensional space of the entire construction scene is divided into fixed-size cubic units, i.e., voxels. The cone-shaped field of view is projected and discretized into a voxel mesh. The set of all voxels located inside the cone constitutes the component-UAV-time ternary field of view, which can be represented as V(c, u, t), containing three-dimensional voxel indices. This reduces the computational complexity of intersection calculations for complex geometry.
[0035] In a further embodiment, occlusion risk is calculated by combining a set of dynamic object trajectories extracted from a pre-stored aligned multi-UAV image sequence, including:
[0036] The predicted positions in the dynamic object trajectory set are transformed into discrete dynamic object voxels.
[0037] In this embodiment, the system handles dynamic interference sources at the construction site. The dynamic object trajectory set contains the position and size information of moving objects such as tower cranes, elevators, and transport vehicles at different points in time. Based on the geometry of each type of object, combined with its position coordinates and orientation at time t, the system generates an occupancy model of that object in three-dimensional space. Similarly, using a voxelization standard, these occupancy models are transformed into a discrete set of dynamic object voxels, denoted as O(t). Preferably, considering the uncertainty of the predicted trajectory, an expansion factor can be introduced when generating voxels. For example, for vehicles moving at high speeds, a certain number of voxels can be expanded along their direction of movement to cover the range of their possible locations. Each type of dynamic object can also be assigned different attribute values, such as occlusion weights, for subsequent risk-weighted calculations.
[0038] The percentage of the overlapping volume between the three-dimensional line-of-sight volume of the component-UAV-time and the dynamic object voxel in a unified spatiotemporal coordinate system.
[0039] Specifically, within a unified spatiotemporal coordinate system, the system performs a set intersection operation on the set of line-of-sight volumes V(c, u, t) and the set of dynamic object voxels O(t). The calculation measures the proportion of voxels in the intersecting portion relative to the total number of voxels in the line-of-sight volume. The overlap volume ratio r(c, u, t) directly reflects the physical degree of line-of-sight obstruction. A ratio of 0 indicates unobstructed line-of-sight; a ratio close to 1 indicates that the component is obstructed by a dynamic object. The system can not only handle physical occlusion but also adapt to different levels of observation precision by adjusting the voxel resolution.
[0040] A quantitative occlusion risk is generated based on the overlapping volume ratio to exclude component observation requests during periods of high occlusion risk.
[0041] In this embodiment, the system converts the geometric occlusion ratio into a risk indicator required for decision-making. For example, the occlusion risk R... occ The evaluation depends not only on the overlap ratio r, but also on the uncertainty σ of the dynamic object and the object category weight w. The specific calculation model can be expressed as a linear weighted or nonlinear function. For example, tower cranes, as large steel structures, have rigid and impenetrable obstructions, thus receiving a higher weight; while construction smoke or small equipment, whose obstructions may be semi-transparent or temporary, receive a lower weight. Preferably, a safety threshold T is set. risk When the calculated occlusion risk R occ When the threshold is exceeded, the current time window is determined to be unavailable for the UAV to observe the component, belonging to a high-risk occlusion period. These high-risk periods will be marked as infeasible in subsequent planning steps and thus excluded from the responsibility candidate set. This ensures that the final acquired image data has extremely high validity and usability.
[0042] like Figure 2 As shown, in an exemplary embodiment, mapping the candidate set of component-view responsibilities to feasible units within the component responsibility grid includes:
[0043] For any component, establish a two-dimensional mesh structure with discrete time steps as columns and UAV numbers as rows.
[0044] In this embodiment, the system constructs an independent two-dimensional coordinate system for each component to be monitored. The horizontal axis corresponds to the time axis of construction monitoring and is discretized into a series of continuous time steps t1, t2, ..., t N The time step depends on the required temporal resolution of the monitoring, such as one node per second or every ten seconds. The vertical axis corresponds to the drone formation participating in the collaborative operation, labeled u1, u2, ..., u M The two-dimensional mesh structure G contains N×M nodes, and each node (t, u) represents a potential decision state: at time t, it is observed by the UAV u.
[0045] By analyzing the time interval information in the candidate set of component responsibility, the component-UAV-time triplet that meets the preset observation requirements is marked as a feasible unit in the two-dimensional grid structure, forming a component responsibility grid.
[0046] Specifically, the discrete candidate set is mapped to a continuous mesh structure. The component-perspective responsibility candidate set is traversed, and for each valid entry (c, u, [t]...)... start , t end]), and the corresponding time interval [t] in grid G. start , t end All nodes (t, u) within the candidate set are marked as feasible units. Conversely, nodes not included in the candidate set, or nodes corresponding to periods of high occlusion risk, are marked as infeasible units or obstacle units. Through a binarization or weighted marking process, traversable regions are plotted in the search space. The resulting component responsibility grid visually displays the distribution of all available observation resources, providing a foundational map for subsequent path optimization.
[0047] like Figure 3 As shown, in a further embodiment, generating a component-perspective responsibility schedule includes:
[0048] The dwell cost is constructed to maintain the same UAV observation in adjacent discrete time steps, and the relay cost is constructed to switch between different UAV observations in adjacent discrete time steps.
[0049] In this embodiment, two types of key costs are introduced to ensure the smoothness and efficiency of the observation mission. The cost of residence, C... stay (t, u) relates to the cost of the UAV maintaining its current observation state. This cost is generally inversely proportional to the observation quality; that is, the better the observation angle and the higher the resolution, the lower the dwell cost. It may also include the hovering energy consumption cost of the UAV. Relay cost C trans (t-1, u) i -> u j The task control authority was defined from the drone. i Transfer to u j The cost of relaying data is significant. When i = j, i.e., continuous observation from the same drone, the relay cost is usually zero or minimal. When i ≠ j, i.e., a cross-drone handover occurs, the relay cost increases significantly. The magnitude of the relay cost depends on the spatial distance between the two drones at the moment of handover, the difference in viewing angle, and the overhead required for communication handshake. For example, if the two drones are far apart, a forced handover may cause a jump in the observation view, thus requiring a very high cost to suppress such handovers.
[0050] Dynamic programming calculations are performed on feasible units in the component responsibility grid along the time dimension, accumulating dwell costs and relay costs to determine the minimum cumulative cost path to each feasible unit.
[0051] For example, let D(t, u) be the minimum cumulative cost when arriving at time step t and being handled by UAV u. Starting from time t = 1, the process progresses towards time t = N. For each feasible unit u at time t, the system iterates through all feasible units v from the previous time step t-1 and calculates the transfer cost:
[0052] Cost temp = D(t-1, v) + Ctrans (t-1, v -> u) + C stay (t, u);
[0053] The system selects to make the transfer cost cost temp The smallest feasible unit v is taken as the optimal predecessor node, and this minimum cost is assigned to D(t, u). Similar to wavefront propagation in a grid, each step expands upon the optimal solution of the previous step. All possible combinations of responsibilities are traversed, and high-cost path branches are implicitly pruned.
[0054] By backtracking the path with the minimum cumulative cost from the end of the discrete time step, the optimal drone responsibility sequence is extracted, and a responsibility timetable from the perspective of components is generated.
[0055] In this embodiment, when the wavefront spread reaches the final time t N At that time, the system corresponds to all nodes D(t) of all drones. N Find the node with the minimum cost in (t1, u) as the endpoint of the path. Based on the recorded optimal predecessor pointer, backtrack to the starting time t1. The backtracking process connects a series of nodes, forming a complete path: (t1, u) a ), (t2, u a ), ..., (t k u b This path represents the globally optimal responsibility allocation scheme. The system transforms it into a structured component-perspective responsibility schedule, clearly indicating when and by which aircraft is responsible, and when a responsibility handover occurs. The component-perspective responsibility schedule serves as the direct instruction basis for subsequent UAV flight control and image data processing.
[0056] According to one aspect of this application, in actual large-scale building construction scenarios, there are often hundreds of components that need to be monitored simultaneously. If optimal planning is only performed independently for each component, it is very easy for a single drone to undertake too many tasks at the same time, exceeding its physical maneuverability or communication bandwidth limitations. Therefore, as Figure 4 As shown, this also includes revisions to the component-perspective responsibility timeline, specifically:
[0057] Summarize the component-perspective responsibility timeline for all components and construct a global load curve reflecting the drone mission density.
[0058] In this embodiment, the system spatiotemporally overlays the component-perspective responsibility timetable generated individually for each component. A global load matrix L is established, with rows corresponding to UAV numbers and columns corresponding to time steps. For each element L(u, t) in the matrix, the total number of component tasks assigned to UAV u at time t is counted.
[0059] Specifically, the definition of load can be more than just the number of tasks. In some preferred embodiments, load calculation can introduce weighting factors. For example, tasks involving components with longer observation distances or requiring high-precision hovering are given higher weights, while simple wide-angle inspection tasks are given lower weights. In this case, L(u,t) equals the sum of the weights of all tasks assigned to UAV u. The generated global load curve visually reflects the busyness of each UAV throughout the entire operation cycle, with peaks representing high-load periods and troughs representing idle periods.
[0060] Based on preset kinematic constraints, conflict periods in the global load curve are detected, and relevant components within the conflict periods are identified.
[0061] In this embodiment, the physical performance indicators of the UAV are used as constraints for conflict screening. For example, the preset maneuver constraints include the maximum number of concurrent target observations by the UAV, the maximum maneuver angular velocity limit, and the data transmission bandwidth limit. The system sets a load threshold L. max When the load value L(u,t) at a certain time t in the global load curve exceeds the threshold, that time segment is marked as a conflict period. Once a conflict period is identified, the task allocation details within that period are further traced back to find all components that caused the load to exceed the limit, i.e., related components. These related components are then sorted according to pre-defined priority rules. For example, main beam components on critical construction paths have higher priority, while decorative components have lower priority. This provides a basis for subsequent replanning decisions.
[0062] During the conflict period, perform local wavefront extension searches on relevant components, reallocate UAV responsibilities, and generate a revised component view responsibility schedule.
[0063] Specifically, local optimization adjustments are performed to eliminate conflicts. The system does not overturn the overall planning scheme, but rather locks the solution for non-conflict periods and only reconstructs the responsibility grid for low-priority related components within the time window of conflict periods. Preferably, a penalty term P is introduced into the cost function. conflict For high-load drone nodes, the dwell cost C stay This penalty term is superimposed, forcing the algorithm to abandon the node during wavefront expansion search and instead search for other drones with lower loads or choose slightly worse but feasible observation perspectives. After this round of local replanning, the system outputs a revised component perspective responsibility schedule, which ensures that the real-time load of each drone is below a safe threshold while meeting the observation needs of all components.
[0064] Building upon this, the system can further generate smooth flight trajectories based on the revised schedule. For example, at the handover point between two consecutive duty segments, the Minimum Snap algorithm is used to generate a smooth curve connecting the two observation points, ensuring the flight stability of the UAV during baton relay maneuvers. This embodiment resolves resource conflicts under multi-task concurrency through a global load coordination mechanism.
[0065] According to one aspect of this application, multi-view observation data is aggregated from aligned multi-UAV image sequences, including:
[0066] By analyzing the responsibility switching moments recorded in the responsibility timeline from the perspective of the components, the aligned multi-UAV image sequence is segmented into responsibility segments continuously handled by different UAVs.
[0067] In this embodiment, the system acts as an intelligent editor. It reads the revised component-perspective responsibility timeline and pinpoints the time when the responsible party for each component changes, i.e., the responsibility switch time T. switch Specifically, for a certain component C, if in the time interval [T1, T2]... switch [The area within] is handled by drone A, while [T] switch Within T2], drone B is responsible for processing the video stream from drone A within the original aligned multi-drone image sequence. switch The video stream was cut off at the point where the preceding segment of the image was extracted; at the same time, the video stream of drone B was also extracted from T. switch The image is truncated at a certain point, and the subsequent segment is extracted. The resulting subset of images is called the responsibility segment. Each responsibility segment has a high degree of internal continuity, meaning it is acquired by the same sensor on a continuous trajectory, and based on the planning when generating the candidate set of responsibility segments for component views, component observation requests during periods of high occlusion risk are excluded. In other words, component C within this segment is always within a good range of visibility.
[0068] For any given component, multi-angle image frames of the same component are extracted and aggregated across the responsible segments to construct multi-view observation data for that component, which is then used for local 3D reconstruction.
[0069] Specifically, the system performs logical aggregation of data. Although the responsible segments are physically dispersed (from different UAVs), logically they collectively describe the same component. The system aggregates all responsible segments belonging to component C into a dedicated data container, removing redundant or blurred frames to form multi-view observation data for that component. Preferably, multi-view geometric reconstruction algorithms, such as Structure of Motion (SfM) or Multi-View Stereo Vision (MVS), are used to process this aggregated data. Since the input data has already eliminated occlusion interference and ensured the complementarity of viewpoints, the reconstruction algorithm can efficiently recover the dense 3D point cloud P of the component.recon After obtaining the point cloud, the system executes the state determination logic. It then calculates and reconstructs the point cloud P. recon To the predefined BIM model surface M bim The unidirectional chamfer distance. Specifically, for each point p in the point cloud, find its nearest point q on the model surface and calculate the distance d(p, q). If the average distance of all points is less than the preset installation accuracy threshold ε1 (e.g., 5 cm), the component is determined to be in the installed state; if the average distance is large but the spatial distribution characteristics of the point cloud match the model in the temporary support structure library, it is determined to be in the temporary fixed state.
[0070] In one embodiment of this application, the method further includes:
[0071] Extract the physical objects actually observed on site from the short-term state sequence of the components, construct a set of physical entity nodes, and establish a set of target component nodes based on predefined BIM component data.
[0072] In this embodiment, a graph model foundation for semantic reasoning is established. The system scans the short-time state sequence of components, extracts all independent objects detected on-site that have stable geometric shapes, and defines them as the physical entity node set V. E These nodes represent what is actually on site. Simultaneously, by parsing the BIM model, all components from the design drawings are defined as the target component node set V. B These nodes represent what should theoretically be present.
[0073] Initialize the candidate correspondence between the physical entity node set and the target component node set, and construct a temporary-permanent hybrid topological semantic graph containing entity nodes, target component nodes and topological constraint nodes.
[0074] In this embodiment, the system constructs a complex graph structure G containing two types of nodes and multiple types of edges. Specifically, it computes each physical entity node e. i Spatial location and each target component node b j The Euclidean distance between the theoretical positions. If this distance is less than the preset search radius R. search (For example, 2 meters), then in e i With b j Establish a candidate edge between them, representing e i It might be b j This is the practical manifestation of the system. All these candidate edges constitute the initial candidate correspondences. In addition, the system also introduces topological constraint nodes. For example, if the BIM model stipulates that component A must be installed on component B, the system will establish a support constraint edge between the corresponding target nodes. The final temporary-permanent hybrid topological semantic graph is not only a set of geometric locations, but also a knowledge graph containing logical constraint relationships.
[0075] In a further embodiment, it also includes:
[0076] Read the temporal behavior features and geometric connectivity relationships in the short-term state sequence of components, and perform multiple rounds of probability updates on the temporary-permanent hybrid topological semantic graph.
[0077] For example, Bayesian inference is used to update the identity probability. The system does not rush to determine the entity's identity in a single frame, but rather accumulates observational evidence over time. Temporal behavioral characteristics, such as an entity remaining in the same position for several consecutive days (suggesting a permanent structure) or frequently moving (suggesting a temporary device), and geometric connections, such as observing a connection between entity A and entity B, are also considered. The system maintains a probability matrix P, where P(i,j) represents the physical entity e. i Corresponding target component b j The confidence level. Whenever new observational evidence E is obtained. new The confidence level is updated according to Bayes' theorem P(H|E) = P(E|H) * P(H) / P(E), where H is the hypothesis, E is the evidence, P(H|E) is the posterior probability, P(E|H) is the likelihood probability, P(H) is the prior probability, and P(E) is the marginal probability of the evidence. For example, if entity e is observed... i It has already been installed in the known b k On the beam, and in BIM b j It was installed in b k This evidence will increase the value of P(i,j).
[0078] Based on the temporary-permanent hybrid topological semantic graph updated by multiple rounds of probability, the candidate correspondence is shrunk by using the consistency check of topological constraint nodes, and the unique entity identity of the physical entity is determined by the delayed binding mechanism.
[0079] In this embodiment, the system utilizes global constraints to eliminate local ambiguity. A delayed binding mechanism allows entities to remain unidentified for a period until accumulated evidence is sufficient to eliminate interference. The system periodically performs consistency checks on graph G. For example, if entities e1 and e2 both correspond to the same target b1 with a very high probability, this violates the physical uniqueness constraint. The system compares the matching degree of their connection relationships, retains the pair with the higher matching degree, and forcibly cuts off the other candidate edge, thereby shrinking the candidate set. After multiple rounds of updates and shrinkage, when the probability matrix P(i,j) exceeds the confirmation threshold (e.g., 0.95) and only one option remains in the candidate set, the system locks the correspondence, confirming the unique entity identity of the physical entity.
[0080] In a further embodiment, it also includes:
[0081] Identify nodes in the temporary-permanent hybrid topological semantic graph that still have multiple identity solutions or topological constraint conflicts after multiple rounds of probability updates, and generate a list of high-uncertainty components.
[0082] In this embodiment, not all inferences converge quickly; some nodes may remain in an ambiguous state for an extended period due to insufficient observation data or overly complex on-site conditions. Optionally, the system quantifies this uncertainty by calculating the information entropy of the probability distribution. If a node e... i The corresponding probability distribution entropy value H(e) i If the value is higher than the preset threshold, it means that the system is still unable to determine what the component is or its progress. These nodes are then aggregated to generate a list of high-uncertainty components.
[0083] By combining the spatial distribution of the list of high-uncertainty components with the weight of construction impact, an observation optimization suggestion set for the next round of data collection is generated.
[0084] Specifically, the system analyzes the sources of uncertainty and formulates countermeasures. If high-uncertainty nodes are concentrated on the north facade of the building, and these nodes are load-bearing columns on the critical path (with high construction impact weight), the system will generate targeted strategies. The set of observation optimization suggestions specifically includes: the coordinates of the spatial areas to be prioritized for observation, the number of observation angles to be increased, and the suggested increase in image resolution.
[0085] The observation optimization suggestion set is fed back into the step of constructing a component-drone-time ternary line-of-sight volume using predefined BIM component data, and is used as an additional weight to adjust the screening and occlusion risk calculation of the next round of component-drone-time ternary line-of-sight volume.
[0086] In this embodiment, a closed loop of data acquisition and data analysis is achieved. The system transforms the aforementioned suggestion set into input parameters for the step of constructing a three-element line-of-sight volume of component-UAV-time using predefined BIM component data. Specifically, for key components in the suggestion set, the system will artificially lower their acceptable occlusion risk threshold T during the next round of occlusion risk calculation. risk (This requires clearer visibility), or assigning a higher weighting factor when calculating the observation quality score. This means that in the next round of flight planning, the drone will prioritize allocating resources to clearly see components that were previously unclear, thereby eliminating uncertainty in the new observation cycle. This gives the entire system the intelligent characteristic of self-optimization as construction progresses.
[0087] This embodiment addresses the complex situations commonly encountered at construction sites, such as the mismatch between physical objects and models and uncertain progress. By introducing graph theory models and probabilistic update mechanisms, it achieves robust inference of component identities and progress, and feeds the inference results back to the front-end planning, forming a closed-loop optimization.
[0088] In one possible embodiment, extracting the dynamic object trajectory set includes:
[0089] Based on pre-stored aligned multi-UAV image sequences, multi-view dynamic object detection and cross-view matching are performed on the original image sequences collected by multiple UAVs.
[0090] In this embodiment, the system performs frame-level processing on the monocular video stream transmitted back by each UAV. A lightweight deep neural network model, such as YOLO or Mask R-CNN, can be used to perform object detection on each frame. The detected object categories include common moving obstacles found at construction sites, such as tower crane booms, concrete mixer trucks, excavators, and construction workers. The output is a two-dimensional bounding box or instance segmentation mask with category labels. Since multiple UAVs observe the same scene from different angles, the same physical object may appear in the images of different cameras. Cross-view matching is performed using calibrated camera extrinsic parameters and a unified timestamp. Specifically, the system uses the epipolar constraint geometry principle to calculate the epipolar line corresponding to a target point in one camera view in another camera view. If the center of the detection box in the other view is located near this epipolar line and the similarity of appearance features (e.g., based on color histograms or ReID feature vectors) is higher than a preset threshold, then the two two-dimensional detection boxes are determined to correspond to the same three-dimensional physical object.
[0091] Three-dimensional localization and trajectory initialization are performed based on cross-view matching results, and a smooth dynamic object trajectory set is generated using a filtering algorithm.
[0092] Specifically, the matched multi-view Figure Two 3D observations are transformed into 3D spatial coordinates. Multi-view is employed. Figure Three The angle measurement algorithm uses the least squares method to solve for the intersection of lines of sight, thereby estimating the three-dimensional center position P of the dynamic object at the current time t. t To construct a continuous trajectory, a multi-target tracking algorithm, such as the Hungarian algorithm, can be used to track the current 3D position P. t Trajectory T from the previous moment (t-1) A correlation is established. The correlation is based on the consistency between spatial distance and velocity vectors. Furthermore, considering the noise and jitter in the original detection data, the system introduces a state estimation filter, preferably a Kalman filter or an extended Kalman filter. A state vector is maintained for each dynamic object, including its position, velocity, and acceleration. The filter continuously corrects the state prediction values based on the observations, outputting a smooth, continuous set of dynamic object trajectories containing short-term prediction information. This trajectory set not only includes the object's current position but also its future motion trend prediction, providing accurate spatiotemporal data support for occlusion risk prediction.
[0093] According to one aspect of this application, the multi-UAV visual collaborative building progress assessment method can also be:
[0094] Acquire multi-source basic data and predefined BIM data from the construction site, generate aligned multi-UAV image sequences in a unified coordinate system, and generate a candidate set of component view responsibility containing occlusion risk information based on component line-of-sight volume analysis.
[0095] A component responsibility grid with time-UAV dimension is constructed based on the component responsibility candidate set. The wavefront expansion algorithm is executed on the component responsibility grid to solve the optimal responsibility path and generate a component responsibility time schedule.
[0096] Based on the responsibility switchover time in the component-perspective responsibility timetable, generate a baton relay flight planning instruction set and a baton relay labeled image sequence with responsibility tags;
[0097] Based on the relay-annotated image sequence, perform component-level local multi-view reconstruction and extract short-time state sequences of components that reflect their physical characteristics.
[0098] A temporary-permanent hybrid topological semantic graph is constructed based on the short-term state sequence of components. A delayed binding mechanism is used to perform joint reasoning on node identity and state, and the component progress evaluation result is output.
[0099] In an optional embodiment of this application, generating a candidate set of component viewpoint responsibilities may further involve: acquiring aligned multi-UAV image sequences and an initial scene geometric framework. Using the same type of target detection and segmentation model in each image stream, candidate dynamic object instances are obtained for each frame, including category and 2D boundary / contour. Geometric consistency matching is performed between different UAV viewpoints at similar times using timestamps and camera extrinsic parameters, for example, through epipolar constraints + appearance feature metrics, to establish cross-viewpoint object correspondences. For objects with established cross-viewpoint matching, triangulation or perspective-n-point (PnP) inverse calculation is performed in a unified coordinate system to obtain 3D position estimates at each time step. Simple trajectory association is performed on the 3D detection results of the same object in consecutive frames, such as 3D intersection-over-union (IoU) / nearest neighbor + velocity consistency constraints, to form initial 3D trajectory segments. Filtering / smoothing models such as Kalman filtering or smoothing splines are applied to the initial trajectory segments to eliminate jitter and occasional false detections. Within the subsequent time window to be considered, i.e., the short-term prediction window, the predicted position and voxel occupancy are extrapolated for each trajectory to assess future occlusion in advance. Output a set of dynamic object trajectories, including the three-dimensional spatiotemporal trajectory and short-term prediction results of each dynamic object, including position, voxel occupancy range, and uncertainty.
[0100] Acquire an initial candidate set of components, an aligned multi-UAV image sequence, an initial scene geometry framework, and a set of UAV maneuver constraint parameters. Discrete the time axis with a fixed time step Δt within a preset short-term time window. Extract the pose (position + attitude) of each UAV at each discrete moment from the aligned multi-UAV image sequence, and, if necessary, combine the UAV maneuver constraint parameter set to predict feasible poses for future short moments. For each component, obtain its 3D bounding box and geometric center, as well as design-recommended observation surfaces, such as exposed surfaces and key connection points, from predefined BIM component data. Select several key points / key surfaces of the components as the target area of the line-of-sight volume to reduce redundant calculations. For each triple (component c, UAV u, time t), construct a cone or truncated cone-shaped line-of-sight volume from the UAV camera center to the component target area, with the field of view and effective observation distance derived from UAV / camera parameters. Clip the line-of-sight volume using the scene boundary in the initial scene geometry framework, restricting it to the construction area. Project each line-of-sight volume into a unified 3D voxel space to generate a set of discrete line-of-sight volume voxels indexed by (c, u, t). Using static structural voxels from the initial scene geometry, a preliminary determination is made as to whether the view volume is permanently occluded over a large area, such as being invisible behind a component. If occluded, the state (c, u, t) is directly marked as statically invisible to avoid redundant calculations for dynamic occlusion later. The output is a set of view volumes for the components, i.e., a voxel data structure indexed by (component c, UAV u, time t), containing the geometric parameters, voxel set, and static occlusion markers for each view volume.
[0101] Read the component view volume set, dynamic object trajectory set, and initial scene geometry. Convert the predicted position of each object in the dynamic object trajectory set at each discrete time step into a voxel representation, such as using bounding boxes plus uncertainty dilation as voxel blocks. Assign an object category (tower crane, vehicle, personnel, etc.) to each voxel for subsequent category-based adjustment of occlusion weights. In a unified voxel coordinate system, perform voxel intersection operations between each (c, u, t) view volume voxel set and the dynamic object voxel set at the same time step t. Calculate the ratio of overlapping voxels to the effective voxels of the view volume to obtain the occlusion ratio r(c, u, t). Simultaneously distinguish between occlusion on the main view ray path and occlusion in the surrounding field of view, giving the former a higher weight. Combine the occlusion ratio r with the object prediction uncertainty σ and the object category weight w to construct an occlusion risk scoring function.
[0102] R occ (c,u,t)=f(r(c,u,t),σ(c,u,t),w class );
[0103] Where R occ(c, u, t) represents the occlusion risk score corresponding to the observation unit u for component c at time t; r(c, u, t) represents the occlusion ratio between the line of sight from observation unit u to component c and the dynamic object at time t; σ(c, u, t) represents the prediction position uncertainty of the dynamic object relative to the observation unit u pointing to component c at time t; w class Here, f represents the category weights of dynamic objects; f() is the occlusion risk scoring function. Temporal smoothing is used to suppress occlusion risk fluctuations caused by brief false detections. The output is the intermediate quantity required to fuse into the occlusion risk time window set: the occlusion percentage r(c,u,t) and the occlusion risk score R for each (c,u,t). occ (c, u, t).
[0104] Read predefined BIM component data, occlusion ratio r(c, u, t), and occlusion risk score R. occ (c, u, t), the set of line-of-sight volumes of the components, and the set of UAV maneuver constraint parameters. For each (c, u, t), based on the geometry of the line-of-sight volumes and the geometry of the components, calculate: the estimated imaging distance d(c, u, t); and the estimated projected area A of the component in the image. img (c, u, t); the angle θ(c, u, t) between the camera optical axis and the component normal. These quantities are compared with the demand thresholds and preference regions in the predefined BIM component data to obtain the observation quality score Q. obs (c, u, t). A comprehensive occlusion risk is considered to form the final overall score:
[0105] S(c, u, t) = g(Q) obs (c, u, t), R occ (c, u, t));
[0106] Where g() is the comprehensive score fusion function. The higher the occlusion risk, the lower the final comprehensive score S. For each (c, u), R is viewed along the time direction. occ The (c, u, t) sequence is used to label high-occlusion-risk intervals and low-occlusion-risk intervals using a threshold + connected interval method. All (c, u) sequences are aggregated to form a set of occlusion-risk time windows organized by component. For each (c, u), continuous intervals satisfying the following conditions are searched on the time axis: S(c, u, t) is higher than the observation quality threshold; it is not within a strong occlusion time window; and the UAV can reach and maintain the required viewpoint within the interval according to the UAV maneuvering constraint parameter set. For each selected continuous interval, its start and end times, average score, and estimated maneuvering costs within the interval are recorded to form a component-UAV viewpoint responsibility candidate segment. A series of (component c, UAV u, time interval [t]) sequences are then used to identify the high-occlusion-risk intervals and low-occlusion-risk intervals. start , t end [) and its comprehensive score and the responsibility candidate set from the perspective of the components of the motor cost.
[0107] In another possible embodiment of this application, generating the component view responsibility timetable can also involve: reading the component view responsibility candidate set, the occlusion risk time window set, the component line-of-sight volume set, aligning multi-UAV image sequences, and the UAV maneuver constraint parameter set. For each component c, based on the time range of all its candidate intervals, a discrete time axis {t1, t2, ..., t...} is defined. T The responsibility planning for all drones on this component is unified along this timeline. A two-dimensional grid G is constructed with time steps as columns and drones as rows. c (i, j), where i is the time index and j is the UAV number. Initially, all units are marked as infeasible. Using the candidate time intervals in the component-perspective responsibility candidate set, the units covering that time interval (t) are... i UAV j Units are marked as candidate feasible, where UAV j For the j-th UAV, referencing the occlusion risk time window set, directly mark elements within strong occlusion zones as infeasible. Combining the UAV maneuver constraint parameter set and UAV pose, check whether the UAV can smoothly transition from its current pose to a candidate responsibility pose in adjacent time steps; if not, change the element to infeasible. Output the component responsibility grid set.
[0108] Read the component responsibility grid set, component line-of-sight volume set, dynamic object trajectory set, and UAV maneuver constraint parameter set. For each grid cell G c (i, j), define the residence cost C stay (c, i, j), usually related to the observation quality score S(c, u) j , t i The negative value of ) is related to motor energy consumption, etc. For two time-adjacent units G c (i, j) and G c (i+1, k), define the transfer cost C trans (c, i, j → k) represents the cost of transferring from unit j at time i to unit k at time i+1, including: if j is not equal to k, then it is the relay cost; based on the dynamic object trajectory set, if there is a potential collision / congestion at the handover moment, an additional collision avoidance penalty is added. The total cost is: C total =∑ i C stay (i, u) i )+∑ i C trans (i, u) i →u i+1 (In the time dimension from t1 to t) T Proceeding sequentially, for each feasible unit G c (i, j), calculate the minimum cumulative cost to reach this state: D(i, j) = mink (D(i-1,k)+C trans (i-1, k→j))+C stay (i, j); Only consider the feasible predecessor units from the previous step. Simultaneously record the optimal predecessor (i-1, k*) that implements D(i, j), forming a predecessor pointer. At the final time t... T First, select the unit (T, j*) with the minimum cumulative cost from all rows of drones. Then, backtrack from (T, j*) to the starting point along the predecessor pointer to obtain the optimal drone responsibility assignment in the time series: {u(t1), u(t2), ..., u(t... T If the responsible drones are the same at adjacent time steps, they naturally form a continuous responsibility segment; if they are different, they are marked as the handover time. Output the initial component view responsibility timetable for each component c, and give the responsible drone and handover time for each time step.
[0109] Read the initial component-perspective responsibility schedule, UAV maneuver constraint parameter set, and component importance weights for all components, such as giving higher weights to critical path components. For each UAV u, project the responsibility schedules of all components onto a unified time axis. Calculate the number of all constructed tasks at each time step and estimate the required maneuver intensity, transmission bandwidth, and other load indicators. Based on the UAV maneuver constraint parameter set, including maximum parallel tasks, multi-task switching time, and maneuver capability limits, mark the overload / time sequence violation periods. Treat the component-time slice-UAV responsibilities involved in the conflict periods as a conflict segment set. Within the time window corresponding to the conflict segment, locally reconstruct the responsibility grid of the relevant components, i.e., rebuild the subgrid only in narrow time windows and on a small number of UAVs. By introducing component importance weights, design a new local cost function, for example: critical component responsibilities are difficult to remove, and non-critical components are given priority in relinquishing resources. Perform wavefront expansion / dynamic programming again on these subgrids to find new responsibility paths to minimize conflicts while keeping the overall cost increment to a minimum. Output the component-perspective responsibility schedule after global consistency coordination.
[0110] Read the component-view responsibility schedule, aligned multi-UAV image sequences, UAV maneuver constraint parameter set, and dynamic object trajectory set after global consistency coordination. For each component responsibility schedule, extract the time step pair (t) from the transition of responsibility from UAV u1 to u2. k , t k+1The handover window is defined as the relay window. Within the handover window, candidate handover locations are determined based on the pose trajectories of the two UAVs and the spatial position of the component, such as within the visible area of the component and where the difference in viewing angle between the two UAVs is not too large. Before and after the handover window, short-range pose adjustment trajectories are planned for UAVs u1 and u2 respectively, ensuring they are simultaneously within a reasonable viewing angle and safe distance range near the handover time. Obstructions and safe distance constraints within the dynamic object trajectory set are considered to prevent the planned trajectories from crossing high-risk areas. Continuous control commands are generated using smooth interpolation or optimal control methods. The final responsibility schedule is mapped back to the aligned multi-UAV image sequence: for each image frame, the responsible UAV tag and responsibility status (before, during, and after the handover) for each component in the current frame are appended. Within the handover window, frames observed jointly by both UAVs are marked as priority data for subsequent multi-view reconstruction. The handover flight planning command set and the handover labeled image sequence are output.
[0111] In summary, the multi-UAV vision-coordinated building progress assessment method includes: constructing a three-dimensional line-of-sight volume based on BIM data (component-UAV-time), calculating the spatiotemporal overlap ratio of line-of-sight voxels and obstacle pixels by combining dynamic object trajectories, and generating a candidate set of component view responsibility; constructing a responsibility grid with time steps and UAVs as dimensions, executing a wavefront expansion search algorithm on the grid, and generating an optimal component view responsibility schedule by calculating and minimizing the cumulative cost of dwelling cost and relay cost; based on this schedule, segmenting and aggregating responsibility segments from multi-source image sequences, performing local 3D reconstruction and state determination, and obtaining a short-term state sequence of components.
[0112] This invention addresses the problem of dynamic line-of-sight occlusion by employing a method of intersection between a ternary line-of-sight volume and dynamic voxels. By explicitly establishing the spatiotemporal relationship between the line of sight and dynamic obstacles (such as tower cranes) during the planning phase, it quantifies and calculates occlusion risk scores, eliminating high-risk periods before flight and resolving the issue of invalid data acquisition due to dynamic occlusion. For the fragmentation problem of multi-drone collaborative observation, it adopts a responsibility grid and wavefront extension relay planning method. By introducing the relay cost as a key parameter and performing dynamic programming, it generates the responsibility switching path with the minimum cost on the time axis. This enables a smooth and orderly handover of observation tasks for the same component by multiple UAVs, ensuring the continuity and consistency of observation perspectives and resolving the problem of data discontinuity caused by disordered switching. Through relay-driven image aggregation and hybrid topological semantic reasoning, high-quality image data is transformed into precise progress status, achieving a closed loop from data acquisition to semantic evaluation and improving the accuracy of the evaluation.
[0113] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. A construction progress assessment method based on multi-UAV visual collaboration, characterized in that, include: Using predefined BIM component data, a three-element line-of-sight volume of component-UAV-time is constructed. Combined with the dynamic object trajectory set extracted from pre-stored aligned multi-UAV image sequences, occlusion risk is calculated, and a candidate set of component view responsibility is generated. Construct a component responsibility grid, map the candidate set of component-perspective responsibilities to feasible units within the component responsibility grid, perform wavefront expansion search on the component responsibility grid, and generate a component-perspective responsibility schedule by calculating the cumulative cost that minimizes the succession cost. Based on the component perspective responsibility schedule, multi-view observation data are aggregated from the aligned multi-UAV image sequence, local 3D reconstruction and state determination are performed, and a short-term state sequence of the component is generated.
2. The method according to claim 1, characterized in that, Mapping the candidate set of responsibility from the component perspective to feasible units within the component responsibility grid includes: For any component, a two-dimensional mesh structure is established with discrete time steps as columns and UAV numbers as rows; By analyzing the time interval information in the candidate set of component responsibility, the component-UAV-time triplet that meets the preset observation requirements is marked as a feasible unit in the two-dimensional grid structure, forming a component responsibility grid.
3. The method according to claim 2, characterized in that, Generate a component-perspective responsibility timeline, including: Construct the dwell cost of maintaining the same UAV observation in adjacent discrete time steps, and the relay cost of switching between different UAV observations in adjacent discrete time steps; Dynamic programming calculations are performed on feasible units in the component responsibility grid along the time dimension, the dwell cost and the relay cost are accumulated, and the minimum cumulative cost path to each feasible unit is determined. By backtracking the path with the minimum cumulative cost from the end of the discrete time step, the optimal UAV responsibility sequence is extracted, and a responsibility schedule from the perspective of components is generated.
4. The method according to claim 1, characterized in that, This also includes revising the component-perspective responsibility timeline, specifically: Summarize the component-perspective responsibility timeline for all components and construct a global load curve reflecting the drone mission density; Based on preset kinematic constraints, detect conflict periods in the global load curve and identify relevant components within the conflict periods; During the conflict period, perform local wavefront extension searches on relevant components, reallocate UAV responsibilities, and generate a revised component view responsibility schedule.
5. The method according to claim 1, characterized in that, Constructing a three-dimensional line-of-sight volume of components, drones, and time using predefined BIM component data, including: Extract the 3D geometric center and target area of the component from the predefined BIM component data, and obtain the pose parameters of the UAV camera from the aligned multi-UAV image sequence; Construct a cone-shaped field of view pointing from the optical center of the UAV camera to the target area, and map the cone-shaped field of view to a unified three-dimensional voxel space to generate a ternary line of view volume indexed by the component, the UAV, and time.
6. The method according to claim 5, characterized in that, Calculate occlusion risk, including: The predicted positions in the dynamic object trajectory set are transformed into discrete dynamic object voxels; The percentage of the overlapping volume between the three-dimensional line-of-sight volume of the component-UAV-time and the dynamic object voxel in a unified spatiotemporal coordinate system is calculated. A quantitative occlusion risk is generated based on the percentage of overlapping volumes.
7. The method according to claim 1, characterized in that, Also includes: Extract the physical objects actually observed on site from the short-time state sequence of the components, construct a set of physical entity nodes, and establish a set of target component nodes based on predefined BIM component data; Initialize the candidate correspondence between the physical entity node set and the target component node set, and construct a temporary-permanent hybrid topological semantic graph containing entity nodes, target component nodes and topological constraint nodes.
8. The method according to claim 7, characterized in that, Also includes: Read the temporal behavior features and geometric connectivity relationships in the short-time state sequence of components, and perform multiple rounds of probability updates on the temporary-permanent hybrid topological semantic graph; Based on the temporary-permanent hybrid topological semantic graph updated by multiple rounds of probability, the candidate correspondence is shrunk by using the consistency check of topological constraint nodes, and the unique entity identity of the physical entity is determined by the delayed binding mechanism.
9. The method according to claim 8, characterized in that, Also includes: Identify nodes in the temporary-permanent hybrid topological semantic graph that still have multiple identity solutions or topological constraint conflicts after multiple rounds of probability updates, and generate a list of high uncertainty components; By combining the spatial distribution of the list of high-uncertainty components with the weight of construction impact, an observation optimization suggestion set for the next round of data collection is generated; The observation optimization suggestion set is fed back into the step of constructing a component-drone-time ternary line-of-sight volume using predefined BIM component data, and is used as an additional weight to adjust the screening and occlusion risk calculation of the next round of component-drone-time ternary line-of-sight volume.
10. The method according to claim 1, characterized in that, Aggregating multi-view observation data from aligned multi-UAV image sequences, including: By analyzing the responsibility switching moments recorded in the responsibility timeline from the component perspective, the aligned multi-UAV image sequence is segmented into responsibility segments continuously handled by different UAVs; For any given component, multi-angle image frames of the same component are extracted and aggregated across the responsible segments to construct multi-view observation data for that component.