Construction site image visual angle alignment method
By constructing a multi-view rendering library and occlusion robust optimization technology, the accuracy and reliability issues of aligning construction site images with BIM models were resolved, achieving efficient and accurate view alignment and a self-correction mechanism.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU HAOLINK INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-04-13
- Publication Date
- 2026-05-12
AI Technical Summary
Existing solutions for aligning construction site images with BIM model perspectives are not very accurate in environments with occlusion, changes in lighting, and repetitive structures. Furthermore, they lack reliability assessment and self-correction mechanisms, leading to false alarms and inefficiency.
By building a multi-view rendering library, we can perform coarse pose estimation using stable structural features, and optimize fine pose based on occlusion robust weights and robust loss functions to shield occlusion interference and improve alignment accuracy.
It achieves high-precision view alignment in complex environments, reduces false alarms, improves the alignment accuracy and reliability of construction site images and BIM models, and reduces the cost of manual intervention.
Smart Images

Figure CN122023490A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of building information processing technology, and in particular to a method for aligning the viewpoints of construction site images. Background Technology
[0002] Currently, in intelligent construction inspection scenarios, a common practice is to compare photos or videos taken at the construction site with a BIM (Building Information Modeling) model to determine missing components, misalignments, or construction non-standard practices. To achieve reliable comparison, the actual site image and the BIM rendering result must first be at the same viewpoint, i.e., the precise pose (including position, orientation, and intrinsic parameters) of the on-site camera relative to the BIM coordinate system must be determined. Currently used viewpoint alignment schemes mainly include: (1) Method based on manual point selection or manual registration: The inspection personnel manually select feature points in the on-site images and register them with the corresponding points in BIM or CAD (Computer-Aided Design). This type of method is inefficient, highly subjective, and has poor error prevention, and is not suitable for large-scale automatic inspection; (2) Methods based on general visual SLAM (Simultaneous Localization and Mapping) or 3D reconstruction: SLAM or SfM (Structure from Motion) is performed on site video to obtain camera trajectory, which is then coarsely aligned with BIM. These methods rely on texture and stable lighting, and are prone to drift or failure due to severe occlusion, dust, reflection, and structural repetition (such as column grid and beam grid) in construction site scenes. (3) Image retrieval or pure geometric ICP (Iterative Closest Point) method: first, search for approximate viewpoints in the pre-rendered BIM viewpoint library, and then perform local optimization. However, since BIM and real-world images are cross-domain data (i.e., geometric rendering and real-world pixels), directly relying on pixel similarity or point cloud ICP often results in mismatches; (4) Lack of reliability assessment of alignment results: Most methods directly use a pose for subsequent difference detection, lacking an automatic identification and correction mechanism for alignment failures caused by occlusion or temporary objects (such as scaffolding, templates, machinery, and people), resulting in widespread false alarms.
[0003] In summary, the main problems or shortcomings of existing viewpoint alignment schemes include: (1) Unstable cross-domain alignment: There are large differences between the appearance of BIM and the image, and it is difficult to stably match the general features; (2) Occlusion sensitivity: Occlusion caused by temporary objects or construction process will lead to position and pose solution deviation and is difficult to automatically identify; (3) Lack of confidence and self-correction loop: alignment failure will not trigger relocation, affecting the credibility of subsequent violation detection. Summary of the Invention
[0004] In view of the above, it is necessary to provide a method for aligning the viewpoints of construction site images, which aims to solve the problem of low viewpoint alignment accuracy in construction sites with complex environmental conditions such as occlusion, changes in lighting, and repetitive structures.
[0005] A method for aligning the viewpoints of construction site images, the method comprising: In response to the image viewpoint alignment trigger signal of the target construction site, the site image data of the target construction site and the stable structural features of each type of BIM component in the corresponding BIM model that can be used for registration are acquired. A multi-view rendering library is constructed based on the stable structural features described above. The on-site image data is processed to obtain two-dimensional structural features and a non-interpretive domain mask; Coarse pose estimation is performed based on the two-dimensional structural features and the multi-view rendering library to obtain the coarse pose. A cross-domain reprojection error objective function is constructed based on occlusion robust weights, and the coarse pose is optimized according to the cross-domain reprojection error objective function to obtain the fine pose.
[0006] A construction site image perspective alignment device, the construction site image perspective alignment device comprising: The acquisition unit is used to acquire the on-site image data of the target construction site and the stable structural features that can be registered for each type of BIM component in the corresponding BIM model in response to the image view alignment trigger signal of the target construction site. A construction unit is used to construct a multi-view rendering library based on the stable structural features; The processing unit is used to process the on-site image data to obtain two-dimensional structural features and non-interpretive domain masks. The estimation unit is used to perform coarse pose estimation based on the two-dimensional structural features and the multi-view rendering library to obtain the coarse pose. The optimization unit is used to construct a cross-domain reprojection error objective function based on occlusion robust weights, and optimize the coarse pose according to the cross-domain reprojection error objective function to obtain the fine pose.
[0007] A computer device, the computer device comprising: A memory that stores at least one instruction; and a processor that executes the instruction stored in the memory to implement the construction site image viewpoint alignment method.
[0008] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the construction site image perspective alignment method.
[0009] As can be seen from the above technical solutions, the present invention can construct a multi-view rendering library covering typical inspection perspectives based on stable structural features, avoiding matching failures caused by viewpoint omissions, and solving the problem of unstable cross-domain alignment between BIM and construction site images; coarse pose estimation is performed based on two-dimensional structural features and the multi-view rendering library, and fine pose is obtained by optimizing the coarse pose based on the cross-domain reprojection error objective function; combined with occlusion robust weights and robust loss functions, abnormal feature interference caused by temporary object occlusion is effectively shielded, improving the accuracy of viewpoint alignment. Attached Figure Description
[0010] Figure 1 This is a flowchart of a preferred embodiment of the construction site image perspective alignment method of the present invention; Figure 2 This is a functional block diagram of a preferred embodiment of the construction site image perspective alignment device of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device that implements the construction site image perspective alignment method of the present invention. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0012] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the construction site image perspective alignment method of the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.
[0013] The construction site image perspective alignment method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0014] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0015] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0016] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0017] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0018] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0019] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0020] S10, in response to the image view alignment trigger signal of the target construction site, acquire the on-site image data of the target construction site, and the stable structural features of each type of BIM component in the corresponding BIM model (Building Information Modeling) that can be registered.
[0021] In this embodiment, the image viewpoint alignment trigger signal can be any signal used to trigger viewpoint alignment, such as a voice signal, a touch signal, etc.
[0022] In this embodiment, the on-site image data can be a single frame of image or a video keyframe.
[0023] The on-site video data can be collected by mobile devices such as smartphones and tablets, or by drones and fixed cameras (such as construction site monitoring cameras).
[0024] The on-site image data may also include coarse pose information collected by IMU (Inertial Measurement Unit), GNSS (Global Navigation Satellite System), odometer, etc.
[0025] In this embodiment, the BIM model may include component-level three-dimensional geometric information (such as size, location, and shape) and semantic information (such as component type: column, beam, wall, slab, opening, axis; component attributes: material, elevation, number, etc.).
[0026] The above embodiments enable the screening of high-value image data and reduce the interference of redundant frames on subsequent processing.
[0027] The BIM model can be preprocessed semantically to extract stable structural features for each type of component that can be used for registration, avoiding easily deformable or temporary features.
[0028] For example, the stable structural features include, but are not limited to, the following: (1) Structural boundary line segments and outline patches: such as the top edge line segment of a beam and the side elevation outline patch of a wall; (2) Main corner points and intersection lines: such as the three-dimensional corner points at the intersection of beams and columns, the vertical intersection lines of wall panels, and the intersection lines of rectangular boundaries of openings; (3) Grid, elevation, control lines: If the BIM model contains building grid (such as X / Y direction axis), floor elevation lines, component center lines, etc., it can be directly extracted as global constraint features.
[0029] Through the above embodiments, core registration features can be extracted from the BIM model to avoid irrelevant information (such as decorative details) from affecting alignment accuracy, providing structured and semantic BIM basic data for subsequent prior library construction and improving the stability of cross-domain matching.
[0030] S11, Construct a multi-view rendering library based on the stable structural features.
[0031] In this embodiment, constructing a multi-view rendering library based on the stable structural features includes: In the BIM coordinate system, the inspection and shooting range constraints are configured according to the construction area of the target construction site and the parameters of the inspection equipment. Based on the inspection and shooting range constrained within the construction area, spatial discretization sampling is performed based on the geometric topological relationship of the construction area to obtain multiple candidate machine locations; A sampling area is generated with each candidate machine site as the center and a preset length as the radius; At each candidate camera position, the camera's line of sight is directed towards the geometric center of the construction area, and multiple deflection angles are sampled within the sampling area corresponding to each candidate camera position to obtain multiple preliminary candidate camera poses; Based on the stable structural features, the visibility of the multiple preliminary candidate camera poses is filtered to obtain multiple intermediate candidate camera poses. The score of each intermediate candidate camera pose is calculated according to a preset dimension, and the intermediate candidate camera poses with scores greater than a preset score threshold are selected from the multiple intermediate candidate camera poses as multiple candidate camera poses. Render each candidate camera pose to obtain the multimodal prior of each candidate camera pose; The multi-view rendering library is constructed based on the multimodal priors of each candidate camera pose.
[0032] The constraints on the inspection and shooting range may include constraints on shooting distance, height, field of view, pitch angle, and safe accessible space.
[0033] Candidate aircraft sites can be sampled at set intervals in the outer enclosure of the construction area or in the walkable or flyable space.
[0034] The sampling area corresponding to each candidate machine site can be the area near each candidate machine site.
[0035] Among these methods, a 3D visibility analysis can be performed based on the BIM model to calculate the visible area ratio of the target component (such as column or beam) under each candidate pose (e.g., under a certain pose, the visible area ratio of the column is 80%, the visible area ratio of the beam is 75%, and the overall visibility rate is 77.5%). Pose with a visibility rate ≥ 60% is retained (50 poses remain after screening). Then, a score is assigned based on "visibility rate × 0.6 + line-of-sight adaptation × 0.3 + pitch angle rationality × 0.1" (e.g., line-of-sight adaptation: 1 point for 8-15m, 0.8 points for 15-20m; pitch angle -10°~10°: 1 point, and the rest: 0.8 points). Pose with a score greater than 0.95 is selected as the final multiple candidate camera poses (or the poses can be sorted from high to low according to the score, and the top 30 poses are selected as the final multiple candidate camera poses).
[0036] After rendering each candidate camera pose, we can obtain an RGB (Red, Green, Blue) rendering map, a depth map, a semantic instance mask, a rendering edge map, and a normal map.
[0037] The RGB rendering map simulates the color rendering of components under realistic lighting (e.g., gray for concrete columns, gray for beams); the depth map represents the three-dimensional distance from the BIM component to the camera for each pixel; the semantic instance mask represents the unique pixel value assigned to each component (e.g., 255 for columns, 128 for beams, 64 for walls), used to achieve pixel-level semantic differentiation; the rendering edge map is used to highlight the boundary contour of the component; and the normal map represents the direction of the surface normal vector of the component (encoded with RGB colors).
[0038] In this embodiment, constructing the multi-view rendering library based on the multimodal prior of each candidate camera pose includes: For each candidate camera pose, a rendered edge map is obtained from the multimodal prior of the candidate camera pose, and features are extracted from the rendered edge map to obtain a structural edge descriptor. Semantic instance masks are obtained from the multimodal priors of the candidate camera poses, and features are extracted from the semantic instance masks to obtain semantic layout vectors. By concatenating the structural edge descriptor with the semantic layout vector, the concatenated feature is obtained; The stitched features are normalized to obtain the first comprehensive feature vector of each candidate camera pose; Obtain the viewpoint number for each candidate camera pose; The first comprehensive feature vector of each candidate camera pose, each candidate camera pose and the number of each candidate camera pose are stored in a vector database, and an index is built using a vector retrieval structure to obtain the multi-view rendering library.
[0039] Specifically, Canny edge detection can be performed on the rendered edge map, dividing the image into a grid. Each grid is used to calculate the edge intensity histogram in each direction. After stitching, the structural contour information of the entire image is encoded into a one-dimensional fixed-length edge descriptor vector, which is used to distinguish the geometric structure distribution under different viewpoints.
[0040] Specifically, by statistically analyzing the pixel proportions of each component category in different image regions within the semantic instance mask, the semantic spatial distribution of the image is encoded into a fixed-length semantic layout vector, thereby using semantic distribution instead of pixel distribution for viewpoint indexing.
[0041] By replacing traditional pixel features with a comprehensive feature of "semantic distribution (corresponding to the semantic instance mask) + geometric structure (corresponding to the rendering edge map)," the problem of cross-domain feature differences between BIM and images is solved.
[0042] The viewpoint number for each candidate camera pose can be generated based on the order of the poses when determining multiple candidate camera poses as described above.
[0043] Each candidate camera pose can be stored using parameters for each candidate camera pose. The parameters for each candidate camera pose can include rotation matrix, translation vector, and intrinsic parameters.
[0044] Through the above embodiments, a priori library covering typical inspection perspectives can be constructed to avoid matching failures caused by missing perspectives; the index design based on the vector database enables fast retrieval of candidate perspectives, which can help improve the efficiency of subsequent coarse pose estimation.
[0045] S12, the on-site image data is processed to obtain two-dimensional structural features and non-interpretive domain masks.
[0046] In this embodiment, the processing of the on-site image data to obtain two-dimensional structural features and a non-interpretive domain mask includes: Pixel-level semantic segmentation is performed on the on-site image data to obtain segmentation results of structurally significant components; Based on the segmentation result, the extraction region of the on-site image data is located as a stable structure region, and two-dimensional component features corresponding to the stable structure features are extracted from the stable structure region to obtain the two-dimensional structure features. The regions in the on-site image data other than the stable structural region are defined as non-interpretation domains; Identify temporary objects and non-BIM objects within the non-interpretation domain, and generate the non-interpretation domain mask based on the temporary objects and the non-BIM objects.
[0047] The non-interpretive domain mask can be used to downsample or remove weights from feature points corresponding to the occluded region when calculating the reprojection error.
[0048] The structurally significant components may include, but are not limited to, columns, beams, walls, openings, floor slab boundaries, etc.
[0049] Accordingly, the segmentation results of the structurally significant components correspond to the pixel positions of columns, beams, walls, openings, floor slab boundaries, etc.
[0050] Among these methods, edge detection can be performed on the segmented stable structural region (i.e., the component region), and boundary line segments with continuous lengths greater than or equal to a specified number of pixels (such as the left edge line segment of a column) can be extracted to obtain a semantic boundary line segment set; corner detection algorithms can be used to extract the corner points of the component boundaries (such as the four corner points at the top of a column) and the intersection points of different components (such as the two-dimensional projection intersection points at the beam-column junction); parallel straight lines on the component surface can also be detected by Hough transform, and the vanishing line direction can be fitted to obtain the vanishing line direction (such as the vanishing line direction of a horizontal beam being 0°), which can then be used as a global geometric constraint.
[0051] In the above embodiments, core registration features are accurately extracted from the on-site images and matched with BIM preprocessing features, laying the foundation for cross-domain matching.
[0052] For the non-interpretive domain, a target detection model can be used to identify temporary objects and non-BIM objects in the image, including scaffolding, formwork, construction machinery (tower cranes, concrete pump trucks), personnel, and material stockpiles (steel bars, cement bags), etc. The pixel values of the identified non-BIM object areas are set to 0, and the remaining areas are set to 255, generating a non-interpretive domain mask for subsequent masking of interference in occluded areas.
[0053] In the above embodiments, non-BIM temporary object areas are marked in advance to avoid interference from occlusion objects on subsequent pose solving and improve the occlusion robustness of the method.
[0054] S13, perform coarse pose estimation based on the two-dimensional structural features and the multi-view rendering library to obtain the coarse pose.
[0055] In this embodiment, the step of performing coarse pose estimation based on the two-dimensional structural features and the multi-view rendering library to obtain the coarse pose includes: Generate the second comprehensive feature vector of the two-dimensional structural features; The second comprehensive feature vector is used to search the multi-view rendering library to obtain multiple candidate BIM views and the image feature similarity corresponding to each candidate BIM view. According to the pre-defined coordinate transformation relationship, the coarse pose information output by the inspection equipment is transformed from the sensor world coordinate system to the BIM coordinate system to obtain the sensor coarse pose in the BIM coordinate system. Obtain the camera pose for each candidate BIM view; For each candidate BIM viewpoint, the position difference is obtained by calculating the Euclidean distance between the camera pose and the sensor coarse pose, and the orientation difference is obtained by calculating the angle between the optical axis direction of the camera pose and the sensor coarse pose. The position difference and the orientation difference are linearly combined according to the first preset weight to obtain the sensor consistency distance; Convert the sensor consistency distance into a sensor consistency score; The image feature similarity corresponding to each candidate BIM viewpoint is normalized to obtain the image similarity score corresponding to each candidate BIM viewpoint. According to the second preset weight, the sensor consistency score and image similarity score corresponding to each candidate BIM view are weighted and fused to obtain the comprehensive score of each candidate BIM view. The camera pose corresponding to the candidate BIM view with the highest comprehensive score among the multiple candidate BIM views is obtained as the coarse pose.
[0056] The second comprehensive feature vector can be generated in the same way as the first comprehensive feature vector was generated earlier, which will not be elaborated here.
[0057] When using the second comprehensive feature vector to perform a search in the multi-view rendering library, candidate BIM views with Top-K (K=5) image feature similarity can be selected.
[0058] The first preset weight and the second preset weight can be optimal values selected based on a large number of experiments.
[0059] The sensor consistency distance can be converted into the sensor consistency score using an exponential function.
[0060] In addition, several perspectives with high comprehensive scores can be retained as a set of initial values for multiple hypotheses to provide redundancy guarantees for low confidence scenarios and improve the robustness of subsequent optimization processes.
[0061] In the above embodiments, cross-domain retrieval can quickly lock the approximate viewpoint, avoiding the inefficiency of global search; the fusion of multi-source coarse pose information (IMU / GNSS) improves the accuracy of the initial coarse pose value, reducing the number of iterations for subsequent fine optimization.
[0062] S14. Construct a cross-domain reprojection error objective function based on occlusion robust weights, and optimize the coarse pose according to the cross-domain reprojection error objective function to obtain the fine pose.
[0063] In this embodiment, the objective function for constructing cross-domain reprojection error based on occlusion robust weights includes: Within the preset deviation range of the coarse pose, a correspondence between BIM three-dimensional semantic geometric features and image two-dimensional semantic features is established based on the stable structural features, the segmentation results of the structurally significant components, the two-dimensional structural features, and the non-interpretive domain mask. Construct the occlusion robust weights, which include non-interpretive domain weights, stable structural region weights, and a robust loss function; Using the occlusion robustness weight as a constraint, the edge distance error, mask cross-union ratio error, and consistency error are calculated according to the correspondence. Obtain a first error corresponding to the edge distance error, a second error corresponding to the mask crossover ratio error, and a third error corresponding to the consistency error; The objective function for cross-domain reprojection error is obtained by weighting the edge distance error, the mask cross-union ratio error, and the consistency error based on the first error, the second error, and the third error. Specifically, the non-interpretive domain weights are used to remove or assign low weights to features falling within the non-interpretive domain mask; the stable structural region weights are used to assign high weights to visible stable structural regions; and the robust loss function is used to suppress outliers. Wherein, the edge distance error represents the distance error between the BIM projection edge and the image semantic edge; the mask intersection-union ratio error represents the intersection-union ratio error between the BIM semantic mask projection and the image segmentation mask; and the consistency error represents the consistency error between the rendering depth and the image estimated depth.
[0064] The preset deviation may include the area near the coarse pose, such as a rotation deviation of ±5° or a translation deviation of ±1m.
[0065] The correspondence may include, but is not limited to: 3D structural boundary / corner point ↔ 2D semantic boundary / corner point; 3D semantic instance projection mask ↔ 2D segmentation mask.
[0066] Specifically, the non-interpretive domain weights can avoid interference from occluded region features; the stable structure region weights can highlight stable structures; and the robust loss function can balance robustness and optimization accuracy.
[0067] In the above embodiments, feature matching based on semantic geometric constraints solves the problem of mismatch between cross-domain data (BIM geometry and on-site image pixels). The combination of occlusion robust weights and robust loss functions effectively shields abnormal feature interference caused by temporary object occlusion.
[0068] In this embodiment, optimizing the coarse pose based on the cross-domain reprojection error objective function to obtain the fine pose includes: With the objective of minimizing the cross-domain reprojection error, gradient descent is performed on the parameter vector corresponding to the coarse pose until convergence. The precise pose is generated based on the currently obtained parameter vector.
[0069] For example, the Perspective-n-Point (PnP) algorithm can be used to solve for the initial fine pose using a certain number of matched 3D and 2D corner points. Then, the local Bundle Adjustment (BA) algorithm is used, with the initial fine pose as the initial value, to minimize the cross-domain reprojection error objective function and iteratively optimize the camera pose and optional camera intrinsic parameters to obtain the fine pose. Of course, if the BA optimization does not converge, differentiable rendering gradient descent can also be used to calculate the gradient of the error function with respect to the camera pose and update the pose parameters according to the gradient direction until convergence, finally outputting the fine pose.
[0070] In the above embodiments, the combined solution strategy of PnP and BA balances solution speed and accuracy, and reduces fine pose error.
[0071] In this embodiment, after obtaining the precise pose, the method further includes: The target confidence level of the precise pose is calculated using the following formula: ; in, Indicates the target confidence level; α represents the attenuation coefficient; This represents the mean residual value obtained after performing residual distribution statistics on the values of the cross-domain reprojection error objective function within the stable structural region. This represents the variance penalty coefficient; This represents the residual variance obtained after performing residual distribution statistics on the values of the cross-domain reprojection error objective function within the stable structural region.
[0072] Furthermore, when the target confidence level is less than the confidence level threshold, the fine pose is self-corrected to obtain the corrected pose; The image view alignment result of the target construction site is generated based on the corrected pose.
[0073] Among them, the residual set of stable component regions such as columns, beams, and walls that are effectively matched can be statistically analyzed.
[0074] The attenuation coefficient can be adjusted according to the project scenario.
[0075] In the above embodiments, residual statistics based on region division can distinguish the error contribution of effective structures and occluded regions, making the confidence assessment more accurate; it can quantify the confidence of pose and provide a preliminary judgment basis for subsequent difference detection; the confidence design takes into account the residual mean (overall accuracy) and variance (stability), which can comprehensively reflect the pose quality.
[0076] The confidence threshold can be adjusted according to the accuracy requirements of the actual project.
[0077] In this embodiment, the step of self-correcting the fine pose to obtain the corrected pose includes: When the on-site image data is video, backtrack to the adjacent frames of the on-site image data to reprocess and obtain the two-dimensional structural features and the non-interpretive domain mask; or when the on-site image data is a single frame, obtain the camera pose corresponding to the candidate BIM view with the second highest comprehensive score from the multiple candidate BIM viewpoints as the backtracked coarse pose. The corrected pose is determined when the target confidence level is detected to be greater than or equal to the confidence threshold, or when the maximum number of backtracking steps is reached.
[0078] For example: If the input is video, it automatically backtracks to the adjacent keyframe (the previous frame or the next frame), and re-executes S12 (on-site image feature extraction) ~ confidence evaluation until a pose with confidence ≥ confidence threshold is obtained, or the number of backtracking frames reaches the upper limit (e.g., 5 frames); if the input is a single frame, it selects the next view with the highest comprehensive score from the Top-K candidate BIM views in S13 (e.g., the original selected view is candidate BIM view A, and the next selected view is candidate BIM view M) as the new coarse pose initial value, and re-executes S14 (fine pose optimization) ~ confidence evaluation until a pose with confidence ≥ confidence threshold is obtained or the number of iterations reaches the upper limit (e.g., 3 times).
[0079] If the confidence level still cannot be met by adjacent frames, the search step size can be increased (e.g., by performing frame skipping back) or the sensor can be recalibrated.
[0080] The above embodiments can automatically repair low-confidence misaligned poses, reduce the cost of manual intervention, and form a closed-loop mechanism of "optimization-evaluation-correction". The combination of video frame backtracking and single-frame multi-candidate iteration can cover the self-correction needs of different input scenarios.
[0081] In this embodiment, generating the image viewpoint alignment result of the target construction site based on the corrected pose includes: Based on the corrected pose, the BIM model is re-rendered to obtain a view-aligned rendering image that is consistent with the viewpoint of the on-site image data. Generate alignment availability markers based on the corrected pose; The image viewpoint alignment result is obtained by combining the corrected pose, the viewpoint alignment rendering, the detected target confidence, and the alignment availability marker.
[0082] The alignment availability flag can be a Boolean value. For example, True indicates that the alignment is valid and can be used for subsequent difference detection; False indicates that the alignment failed and requires manual intervention.
[0083] Usability markers can clearly distinguish alignment effects, providing a reliable entry point for subsequent difference detection and violation judgment, and improving the credibility of the overall inspection process.
[0084] The above embodiments can provide comprehensive image viewpoint alignment results.
[0085] As can be seen from the above technical solutions, the present invention can construct a multi-view rendering library covering typical inspection perspectives based on stable structural features, avoiding matching failures caused by viewpoint omissions, and solving the problem of unstable cross-domain alignment between BIM and construction site images; coarse pose estimation is performed based on two-dimensional structural features and the multi-view rendering library, and fine pose is obtained by optimizing the coarse pose based on the cross-domain reprojection error objective function; combined with occlusion robust weights and robust loss functions, abnormal feature interference caused by temporary object occlusion is effectively shielded, improving the accuracy of viewpoint alignment.
[0086] like Figure 2 The diagram shown is a functional block diagram of a preferred embodiment of the construction site image perspective alignment device of the present invention. The construction site image perspective alignment device 11 includes an acquisition unit 110, a construction unit 111, a processing unit 112, an estimation unit 113, and an optimization unit 114. The module / unit referred to in this invention refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0087] The acquisition unit 110 is used to acquire the on-site image data of the target construction site and the stable structural features that can be registered for each type of BIM component in the corresponding BIM model in response to the image view alignment trigger signal of the target construction site. The construction unit 111 is used to construct a multi-view rendering library based on the stable structural features; The processing unit 112 is used to process the on-site image data to obtain two-dimensional structural features and non-interpretive domain masks. The estimation unit 113 is used to perform coarse pose estimation based on the two-dimensional structural features and the multi-view rendering library to obtain a coarse pose. The optimization unit 114 is used to construct a cross-domain reprojection error objective function based on occlusion robust weights, and optimize the coarse pose according to the cross-domain reprojection error objective function to obtain the fine pose.
[0088] As can be seen from the above technical solutions, the present invention can construct a multi-view rendering library covering typical inspection perspectives based on stable structural features, avoiding matching failures caused by viewpoint omissions, and solving the problem of unstable cross-domain alignment between BIM and construction site images; coarse pose estimation is performed based on two-dimensional structural features and the multi-view rendering library, and fine pose is obtained by optimizing the coarse pose based on the cross-domain reprojection error objective function; combined with occlusion robust weights and robust loss functions, abnormal feature interference caused by temporary object occlusion is effectively shielded, improving the accuracy of viewpoint alignment.
[0089] like Figure 3 The diagram shown is a schematic diagram of the computer device for implementing the construction site image perspective alignment method of the present invention.
[0090] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a construction site image perspective alignment program.
[0091] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.
[0092] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.
[0093] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as the code for a construction site image perspective alignment program, but also to temporarily store data that has been output or will be output.
[0094] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., executing a construction site image viewpoint alignment program) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.
[0095] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes these applications to implement the steps in the various construction site image perspective alignment method embodiments described above, for example... Figure 1 The steps are shown.
[0096] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into an acquisition unit 110, a construction unit 111, a processing unit 112, an estimation unit 113, and an optimization unit 114.
[0097] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the construction site image perspective alignment method described in the various embodiments of this invention.
[0098] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.
[0099] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.
[0100] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.
[0101] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0102] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.
[0103] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0104] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the computer device 1 and other computer devices.
[0105] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.
[0106] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0107] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0108] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement a construction site image viewpoint alignment method, and the processor 13 can execute the multiple instructions to achieve the following: In response to the image viewpoint alignment trigger signal of the target construction site, the site image data of the target construction site and the stable structural features of each type of BIM component in the corresponding BIM model that can be used for registration are acquired. A multi-view rendering library is constructed based on the stable structural features described above. The on-site image data is processed to obtain two-dimensional structural features and a non-interpretive domain mask; Coarse pose estimation is performed based on the two-dimensional structural features and the multi-view rendering library to obtain the coarse pose. A cross-domain reprojection error objective function is constructed based on occlusion robust weights, and the coarse pose is optimized according to the cross-domain reprojection error objective function to obtain the fine pose.
[0109] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0110] It should be noted that all the data involved in this case was legally obtained.
[0111] If any AI models, software tools, or components not belonging to this company appear in the embodiments of this invention, they are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this invention has been obtained by an entity authorized (with the knowledge and consent) or fully authorized by all parties through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0112] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0113] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0114] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0115] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0116] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0117] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0118] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for aligning the viewpoints of construction site images, characterized in that, The method for aligning the viewpoints of construction site images includes: In response to the image viewpoint alignment trigger signal of the target construction site, the site image data of the target construction site and the stable structural features of each type of BIM component in the corresponding BIM model that can be used for registration are acquired. A multi-view rendering library is constructed based on the stable structural features described above. The on-site image data is processed to obtain two-dimensional structural features and a non-interpretive domain mask; Coarse pose estimation is performed based on the two-dimensional structural features and the multi-view rendering library to obtain the coarse pose. A cross-domain reprojection error objective function is constructed based on occlusion robust weights, and the coarse pose is optimized according to the cross-domain reprojection error objective function to obtain the fine pose.
2. The construction site image perspective alignment method as described in claim 1, characterized in that, The construction of the multi-view rendering library based on the stable structural features includes: In the BIM coordinate system, the inspection and shooting range constraints are configured according to the construction area of the target construction site and the parameters of the inspection equipment. Based on the inspection and shooting range constrained within the construction area, spatial discretization sampling is performed based on the geometric topological relationship of the construction area to obtain multiple candidate machine locations; A sampling area is generated with each candidate machine site as the center and a preset length as the radius; At each candidate camera position, the camera's line of sight is directed towards the geometric center of the construction area, and multiple deflection angles are sampled within the sampling area corresponding to each candidate camera position to obtain multiple preliminary candidate camera poses; Based on the stable structural features, the visibility of the multiple preliminary candidate camera poses is filtered to obtain multiple intermediate candidate camera poses. The score of each intermediate candidate camera pose is calculated according to a preset dimension, and the intermediate candidate camera poses with scores greater than a preset score threshold are selected from the multiple intermediate candidate camera poses as multiple candidate camera poses. Render each candidate camera pose to obtain the multimodal prior of each candidate camera pose; The multi-view rendering library is constructed based on the multimodal priors of each candidate camera pose.
3. The construction site image perspective alignment method as described in claim 2, characterized in that, The construction of the multi-view rendering library based on the multimodal prior of each candidate camera pose includes: For each candidate camera pose, a rendered edge map is obtained from the multimodal prior of the candidate camera pose, and features are extracted from the rendered edge map to obtain a structural edge descriptor. Semantic instance masks are obtained from the multimodal priors of the candidate camera poses, and features are extracted from the semantic instance masks to obtain semantic layout vectors. By concatenating the structural edge descriptor with the semantic layout vector, the concatenated feature is obtained; The stitched features are normalized to obtain the first comprehensive feature vector of each candidate camera pose; Obtain the viewpoint number for each candidate camera pose; The first comprehensive feature vector of each candidate camera pose, each candidate camera pose and the number of each candidate camera pose are stored in a vector database, and an index is built using a vector retrieval structure to obtain the multi-view rendering library.
4. The construction site image perspective alignment method as described in claim 2, characterized in that, The processing of the on-site image data to obtain two-dimensional structural features includes: Pixel-level semantic segmentation is performed on the on-site image data to obtain segmentation results of structurally significant components; Based on the segmentation results, the extraction region of the on-site image data is located as a stable structural region, and two-dimensional component features corresponding to the stable structural features are extracted from the stable structural region to obtain the two-dimensional structural features.
5. The construction site image perspective alignment method as described in claim 4, characterized in that, The step of performing coarse pose estimation based on the two-dimensional structural features and the multi-view rendering library to obtain the coarse pose includes: Generate the second comprehensive feature vector of the two-dimensional structural features; The second comprehensive feature vector is used to search the multi-view rendering library to obtain multiple candidate BIM views and the image feature similarity corresponding to each candidate BIM view. According to the pre-defined coordinate transformation relationship, the coarse pose information output by the inspection equipment is transformed from the sensor world coordinate system to the BIM coordinate system to obtain the sensor coarse pose in the BIM coordinate system. Obtain the camera pose for each candidate BIM view; For each candidate BIM viewpoint, the position difference is obtained by calculating the Euclidean distance between the camera pose and the sensor coarse pose, and the orientation difference is obtained by calculating the angle between the optical axis direction of the camera pose and the sensor coarse pose. The position difference and the orientation difference are linearly combined according to the first preset weight to obtain the sensor consistency distance; Convert the sensor consistency distance into a sensor consistency score; The image feature similarity corresponding to each candidate BIM viewpoint is normalized to obtain the image similarity score corresponding to each candidate BIM viewpoint. According to the second preset weight, the sensor consistency score and image similarity score corresponding to each candidate BIM view are weighted and fused to obtain the comprehensive score of each candidate BIM view. The camera pose corresponding to the candidate BIM view with the highest comprehensive score from the multiple candidate BIM views is obtained as the coarse pose.
6. The construction site image perspective alignment method as described in claim 5, characterized in that, The objective function for constructing cross-domain reprojection error based on occlusion robust weights includes: Within the preset deviation range of the coarse pose, a correspondence between BIM three-dimensional semantic geometric features and image two-dimensional semantic features is established based on the stable structural features, the segmentation results of the structurally significant components, the two-dimensional structural features, and the non-interpretive domain mask. Construct the occlusion robust weights, which include non-interpretive domain weights, stable structural region weights, and a robust loss function; Using the occlusion robustness weight as a constraint, the edge distance error, mask crossover ratio error, and consistency error are calculated according to the correspondence. Obtain a first error corresponding to the edge distance error, a second error corresponding to the mask crossover ratio error, and a third error corresponding to the consistency error; The objective function for cross-domain reprojection error is obtained by weighting the edge distance error, the mask cross-union ratio error, and the consistency error based on the first error, the second error, and the third error. Specifically, the non-interpretive domain weights are used to remove or assign low weights to features falling within the non-interpretive domain mask; the stable structural region weights are used to assign high weights to visible stable structural regions; and the robust loss function is used to suppress outliers. Wherein, the edge distance error represents the distance error between the BIM projection edge and the image semantic edge; the mask intersection-union ratio error represents the intersection-union ratio error between the BIM semantic mask projection and the image segmentation mask; and the consistency error represents the consistency error between the rendering depth and the image estimated depth.
7. The construction site image perspective alignment method as described in claim 6, characterized in that, The step of optimizing the coarse pose based on the cross-domain reprojection error objective function to obtain the fine pose includes: With the objective of minimizing the cross-domain reprojection error, gradient descent is performed on the parameter vector corresponding to the coarse pose until convergence. The precise pose is generated based on the currently obtained parameter vector.
8. The construction site image perspective alignment method as described in claim 7, characterized in that, After obtaining the precise pose, the method further includes: The target confidence level of the precise pose is calculated using the following formula: ; in, Indicates the target confidence level; α represents the attenuation coefficient; This represents the mean residual value obtained after performing residual distribution statistics on the values of the cross-domain reprojection error objective function within the stable structural region. This represents the variance penalty coefficient; This represents the residual variance obtained after performing residual distribution statistics on the values of the cross-domain reprojection error objective function within the stable structural region; When the target confidence level is less than the confidence threshold, the fine pose is self-corrected to obtain the corrected pose; The image view alignment result of the target construction site is generated based on the corrected pose.
9. The construction site image perspective alignment method as described in claim 8, characterized in that, The step of self-correcting the fine pose to obtain the corrected pose includes: When the on-site image data is video, backtrack to the adjacent frames of the on-site image data to reprocess and obtain the two-dimensional structural features and the non-interpretive domain mask; or when the on-site image data is a single frame, obtain the camera pose corresponding to the candidate BIM view with the second highest comprehensive score from the multiple candidate BIM viewpoints as the backtracked coarse pose. The corrected pose is determined when the target confidence level is detected to be greater than or equal to the confidence threshold, or when the maximum number of backtracking steps is reached.
10. The construction site image perspective alignment method as described in claim 9, characterized in that, The step of generating the image view alignment result of the target construction site based on the corrected pose includes: Based on the corrected pose, the BIM model is re-rendered to obtain a view-aligned rendering image that is consistent with the viewpoint of the on-site image data. Generate alignment availability markers based on the corrected pose; The image viewpoint alignment result is obtained by combining the corrected pose, the viewpoint alignment rendering, the detected target confidence, and the alignment availability marker.