Stacked material mass identification and generation method fusing single-view 3D reconstruction and BIM calibration

By integrating single-view 3D reconstruction and BIM calibration methods, using deep learning and geometric computing technology, the scale uncertainty problem in single-view reconstruction is solved, and efficient, accurate measurement and automated management of stack volume is achieved, which is suitable for digital management and real-time monitoring of construction sites.

CN120374885AActive Publication Date: 2025-07-25XIAMEN UNIV OF TECH

Patent Information

Application Number
CN202510827984.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-25
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The existing stack volume measurement technology has limitations in terms of accuracy, cost and degree of automation, and it is difficult to meet the needs of modern engineering construction for efficient, accurate and automated measurements. Especially in traditional single-view reconstruction technology, there is a problem of scale uncertainty, and it is impossible to directly obtain the real physical size and volume of the object.

Method used

The stack material volume recognition and generation method is integrated with single-view 3D reconstruction and BIM calibration. By acquiring two-dimensional images for rotational object detection and classification, ray-plane intersection calculation is performed by combining the internal parameters and external parameter matrix of the preset camera, a scaleless three-dimensional grid is generated using the OccNet model, and a linear scale factor is calculated through pinhole projection to realize the visual integration of the real volume of the stack material.

Benefits of technology

It realizes rapid and accurate measurement of stacking volumes on the construction site, reduces the cost of measurement equipment, improves measurement efficiency and accuracy, and reduces manual intervention through automated processes to support material management and cost control at the construction site.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374885A_ABST
    Figure CN120374885A_ABST
Patent Text Reader

Abstract

The invention provides a stacked material volume identification and generation method fusing single view 3D reconstruction and BIM calibration, and relates to the technical field of stacked material volume generation, and the method employs a low-cost monocular camera to collect images, and generates a three-dimensional model of stacked materials by means of an advanced deep learning model and a three-dimensional reconstruction algorithm. Through scale calibration with a virtual camera in the BIM model, the problem of scale uncertainty which is difficult to overcome in traditional single-view reconstruction is solved, so that the real volume of the stacked material can be accurately calculated. In addition, through automatic target detection, three-dimensional coordinate calculation and volume calibration processes, the measurement efficiency is greatly improved, and manual intervention is reduced. The method is not only suitable for volume measurement of static stacked materials, but also can be expanded to real-time monitoring of dynamic construction scenes, and provides powerful technical support for construction management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bulk material volume generation, and specifically relates to a method for identifying and generating bulk material volume by integrating monocular Figure 3 D reconstruction and BIM calibration. Background Art

[0002] In the fields of construction and project management, accurately measuring the volume of bulk materials at the construction site is a key task. It not only relates to the rational use of materials and cost control, but also directly affects the construction progress and safety management. However, the current technical means for measuring the volume of bulk materials have many deficiencies and are difficult to meet the requirements of modern engineering construction for efficient, accurate, and automated measurement.

[0003] Traditional methods for measuring the volume of bulk materials mainly rely on manual operations. For example, a tape measure or a measuring wheel is used to measure the base dimensions and height of the bulk materials, and the volume is estimated. This method is not only inefficient, but also the measurement results are extremely vulnerable to the experience and technical level of the operators and the interference of the on-site environment, resulting in unstable measurement accuracy and difficulty in ensuring the accuracy and consistency of the measurement results.

[0004] With the development of technology, multi-view photogrammetry and laser detection and ranging (LiDAR) scanning technology have been introduced into the field of bulk material volume measurement. These technologies can provide relatively accurate three-dimensional point cloud data, thus making it possible for high-precision volume calculation. However, the application of these technologies also faces many challenges. First of all, these technologies have high requirements for environmental conditions. For example, good lighting conditions, stable weather conditions, and accurate ground control points are required. Once the environmental conditions do not meet the requirements, the accuracy and reliability of the measurement results will be greatly reduced. In addition, from data acquisition to point cloud processing, three-dimensional modeling and then to the final volume calculation, the entire process is complex and time-consuming, and it is difficult to meet the need for rapid measurement and cannot provide accurate decision-making basis for construction management in a timely manner.

[0005] In recent years, monocular Figure 3 dimensional reconstruction technology, as a new technical means, has gradually attracted attention. This technology predicts the three-dimensional shape, orientation, and size of an object through a single color image (RGB image), and has the advantages of low cost and convenient deployment. However, due to the lack of depth information in monocular images, the reconstruction results often have scale uncertainty, that is, the true physical size and volume of the object cannot be directly obtained. This means that although the monocular Figure 3The 3D reconstruction technology can quickly generate a 3D model of an object, but the size of the model is only relative and cannot be directly applied to actual volume measurement. In addition, under different camera parameters, shooting distances, angles, etc., the reusability of the model is poor, and it is difficult to achieve standardized and automated scale recovery, which limits its wide application in actual engineering.

[0006] At the same time, Building Information Modeling (BIM) technology has been widely used in the field of engineering construction. BIM can provide accurate spatial geometric information and depth prior knowledge, providing strong support for the whole life cycle management of engineering projects. However, there is currently no mature technical solution that can effectively combine BIM technology with monocular 3D reconstruction technology to solve the scale ambiguity problem in monocular reconstruction and achieve accurate measurement of the volume of stacked materials.

[0007] Generally speaking, the existing technologies for measuring the volume of stacked materials have different degrees of limitations in terms of accuracy, cost, and automation, and it is difficult to meet the requirements of modern engineering construction for efficient, accurate, and automated measurement. In view of this, this application is proposed. Summary of the Invention

[0008] The present invention provides a method for identifying and generating the volume of stacked materials by integrating monocular Figure 3 3D reconstruction and BIM calibration, which can at least partially improve the above problems.

[0009] To achieve the above object, the present invention adopts the following technical solutions: A method for identifying and generating the volume of stacked materials by integrating monocular Figure 3 3D reconstruction and BIM calibration, which includes: Obtain the two-dimensional image collected by the preset camera, and use the preset multi-task detection model to process the two-dimensional image for rotating object detection and stacked material category classification tasks to obtain the contact points between the stacked materials and the ground; Obtain the internal parameter matrix and external parameter matrix of the preset camera. In the visual programming tool, reproduce the configuration of the internal parameter matrix and external parameter matrix, calculate the ray-plane intersection of the constructed ray from the midpoint of the bottom edge and the BIM ground plane equation, and obtain the real-space depth corresponding to the midpoint of the bottom edge in the image; Call the pre-trained OccNet model to calculate the two-dimensional image to generate a scale-free 3D mesh of the construction stacked materials; Use the visual programming tool to extract the scale-free 3D mesh to obtain the pixel coordinates of the four endpoints of the bottom surface of the three-dimensional minimum oriented bounding box, and calculate the Euclidean distance between the pixel coordinates and the contact points to obtain the final projected width; Under the pinhole camera model, the linear scale factor is calculated according to the final projection width, the real volume of the pile is calculated according to the linear scale factor, and the real volume of the pile is visualized and integrated using a visual programming tool.

[0010] In summary, the fusion single-view Figure 3 The method of identifying and generating the volume of piled materials by combining single-view Figure 3 The 3D reconstruction technology and the calibration function of the building information model (BIM) have realized the rapid and accurate measurement of the volume of the piled materials on the construction site. It can automatically complete target detection, coordinate calculation, 3D reconstruction, scale recovery and volume calibration in a single construction site image, and map the actual data of these construction piled materials into BIM, thereby realizing the real-time update of the BIM model.

[0011] Specifically, this method uses images taken by a monocular camera, with the help of advanced image processing and deep learning algorithms, to extract the geometric features of the piled materials and generate a three-dimensional model. By calibrating with the virtual camera in the BIM model and introducing real depth information, the scale uncertainty problem existing in the traditional single-view reconstruction technology is solved, so that the real volume of the piled materials can be accurately calculated. This method not only greatly reduces the cost of measuring equipment, but also reduces manual intervention through automated processes, thereby improving measurement efficiency and accuracy. In addition, this technology can be seamlessly integrated with the existing BIM system to achieve real-time data update and visual display, providing strong support for material management, cost control and safety monitoring at the construction site, and has significant engineering application value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 The fusion monocular vision provided by the embodiment of the present invention Figure 3 Schematic diagram of the process of identifying and generating the volume of piled materials based on D reconstruction and BIM calibration; Figure 2 is a schematic diagram of ray-plane intersection provided by an embodiment of the present invention; Figure 3 is a schematic diagram of an OBB local coordinate system provided by an embodiment of the present invention; Figure 4 It is a schematic diagram of pinhole projection provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0013] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0014] refer to Figure 1 As shown, the first embodiment of the present invention discloses a fusion single-view Figure 3Method for identifying and generating stacked material volume by 3D reconstruction and BIM calibration, which can be performed by a recognition and generation device for stacked material volume by 3D reconstruction and BIM calibration (hereinafter referred to as the recognition and generation device). Specifically, it is performed by one or more processors in the recognition and generation device to implement the following method: Figure 3 S1, obtain a two-dimensional image collected by a preset camera, and use a preset multi-task detection model to perform rotation target detection and stacked material category classification tasks on the two-dimensional image to obtain the contact points between the stacked materials and the ground; Specifically, step S1 further includes: obtaining a two-dimensional image of the construction stacked materials collected by a camera configured at the construction site, and preprocessing the two-dimensional image; Input the preprocessed two-dimensional image into a multi-task detection model constructed based on YOLOv8-OBB to perform rotation target detection and stacked material category classification tasks, and output information results. Among them, the information results include the four vertex pixel coordinates, confidence level, and category label of the rotation bounding box of each stacked material target; Screen out the information results with confidence values lower than the preset value, sort the four vertex pixel coordinates in descending order according to the ordinate of the pixel coordinates, and select the first two vertex pixel coordinates in the sequence as the endpoints of the bottom edge of the stacked material, marked as (u1, v1) and (u2, v2), corresponding to the contact points between the stacked material and the ground.

[0015] In this embodiment, at the construction site, a camera is pre-configured, and its position and angle are carefully set to ensure that the construction stacked materials can be clearly photographed. When the volume measurement of the stacked materials is required, first obtain the two-dimensional image of the construction stacked materials collected by this camera. Since the image may be affected by factors such as light and noise during the collection process, it needs to be preprocessed before being input into the detection model. The preprocessing process includes operations such as graying, binarization, and filtering of the image to remove noise and enhance the image contrast, thereby improving the accuracy of subsequent detection.

[0016] Subsequently, input the preprocessed two-dimensional image into a multi-task detection model constructed based on YOLOv8-OBB. YOLOv8-OBB is an advanced target detection model that can quickly and accurately detect rotation targets in an image and classify them. In this embodiment, the model is specially trained to enable it to identify the categories of construction stacked materials and detect the rotation bounding boxes of the stacked materials. The information results output by the model include the four vertex pixel coordinates, confidence level, and category label of the rotation bounding box of each stacked material target. The confidence level reflects the confidence degree of the model in this detection result. By setting a preset value, the information results with confidence levels lower than this value (set to 0.5 here) can be screened out, thereby improving the reliability of the detection results. ​

[0017] In addition to using the YOLOv8-OBB model, instance segmentation models such as Mask R-CNN can also be used to detect the contour of the stacked materials, and then the minimum bounding rectangle is used to fit the segmentation mask to obtain the rotated bounding box. This method can replace the two-dimensional image object detection and rotated box extraction method.

[0018] For the detection results screened by confidence, the pixel coordinates of the four vertices are sorted in descending order according to the ordinate of the pixel coordinates. Since the part of the stacked materials in contact with the ground is usually located at the bottom of the image, the first two vertex pixel coordinates in the sequence are selected as the endpoints of the bottom edge of the stacked materials, which are marked as (u1, v1) and (u2, v2) respectively. These two endpoints can accurately correspond to the contact points between the stacked materials and the ground, providing important basic data for subsequent three-dimensional coordinate calculation.

[0019] Through this step, the position and category of the stacked materials can be quickly and accurately detected in a single two-dimensional image, and the contact points between the stacked materials and the ground can be determined. This not only provides key information for subsequent three-dimensional reconstruction and volume calculation, but also due to the use of the advanced YOLOv8-OBB model, the entire detection process has high efficiency and accuracy, can process a large amount of image data in a short time, and meet the real-time requirements of the stacked materials volume measurement at the construction site. At the same time, through confidence screening and reasonable endpoint selection, the reliability and stability of the detection results are further improved, laying a solid foundation for the smooth progress of subsequent steps.

[0020] Please refer to Figure 2 , where Figure 2 E in represents the actual position of the camera, O represents the origin, H represents the actual height of the camera from the ground, represents the camera tilt angle, and d is the direction vector. S2. Obtain the internal parameter matrix and external parameter matrix of the preset camera. In the visual programming tool, reproduce the configuration of the internal parameter matrix and external parameter matrix, and perform ray-plane intersection calculation on the ray constructed by the midpoint of the bottom edge and the BIM ground plane equation to obtain the real space depth corresponding to the midpoint of the bottom edge in the image; Specifically, step S2 further includes: calibrating the preset camera to obtain its internal parameter matrix K and external parameter matrix [R|T]; , where the internal parameter matrix K is used to describe the internal optical characteristics of the camera, including the focal length, principal point coordinates, and distortion coefficients, and its matrix form is: is the focal length of the camera in the x direction, is the abscissa of the camera principal point, is the ordinate of the camera principal point; The external parameter matrix [R∣T] is used to represent the position and orientation of the camera in the world coordinate system. It consists of the rotation matrix R and the translation vector t, and its matrix form is: , is the translation component of the camera optical center on the x-axis of the world coordinate system, is the translation component of the camera optical center on the y-axis of the world coordinate system, is the translation component of the camera optical center on the z-axis of the world coordinate system, is the projection of the x-axis of the camera coordinate system in the x-axis direction of the world coordinate system, is the projection of the y-axis of the camera coordinate system in the x-axis direction of the world coordinate system, is the projection of the z-axis of the camera coordinate system in the x-axis direction of the world coordinate system, is the projection of the x-axis of the camera coordinate system in the y-axis direction of the world coordinate system, is the projection of the y-axis of the camera coordinate system in the y-axis direction of the world coordinate system, is the projection of the z-axis of the camera coordinate system in the y-axis direction of the world coordinate system, is the projection of the x-axis of the camera coordinate system in the z-axis direction of the world coordinate system, is the projection of the y-axis of the camera coordinate system in the z-axis direction of the world coordinate system, is the projection of the z-axis of the camera coordinate system in the z-axis direction of the world coordinate system.

[0021] Combined with the Revit software and the Dynamo visual programming tool, create a new perspective view, and set the view point position and orientation of this perspective view according to the internal parameter matrix and external parameter matrix of the preset camera to ensure that the virtual view point coincides exactly with the preset camera in the three-dimensional space; According to the contact point, calculate the midpoint coordinates (u m , v m ) of the bottom edge endpoints of the stacked materials. The formula is: , and use the inverse matrix K -1 of the internal parameter matrix K to convert the midpoint coordinates (u m , v m ) to the direction vector in the camera coordinate system. The formula is: , is the direction vector from the camera optical center to the current pixel point on the image, is the transpose operation; According to the formula Convert the direction vector in the camera coordinate system to the BIM world coordinate system, and in the BIM world coordinate system, starting from the camera center C, generate a ray along the direction vector to achieve the back-projection of the pixel point into the three-dimensional space. Among them, is the transpose of the camera rotation matrix; Ground plane equations in BIM models , and rays Perform the combined operation to obtain the ray-plane intersection parameters ,in, is the plane normal vector, is the offset of the plane relative to the origin of the world coordinate system, is any point in three-dimensional space; Set the ray-plane intersection parameters Substitute the intersection point P hit In the figure, the three-dimensional coordinates of the construction material in the world coordinate system are marked, and the formula is: and the intersection point P hit Remap to the camera coordinate system to get the vector , extract the third component of this vector As the coordinates of the midpoint of the bottom edge of the image (u m ,v m ) corresponds to the true depth of the 3D point relative to the camera, is the first component, is the second component, Represents a ray-plane intersection.

[0022] In this embodiment, first, the preset camera is calibrated, which is the basis of the whole process. Through calibration, the intrinsic parameter matrix and extrinsic parameter matrix of the camera can be obtained. The intrinsic parameter matrix describes the internal optical properties of the camera, including focal length, principal point coordinates, and distortion coefficients. The extrinsic parameter matrix describes the position and posture of the camera in the world coordinate system, which is composed of a rotation matrix and a translation vector.

[0023] After obtaining the camera's intrinsic and extrinsic matrix, use Revit software and Dynamo visual programming tools to create a new perspective view. Set the viewpoint position and orientation of the perspective view according to the preset camera's intrinsic and extrinsic matrix to ensure that the virtual viewpoint completely overlaps with the preset camera in three-dimensional space. This process provides an accurate virtual environment for subsequent three-dimensional space calculations, so that calculations performed in the virtual environment can correspond one-to-one with the actual environment. In the camera properties panel of this view, set the focal length f and the principal point coordinates (u0, v0) to the f corresponding to the known intrinsic matrix K. x ,f y ,u0,v0, ensure that the virtual projection has no deviation from the physical lens.

[0024] Next, based on the pixel coordinates (u1, v1) and (u2, v2) of the point where the pile of wood touches the ground obtained in the previous step, calculate the midpoint coordinates (u m ,vm ). Using the inverse matrix of the internal parameter matrix, the midpoint coordinates (u m ,v m ) is converted to the direction vector in the camera coordinate system. This conversion process is the key step in mapping the pixel points in the two-dimensional image to the three-dimensional space. Through the inverse matrix of the intrinsic parameter matrix, the pixel coordinates can be converted to the direction vector in the camera coordinate system, thus providing a basis for subsequent three-dimensional space calculations.

[0025] Subsequently, the direction vector in the camera coordinate system is converted to the BIM world coordinate system according to the formula, and a ray is generated along the direction vector in the BIM world coordinate system with the camera center C as the starting point. This ray represents the direction from the camera optical center to the pixel point on the image, extending in three-dimensional space. By combining the ray with the ground plane equation in the BIM model, the intersection parameters of the ray and the ground plane can be obtained. In order to determine the actual depth of the bottom of the pile, the present invention adopts a method for calculating the intersection of the ray and the ground plane.

[0026] Specifically, the ray-plane intersection parameters are substituted into the intersection point to obtain the three-dimensional coordinates of the construction pile in the world coordinate system. This coordinate point is calculated by the intersection of the ray and the ground plane, and it accurately represents the position of the midpoint of the bottom edge of the pile in three-dimensional space. Finally, the intersection point is remapped to the camera coordinate system to obtain a vector, and the third component of the vector is extracted as the true depth of the three-dimensional point corresponding to the coordinates of the midpoint of the bottom edge in the image relative to the camera.

[0027] Through the above steps, the pixels in the two-dimensional image are successfully mapped to the three-dimensional space, and their true spatial depth is obtained. This process not only makes full use of the camera's intrinsic and extrinsic matrix, but also uses the precise geometric information of the BIM model to make the calculation results more accurate and reliable. In addition, by performing calculations in a virtual environment, the difficulty of performing complex measurements directly in the actual environment is avoided, the measurement efficiency is improved, and the measurement cost is reduced. This technical feature provides important three-dimensional spatial information for the subsequent calculation of the volume of the pile of wood, making the entire volume measurement process of the pile of wood more efficient and accurate.

[0028] Preferably, before calling the pre-trained OccNet model to calculate the two-dimensional image, the method further includes: Get the ShapeNetCore dataset and pre-train the initial OccNet model based on the ShapeNetCore dataset to provide initial implicit field parameters; Obtain the RGB image of the wood piling scene on the construction site collected by the preset UAV and the corresponding laser scanning point cloud data, perform surface reconstruction on the corresponding laser scanning point cloud data, and generate a triangular mesh model as the ground truth of the implicit field; Use Blender software to construct synthetic scenes of various heap material geometries and materials, randomly adjust lighting, camera poses, and textures, expand the dataset, and divide the dataset into a training set, a validation set, and a test set according to a certain proportion; Continue to train the initial OccNet model based on this dataset, and verify and evaluate the trained initial OccNet model based on two metrics, namely IoU and Chamfer distance, until the model output reaches the preset value, and obtain the trained OccNet model.

[0029] The OccNet model adopts an encoder-decoder architecture. Among them, the encoder is a ResNet-18 convolutional network, which is used to extract 256-dimensional feature vectors of the input image, and the decoder consists of 5 fully connected blocks.

[0030] In this embodiment, a large amount of 3D model data is first obtained from the ShapeNetCore dataset, and these data are used to pre-train the initial OccNet model (Occupancy Networks). ShapeNetCore is a widely used 3D model dataset that contains a rich variety of object shapes and structures. By pre-training it, initial implicit field parameters can be provided for the model, thus laying a good foundation for subsequent training. The beneficial effect of this pre-training step is that it can significantly improve the model's adaptability to different shapes and structures, and reduce the computational amount and time cost in the subsequent training process.

[0031] Immediately afterwards, RGB images of the heap material scene collected by a preset drone on the construction site and the corresponding laser scan point cloud data are obtained. The laser scan point cloud data can provide high-precision 3D information. By performing surface reconstruction processing on these point cloud data, triangular mesh models are generated, and these models are used as the ground truth of the implicit field. The beneficial effect of this process is that it can provide training data highly relevant to the actual application scenario for the model, thereby improving the applicability and accuracy of the model in the actual construction scenario.

[0032] To further expand the dataset and enhance the robustness of the model, synthetic scenes with various stacking geometries and materials were constructed using Blender software. Blender is a powerful 3D modeling and animation software that can generate highly realistic 3D scenes. In these synthetic scenes, parameters such as lighting, camera pose, and texture were randomly adjusted to generate more diverse and complex training samples. Subsequently, these synthetic data were mixed with the actually collected data and divided into a training set, a validation set, and a test set according to a certain ratio (here, an 80% / 10% / 10% ratio was adopted). The beneficial effect of this data augmentation step is that it can significantly improve the model's adaptability to different environmental conditions and stacking morphologies and enhance the model's generalization performance.

[0033] During the model training process, the OccNet model with an encoder-decoder architecture was adopted. Among them, the encoder part used a ResNet-18 convolutional network, which could extract 256-dimensional feature vectors from the input two-dimensional images. These feature vectors contained the key information in the images and provided the basis for subsequent 3D reconstruction. The decoder part consisted of 5 fully connected blocks (FC-ResNet). These fully connected layers could further process the feature vectors extracted by the encoder and finally output the occupancy probability of each point in 3D space. The beneficial effect of this model architecture is that it can effectively capture the two-dimensional information in the images and convert it into accurate 3D geometric information, thus achieving high-quality 3D reconstruction.

[0034] After the model training was completed, the trained model was verified and evaluated based on two metrics. These two metrics were IoU (Intersection over Union) and Chamfer distance. The IoU metric is the intersection over union metric, which is used to measure the ratio of the overlap between the predicted mesh and the ground truth mesh to their union, while the Chamfer distance is used to measure the geometric surface error between two meshes, that is, a certain number of points are randomly sampled from each of the two meshes, and the average value of the bidirectional nearest neighbor distances is calculated to measure the geometric surface error. Through the comprehensive evaluation of these two metrics, the performance of the model can be comprehensively understood, and the model can be further optimized as needed. Finally, when the output of the model reached the preset performance metrics, a trained OccNet model was obtained. This model can not only accurately reconstruct the 3D mesh of the stacking material from the two-dimensional image but also has high robustness and generalization ability and can adapt to different construction scenes and stacking morphologies.

[0035] S3, call the pre-trained OccNet model to calculate the two-dimensional image to generate a scale-free 3D mesh of the construction stacking material; Specifically, step S3 further includes: extracting the two-dimensional image using the encoder of the OccNet model to obtain a 256-dimensional feature vector of the two-dimensional image; Normalizing the 256-dimensional feature vector using the decoder of the OccNet model to fuse image features and output the occupancy probability of any three-dimensional point (x, y, z); Using the multi-resolution isosurface extraction algorithm to recursively subdivide only active voxels in the octree structure, gradually refining to the target resolution, and applying the Marching Cubes algorithm to convert the occupied voxels into a continuous triangular mesh to obtain a scale-free three-dimensional mesh.

[0036] In this embodiment, the preprocessed two-dimensional image is input into the OccNet model. The OccNet model adopts an encoder-decoder architecture, where the encoder part is based on the ResNet-18 convolutional network and can extract features from the input two-dimensional image. Through multi-layer convolutional operations, the encoder compresses the complex information in the image into a 256-dimensional feature vector. This feature vector contains the key information about the shape, texture, and structure of the stacked materials in the image, providing a basis for subsequent 3D reconstruction. This efficient feature extraction method can significantly reduce the computational amount while retaining the important information in the image, providing a solid foundation for subsequent 3D reconstruction.

[0037] Next, the decoder part of the OccNet model further processes the extracted 256-dimensional feature vector. The decoder consists of 5 fully connected blocks, and each layer fuses the image features into the three-dimensional space through conditional batch normalization (CBN) technology. In this way, the decoder can output the occupancy probability of any three-dimensional point (x, y, z), that is, whether the point belongs to the interior or surface of the stacked materials. This probability-based output method enables the model to flexibly handle complex three-dimensional shapes and adapt to input images with different perspectives and resolutions.

[0038] To convert these occupancy probabilities into an actual three-dimensional mesh, the multi-resolution isosurface extraction algorithm (MISE) is adopted. This algorithm recursively subdivides only the "active" voxels in the octree structure and gradually refines to the target resolution. In this way, the model can efficiently generate a high-resolution three-dimensional mesh while avoiding unnecessary calculations. Finally, the Marching Cubes algorithm is used to convert the occupied voxels into a continuous triangular mesh. The Marching Cubes algorithm is a classic three-dimensional mesh generation algorithm that can efficiently convert voxel data into a smooth triangular mesh, thereby obtaining a scale-free three-dimensional mesh of the construction stacked materials.

[0039] Please refer to Figure 3 、 Figure 4, S4, use a visual programming tool to extract and process the scale-free three-dimensional grid, obtain the pixel coordinates of the four endpoints of the bottom surface of the three-dimensional minimum oriented bounding box, and calculate the Euclidean distance between the pixel coordinates and the contact point to obtain the final projection width; Specifically, step S4 further includes: importing the scale-free three-dimensional grid into the Dynamo visual programming tool, and reading the grid through the Open3D data processing library at the CPython node of the Dynamo visual programming tool; In the Open3D environment, perform geometric estimation based on PCA (Principal Component Analysis) on the vertex set of the scale-free three-dimensional grid to obtain the minimum oriented bounding box enclosing the grid and its attributes, where the attributes include the three-dimensional coordinates of the center point c of the bounding box, the rotation matrix of the minimum oriented bounding box and the size vector (l, w, h); In the local coordinate system of the minimum oriented bounding box, the bottom surface corresponds to the negative direction of the local height axis e z and its normal vector is -e z , and there are four vertices on this plane, and their coordinate representations in the local coordinate system are: , the local coordinates of the four endpoints of the bottom surface are expressed as: , is the length of the bounding box, is the width of the three-dimensional bounding box, is the height of the three-dimensional bounding box; Convert the local coordinates to the world coordinate system, and the formula is: , and convert the coordinates in the world coordinate system to the camera coordinate system, and the formula is: ; Perform pinhole projection according to the intrinsic matrix K, and its formula is: , and through normalization processing, homogenize the projection to obtain the corresponding pixel coordinates , where is the horizontal homogeneous coordinate, that is, the first row component of, is the vertical homogeneous coordinate, that is, the second row component of, is the depth component, that is, the third row component of; After projecting the four vertices of the three-dimensional minimum oriented bounding box onto the image plane respectively to obtain the corresponding pixel coordinates, calculate the Euclidean distance between it and the contact point, and its formula is: , is the pixel abscissa of the j-th bottom point of the two-dimensional bounding box on the image, is the pixel vertical coordinate of the j-th bottom point of the two-dimensional bounding box on the image; Select two endpoints closest to the bottom edge of the two-dimensional minimum oriented bounding box from the four candidate pixel points, and calculate their Euclidean distance on the pixel plane to obtain the final projected width. The formula is: , where is the final projected width, indicating the pixel scale of the width direction of the three-dimensional minimum oriented bounding box projected onto the image plane at the third component It is output through the node to the Dynamo visualization programming tool, is the abscissa of the first endpoint closest to the bottom edge of the two-dimensional minimum oriented bounding box among the four candidate pixel points of the three-dimensional bounding box, is the ordinate of the first endpoint closest to the bottom edge of the two-dimensional minimum oriented bounding box among the four candidate pixel points of the three-dimensional bounding box, is the abscissa of the second endpoint closest to the bottom edge of the two-dimensional minimum oriented bounding box among the four candidate pixel points of the three-dimensional bounding box, is the ordinate of the second endpoint closest to the bottom edge of the two-dimensional minimum oriented bounding box among the four candidate pixel points of the three-dimensional bounding box.

[0040] In this embodiment, the generated scale-free three-dimensional mesh is imported into the Dynamo visualization programming tool. Dynamo is a powerful visualization programming tool that can intuitively display the data processing process and support the extension of multiple programming languages. At the CPython node of Dynamo, the mesh is read through the Open3D data processing library. Open3D is an open-source modern 3D data processing library that provides rich geometric processing functions and can efficiently operate on three-dimensional meshes.

[0041] In the Open3D environment, geometric estimation based on principal component analysis (PCA) is performed on the mesh vertex set of the scale-free three-dimensional mesh. PCA is a commonly used statistical method that can extract important features from high-dimensional data and reduce the data dimension. Through PCA, the minimum oriented bounding box (OBB, Oriented Bounding Box) enclosing the mesh and its attributes can be obtained. These attributes include the three-dimensional coordinates of the center point of the bounding box, the rotation matrix of the minimum oriented bounding box, and the size vector. These attributes provide important basic information for subsequent geometric calculations. In addition, the Python Trimesh library can also be used to obtain the minimum volume OBB, which can achieve a similar effect to Open3D. Similarly, the minimum volume bounding box algorithm in CGAL can also be used, or the rotating calipers method can be applied to each edge of the convex hull to solve the 3D OBB; an alternative to the method for extracting the oriented bounding box of the scale-free mesh is implemented.

[0042] In the local coordinate system of the minimum oriented bounding box, the bottom face corresponds to the negative direction of the local height axis, and its normal vector is -e z Four vertices are included on this plane, and their coordinates in the local coordinate system can be calculated through simple geometric relationships. Next, the local coordinates are transformed into the world coordinate system. Subsequently, the coordinates in the world coordinate system are transformed into the camera coordinate system.

[0043] To project a three-dimensional point onto a two-dimensional image plane, a pinhole projection is performed according to the intrinsic matrix. The pinhole projection model is a commonly used projection model in computer vision, which can map points in three-dimensional space onto a two-dimensional image plane. Through normalization, the corresponding pixel coordinates are obtained.

[0044] After projecting the four vertices of the three-dimensional minimum oriented bounding box onto the image plane respectively, the corresponding pixel coordinates are obtained. Next, the Euclidean distances between these projected points and the contact point are calculated. The distances between each projected point and the contact point are calculated through formulas. From the four candidate pixel points, the two endpoints closest to the bottom edge of the two-dimensional minimum oriented bounding box are selected, and their Euclidean distances in the pixel plane are calculated to obtain the final projected width (in pixels). The projected width represents the pixel scale of the width direction of the three-dimensional minimum oriented bounding box projected onto the image plane at depth Z real and this result is output through the Dynamo visual programming tool.

[0045] The beneficial effect of this process is that it can accurately map the geometric information in three-dimensional space onto a two-dimensional image plane, and through the association with the contact point, calculate the actual projected width. This accurate mapping and calculation provide key geometric constraints for subsequent scale recovery, enabling the real size of the stacked materials to be recovered under the condition of a single two-dimensional image. Through the combined use of Dynamo and Open3D, not only the efficiency of data processing is improved, but also the scalability and flexibility of the entire system are enhanced.

[0046] S5. Under the pinhole camera model, calculate the linear scale factor according to the final projected width, calculate the real volume of the stacked materials according to the linear scale factor, and use the visual programming tool to perform visual integration on the real volume of the stacked materials.

[0047] Specifically, step S5 further includes: under the framework of the pinhole camera model, according to the geometric relationship of the image projection of the bounding box width, calculate the real width according to the final projected width , where f is the focal length; Define the linear scale factor according to the distance w between the two endpoints of the three-dimensional minimum oriented bounding box in the local coordinate system and calculate the real volume of the stacked materials in combination with the linear scale factor , where V is the volume of the generated scale-free three-dimensional grid model, which is calculated by Dynamo.

[0048] The Dynamo visual programming tool is used to map the actual volume of the stacked materials into the BIM model to achieve real-time updating of the BIM model.

[0049] In this embodiment, under the framework of the pinhole camera model, the final projection width is used to calculate the actual width of the stacked materials. The pinhole camera model is a simplified camera imaging model that assumes that light passes through the optical center of the camera and projects onto the image plane. According to the geometric relationship of the image projection of the bounding box width, the actual width can be calculated by a formula. The beneficial effect of this calculation process is that it can convert the pixel scale in the two-dimensional image into the actual physical scale, thus providing an accurate geometric basis for subsequent volume calculations.

[0050] Subsequently, according to the distance between the two endpoints of the three-dimensional minimum oriented bounding box in the local coordinate system, the linear scale factor is defined. The linear scale factor is a key parameter for converting the size of the scale-free three-dimensional grid into the actual physical size. Through this scale factor, the volume of the scale-free three-dimensional grid can be converted into the actual volume of the stacked materials. The beneficial effect of this process is that it can accurately recover the actual volume of the stacked materials, solve the scale ambiguity problem in monocular 3D reconstruction, and make it possible to directly measure the volume from a two-dimensional image.

[0051] To achieve the visual display of the actual volume of the stacked materials and the real-time updating of the BIM model, the Dynamo visual programming tool is adopted. Dynamo is a visual programming tool that is closely integrated with BIM software (such as Revit), and it can conveniently map the calculation results into the BIM model. Through Dynamo, the calculated actual volume of the stacked materials and related three-dimensional geometric information (such as position, size, etc.) are mapped into the BIM model. This process not only enables the volume information of the construction stacked materials to be visually displayed in the BIM model but also realizes the real-time updating of the BIM model, providing real-time and accurate data support for material management, cost control, and progress monitoring at the construction site.

[0052] Simply put, the category, three-dimensional coordinates, and grid model of the construction stacked materials are input into the BIM model using the Dynamo visual programming language. Then, Dynamo calculates the volume of the grid model. Combining the scale factor, the actual volume of the stacked materials can finally be calculated. Mapping these actual data of the construction stacked materials into BIM realizes the real-time updating of the BIM model.

[0053] The beneficial effects of this process are that it not only improves the accuracy and efficiency of the measurement of the volume of stacked materials, but also enhances the digital management level of the construction site through integration with the BIM model. Through the Dynamo tool, complex calculation results can be presented to construction management personnel in an intuitive way, enabling them to understand the material usage situation at the construction site in real time, so as to make more scientific and reasonable decisions. In addition, this real-time updated BIM model can also support the dynamic monitoring of the construction progress, further enhancing the digital management level of the construction site.

[0054] Briefly speaking, the method for identifying and generating the volume of stacked materials by fusing monocular Figure 3 D reconstruction and BIM calibration aims at the problem of scale ambiguity existing in monocular Figure 3 D reconstruction. It introduces the BIM virtual camera environment to provide real depth prior, and automatically restores the real geometric dimensions of the construction stacked materials and accurately calculates the volume by combining YOLOv8-OBB target recognition and rotation detection with Open3D scale-free mesh OBB reconstruction and pinhole projection model. This method does not require multi-view or laser scanning equipment. Relying only on a fixed camera and the existing BIM model, it can achieve target detection, coordinate calculation, three-dimensional reconstruction, scale recovery and volume calibration under a single image at the construction site. The whole process is automated and has high precision and engineering practicability. It aims to achieve high-efficiency, automation and high-precision measurement of the volume of stacked materials without manual distance measurement in a complex construction environment; not relying on expensive multi-view or laser scanning equipment, and only relying on a single fixed camera, it can also complete the measurement of the volume of stacked materials and has strong adaptability to complex sites; introducing real depth prior under the condition of a single image can automatically eliminate the scale ambiguity of monocular reconstruction and realize the function of restoring real geometric dimensions.

[0055] To sum up, the method for identifying and generating the volume of stacked materials by fusing monocular Figure 3 D reconstruction and BIM calibration realizes the efficient and accurate conversion from a two-dimensional image to a real three-dimensional volume by fusing monocular Figure 3 D reconstruction technology and building information model (BIM) calibration. The core of this method is to use a low-cost monocular camera to collect images, and combined with deep learning models and advanced geometric calculation techniques, it solves the problems of insufficient accuracy, high cost and low automation degree existing in traditional measurement methods.

[0056] Specifically, this method uses the pre-trained deep learning model YOLOv8-OBB to detect and classify the stacked materials in the two-dimensional image, and quickly extracts the contact points between the stacked materials and the ground. This process not only improves the detection efficiency, but also ensures the high reliability of the detection results through confidence screening and reasonable endpoint selection. Subsequently, using the internal and external parameter calibration of the camera and the virtual camera environment in the BIM model, the true spatial depth of the midpoint of the bottom edge of the stacked material is restored through ray-plane intersection calculation. This step makes full use of the accurate geometric information of the BIM model and provides a key depth prior for subsequent 3D reconstruction.

[0057] It should be noted that in addition to the BIM ray-plane intersection, that is, using the reference ground in the BIM and combining ray intersection to calculate the depth, the scale factor can also be calculated based on known size markers. By placing ArUco / AprilTag markers on-site and directly calculating the scale factor using the ratio of the true size to the projected size of the markers. It is also possible to use known object sizes, such as standardized components like safety fences and construction machinery, and combine their CAD sizes with the pixel projection widths to invert the scale factor and calibrate the grid. To achieve an alternative to the scale factor recovery and calibration method.

[0058] In the 3D reconstruction stage, a monocular Figure 3 D reconstruction model based on OccNet is adopted. Through the encoder-decoder architecture and the multi-resolution isosurface extraction algorithm, a scale-free 3D mesh is generated from the two-dimensional image. This process not only retains the key information in the image, but also significantly improves the speed and accuracy of the reconstruction through efficient algorithm design. In addition, through Open3D and PCA techniques, the minimum oriented bounding box (OBB) of the 3D mesh is further extracted, and its projected width on the image plane is calculated. The calculation of this projected width provides an important geometric constraint for scale recovery.

[0059] Finally, under the pinhole camera model, this method calculates the linear scale factor through the geometric relationship between the projected width and the true width, and accordingly restores the true volume of the stacked material. This process not only solves the scale ambiguity problem in monocular reconstruction, but also realizes the visual display of the measurement results and the dynamic update of the BIM model through real-time integration with the BIM model. This real-time updated BIM model not only provides accurate data support for material management at the construction site, but also lays a foundation for the dynamic monitoring of the construction progress and the application of digital twin technology.

[0060] Compared with the existing technology, this method Figure 3The organic combination of 3D reconstruction and BIM virtual camera scale calibration has brought significant technological improvements and application values in the field of measuring the volume of stacked materials at construction sites, which are mainly reflected in the following aspects: 1. Only one fixed monocular camera perspective photogrammetry device is required, greatly reducing equipment procurement and maintenance costs; 2. The YOLOv8-OBB model is used to complete the real-time detection of stacked material targets on a single image, and the Open3D is combined to automatically extract the unscaled oriented bounding box. The entire detection-reconstruction-calibration process can be executed with one key through scripts and Dynamo visual programming, greatly improving the measurement efficiency and reducing manual intervention; 3. With the help of the internal and external parameters of the virtual camera configured in the BIM environment, the centimeter-level true depth prior is obtained through the ray-plane intersection algorithm, and the linear scale factor is inverted in combination with the pinhole projection model, completely solving the scale ambiguity problem of monocular 3D reconstruction and realizing the accurate reconstruction of real physical dimensions and volumes; 4. With the help of the rotation detection ability of YOLOv8-OBB, it has strong robustness to inclined and partially occluded targets. The Open3D PCA-based oriented bounding box algorithm can stably extract geometric features under the irregular grid shape, enabling the method to maintain high-precision measurement in complex construction environments; 5. The real camera parameters can be directly loaded in Revit through Dynamo or API, automatically mapping the image pixel coordinates to the BIM three-dimensional space, realizing the end-to-end SOP (Standard Operating Procedure) operation process, which is convenient for popularization to daily project management; 6. It can not only be used for static stacked material volume measurement, but also be combined with an automatic progress monitoring scheme based on BIM geometric constraints and point cloud matching to realize the digital twin and progress comparison of the construction dynamic scene, and has the potential for more lightweight and real-time engineering applications.

[0061] The above is the preferred implementation mode of the present invention. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A method for identifying and generating the volume of stacked materials by integrating single-view 3D reconstruction and BIM calibration, characterized in that, Including: Obtain the two-dimensional image collected by the preset camera, and use the preset multi-task detection model to perform rotation target detection and stacking material category classification tasks on the two-dimensional image to obtain the contact point between the stacking material and the ground; Obtain the internal parameter matrix and external parameter matrix of the preset camera. In the visual programming tool, reproduce the configuration of the internal parameter matrix and external parameter matrix, construct a ray for the midpoint of the bottom edge and perform ray-plane intersection calculation with the BIM ground plane equation to obtain the real-space depth corresponding to the midpoint of the bottom edge in the image; Call the pre-trained OccNet model to calculate the two-dimensional image and generate a scale-free three-dimensional mesh of the construction stacking material; Use the visual programming tool to extract the scale-free three-dimensional mesh to obtain the pixel coordinates of the four endpoints of the bottom surface of the three-dimensional minimum oriented bounding box, and calculate the Euclidean distance between the pixel coordinates and the contact point to obtain the final projection width; Under the pinhole camera model, calculate the linear scale factor according to the final projection width, calculate the real volume of the stacking material according to the linear scale factor, and use the visual programming tool to perform visual integration on the real volume of the stacking material; Before calling the pre-trained OccNet model to calculate the two-dimensional image, it also includes: Obtain the ShapeNetCore dataset, and pre-train the initial OccNet model according to the ShapeNetCore dataset to provide initial implicit field parameters; Obtain the RGB image of the stacking material scene on the construction site collected by the preset unmanned aerial vehicle and the corresponding laser scanning point cloud data, and perform surface reconstruction processing on the corresponding laser scanning point cloud data to generate a triangular mesh model as the ground truth of the implicit field; Use Blender software to construct a synthetic scene of various stacking material geometries and materials, randomly adjust the lighting, camera pose and texture, expand the dataset, and divide the dataset into a training set, a validation set, and a test set according to a ratio; Continue to train the initial OccNet model according to this dataset, and verify and evaluate the trained initial OccNet model based on two metrics, IoU and Chamfer distance, until the model output reaches the preset value to obtain the trained OccNet model.

2. The method for identifying and generating the piled material volume by integrating single-view 3D reconstruction and BIM calibration according to claim 1, wherein Obtain the two-dimensional image collected by the preset camera, and use the preset multi-task detection model to perform rotation target detection and stacking material category classification tasks on the two-dimensional image to obtain the contact point between the stacking material and the ground. Specifically: Obtain the two-dimensional image of the construction stacking material collected by the camera configured at the construction site, and preprocess the two-dimensional image; Input the preprocessed two-dimensional image into the multi-task detection model constructed based on YOLOv8-OBB to perform rotation target detection and stacking material category classification tasks, and output information results. Among them, the information results include the four vertex pixel coordinates, confidence level, and category label of the rotation bounding box of each stacking material target; Filter out the information results with confidence values lower than the preset value. Arrange the pixel coordinates of the four vertices in descending order according to the ordinate of the pixel coordinates, and select the first two vertex pixel coordinates in the sequence as the endpoints of the bottom edge of the stacked material, marked as (u1, v1) and (u2, v2), corresponding to the contact points of the stacked material with the ground.

3. The method for identifying and generating the stacking volume by fusing single-view 3D reconstruction and BIM calibration according to claim 2, characterized in that Obtain the internal parameter matrix and external parameter matrix of the preset camera, specifically: Calibrate the preset camera to obtain its internal parameter matrix K and external parameter matrix [R∣T]; Among them, the intrinsic parameter matrix K is used to represent the internal optical characteristics of the camera, including the focal length, the coordinates of the principal point, and the distortion coefficients. Its matrix form is: , is the focal length of the camera in the x direction, is the focal length of the camera in the y direction, is the abscissa of the principal point of the camera, is the ordinate of the principal point of the camera; The external parameter matrix [R∣T] is used to represent the position and orientation of the camera in the world coordinate system. It consists of the rotation matrix R and the translation vector t, and its matrix form is: , is the translation component of the camera optical center on the x-axis of the world coordinate system, is the translation component of the camera optical center on the y-axis of the world coordinate system, is the translation component of the camera optical center on the z-axis of the world coordinate system, is the projection of the x-axis of the camera coordinate system in the x-axis direction of the world coordinate system, is the projection of the y-axis of the camera coordinate system in the x-axis direction of the world coordinate system, is the projection of the z-axis of the camera coordinate system in the x-axis direction of the world coordinate system, is the projection of the x-axis of the camera coordinate system in the y-axis direction of the world coordinate system, is the projection of the y-axis of the camera coordinate system in the y-axis direction of the world coordinate system, is the projection of the z-axis of the camera coordinate system in the y-axis direction of the world coordinate system, is the projection of the x-axis of the camera coordinate system in the z-axis direction of the world coordinate system, is the projection of the y-axis of the camera coordinate system in the z-axis direction of the world coordinate system, is the projection of the z-axis of the camera coordinate system in the z-axis direction of the world coordinate system.

4. The method for identifying and generating the stacking volume by fusing single-view 3D reconstruction and BIM calibration according to claim 3, wherein Obtain the internal parameter matrix and external parameter matrix of the preset camera. In the visual programming tool, reproduce the configuration of the internal parameter matrix and external parameter matrix, construct a ray for the midpoint of the bottom edge and perform ray-plane intersection calculation with the BIM ground plane equation to obtain the real-space depth corresponding to the midpoint of the bottom edge in the image, specifically: Combine the Revit software and the Dynamo visual programming tool to create a new perspective view, and set the viewpoint position and orientation of this perspective view according to the internal parameter matrix and external parameter matrix of the preset camera to ensure that the virtual viewpoint coincides exactly with the preset camera in the three-dimensional space; Calculate the midpoint coordinates (u m , v m ) of the bottom edge endpoints of the stacked materials based on the contact points. The formula is: , and use the inverse matrix K -1 of the internal parameter matrix K to convert the midpoint coordinates (u m , v m ) into the direction vector in the camera coordinate system. The formula is: , is the direction vector from the camera optical center to the current pixel point on the image, is the transpose operation; According to the formula Convert the direction vector in the camera coordinate system to the BIM world coordinate system, and in the BIM world coordinate system, starting from the camera center C, along the direction vector generate a ray , to achieve the back-projection of pixel points into three-dimensional space, where is the transpose of the camera rotation matrix; Ground plane equations in BIM models , and rays Perform the combined operation to obtain the ray-plane intersection parameters ,in, is the plane normal vector, is the offset of the plane relative to the origin of the world coordinate system, is any point in three-dimensional space; Substitute the ray-plane intersection parameter into the intersection point P hit . Mark it as the three-dimensional coordinates of the construction stockpile in the world coordinate system. The formula is: , and remap the intersection point P hit to the camera coordinate system to obtain the vector . Extract the third component of this vector as the true depth of the three-dimensional point corresponding to the midpoint coordinates (u m , v m ) of the bottom edge in the image relative to the camera. is the first component, is the second component, represents the ray-plane intersection point.

5. The method for identifying and generating the stacking volume by fusing single-view 3D reconstruction and BIM calibration according to claim 1, wherein The OccNet model adopts an encoder-decoder architecture. Among them, the encoder is a ResNet-18 convolutional network, which is used to extract a 256-dimensional feature vector of the input image, and the decoder consists of 5 fully connected blocks.

6. The method for identifying and generating the stacking volume by fusing single-view 3D reconstruction and BIM calibration according to claim 5, wherein Call the pre-trained OccNet model to calculate the two-dimensional image and generate a scale-free three-dimensional mesh of the construction stacked material, specifically: Use the encoder of the OccNet model to extract the two-dimensional image to obtain a 256-dimensional feature vector of the two-dimensional image; Use the decoder of the OccNet model to normalize the 256-dimensional feature vector to fuse the image features and output the occupancy probability of any three-dimensional point (x, y, z); Use the multi-resolution voxel extraction algorithm to recursively subdivide only active voxels in the octree structure, gradually refine to the target resolution, and use the Marching Cubes algorithm to convert the occupied voxels into a continuous triangular mesh to obtain a scale-free three-dimensional mesh.

7. The method for identifying and generating the volume of stacked materials by fusing single-view 3D reconstruction and BIM calibration according to claim 3, characterized in that Use the visual programming tool to extract and process the scale-free three-dimensional mesh to obtain the pixel coordinates of the four endpoints of the bottom surface of the three-dimensional minimum oriented bounding box, and calculate the Euclidean distance between this pixel coordinate and the contact point to obtain the final projection width, specifically: Import the scale-free three-dimensional mesh into the Dynamo visual programming tool, and read the mesh through the Open3D data processing library at the CPython node of the Dynamo visual programming tool; In the Open3D environment, perform PCA-based geometric estimation on the mesh vertex set of the scale-free 3D mesh to obtain the minimum oriented bounding box enclosing the mesh and its attributes, where the attributes include the three-dimensional coordinates of the center point c of the bounding box and the rotation matrix of the minimum oriented bounding box and the size vector (l, w, h); In the local coordinate system of the minimum oriented bounding box, the bottom surface corresponds to the negative direction of the local height axis e z and its normal vector is -e z . There are four vertices on this plane, and their coordinate representations in the local coordinate system are: . The local coordinate representations of the four endpoints of the bottom surface are: , is the length of the bounding box, is the width of the three-dimensional bounding box, is the height of the three-dimensional bounding box; Convert the local coordinates to the world coordinate system. The formula is: , and convert the coordinates in the world coordinate system to the camera coordinate system. The formula is: ; Perform pinhole projection according to the internal parameter matrix K, and its formula is: , through normalization processing, homogenize the projection to obtain the corresponding pixel coordinates , where is the horizontal homogeneous coordinate, is the vertical homogeneous coordinate, is the depth component; After projecting the four vertices of the three-dimensional minimum oriented bounding box onto the image plane respectively to obtain the corresponding pixel coordinates, calculate the Euclidean distance between them and the contact point. The formula is as follows: , is the pixel abscissa of the j-th bottom point of the two-dimensional bounding box on the image, is the pixel ordinate of the j-th bottom point of the two-dimensional bounding box on the image; Select two endpoints closest to the bottom edge of the two-dimensional minimum oriented bounding box from the four candidate pixel points, and calculate their Euclidean distance in the pixel plane to obtain the final projected width. The formula is as follows: , where is the final projected width, representing the pixel scale of the width direction of the three-dimensional minimum oriented bounding box projected onto the image plane at the third component . It is output to the Dynamo visual programming tool through the node, is the abscissa of the first endpoint closest to the bottom edge of the two-dimensional minimum oriented bounding box among the four candidate pixel points of the three-dimensional bounding box, is the ordinate of the first endpoint closest to the bottom edge of the two-dimensional minimum oriented bounding box among the four candidate pixel points of the three-dimensional bounding box, is the abscissa of the second endpoint closest to the bottom edge of the two-dimensional minimum oriented bounding box among the four candidate pixel points of the three-dimensional bounding box, is the ordinate of the second endpoint closest to the bottom edge of the two-dimensional minimum oriented bounding box among the four candidate pixel points of the three-dimensional bounding box.

8. The method for identifying and generating the piled material volume by fusing single-view 3D reconstruction and BIM calibration according to claim 7, wherein Under the pinhole camera model, calculate the linear scale factor according to the final projection width, and calculate the real volume of the stacked material according to the linear scale factor, specifically: Under the framework of the pinhole camera model, based on the image projection geometric relationship of the bounding box width, calculate the real width according to the final projected width , where f is the focal length; Define a linear scale factor based on the distance w between the two endpoints of the three-dimensional minimum oriented bounding box in the local coordinate system and calculate the true volume of the stacked materials in combination with the linear scale factor , where is the volume of the generated dimensionless three-dimensional mesh model, which is calculated by Dynamo.

9. The method for identifying and generating the volume of stacked materials by integrating single-view 3D reconstruction and BIM calibration according to claim 1, characterized in that, Use the visual programming tool to perform visual integration on the real volume of the stacked material, specifically: use the Dynamo visual programming tool to map the real volume of the stacked material to the BIM model to realize the real-time update of the BIM model.

Citation Information

Patent Citations

  • Directional calibration target for camera inner and outer parameter calibration

    CN104867160A

  • Three-dimensional reconstruction method for vehicle target in road scene based on monocular vision

    CN113129348A

  • Three-dimensional measurement method based on single aerial picture of unmanned aerial vehicle

    CN116309844A

  • BIM three-dimensional reconstruction method for steel structure factory building

    CN116597108A

  • Target rapid three-dimensional modeling system based on unmanned aerial vehicle image

    CN118736128A

Cited By

  • Adaptive merging method and system based on multiple single BIM models

    CN120747435A