A method for identifying and generating stacked material volumes by integrating single-view 3D reconstruction and BIM calibration

By integrating single-view 3D reconstruction and BIM calibration technology, and utilizing camera images and deep learning models, the true volume of the piled materials is automatically calculated, solving the accuracy and automation issues in traditional measurement methods, and achieving efficient and accurate volume measurement of piled materials and real-time updating of BIM models.

CN120374885BActive Publication Date: 2025-09-09XIAMEN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510827984.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-09
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing wood volume measurement technology has limitations in accuracy, cost and degree of automation, and cannot meet the needs of modern engineering construction for efficient, accurate and automated measurement.

Method used

A method for identifying and generating the volume of piled materials that integrates single-view 3D reconstruction and BIM calibration acquires camera images, uses a multi-task detection model for target detection and classification, and combines the camera's internal and external parameter matrices with the BIM model to calculate the true depth and volume of the piled materials, achieving automated 3D reconstruction and volume calibration.

Benefits of technology

It achieves fast and accurate measurement of the volume of piled materials at construction sites, reduces the cost of measurement equipment, improves measurement efficiency and accuracy, and supports real-time updating and data visualization of BIM models, making it suitable for material management and cost control at construction sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374885B_ABST
    Figure CN120374885B_ABST
Patent Text Reader

Abstract

The present invention provides a method for identifying and generating the volume of piled materials by integrating single-view 3D reconstruction and BIM calibration, which relates to the technical field of volume generation of piled materials. The method uses a low-cost monocular camera to capture images, and with the help of advanced deep learning models and three-dimensional reconstruction algorithms, generates a three-dimensional model of the piled materials. By performing scale calibration with the virtual camera in the BIM model, the scale uncertainty problem that is difficult to overcome in traditional single-view reconstruction is solved, so that the true volume of the piled materials can be accurately calculated. In addition, the method also greatly improves measurement efficiency and reduces manual intervention through automated target detection, three-dimensional coordinate calculation and volume calibration processes. It is not only suitable for volume measurement of static piled materials, but can also be extended to real-time monitoring of dynamic construction scenes, providing strong technical support for construction management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of material volume generation, in particular to a fusion single-view Figure 3 A method for identifying and generating piled material volumes based on 3D reconstruction and BIM calibration. Background Art

[0002] In the fields of construction and project management, accurately measuring the volume of materials on construction sites is a critical task. It not only affects the rational use of materials and cost control, but also directly impacts construction progress and safety management. However, current technologies for measuring material volume have numerous shortcomings, making them unable to meet the demands of modern engineering construction for efficient, accurate, and automated measurement.

[0003] Traditional methods for measuring pile volume rely primarily on manual labor, such as using a tape measure or measuring wheel to measure the base size and height of the pile and then estimating the volume. This method is not only inefficient, but the results are also easily affected by operator experience, technical skills, and the on-site environment, resulting in unstable measurement accuracy and difficulty ensuring accurate and consistent results.

[0004] With technological advancements, multi-view photogrammetry and LiDAR (Light Detection and Ranging) scanning technologies have been introduced to the field of wood pile volume measurement. These technologies provide relatively accurate 3D point cloud data, enabling high-precision volume calculations. However, their application also faces numerous challenges. First, these technologies require strict environmental conditions, such as good lighting, stable weather conditions, and precise ground control points. If these conditions are not met, the accuracy and reliability of the measurement results will be significantly compromised. Furthermore, the entire process, from data collection to point cloud processing, 3D modeling, and final volume calculation, is complex and time-consuming, making it difficult to meet the demand for rapid measurement and providing timely and accurate decision-making for construction management.

[0005] In recent years, single vision Figure 3 As an emerging technical means, 3D reconstruction technology has gradually attracted attention. This technology uses a single color image (RGB image) to predict the 3D shape, direction and size of an object, with the advantages of low cost and easy deployment. However, due to the lack of depth information in monocular images, the reconstruction results often have scale uncertainty, that is, the true physical size and volume of the object cannot be directly obtained. This means that although it can be obtained through single-view Figure 3While 3D reconstruction techniques rapidly generate 3D models of objects, the scale of the models is only relative and cannot be directly applied to actual volume measurements. Furthermore, the model's reusability is poor under varying camera parameters, shooting distances, and angles, making standardized and automated scale restoration difficult to achieve, limiting its widespread application in practical engineering applications.

[0006] At the same time, Building Information Modeling (BIM) technology has been widely adopted in the construction industry. BIM provides precise spatial geometry and in-depth prior knowledge, providing strong support for the full lifecycle management of engineering projects. However, there is currently no mature technical solution that effectively combines BIM with monocular 3D reconstruction technology to resolve scale ambiguity in monocular reconstruction and achieve accurate measurement of stockpiled material volumes.

[0007] In general, existing pile volume measurement technologies have varying degrees of limitations in terms of accuracy, cost, and automation, making it difficult to meet the demands of modern engineering construction for efficient, accurate, and automated measurement. In light of this, this application is filed. Summary of the Invention

[0008] The present invention provides a fusion single-view Figure 3 The method of identifying and generating the volume of piled materials based on 3D reconstruction and BIM calibration can at least partially improve the above problems.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] A fusion monocular Figure 3 The method for identifying and generating the volume of piled materials based on 3D reconstruction and BIM calibration includes:

[0011] Obtain a two-dimensional image captured by a preset camera, use a preset multi-task detection model to perform rotating target detection and pile material classification on the two-dimensional image, and obtain the contact points between the pile and the ground;

[0012] Obtain the intrinsic and extrinsic matrix of the preset camera, reproduce the configuration of the intrinsic and extrinsic matrix in a visual programming tool, perform ray-plane intersection calculation on the ray constructed at the midpoint of the base and the BIM ground plane equation, and obtain the real space depth corresponding to the midpoint of the base in the image;

[0013] Calling a pre-trained OccNet model to perform calculations on the two-dimensional image to generate a scale-free three-dimensional grid of the construction material pile;

[0014] Use a visual programming tool to extract and process the scale-free 3D grid to obtain the pixel coordinates of the four endpoints of the bottom surface of the 3D minimum oriented bounding box, and calculate the Euclidean distance between the pixel coordinates and the contact point to obtain the final projection width;

[0015] Under the pinhole camera model, the linear scale factor is calculated according to the final projection width, the real volume of the pile is calculated according to the linear scale factor, and the real volume of the pile is visualized and integrated using visual programming tools.

[0016] In summary, the fusion single view Figure 3 The method of identifying and generating the volume of piled materials by combining single-view Figure 3 The 3D reconstruction technology and Building Information Model (BIM) calibration function enable fast and accurate volume measurement of construction site materials. It can automatically complete target detection, coordinate calculation, 3D reconstruction, scale recovery, and volume calibration within a single construction site image. It then maps the actual data of these construction materials into the BIM, enabling real-time updates of the BIM model.

[0017] Specifically, this method uses images captured by a monocular camera, with the help of advanced image processing and deep learning algorithms, to extract the geometric features of the piled materials and generate a three-dimensional model. By calibrating with the virtual camera in the BIM model and introducing real depth information, the scale uncertainty problem existing in traditional single-view reconstruction technology is solved, so that the true volume of the piled materials can be accurately calculated. This method not only significantly reduces the cost of measuring equipment, but also reduces manual intervention through automated processes, thereby improving measurement efficiency and accuracy. In addition, this technology can be seamlessly integrated with existing BIM systems to achieve real-time data updates and visual display, providing strong support for material management, cost control and safety monitoring at the construction site, and has significant engineering application value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 The embodiment of the present invention provides a fusion monocular Figure 3 Flowchart of the method for identifying and generating the volume of piled materials based on 3D reconstruction and BIM calibration;

[0019] Figure 2 is a schematic diagram of ray-plane intersection provided by an embodiment of the present invention;

[0020] Figure 3 Schematic diagram of the OBB local coordinate system provided by an embodiment of the present invention;

[0021] Figure 4 Schematic diagram of pinhole projection provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0023] refer to Figure 1 As shown, the first embodiment of the present invention discloses a fusion single-view Figure 3 The method of identifying and generating the volume of piled materials based on 3D reconstruction and BIM calibration can be achieved by fusing single-view Figure 3 The method is executed by a material volume recognition and generation device for 3D reconstruction and BIM calibration (hereinafter referred to as the recognition and generation device), and in particular, is executed by one or more processors in the recognition and generation device to implement the following method:

[0024] S1: Obtain a two-dimensional image captured by a preset camera, use a preset multi-task detection model to perform rotating target detection and pile material classification on the two-dimensional image, and obtain the contact points between the pile material and the ground;

[0025] Specifically, step S1 further includes: acquiring a two-dimensional image of the construction materials piled up, captured by a camera configured at the construction site, and preprocessing the two-dimensional image;

[0026] The preprocessed 2D image is fed into a multi-task detection model based on YOLOv8-OBB to perform rotation target detection and wood pile classification tasks, and output information results, where the information results include the pixel coordinates of the four vertices of the rotation bounding box of each wood pile target, the confidence score, and the category label;

[0027] The information results with confidence values ​​lower than the preset value are filtered out, and the four vertex pixel coordinates are arranged in descending order according to the vertical coordinates of the pixel coordinates. The first two vertex pixel coordinates in the sequence are selected as the endpoints of the bottom edge of the pile, marked as (u1, v1) and (u2, v2) to correspond to the contact points between the pile and the ground.

[0028] In this embodiment, a camera is pre-positioned at the construction site, carefully positioned and angled to ensure a clear view of the construction material pile. When volume measurement is required, a two-dimensional image of the material pile, captured by the camera, is first acquired. Because image acquisition can be affected by factors such as lighting and noise, it requires preprocessing before inputting it into the detection model. This preprocessing involves grayscaling, binarization, and filtering to remove noise and enhance image contrast, thereby improving the accuracy of subsequent detection.

[0029] Subsequently, the preprocessed two-dimensional image is input into a multi-task detection model built based on YOLOv8-OBB. YOLOv8-OBB is an advanced target detection model that can quickly and accurately detect and classify rotating targets in images. In this embodiment, the model is specially trained to enable it to identify the category of construction piles and detect the rotated bounding box of the piles. The information output by the model includes the pixel coordinates of the four vertices of the rotated bounding box of each pile target, the confidence level, and the category label. The confidence level reflects the model's confidence in the detection result. By setting a preset value, information results with a confidence level lower than this value (here set to 0.5) can be filtered out, thereby improving the reliability of the detection results.

[0030] In addition to using the YOLOv8-OBB model, you can also use instance segmentation models such as Mask R-CNN to detect the outline of the pile of wood, and then fit the segmentation mask with the minimum rotated rectangle to obtain the rotated bounding box. This can replace the 2D image object detection and rotated bounding box extraction methods.

[0031] For detection results that pass the confidence filter, the four vertex pixel coordinates are sorted in descending order based on their vertical coordinates. Since the area where the pile of wood meets the ground is typically at the bottom of the image, the first two vertex pixel coordinates in the sequence are selected as the endpoints of the pile's bottom edge, labeled (u1, v1) and (u2, v2), respectively. These two endpoints accurately correspond to the points of contact between the wood and the ground, providing crucial foundational data for subsequent 3D coordinate calculations.

[0032] This step enables rapid and accurate detection of the location and type of piled materials within a single 2D image, as well as pinpointing the points of contact between the materials and the ground. This not only provides critical information for subsequent 3D reconstruction and volume calculation, but also, thanks to the advanced YOLOv8-OBB model, ensures high efficiency and accuracy throughout the detection process. This allows for rapid processing of large amounts of image data, meeting the real-time volume measurement requirements of construction sites. Furthermore, confidence level screening and appropriate endpoint selection further enhance the reliability and stability of detection results, laying a solid foundation for the smooth implementation of subsequent steps.

[0033] See also Figure 2 ,in, Figure 2 E represents the actual position of the camera, O represents the origin, and H represents the actual height between the camera and the ground. Represents the camera tilt angle, and d is the direction vector. S2, obtain the intrinsic parameter matrix and extrinsic parameter matrix of the preset camera, reproduce the configuration of the intrinsic parameter matrix and extrinsic parameter matrix in the visual programming tool, calculate the ray-plane intersection of the constructed ray at the midpoint of the base and the BIM ground plane equation, and obtain the real space depth corresponding to the midpoint of the base in the image;

[0034] Specifically, step S2 further includes: calibrating the preset camera to obtain its intrinsic parameter matrix K and extrinsic parameter matrix [R|T];

[0035] Among them, the intrinsic parameter matrix K is used to describe the internal optical characteristics of the camera, including focal length, principal point coordinates and distortion coefficients, and its matrix form is: , is the focal length of the camera in the x direction, is the focal length of the camera in the y direction, is the horizontal coordinate of the camera principal point, is the ordinate of the camera's principal point;

[0036] The extrinsic matrix [R|T] is used to describe the position and posture of the camera in the world coordinate system. It consists of the rotation matrix R and the translation vector t. Its matrix form is: , is the translation component of the camera’s optical center on the x-axis of the world coordinate system, is the translation component of the camera’s optical center on the y-axis of the world coordinate system, is the translation component of the camera’s optical center on the z-axis of the world coordinate system, is the projection of the camera coordinate system x-axis in the world coordinate system x-axis direction, is the projection of the camera coordinate system y-axis on the world coordinate system x-axis, is the projection of the camera coordinate system z-axis on the world coordinate system x-axis, is the projection of the camera coordinate system x-axis on the world coordinate system y-axis, is the projection of the y-axis of the camera coordinate system on the y-axis of the world coordinate system, is the projection of the camera coordinate system z-axis on the world coordinate system y-axis, is the projection of the camera coordinate system x-axis on the world coordinate system z-axis, is the projection of the camera coordinate system y-axis on the world coordinate system z-axis, It is the projection of the camera coordinate system z-axis onto the world coordinate system z-axis.

[0037] Combine Revit software and the Dynamo visual programming tool to create a new perspective view. Set the viewpoint position and orientation of the perspective view according to the intrinsic and extrinsic matrix of the preset camera to ensure that the virtual viewpoint completely overlaps with the preset camera in three-dimensional space.

[0038] According to the contact point, calculate the midpoint coordinates of the bottom end point of the pile (u m ,v m ), the formula is: , and use the inverse matrix K of the internal parameter matrix K-1 , the midpoint coordinates (u m ,v m ) is converted to the direction vector in the camera coordinate system, and its formula is: , is the direction vector from the camera optical center to the current pixel on the image, is the transpose operation;

[0039] According to the formula The direction vector in the camera coordinate system Convert to the BIM world coordinate system, and in the BIM world coordinate system, take the camera center C as the starting point and follow the direction vector Generate a ray , to realize the back projection of pixel points into three-dimensional space, where, is the transpose of the camera rotation matrix;

[0040] Ground plane equations in BIM models , and rays Perform the combined equation to obtain the ray-plane intersection parameters ,in, is the plane normal vector, is the offset of the plane relative to the origin of the world coordinate system, is any point in three-dimensional space;

[0041] Set the ray-plane intersection parameters Substitute the intersection point P hit In the figure, the mark is the three-dimensional coordinate of the construction material in the world coordinate system, and the formula is: and the intersection point P hit Remap to the camera coordinate system to get the vector , extract the third component of the vector As the coordinates of the midpoint of the bottom edge of the image (u m ,v m ) corresponds to the true depth of the 3D point relative to the camera, is the first component, is the second component, Represents a ray-plane intersection point.

[0042] In this embodiment, the preset camera is first calibrated, which is the foundation of the entire process. Through calibration, the camera's intrinsic parameter matrix and extrinsic parameter matrix can be obtained. The intrinsic parameter matrix describes the camera's internal optical properties, including focal length, principal point coordinates, and distortion coefficients. The extrinsic parameter matrix describes the camera's position and posture in the world coordinate system and is composed of a rotation matrix and a translation vector.

[0043] After obtaining the camera's intrinsic and extrinsic matrix, create a new perspective view using Revit software and Dynamo visual programming tools. Set the viewpoint position and orientation of the perspective view according to the preset camera's intrinsic and extrinsic matrix to ensure that the virtual viewpoint completely overlaps with the preset camera in three-dimensional space. This process provides an accurate virtual environment for subsequent three-dimensional space calculations, so that the calculations performed in the virtual environment can correspond one-to-one with the situation in the actual environment. In the camera properties panel of this view, set the focal length f and the principal point coordinates (u0, v0) to the f corresponding to the known intrinsic matrix K. x ,f y ,u0,v0, ensure that the virtual projection has no deviation from the physical lens.

[0044] Next, based on the pixel coordinates (u1, v1) and (u2, v2) of the contact point between the pile and the ground obtained in the previous step, calculate the midpoint coordinates of the bottom edge of the pile (u m ,v m ). Use the inverse matrix of the internal parameter matrix to convert the midpoint coordinates (u m ,v m ) to the direction vector in the camera coordinate system. This conversion process is a key step in mapping the pixels in the two-dimensional image to the three-dimensional space. Through the inverse matrix of the intrinsic parameter matrix, the pixel coordinates can be converted to the direction vector in the camera coordinate system, thus providing the basis for subsequent three-dimensional space calculations.

[0045] The direction vector in the camera coordinate system is then converted to the BIM world coordinate system according to a formula. In the BIM world coordinate system, a ray is generated along the direction vector, starting from the camera center C. This ray represents the direction from the camera's optical center to the pixel in the image, extending in three-dimensional space. By combining the ray with the ground plane equation in the BIM model, the parameters of the intersection point between the ray and the ground plane can be obtained. To determine the actual depth of the piled material bottom, the present invention uses a method to calculate the intersection point of the ray and the ground plane.

[0046] Specifically, the ray-plane intersection parameters are substituted into the intersection point to obtain the 3D coordinates of the construction pile in the world coordinate system. This coordinate point, calculated from the intersection of the ray and the ground plane, accurately represents the position of the midpoint of the bottom edge of the pile in 3D space. Finally, the intersection point is remapped to the camera coordinate system to obtain a vector. The third component of this vector is extracted as the true depth of the 3D point corresponding to the coordinates of the bottom edge midpoint in the image relative to the camera.

[0047] Through the above steps, the pixels in the two-dimensional image were successfully mapped into three-dimensional space, and their true spatial depth was obtained. This process not only fully utilizes the camera's intrinsic and extrinsic parameter matrices, but also draws on the precise geometric information of the BIM model, resulting in highly accurate and reliable calculation results. Furthermore, performing calculations in a virtual environment avoids the difficulties of performing complex measurements directly in the real environment, improving measurement efficiency and reducing measurement costs. This technical feature provides important three-dimensional spatial information for subsequent wood volume calculations, making the entire wood volume measurement process more efficient and accurate.

[0048] Preferably, before calling the pre-trained OccNet model to calculate the two-dimensional image, the method further includes:

[0049] Obtain the ShapeNetCore dataset and pre-train the initial OccNet model based on the ShapeNetCore dataset to provide initial implicit field parameters;

[0050] Obtain the RGB image of the wood piling scene on the construction site and the corresponding laser scanning point cloud data collected by the preset drone, and perform surface reconstruction on the corresponding laser scanning point cloud data to generate a triangular mesh model as the ground truth value of the implicit field;

[0051] Use Blender software to build synthetic scenes with various stacking geometry and materials, randomly adjust lighting, camera pose, and texture, expand the dataset, and divide the dataset into training, validation, and test sets in proportion.

[0052] The initial OccNet model is trained based on the data set, and the trained initial OccNet model is verified and evaluated based on the two indicators of IoU and Chamfer distance until the model output reaches the preset value, and the trained OccNet model is obtained.

[0053] The OccNet model adopts an encoder-decoder architecture, where the encoder is a ResNet-18 convolutional network used to extract the 256-dimensional feature vector of the input image, and the decoder consists of 5 layers of fully connected blocks.

[0054] In this example, a large amount of 3D model data was first obtained from the ShapeNetCore dataset and used to pre-train the initial OccNet model (Occupancy Networks). ShapeNetCore is a widely used 3D model dataset that contains a rich variety of object shapes and structures. Pre-training on this dataset provides the model with initial implicit field parameters, laying a solid foundation for subsequent training. This pre-training step significantly improves the model's adaptability to diverse shapes and structures, reducing the computational effort and time required for subsequent training.

[0055] Next, the system acquires RGB images of the wood piling scene captured by a pre-set drone on the construction site, along with the corresponding laser scanning point cloud data. Laser scanning point cloud data provides highly accurate 3D information. Surface reconstruction of this point cloud data generates triangular mesh models, which serve as the ground truth for the implicit field. This process is beneficial in providing the model with training data that is highly relevant to real-world application scenarios, thereby improving the model's applicability and accuracy in real-world construction scenarios.

[0056] To further expand the dataset and enhance the robustness of the model, synthetic scenes with various wood pile geometries and materials were constructed using Blender. Blender is a powerful 3D modeling and animation software capable of generating highly realistic 3D scenes. Within these synthetic scenes, parameters such as lighting, camera pose, and texture were randomly adjusted to generate more diverse and complex training samples. This synthetic data was then mixed with real-world data and divided into training, validation, and test sets in a ratio of 80% / 10% / 10% (here). This data augmentation step significantly improves the model's adaptability to diverse environmental conditions and wood pile configurations, enhancing its generalization performance.

[0057] During model training, the OccNet model with an encoder-decoder architecture was employed. The encoder utilizes a ResNet-18 convolutional network, which extracts 256-dimensional feature vectors from the input 2D image. These feature vectors contain key information from the image and provide the basis for subsequent 3D reconstruction. The decoder consists of five fully connected layers (FC-ResNet). These layers further process the feature vectors extracted by the encoder, ultimately outputting the occupancy probability of each point in 3D space. The beneficial effect of this model architecture is that it effectively captures the 2D information in the image and converts it into accurate 3D geometric information, thereby achieving high-quality 3D reconstruction.

[0058] After model training was complete, we validated and evaluated the trained model using two metrics: Intersection over Union (IoU) and Chamfer distance. The IoU metric measures the ratio of the overlap between the predicted and true meshes to their union, while the Chamfer distance measures the geometric surface error between the two meshes. This error is measured by randomly sampling a certain number of points from each mesh and calculating the average of the two-way nearest neighbor distances. This combined evaluation of these two metrics provides a comprehensive understanding of the model's performance and allows for further optimization as needed. Ultimately, when the model output meets the pre-defined performance metrics, a trained OccNet model is obtained. This model not only accurately reconstructs the 3D mesh of the piled materials from 2D images but also exhibits high robustness and generalization capabilities, adapting to diverse construction scenarios and material pile configurations.

[0059] S3, calling the pre-trained OccNet model to calculate the two-dimensional image and generate a scale-free three-dimensional grid of the construction material pile;

[0060] Specifically, step S3 further includes: extracting the two-dimensional image using an encoder of the OccNet model to obtain a 256-dimensional feature vector of the two-dimensional image;

[0061] The decoder of the OccNet model is used to normalize the 256-dimensional feature vector to fuse image features and output the occupancy probability of any 3D point (x, y, z);

[0062] A multi-resolution isosurface extraction algorithm is used to recursively subdivide only active voxels in an octree structure, gradually refining them to the target resolution. The occupied voxels are converted into continuous triangular meshes using the MarchingCubes algorithm to obtain a scale-free three-dimensional mesh.

[0063] In this embodiment, a preprocessed two-dimensional image is input into the OccNet model. The OccNet model uses an encoder-decoder architecture, in which the encoder part is based on the ResNet-18 convolutional network and can extract features from the input two-dimensional image. Through multi-layer convolution operations, the encoder compresses the complex information in the image into a 256-dimensional feature vector. This feature vector contains key information about the shape, texture, and structure of the pile in the image, providing a basis for subsequent three-dimensional reconstruction. This efficient feature extraction method can significantly reduce the amount of computation while retaining important information in the image, providing a solid foundation for subsequent three-dimensional reconstruction.

[0064] Next, the decoder portion of the OccNet model further processes the extracted 256-dimensional feature vector. The decoder consists of five layers of fully connected blocks, each of which fuses image features into a three-dimensional space using conditional batch normalization (CBN). This allows the decoder to output the occupancy probability of any 3D point (x, y, z), indicating whether the point belongs inside or on the surface of the pile. This probabilistic output allows the model to flexibly handle complex 3D shapes and adapt to input images of varying viewpoints and resolutions.

[0065] To convert these occupancy probabilities into an actual 3D mesh, the Multiresolution Isosurface Extraction (MISE) algorithm was employed. This algorithm recursively subdivides only the "active" voxels in an octree structure, gradually refining them to the target resolution. In this way, the model efficiently generates a high-resolution 3D mesh while avoiding unnecessary computations. Finally, the Marching Cubes algorithm is used to convert the occupied voxels into a continuous triangular mesh. The Marching Cubes algorithm is a classic 3D mesh generation algorithm that efficiently converts voxel data into a smooth triangular mesh, resulting in a scale-free 3D mesh of the construction material.

[0066] See also Figure 3 、 Figure 4 ,S4, use visual programming tools to extract and process the scale-free three-dimensional grid, obtain the pixel coordinates of the four endpoints of the bottom surface of the three-dimensional minimum oriented bounding box, and calculate the Euclidean distance between the pixel coordinates and the contact point to obtain the final projection width;

[0067] Specifically, step S4 further includes: importing the scale-free three-dimensional grid into the Dynamo visual programming tool, and reading the grid through the Open3D data processing library at the CPython node of the Dynamo visual programming tool;

[0068] In the Open3D environment, the geometry of the mesh vertex set of the scale-free 3D mesh is estimated based on PCA (Principal Component Analysis) to obtain the minimum oriented bounding box surrounding the mesh and its attributes, where the attributes include the 3D coordinates of the center point c of the bounding box and the rotation matrix of the minimum oriented bounding box. and the size vector (l, w, h);

[0069] In the local coordinate system of the minimum oriented bounding box, the bottom plane corresponds to the local height axis e z The negative direction of the normal vector is -e z , the plane contains four vertices, whose coordinates in the local coordinate system are expressed as: , the local coordinates of the four endpoints of the bottom surface are expressed as: , is the length of the bounding box, is the width of the 3D bounding box, is the height of the 3D bounding box;

[0070] The local coordinates Converted to the world coordinate system, the formula is: , the coordinates in the world coordinate system Converted to the camera coordinate system, the formula is: ;

[0071] Pinhole projection is performed according to the intrinsic parameter matrix K, and its formula is: , through normalization, the projection is homogenized to obtain the corresponding pixel coordinates ,in, is the horizontal homogeneous coordinate, that is The first row component of is the longitudinal homogeneous coordinate, that is The second row component of is the depth component, i.e. The third row component of ;

[0072] Project the four vertices of the 3D minimum oriented bounding box onto the image plane. After obtaining the corresponding pixel coordinates, calculate the Euclidean distance between them and the contact point. The formula is: , is the pixel horizontal coordinate of the jth bottom point of the two-dimensional bounding box on the image, is the pixel ordinate of the jth bottom point of the two-dimensional bounding box on the image;

[0073] The two endpoints closest to the bottom edge of the two-dimensional minimum oriented bounding box are selected from the four candidate pixels, and their Euclidean distance on the pixel plane is calculated to obtain the final projection width. The formula is: ,in, is the final projection width, which represents the width of the three-dimensional minimum oriented bounding box in the third component The pixel scale projected onto the image plane is output to the Dynamo visual programming tool through the node. is the horizontal coordinate of the first endpoint of the four candidate pixels of the 3D bounding box that is closest to the bottom edge of the 2D minimum oriented bounding box. is the ordinate of the first endpoint of the four candidate pixels of the 3D bounding box that is closest to the bottom edge of the 2D minimum oriented bounding box. is the horizontal coordinate of the second endpoint of the four candidate pixels of the 3D bounding box that is closest to the bottom edge of the 2D minimum oriented bounding box. It is the ordinate of the second endpoint of the four candidate pixels of the 3D bounding box that is closest to the bottom edge of the 2D minimum oriented bounding box.

[0074] In this example, the generated scale-free 3D mesh is imported into the Dynamo visual programming tool. Dynamo is a powerful visual programming tool that intuitively displays data processing flows and supports extensions in multiple programming languages. The mesh is read into the Dynamo CPython node using the Open3D data processing library. Open3D is an open-source, modern 3D data processing library that provides a rich set of geometry processing capabilities and enables efficient manipulation of 3D meshes.

[0075] In the Open3D environment, principal component analysis (PCA) is used to perform geometric estimation on the vertex set of a scale-free 3D mesh. PCA is a commonly used statistical method for extracting important features from high-dimensional data and reducing its dimensionality. PCA can be used to obtain the minimum oriented bounding box (OBB) enclosing the mesh and its properties. These properties include the 3D coordinates of the bounding box's center point, the minimum oriented bounding box's rotation matrix, and the size vector. These properties provide important foundational information for subsequent geometric calculations. Alternatively, the Python Trimesh library can be used to obtain the minimum volume OBB, achieving similar results to Open3D. Similarly, the minimum volume bounding box algorithm in CGAL or the rotating caliper method applied to the edges of the convex hull can be used to solve the 3D OBB, providing an alternative to extracting oriented bounding boxes from scale-free meshes.

[0076] In the local coordinate system of the minimum oriented bounding box, the bottom plane corresponds to the negative direction of the local height axis, and its normal vector is -e z The plane contains four vertices, whose coordinates in the local coordinate system can be calculated using simple geometric relationships. Next, the local coordinates are transformed into the world coordinate system. Subsequently, the coordinates in the world coordinate system are transformed into the camera coordinate system.

[0077] To project a 3D point onto a 2D image plane, pinhole projection is performed based on the intrinsic parameter matrix. The pinhole projection model is a commonly used projection model in computer vision, mapping points in 3D space onto a 2D image plane. Through normalization, the corresponding pixel coordinates are obtained.

[0078] After projecting the four vertices of the 3D minimum oriented bounding box onto the image plane, the corresponding pixel coordinates are obtained. Next, the Euclidean distance between these projection points and the contact point is calculated. The distance between each projection point and the contact point is calculated using the formula. From the four candidate pixel points, the two endpoints closest to the bottom edge of the 2D minimum oriented bounding box are selected, and their Euclidean distance on the pixel plane is calculated to obtain the final projection width (in pixels). The projection width represents the width of the 3D minimum oriented bounding box in the depth Z direction. real The pixel scale of the image is projected onto the image plane, and the result is output through the Dynamo visual programming tool.

[0079] The beneficial effect of this process is that it accurately maps geometric information in three-dimensional space onto the two-dimensional image plane and, by associating it with the contact points, calculates the actual projected width. This precise mapping and calculation provides key geometric constraints for subsequent scale recovery, enabling the true dimensions of the pile to be restored from a single two-dimensional image. The combined use of Dynamo and Open3D not only improves data processing efficiency but also enhances the scalability and flexibility of the entire system.

[0080] S5, under the pinhole camera model, calculates the linear scale factor according to the final projection width, calculates the real volume of the pile according to the linear scale factor, and uses visual programming tools to visualize and integrate the real volume of the pile.

[0081] Specifically, step S5 further includes: in the pinhole camera model framework, according to the image projection geometric relationship of the bounding box width, calculating the actual width according to the final projection width , where f is the focal length;

[0082] Define the linear scale factor based on the distance w between the two endpoints of the 3D minimum oriented bounding box in the local coordinate system , and calculate the actual volume of the pile by combining the linear scale factor ,in, is the volume of the generated scale-free 3D mesh model, which is calculated by Dynamo.

[0083] The Dynamo visual programming tool is used to map the actual volume of the stacked materials into the BIM model, thereby achieving real-time updating of the BIM model.

[0084] In this embodiment, the final projected width is used to calculate the true width of the pile within the framework of the pinhole camera model. The pinhole camera model is a simplified camera imaging model that assumes that light passes through the camera's optical center and is projected onto the image plane. Based on the image projection geometry of the bounding box width, the true width can be calculated using a formula. This calculation process has the beneficial effect of converting the pixel scale in the two-dimensional image to the actual physical scale, providing an accurate geometric foundation for subsequent volume calculations.

[0085] Next, a linear scale factor is defined based on the distance between the two endpoints of the 3D minimum oriented bounding box in the local coordinate system. The linear scale factor is a key parameter for converting the dimensions of the scale-free 3D mesh to its true physical dimensions. Using this scale factor, the volume of the scale-free 3D mesh can be converted to the true volume of the pile. This process is beneficial in that it accurately recovers the true volume of the pile, resolving the scale ambiguity issue in monocular 3D reconstruction and enabling direct volume measurement from 2D images.

[0086] To visualize the actual volume of the piled materials and update the BIM model in real time, the Dynamo visual programming tool was employed. Dynamo is a visual programming tool tightly integrated with BIM software (such as Revit), enabling easy mapping of calculation results into the BIM model. Using Dynamo, the calculated actual volume of the piled materials, along with associated 3D geometric information (such as position and dimensions), was mapped into the BIM model. This process not only enabled intuitive display of the volumetric information of the construction materials within the BIM model but also enabled real-time updates of the BIM model, providing real-time, accurate data support for material management, cost control, and progress monitoring on the construction site.

[0087] Simply put, the Dynamo visual programming language is used to input the type, 3D coordinates, and mesh model of the construction material into the BIM model. Dynamo then calculates the volume of the mesh model, and combined with the scale factor, the actual volume of the material can be calculated. By mapping this actual material data into the BIM, the BIM model can be updated in real time.

[0088] The beneficial effect of this process is that it not only improves the accuracy and efficiency of material volume measurement but also enhances the digital management of construction sites through integration with BIM models. The Dynamo tool can intuitively present complex calculation results to construction managers, enabling them to understand material usage on the construction site in real time and make more scientific and reasonable decisions. Furthermore, this real-time updated BIM model supports dynamic monitoring of construction progress, further enhancing the digital management of construction sites.

[0089] In simple terms, the fusion monocular Figure 3 A method for identifying and generating stacked material volumes based on 3D reconstruction and BIM calibration for single-view Figure 3 To address the scale ambiguity inherent in 3D reconstruction, a BIM virtual camera environment is introduced to provide a true depth prior. By combining YOLOv8-OBB object recognition and rotation detection, Open3D scale-free mesh OBB reconstruction, and a pinhole projection model, the method automatically recovers the true geometric dimensions of construction materials and accurately calculates their volume. This method, which relies solely on a fixed camera and an existing BIM model, eliminates the need for multi-view or laser scanning equipment and performs object detection, coordinate calculation, 3D reconstruction, scale recovery, and volume calibration from a single image of the construction site. This fully automated process offers high accuracy and engineering practicality. This approach aims to eliminate manual ranging and achieve efficient, automated, and highly accurate volume measurement of construction materials in complex construction environments. It also eliminates the need for expensive multi-view or laser scanning equipment, enabling volume measurement with a single fixed camera and strong adaptability to complex sites. By incorporating a true depth prior within a single image, the method automatically eliminates scale ambiguity in monocular reconstruction and achieves true geometric dimension recovery.

[0090] In summary, fusion monocular Figure 3 The method of identifying and generating the volume of stacked materials by 3D reconstruction and BIM calibration is achieved by fusing single-view Figure 3 3D reconstruction technology and Building Information Model (BIM) calibration achieve efficient and accurate conversion from 2D images to true 3D volumes. The core of this method is to use low-cost monocular cameras to capture images, combined with deep learning models and advanced geometric computing technology, to solve the problems of insufficient accuracy, high cost and low automation in traditional measurement methods.

[0091] Specifically, this method uses the pre-trained deep learning model YOLOv8-OBB to detect and classify the piled wood in the 2D image, rapidly extracting the contact points between the wood and the ground. This process not only improves detection efficiency but also ensures high reliability through confidence screening and appropriate endpoint selection. Subsequently, using the camera's internal and external calibration parameters and the virtual camera environment in the BIM model, the real-world depth of the midpoint of the pile's base is recovered through ray-plane intersection calculation. This step fully utilizes the precise geometric information of the BIM model, providing a critical depth prior for subsequent 3D reconstruction.

[0092] It's important to note that in addition to BIM ray-plane intersection, which uses the BIM reference ground plane and ray intersection to calculate depth, scale factors can also be calculated based on known-size markers. This involves placing ArUco / AprilTag markers on-site and directly calculating the scale factor using the ratio of their true size to their projected size. Alternatively, known object sizes, such as standardized components like safety guardrails and construction machinery, can be used to invert the scale factor and calibrate the grid by combining their CAD dimensions with their pixel projected width. This provides an alternative to scale factor recovery and calibration methods.

[0093] In the 3D reconstruction stage, a single view based on OccNet is used. Figure 3 The 3D reconstruction model generates a scale-free 3D mesh from a 2D image using an encoder-decoder architecture and a multi-resolution isosurface extraction algorithm. This process not only preserves key information in the image but also significantly improves the speed and accuracy of reconstruction through efficient algorithm design. Furthermore, using Open3D and PCA techniques, the minimum oriented bounding box (OBB) of the 3D mesh is extracted and its projected width on the image plane is calculated. This projection width calculation provides important geometric constraints for scale recovery.

[0094] Ultimately, using the pinhole camera model, this method calculated the linear scale factor based on the geometric relationship between the projected width and the true width, thereby restoring the true volume of the piled materials. This process not only resolves the scale ambiguity in monocular reconstruction but also enables the visualization of measurement results and dynamic updating of the BIM model through real-time integration with the BIM model. This real-time BIM model not only provides accurate data support for material management on the construction site but also lays the foundation for dynamic monitoring of construction progress and the application of digital twin technology.

[0095] Compared with the existing technology, this method uses a single-view Figure 3The organic combination of 3D reconstruction and BIM virtual camera scale calibration has brought significant technical improvements and application value in the field of volume measurement of piled materials on construction sites, mainly reflected in: 1. Only one fixed monocular camera perspective photogrammetry device is required, which greatly reduces equipment procurement and maintenance expenses; 2. The YOLOv8-OBB model is used to complete real-time piled material target detection on a single image, and combined with Open3D to automatically extract scale-free oriented bounding boxes. The entire detection-reconstruction-calibration process can be executed with one click through scripts and Dynamo visual programming, greatly improving measurement efficiency and reducing manual intervention; 3. With the help of the virtual camera intrinsic parameter matrix and extrinsic parameters configured in the BIM environment, the centimeter is obtained through the ray-plane intersection algorithm The method uses a meter-level true depth prior, combined with a pinhole projection model to invert the linear scale factor, to completely resolve the scale ambiguity problem in monocular 3D reconstruction and achieve accurate reconstruction of true physical size and volume. 4. Leveraging YOLOv8-OBB's rotation detection capability, it is highly robust to tilted and partially occluded targets. Open3D's PCA-based oriented bounding box algorithm can stably extract geometric features from irregular grid shapes, enabling the method to maintain high-precision measurements even in complex construction environments. 5. Real-world camera parameters can be directly loaded into Revit via Dynamo or the API, automatically mapping image pixel coordinates to BIM 3D space, achieving an end-to-end SOP (Standard Operating Procedure) workflow and facilitating its application to daily project management. 6. It can be used not only for static material volume measurement but also in combination with an automatic progress monitoring solution based on BIM geometric constraints and point cloud matching to achieve digital twins and progress comparisons in dynamic construction scenes, demonstrating its potential for lighter-weight and more real-time engineering applications.

[0096] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for identifying and generating piled material volume by integrating single-view 3D reconstruction and BIM calibration, characterized in that: include: Obtain a two-dimensional image captured by a preset camera, use a preset multi-task detection model to perform rotating target detection and pile material classification on the two-dimensional image, and obtain the contact points between the pile and the ground; Obtain the intrinsic parameter matrix and extrinsic parameter matrix of the preset camera, reproduce the configuration of the intrinsic parameter matrix and the extrinsic parameter matrix in a visual programming tool, perform ray-plane intersection calculation on the ray constructed at the midpoint of the base and the BIM ground plane equation, and obtain the real space depth corresponding to the midpoint of the base in the image. Specifically, substitute the ray-plane intersection parameters into the intersection to obtain the three-dimensional coordinates of the construction pile in the world coordinate system, remap the intersection to the camera coordinate system to obtain a vector, and extract the third component of the vector as the real depth of the three-dimensional point corresponding to the coordinates of the midpoint of the base in the image relative to the camera; Calling a pre-trained OccNet model to perform calculations on the two-dimensional image to generate a scale-free three-dimensional grid of the construction material pile; Use a visual programming tool to extract and process the scale-free 3D grid to obtain the pixel coordinates of the four endpoints of the bottom surface of the 3D minimum oriented bounding box, and calculate the Euclidean distance between the pixel coordinates and the contact point to obtain the final projection width; Under the pinhole camera model, the linear scale factor is calculated according to the final projection width, the real volume of the pile is calculated based on the linear scale factor, and the real volume of the pile is visualized and integrated using visual programming tools; Before calling the pre-trained OccNet model to calculate the two-dimensional image, the following steps are also included: Obtain the ShapeNetCore dataset and pre-train the initial OccNet model based on the ShapeNetCore dataset to provide initial implicit field parameters; Obtain the RGB image of the wood piling scene on the construction site and the corresponding laser scanning point cloud data collected by the preset drone, and perform surface reconstruction on the corresponding laser scanning point cloud data to generate a triangular mesh model as the ground truth value of the implicit field; Use Blender software to build synthetic scenes with various stacking geometry and materials, randomly adjust lighting, camera pose, and texture, expand the dataset, and divide the dataset into training, validation, and test sets in proportion. The initial OccNet model is trained based on the data set, and the trained initial OccNet model is verified and evaluated based on the two indicators of IoU and Chamfer distance until the model output reaches the preset value, and the trained OccNet model is obtained.

2. The method for identifying and generating piled material volume by integrating single-view 3D reconstruction and BIM calibration according to claim 1 is characterized in that: Obtain the 2D image captured by the preset camera, use the preset multi-task detection model to perform rotating target detection and pile material classification on the 2D image, and obtain the contact points between the pile and the ground, specifically: Acquire a two-dimensional image of the construction materials piled up, captured by a camera configured at the construction site, and pre-process the two-dimensional image; The preprocessed 2D image is fed into a multi-task detection model based on YOLOv8-OBB to perform rotation target detection and wood pile classification tasks, and output information results, where the information results include the pixel coordinates of the four vertices of the rotation bounding box of each wood pile target, the confidence score, and the category label; The information results with confidence values ​​lower than the preset value are filtered out, and the four vertex pixel coordinates are arranged in descending order according to the vertical coordinates of the pixel coordinates. The first two vertex pixel coordinates in the sequence are selected as the endpoints of the bottom edge of the pile, marked as (u1, v1) and (u2, v2) to correspond to the contact points between the pile and the ground.

3. The method for identifying and generating stacked material volume by integrating single-view 3D reconstruction and BIM calibration according to claim 2 is characterized in that: Obtain the intrinsic parameter matrix and extrinsic parameter matrix of the preset camera, specifically: Calibrate the preset camera to obtain its intrinsic parameter matrix K and extrinsic parameter matrix [R|T]; Among them, the intrinsic parameter matrix K is used to describe the internal optical characteristics of the camera, including focal length, principal point coordinates and distortion coefficients, and its matrix form is: , is the focal length of the camera in the x direction, is the focal length of the camera in the y direction, is the horizontal coordinate of the camera principal point, is the ordinate of the camera's principal point; The extrinsic matrix [R|T] is used to describe the position and posture of the camera in the world coordinate system. It consists of the rotation matrix R and the translation vector t. Its matrix form is: , is the translation component of the camera’s optical center on the x-axis of the world coordinate system, is the translation component of the camera’s optical center on the y-axis of the world coordinate system, is the translation component of the camera’s optical center on the z-axis of the world coordinate system, is the projection of the camera coordinate system x-axis on the world coordinate system x-axis, is the projection of the camera coordinate system y-axis on the world coordinate system x-axis, is the projection of the camera coordinate system z-axis on the world coordinate system x-axis, is the projection of the camera coordinate system x-axis on the world coordinate system y-axis, is the projection of the y-axis of the camera coordinate system on the y-axis of the world coordinate system, is the projection of the camera coordinate system z-axis on the world coordinate system y-axis, is the projection of the camera coordinate system x-axis on the world coordinate system z-axis, is the projection of the camera coordinate system y-axis on the world coordinate system z-axis, It is the projection of the camera coordinate system z-axis onto the world coordinate system z-axis.

4. The method for identifying and generating piled material volume by integrating single-view 3D reconstruction and BIM calibration according to claim 3 is characterized in that: Obtain the intrinsic and extrinsic matrix of the preset camera. Recreate the configuration of the intrinsic and extrinsic matrix in a visual programming tool. Calculate the ray-plane intersection between the constructed ray at the midpoint of the base and the BIM ground plane equation to obtain the real-world depth corresponding to the midpoint of the base in the image. Specifically, Combine Revit software and Dynamo visual programming tools to create a new perspective view. Set the viewpoint position and orientation of the perspective view according to the intrinsic and extrinsic matrix of the preset camera to ensure that the virtual viewpoint completely overlaps with the preset camera in three-dimensional space. According to the contact point, calculate the midpoint coordinates of the bottom end point of the pile (u m ,v m ), the formula is: , and use the inverse matrix K of the internal parameter matrix K -1 , the midpoint coordinates (u m ,v m ) is converted to the direction vector in the camera coordinate system, and its formula is: , is the direction vector from the camera optical center to the current pixel on the image, is the transpose operation; According to the formula The direction vector in the camera coordinate system Convert to the BIM world coordinate system, and in the BIM world coordinate system, take the camera center C as the starting point and follow the direction vector Generate a ray , to realize the back projection of pixel points into three-dimensional space, where, is the transpose of the camera rotation matrix; Ground plane equations in BIM models , and rays Perform the combined equation to obtain the ray-plane intersection parameters ,in, is the plane normal vector, is the offset of the plane relative to the origin of the world coordinate system, is any point in three-dimensional space; Set the ray-plane intersection parameters Substitute the intersection point P hit In the figure, the mark is the three-dimensional coordinate of the construction material in the world coordinate system, and the formula is: and the intersection point P hit Remap to the camera coordinate system to get the vector , extract the third component of the vector As the coordinates of the midpoint of the bottom edge of the image (u m ,v m ) corresponds to the true depth of the 3D point relative to the camera, is the first component, is the second component, Represents a ray-plane intersection point.

5. The method for identifying and generating piled material volume by integrating single-view 3D reconstruction and BIM calibration according to claim 1 is characterized in that: The OccNet model adopts an encoder-decoder architecture, where the encoder is a ResNet-18 convolutional network used to extract the 256-dimensional feature vector of the input image, and the decoder consists of 5 layers of fully connected blocks.

6. The method for identifying and generating stacked material volume by integrating single-view 3D reconstruction and BIM calibration according to claim 5 is characterized in that: The pre-trained OccNet model is called to calculate the two-dimensional image and generate a scale-free three-dimensional grid of the construction materials. Specifically: Extract the two-dimensional image using the encoder of the OccNet model to obtain a 256-dimensional feature vector of the two-dimensional image; The decoder of the OccNet model is used to normalize the 256-dimensional feature vector to fuse image features and output the occupancy probability of any 3D point (x, y, z); A multi-resolution isosurface extraction algorithm is used to recursively subdivide only active voxels in an octree structure, gradually refining them to the target resolution. The occupied voxels are converted into continuous triangular meshes using the MarchingCubes algorithm to obtain a scale-free three-dimensional mesh.

7. The method for identifying and generating piled material volume by integrating single-view 3D reconstruction and BIM calibration according to claim 3 is characterized in that: Use a visual programming tool to extract and process the scale-free 3D grid to obtain the pixel coordinates of the four endpoints of the bottom surface of the 3D minimum oriented bounding box. Then calculate the Euclidean distance between the pixel coordinates and the contact point to obtain the final projection width, which is: Importing the scale-free three-dimensional grid into the Dynamo visual programming tool, and reading the grid through the Open3D data processing library at the CPython node of the Dynamo visual programming tool; In the Open3D environment, a PCA-based geometric estimation is performed on the mesh vertex set of a scale-free 3D mesh to obtain the minimum oriented bounding box surrounding the mesh and its attributes, where the attributes include the 3D coordinates of the center point c of the bounding box and the rotation matrix of the minimum oriented bounding box. and the size vector (l, w, h); In the local coordinate system of the minimum oriented bounding box, the bottom plane corresponds to the local height axis e z The negative direction of the normal vector is -e z , the plane contains four vertices, whose coordinates in the local coordinate system are expressed as: , the local coordinates of the four endpoints of the bottom surface are expressed as: , is the length of the bounding box, is the width of the 3D bounding box, is the height of the 3D bounding box; The local coordinates Converted to the world coordinate system, the formula is: , the coordinates in the world coordinate system Converted to the camera coordinate system, the formula is: ; Pinhole projection is performed according to the intrinsic parameter matrix K, and its formula is: , through normalization, the projection is homogenized to obtain the corresponding pixel coordinates ,in, is the horizontal homogeneous coordinate, is the vertical homogeneous coordinate, is the depth component; Project the four vertices of the 3D minimum oriented bounding box onto the image plane. After obtaining the corresponding pixel coordinates, calculate the Euclidean distance between them and the contact point. The formula is: , is the pixel horizontal coordinate of the jth bottom point of the two-dimensional bounding box on the image, is the pixel ordinate of the jth bottom point of the two-dimensional bounding box on the image; The two endpoints closest to the bottom edge of the two-dimensional minimum oriented bounding box are selected from the four candidate pixels, and their Euclidean distance on the pixel plane is calculated to obtain the final projection width. The formula is: ,in, is the final projection width, which represents the width of the three-dimensional minimum oriented bounding box in the third component The pixel scale projected onto the image plane is output to the Dynamo visual programming tool through the node. is the horizontal coordinate of the first endpoint of the four candidate pixels of the 3D bounding box that is closest to the bottom edge of the 2D minimum oriented bounding box. is the ordinate of the first endpoint of the four candidate pixels of the 3D bounding box that is closest to the bottom edge of the 2D minimum oriented bounding box. is the horizontal coordinate of the second endpoint of the four candidate pixels of the 3D bounding box that is closest to the bottom edge of the 2D minimum oriented bounding box. It is the ordinate of the second endpoint of the four candidate pixels of the 3D bounding box that is closest to the bottom edge of the 2D minimum oriented bounding box.

8. The method for identifying and generating piled material volume by integrating single-view 3D reconstruction and BIM calibration according to claim 7 is characterized in that: Under the pinhole camera model, the linear scale factor is calculated based on the final projection width, and the actual volume of the pile is calculated based on the linear scale factor. Specifically: In the pinhole camera model framework, the true width is calculated based on the image projection geometry of the bounding box width and the final projection width. , where f is the focal length; Define the linear scale factor based on the distance w between the two endpoints of the 3D minimum oriented bounding box in the local coordinate system , and calculate the actual volume of the pile by combining the linear scale factor ,in, is the volume of the generated scale-free 3D mesh model, which is calculated by Dynamo.

9. The method for identifying and generating piled material volume by integrating single-view 3D reconstruction and BIM calibration according to claim 1, characterized in that: The real volume of the piled materials is visualized and integrated using a visual programming tool. Specifically, the real volume of the piled materials is mapped into the BIM model using the Dynamo visual programming tool to achieve real-time updating of the BIM model.

Citation Information

Patent Citations

  • Directional calibration target for camera inner and outer parameter calibration

    CN104867160A

  • Three-dimensional reconstruction method for vehicle target in road scene based on monocular vision

    CN113129348A