Open world construction scene reconstruction method based on video and point cloud bimodal data

By combining dual-modal data from LiDAR and industrial cameras with open-world real-time object recognition technology, the problems of information loss and dynamic adaptability in 3D construction scene reconstruction are solved, achieving efficient and automated 3D reconstruction and real-time monitoring, adapting to the dynamic changes of complex construction scenes.

CN121962467APending Publication Date: 2026-05-01SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2026-02-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing 3D construction scene reconstruction methods suffer from problems such as missing information, long reconstruction time, and large computational load in complex environments. They are difficult to adapt to dynamic changes in open worlds. Traditional methods cannot effectively handle the complex and ever-changing environment and objects on site, resulting in inaccurate and incomplete reconstruction.

Method used

A method based on dual-modal data of video and point cloud is adopted. By assembling LiDAR and industrial camera, three-dimensional point cloud and two-dimensional color RGB images are acquired simultaneously. Combined with automatic pixel-level calibration algorithm and open-world real-time object recognition technology, real-time visualization and object instance segmentation are achieved. Extrinsic parameter matrix is ​​used for accurate mapping to improve reconstruction accuracy and completeness.

Benefits of technology

It achieves high-precision, fully automated reconstruction of 3D construction scenes, adapts to dynamic changes in the open world, improves reconstruction efficiency and intelligence, meets the needs of real-time monitoring and rapid response, and has good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962467A_ABST
    Figure CN121962467A_ABST
Patent Text Reader

Abstract

The invention discloses an open world construction scene reconstruction method based on video and point cloud bimodal data, and belongs to the technical field of machine vision and artificial intelligence combination, and the method comprises the steps: carrying out the synchronous collection of the bimodal data of a three-dimensional point cloud and a two-dimensional color RGB image in a construction scene through assembling a laser radar and an industrial camera and anchoring a single visual angle; calibrating an external parameter matrix for the collected point cloud and image by using an automatic pixel-level calibration algorithm, and visualizing the recorded and registered color point cloud in real time; segmenting a specific object in the construction scene for the color point cloud instance captured in real time based on an input open world text cue word; and accurately mapping the segmented object region to a three-dimensional point cloud through an external parameter matrix to realize accurate analysis of a three-dimensional space. According to the method, the precision and integrity of three-dimensional construction scene reconstruction are improved, the intelligent level of construction scene reconstruction is improved, and the method can adapt to dynamically changing environments and novel objects in the open world.
Need to check novelty before this filing date? Find Prior Art

Description

Open-world construction scene reconstruction method based on video and point cloud dual-modal data Technical Field

[0001] This invention relates to the field of machine vision and artificial intelligence, and in particular to a method for reconstructing open-world construction scenes based on dual-modal data of video and point cloud. Background Technology

[0002] In the field of computer vision, rapid 3D construction scene reconstruction technology is one of the key research topics. Currently, common 3D reconstruction methods mainly include image rendering and compositing, surveying and modeling, and radar point cloud scanning. However, with the rapid development of social production technology, the demand for applications that can quickly scan and reconstruct 3D construction scenes in various engineering projects is increasing. In practical engineering applications, due to the long reconstruction time, large computational load, and time-consuming and labor-intensive nature of image rendering and compositing and surveying and modeling, radar point cloud scanning is usually more commonly used to reconstruct 3D construction scenes.

[0003] With the rapid development of computer vision and artificial intelligence technologies, 3D construction scene reconstruction technology has been widely applied in many fields such as building construction, industrial inspection, and virtual reality. Traditional 3D construction scene reconstruction methods mainly rely on single-modal data acquisition, such as LiDAR or camera images. While these methods can achieve 3D construction scene reconstruction to a certain extent, they have many limitations. For example, when using only LiDAR data for reconstruction, although relatively accurate depth information can be obtained, data loss or noise interference is likely to occur when dealing with complex construction scenes; while relying solely on camera image data for reconstruction can provide rich texture information, it lacks depth information, resulting in an inaccurate and incomplete 3D construction scene reconstruction.

[0004] In recent years, with the continuous development of deep learning technology, instance segmentation technology has made significant progress in the field of image processing. Instance segmentation technology can accurately segment each object in an image and assign it a unique identifier, providing new ideas and methods for solving the problem of 3D construction scene reconstruction in open-world construction scenarios. However, most existing instance segmentation technologies are concentrated in the field of 2D images, and there are still many technical challenges in processing 3D point cloud data. How to effectively apply instance segmentation technology to 3D point cloud data and achieve rapid reconstruction of open-world construction scenes has become one of the important problems that urgently need to be solved in the field of computer vision and artificial intelligence.

[0005] Meanwhile, the development of open-world real-time object recognition technology has provided a new solution to this problem. Open-world recognition refers to the system's ability to identify novel objects not seen during the training phase in a constantly changing environment. This technology is particularly important in field applications because the environment and objects on-site are often in dynamic change, and traditional closed-world recognition methods are difficult to adapt to such dynamic changes. Therefore, researching and developing a rapid reconstruction system for open-world 3D construction scenes based on dual-modal data of video and point clouds is of great significance for improving on-site safety and monitoring efficiency, and has broad application prospects. Summary of the Invention

[0006] The purpose of this invention is to provide an open-world construction scene reconstruction method based on dual-modal data of video and point cloud. This method addresses the problems of traditional construction scene manual inspection using only simple methods such as two-dimensional video streams, which require inspectors to maintain high concentration for extended periods and are prone to fatigue or negligence. Furthermore, traditional methods cannot effectively handle the complex and ever-changing environment and objects on-site, making them unsuitable for real-time object recognition in open-world construction scenarios. This invention aims to effectively improve the efficiency of object instance segmentation by integrating the advantages of dual-modal data from LiDAR and industrial cameras, combined with open-world real-time object recognition technology.

[0007] To achieve the above objectives, this invention provides an open-world construction scene reconstruction method based on dual-modal data of video and point cloud, comprising the following steps: S1, simultaneously acquiring dual-modal data of 3D point cloud and 2D color RGB image within the construction scene by assembling a LiDAR and an industrial camera and anchoring a single viewpoint; S2, extracting features from the acquired construction scene point cloud and image, accurately calibrating the extrinsic parameter matrix using an automatic pixel-level calibration algorithm, and achieving real-time visualization of the recorded and registered color point cloud; S3, segmenting specific objects in the construction scene based on input open-world text prompts; S4, accurately mapping the segmented object regions to the 3D point cloud using the extrinsic parameter matrix to achieve accurate 3D spatial analysis.

[0008] Preferably, in S1, the dual-modal data synchronous acquisition includes the following steps: S11, ensuring that the point cloud data acquired by the LiDAR corresponds to the RGB image data captured by the industrial camera in time, by recording and matching the timestamps of both; S12, controlling the LiDAR and the industrial camera to start and stop data acquisition simultaneously to ensure the synchronization of data acquisition; S13, setting the LiDAR and the industrial camera to acquire data at the same time interval to ensure that the point cloud and RGB image of each frame are consistent in time; S14, realizing a one-to-one correspondence between point cloud data and RGB image data within the same frame, providing accurate data pairs for subsequent dual-modal data registration; S15, using the synchronously acquired data above, preparing a benchmark dataset for dual-modal data registration to support high-precision point cloud and image fusion.

[0009] Preferably, in S2, feature extraction includes the following steps: S201, extracting the construction scene point cloud data and RGB image data recorded in S1; S202, downsampling the collected raw point cloud data to reduce the amount of data and improve the efficiency of subsequent processing, while retaining key geometric features; S203, applying statistical methods to identify and remove outliers in the point cloud, and the noise reduction process helps to improve the quality of the point cloud data; S204, determining the calibration range of the point cloud data, i.e., the region of interest, focusing on analysis and eliminating interference from irrelevant data.

[0010] Preferably, in S2, the specific process includes the following steps: S21, using a calibration board to capture images to obtain the intrinsic parameter matrix of the industrial camera for calibration of the construction scene; S22, applying an automatic pixel-level calibration algorithm to extract edge feature information and calibrate the extrinsic parameter matrix of the collected construction scene dataset; S23, applying the calibrated extrinsic parameter matrix, and combining the timestamp matching in S1 on-site, merging the recorded LiDAR point cloud and the RGB image recorded by the industrial camera in the same frame to achieve real-time four-dimensional color point cloud visualization of the on-site recorded data, where four-dimensional specifically refers to three-dimensional coordinates plus one-dimensional color information.

[0011] Preferably, step S3 specifically includes the following steps: S31, using prompt words from an open-source text set to construct a semantic understanding model based on an advanced semantic understanding architecture; to achieve efficient semantic parsing and understanding of text data related to the construction scene; the semantic understanding architecture is used to process open-world text data, including but not limited to the Transformer architecture, the BERT autoencoder semantic understanding model, a semantic understanding model based on graph neural networks, and a semantic understanding model based on knowledge graphs; S32, inputting open-world text data related to the construction scene into the semantic understanding model; S33, based on the collected low-latency real-time four-dimensional color point cloud data from the site, applying the input open-world text prompt words in real time to perform object instance segmentation on each frame of the image.

[0012] Preferably, in S4, specifically: using the calibrated extrinsic parameter matrix and four-dimensional color point cloud data, the image results after segmentation of the construction scene object instance are mapped to the color point cloud of the corresponding frame in real time.

[0013] This invention also provides an open-world construction scene reconstruction system based on dual-modal data of video and point cloud, comprising: an assembly and acquisition module: simultaneously acquiring dual-modal data of 3D point cloud and 2D color RGB image within the construction scene by assembling a LiDAR and an industrial camera and anchoring a single viewpoint; a recording and visualization module: accurately calibrating the extrinsic parameter matrix using an automatic pixel-level calibration algorithm for the acquired construction scene point cloud and image, and realizing real-time visualization of the recorded and registered color point cloud; an open-world recognition module: segmenting specific objects in the construction scene based on input open-world text prompts; and a 3D scene mapping module: accurately mapping the segmented object region to the 3D point cloud through the extrinsic parameter matrix to achieve accurate 3D spatial analysis.

[0014] Preferably, the lidar acquires point cloud data in real time at high frequency, providing sufficient data density for rapid acquisition; the lidar and industrial camera acquire data through a time synchronization mechanism to ensure the consistency of dual-modal data in time, providing a guarantee for high-precision fusion; the lidar is connected to the edge computing device through a standard interface and integrated into the system of this invention; the industrial camera provides high-resolution color images, capturing detailed information on site; it is designed with a wide dynamic range to maintain image quality under different lighting conditions; it supports high frame rate acquisition, providing a smooth video stream for real-time construction scene segmentation and reconstruction; it is connected to the edge computing device through a standard interface and integrated into the system of this invention.

[0015] This invention provides an edge computing device, comprising: a processor designed specifically for edge computing to handle complex computing tasks, including machine learning and computer vision; a memory storing a computer-executable program, which, when executed by the processor, enables the processor to perform an open-world construction scene reconstruction method based on bimodal video and point cloud data; a communication interface supporting multiple communication protocols to achieve efficient data transmission between devices; power management including a power management system to adapt to energy constraints in edge computing environments; and an expansion interface to support the connection and expansion of external devices, including LiDAR, industrial cameras, mice, and keyboards.

[0016] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements an open-world construction scene reconstruction method based on dual-modal data of video and point cloud.

[0017] Therefore, the above-mentioned open-world construction scene reconstruction method based on video and point cloud dual-modal data has the following beneficial effects: 1) Improve reconstruction accuracy and completeness. By fusing dual-modal data of lidar point cloud and industrial camera images, the advantages of the two types of data are fully utilized, effectively making up for the information loss problem of single modality in complex construction scenes, improving the accuracy and completeness of three-dimensional construction scene reconstruction, and more realistically and accurately restoring the three-dimensional structure and details of the construction scene.

[0018] 2) Improve automation and intelligence by introducing open-world real-time object recognition technology. Based on the input open-world text prompts, the real-time captured color point cloud is segmented into instances and accurately mapped to the 3D point cloud to achieve accurate analysis of 3D space. This enables full-process automation from data acquisition to object recognition, segmentation and 3D reconstruction, improving the efficiency and automation of construction scene reconstruction, while adapting to the dynamically changing environment and new objects in the open world.

[0019] 3) Real-time performance and efficiency: Optimize data acquisition, processing, and analysis processes to achieve low-latency real-time reconstruction and visualization. High-frequency data acquisition, combined with automatic pixel-level calibration algorithms and the powerful processing capabilities of edge computing devices, ensures real-time processing of large amounts of dual-modal data and rapid generation of four-dimensional color point clouds for visualization, meeting the requirements of real-time monitoring and rapid response at construction sites.

[0020] 4) Excellent scalability: The system's software architecture is also easy to expand and upgrade, allowing for the convenient addition of new functional modules or optimization and improvement of existing modules to adapt to ever-changing application needs and technological developments.

[0021] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0022] Figure 1 is a flowchart of an embodiment of the present invention; Figure 2 is an overall logical framework diagram of an embodiment of the present invention; Figure 3 is a system composition block diagram of an embodiment of the present invention; Figure 4 is a schematic diagram of the hardware structure of the edge computing device of an embodiment of the present invention; Figure 5 is a visualization result of the original point cloud collected by Livox lidar at the construction site in an embodiment of the present invention; Figure 6 is a visualization result of the original color RGB image collected by MVS camera at the construction site in an embodiment of the present invention; Figure 7 is a registration intrinsic and extrinsic parameter matrix obtained after registering a single-frame laboratory-acquired image and point cloud using the edge feature extraction algorithm in an embodiment of the present invention; Figure 8 is a visualization of a single-frame color point cloud at the construction site (corresponding to Figures 4 and 5) obtained after applying the registration matrix in an embodiment of the present invention; Figure 9 is a two-dimensional image instance segmentation mask extracted after applying the Groungding-DINO open-world set prompt input text "person" (construction worker) and "excavator" (excavator) in an embodiment of the present invention, including accuracy indicators; a is a two-dimensional image, b is the obtained segmentation mask; Figure 10 is a visualization of the three-dimensional segmentation mask result obtained by reprojecting the two-dimensional image mask back into the original point cloud in an embodiment of the present invention. Detailed Implementation

[0023] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0024] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0025] Example 1: This invention provides an open-world construction scene reconstruction method based on dual-modal data of video and point cloud. The process is shown in Figure 1. Figure 2 is the overall logical framework diagram of this embodiment, which integrates the process in Figure 1 at the system level and complements Figure 1, for understanding both the method process and the system logic. Specifically, it includes the following steps: S1, synchronously acquiring dual-modal data of three-dimensional point cloud and two-dimensional color RGB image within the construction scene by assembling a LiDAR and an industrial camera and anchoring a single viewpoint. The dual-modal data synchronous acquisition includes the following steps: S11, ensuring that the point cloud data acquired by the LiDAR corresponds to the RGB image data captured by the industrial camera in time, achieved by recording and matching their timestamps; S12, controlling the LiDAR and industrial camera to start and stop data acquisition simultaneously to ensure data acquisition synchronization; S13, setting the LiDAR and industrial camera to acquire data at the same time interval to ensure that the point cloud and RGB image in each frame are consistent in time; S14, achieving a one-to-one correspondence between point cloud data and RGB image data within the same frame, providing accurate data pairs for subsequent dual-modal data registration; S15, using the synchronously acquired data to prepare a benchmark dataset for dual-modal data registration to support high-precision point cloud and image fusion.

[0026] In this embodiment, the Livox Avia lidar and the HikRobot MVS industrial camera are fixed at the same viewing angle using a rigid bracket to ensure that their fields of view are aligned and the overlap area is greater than 80%. Microsecond-level time synchronization is achieved via the PTP (Precision Time Protocol), with the acquisition frequency set to 10Hz. Each frame of data includes: a 3D point cloud: scanned in real-time, covering the construction scene area within the lidar's field of view; and an RGB image: 2448×2880 resolution, representing the real-time construction scene corresponding to the 3D point cloud.

[0027] Recorded 3D point cloud and RGB image data are published via ROS2 topics. The recording topics for the Livox Avia LiDAR and HikRobot MVS industrial camera are: / livox / lidar (sensor_msgs / PointCloud2) / camera / color / image_raw (sensor_msgs / Image), respectively. Frame-level synchronization is achieved through a timestamp alignment mechanism. Before system startup, a parallel shell script (sync_record.sh) sends a "start" command to both the LiDAR and industrial camera nodes. Internally, the script calls the roslaunch command to simultaneously start the driver and synchronously terminates recording upon receiving the SIGINT signal, thus ensuring that the time deviation between the first and last frames is <1 ms. The raw data generated by the recording is centrally stored in ROS.bag format, with the file name including the start timestamp for easy offline retrieval later. During the playback phase, the `rosbag play` tool publishes the topics ` / livox / lidar` and ` / camera / color / image_raw` in the original timestamp order. The ROS Time Synchronizer filter subscribes to these two topics, performs frame alignment with a 33 ms time window, and outputs synchronized PointCloud2 and Image message pairs. Finally, the script `extract_sync_frames.sh` writes the matched frames sequentially to the corresponding folders (` / dataset / YYYY-MM-DD / pcd` and ` / rgb`), and generates a timestamped index file, forming a construction scene dataset that can be directly used for extrinsic parameter calibration and network training. In this embodiment, Figure 5 shows the visualization result of the original point cloud collected by the Livox LiDAR at the construction site; Figure 6 shows the visualization result of the original color RGB image collected by the MVS camera at the construction site.

[0028] S2. Feature extraction is performed on the collected construction scene point cloud and images. An automatic pixel-level calibration algorithm is used to accurately calibrate the extrinsic parameter matrix, enabling real-time visualization of the recorded and registered color point cloud. Specifically, this includes the following steps: S21. Images are captured using a calibration board to obtain the intrinsic parameter matrix of the industrial camera for construction scene calibration; S22. An automatic pixel-level calibration algorithm is applied to extract edge feature information and calibrate the extrinsic parameter matrix of the collected construction scene dataset; S23. Using the calibrated extrinsic parameter matrix, and combining it with the timestamp matching from S1, the recorded LiDAR point cloud and the RGB image recorded by the industrial camera are merged in the same frame to achieve real-time four-dimensional color point cloud visualization of the recorded data. The four-dimensional aspect specifically refers to three-dimensional coordinates plus one-dimensional color information.

[0029] The feature extraction specifically includes the following steps: S201, extracting the point cloud data and RGB image data of the construction scene recorded in S1; S202, downsampling the collected raw point cloud data to reduce the amount of data and improve the efficiency of subsequent processing, while retaining key geometric features; S203, applying statistical methods to identify and remove outliers in the point cloud, and the noise reduction process helps to improve the quality of the point cloud data; S204, determining the calibration range of the point cloud data, i.e., the region of interest, focusing on analysis and eliminating interference from irrelevant data.

[0030] The specific operation in this embodiment is as follows: First, PCL is applied to process the original point cloud as follows: the point cloud is downsampled using a VoxelGrid filter with a voxel size of 0.05m; outlier removal is performed using StatisticalOutlierRemoval with a neighbor count of 50 and a standard deviation factor of 1.0; the point cloud after the above two steps is then cropped, with the foresee direction as the center and the cropping range being x∈[0,30m], y∈[-15m,15m], z∈[-2m,5m]. This retains the data of the construction scene of interest while removing interference from edge construction scenes and distant construction scenes, reducing the amount of data while accelerating point cloud registration and processing efficiency.

[0031] Next, the camera intrinsic parameters are calibrated using Zhang Zhengyou's chessboard method to obtain the intrinsic parameter matrix K: ; where, (f x ,f y (c) is the focal length (unit: pixels), x ,c y () are the coordinates of the main point.

[0032] An automatic pixel-level calibration algorithm based on edge feature matching is adopted, eliminating the need for extrinsic parameter matrix calibration using a calibration board. The specific process is as follows: voxel segmentation is performed on the LiDAR point cloud, and RANSAC is used to fit the local plane to extract depth-continuous edges; the Canny algorithm is used to extract two-dimensional edges from the RGB image and a kD tree is constructed; the LiDAR edge points are projected onto the image plane to construct the residual function. Where π represents the pinhole projection model, and f is the distortion correction function. and Let L be the normal vector and position point of the image edge, and let L be the coordinate system of the lidar, C be the coordinate system of the camera, and C be the two-dimensional coordinate system of the image plane.

[0033] This represents the external parameters between the lidar and the camera to be calibrated.

[0034] The extrinsic parameter matrix typically contains rotation and translation information, used to describe the spatial position and orientation of one coordinate system relative to another. This represents a 3×3 rotation matrix from the lidar coordinate system L to the camera coordinate system C. This represents a 3×1 translation vector from the origin of the lidar coordinate system L to the origin of the camera coordinate system C. denoted as a special Euclidean group, which is a group consisting of all rotation matrices and translation vectors used to describe rigid body transformations in three-dimensional space.

[0035] During the calibration process, the goal of this embodiment is to accurately determine the parameters so that point cloud data and image data can be accurately converted between different coordinate systems. By minimizing reprojection error or other optimization objectives, the optimal rotation matrix and translation vector are solved, thereby achieving precise alignment between the LiDAR and the camera.

[0036] In this embodiment, the Levenberg-Marquardt algorithm is used to solve for the optimal extrinsic parameters, thereby minimizing the residual function.

[0037] Finally, the recorded and registered color point cloud was visualized in real time. The point cloud and the image were fused into a four-dimensional color point cloud (x, y, z, RGB) by applying the calibrated extrinsic parameter matrix, and then visualized in real time using RViz2 software. Figure 7 shows the registration extrinsic and extrinsic parameter matrix obtained after registering a single-frame laboratory image and point cloud using the edge feature extraction algorithm; Figure 8 shows the visualization of a single-frame construction site color point cloud obtained after applying the registration matrix, corresponding to Figures 4 and 5.

[0038] S3. Segment specific objects in the construction scene based on real-time captured color point cloud instances using input open-world text prompts. Specifically, this includes the following steps: S31. Construct a semantic understanding model based on an advanced semantic understanding architecture using prompts from an open-source text collection; this enables efficient semantic parsing and understanding of text data related to the construction scene. The semantic understanding architecture is used to process open-world text data, including the Transformer architecture, the BERT autoencoder semantic understanding model, a graph neural network-based semantic understanding model, and a knowledge graph-based semantic understanding model; S32. Input the open-world text data related to the construction scene into the semantic understanding model; S33. Based on the collected low-latency real-time four-dimensional color point cloud data from the site, apply the input open-world text prompts in real-time to segment objects in each frame of the image.

[0039] In this embodiment, the algorithm for segmenting specific objects in a construction scene based on real-time captured color point cloud instances using input open-world text prompts employs the open-set object detection algorithm GroundingDINO1.5 Edge model.

[0040] The GroundingDINO1.5 model is based on a Transformer architecture and is used to process image and language data. Its deep early fusion strategy combines language and image features through a cross-attention mechanism during the feature extraction stage. Large-scale pre-training and self-training are performed on the Grounding-20M dataset, and self-training is conducted using pseudo-labeled data. The loss function optimizes bounding box regression and enhances the classification ability between predicted objects and language labels during training. GroundingDINO1.5 Edge is optimized for edge devices, using EfficientViT-L1 as the image backbone for fast multi-scale feature extraction.

[0041] Specifically, the GroundingDINO1.5 Edge model takes as input RGB images of a construction scene captured by an industrial camera and natural language prompts, and outputs bounding boxes of different categories of objects to be segmented, category confidence scores, and pixel-level masks. Figure 9 shows the two-dimensional image instance segmentation mask extracted after applying the Groungding-DINO open-world set prompts input text "person" (construction worker) and "excavator" (excavator) in this embodiment, including accuracy metrics.

[0042] Based on the input open-world text prompts, the real-time captured images of specific objects in the construction scene are segmented from real-time captured color point cloud instances. The real-time captured images are the RGB image stream of the real-time construction scene recorded by an industrial camera in S1.

[0043] S4. The segmented object region is accurately mapped to the 3D point cloud using an extrinsic parameter matrix to achieve accurate 3D spatial analysis. Specifically, using the calibrated extrinsic parameter matrix and 4D color point cloud data, the image results of the segmented construction scene object instance are mapped to the corresponding frame's color point cloud in real time.

[0044] In this embodiment, the image mask obtained in S3 is projected onto the point cloud space through an extrinsic parameter matrix. Specifically, for each pixel (u,v) within the mask, it is back-projected onto the spatial ray P. c : Further transformation to the lidar coordinate system P l :P l =( ) -1 ·P c Match the nearest point in the point cloud and assign a semantic label.

[0045] The output is a semantic color point cloud, supporting real-time rendering and interaction. The semantic labels of the point cloud include various objects required for segmentation by the open-world text prompts input in S3. Figure 10 is a visualization of the 3D segmentation mask result obtained by reprojecting the 2D image mask back into the original point cloud in this embodiment.

[0046] Figure 3 is a block diagram of the open world construction scene reconstruction system based on video and point cloud dual-modal data, including assembly and acquisition module 310, recording and visualization module 320, open world recognition module 330, and 3D scene mapping module 340.

[0047] The assembly and acquisition module 310 performs operation S1, simultaneously acquiring dual-modal data of 3D point cloud and 2D color RGB image within the construction scene by assembling a LiDAR and an industrial camera and anchoring them in a single-viewpoint configuration. The LiDAR acquires point cloud data at high frequency in real time, providing sufficient data density for rapid acquisition. The LiDAR and industrial camera acquire data through a time synchronization mechanism, ensuring temporal consistency of the dual-modal data and guaranteeing high-precision fusion. The LiDAR connects to the edge computing device via a standard interface and is integrated into the system of this invention. The industrial camera provides high-resolution color images, capturing detailed information from the scene; it is designed with a wide dynamic range to maintain image quality under different lighting conditions; it supports high frame rate acquisition, providing a smooth video stream for real-time construction scene segmentation and reconstruction; and it connects to the edge computing device via a standard interface and is integrated into the system of this invention.

[0048] The recording and visualization module 320 performs the S2 operation, which uses an automatic pixel-level calibration algorithm to accurately calibrate the extrinsic parameter matrix for the collected construction scene point cloud and images, and realizes real-time visualization of the recorded and registered color point cloud.

[0049] The open world recognition module 330 performs the S3 operation, segmenting specific objects in the construction scene based on the real-time captured color point cloud instances using the input open world text prompts.

[0050] The 3D scene mapping module 340 executes the S4 operation to accurately map the segmented object region to the 3D point cloud through the extrinsic parameter matrix to achieve accurate analysis of the 3D space.

[0051] This embodiment also illustrates an edge computing device, as shown in Figure 4. The device 400 includes a processor 410, a readable storage medium 420, and a communication interface 430. This device 400 can execute the aforementioned method for rapid reconstruction of open-world 3D construction scenes based on dual-modal data of video and point clouds.

[0052] The processor 410 includes an NVIDIA Jetson Orin Nano system-on-a-module (SoM) that integrates an ARM Cortex-A78AE CPU and an Ampere GPU, providing up to 40 TOPS of AI computing power and supporting TensorRT, CUDA, and OpenCV acceleration libraries. It also features 8 GB of onboard LPDDR5 memory for caching point clouds, images, and intermediate features. The processor 410 can function as a single processing unit or be expanded into a multi-node cluster via NVLink or Ethernet to perform various tasks in parallel, such as LiDAR point cloud preprocessing, industrial camera image inference, and security risk assessment.

[0053] The readable storage medium 420 can be any medium that contains, stores, transmits, propagates, or transfers instructions. It includes a 256 GB NVMe SSD (PCIe 3.0 x4) and a 64 GB eMMC 5.1 flash memory, used for storing historical video recordings, GroundingDINO 1.5 Edge weight files, and Docker images, respectively; an M.2 2242 slot is also provided for quick on-site replacement of larger capacity SSDs. The storage medium 420 can also be a vibration-resistant industrial-grade SD card or a SATA DOM to adapt to high-dust, high-vibration environments.

[0054] The communication interface 430 supports Gigabit Ethernet, 802.11ac Wi-Fi, and LTE / 5G daughter cards for remotely pushing alarm JSON to the monitoring platform via the MQTT(S) protocol. It also provides 4×USB 3.2 Gen1 ports, 1×HDMI 2.0 port, and 2×RS-485 ports, allowing connection to LiDAR, industrial cameras, audible and visual alarm lights, and PLCs, enabling heterogeneous access for multiple sensors. The interface driver layer incorporates a TSN (Time-Sensitive Networking) protocol stack, ensuring that the time error between point cloud and image frames on the transmission link is <1 ms, providing a guarantee for subsequent pixel-level fusion.

[0055] The edge computing unit used in this embodiment is an NVIDIA Jetson Orin Nano 8GB, running Ubuntu 22.04 with a kernel version of Linux 6.2.0-jetson. It supports ROS2, OpenCV, TensorRT, and PyTorch, and is started as a systemd service, supporting breakpoint resume and logging.

[0056] Computer program 421 can be configured to have one or more program modules, such as 421A acquisition module, 421B calibration module, 421C open-world recognition module, and 421D mapping module. When these program modules are executed by processor 410, the device 400 completes the entire process of lidar-industrial camera time synchronization, automatic extrinsic parameter calibration, open-world text instance segmentation and reasoning, and semantic point cloud mapping. The module division and number are not fixed. Those skilled in the art can add, delete, or merge functional modules according to on-site needs. As long as the method flow shown in Figures 1 and 2 can still be implemented, it falls within the protection scope of this invention.

[0057] Therefore, this invention adopts the above-mentioned open-world construction scene reconstruction method based on dual-modal data of video and point cloud. By fusing the precise depth information of LiDAR point cloud with the rich texture information of industrial camera images, it effectively compensates for the information loss problem of single-modal data in complex construction scenes, significantly improving the accuracy and completeness of 3D construction scene reconstruction, and can more realistically and accurately restore the 3D structure and details of the scene. It introduces open-world real-time object recognition technology, which can segment objects in real-time color point cloud based on input text prompts, and accurately map the segmentation results to 3D point cloud with the help of calibrated extrinsic matrix to achieve spatial analysis. It realizes full-process automation from data acquisition, object recognition to 3D reconstruction, greatly improving the intelligence level of construction scene reconstruction, and can adapt to the dynamically changing environment and new objects in the open world. At the same time, by optimizing the data acquisition and processing process and combining the high-efficiency computing power of edge computing devices, it achieves low-latency real-time reconstruction and 4D color point cloud visualization, meeting the needs of construction sites for real-time monitoring and rapid response. In addition, the system software architecture has good scalability, which facilitates the addition of new functional modules or optimization of existing modules to adapt to the ever-changing application needs and technological developments.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for reconstructing open-world construction scenes based on dual-modal data of video and point cloud, characterized in that, Includes the following steps: S1. Simultaneously acquire dual-modal data of 3D point cloud and 2D color RGB image within the construction scene by assembling a LiDAR and an industrial camera and anchoring a single viewpoint; S2. Extract features from the acquired construction scene point cloud and images, accurately calibrate the extrinsic parameter matrix using an automatic pixel-level calibration algorithm, and achieve real-time visualization of the recorded and registered color point cloud; S3. Segment specific objects in the construction scene based on input open-world text prompts; S4. Accurately map the segmented object regions to the 3D point cloud using the extrinsic parameter matrix to achieve precise 3D spatial analysis.

2. The method for reconstructing open-world construction scenes based on dual-modal data of video and point cloud as described in claim 1, characterized in that, In S1, the dual-modal data synchronous acquisition includes the following steps: S11, ensuring that the point cloud data acquired by the LiDAR corresponds to the RGB image data captured by the industrial camera in time, achieved by recording and matching their timestamps; S12, controlling the LiDAR and industrial camera to start and stop data acquisition simultaneously to ensure data acquisition synchronization; S13, setting the LiDAR and industrial camera to acquire data at the same time interval to ensure that the point cloud and RGB image in each frame are consistent in time; S14, realizing a one-to-one correspondence between point cloud data and RGB image data within the same frame, providing accurate data pairs for subsequent dual-modal data registration; S15, using the synchronously acquired data above, preparing a benchmark dataset for dual-modal data registration to support high-precision point cloud and image fusion.

3. The method for reconstructing open-world construction scenes based on dual-modal data of video and point cloud as described in claim 2, characterized in that, In S2, feature extraction includes the following steps: S201, extracting the construction scene point cloud data and RGB image data recorded in S1; S202, downsampling the collected raw point cloud data to reduce the amount of data and improve the efficiency of subsequent processing, while retaining key geometric features; S203, applying statistical methods to identify and remove outliers in the point cloud, and the noise reduction process helps to improve the quality of the point cloud data; S204, determining the calibration range of the point cloud data, i.e., the region of interest, focusing on analysis and eliminating interference from irrelevant data.

4. The method for reconstructing open-world construction scenes based on dual-modal data of video and point cloud as described in claim 3, characterized in that, In S2, the specific process The process includes the following steps: S21, using a calibration board to capture images to obtain the intrinsic parameter matrix of the industrial camera for calibration of the construction scene; S22, applying an automatic pixel-level calibration algorithm to extract edge feature information and calibrate the extrinsic parameter matrix of the collected construction scene dataset; S23, applying the calibrated extrinsic parameter matrix and combining it with the timestamp matching in S1 on-site, merging the recorded LiDAR point cloud and the RGB image recorded by the industrial camera in the same frame to achieve real-time four-dimensional color point cloud visualization of the on-site recorded data, where four-dimensional specifically refers to three-dimensional coordinates plus one-dimensional color information.

5. The method for reconstructing open-world construction scenes based on dual-modal data of video and point cloud as described in claim 4, characterized in that, S3 specifically includes the following steps: S31, using prompt words from an open-source text set to construct a semantic understanding model based on an advanced semantic understanding architecture; to achieve efficient semantic parsing and understanding of text data related to construction scenarios; the semantic understanding architecture is used to process open-world text data, including but not limited to the Transformer architecture, the BERT autoencoder semantic understanding model, a semantic understanding model based on graph neural networks, and a semantic understanding model based on knowledge graphs; S32, inputting open-world text data related to construction scenarios into the semantic understanding model; S33, based on the collected low-latency real-time four-dimensional color point cloud data from the site, applying the input open-world text prompt words in real time to perform object instance segmentation on each frame of the image.

6. The method for reconstructing open-world construction scenes based on dual-modal data of video and point cloud as described in claim 5, characterized in that, In S4, specifically: using the calibrated extrinsic parameter matrix and four-dimensional color point cloud data, the image results after segmentation of construction scene object instances are mapped in real time to the color point cloud of the corresponding frame.

7. An open-world construction scene reconstruction system based on dual-modal data of video and point cloud, characterized in that, include: Assembly and acquisition module: Simultaneously acquire dual-modal data of 3D point cloud and 2D color RGB image within the construction scene by assembling a lidar and an industrial camera and anchoring a single viewpoint; Recording and Visualization Module: For the collected construction scene point clouds and images, an automatic pixel-level calibration algorithm is used to accurately calibrate the extrinsic parameter matrix, enabling real-time visualization of the recorded and registered color point clouds; Open World Recognition Module: Based on input open world text prompts, specific objects in the construction scene are segmented from real-time captured color point cloud instances; 3D Scene Mapping Module: The segmented object regions are accurately mapped to the 3D point cloud through the extrinsic parameter matrix to achieve accurate 3D spatial analysis.

8. The open-world construction scene reconstruction system based on dual-modal data of video and point cloud as described in claim 7, characterized in that, include: The lidar acquires point cloud data in real time at high frequency, providing sufficient data density for rapid acquisition. The lidar and industrial camera acquire data through a time synchronization mechanism, ensuring temporal consistency of the dual-modal data and guaranteeing high-precision fusion. The lidar connects to edge computing devices via a standard interface and is integrated into the system of this invention. The industrial camera provides high-resolution color images, capturing detailed information from the scene. It is designed with a wide dynamic range to maintain image quality under different lighting conditions. It supports high frame rate acquisition, providing a smooth video stream for real-time construction scene segmentation and reconstruction. It connects to edge computing devices via a standard interface and is integrated into the system of this invention.

9. An edge computing device, characterized in that, include: The processor is designed for edge computing and handles complex computational tasks, including machine learning and computer vision; the memory stores computer-executable programs, which, when executed by the processor, enable the processor to perform an open-world construction scene reconstruction method based on bimodal video and point cloud data; The communication interface supports multiple communication protocols to achieve efficient data transmission between devices; power management includes a power management system to adapt to energy constraints in edge computing environments; The expansion interface is used to support the connection and expansion of external devices, including LiDAR, industrial cameras, mice and keyboards.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the open-world construction scene reconstruction method based on video and point cloud dual-modal data as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Three-dimensional building fine geometric reconstruction method integrating airborne and vehicle-mounted three-dimensional laser point clouds and streetscape images

    CN111815776A

  • Semantic segmentation and three-dimensional reconstruction-based clamped building block cutting system and method

    CN116310342A

  • High-place operation risk real-time judgment method based on image and point cloud bimodal data

    CN121459296A