Precast pile construction information real-time detection method based on cross-dimensional multi-modal sensing fusion

By employing a real-time detection method for precast pile construction information through cross-dimensional multimodal perception fusion, and utilizing deep learning and inverse projection algorithms, the automated and high-precision measurement of pile foundation construction parameters was achieved. This solved the problem of real-time detection in pile foundation construction and improved construction quality and safety.

CN121147833APending Publication Date: 2025-12-16STATE GRID JIANGSU ELECTRIC POWER ENG CONSULTING CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511048440.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing technologies cannot achieve real-time and accurate detection of pile length and inclination in pile foundation construction, which makes it difficult to guarantee construction quality and safety. Moreover, the detection methods are cumbersome and cannot meet the requirements of intelligent construction.

Method used

A real-time detection method for precast pile construction information using cross-dimensional multimodal perception fusion is adopted. By constructing a cross-dimensional multimodal perception device, combining an image acquisition module and a point cloud acquisition module, and using deep learning and inverse projection algorithms for data processing, a two-dimensional positioning mask and a cross-dimensional spatiotemporal alignment matrix are generated to achieve accurate reconstruction and parameter measurement of the target point cloud cluster of the precast pile.

Benefits of technology

It enables intelligent real-time detection of the precast pile construction process, accurately measures pile length and inclination with centimeter-level measurement accuracy, can stably detect in extreme environments, and reduces computational complexity, thereby improving the timeliness of construction quality control and on-site detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121147833A_ABST
    Figure CN121147833A_ABST
Patent Text Reader

Abstract

The invention discloses a precast pile construction information real-time detection method based on cross-dimensional multi-modal sensing fusion, which comprises the following steps: constructing a cross-dimensional multi-modal sensing device on a precast pile construction site, and arranging the device on the precast pile construction site; processing the acquired data by using a precast pile two-dimensional lightweight detection model based on deep learning to generate a two-dimensional positioning mask; generating a cross-dimension space-time alignment matrix based on rigid coupling calibration of the camera and the laser radar; mapping the acquired data to a camera coordinate system according to the cross-dimension space-time alignment matrix by utilizing a reverse projection driving algorithm, and performing point cloud feature screening through visual positioning information to obtain a precast pile target point cloud cluster; and spatial vector analysis is conducted on the precast pile target point cloud cluster, and the length of the precast pile is obtained. According to the method, the precision of pile length and gradient detection can be improved, direct recognition of segmentation point cloud data is avoided, and the detection efficiency is improved to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of target detection and image processing technology, specifically to a real-time detection method for precast pile construction information based on cross-dimensional multimodal perception fusion. Background Technology

[0002] In construction engineering, pile foundation engineering plays a crucial role in ensuring project safety. It transfers the load of the superstructure to the piles through pile caps or beams, and then to the deeper, more resilient soil (rock) layers, or compacts weak soil layers to improve the bearing capacity and density of the foundation soil. In this process, the pile bearing capacity is primarily borne by the pile's side skin friction and end resistance within the pile's driving depth; therefore, pile driving depth is one of the most important factors affecting the foundation's bearing capacity.

[0003] However, in actual pile foundation construction, due to varying skill levels among construction workers and the tight schedule and heavy workload, incorrect piles are frequently used and pile driving depths often fail to meet design requirements, seriously threatening building quality and safety. To solve this problem, a significant amount of manpower is typically required to monitor pile length changes in real time during construction. Construction technicians calculate the pile driving depth before driving, obtain approval from the supervisor, and then hand it over to the operator. Appropriate markings are made on the pile or pile driver, and on-site surveyors track and monitor the process. The pile driving depth is calculated based on the design pile top elevation, pile anchorage length, and the natural ground elevation at the pile location, ensuring that the pile top elevation deviation is controlled within -50 to +100 mm. While this method can accurately obtain the pile top elevation, it is still somewhat cumbersome in practical applications, cannot meet the requirements of intelligent construction, and cannot obtain pile length and inclination in real time during pile driving. Summary of the Invention

[0004] The purpose of this invention is to provide a real-time detection method for precast pile construction information based on cross-dimensional multimodal perception fusion, which can quickly realize the reconstruction of two-dimensional and three-dimensional space at the construction site, accurately fuse multimodal data to achieve automatic alignment of cross-dimensional information of precast piles, thereby realizing intelligent real-time detection of the entire pile foundation construction process.

[0005] This invention adopts the following technical solution: a real-time detection method for precast pile construction information based on cross-dimensional multimodal perception fusion, comprising the following steps:

[0006] S1. Construct a multi-dimensional multimodal sensing device for the precast pile construction site. The device includes a shell, an image acquisition module, a point cloud acquisition module, a data processing module, and a power supply. The device is set up on a tripod at a position 4m away from the pile driving point to be detected and 1m above the ground.

[0007] S2. In the data processing module, a two-dimensional lightweight detection model for precast piles based on deep learning is used to process the data collected in the image acquisition module and generate a two-dimensional positioning mask.

[0008] S3. Generate a cross-dimensional spatiotemporal alignment matrix based on rigid coupling calibration of camera and lidar.

[0009] S4. Using the inverse projection driving algorithm, the data collected in the point cloud acquisition module is mapped to the camera coordinate system according to the cross-dimensional spatiotemporal alignment matrix. Point cloud features are filtered through visual positioning information to obtain the precast pile target point cloud cluster, thus completing the processing in the data processing module.

[0010] S5. Perform spatial vector analysis on the target point cloud cluster of the precast pile to obtain the length of the precast pile.

[0011] Furthermore, in step S1, the outer shell is a rigid cuboid with a hollow internal structure, and the interior of the outer shell is divided into upper and lower parts by a transverse partition; the image acquisition module is fixed to the upper part of the outer shell with screws; the point cloud acquisition module is fixed to the lower part of the outer shell with screws; and the data processing module is located inside the outer shell.

[0012] The image acquisition module is connected to the data processing module via a USB data cable, the point cloud acquisition module is connected to the data processing module via a USB data cable, the point cloud acquisition module is connected to the power supply via a first power cable, and the data processing module is connected to the power supply via a second power cable.

[0013] The image acquisition module includes a camera and a USB data cable, used to acquire two-dimensional image data of precast piles at the construction site. The camera must be a fixed-focus lens.

[0014] The point cloud acquisition module includes a lidar, a first data cable, and a first power cable, and is used to acquire 3D point cloud data corresponding to images of precast piles at the construction site. The lidar can be a mechanical lidar or a solid-state lidar.

[0015] The data processing module includes an industrial motherboard, a display screen, a second data cable, and a second power cable. It receives data from the image acquisition module and the point cloud acquisition module through the second data cable, processes the data, and obtains the point cloud cluster of the precast pile target.

[0016] Furthermore, in step S2, generating the two-dimensional positioning mask includes the following:

[0017] The deep learning-based two-dimensional lightweight detection model for precast piles includes a lightweight feature extraction backbone network, a feature fusion and adaptive receptive field module, and a pixel-level localization mask prediction head.

[0018] The lightweight feature extraction backbone network includes depthwise separable convolutional layers; the feature fusion and adaptive receptive field module includes a cross-scale feature fusion unit and an adaptive receptive field enhancement unit; the adaptive receptive field enhancement unit includes a first 1×1 convolutional layer, three max pooling layers and a second 1×1 convolutional layer connected in sequence.

[0019] The two-dimensional image data from step S1 is input into a lightweight two-dimensional detection model for precast piles based on deep learning. After passing through a lightweight feature extraction backbone network, feature extraction is performed using depthwise separable convolutional layers to obtain multi-scale feature maps that include spatial and semantic information at different levels. This feature map is then processed by a feature fusion and adaptive receptive field module. A cross-scale feature fusion unit is used to transfer high-level semantic features to low-level features and low-level detail features to high-level features, and laterally connects and integrates features of different scales to obtain a cross-scale fused enhanced feature map. The first convolutional layer of the adaptive receptive field enhancement unit adjusts the channels of this feature map, and three max-pooling layers are used to extract features at different scales to obtain feature maps of different scales. The results are then fused using the second convolutional layer to obtain the final enhanced feature map. The feature map is then processed by a pixel-level localization mask prediction head using a small convolutional neural network, and the Sigmoid function is used to predict the probability that each pixel belongs to the precast pile region. Pixels with a probability greater than 0.5 are used as two-dimensional localization masks. The resolution of this mask is consistent with the resolution of the input image, and the position, shape, and contour of the precast pile in the image are identified.

[0020] Furthermore, in step S3, generating the cross-dimensional spatiotemporal alignment matrix includes the following:

[0021] A multi-dimensional, multimodal sensing device is used to photograph and scan a calibration target with known physical dimensions to obtain camera images and point clouds of the calibration target. The camera images are then subjected to distortion correction using the camera intrinsic parameter matrix and contrast enhancement to obtain processed camera images. Outlier filtering is applied to the point clouds, and motion distortion is compensated using IMU data to obtain processed point clouds.

[0022] 2D feature points are extracted from the processed camera image. 3D feature points are extracted from the processed point cloud using reflection intensity segmentation and RANSAC fitting. Matched feature point pairs are obtained through similarity matching.

[0023] Based on the matched feature point pairs, the initial extrinsic parameter matrix between the camera and the lidar is obtained using a closed-form solution.

[0024] A nonlinear optimization model is constructed, whose error loss function is a composite loss function of reprojection error and point-to-surface distance error. The initial extrinsic parameter matrix is ​​iterated a set number of times using the nonlinear optimization model to obtain the spatiotemporal alignment matrix.

[0025] The spatiotemporal alignment matrix is ​​fused with the timestamp offset to generate a 4×4 cross-dimensional spatiotemporal alignment matrix. This matrix is ​​then validated by point cloud projection to obtain the fusion result of the projected point cloud and the camera image. The validation is divided into static and dynamic tests: In static validation, the point cloud is projected onto the camera image. If the alignment accuracy between the object in the projected image and the edge of the projected point cloud is within 3 pixels, the validation is passed. In dynamic validation, the timing synchronization error is analyzed by tracking a moving target. If the error is within 100ms, the validation is passed.

[0026] Furthermore, in step S4, the precast pile target point cloud cluster obtained includes the following:

[0027] Parallel processing accelerated by CUDA (Compute Unified Device Architecture) transforms the 3D point cloud data in step S1 to the camera coordinate system through a cross-dimensional spatiotemporal alignment matrix. Perspective projection is performed based on the pinhole model to obtain the corresponding projection results. Bayesian fusion is performed on the projection results using a probabilistic depth filter to obtain image data with fused point cloud information. A 2D positioning mask is used as a guiding mask to filter out key points. Based on these key points, the projection effect of the precast pile point cloud is further optimized to obtain the target point cloud cluster of the precast pile.

[0028] Furthermore, in step S5, the length of the precast pile is obtained including the following:

[0029] Principal component analysis was used to extract the principal direction vector of the target point cloud cluster of the precast pile, and this vector was defined as the axial reference vector v of the pile body space. The RANSAC (Random Sample Consensus) algorithm was used to fit the central axis of the precast pile body.

[0030] When the angle between the fitted precast pile center axis and the vertical direction is greater than 3°, the point cloud cluster is aligned to the vertical coordinate system by a rotation matrix.

[0031] The maximum projection point P of the point cloud cluster along the v-axis onto the central axis of the precast pile body is calculated. max and the projected minimum point P min The Euclidean distance between the extreme points is calculated as the length L of the precast pile. The specific formula is as follows:

[0032] L = ||P max -P min ‖2;

[0033] The length and inclination of the precast piles are displayed in real time on the screen of the data processing module.

[0034] Furthermore, the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the real-time detection method for precast pile construction information based on cross-dimensional multimodal perception fusion.

[0035] Furthermore, the present invention also proposes a computer-readable storage medium storing a computer program, which is executed by a processor to perform the aforementioned method for real-time detection of precast pile construction information based on cross-dimensional multimodal perception fusion.

[0036] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects:

[0037] 1. This invention integrates cross-dimensional multimodal sensing devices with intelligent algorithms and ensures data consistency through spatiotemporal synchronous acquisition; the lightweight detection module based on deep learning and the inverse projection algorithm work together to realize automated and high-precision measurement of precast pile construction parameters, which has significant technical advantages and engineering value.

[0038] 2. This invention overcomes the limitations of the construction environment on detection accuracy. Through a multimodal data complementarity mechanism, it can maintain centimeter-level measurement accuracy even under direct sunlight, nighttime operation, or rainy / foggy weather.

[0039] 3. This invention, combined with an adaptive filtering algorithm, effectively suppresses vibration interference generated by pile driver operation, allowing the measurement process to proceed without interrupting construction and achieving true online real-time monitoring. Simultaneously, measurement data is uploaded to a cloud management platform in real-time via 4G / 5G networks, significantly improving the timeliness of construction quality control.

[0040] 4. This invention achieves centimeter-level detection accuracy and stable detection performance even in extreme construction environments through cross-modal collaboration of depth information and visual features. It deeply integrates the precise 3D geometric information provided by LiDAR with the rich texture features captured by the camera, enabling cross-modal collaboration between depth information and visual features. Furthermore, it boasts high detection accuracy and strong anti-interference capabilities, overcoming challenges such as complex lighting conditions, dust interference, and rain / fog obstruction at construction sites that could lead to single-modal failure.

[0041] 5. This invention converts 3D computation into 2D processing using a spatiotemporal alignment matrix, reducing computational complexity by two orders of magnitude. Simultaneously, it only requires processing key point cloud clusters after visual mask filtering, avoiding large-scale point cloud computing and improving on-site detection speed and efficiency. Attached Figure Description

[0042] Figure 1 This is a flowchart illustrating the overall implementation of the present invention.

[0043] Figure 2 This is a rear view of the housing of the multi-dimensional multimodal sensing device of the present invention.

[0044] Figure 3 This is the processed point cloud result image in the embodiment of the present invention.

[0045] Figure 4 This is a schematic diagram illustrating the acquisition of the cross-dimensional spatiotemporal alignment matrix according to the present invention.

[0046] Figure 5 This is a projection effect diagram of static verification in an embodiment of the present invention.

[0047] Figure 6 This is a projection diagram of the static verification of the precast pile construction site in an embodiment of the present invention. Detailed Implementation

[0048] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0049] To achieve the above objectives, this invention proposes a real-time detection method for precast pile construction information based on cross-dimensional multimodal perception fusion, such as... Figure 1 As shown, the specific steps are as follows:

[0050] S1. Construct a multi-dimensional, multi-modal sensing device for the precast pile construction site. This device includes a casing, an image acquisition module, a point cloud acquisition module, a data processing module, and a power supply. The device is deployed on a tripod at a distance of 4m from the pile driving point to be inspected and 1m above the ground. Throughout the inspection process, the multi-dimensional, multi-modal sensing device must be kept directly facing the precast pile, and obstructions to the pile should be avoided to prevent affecting the accuracy and effectiveness of the inspection.

[0051] Among them, such as Figure 2 As shown, the outer shell is a rigid cuboid with a hollow internal structure. The interior of the outer shell is divided into upper and lower parts by a horizontal partition. The image acquisition module is fixed to the upper part of the outer shell with screws. The point cloud acquisition module is fixed to the lower part of the outer shell with screws. The data processing module is located inside the outer shell.

[0052] The image acquisition module is connected to the data processing module via a USB data cable, the point cloud acquisition module is connected to the data processing module via a USB data cable, the point cloud acquisition module is connected to the power supply via a first power cable, and the data processing module is connected to the power supply via a second power cable.

[0053] The image acquisition module includes a camera and a USB data cable, used to acquire two-dimensional image data of precast piles at the construction site. The camera must be a fixed-focus lens, including a 2-megapixel industrial lens.

[0054] The point cloud acquisition module includes a lidar, a first data cable, and a first power cable, and is used to acquire 3D point cloud data corresponding to images of precast piles at the construction site. The lidar can be a mechanical lidar or a solid-state lidar.

[0055] The data processing module includes an industrial motherboard, a display screen, a second data cable, and a second power cable. It receives data from the image acquisition module and the point cloud acquisition module through the second data cable, processes the data, and obtains the point cloud cluster of the precast pile target.

[0056] By maintaining a time deviation of less than 1ms between camera exposure and lidar scanning, and by adding a unified timestamp to the collected data, spatiotemporal cross-dimensional reconstruction of the precast pile construction site can be achieved.

[0057] S2. In the data processing module, a deep learning-based two-dimensional lightweight detection model for precast piles is used to process the data acquired in the image acquisition module, generating a two-dimensional positioning mask with pixel-level accuracy for the precast piles. Specifically:

[0058] The deep learning-based two-dimensional lightweight detection model for precast piles includes a lightweight feature extraction backbone network, a feature fusion and adaptive receptive field module, and a pixel-level localization mask prediction head.

[0059] The lightweight feature extraction backbone network includes depthwise separable convolutional layers; the feature fusion and adaptive receptive field module includes a cross-scale feature fusion unit and an adaptive receptive field enhancement unit; the adaptive receptive field enhancement unit includes a first 1×1 convolutional layer, three max pooling layers and a second 1×1 convolutional layer connected in sequence.

[0060] The two-dimensional image data from step S1 is input into a deep learning-based two-dimensional lightweight detection model for precast piles. After passing through a lightweight feature extraction backbone network, feature extraction is performed using depthwise separable convolutional layers to obtain multi-scale feature maps containing spatial and semantic information at different levels. These feature maps are then processed by a feature fusion and adaptive receptive field module. A cross-scale feature fusion unit transfers deeper (higher) semantic features (such as object categories) closer to the network output from the multi-scale feature map to shallower (lower) features closer to the network input, while transferring lower-level detailed features (such as object edges and corners) to higher levels. Furthermore, features at different scales are horizontally connected and integrated, enhancing the spatial detail information of shallow features and the semantic information expression capability of deep features, particularly strengthening the representation of precast pile edges. The key regions, such as edges and ends, are characterized to obtain enhanced feature maps fused across scales. The first convolutional layer of the adaptive receptive field enhancement unit adjusts the channels of this feature map, and three max pooling layers are used to extract features at different scales to obtain feature maps of different scales. The results are then fused using the second convolutional layer to obtain the final enhanced feature map with adaptive context awareness. The feature map is then processed by a pixel-level localization mask prediction head using a small convolutional neural network that includes transposed convolution or upsampling operations. The Sigmoid function is used to predict the probability that each pixel belongs to the precast pile region. Pixels with a probability greater than 0.5 are used as a two-dimensional localization mask with pixel-level accuracy for the precast pile. The resolution of this mask is consistent with the resolution of the input image, which identifies the position, shape, and contour of the precast pile in the image.

[0061] S3. Generate a cross-dimensional spatiotemporal alignment matrix based on rigid coupling calibration between the camera and LiDAR. Specifically:

[0062] like Figure 3 As shown, a multi-dimensional multimodal sensing device is used to capture and scan a calibration target of known physical size to obtain camera images and point clouds of the calibration target; distortion correction of the camera images is performed using the camera intrinsic parameter matrix, and contrast is enhanced by histogram equalization or HDR fusion to obtain the processed camera images; outlier filtering is performed on the point clouds, and motion distortion is compensated by IMU data to obtain the processed point clouds.

[0063] 2D feature points in the processed camera image are extracted using pixel-level corner detection or deep learning keypoint localization. 3D feature points in the processed point cloud are extracted using reflection intensity segmentation and RANSAC fitting. Similarity matching is used to ensure spatial consistency across modal data and obtain matched feature point pairs.

[0064] like Figure 4 As shown, based on the matched feature point pairs, the initial extrinsic parameter matrix between the camera and the lidar is obtained using a closed-form solution.

[0065] A nonlinear optimization model is constructed, whose error loss function is a composite loss function of reprojection error (weight 0.6) and point-to-surface distance error (weight 0.4). The initial extrinsic parameter matrix is ​​iteratively optimized 50 times using this model to obtain a high-precision spatiotemporal alignment matrix. This high-precision spatiotemporal alignment matrix is ​​then fused with the timestamp offset to generate a 4×4 cross-dimensional spatiotemporal alignment matrix.

[0066] Figure 4 In this context, CAMERA represents the camera coordinate system, LIDAR represents the lidar coordinate system, and the combination of R and T is the desired cross-dimensional spatiotemporal alignment matrix, which represents the transformation relationship required from the lidar coordinate system to the camera coordinate system.

[0067] Point cloud projection verification was performed on the cross-dimensional spatiotemporal alignment matrix to obtain the fusion result of the projected point cloud and the camera image. The verification consisted of static and dynamic tests: in static verification, the point cloud was projected onto the camera image. If the alignment accuracy between the object in the projected image and the edge of the projected point cloud was within 3 pixels, the verification was passed, and the indoor projection effect was as follows: Figure 5 As shown; in dynamic verification, the timing synchronization error is analyzed by tracking the moving target. The error varies depending on the camera shutter time of the captured image. If the error is within 100ms, the verification is passed.

[0068] Figure 6 The results of the point cloud projection verification on site show that the scanned point cloud of the precast piles all landed on the pile body, indicating that the calibration matrix is ​​effective.

[0069] S4. Using the inverse projection-driven algorithm, the data collected in the point cloud acquisition module is mapped to the camera coordinate system based on the cross-dimensional spatiotemporal alignment matrix. Point cloud features are then filtered using visual positioning information to obtain the precast pile target point cloud cluster, completing the processing in the data processing module. Specifically:

[0070] Parallel processing accelerated by CUDA (Compute Unified Device Architecture) transforms the 3D point cloud data in step S1 to the camera coordinate system through a cross-dimensional spatiotemporal alignment matrix. Perspective projection is then performed based on the pinhole model to obtain the corresponding projection results. Bayesian fusion of the projection results is performed using a probabilistic depth filter to improve the mapping accuracy of the object edge region after projection, resulting in image data with fused point cloud information. A 2D positioning mask is used as a guiding mask to filter out key points with high visual-geometric consistency, such as the top and outline of the precast pile. Based on these key points, the projection effect of the precast pile point cloud is further optimized to obtain the target point cloud cluster of the precast pile.

[0071] S5. Perform spatial vector analysis on the target point cloud cluster of the precast pile to obtain the pile length. Specifically:

[0072] Principal component analysis was used to extract the principal direction vector of the target point cloud cluster of the precast pile, and this vector was defined as the axial reference vector v of the pile body space. The RANSAC (Random Sample Consensus) algorithm was used to fit the central axis of the precast pile body to eliminate noise interference caused by surface attachments or local point cloud missingness.

[0073] When the angle between the fitted precast pile center axis and the vertical direction is greater than 3°, the point cloud cluster is aligned to the vertical coordinate system by a rotation matrix to eliminate the measurement error caused by the pile driving angle.

[0074] The maximum projection point P of the point cloud cluster along the v-axis onto the central axis of the precast pile body is calculated. max and the projected minimum point P min The Euclidean distance between the extreme points is calculated as the length L of the precast pile. The specific formula is as follows:

[0075] L = ||P max -P min ‖2;

[0076] The length and inclination of the precast piles are displayed in real time on the screen of the data processing module.

[0077] This invention integrates cross-dimensional multimodal sensing devices with intelligent algorithms. At the hardware level, a dual-modal sensor integrated with a rigid shell ensures data consistency through spatiotemporal synchronous acquisition. At the software level, a lightweight detection module based on deep learning and an inverse projection algorithm work together to achieve automated and high-precision measurement of precast pile construction parameters, which has significant technical advantages and engineering value.

[0078] For camera image recognition, this invention innovatively introduces a LiDAR device to construct a multimodal fusion detection system. Through cross-modal collaboration of depth information and visual features, it deeply fuses the precise three-dimensional geometric information provided by the LiDAR with the rich texture features captured by the camera, achieving centimeter-level detection accuracy and stable detection performance even in extreme construction environments. Simultaneously, it boasts high detection accuracy and strong anti-interference capabilities, adapting to the challenges of single-modal failure caused by complex lighting conditions, dust interference, and rain / fog obstruction at construction sites.

[0079] Meanwhile, for point cloud recognition, segmentation, and detection, the inverse projection-driven algorithm proposed in this invention converts 3D computation into 2D processing through a spatiotemporal alignment matrix, reducing computational complexity by two orders of magnitude. Furthermore, the system only needs to process key point cloud clusters after visual mask filtering, avoiding large-scale point cloud computing compared to traditional full point cloud processing methods, thus improving on-site detection speed and efficiency.

[0080] This invention also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. It should be noted that when the processor executes the computer program, it corresponds to the specific steps of the method provided in this invention, possessing the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in this invention.

[0081] This invention also proposes a computer-readable storage medium storing a computer program. It should be noted that when the computer program is executed by a processor, it corresponds to the specific steps of the method provided in this invention, possessing the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in this invention.

[0082] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for real-time detection of precast pile construction information based on cross-dimensional multimodal perception fusion, characterized in that, include: S1. Construct a multi-dimensional multimodal sensing device for the precast pile construction site. The device includes a shell, an image acquisition module, a point cloud acquisition module, a data processing module, and a power supply. Set up the device at the precast pile construction site. S2. In the data processing module, a two-dimensional lightweight detection model for precast piles based on deep learning is used to process the data collected in the image acquisition module and generate a two-dimensional positioning mask. S3. Generate a cross-dimensional spatiotemporal alignment matrix based on rigid coupling calibration of camera and lidar; S4. Using the reverse projection driving algorithm, the data collected in the point cloud acquisition module is mapped to the camera coordinate system according to the cross-dimensional spatiotemporal alignment matrix. The point cloud features are filtered through visual positioning information to obtain the precast pile target point cloud cluster, thus completing the processing in the data processing module. S5. Perform spatial vector analysis on the target point cloud cluster of the precast pile to obtain the length of the precast pile.

2. The method for real-time detection of precast pile construction information based on cross-dimensional multimodal perception fusion according to claim 1, characterized in that, In step S1, the outer shell is a rigid cuboid with a hollow internal structure. The interior of the outer shell is divided into upper and lower parts by a horizontal partition. The image acquisition module is fixed to the upper part of the outer shell with screws. The point cloud acquisition module is fixed to the lower part of the outer shell with screws. The data processing module is located inside the outer shell. The image acquisition module is connected to the data processing module via a USB data cable, the point cloud acquisition module is connected to the data processing module via a USB data cable, the point cloud acquisition module is connected to the power supply via a first power cable, and the data processing module is connected to the power supply via a second power cable. The image acquisition module includes a camera and a USB data cable, used to acquire two-dimensional image data of precast piles at the construction site; The point cloud acquisition module includes a lidar, a first data cable, and a first power cable, and is used to acquire 3D point cloud data corresponding to images of precast piles at the construction site. The data processing module includes an industrial motherboard, a display screen, a second data cable, and a second power cable. It receives data from the image acquisition module and the point cloud acquisition module through the second data cable, processes the data, and obtains the point cloud cluster of the precast pile target.

3. The real-time detection method for precast pile construction information based on cross-dimensional multimodal perception fusion according to claim 2, characterized in that, In step S2, generating the two-dimensional positioning mask includes the following: The deep learning-based two-dimensional lightweight detection model for precast piles includes a lightweight feature extraction backbone network, a feature fusion and adaptive receptive field module, and a pixel-level localization mask prediction head. The lightweight feature extraction backbone network includes depthwise separable convolutional layers; the feature fusion and adaptive receptive field module includes a cross-scale feature fusion unit and an adaptive receptive field enhancement unit. The adaptive receptive field enhancement unit consists of a first 1×1 convolutional layer, three max pooling layers, and a second 1×1 convolutional layer connected in sequence. The two-dimensional image data in step S1 is input into a two-dimensional lightweight detection model for precast piles based on deep learning. After passing through a lightweight feature extraction backbone network, feature extraction is performed using depthwise separable convolutional layers to obtain multi-scale feature maps that include spatial and semantic information at different levels. This feature map is then processed by a feature fusion and adaptive receptive field module. The cross-scale feature fusion unit transfers high-level semantic features to low-level features and low-level detail features to high-level features, and horizontally connects and integrates features of different scales to obtain a cross-scale fused enhanced feature map. The first convolutional layer of the adaptive receptive field enhancement unit adjusts the channels of this feature map, and three max pooling layers are used to extract features at different scales to obtain feature maps of different scales. The results are then fused using the second convolutional layer to obtain the final enhanced feature map. The feature map is processed by a pixel-level localization mask prediction head and a small convolutional neural network. The Sigmoid function is used to predict the probability that each pixel belongs to the precast pile area. Pixels with a probability greater than 0.5 are used as two-dimensional localization masks. The resolution of the mask is consistent with the resolution of the input image, and the position, shape and contour of the precast pile in the image are identified.

4. The method for real-time detection of precast pile construction information based on cross-dimensional multimodal perception fusion according to claim 2, characterized in that, In step S3, generating the cross-dimensional spatiotemporal alignment matrix includes the following: A multi-dimensional multimodal sensing device is used to photograph and scan a calibration target with known physical dimensions to obtain camera images and point clouds of the calibration target; distortion correction of the camera images is performed using the camera intrinsic parameter matrix, and then the contrast is enhanced to obtain the processed camera images; outlier filtering is performed on the point clouds, and motion distortion is compensated using IMU data to obtain the processed point clouds. 2D feature points are extracted from the processed camera image. 3D feature points are extracted from the processed point cloud using reflection intensity segmentation and RANSAC fitting. Matched feature point pairs are obtained through similarity matching. Based on the matched feature point pairs, the initial extrinsic parameter matrix between the camera and the lidar is obtained using a closed-form solution method. A nonlinear optimization model is constructed, whose error loss function is a composite loss function of reprojection error and point-to-surface distance error. The initial extrinsic parameter matrix is ​​iterated a set number of times using the nonlinear optimization model to obtain the spatiotemporal alignment matrix. The spatiotemporal alignment matrix is ​​fused with the timestamp offset to generate a 4×4 cross-dimensional spatiotemporal alignment matrix. This matrix is ​​then validated by point cloud projection to obtain the fusion result of the projected point cloud and the camera image. The validation is divided into static and dynamic tests: In static validation, the point cloud is projected onto the camera image. If the alignment accuracy between the object in the projected image and the edge of the projected point cloud is within 3 pixels, the validation is passed. In dynamic validation, the timing synchronization error is analyzed by tracking a moving target. If the error is within 100ms, the validation is passed.

5. The method for real-time detection of precast pile construction information based on cross-dimensional multimodal perception fusion according to claim 2, characterized in that, In step S4, the target point cloud cluster of the precast piles obtained includes the following: The 3D point cloud data in step S1 is transformed to the camera coordinate system using a cross-dimensional spatiotemporal alignment matrix through parallel processing accelerated by CUDA. Perspective projection is performed based on the pinhole model to obtain the corresponding projection results. Bayesian fusion is performed on the projection results using a probabilistic depth filter to obtain image data with fused point cloud information. A 2D positioning mask is used as a guiding mask to filter out key points. Based on these key points, the projection effect of the precast pile point cloud is further optimized to obtain the target point cloud cluster of the precast pile.

6. The method for real-time detection of precast pile construction information based on cross-dimensional multimodal perception fusion according to claim 1, characterized in that, In step S5, the length of the precast pile is obtained including the following: The principal direction vector of the target point cloud cluster of the precast pile is extracted using principal component analysis, and this vector is defined as the spatial axial reference vector v of the pile body; the central axis of the precast pile body is fitted using the RANSAC algorithm. When the angle between the fitted precast pile center axis and the vertical direction is greater than 3°, the point cloud cluster is aligned to the vertical coordinate system by a rotation matrix. The maximum projection point P of the point cloud cluster along the v-axis onto the central axis of the precast pile body is calculated. max and the projected minimum point P min The Euclidean distance between the extreme points is calculated as the length L of the precast pile. The specific formula is as follows: L=‖P max -P min ‖2; The length and inclination of the precast piles are displayed in real time on the screen of the data processing module.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the real-time detection method for precast pile construction information based on cross-dimensional multimodal perception fusion as described in any one of claims 1 to 6.

8. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by the processor, performs the real-time detection method for precast pile construction information based on cross-dimensional multimodal perception fusion as described in any one of claims 1 to 6.