A crane point cloud completion method based on multi-modal data
Patent Information
- Application Number
- CN202610666182.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-09-15
AI Technical Summary
1、以充分利用图像与激光点云的互补信息,增强对吊臂、支腿及连接节点等关键结构区域的特征表达能力,并提高补全结果的完整性、连续性和空间分布均匀性,从而为后续起重机姿态估计、关键参数计算及倾覆风险评估提供可靠的三维数据基础。
Smart Images

Figure CN122760901A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a crane point cloud completion method, and more particularly to a crane point cloud completion method based on multimodal data. Background Technology
[0002] With the rapid advancement of infrastructure construction, engineering projects have broadly covered multiple fields such as highways, urban roads, rail transit, bridges, tunnels, and ports. These projects are typically large-scale and technically complex, placing higher demands on the reliability, flexibility, and operational efficiency of construction machinery. Against this backdrop, lifting machinery, as a key auxiliary equipment, plays an irreplaceable and vital role. Among them, mobile truck cranes, due to their reasonable structure and convenient operation, are widely used in various construction scenarios. Truck cranes mount the crane body on a general-purpose or special-purpose vehicle chassis, possessing excellent mobility and traversability, allowing them to move flexibly on urban roads, construction sites, and trackless surfaces. Therefore, truck cranes have become an important supporting equipment for promoting the intelligent, modular, and efficient development of the construction process.
[0003] However, with the expansion of crane applications and the increase in usage frequency, the difficulty of safety management also increases. Therefore, accurately calculating the attitude parameters of cranes during operation and designing reasonable stability assessment methods have become key issues in achieving safety monitoring.
[0004] Existing crane monitoring methods mainly include traditional monitoring technologies, IoT-based monitoring technologies, and computer vision-based monitoring technologies. Traditional monitoring technologies typically measure parameters such as lifting weight, lever arm amplitude, and lifting height by installing weight sensors, angle sensors, or absolute encoders on key components such as the hook and boom, thereby achieving operational status monitoring. However, these methods generally suffer from poor portability, difficulties in periodic equipment calibration and adjustment, complex management, and are susceptible to installation errors, sensor drift, and environmental interference.
[0005] IoT-based monitoring technology primarily achieves remote monitoring of crane operating status by working in conjunction with crane torque sensors. Compared to traditional monitoring methods, this technology has certain advantages in remote data acquisition, status management, and collaborative monitoring. However, it still relies on traditional sensors, thus limiting system adaptability and monitoring reliability.
[0006] Computer vision-based monitoring technology can acquire image information from construction sites through visual sensors. Compared with traditional torque or IoT monitoring methods, it can obtain richer spatial information and provide intuitive and visual analysis of on-site operations. Currently, this type of method is mainly used for worker safety identification, load identification, and equipment status visualization, but its application in the overall stability monitoring of cranes remains relatively limited.
[0007] Overall, current crane stability analysis still primarily relies on onboard sensors or IoT remote monitoring, but these methods have limitations in engineering applications. To avoid issues such as sensor installation errors, drift calibration difficulties, and manual sensor adjustments, a non-contact sensing method that integrates vision and LiDAR offers significant advantages for crane overturning monitoring. In this type of non-contact sensing method, sparse point cloud completion is a crucial foundation for subsequent point cloud extraction of key components, 3D parameter calculation, and stability assessment. Since single-frame LiDAR point clouds are often affected by occlusion, sparse sampling, and noise interference, resulting in structural gaps and local incompleteness, it is necessary to research point cloud completion methods suitable for crane scenarios to improve the completeness and reliability of 3D perception results.
[0008] Sparse laser point cloud completion methods are mainly divided into traditional point cloud completion methods, deep learning-based single-modal point cloud completion methods, and multi-modal point cloud completion methods. Traditional point cloud completion methods mainly include geometry-based completion methods and template-based completion methods. Geometry-based methods typically utilize the visible parts of the point cloud and predict the missing parts through local interpolation or prior geometric assumptions; template-based methods restore the missing regions by matching the incomplete point cloud with complete models in the database. These methods have certain feasibility in theoretical research and small-scale data scenarios, but they still have obvious limitations: geometric methods can usually only handle small-scale missing areas and have limited ability to complete large-area missing areas or complex-shaped point clouds; template methods rely on large-scale databases, have high matching and optimization costs, and are sensitive to noise, making it difficult to meet the efficiency and robustness requirements of complex engineering scenarios.
[0009] With the development of deep learning, network-based unimodal point cloud completion methods have gradually become the mainstream research approach. These methods directly encode features and decode structures from incomplete point clouds to achieve overall shape restoration and local detail completion. Compared to traditional methods, unimodal deep learning methods have advantages in completion accuracy and expressive power. However, when dealing with sparse, severely occluded, or large missing point clouds, they still tend to suffer from insufficient restoration of local structural details, incomplete overall shape, and uneven point cloud distribution. Especially for targets with elongated structures and complex connectivity, unimodal point cloud features often fail to provide sufficient prior constraints, limiting the completion performance.
[0010] To fully utilize the texture and structural information provided by images, researchers have further proposed image-guided multimodal point cloud completion methods. These methods, by fusing point cloud and image features, improve the recovery of missing regions to some extent and enhance the overall structural reconstruction quality. However, existing multimodal point cloud completion methods still suffer from some common problems: significant differences in features between different modalities make direct fusion prone to introducing redundant information and noise; a balance between local detail recovery and global structural consistency is often difficult to achieve; and in scenarios with sparse point clouds, severely occluded regions, and complex, slender components, inaccurate structural recovery, discontinuous details, and uneven spatial distribution are still likely to occur.
[0011] Furthermore, most existing multimodal point cloud completion methods are designed for public datasets or general object categories, lacking specific research for engineering equipment such as truck cranes with long straight booms, outrigger structures, and complex connection nodes. Crane point clouds are characterized by obvious slender component features, large local structural differences, and severe occlusion, making it difficult for existing methods to achieve high-precision, structurally complete point cloud reconstruction in such scenarios. Summary of the Invention
[0012] The purpose of this invention is to provide a crane point cloud completion method based on multimodal data, so as to make full use of the complementary information of image and laser point cloud, enhance the feature representation ability of key structural areas such as boom, outriggers and connecting nodes, and improve the completeness, continuity and spatial distribution uniformity of the completion results, thereby providing a reliable three-dimensional data foundation for subsequent crane attitude estimation, key parameter calculation and overturning risk assessment.
[0013] The objective of this invention can be achieved through the following technical solutions: A crane point cloud completion method based on multimodal data includes: Step S1: Acquire images and sparse laser point clouds of the crane operation scene; Step S2: Based on sparse laser point clouds, multi-scale hierarchical features of the point cloud are extracted using hybrid sampling and local feature aggregation. Step S3: Extract multi-scale hierarchical features from images of crane operation scenarios; Step S4: Based on the multi-scale hierarchical features of the point cloud and the multi-scale hierarchical features of the image, the multi-scale hierarchical features of the point cloud after cross-modal fusion are obtained by utilizing cross-modal attention and feature transfer constraints. Step S5: Based on the point cloud features after cross-modal fusion, a coarsely completed point cloud is obtained by using multi-branch mapping and adaptive weight fusion; Step S6: Obtain the completed point cloud based on the coarsely completed point cloud and the sparse laser point cloud.
[0014] Step S1 also includes: preprocessing the image of the crane operation scene and the sparse laser point cloud.
[0015] Step S2 includes: Step S2-1: Obtain the set of downsampling center points using hybrid sampling; Step S2-2: Assign weights to each point in the set of downsampled center points through structure-aware sampling; Step S2-3: Select a local neighborhood for each point in the downsampled center point set, and fuse the center point features with the relative positions and feature information of the neighborhood points to form the second feature of each point in the downsampled center point set: in: For point The second characteristic, For point The second characteristic, For point The neighborhood, It is a multilayer perceptron. This is a pooling operation; Step S2-4: Repeat steps S2-1 to S2-3 multiple times to obtain multi-scale hierarchical features of the point cloud.
[0016] Step S2-1 includes: Step S2-1-1: Generate a set of spatially uniformly distributed center points by sampling from the farthest point, and expand it to obtain a set of basic sampling points; Step S2-1-2: For each point in the basic sampling point set, construct its neighborhood using the KNN algorithm to obtain the neighborhood relative coordinate distribution matrix: in: For point The neighborhood relative coordinate distribution matrix, for midpoint; Step S2-1-3: Construct the local covariance matrix: in: for The local covariance matrix, k The number of points in the neighborhood; Step S2-1-4: Perform eigenvalue decomposition on the covariance matrix to obtain the three first eigenvalues arranged in ascending order. Define local curvature and linearity indices based on the eigenvalue relationships, and further obtain the structural score. in: for Structural scoring, for The local curvature index, As the weight of the local curvature index, for linearity index, The weights of the linearity index, , , These are the three first eigenvalues arranged in ascending order; Step S2-1-5: Strengthen the weight distribution of key structural regions through an exponential modulation mechanism; Step S2-1-6: Obtain the structural score distribution based on the structural score after strengthening the weight distribution of the key structural regions, and perform probability sampling based on the normalized structural score distribution to obtain the set of structural guidance sampling points. The union of the set of center points and the set of structural guidance sampling points is used as the set of downsampling center points.
[0017] Step S3 includes: Step S3-1: Use a pre-trained first convolutional neural network to extract multi-scale feature maps of different levels of the image; Step S3-2: Convert the multi-scale feature map into a serialized token; Step S3-3: Integrate the tokens from different levels into a multi-scale visual token sequence as a multi-scale hierarchical feature of the image.
[0018] In step S4, a cross-modal attention module is constructed using point cloud multi-scale hierarchical features as queries and image multi-scale hierarchical features as keys and values. Image structural information is introduced into the point cloud representation through a cross-attention mechanism to obtain cross-modal fused point cloud multi-scale hierarchical features. The feature transfer constraint minimizes the cross-modal feature transfer loss: in: For cross-modal feature transfer loss, This is the initial migration loss. To maintain constraints on content, For the first l The first point cloud feature at layer -1 For the first l The first point cloud feature of the layer, For Gram matrices, For the first l The first image feature of the layer, the multi-scale hierarchical feature of the point cloud is composed of the first point cloud features of all layers, and the multi-scale hierarchical feature of the image is composed of the first image features of all layers.
[0019] Step S5 includes: Step S5-1: Apply multiple independent 1D convolutional layer branches to the multi-scale hierarchical features of the point cloud after cross-modal fusion for dimensionality upscaling; Step S5-2: Reshape the result of each branch after dimensionality increase into point cloud coordinate format to obtain the original branch point cloud corresponding to each branch; Step S5-3: Obtain the weighted branch point cloud based on each original branch point cloud and the branch weight matrix; Step S5-4: Summarize the weighted branch point clouds of all branches along the point dimension to obtain a coarsely completed point cloud.
[0020] Step S6 includes: Step S6-1: Stitch the coarsely completed point cloud with representative observation points to obtain a coarse-grained point cloud: in: For coarse-grained point clouds, To roughly complete the point cloud, The representative observation points were obtained by sampling the farthest point from the sparse laser point cloud. Step S6-2: Construct a dynamic KNN neighborhood for each center point in the coarse-grained point cloud; Step S6-3: Further construct the local relation descriptor: in: The center point in a coarse-grained point cloud The neighborhood, for neighborhood points, for and Local relation descriptors, Describe neighborhood points Relative to the center point Local geometric offset; Step S6-4: Extract local edge features from the multi-layer convolutional mapping network that shares the input of the local relation descriptors. ; Step S6-5: Based on all local edge features of the center point, calculate the max pooling feature and average pooling feature respectively; Step S6-6: Based on the max pooling feature and average pooling feature of each center point, obtain the fused local features: in: Center point The fusion of local features, Center point Max pooling characteristics, Center point Average pooling characteristics This is a channel fusion mapping based on 1×1 convolution; Step S6-7: Based on the fusion of local features, predict the three-dimensional residual offset of each point, and update the coordinates of the center point using the three-dimensional residual offset; Step S6-8: Use all the updated center points as the completed point cloud.
[0021] A crane point cloud completion device based on multimodal data includes a memory, a processor, and a program stored in the memory. When the processor executes the program, it implements the method described above.
[0022] A storage medium having a program stored thereon, which, when executed, implements the method described above.
[0023] Compared with the prior art, the present invention has the following beneficial effects: 1. By fully utilizing the complementary information of images and laser point clouds, the feature representation ability of key structural areas such as booms, outriggers and connecting nodes is enhanced, and the completeness, continuity and spatial distribution uniformity of the completion results are improved, thereby providing a reliable three-dimensional data foundation for subsequent crane attitude estimation, key parameter calculation and overturning risk assessment.
[0024] 2. Effectively eliminates noise, distortion, and calibration errors in the original data acquisition process, improves the quality and consistency of input data, provides a more reliable input basis for subsequent feature extraction and completion, and reduces error propagation.
[0025] 3. Taking into account both global structure and local details, features are gradually abstracted through multi-level downsampling, which reduces computational complexity and enhances the ability to represent complex geometries such as the slender structure and connection nodes of cranes, avoiding the loss of key information by single-scale features.
[0026] 4. Targeted enhancement of sampling weights for key structural areas such as booms and outriggers to address the problem that slender components in crane point clouds are easily ignored by traditional uniform sampling, thereby improving the ability to represent the features of important structures and the accuracy of completion.
[0027] 5. Fully utilize the general visual semantic information of the pre-trained model, while retaining image semantics and texture details at different levels through multi-scale features, providing richer and more hierarchical image priors for cross-modal fusion and compensating for the lack of sparsity in point clouds.
[0028] 6. Accurately align image and point cloud features, effectively transferring structural and texture information from the image to the point cloud representation; at the same time, content constraints prevent the loss of the original point cloud geometric structure during the fusion process, balancing cross-modal consistency and geometric realism, and reducing redundant information and noise interference.
[0029] 7. Different branches can focus on the structural features of the crane at different scales, including the overall outline and local connections. The contribution of each branch is dynamically adjusted by adaptive weights to improve the coarse completion of the point cloud to adapt to complex shapes, making the completion result more complete and the distribution more uniform.
[0030] 8. By modeling local geometric relationships, the position of the point cloud is finely adjusted, effectively repairing the local discontinuities and structural deviations in the coarse completion stage; combined with representative observation points sampled from the farthest point, the completion result is ensured to be strictly consistent with the original sparse point cloud in space, improving the geometric accuracy and continuity of the final point cloud. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the main steps of the method of the present invention; Figure 2 This is a schematic diagram of the technical route of the present invention; Figure 3 A framework diagram for point cloud encoding; Figure 4 This is a diagram showing the sampling results of a hybrid structure with three-layer downsampling. Figure 5 This is a schematic diagram of hierarchical feature sampling and aggregation; Figure 6 This is a schematic diagram of a multi-branch structure aggregation and residual refinement decoder; Figure 7 This is a schematic diagram of max pooling and average pooling; Figure 8 A qualitative analysis of the point cloud completion methods on the simulation dataset was conducted. Figure 9 The qualitative analysis results compare point cloud completion methods on real datasets. Detailed Implementation
[0032] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0033] To address the challenges of sparse point clouds, missing local structures, and difficulties in multimodal information coordination in crane scenarios, this invention proposes a Structure-Guided Multimodal Completion Network (SGMCNet). This network uses both image and laser point cloud as bimodal inputs and achieves step-by-step reconstruction of dense 3D structures in complex crane scenes through structure-aware encoding, cross-modal information fusion, and geometric constraint decoding. The network consists of a multimodal structure-aware encoding module, a cross-modal information interaction fusion module, and a multi-branch structure aggregation and residual refinement decoding module. During the input phase, the image is convolved to form visual features that describe the crane outline and structural edges; the point cloud uses 3D coordinates and its local geometric neighborhood features as input, providing spatial information for structural modeling.
[0034] A crane point cloud completion method based on multimodal data, such as Figure 1 and Figure 2 As shown, it includes: Step S1: Acquire images and sparse laser point clouds of the crane operation scene; The images of the crane operation scene are acquired using an RGB camera, and the sparse laser point cloud is acquired using a lidar. Step S1 also includes: preprocessing the images of the crane operation scene and the sparse laser point cloud to remove outlier noise points and retain the three-dimensional coordinates and local neighborhood features.
[0035] Step S2: Based on sparse laser point clouds, multi-scale hierarchical features of the point cloud are extracted using a hybrid sampling and local feature aggregation method, including: Step S2-1: Hybrid sampling is used to obtain the set of downsampled center points. Although FPS can guarantee good spatial coverage, it only samples based on distance distribution and does not consider the differences in the complexity and importance of local geometric structures in the point cloud. In the crane scene, a large number of point cloud regions correspond to flat structural surfaces or background areas, while key structural information is mainly concentrated in long straight booms, joint connections, and geometric abrupt changes. The uniform sampling mechanism tends to allocate too many sampling resources to low-information areas, thereby weakening the expressive power of the key structural skeleton. The above problem is particularly obvious in single-frame sparse crane point cloud scenes. Affected by occlusion effects and ranging errors, the limited sampling budget is more likely to be allocated to areas with insufficient structural information, further compressing the representation ratio of key components in the hierarchical abstraction process.
[0036] To address the aforementioned shortcomings, this application provides a structure-guided hierarchical point cloud encoder. While maintaining the global coverage advantage of FPS, it introduces a local geometric structure awareness mechanism to achieve adaptive sampling enhancement for structurally significant regions. The overall framework diagram of the point cloud encoder is shown below. Figure 3 As shown, it includes: Step S2-1-1: Generate a set of spatially uniformly distributed center points by sampling from the farthest point, and expand it to obtain a set of basic sampling points; Step S2-1-2: For each point in the basic sampling point set, construct its neighborhood using the KNN algorithm to obtain the neighborhood relative coordinate distribution matrix: in: For point The neighborhood relative coordinate distribution matrix, for midpoint; Step S2-1-3: Construct the local covariance matrix: in: for The local covariance matrix, k The number of points in the neighborhood; Step S2-1-4: Perform eigenvalue decomposition on the covariance matrix to obtain the three first eigenvalues arranged in ascending order. Define local curvature and linearity indices based on the eigenvalue relationships, and further obtain the structural score. in: for Structural scoring, for The local curvature index, As the weight of the local curvature index, for linearity index, The weights of the linearity index, , , These are the three first eigenvalues arranged in ascending order; Step S2-1-5: Strengthen the weight distribution of key structural regions through an exponential modulation mechanism; Step S2-1-6: Obtain the structural score distribution based on the structural score after strengthening the weight distribution of the key structural regions, and perform probability sampling based on the normalized structural score distribution to obtain the set of structural guidance sampling points. The union of the set of center points and the set of structural guidance sampling points is used as the set of downsampling center points.
[0037] Step S2-2: Assign weights to each point in the set of downsampled center points through structure-aware sampling; Step S2-3: Select a local neighborhood for each point in the downsampled center point set, and fuse the center point features with the relative positions and feature information of the neighborhood points to form the second feature of each point in the downsampled center point set: in: For point The second characteristic, For point The second characteristic, For point The neighborhood, It is a multilayer perceptron. This is a pooling operation; Step S2-4: Repeat steps S2-1 to S2-3 multiple times to obtain multi-scale hierarchical features of the point cloud.
[0038] Figure 4 This demonstrates the three-layer downsampling results of the sampling encoder on a single-frame input point cloud. The first layer of sampling, based on a strategy using FPS and a structure-aware candidate set, ensures global spatial coverage while highlighting local geometric key points, such as edges, sharp corners, and rod-like structures, with red dots. The second layer of sampling further downsamples based on the first layer, reducing redundant points while preserving local key structural information, making key geometric points more concentrated. The third layer of sampling, the final token layer, has the fewest points and primarily covers the most significant structural key points, providing an efficient representation for subsequent feature encoding and multimodal fusion.
[0039] Through the step-by-step processing of multi-layer convolution and pooling, a hierarchical abstract representation of local geometric features can be achieved. The number of sampling points is gradually reduced during each downsampling process, while retaining the information of key structural points, thereby realizing a hierarchical feature representation from local to global.
[0040] A schematic diagram of hierarchical feature sampling and aggregation is shown below. Figure 5 As shown in the figure, the red dots represent the center point, the yellow dots represent its neighborhood points in the second layer, and the arrows represent the KNN neighborhood aggregation relationship.
[0041] The third-layer sampling points are further downsampled based on the second-layer sampling results. After selecting the center point, the relative coordinates and feature representations of its neighboring points in the second layer are obtained through KNN. Feature aggregation is then performed to generate multi-scale abstract features of the third-layer center point, realizing hierarchical representation of the point cloud from local to global.
[0042] From a probabilistic sampling perspective, the structure-guided mechanism introduced by SGH-Encoder can be viewed as a geometrically saliency-driven importance sampling process. This mechanism increases the weight of structurally complex regions in the sampling, allowing key components to be more fully expressed during hierarchical abstraction, while avoiding sampling centralization by retaining FPS sampling points.
[0043] Since crane point clouds are mainly composed of long, straight beam-like components and complex connection areas, their geometry naturally exhibits high linearity and abrupt changes in local curvature. SGH-Encoder, through a curvature- and linearity-driven structure-aware sampling mechanism, enables the network to prioritize capturing the crane's geometric skeleton and key connection points, thereby obtaining a more stable and discriminative multi-scale feature representation.
[0044] Step S3: Based on the image of the crane operation scene, extract multi-scale hierarchical features of the image, including: Step S3-1: Use a pre-trained first convolutional neural network to extract multi-scale feature maps of different levels of the image; Step S3-2: Convert the multi-scale feature map into a serialized token; Step S3-3: Integrate the tokens from different levels into a multi-scale visual token sequence as a multi-scale hierarchical feature of the image.
[0045] Step S4: Based on the multi-scale hierarchical features of the point cloud and the multi-scale hierarchical features of the image, the multi-scale hierarchical features of the point cloud after cross-modal fusion are obtained by utilizing cross-modal attention and feature transfer constraints. In step S4, a cross-modal attention module is constructed using point cloud multi-scale hierarchical features as queries and image multi-scale hierarchical features as keys and values. Through a cross-attention mechanism, image structural information is introduced into the point cloud representation to obtain the cross-modal fused point cloud multi-scale hierarchical features. This process can be represented as: in: For the first l The first point cloud feature of layer +1.
[0046] Feature transfer constraints minimize cross-modal feature transfer loss: in: For cross-modal feature transfer loss, This is the initial migration loss. To maintain constraints on content, For the first lThe first point cloud feature at layer -1 For the first l The first point cloud feature of the layer, For Gram matrices, For the first l The first image feature of the layer, the point cloud multi-scale hierarchical feature is composed of the first point cloud features of all layers, and the image multi-scale hierarchical feature is composed of the first image features of all layers.
[0047] Step S5: Based on the point cloud features after cross-modal fusion, a coarsely completed point cloud is obtained by using multi-branch mapping and adaptive weight fusion.
[0048] In crane point cloud completion tasks, the decoding stage not only needs to restore the overall geometric contour but also needs to ensure the topological continuity and structural consistency between components. Cranes are typically composed of multi-scale components such as long-scale continuous booms, rigid connection nodes, outrigger structures, and steel cables, and their geometric structures have obvious directional and engineering constraint characteristics. Therefore, traditional single-branch MLP or Folding-type decoders are prone to problems such as generation distribution collapse, long-scale structural oscillations, and component connection area breaks in this type of scenario, making it difficult to meet the requirements of engineering-level structural reconstruction. To address these issues, this invention proposes a multi-branch structural aggregation and residual refinement decoder. A schematic diagram of the multi-branch structural aggregation and residual refinement decoder is shown below. Figure 6 As shown.
[0049] Step S5 includes: Step S5-1: Apply multiple independent 1D convolutional layer branches to the multi-scale hierarchical features of the point cloud after cross-modal fusion for dimensionality upscaling; Step S5-2: Reshape the result of each branch after dimensionality increase into point cloud coordinate format to obtain the original branch point cloud corresponding to each branch; Step S5-3: Obtain the weighted branch point cloud based on each original branch point cloud and the branch weight matrix; Step S5-4: Summarize the weighted branch point clouds of all branches along the point dimension to obtain a coarsely completed point cloud.
[0050] Step S6: Obtain the completed point cloud based on the coarsely completed point cloud and the sparse laser point cloud.
[0051] To further improve the local geometric accuracy of the generated point cloud, this invention introduces a residual refinement module based on the coarse point cloud. The main reason for this is that although point clouds generated based on global features can recover the overall structure, relying solely on one-time generation makes it difficult to simultaneously ensure the consistency of the overall structure and the accuracy of local details, especially in boundary and detail areas where deviations and blurring are prone to occur. Therefore, by using residual learning to perform fine-grained correction of point positions, local geometric accuracy can be effectively improved while maintaining overall consistency.
[0052] Step S6 includes: Step S6-1: Stitch the coarsely completed point cloud with representative observation points to obtain a coarse-grained point cloud: in: For coarse-grained point clouds, To roughly complete the point cloud, The representative observation points were obtained by sampling the farthest point from the sparse laser point cloud. Step S6-2: Construct a dynamic KNN neighborhood for each center point in the coarse-grained point cloud; Step S6-3: Further construct the local relation descriptor: in: The center point in a coarse-grained point cloud The neighborhood, for neighborhood points, for and Local relation descriptors, Describe neighborhood points Relative to the center point Local geometric offset; Step S6-4: Extract local edge features from the multi-layer convolutional mapping network that shares the input of the local relation descriptors. ; Step S6-5: Based on all local edge features of the center point, calculate the max pooling feature and average pooling feature respectively; Step S6-6: Based on the max pooling feature and average pooling feature of each center point, obtain the fused local features: in: Center point The fusion of local features, Center point Max pooling characteristics, Center point Average pooling characteristics This is a channel fusion mapping based on 1×1 convolution; Max pooling emphasizes the most significant response in the neighborhood, which is beneficial for capturing boundary and locally abrupt structures; mean pooling, on the other hand, reflects the overall distribution characteristics of the neighborhood, helping to maintain the continuity and stability of long-scale components. A schematic diagram of the feature aggregation of max pooling and mean pooling is shown below. Figure 7As shown, the color intensity represents the characteristic response strength after pooling. Max pooling emphasizes the most significant local response, while mean pooling reflects the overall structural trend of the neighborhood.
[0053] Step S6-7: Based on the fusion of local features, predict the three-dimensional residual offset of each point, and update the coordinates of the center point using the three-dimensional residual offset; In the decoding stage, this application designs a multi-branch structure aggregation and residual refinement decoding module to optimize the structure of the coarsely completed full point cloud. This module generates an initial 3D point set through multi-branch structure mapping, enabling components of different scales to obtain balanced representation capabilities; subsequently, a dynamic graph neighborhood structure is constructed, and structure aggregation is performed through the statistical features of neighborhood extrema and mean values, combined with a residual update method to achieve progressive geometric reconstruction, thereby improving the continuity of local structures and the overall geometric consistency.
[0054] During network training, traditional point cloud completion methods mainly rely on constraints such as geometric errors, lacking constraints on local point distribution, which easily leads to point clustering and uneven sampling. For slender structures and complex connection areas in cranes, this invention introduces a uniformity constraint based on an inter-point repulsion mechanism, in addition to geometric reconstruction loss and cross-modal structural constraints, to improve the spatial distribution uniformity and continuity of the generated point cloud.
[0055] Step S6-8: Use all the updated center points as the completed point cloud.
[0056] The innovative aspects of this application are as follows: 1. Introduce a multimodal collaborative completion mechanism for images and laser point clouds. Compared to methods that rely solely on single-frame laser point clouds for completion, this invention combines synchronously acquired images and point cloud information to complete and reconstruct crane point clouds. This allows the texture and structural information in the images to guide the missing areas of the sparse point cloud, thereby improving the problem of incomplete structure of a single laser point cloud under conditions of occlusion, sparse sampling, and environmental interference.
[0057] 2. Design a structure-guided hierarchical point cloud coding mechanism Compared to point cloud encoding methods that rely solely on uniform sampling and feature extraction based on spatial distance, this invention introduces a local geometric complexity awareness mechanism during hierarchical point cloud encoding. It uses geometric statistics such as curvature and linearity to score the structure of sampling points and combines a hybrid sampling strategy to enhance the feature representation capabilities of key structural regions such as booms, outriggers, and connecting nodes, thereby improving the retention of key components in the hierarchical abstraction process.
[0058] 3. Design of a multi-branch structure aggregation decoder and residual refinement mechanism for reconstruction. Compared to traditional single-branch decoding or completion methods that rely solely on geometric error constraints, this invention employs a multi-branch structure aggregation and residual refinement mechanism in the decoding stage to progressively restore the dense three-dimensional structure.
[0059] 4. Design a point-to-point exclusion constraint optimization module Furthermore, a point exclusion constraint optimization module is introduced to explicitly constrain the local spatial distribution of the generated point cloud, thereby reducing point clustering and uneven distribution, and improving the spatial uniformity, continuity and overall geometric consistency of the generated point cloud.
[0060] In simulation experiments, this invention constructed a multimodal completion dataset for dynamic crane operation scenarios and verified the completion effect of the proposed method based on this dataset. The visualization effect of the point cloud completion method comparison is shown below. Figure 8 As shown, the method of the present invention also exhibits better point cloud completion effect in the crane boom, body and outrigger areas.
[0061] In real-world point cloud completion tasks, different methods exhibit significant differences in structural restoration performance. Traditional methods are prone to structural breaks, missing slender components, and geometric distortion under sparse observation conditions. Furthermore, they are susceptible to localized point cloud clustering or uneven distribution under complex environmental noise interference. In contrast, the method of this invention can better restore slender structures such as crane booms, hooks, and support rods, resulting in more continuous and complete completion results in both overall shape and local details. Figure 9 As shown in the figure. Furthermore, in the presence of noise and uneven sampling, the proposed method outperforms the other methods in terms of denoising capability and point cloud distribution uniformity. The completed point cloud distribution is more uniform, effectively avoiding local clustering and density unevenness, thereby improving the stability of the geometric representation. For structurally complex or severely missing regions, this method can utilize multimodal information to compensate for insufficient geometric information, suppressing noise interference while maintaining structural integrity, and obtaining more stable 3D reconstruction results.
[0062] This invention utilizes multimodal information from images and laser point clouds to complete and reconstruct sparse crane point clouds. Even when single-frame laser point clouds are affected by occlusion, sparse sampling, and environmental interference, resulting in structural gaps and local incompleteness, it better recovers key structures such as the crane boom, hook, outriggers, and connecting areas, thereby improving the completeness, continuity, and structural consistency of the point cloud reconstruction results. In particular, through structure-guided hierarchical point cloud encoding, cross-modal attention and feature transfer, multi-branch aggregation, and residual refinement techniques, it enhances the ability to represent and recover slender components and complex connecting regions, making the completed point cloud more stable and accurate in both overall geometry and local details.
[0063] Furthermore, this invention optimizes the local spatial distribution of the generated point cloud by introducing inter-point exclusion constraints, effectively suppressing point aggregation and uneven distribution, and improving the spatial uniformity and surface continuity of the output point cloud. Experimental results combining simulation scenarios and real-world conditions demonstrate that this invention maintains good reconstruction accuracy, structural integrity, and robustness even in complex engineering environments, providing a more reliable 3D data foundation for subsequent crane attitude estimation, key parameter calculation, and overturning risk assessment.
[0064] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for crane point cloud completion based on multi-modal data, characterized in that, include: Step S1: Acquire images and sparse laser point clouds of the crane operation scene; Step S2: Based on sparse laser point clouds, multi-scale hierarchical features of the point cloud are extracted using hybrid sampling and local feature aggregation. Step S3: Extract multi-scale hierarchical features from images of crane operation scenarios; Step S4: Based on the multi-scale hierarchical features of the point cloud and the multi-scale hierarchical features of the image, the multi-scale hierarchical features of the point cloud after cross-modal fusion are obtained by utilizing cross-modal attention and feature transfer constraints. Step S5: Based on the point cloud features after cross-modal fusion, a coarsely completed point cloud is obtained by using multi-branch mapping and adaptive weight fusion; Step S6: Obtain the completed point cloud based on the coarsely completed point cloud and the sparse laser point cloud.
2. The crane point cloud completion method based on multi-modal data according to claim 1, characterized in that, Step S1 also includes: preprocessing the image of the crane operation scene and the sparse laser point cloud.
3. The crane point cloud completion method based on multi-modal data according to claim 1, characterized in that, Step S2 includes: Step S2-1: Obtain the set of downsampling center points using hybrid sampling; Step S2-2: Assign weights to each point in the set of downsampled center points through structure-aware sampling; Step S2-3: Select a local neighborhood for each point in the downsampled center point set, and fuse the center point features with the relative positions and feature information of the neighborhood points to form the second feature of each point in the downsampled center point set: wherein: is a second feature of the point , is a second feature of the point , is a neighborhood of the point , is a multi-layer perceptron, is a pooling operation; Step S2-4: Repeat steps S2-1 to S2-3 multiple times to obtain multi-scale hierarchical features of the point cloud.
4. The crane point cloud completion method based on multimodal data according to claim 3, characterized in that, Step S2-1 includes: Step S2-1-1: Generate a set of spatially uniformly distributed center points by sampling from the farthest point, and expand it to obtain a set of basic sampling points; Step S2-1-2: For each point in the basic sampling point set, construct its neighborhood using the KNN algorithm to obtain the neighborhood relative coordinate distribution matrix: in: For point The neighborhood relative coordinate distribution matrix, for midpoint; Step S2-1-3: Construct the local covariance matrix: in: for The local covariance matrix, k The number of points in the neighborhood; Step S2-1-4: Perform eigenvalue decomposition on the covariance matrix to obtain the three first eigenvalues arranged in ascending order. Define local curvature and linearity indices based on the eigenvalue relationships, and further obtain the structural score. in: for Structural scoring, for The local curvature index, As the weight of the local curvature index, for linearity index, The weights of the linearity index, , , These are the three first eigenvalues arranged in ascending order; Step S2-1-5: Strengthen the weight distribution of key structural regions through an exponential modulation mechanism; Step S2-1-6: Obtain the structural score distribution based on the structural score after strengthening the weight distribution of the key structural regions, and perform probability sampling based on the normalized structural score distribution to obtain the set of structural guidance sampling points. The union of the set of center points and the set of structural guidance sampling points is used as the set of downsampling center points.
5. The crane point cloud completion method based on multimodal data according to claim 1, characterized in that, Step S3 includes: Step S3-1: Use a pre-trained first convolutional neural network to extract multi-scale feature maps of different levels of the image; Step S3-2: Convert the multi-scale feature map into a serialized token; Step S3-3: Integrate the tokens from different levels into a multi-scale visual token sequence as a multi-scale hierarchical feature of the image.
6. The crane point cloud completion method based on multimodal data according to claim 1, characterized in that, In step S4, a cross-modal attention module is constructed using point cloud multi-scale hierarchical features as queries and image multi-scale hierarchical features as keys and values. Image structural information is introduced into the point cloud representation through a cross-attention mechanism to obtain cross-modal fused point cloud multi-scale hierarchical features. The feature transfer constraint minimizes the cross-modal feature transfer loss: in: For cross-modal feature transfer loss, This is the initial migration loss. To maintain constraints on content, For the first l The first point cloud feature at layer -1 For the first l The first point cloud feature of the layer, For Gram matrices, For the first l The first image feature of the layer, the multi-scale hierarchical feature of the point cloud is composed of the first point cloud features of all layers, and the multi-scale hierarchical feature of the image is composed of the first image features of all layers.
7. The crane point cloud completion method based on multimodal data according to claim 1, characterized in that, Step S5 includes: Step S5-1: Apply multiple independent 1D convolutional layer branches to the multi-scale hierarchical features of the point cloud after cross-modal fusion for dimensionality upscaling; Step S5-2: Reshape the result of each branch after dimensionality increase into point cloud coordinate format to obtain the original branch point cloud corresponding to each branch; Step S5-3: Obtain the weighted branch point cloud based on each original branch point cloud and the branch weight matrix; Step S5-4: Summarize the weighted branch point clouds of all branches along the point dimension to obtain a coarsely completed point cloud.
8. A crane point cloud completion method based on multimodal data according to claim 1, characterized in that, Step S6 includes: Step S6-1: Stitch the coarsely completed point cloud with representative observation points to obtain a coarse-grained point cloud: in: For coarse-grained point clouds, To roughly complete the point cloud, The representative observation points were obtained by sampling the farthest point from the sparse laser point cloud. Step S6-2: Construct a dynamic KNN neighborhood for each center point in the coarse-grained point cloud; Step S6-3: Further construct the local relation descriptor: in: The center point in a coarse-grained point cloud The neighborhood, for neighborhood points, for and Local relation descriptors, Describe neighborhood points Relative to the center point Local geometric offset; Step S6-4: Extract local edge features from the multi-layer convolutional mapping network that shares the input of the local relation descriptors. ; Step S6-5: Based on all local edge features of the center point, calculate the max pooling feature and average pooling feature respectively; Step S6-6: Based on the max pooling feature and average pooling feature of each center point, obtain the fused local features: in: Center point The fusion of local features, Center point Max pooling characteristics, Center point Average pooling characteristics This is a channel fusion mapping based on 1×1 convolution; Step S6-7: Based on the fusion of local features, predict the three-dimensional residual offset of each point, and update the coordinates of the center point using the three-dimensional residual offset; Step S6-8: Use all the updated center points as the completed point cloud.
9. A crane point cloud completion device based on multimodal data, comprising a memory, a processor, and a program stored in the memory, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-8.
10. A storage medium having a program stored thereon, characterized in that, When the program is executed, it implements the method as described in any one of claims 1-8.