Crane operation monitoring method based on monocular depth estimation and three-dimensional restoration

CN122157174BActive Publication Date: 2026-09-25CHINA RAILWAY DESIGN GRP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610636921.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-09-25
Estimated Expiration
2046-05-11

AI Technical Summary

Benefits of technology

[0058]1. 高精度吊臂检测:通过深度图生成技术,本发明的方法能够在复杂施工环境中实时、精确地检测吊臂位置。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157174B_ABST
    Figure CN122157174B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on monocular depth estimation and three-dimensional restoration's boom operation monitoring method, comprising the following steps: S1, monocular depth estimation model is constructed;S2, monocular depth estimation model is trained;S3, obtains calibration plate image and depth control point;Real-time acquisition crane operation scene photo and video;S4, based on calibration plate image, obtain internal and external parameter matrix to photo and video are calibrated and corrected, obtain crane arm position plane coordinate and railway position annotation file;S5, photo and video are input monocular depth estimation model to obtain disparity prediction graph, obtain depth prediction graph;S6, to depth prediction graph geometric correction;S7, generate three-dimensional point cloud;S8, based on crane arm position plane coordinate and railway position annotation file, utilize three-dimensional point cloud to obtain crane arm and railway real world position, the posture of crane boom operation is monitored and risk early warning.The method of the application is high in precision, low in cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and deep learning technology, specifically to a method for monitoring crane operations based on monocular depth estimation and 3D reconstruction. Background Technology

[0002] With the advancement of urban infrastructure construction and modern engineering projects, accurately monitoring whether cranes are within safe operating ranges is becoming increasingly important in the construction of railway lines, tracks, and bridges. The boom length, operating position, and angle with the ground are key factors in assessing whether a crane is within a safe operating range.

[0003] Existing monitoring methods include lidar monitoring, monocular ranging, and binocular ranging. Among them, lidar monitoring is widely used, but it is costly, complex to install, and has poor adaptability in certain special environments. Monocular ranging and binocular ranging are only suitable for near-field targets and cannot fully meet the accuracy requirements of boom monitoring in practical applications. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a highly adaptable, accurate, and low-cost method for monitoring boom operations based on monocular depth estimation and 3D reconstruction.

[0005] Therefore, the present invention adopts the following technical solution:

[0006] A method for monitoring crane operations based on monocular depth estimation and 3D reconstruction includes the following steps:

[0007] S1, Construct a monocular depth estimation model, including a DepthAnythingV2-L pre-trained model, multiple upsampling modules, and multiple fusion enhancement modules;

[0008] S2, obtain image-disparity value image pairs based on public datasets, and use them to train the monocular depth estimation model;

[0009] S3: Acquire calibration board images captured by each camera in the crane operation scene; acquire depth control points; acquire photos and videos of the crane operation scene in real time through the cameras and store them according to the naming rule of "date-time-operation scene";

[0010] S4, based on the calibration board image acquired by S3, obtains the camera intrinsic and extrinsic parameter matrix, which is used to calibrate and correct the photos and videos acquired in real time by S3, and then obtains the planar coordinates of the crane boom position; based on the photos and videos of the crane operation scene acquired in real time by S3, the corresponding railway position annotation file is generated, wherein the video is processed in the form of video frames;

[0011] S5 takes the calibrated and corrected photos and videos from S4 and inputs them into the trained monocular depth estimation model. The output disparity prediction map is then processed by performing a reciprocal operation to obtain the depth prediction map and the depth prediction value for the entire image. When the input is video, it is input into the monocular depth estimation model in the form of video frames for processing.

[0012] S6, using the depth control points obtained in S3, perform geometric correction on the depth prediction map to obtain the corrected depth map and the depth prediction value. ;

[0013] S7 generates a high-precision 3D point cloud through inverse projection based on the corrected depth map;

[0014] S8, based on the planar coordinates of the crane boom position and the railway position annotation file obtained in S4, uses the three-dimensional point cloud to obtain the real-world positions of the crane boom and the railway, establishes an electronic fence, and monitors and provides risk warnings for the crane boom's operating posture in conjunction with crane operation constraints.

[0015] In step S1 above:

[0016] The input to the monocular depth estimation model is The size is A three-channel image with an output dimension of Feature map The upsampling modules consist of five identical modules: RAFB1, RAFB2, RAFB3, RAFB4, and RAFB5. The fusion enhancement modules consist of six modules: FE1, FE2, FE3, FE4, FE5, and FE6.

[0017] In step S1 above:

[0018] The monocular depth estimation model is based on the input... The size is The processing procedure for the three-channel image is as follows:

[0019] The image input to the monocular depth estimation model is processed using the DepthAnythingV2-L pre-trained model to obtain the dimension. Feature map ;

[0020] Feature map These are the inputs to the upsampling module RAFB1 and the fusion enhancement module FE1, respectively. The output of the upsampling module RAFB1 is a module with dimension [missing information]. Feature map The output of the fusion enhancement module FE1 is a dimension of Feature map Among them, the fusion enhancement module FE1 will input dimension as... Feature map Expanded to a 5-dimensional feature map, with dimensions of Then, the output feature map is obtained through deep attention feature enhancement structure, feature map restoration and deconvolution processing. ;

[0021] The input to the upsampling module RAFB2 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE2 is the feature map. and The output is Feature map ;

[0022] The input to the upsampling module RAFB3 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE3 is the feature map. and The output is Feature map ;

[0023] The input to the upsampling module RAFB4 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE4 is the feature map. and The output is Feature map ;

[0024] The input to the upsampling module RAFB5 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE5 is the feature map. and The output is Feature map ;

[0025] The fusion enhancement modules FE2 to FE5 have the same internal structure and the same process for processing the two input feature maps: first, the two input feature maps are superimposed, and then the 4-dimensional feature map is expanded into a 5-dimensional feature map through reshaping. Then, the feature map is enhanced by a deep attention feature enhancement structure. The feature map is obtained through processing. Then, the feature map is reshaped. The dimensions are restored to be the same as the two input feature maps; finally, through deconvolution, the length and width of the output feature map are expanded to twice the original, and the number of channels is compressed to half the original.

[0026] The input to the fusion enhancement module FE6 is a feature map. and The output is a feature map. Among them, the fusion enhancement module FE6 will input dimension as Feature map and The features are stacked, and then processed through feature map expansion, deep attention feature enhancement structure, and feature map restoration to obtain a dimension of [dimensional value missing]. Feature maps; finally, through linear layers Compressing the feature map channel number to 1 results in a dimension of Feature map .

[0027] The upsampling module includes two adaptive feature enhancement layers and one deconvolution layer. The upsampling module will process the input dimension as follows: The feature map's length and width are doubled, while the number of channels is halved, resulting in a dimension of... The feature map, where:

[0028] The adaptive feature enhancement layer first passes through a... Convolution converts the input dimension to 1. The number of channels in the feature map is compressed to one-quarter of the original, resulting in a dimension of The feature map is then fed into a 3D attention-modulated convolution, where the dimensions of the output feature map remain unchanged; finally, it passes through a... Convolution expands the number of channels in the feature map to its original number, ensuring that the input and output feature maps of the adaptive feature enhancement layer have the same dimension; the deconvolution layer is used to expand the dimension of the feature map and compress the number of channels in the feature map.

[0029] The three-dimensional attention-modulated convolution includes three different convolutions:

[0030] For spatial attention convolution, the Gather-Excite algorithm is used;

[0031] For channel attention convolution, the Squeeze-Excitation algorithm is used;

[0032] For filter-dimensional attention mechanism convolutions, weight adjustment is performed using convolution kernels based on SE (Search Engine) principles.

[0033] The deep attention feature enhancement structure described in step S1 above operates on the input as follows:

[0034] First, for dimension... Feature map Perform a reshape operation to convert its dimensions to... ,in, This represents the total number of spatial locations on the feature map, and is then used as input to two branches:

[0035] The first branch is a SEBlock connected to a linear layer. The output dimension is Feature map ;

[0036] In the second branch, first use a step size of The 3×3 convolution is downsampled to obtain a dimension of The feature map is used as input to the linear layer. The output dimension is Feature map Simultaneously, the feature map obtained from the downsampling process is used as input into a cascaded convolutional layer and a linear layer. The output dimension is Feature map ;

[0037] The obtained feature map Attention is calculated according to the following formula:

[0038] ,

[0039] in, For splicing operations; For activation functions; ; Represents the vector dimension. As a scaling factor;

[0040] The result after attention calculation is passed through a linear layer. Processing, output dimension is Feature map .

[0041] The specific operation of geometric correction in step S6 above is as follows:

[0042] A polynomial fitting correction method is used to adjust each control point. depth prediction value Use this as input to establish a correction function:

[0043] ,

[0044] in, Based on depth prediction Calculated correction depth value; , , and The correction coefficients to be solved are;

[0045] The data from at least four control points are fitted using the least squares method to minimize the following objective function:

[0046] ,

[0047] in, This is the actual measured depth value. The number of control points is the basis for solving this optimization problem, which yields the optimal correction coefficients. , , and ;

[0048] The depth prediction value described in S5 Substituting into the correction function, the corrected depth value is calculated. Generate a corrected depth map;

[0049] Two control points that were not involved in the coefficient fitting were randomly selected, and the error between their corrected depth values ​​and actual depth values ​​was calculated to ensure that the average absolute error was less than 0.1 meters. If the error exceeded the standard, control points were reselected and the correction coefficients were fitted.

[0050] The specific operation of step S7 above is as follows:

[0051] The two-dimensional coordinates on the corrected depth map are converted to a normalized camera coordinate system using the inverse of the intrinsic parameter matrix, and then multiplied by the depth value to obtain the actual three-dimensional coordinates. The three-dimensional point cloud of the entire image is generated using the three-dimensional coordinates calculated pixel by pixel, as expressed by the formula:

[0052] ,

[0053] Among them, image coordinates The corresponding three-dimensional coordinates of the camera coordinate system are ( ), Camera intrinsic parameter matrix The inverse matrix.

[0054] In step S8 above:

[0055] Calculate the shortest straight-line distance L from the key part of the crane boom to the railway area. If L < 5 meters, it is judged as "Level 1 Warning"; if L < 2 meters, it is judged as "Level 2 Warning"; if it violates the limit, a red warning light will be displayed directly in the pop-up window.

[0056] Based on the three-dimensional points at the top and bottom of the boom, combined with the vertical orientation of the camera, the boom extension length, working radius, and working angle are calculated. Corresponding constraints are set according to the vehicle tonnage, load weight, and construction site requirements. When the constraints are exceeded, an audible and visual warning is issued.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] 1. High-precision boom detection: Through depth map generation technology, the method of this invention can detect the boom position in real time and accurately in complex construction environments.

[0059] 2. Reconstructing 3D information in a purely visual plane: This invention generates and corrects depth maps to further monitor the working posture of the boom, such as boom length and working azimuth angle. It can provide timely warnings when the boom intrudes into the operating limits, ensuring that the boom is within the safe operating range.

[0060] 3. Low cost: Compared with lidar sensor solutions, the method of this invention uses a high-definition camera for real-time monitoring, avoiding the use of high-cost equipment, and has greater adaptability and lower implementation cost. Attached Figure Description

[0061] Figure 1 This is a flowchart of the crane operation monitoring method in an embodiment of the present invention;

[0062] Figure 2 This is a structural diagram of the monocular depth estimation model in an embodiment of the present invention;

[0063] Figure 3 The diagram shows the upsampling module and its components in the monocular depth estimation model in this embodiment of the invention. (a) is a structural diagram of the upsampling module, (b) is a structural diagram of the adaptive feature enhancement layer in the upsampling module, and (c) is a structural diagram of the three-dimensional attention modulation convolution in the adaptive feature enhancement layer.

[0064] Figure 4 This is a structural diagram of the depth attention feature enhancement structure in the monocular depth estimation model of this invention.

[0065] Figure 5 The following is an example of two pairs of "image-disparity value" pairs used for training the monocular depth estimation model in an embodiment of the present invention, wherein (a) and (b) are a pair of "image-disparity value" images, and (c) and (d) are another pair of "image-disparity value" images;

[0066] Figure 6 Figures (a) and (b) in the figure are the input image and the output single-channel disparity prediction image of the monocular depth estimation model in the embodiment of the present invention, respectively.

[0067] Figure 7 The above are examples of three-dimensional point cloud reconstructions of scenes in embodiments of the present invention, wherein (a) is the original image, (b) is the obtained three-dimensional point cloud map, and (c) and (d) show the morphology of the three-dimensional point cloud from different perspectives. Detailed Implementation

[0068] The technical solution of the invention will be clearly and completely described below with reference to the accompanying drawings and embodiments. Obviously, the following embodiments are only some embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0069] Example

[0070] like Figure 1 As shown in the figure, the crane operation monitoring method based on monocular depth estimation and 3D reconstruction in this embodiment has the following specific steps:

[0071] S1, Construct a monocular depth estimation model.

[0072] The constructed monocular depth estimation model is as follows: Figure 2 As shown, it includes a DepthAnythingV2-L pre-trained model, 5 upsampling modules (RAFB Groups), and 6 fusion enhancement modules. The model input is... The size is A three-channel image with an output dimension of Feature map ,yes The size is The single-channel image. The five upsampling modules have the same structure, namely RAFB1, RAFB2, RAFB3, RAFB4, and RAFB5; the six fusion enhancement modules are FE1, FE2, FE3, FE4, FE5, and FE6. Specifically, the image input to the monocular depth estimation model is processed by the DepthAnythingV2-L pre-trained model to obtain a dimension of Feature map ;

[0073] Feature map These are the inputs to the upsampling module RAFB1 and the fusion enhancement module FE1, respectively. The output of the upsampling module RAFB1 is a module with dimension [missing information]. Feature map The output of the fusion enhancement module FE1 is a dimension of Feature map Among them, the fusion enhancement module FE1 will input dimension as... Feature map Expanded to a 5-dimensional feature map, with dimensions of Then, feature maps are obtained through deep attention feature enhancement structure, feature map restoration and deconvolution processing. .

[0074] The input to the upsampling module RAFB2 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE2 is the feature map. and The output is Feature map ;

[0075] The input to the upsampling module RAFB3 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE3 is the feature map. and The output is Feature map ;

[0076] The input to the upsampling module RAFB4 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE4 is the feature map. and The output is Feature map ;

[0077] The input to the upsampling module RAFB5 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE5 is the feature map. and The output is Feature map ;

[0078] Among them, the internal structure of the fusion enhancement modules FE2 to FE5 is the same, and the process of processing the two input feature maps is the same. Taking the fusion enhancement module FE2 as an example, it first processes the input feature maps of dimension 1. Feature map and feature map The features are stacked and then reshaped to expand the 4D feature map into a 5D feature map. ,in, , representing the total number of spatial locations on the feature map; then, the feature map is enhanced using a deep attention feature enhancement structure. The feature map is obtained through processing. Then, the feature map is reshaped. By reconstructing the feature map, we obtain a dimension of The feature map is then processed; finally, deconvolution is used to double the length and width of the feature map, while halving the number of channels, resulting in a feature map with dimensions of [insert dimensions here]. Feature map .

[0079] The input to the fusion enhancement module FE6 is a feature map. and The output is a feature map. Among them, the fusion enhancement module FE6 will input a dimension of... Feature map and The features are stacked, and then processed through feature map expansion, deep attention feature enhancement structure, and feature map restoration to obtain a dimension of [dimensional value missing]. Feature maps; finally, through linear layers Compressing the feature map channel number to 1 results in a dimension of Feature map .

[0080] The following describes the specific structure of the deep attention feature enhancement structure in the upsampling module and the fusion enhancement module.

[0081] like Figure 3 As shown in Figure (a), the upsampling module includes two adaptive feature enhancement layers and one deconvolution layer (with a kernel size of 3×3). The adaptive feature enhancement layers do not change the dimension of the input feature map; the deconvolution layer expands the dimension of the feature map and compresses the number of channels. The upsampling module upsamples the input feature map with a dimension of... The feature map length and width are doubled, the number of channels is halved, and the output dimension is... The feature map, where, Represents batch processing volume. Represents the number of channels. and These represent the length and width of the feature map, respectively.

[0082] Among them, such as Figure 3 As shown in Figure (b), the adaptive feature enhancement layer in the upsampling module first uses a 1×1 convolution to enlarge the input dimension. The feature map channel number is compressed to one-quarter of the original, resulting in a dimension of The feature map is then fed into a 3D attention-modulated convolution, where the dimensions of the output feature map remain unchanged. Finally, it passes through a... Convolution expands the number of channels in the feature map to the original number, ensuring that the input and output feature maps of the adaptive feature enhancement layer have the same dimension.

[0083] like Figure 3 As shown in Figure (c), the 3D attention modulation convolution in the adaptive feature enhancement layer includes three different convolutions, wherein:

[0084] For spatial attention convolution, the Gather-Excite algorithm is used to dynamically adjust the weights of the convolution kernel at different spatial locations, enabling the convolution operation to focus more precisely on specific regions and adapt to changes in the spatial structure of the input image.

[0085] For channel attention convolution, the Squeeze-Excitation algorithm is used to dynamically adjust the weights of each channel of the convolution kernel according to the channel information of the input feature map, so that the network can focus on important features according to the importance of specific channels.

[0086] For the filter dimension attention mechanism convolution, a convolution kernel based on the SE idea is used to adjust the weights. The filter weights of the convolution kernel are adjusted according to the different feature types of the input features (such as texture, shape, etc.), thereby improving the recognition ability of specific feature types.

[0087] like Figure 4 As shown, the deep attention feature enhancement structure first processes the feature map... Perform a reshape operation to convert its dimensions to... Then, it is used as input into two branches:

[0088] The first branch is a linear layer connected in series with an SEBlock (Squeeze-and-Excitation Block). The output dimension is Feature map ,in for .

[0089] In the second branch, a 3×3 convolution (with a stride of 1) is first applied. Perform downsampling to obtain dimension . The feature map obtained by the downsampling process is used as input into the linear layer. The output dimension is Feature map Simultaneously, the feature map obtained from the downsampling process is used as input into a cascaded convolutional layer and a linear layer. The output dimension is Feature map .

[0090] Then, the obtained feature map Attention is calculated according to the following formula:

[0091] ,

[0092] in, For splicing operations; For activation functions; ; Represents the vector dimension. As a scaling factor, it prevents the inner product value from becoming too large, which could lead to gradient vanishing or exploding.

[0093] The result after attention calculation is passed through a linear layer. Processing, output dimension is Feature map .

[0094] In the above operations:

[0095] The feature map Responsible for feature selection. SEBlock is an enhanced feature representation mechanism that automatically learns the weight of each channel in the feature map through channel attention, allowing the model to adjust according to the importance of the input features. This process helps the feature map... To more accurately represent the model's focus.

[0096] The feature map To carry semantic associations, it is only necessary to use it with feature maps. To match and calculate similarity, a linear transformation is used to map it to a suitable space so that it can be compared with the feature map. and feature map To achieve effective matching and interaction.

[0097] The feature map The content it carries includes not only simple features, but also information from different locations. The transmission of this information is influenced by the feature map. and feature map The impact of feature maps. Having more information is to ensure that as much useful content as possible is retained and transmitted during the weighting and transmission process.

[0098] Convolution operations can aggregate spatial features of input features and improve the efficiency of information flow.

[0099] In summary, the deep attention feature enhancement structure improves computational efficiency through effective feature compression and information flow, while ensuring that the model can accurately capture key information in the attention mechanism.

[0100] S2 trains the monocular depth estimation model constructed by S1.

[0101] To address the depth estimation requirements of construction site crane monitoring scenarios, and considering the core features of the scenario (open outdoor environment, large mechanical targets, complex lighting / weather interference, and the need for near-field and mid-to-far-field depth perception), three publicly available image-disparity image pair datasets were selected: TartanAir, MegaDepth, and HRWSI.

[0102] Among them, TartanAir's 306,000 images can provide simulated outdoor scenes similar to construction sites (such as open spaces, complex weather and lighting changes) and multimodal labels, helping the model built by S1 to adapt to the changing environment of construction sites; MegaDepth's 128,000 outdoor scene images can supplement the learning of mid-to-far field depth patterns during crane operations; HRWSI's 20,000 high-resolution images and rich mask information can accurately capture the depth features of details such as the crane body and boom, and filter out invalid interference information in the construction site background.

[0103] The three publicly available datasets contain a total of 454,000 labeled images. These images are divided into a training set (317,800 images), a test set (90,800 images), and a validation set (45,400 images) in a 7:2:1 ratio. This approach maintains dataset diversity while leveraging the characteristics of different data sources to enhance the model's generalization ability. It ensures both the sufficiency and scenario diversity of the training data, and allows for accurate evaluation of the model's depth estimation accuracy and robustness in construction site crane monitoring scenarios through independent test and validation sets.

[0104] The monocular depth estimation model constructed in S1 was trained using the training set. During training, the weight parameters of the pre-trained DepthAnythingV2-L model were retained, and its top-level features were fine-tuned. The model's specificity was improved by optimizing labeled data from crane operation scenarios and long-distance scenarios. The training of the monocular depth estimation model is essentially a regression task under supervised learning, using labeled data such as... Figure 5The image-disparity value pairs shown are training samples (two pairs of samples are shown in the figure, where Figure (a) and Figure (b) are one pair of "image-disparity value" images, and Figure (c) and Figure (d) are another pair of "image-disparity value" images). The model parameters are continuously adjusted through backpropagation to minimize the error between the predicted disparity output by the model and the true disparity.

[0105] In this invention, the training loss function of the constructed monocular depth estimation model is the sum of the mean squared error loss function (MSELoss) and the structural similarity loss (SSIM). Training stops if the loss function does not decrease within 20 iterations. After training, a trained monocular depth estimation model is obtained. After each training round, the model's performance is evaluated based on the validation set, and the training process is monitored according to changes in performance evaluation metrics. The test set is used to evaluate the performance of the trained monocular depth estimation model.

[0106] The monocular depth estimation model of the present invention was compared with the performance of several existing depth estimation models (Semantic-Mono-Depth, Monodepth, Refine-and-Distill and DepthAnythingV2) at a distance of 80 meters. The performance evaluation indicators used included absolute relative error, root mean square error and logarithmic root mean square error. The experimental results are shown in Table 1.

[0107] Table 1

[0108]

[0109] As can be seen from the experimental results in Table 1, the monocular depth estimation model of this invention outperforms other models in all performance evaluation metrics, demonstrating its superior performance in depth estimation tasks. In particular, the absolute relative error is 0.087, and the root mean square error is reduced to 3.023, significantly improving the accuracy and reliability of depth estimation.

[0110] S3, Data Acquisition, the specific process is as follows:

[0111] S31. Acquire the calibration board image captured by the camera. Prepare a checkerboard calibration board (12×9 squares, 45mm side length, 600×450mm overall dimensions) and a fixing rod. Keeping the camera position unchanged, gradually change the placement angle and distance of the calibration board (adjusted sequentially to 1m, 1.5m, 2.5m, 3m). Take 5-10 clear images of the calibration board at each angle and distance as calibration images for the camera's internal and external parameters, ensuring that the calibration board in the images is free of blur and distortion.

[0112] S32, Obtain Depth Control Points. Depth control points are clearly visible, salient object feature points within the monitored scene. Their image coordinates and the actual depth (i.e., straight-line distance, referred to as depth) from the point to the camera must be recorded simultaneously. When selecting control points, priority should be given to uniformly distributed, staggered, and stable fixed object points (such as the base of fixed streetlights, building walls, railway support bases, etc.), and at least six should be selected. Simultaneously, the actual depth from each control point to the camera should be measured and recorded. There are two methods for obtaining the depth control points:

[0113] 3D model measurement method: Based on the construction site area, an industrial-grade drone equipped with a dual-lens aerial photography module (4K resolution) and a GPS positioning system was selected. Ground station software was used to plan the drone's flight path, employing a grid-like flight mode. The flight altitude was set at 50m, with a horizontal overlap rate of 80% and a vertical overlap rate of 70%, ensuring effective overlap of adjacent aerial images to meet the requirements of 3D modeling. Effective images were categorized and stored in a 3D modeling database. The scene was modeled using the 3D modeling software Pix4D. In the modeled 3D scene, distance measurement tools were used to obtain the distance from control points to the camera positions.

[0114] On-site measurement method: High-precision measuring instruments, such as a total station (Leica TS09, measurement accuracy ±1mm +1.5ppm) or a handheld laser rangefinder (measurement range 0.1-100m, accuracy ±0.2mm), are used for direct data acquisition. First, a measurement reference point is established at the camera installation location. The three-dimensional coordinates (X, Y, Z) of the camera are calibrated using a total station. Then, a prominent object point is selected in the monitoring screen, and the surveyor locates the corresponding physical point on-site. The straight-line distance between this point and the camera reference point is measured using a total station; this distance is the depth value. A handheld laser rangefinder can be used for rapid data acquisition, improving efficiency.

[0115] S33 can capture photos and videos of crane operation scenes in real time.

[0116] Industrial-grade high-precision cameras are used to record videos of different crane operation scenarios (including no-load operation, light-load operation, heavy-load operation, slewing motion, luffing motion, and telescopic motion). A photo is automatically taken every 5 seconds. The captured photos and videos are categorized and stored according to the naming convention of "date-time-operation scenario," and backed up to a local server and a cloud database. Industrial-grade high-precision cameras with a resolution of at least 4K are selected, paired with a gimbal stabilization system and a dustproof and waterproof housing. The frame rate is adjusted to 30fps, and exposure, white balance, and ISO are automatically adjusted according to ambient light to ensure image / video clarity.

[0117] S4, Data Preprocessing, the specific operations are as follows:

[0118] S41, based on the calibration board image acquired by S31, obtains the camera's intrinsic and extrinsic parameter matrix. The images acquired by the camera are affected by the camera's intrinsic parameters (focal length, principal point) and extrinsic parameters (camera pose relative to the world coordinate system), therefore, camera calibration is required before 3D reconstruction.

[0119] Intrinsic parameter matrix The formula is:

[0120] ,

[0121] in, and These are the focal lengths of the camera along the y-axis and x-axis, respectively. and The coordinates of the main point.

[0122] The extrinsic parameter matrix describes the position of the camera coordinate system relative to the world coordinate system, and the formula is as follows:

[0123] ,

[0124] in, For rotation matrix, It is a displacement vector. Three-dimensional coordinates in the world coordinate system. These are the three-dimensional coordinates in the camera image coordinate system.

[0125] S42, obtain the planar coordinates of the crane boom:

[0126] Based on the camera intrinsic and extrinsic parameter matrix calculated by S41, the photos and real-time videos of the crane operation scene collected in real time by S33 are calibrated and corrected to obtain the planar coordinates of the crane boom position.

[0127] There are two optional methods for obtaining the planar coordinates of the crane boom position. The appropriate method can be chosen flexibly as needed to ensure the integrity and accuracy of the position data, providing standardized coordinate input for subsequent 3D positioning. The details are as follows:

[0128] (1) Automatic detection and extraction: Based on the pre-trained yolov11 crane boom target detection algorithm, the calibrated and corrected video or image is automatically detected, the crane boom target in the image is identified and selected, and its two-dimensional position coordinate information is extracted simultaneously and accurately.

[0129] (2) Manual annotation and completion: When the automatic detection algorithm fails to identify the crane arm due to occlusion, complex working conditions (such as interference from the construction site background, special crane arm posture, etc.), manual interactive annotation is adopted. The annotator manually selects the crane arm in the image to ensure that the target boundary is accurately defined, and clearly records its two-dimensional position coordinates in the label.

[0130] S43, railway location marker:

[0131] For the photos and videos captured by each deployed industrial-grade high-precision camera, a manual interactive annotation method is adopted. 5-10 railway feature points (such as track edges, sleeper endpoints, etc.) are manually clicked in the image to accurately locate the railway position. Finally, each image generates a corresponding rail_points.csv file, which is a railway position annotation file. The video is processed in the form of video frames.

[0132] S5, Real-time Scene Depth Prediction:

[0133] The calibrated and corrected photos and videos obtained from S42 are input into the monocular depth estimation model trained from S2. The model outputs a single-channel disparity prediction map of the entire image, such as... Figure 6 As shown (where (a) is the input image of the monocular depth estimation model, and (b) is the single-channel disparity prediction map output by the monocular depth estimation model), when the input is video, it is input into the monocular depth estimation model in the form of video frames for processing. Since the depth model outputs disparity values, the pixel value is larger the closer to the camera position. After performing a reciprocal operation on the disparity map output by the model, the depth prediction map is obtained, thus obtaining the full-image depth prediction value. .

[0134] S6. Use the depth control points obtained in S32 to perform geometric correction on the depth prediction map obtained in S5. The specific operation is as follows:

[0135] A polynomial fitting correction method is used to adjust each control point. depth prediction Establish the correction function as input:

[0136] ,

[0137] in, Based on depth prediction Calculated correction depth value; , , and The correction coefficients to be solved are denoted as .

[0138] The control point (at least 4) data are fitted using the least squares method to minimize the following objective function:

[0139] ,

[0140] in, Here, n is the actual measured depth value, and n is the number of control points. After solving this optimization problem, the optimal correction coefficient is obtained. , , and .

[0141] The full-map depth prediction value output by S5 Substitute the values ​​into the correction function to calculate the corrected depth value. This generates a corrected depth map.

[0142] Randomly select two control points that were not involved in the coefficient fitting, and calculate the error between their corrected depth values ​​and actual depth values, ensuring that the MAE (mean absolute error) is less than 0.1 meters. If the error exceeds the standard, it is necessary to reselect control points and refit the correction coefficients.

[0143] S7 generates a high-precision 3D point cloud based on the corrected depth map through inverse projection:

[0144] The corrected depth map provides the position and corresponding depth information of each pixel on the image plane. The camera's intrinsic parameter matrix converts the two-dimensional coordinates on the image plane into three-dimensional coordinates in the camera coordinate system. Through inverse projection, a high-precision three-dimensional point cloud is reconstructed from the corrected depth map. Specifically, each pixel on the depth map corresponds to a three-dimensional point, representing a location in the scene. The two-dimensional coordinates on the image plane are converted to a normalized camera coordinate system using the inverse of the intrinsic parameter matrix, and then multiplied by the depth value to obtain the actual three-dimensional coordinates. The three-dimensional point cloud of the entire image is generated by calculating the three-dimensional coordinates pixel by pixel, as expressed by the formula:

[0145] ,

[0146] Among them, image coordinates The corresponding three-dimensional coordinates of the camera coordinate system are ( ), Camera intrinsic parameter matrix The inverse matrix.

[0147] The three-dimensional point cloud generated in one embodiment of the present invention, such as Figure 7 As shown, (a) is the original image, (b) is the obtained 3D point cloud, and (c) and (d) show the morphology of the 3D point cloud from different perspectives, which are used to intuitively verify the spatial structural integrity and geometric consistency of the generated 3D point cloud.

[0148] S8, Crane boom operation posture monitoring and risk warning:

[0149] Based on the planar coordinates of the crane boom position obtained in S42 and the railway position annotation file obtained in S43, the real-world positions of the crane boom and the railway are obtained using the 3D point cloud obtained in S7. An electronic fence is then established, and combined with crane operation constraints, the crane boom's operational posture in real-time scenarios can be monitored and warned. This includes:

[0150] Railway encroachment warning: Calculate the shortest straight-line distance L from the key parts of the crane boom (top, hook) to the railway area. If L < 5 meters, it is judged as "Level 1 warning"; if L < 2 meters, it is judged as "Level 2 warning"; if it encroaches on the railway limit, a red light warning will be displayed directly in the pop-up window.

[0151] Crane boom overlength warning: Based on the three-dimensional points at the top and bottom of the boom, combined with the vertical orientation of the camera, the boom extension length, working radius, and working angle are calculated. Corresponding constraints are set according to the vehicle tonnage, load weight, and construction site requirements. When the constraints are exceeded, an audible and visual warning is issued.

[0152] In one embodiment of the present invention, the actual measured boom length and working radius on site are compared with the boom length and working radius detected by the method of the present invention, and the error is calculated. The results are shown in Table 2.

[0153] Table 2

[0154]

[0155] As can be seen from Table 2, the error between the boom length and working radius detected by the present invention and the actual measurement results is very low, proving that the method of the present invention has high boom detection accuracy in actual boom operation scenarios.

Claims

1. A method for monitoring crane operation based on monocular depth estimation and 3D reconstruction, characterized in that, Includes the following steps: S1, Construct a monocular depth estimation model, including a DepthAnythingV2-L pre-trained model, multiple upsampling modules, and multiple fusion enhancement modules; where: The upsampling module includes two adaptive feature enhancement layers and one deconvolution layer. The upsampling module will process the input dimension as follows: The feature map's length and width are doubled, while the number of channels is halved, resulting in a dimension of... Feature maps, where: The adaptive feature enhancement layer first passes through a... Convolution converts the input dimension to 1. The number of channels in the feature map is compressed to one-quarter of the original, resulting in a dimension of The feature map is then fed into a 3D attention-modulated convolution, where the dimensions of the output feature map remain unchanged; finally, it passes through a... Convolution expands the number of channels in the feature map to its original number, ensuring that the input and output feature maps of the adaptive feature enhancement layer have the same dimension; the deconvolution layer is used to expand the dimension of the feature map and compress the number of channels in the feature map. S2, obtain image-disparity value image pairs based on public datasets, and use them to train the monocular depth estimation model; S3: Acquire calibration board images captured by each camera in the crane operation scene; acquire depth control points; acquire photos and videos of the crane operation scene in real time through the cameras and store them according to the naming rule of "date-time-operation scene"; S4, based on the calibration board image acquired by S3, obtains the camera intrinsic and extrinsic parameter matrix, which is used to calibrate and correct the photos and videos acquired in real time by S3, and then obtains the planar coordinates of the crane boom position; based on the photos and videos of the crane operation scene acquired in real time by S3, the corresponding railway position annotation file is generated, wherein the video is processed in the form of video frames; S5 takes the calibrated and corrected photos and videos from S4 and inputs them into the trained monocular depth estimation model. The output disparity prediction map is then processed by performing a reciprocal operation to obtain the depth prediction map and the depth prediction value for the entire image. When the input is video, it is input into the monocular depth estimation model in the form of video frames for processing. S6, using the depth control points obtained in S3, perform geometric correction on the depth prediction map to obtain the corrected depth map and the depth prediction value. ; S7 generates a high-precision 3D point cloud through inverse projection based on the corrected depth map; S8, based on the planar coordinates of the crane boom position and the railway position annotation file obtained in S4, uses the three-dimensional point cloud to obtain the real-world positions of the crane boom and the railway, establishes an electronic fence, and monitors and provides risk warnings for the crane boom's operating posture in conjunction with crane operation constraints.

2. The crane operation monitoring method according to claim 1, characterized in that, In S1: The input to the monocular depth estimation model is The size is A three-channel image with an output dimension of Feature map The upsampling modules consist of five identical modules: RAFB1, RAFB2, RAFB3, RAFB4, and RAFB5. The fusion enhancement modules consist of six modules: FE1, FE2, FE3, FE4, FE5, and FE6.

3. The crane operation monitoring method according to claim 2, characterized in that, In S1: The monocular depth estimation model is based on the input... The size is The processing procedure for the three-channel image is as follows: The image input to the monocular depth estimation model is processed using the DepthAnythingV2-L pre-trained model to obtain the dimension. Feature map ; Feature map These are the inputs to the upsampling module RAFB1 and the fusion enhancement module FE1, respectively. The output of the upsampling module RAFB1 is a module with dimension [missing information]. Feature map The output of the fusion enhancement module FE1 is a dimension of Feature map Among them, the fusion enhancement module FE1 will input dimension as... Feature map Expanded to a 5-dimensional feature map, with dimensions of Then, the output feature map is obtained through deep attention feature enhancement structure, feature map restoration and deconvolution processing. ; The input to the upsampling module RAFB2 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE2 is the feature map. and The output is Feature map ; The input to the upsampling module RAFB3 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE3 is the feature map. and The output is Feature map ; The input to the upsampling module RAFB4 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE4 is the feature map. and The output is Feature map ; The input to the upsampling module RAFB5 is the feature map. The output is a dimension of Feature map The input to the fusion enhancement module FE5 is the feature map. and The output is Feature map ; The fusion enhancement modules FE2 to FE5 have the same internal structure and the same process for processing the two input feature maps: first, the two input feature maps are superimposed, and then the 4-dimensional feature map is expanded into a 5-dimensional feature map through reshaping. Then, the feature map is enhanced by a deep attention feature enhancement structure. The feature map is obtained through processing. Then, the feature map is reshaped. The dimensions are restored to be the same as the two input feature maps; finally, through deconvolution, the length and width of the output feature map are expanded to twice the original, and the number of channels is compressed to half the original. The input to the fusion enhancement module FE6 is a feature map. and The output is a feature map. Among them, the fusion enhancement module FE6 will input dimension as Feature map and The features are stacked, and then processed through feature map expansion, deep attention feature enhancement structure, and feature map restoration to obtain a dimension of [dimensional value missing]. Feature maps; finally, through linear layers Compressing the feature map channel number to 1 results in a dimension of Feature map .

4. The crane operation monitoring method according to claim 3, characterized in that, The three-dimensional attention-modulated convolution includes three different convolutions: For spatial attention convolution, the Gather-Excite algorithm is used; For channel attention convolution, the Squeeze-Excitation algorithm is used; For filter-dimensional attention mechanism convolutions, weight adjustment is performed using convolution kernels based on SE (Search Engine) principles.

5. The crane operation monitoring method according to claim 4, characterized in that, The deep attention feature enhancement structure described in S1 operates on the input as follows: First, for dimension... Feature map Perform a reshape operation to convert its dimensions to... ,in, This represents the total number of spatial locations on the feature map, and is then used as input to two branches: The first branch is a SEBlock connected to a linear layer. The output dimension is Feature map ; In the second branch, first use a step size of The 3×3 convolution is downsampled to obtain a dimension of The feature map is used as input to the linear layer. The output dimension is Feature map Simultaneously, the feature map obtained from the downsampling process is used as input into a cascaded convolutional layer and a linear layer. The output dimension is Feature map ; The obtained feature map Attention is calculated according to the following formula: , in, For splicing operations; For activation functions; ; Represents the vector dimension. As a scaling factor; The result after attention calculation is passed through a linear layer. Processing, output dimension is Feature map .

6. The crane operation monitoring method according to claim 1, characterized in that, The specific operations for geometric correction described in S6 are as follows: A polynomial fitting correction method is used to adjust each control point. depth prediction value Establish the correction function as input: , in, Based on depth prediction Calculated correction depth value; , , and The correction coefficients to be solved are; The data from at least four control points are fitted using the least squares method to minimize the following objective function: , in, This is the actual measured depth value. The number of control points is the basis for solving this optimization problem, which yields the optimal correction coefficients. , , and ; The depth prediction value described in S5 Substituting into the correction function, the corrected depth value is calculated. Generate a corrected depth map; Two control points that were not involved in the coefficient fitting were randomly selected, and the error between their corrected depth values ​​and actual depth values ​​was calculated to ensure that the average absolute error was less than 0.1 meters. If the error exceeded the standard, control points were reselected and the correction coefficients were fitted.

7. The crane operation monitoring method according to claim 6, characterized in that, The specific operation of S7 is as follows: The two-dimensional coordinates on the corrected depth map are converted to a normalized camera coordinate system using the inverse of the intrinsic parameter matrix, and then multiplied by the depth value to obtain the actual three-dimensional coordinates. The three-dimensional point cloud of the entire image is generated using the three-dimensional coordinates calculated pixel by pixel, as expressed by the formula: , Among them, image coordinates The corresponding three-dimensional coordinates of the camera coordinate system are ( ), Camera intrinsic parameter matrix The inverse matrix.

8. The crane operation monitoring method according to claim 7, characterized in that, In S8: Calculate the shortest straight-line distance L from the critical part of the crane boom to the railway area. If L < 5 meters, it is judged as "Level 1 Warning"; if L < 2 meters, it is judged as "Level 2 Warning"; if it violates the limit, a red warning light will be displayed directly in the pop-up window. Based on the three-dimensional points at the top and bottom of the boom, combined with the vertical orientation of the camera, the boom extension length, working radius, and working angle are calculated. Corresponding constraints are set according to the vehicle tonnage, load weight, and construction site requirements. When the constraints are exceeded, an audible and visual warning is issued.

Citation Information

Patent Citations

  • Method and device for detecting personnel under suspension arm based on monocular image and computer equipment

    CN119479015A