Tower crane inspection method, system, device and medium
By using drones to collect multimodal perception data and generate 3D spatial maps and perform cross-modal semantic recognition, the problems of high manpower consumption, low efficiency, and inconsistent accuracy in tower crane inspection have been solved, achieving efficient and accurate tower crane inspection.
Patent Information
- Application Number
- CN202511187398.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-08-25
AI Technical Summary
Existing tower crane inspection methods are labor-intensive, inefficient, and inconsistent in accuracy, relying heavily on the experience of inspection personnel.
The system uses a drone equipped with a dual-channel coaxial imaging device and a lidar to collect multimodal perception data. Through spatial registration, a three-dimensional spatial map is generated. Combined with feature extraction and cross-modal semantic recognition, the system enables automated defect detection and risk assessment of tower crane inspection areas.
It has automated tower crane inspection, reduced manpower consumption, improved inspection efficiency and accuracy, and significantly increased the detection rate of defects in inspected components.
Smart Images

Figure CN120673296B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a tower crane inspection method, system, equipment and medium. Background Technology
[0002] Tower cranes are the most commonly used lifting equipment on construction sites. Also known as tower hoists, they are used to lift materials such as steel bars, timber, concrete, and steel pipes used in construction. They are an indispensable piece of equipment on construction sites.
[0003] The current tower crane inspection method usually involves manual inspection of various parts of the tower crane. This method has high manpower consumption and low inspection efficiency due to the large number of inspection parts. In addition, it relies on the inspection personnel's own experience, resulting in inconsistent accuracy of inspection results. Summary of the Invention
[0004] This application provides a tower crane inspection method, system, equipment, and medium to address the problems of high manpower consumption, low inspection efficiency, and inconsistent accuracy in existing tower crane inspection methods. The technical solution provided in this application is as follows:
[0005] On the one hand, this application provides a tower crane inspection method, including:
[0006] The system acquires multimodal perception data of a target tower crane inspection scenario collected by a drone equipped with a dual-channel coaxial imaging device and a lidar, as well as the drone's flight trajectory data; the multimodal perception data includes RTSP streams and laser point cloud sets.
[0007] A 3D spatial map of the target tower crane inspection scene is obtained by spatially registering the laser point cloud set and flight trajectory data; the bounding boxes of each inspection part of the target tower crane are projected onto the 3D spatial map to obtain local point cloud slices of each inspection part; the geometric measurement values of each inspection part are calculated based on the local point cloud slices of each inspection part; wherein, the bounding boxes of each inspection part are obtained through the BIM model of the target tower crane.
[0008] Preprocessing of RTSP streams and laser point cloud sets yields timestamped visible light image sequences, infrared thermal image sequences, and laser depth image sequences. Feature extraction is performed based on these sequences to obtain visible light texture features, infrared temperature features, and laser depth features. Cross-modal semantic recognition is then performed based on these features to obtain the defect category and confidence level of each inspection location, as well as the defect center location and defect physical quantity regression values.
[0009] The inspection result of each inspection site is obtained by joint risk determination based on the geometric measurement value, defect category and confidence of each inspection site, and the defect center position and defect physical quantity regression value.
[0010] Optionally, the laser point cloud set and flight trajectory data are spatially registered to obtain a three-dimensional space map of the target tower crane inspection scene, including:
[0011] The collection time stamp of each frame of laser point cloud data in the laser point cloud set is aligned with the UTC time stamp of each flight state data in the flight trajectory data to obtain a frame-by-frame correspondence relationship between the laser point cloud data and the flight state data.
[0012] For the laser point cloud data and the flight state data with the frame-by-frame correspondence relationship, the laser point cloud data is transformed into a global coordinate system with the flight state data as the initial value, forming a three-dimensional space map composed of all laser point clouds of the target tower crane and its surrounding scene.
[0013] Optionally, the RTSP stream and the laser point cloud set are preprocessed to obtain a timestamp-synchronized visible light image sequence, an infrared thermal image sequence and a laser depth image sequence, including:
[0014] The RTSP stream is decoded to obtain a timestamp-consistent visible light image sequence and an infrared thermal image sequence.
[0015] Based on the timestamp, an instantaneous laser point cloud slice corresponding to each frame of visible light image in the visible light image sequence is extracted from the laser point cloud set;
[0016] Each instantaneous laser point cloud slice corresponding to each frame of visible light image is converted into a laser depth image to obtain a laser depth image sequence.
[0017] Optionally, feature extraction is performed based on the visible light image sequence, the infrared thermal image sequence and the laser depth image sequence to obtain visible light texture features, infrared temperature features and laser depth features, including:
[0018] A ResNet-50 branch is used to extract features from the visible light image sequence to obtain a visible light texture vector and a visible light texture feature map as visible light texture features;
[0019] A lightweight 3-CNN branch is used to extract features from the infrared thermal image sequence to obtain an infrared temperature vector and an infrared thermal temperature feature map as infrared temperature features;
[0020] A Depth-CNN branch is used to extract features from the laser depth image sequence to obtain a laser depth vector and a laser depth feature map as laser depth features.
[0021] Optionally, cross-modal semantic recognition based on visible light texture features, infrared temperature features, and laser depth features obtains defect categories and confidence of each inspection site, and defect center position and defect physical quantity regression value, including:
[0022] A double-stream encoding mechanism is adopted to perform weighted splicing on visible light texture vectors, infrared temperature vectors, and laser depth vectors in the channel dimension, then perform full connection processing to obtain global semantic feature tokens, and perform weighted splicing on visible light texture feature maps, infrared temperature feature maps, and laser depth feature maps in the channel dimension, then perform slicing processing to obtain each feature image block, and perform linear mapping on each feature image block, then perform 2D position encoding to obtain image block feature token sequences;
[0023] A bidirectional cross-attention mechanism is adopted to perform bidirectional updating on global semantic feature tokens and image block feature token sequences to obtain updated global semantic feature tokens and updated image block feature token sequences;
[0024] A classification head is adopted to predict defect categories and confidence of defect categories of each inspection site based on updated global semantic feature tokens;
[0025] A detection head is adopted to predict defect center positions of each inspection site based on updated image block feature token sequences;
[0026] A regression head is adopted to predict defect physical quantity regression values of each inspection site based on defect center positions of each inspection site.
[0027] Optionally, joint risk determination based on geometric measurement values, defect categories and confidence, and defect center positions and defect physical quantity regression values of each inspection site obtains inspection results of each inspection site, including:
[0028] A rule determination engine is adopted to perform clause-level Boolean logic combination operation on geometric measurement values, defect categories and confidence, and defect center positions and defect physical quantity regression values of each inspection site to obtain inspection results of each inspection site.
[0029] Optionally, the tower crane inspection method provided in the application further includes:
[0030] Based on the inspection results of each inspection site, when there is a dangerous inspection site in each inspection site, the target tower crane is immediately stopped, and a safety hazard early warning message is generated based on the inspection result of the dangerous inspection site and pushed to the terminal of the person in charge;
[0031] Based on the inspection results of each inspection site, when there is an abnormal inspection site in each inspection site, a maintenance work order is generated based on the inspection result of the abnormal inspection site and pushed to the terminal of the maintenance person.
[0032] In another aspect, the application provides a tower inspection system, comprising:
[0033] A data acquisition module is configured to acquire multi-modal perception data of a target tower inspection scene collected by a UAV installed with a dual-path coaxial imaging device and a laser radar and flight trajectory data of the UAV, wherein the multi-modal perception data comprises an RTSP stream and a laser point cloud set.
[0034] A geometry determination module is configured to perform spatial registration on the laser point cloud set and the flight trajectory data to obtain a three-dimensional spatial map of the target tower inspection scene, project bounding boxes of each inspection part of the target tower into the three-dimensional spatial map to obtain local point cloud slices of each inspection part, and calculate geometric measurement values of each inspection part based on the local point cloud slices of each inspection part, wherein the bounding boxes of each inspection part are obtained through a BIM model of the target tower.
[0035] A defect identification module is configured to pre-process the RTSP stream and the laser point cloud set to obtain a timestamp-synchronized visible light image sequence, an infrared thermal image sequence and a laser depth image sequence, perform feature extraction based on the visible light image sequence, the infrared thermal image sequence and the laser depth image sequence to obtain visible light texture features, infrared temperature features and laser depth features, and perform cross-modal semantic recognition based on the visible light texture features, the infrared temperature features and the laser depth features to obtain defect categories and confidence levels of each inspection part as well as defect center positions and defect physical quantity regression values.
[0036] A rule determination module is configured to perform joint risk determination based on the geometric measurement values of each inspection part, the defect categories and confidence levels of each inspection part as well as the defect center positions and the defect physical quantity regression values to obtain inspection results of each inspection part.
[0037] In another aspect, the application provides an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned tower inspection method when executing the computer program.
[0038] In another aspect, the application provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are executed by a processor to implement the above-mentioned tower inspection method.
[0039] The application has the following beneficial effects:
[0040] The application can realize automatic inspection of the target tower crane inspection scene by collecting RTSP stream, laser point cloud set and flight trajectory data by using the unmanned aerial vehicle, and performing defect detection and risk judgment by using the RTSP stream, laser point cloud set and flight trajectory data, and can reduce human consumption and improve inspection efficiency, and by spatial registration of the laser point cloud set and the flight trajectory data, a high-precision three-dimensional space map can be obtained, so that high-precision geometric measurement values can be obtained when calculating the geometric measurement values of each inspection part based on the three-dimensional space map, and then high-quality data basis can be provided for subsequent joint risk judgment, in addition, by performing cross-modal semantic recognition after pre-processing and feature extraction of the RTSP stream and the laser point cloud set, the visible light texture feature, the infrared temperature feature and the laser depth feature three modal features can be complementary, so that the detection rate of the inspection part defect can be significantly improved, and the accuracy of the inspection part defect detection can be improved.
[0041] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and achieved by the structure particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0042] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0043] Figure 1 A schematic diagram of the general process of the tower crane inspection method in the embodiments of the present application is shown;
[0044] Figure 2 A functional structure schematic diagram of the tower crane inspection system in the embodiments of the present application is shown;
[0045] Figure 3 A hardware structure schematic diagram of the electronic device in the embodiments of the present application is shown. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical scheme and beneficial effects of the present application more clear, the technical scheme in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0047] The embodiment of the present application provides a tower machine inspection method, referring to Figure 1 The embodiment of the present application provides a tower machine inspection method, referring to
[0048] Step 101: Obtain the multi-modal perception data of the target tower machine inspection scene collected by the unmanned aerial vehicle and the flight trajectory data of the unmanned aerial vehicle, wherein the multi-modal perception data includes RTSP stream and laser point cloud set.
[0049] In the embodiment of the present application, the visible light camera and the infrared thermal imaging camera are installed on the same rigid support of the unmanned aerial vehicle gimbal, ensuring that the optical axes of the two are completely coincident. The visible light camera is used to collect visible light images related to the appearance of the tower machine components, and the infrared thermal imaging camera is used to collect infrared thermal images related to the temperature of the tower machine components. In addition, the laser radar is installed below the body of the unmanned aerial vehicle, which is used to collect three-dimensional point cloud data of the tower machine and its surrounding environment. The scanning frequency of the laser radar is matched with the flight speed of the unmanned aerial vehicle to ensure the continuity and integrity of the data. In addition, the flight control system arranged inside the unmanned aerial vehicle realizes high-precision positioning and attitude measurement function through the integration of GNSS / INS module and RTK differential module, which can record detailed flight trajectory data, including but not limited to position (longitude, latitude, height), attitude (roll angle, pitch angle, yaw angle) and other 6DoF information. In this way, the RTSP stream and laser point cloud set of the target tower machine inspection scene collected by the unmanned aerial vehicle and the flight trajectory data of the unmanned aerial vehicle can be obtained by controlling the unmanned aerial vehicle to fly. For example, first, a unmanned aerial vehicle with RTK differential module and GNSS / INS module is arranged on site, a laser radar is installed below the body of the unmanned aerial vehicle, and a high-resolution visible light camera and an infrared thermal imaging camera are installed on the same rigid support. The unmanned aerial vehicle hovers at a position about eight meters away from the tower machine, the laser radar emits laser pulses to the tower machine components at a density of one hundred thousand laser points per second to obtain high-density three-dimensional point cloud data, that is, the laser point cloud set, at the same time, the visible light camera and the infrared thermal imaging camera shoot visible light images and infrared thermal images to form RTSP stream at the same time according to the scanning frequency of the laser radar, at the same time, the GNSS / INS module records the flight trajectory data of the unmanned aerial vehicle, the reference station and the mobile station of the RTK differential module communicate in real time to correct the flight trajectory data of the unmanned aerial vehicle in real time, improve the accuracy of the flight trajectory data, so as to obtain the RTSP stream and laser point cloud set of the target tower machine inspection scene and the flight trajectory data of the unmanned aerial vehicle.
[0050] Step 102: spatially register the laser point cloud set and the flight trajectory data to obtain a three-dimensional space map of the target tower crane inspection scene; project the bounding box of each inspection part of the target tower crane into the three-dimensional space map to obtain a local point cloud slice of each inspection part; calculate the geometric measurement value of each inspection part based on the local point cloud slice of each inspection part; wherein the bounding box of each inspection part is obtained through the BIM model of the target tower crane.
[0051] In the embodiments of the present application, when the laser point cloud set and the flight trajectory data are spatially registered to obtain a three-dimensional space map of the target tower crane inspection scene, the following methods can be used, but are not limited to:
[0052] First, align the collection time stamp of each frame of laser point cloud data in the laser point cloud set with the UTC time stamp of each flight state data in the flight trajectory data to obtain the frame-by-frame correspondence between the laser point cloud data and the flight state data.
[0053] Then, for the laser point cloud data and flight state data with frame-by-frame correspondence, the laser point cloud data is transformed into the global coordinate system with the flight state data as the initial value, forming a three-dimensional space map composed of all laser point clouds of the target tower crane and its surrounding scene. Specifically, it includes an initial registration stage and a dynamic optimization stage; in the initial registration stage, the flight state data (i.e. the 6DoF information corrected by the RTK differential module) is used as the initial value to transform each frame of laser point cloud data from the laser radar coordinate system to the global coordinate system. The position information and attitude information in the flight state data are used to calculate the transformation matrix of each frame of laser point cloud data, and the initial registration of each frame of laser point cloud data is performed according to the transformation matrix; in the dynamic optimization stage, after the RTSP stream is preprocessed to obtain the timestamp-synchronized visible light image sequence and infrared thermal image sequence, texture features and temperature features are extracted from the visible light image sequence and infrared thermal image sequence to generate a two-dimensional feature map. The two-dimensional feature map is spatially aligned with the laser point cloud data to form multi-modal feature enhanced laser point cloud data; the ICP (Iterative Closest Point) algorithm is used to perform iterative closest point registration based on the multi-modal feature enhanced laser point cloud data to align the multi-modal feature enhanced laser point cloud data in the same global coordinate system to form spatially aligned laser point cloud data and a high-precision high-density initial three-dimensional space map; the spatially aligned laser point cloud data is used as the node of the graph, the relative pose between the spatially aligned laser point cloud data is used as the edge of the graph, and the pre-constructed geometric constraint, visual constraint and prior constraint are used as the weight of the edge to construct an initial graph model. A nonlinear optimization algorithm (such as CeresSolver) is used to take the initial three-dimensional space map as the starting point of nonlinear optimization, and the initial graph model is nonlinearly optimized to minimize the error of all constraints. Through multiple iterations, the pose of each frame of laser point cloud data is gradually adjusted until the error converges to obtain the final three-dimensional space map of the target tower crane inspection scene; wherein the geometric constraint includes the geometric relationship between the laser point cloud data, such as plane, edge, curvature, etc.; the visual constraint includes the matching relationship between the texture features and the temperature features, such as the correspondence between the feature points; the prior constraint includes prior knowledge such as the position and size of the tower crane components, the position of the surrounding buildings, etc.
[0054] In the embodiments of the present application, when the bounding box of each inspection part of the target tower crane is projected into the three-dimensional space map to obtain the local point cloud slice of each inspection part, the following methods can be used, but are not limited to:
[0055] Firstly, the analysis data of each inspection part is extracted from the BIM model of the target tower crane, including but not limited to the geometric shape, position and size of each inspection part, and based on the analysis data of each inspection part, the bounding box of each inspection part is extracted; wherein the BIM model is usually stored in the IFC (Industry Foundation Classes) format, and tools such as ifcopenshell can be used for analysis to obtain the analysis data of each inspection part; the bounding box is a smallest rectangular frame that can completely contain the geometric shape of the inspection part, and the parameters of the bounding box include the minimum point coordinates and the maximum point coordinates.
[0056] Then, the parameters (minimum point and maximum point coordinates) of the bounding box of each inspection part extracted from the BIM model are mapped to the global coordinate system of the three-dimensional space map, and for each inspection part, the laser point cloud data located within the bounding box of the inspection part is extracted from the three-dimensional space map to form a local point cloud slice of the inspection part.
[0057] Finally, the local point cloud slice of each inspection part is denoised to remove outliers and noise; wherein if the density of the local point cloud slice is not uniform, the point cloud density can be adjusted by voxel filtering or interpolation method to make it more suitable for subsequent analysis and processing.
[0058] In the embodiments of the present application, the geometric measurement values of each inspection part are flexibly set according to actual inspection requirements, such as the geometric measurement values of each inspection part including but not limited to length, diameter, area, volume, deflection, etc.; based on this, when calculating the geometric measurement values of each inspection part based on the local point cloud slice of each inspection part, the following methods can be used but are not limited to:
[0059] (1) For the linear structure inspection part (such as the jib of the tower crane, the pull rod, etc.), the length thereof is calculated: first, the principal component analysis (PCA) is used to extract the principal axis direction of the laser point cloud data, and then the maximum projection length of the laser point cloud data along the principal axis direction is calculated, which is the length of the inspection part.
[0060] (2) For the circular or approximately circular structure inspection part (such as the standard section of the tower crane, the hook, etc.), the diameter thereof is calculated: first, a least squares cylinder is fitted in the laser point cloud data, and the cylinder parameters (center axis direction vector, center point, radius) of the least squares cylinder are obtained, and then the diameter of the inspection part is calculated based on the cylinder parameters (center axis direction vector, center point, radius).
[0061] (3) For the inspection components of flat or approximately flat structure (such as the platform of the tower crane, the attachment frame, etc.), the area thereof is calculated: firstly, a least square plane is fitted in the laser point cloud data to obtain a plane normal vector and an intercept, then, based on the plane normal vector and the intercept, the laser point cloud data is projected onto the least square plane to form a two-dimensional point set, and finally, the area of the inspection component is calculated by using the convex hull or polygon fitting of the two-dimensional point set.
[0062] (4) For the inspection components of three-dimensional structure (such as the components of the tower crane, the concrete structure, etc.), the volume thereof is calculated: firstly, a three-dimensional model (such as a convex hull, a mesh model, etc.) is fitted in the laser point cloud data, and then, based on the fitted three-dimensional model, the volume of the inspection component is calculated.
[0063] (5) For the inspection components with deformation structure (such as the jib of the tower crane, the counterweight arm, etc.), the deflection thereof is calculated: firstly, the geometric features for describing the shape and structure of the inspection component are extracted from the laser point cloud data, including but not limited to the principal direction of the laser point cloud data extracted by PCA, which is helpful to determine the center line of the structure; the curvature of each laser point in the laser point cloud data is calculated, which is helpful to reflect the bending degree of the local surface; the normal vector of each laser point in the laser point cloud data is calculated, which is helpful to determine the orientation of the surface; the point cloud density of the laser point cloud data is calculated, which is helpful to reflect the distribution of the point cloud to identify abnormal areas; then, the center line of the inspection component is extracted based on the geometric features; finally, the maximum deviation distance of each point on the center line to the theoretical straight line is calculated, which is the deflection of the inspection component; wherein, the theoretical straight line refers to the straight line form that the center line of the inspection component (such as the jib of the tower crane, the counterweight arm, etc.) should present in the ideal case without deformation, which can be obtained by measurement or a straight line fitted using the point cloud data as the theoretical straight line, for example, a straight line fitted by extracting the principal direction of the laser point cloud data by PCA as the theoretical straight line.
[0064] Step 103: The RTSP stream and the laser point cloud set are preprocessed to obtain the timestamp-synchronized visible light image sequence, the infrared thermal image sequence and the laser depth image sequence; feature extraction is performed based on the visible light image sequence, the infrared thermal image sequence and the laser depth image sequence to obtain visible light texture features, infrared temperature features and laser depth features; cross-modal semantic recognition is performed based on the visible light texture features, the infrared temperature features and the laser depth features to obtain the defect category and confidence of each inspection component, and the defect center position and defect physical quantity regression value.
[0065] In the embodiments of the present application, when the RTSP stream and the laser point cloud set are preprocessed to obtain the timestamp-synchronized visible light image sequence, the infrared thermal image sequence and the laser depth image sequence, the following methods can be used, but are not limited to:
[0066] First, the RTSP stream is decoded to obtain a timestamp-consistent visible light image sequence and an infrared thermal image sequence.
[0067] Then, based on the timestamp, an instantaneous laser point cloud slice corresponding to each frame of the visible light image sequence is extracted from the laser point cloud set.
[0068] Finally, the instantaneous laser point cloud slice corresponding to each frame of the visible light image sequence is converted into a laser depth image to obtain a laser depth image sequence.
[0069] In the embodiments of the present application, when the visible light texture feature, the infrared temperature feature and the laser depth feature are extracted based on the visible light image sequence, the infrared thermal image sequence and the laser depth image sequence, a three-channel feature extraction network can be used, which includes a ResNet-50 branch, a lightweight 3-CNN branch and a Depth-CNN branch, and specifically includes:
[0070] (1) The ResNet-50 branch is used to extract features from the visible light image sequence to obtain a visible light texture vector and a visible light texture feature map as the visible light texture feature; the visible light image sequence is first rapidly down-sampled through a 7x7 large convolution kernel to expand the RGB three-channel input into a 64-channel low-level texture map; then it is continuously passed through 4 groups of residual blocks, each of which is composed of a 3x3 convolution, a batch normalization and a ReLU, and presents an increasing ladder of 64-256-512-1024-2048 in the number of channels. After passing through each group of residual blocks, the spatial size of the feature map is halved and the number of channels is doubled, gradually capturing edges, textures, component outlines and high-level semantics; finally, the 2048-channel feature map is compressed into a 2048-dimensional compact vector through global average pooling, which is the visible light texture feature, and a two-dimensional visible light texture feature map is also output. During the whole process, the residual connection ensures smooth gradient, and the ResNet-50 branch network can not only maintain detailed texture, but also has sufficient abstract ability.
[0071] (2) The infrared thermal image sequence is extracted by a lightweight 3-CNN branch to obtain an infrared temperature vector and an infrared thermal temperature feature map as infrared temperature features. The infrared thermal image sequence is first compressed to 16 channels by 3x3 convolution, and then enters three consecutive micro residual bottlenecks. Each bottleneck first reduces dimension by 1x1, then extracts temperature gradient by 3x3 depth separable convolution, and then increases dimension by 1x1, so that the channel number always maintains a bottleneck shape of 32-64-32, thereby greatly reducing the computational load. An ECANet channel attention is inserted after the third bottleneck, that is, 32 channel weights are obtained by global average pooling, then the cross-channel relationship is learned by one-dimensional convolution, the feature map is multiplied by the weight, and the high-temperature abnormal area is highlighted. Finally, the 32-channel feature map is compressed into a 512-dimensional compact vector by global average pooling, that is, the infrared temperature feature, and a two-dimensional infrared temperature feature map is output, which retains the details of the thermal anomaly and meets the low-power requirement of the edge.
[0072] (3) The laser depth image sequence is extracted by a Depth-CNN branch to obtain a laser depth vector and a laser depth feature map as laser depth features. The laser depth image sequence is first subjected to 3x3 convolution, and the convolution kernel weight is initialized as a Sobel shape to make the network sensitive to depth edges. Then it enters 4 groups of Depth-aware residual bottlenecks, each of which is reduced in dimension by 1x1, then subjected to 3x3 dilated convolution (dilation=2), and then increased in dimension by 1x1. The residual branch retains the features before the dilated convolution to prevent excessive smoothing. A multi-scale branch is inserted after the second and third bottlenecks to perform 2x2 max pooling and 3x3 average pooling, and then upsampled to the original size and added to the main branch to capture high and low frequency geometric information. Finally, the 512-channel feature map is globally averaged to obtain a 512-dimensional laser depth feature, and a two-dimensional laser depth feature map is output, which completely retains the edge, plane normal and curvature clues, providing accurate geometric descriptions for subsequent cross-modal fusion.
[0073] In the embodiment of the application, when performing cross-modal semantic recognition based on visible light texture features, infrared temperature features and laser depth features to obtain the defect category and confidence of each inspection part and the defect center position and defect physical quantity regression value, a dual-flow Transformer-Decoder multi-task network can be used. The dual-flow Transformer-Decoder multi-task network includes a dual-flow encoder, a cross-modal decoder and a multi-task head, and specifically includes:
[0074] Firstly, the visible light texture vector, the infrared temperature vector and the laser depth vector are weighted and spliced according to the channel dimension to obtain a global semantic feature token, and the visible light texture feature map, the infrared thermal temperature feature map and the laser depth feature map are weighted and spliced according to the channel dimension to obtain each feature image block, and then the linear mapping is performed on each feature image block to obtain an image block feature token sequence. In the embodiment of the application, the visible light texture vector, the infrared temperature vector and the laser depth vector three global feature vectors (2048-d, 512-d, 512-d) are first spliced according to the channel and then linearly projected to generate a cross-modal semantic token, that is, a global token; the visible light texture feature map, the infrared thermal temperature feature map and the laser depth feature map three feature maps are first spliced according to the channel and then cut into fixed size feature image blocks to obtain a feature image block sequence, that is, a patch token sequence. Figure 3 The visible light texture vector, the infrared temperature vector and the laser depth vector three global feature vectors (2048-d, 512-d, 512-d) are first spliced according to the channel and then linearly projected to generate a cross-modal semantic token, that is, a global token; the visible light texture feature map, the infrared thermal temperature feature map and the laser depth feature
[0075] Then, a bidirectional cross-attention mechanism is used to update the global semantic feature token and the image block feature token sequence in a bidirectional manner to obtain an updated global semantic feature token and an updated image block feature token sequence. In the embodiment of the application, the Cross-Modal Transformer Decoder is used to perform bidirectional cross-attention update of the global token ↔ patch tokens by using the Transformer Decoder layer, so as to realize global semantic guiding local details and local details feeding back to global judgment, thereby improving the feature accuracy.
[0076] Finally, a classification head is adopted to predict the defect category and the confidence of the defect category of each inspection part based on the updated global semantic feature token; a detection head is adopted to predict the defect center position of each inspection part based on the updated image block feature token sequence; and a regression head is adopted to predict the defect physical quantity regression value of each inspection part based on the defect center position of each inspection part. In the embodiment of the application, the multi-task head includes the classification head, the detection head and the regression head, the classification head, the detection head and the regression head are connected in parallel to the output end of the Cross-Modal Transformer Decoder, the classification head can output the defect category and the confidence by sequentially passing the global token through a fully connected layer (FC) and an activation function layer (Softmax); the detection head can output the defect center position and the defect width and height by sequentially passing the patch token through a lightweight feature pyramid network (Feature Pyramid Network, FPN) and an anchor-free detection head (Anchor-Free Detection Head); and the regression head can output the defect physical quantity regression value by sequentially passing the patch token of the defect center position given by the detection head through a fully connected layer (FC) and an activation function layer (RectifiedLinear Unit, ReLU).
[0077] Step 104: Joint risk judgment is performed based on the geometric measurement value, the defect category and the confidence, and the defect center position and the defect physical quantity regression value of each inspection part to obtain the inspection result of each inspection part.
[0078] In the embodiment of the application, a rule judgment engine is adopted to perform clause-level Boolean logic combination operation on the geometric measurement value, the defect category and the confidence, and the defect center position and the defect physical quantity regression value of each inspection part to obtain the inspection result of each inspection part. In specific implementation, the following modes can be adopted but are not limited to the following modes:
[0079] Firstly, an itemized rule library is constructed. Specifically, the specification clauses of each inspection part (tower body standard section, hoisting arm, balance arm, slewing support, attachment frame, etc.) of the tower crane are disassembled into atomic judgment items, each atomic judgment item combines the geometric measurement value threshold, the defect category confidence threshold and the physical quantity threshold into three factors, a Boolean expression template is defined, and the Boolean expression template is stored in the rule judgment engine in an XML / JSON structure to form a clause-level Boolean logic rule library; wherein the specification clause refers to a risk judgment rule for evaluating whether the geometric measurement value, the defect category and the confidence, and the physical quantity regression value of the tower machine part meet the safety requirements, which is defined in advance based on industry standards, safety specifications, design requirements, etc.
[0080] Then, the geometric measurement values (length, diameter, deflection, etc.) and the defect category confidence, the defect position coordinates, and the physical quantity regression values are uniformly mapped to the dimensionless interval [0, 1], and based on the geometric measurement values (length, diameter, deflection, etc.) and the defect category confidence, the defect position coordinates, and the physical quantity regression values mapped to the dimensionless interval [0, 1], a linear normalization or an expert scoring function is used to eliminate the dimensional difference.
[0081] Secondly, the inspection part corresponding clause-level Boolean logic rules are extracted from the clause-level Boolean logic rule library according to the inspection part index, and the AND / OR / NOT nodes in the clause-level Boolean logic rules are evaluated layer by layer to form a Boolean result tree; if the Boolean result tree output is true, the conclusion label (such as “over-limit”, “immediate repair”, “observation”, etc.) of the corresponding clause is triggered.
[0082] Finally, the Boolean result is mapped to a machine-readable enumeration value, and the triggered clause number is attached, so as to generate an inspection part-triggered clause-inspection conclusion triple as the inspection result.
[0083] In the embodiments of the present application, when the inspection results of each inspection part are obtained based on the geometric measurement values, the defect category and confidence, and the defect center position and the defect physical quantity regression value, a Bayesian network or a decision tree can also be used instead of the clause-level Boolean logic combination operation to realize probabilistic judgment. Specifically, it includes:
[0084] Firstly, a Bayesian network is constructed with the geometric measurement value node, the defect confidence node, and the physical quantity regression value node as parent nodes and the risk level node as a child node.
[0085] Then, the real-time geometric measurement value, the real-time defect category confidence, and the real-time physical quantity regression value are input into the Bayesian network to obtain the posterior probability distribution of the risk level.
[0086] Secondly, the posterior probability distribution is input into the decision tree model to obtain the risk level and confidence.
[0087] Finally, the risk level and confidence are mapped to the specification clause to generate the inspection result; wherein the specification clause refers to the risk judgment rules for evaluating whether the geometric measurement value, the defect category and its confidence, and the physical quantity regression value of the tower crane component meet the safety requirements, which are predefined based on industry standards, safety specifications, design requirements, etc.
[0088] Further, after the joint risk judgment based on the geometric measurement value of each inspection site, the defect category and confidence, and the defect center position and defect physical quantity regression value is performed to obtain the inspection result of each inspection site, the target tower crane can be immediately stopped when a dangerous inspection site exists in each inspection site based on the inspection result of each inspection site, and a safety hazard warning message can be generated based on the inspection result of the dangerous inspection site and pushed to the terminal of the person in charge; when an abnormal inspection site exists in each inspection site based on the inspection result of each inspection site, a maintenance work order can be generated based on the inspection result of the abnormal inspection site and pushed to the terminal of the maintenance person. In this way, by timely controlling the tower crane to stop or generating a maintenance work order for maintenance, the tower crane operation risk can be effectively reduced, and the safety of the tower crane operation can be improved.
[0089] Based on the above embodiment, the embodiment of the present application provides a tower crane inspection system, as shown in Figure 2 The tower crane inspection device 200 provided by the embodiment of the present application at least includes:
[0090] The data acquisition module 201 is configured to acquire multi-modal perception data of a target tower crane inspection scene collected by a UAV installed with a dual-path coaxial imaging device and a laser radar and flight trajectory data of the UAV; wherein the multi-modal perception data includes an RTSP stream and a laser point cloud set;
[0091] The geometric determination module 202 is configured to perform spatial registration on the laser point cloud set and the flight trajectory data to obtain a three-dimensional space map of the target tower crane inspection scene, project a bounding box of each inspection site of the target tower crane into the three-dimensional space map to obtain a local point cloud slice of each inspection site, and calculate a geometric measurement value of each inspection site based on the local point cloud slice of each inspection site; wherein the bounding box of each inspection site is obtained through a BIM model of the target tower crane.
[0092] The defect identification module 203 is configured to pre-process the RTSP stream and the laser point cloud set to obtain a time-stamped visible light image sequence, an infrared thermal image sequence and a laser depth image sequence; perform feature extraction based on the visible light image sequence, the infrared thermal image sequence and the laser depth image sequence to obtain visible light texture features, infrared temperature features and laser depth features; and perform cross-modal semantic recognition based on the visible light texture features, the infrared temperature features and the laser depth features to obtain a defect category and confidence of each inspection site and a defect center position and defect physical quantity regression value.
[0093] The rule judgment module 204 is configured to perform joint risk judgment based on the geometric measurement value of each inspection site, the defect category and confidence, and the defect center position and defect physical quantity regression value to obtain an inspection result of each inspection site.
[0094] In a possible implementation, the geometry determination module 202 is configured to align the collection time stamp of each frame of laser point cloud data in the laser point cloud set with the UTC time stamp of each piece of flight state data in the flight trajectory data, to obtain a frame-by-frame correspondence between the laser point cloud data and the flight state data; and transform the laser point cloud data into a global coordinate system by taking the flight state data as an initial value, to form a three-dimensional space map composed of all laser point clouds of the target tower crane and the surrounding scene.
[0095] In a possible implementation, the defect identification module 203 is configured to decode the RTSP stream to obtain a time-stamped visible light image sequence and an infrared thermal image sequence; extract, from the laser point cloud set, an instantaneous laser point cloud slice corresponding to each frame of visible light image in the visible light image sequence, based on the time stamp; and convert each instantaneous laser point cloud slice corresponding to each frame of visible light image into a laser depth image to obtain a laser depth image sequence.
[0096] In a possible implementation, the defect identification module 203 is configured to perform feature extraction on the visible light image sequence by using a ResNet-50 branch to obtain a visible light texture vector and a visible light texture feature map as visible light texture features; perform feature extraction on the infrared thermal image sequence by using a lightweight 3-CNN branch to obtain an infrared temperature vector and an infrared thermal temperature feature map as infrared temperature features; and perform feature extraction on the laser depth image sequence by using a Depth-CNN branch to obtain a laser depth vector and a laser depth feature map as laser depth features.
[0097] In a possible implementation, the rule determination module 204 is configured to perform full connection processing on the visible light texture vector, the infrared temperature vector and the laser depth vector after weighted splicing according to a channel dimension by using a double-flow encoding mechanism, to obtain a global semantic feature token, and perform slicing processing on the visible light texture feature map, the infrared thermal temperature feature map and the laser depth feature map after weighted splicing according to a channel dimension, to obtain each feature image block, and perform linear mapping on each feature image block and then perform 2D position encoding to obtain an image block feature token sequence; perform bidirectional updating on the global semantic feature token and the image block feature token sequence by using a bidirectional cross-attention mechanism, to obtain an updated global semantic feature token and an updated image block feature token sequence; predict, by using a classification head, a defect class and a confidence of the defect class of each inspection part based on the updated global semantic feature token; predict, by using a detection head, a defect center position of each inspection part based on the updated image block feature token sequence; and predict, by using a regression head, a defect physical quantity regression value of each inspection part based on the defect center position of each inspection part.
[0098] In a possible implementation, the rule determination module 204 is configured to use a rule determination engine to perform clause-level Boolean logical combination operation on the geometric measurement value, the defect category and the confidence level, and the defect center position and the defect physical quantity regression value of each inspection site, to obtain the inspection result of each inspection site.
[0099] In a possible implementation, the crane inspection system 200 provided by the embodiment of the present application further includes:
[0100] The safety control unit 205 is configured to determine, based on the inspection result of each inspection site, that there is a dangerous inspection site in each inspection site, control the target crane to stop immediately, and generate a safety hazard early warning message based on the inspection result of the dangerous inspection site and push the safety hazard early warning message to the terminal of the person in charge; determine, based on the inspection result of each inspection site, that there is an abnormal inspection site in each inspection site, generate a maintenance work order based on the inspection result of the abnormal inspection site, and push the maintenance work order to the terminal of the maintenance person.
[0101] It should be noted that the principle of the tower crane inspection device provided by the above-mentioned embodiment of the present application for solving the technical problems is similar to the tower crane inspection method provided by the embodiment of the present application, and therefore, the implementation of the tower crane inspection device provided by the embodiment of the present application can be referred to the implementation of the tower crane inspection method provided by the embodiment of the present application, and the repeated parts will not be described herein.
[0102] Next, the electronic device provided by the embodiment of the present application is briefly introduced. The electronic device can be a computer, a tablet computer, a mobile phone, or the like, which is a tower crane inspection device. As shown in FIG. 1, the electronic device 300 provided by the embodiment of the present application at least includes a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301, and the processor 301 implements the above-mentioned tower crane inspection method provided by the embodiment of the present application when executing the computer program. Figure 3
[0103] The electronic device 300 provided by the embodiment of the present application can further include a bus 303 connected to different components (including the processor 301 and the memory 302). The bus 303 represents one or more of several bus structures, including a memory bus, a peripheral bus, a local bus, and the like.
[0104] The memory 302 can include a readable medium in the form of volatile memory, such as a random access memory (RAM) 3021 and / or cache memory 3022, and can further include a read-only memory (ROM) 3023. The memory 302 can also include a program tool 3025 having a set of one or more program modules 3024, including but not limited to an operating system, one or more applications, other program modules, and program data, each of which can include implementation of a network environment, alone or in some combination.
[0105] The processor 301 can be one processing element or a collective term for a plurality of processing elements, for example, the processor 301 can be a microcontroller unit (MCU), or a central processing unit (CPU), or one or more integrated circuits configured to implement the above-mentioned tower crane inspection method provided by the embodiments of the present application. Specifically, the processor 301 can be a general-purpose processor, including but not limited to a CPU, an application specific integrated circuit (ASIC), a ready-to-program gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.
[0106] The electronic device 300 can also communicate with one or more devices that enable a user to interact with the electronic device 300 (for example, a mobile phone, a computer, etc.), and / or with various external devices 304 that enable the electronic device 300 to communicate with one or more other electronic devices (for example, a router, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 305. Furthermore, the electronic device 300 can also communicate with one or more networks (for example, a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) through a network adapter 306. As Figure 3 shown, the network adapter 306 communicates with other modules of the electronic device 300 through the bus 303. It should be understood that although Figure 3Other hardware and / or software modules can be used in conjunction with the electronic device 300, as shown, including, but not limited to, microcode, device drivers, redundant processors, external disk drive arrays, Redundant Arrays of Independent Disks (RAID) subsystems, tape drives, and data backup storage subsystems, etc.
[0107] It should be noted that, Figure 3 The electronic device 300 shown is merely an example and should not limit the function and use range of the embodiments of the present application.
[0108] In addition, the embodiments of the present application further provide a computer readable storage medium, which stores computer instructions, and the computer instructions are executed by a processor to realize the tower crane inspection method provided by the embodiments of the present application. Specifically, the computer instructions can be built-in or installed in the processor, so that the processor can realize the tower crane inspection method provided by the embodiments of the present application by executing the built-in or installed computer instructions.
[0109] Moreover, the tower crane inspection method provided by the embodiments of the present application can also be realized as a program product, which includes program codes, and the program codes are executed by a processor to realize the tower crane inspection method provided by the embodiments of the present application.
[0110] The program product provided by the embodiments of the present application can adopt any combination of one or more readable media, wherein the readable media can be a readable signal medium or a readable storage medium, and the readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above, and more specifically, the more specific examples (non-exhaustive list) of the readable storage medium include an electrical connection with one or more conductive wires, a portable disk, a hard disk, a RAM, a ROM, an Erasable Programmable Read Only Memory (EPROM), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0111] The program product provided by the embodiments of the present application can adopt a CD-ROM and include program codes, and can also run on an electronic device. However, the program product provided by the embodiments of the present application is not limited to this, and in the embodiments of the present application, the readable storage medium can be any tangible medium containing or storing programs, which can be used or combined with an instruction execution system, device or component.
[0112] It should be noted that, although several units or sub-units of the apparatus are mentioned in the above detailed description, such division is merely exemplary and not mandatory. Indeed, according to an embodiment of the present application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided into several units to be embodied.
[0113] Moreover, although the operations of the method(s) herein are described in a particular, sequential order, this order is not meant to be a limitation and is not intended to imply that
[0114] Although preferred embodiments of the application have been described herein, those skilled in the art will readily devise additional variations of these preferred embodiments that fall within the scope of the present application. Accordingly, the appended claims are intended to encompass all such variations as falling within the scope of the present application.
[0115] Obviously, numerous modifications and variations of the present embodiments are possible in light of the above teachings. It is therefore to be understood that within the scope of the claims and their equivalents, the present application can be practiced otherwise than as specifically described herein.
Claims
1. A tower crane inspection method, characterized in that, include: The system acquires multimodal perception data of a target tower crane inspection scenario collected by a drone equipped with a dual-channel coaxial imaging device and a lidar, as well as the drone's flight trajectory data; the multimodal perception data includes RTSP streams and laser point cloud sets. Spatial registration is performed between the laser point cloud set and the flight trajectory data to obtain a three-dimensional spatial map of the target tower crane inspection scene; the bounding boxes of each inspection part of the target tower crane are projected onto the three-dimensional spatial map to obtain local point cloud slices of each inspection part; and geometric measurement values of each inspection part are calculated based on the local point cloud slices of each inspection part. The bounding boxes of each of the inspection locations are obtained through the BIM model of the target tower crane; The RTSP stream and the laser point cloud set are preprocessed to obtain a timestamp-synchronized visible light image sequence, infrared thermal image sequence, and laser depth image sequence. A three-channel feature extraction network is used to extract visible light texture features, infrared temperature features, and laser depth features based on the visible light image sequence, the infrared thermal image sequence, and the laser depth image sequence. The three-channel feature extraction network includes: a ResNet-50 branch for extracting visible light texture features from the visible light image sequence, a lightweight 3-CNN branch for extracting infrared temperature features from the infrared thermal image sequence, and a Depth-CNN branch for extracting laser depth features from the laser depth image sequence. A dual-stream Transformer-Decoder multi-task network is employed to perform cross-modal semantic recognition based on the visible light texture features, the infrared temperature features, and the laser depth features to obtain the defect category, confidence level, defect center location, and defect physical quantity regression value for each of the inspected locations. The dual-stream Transformer-Decoder multi-task network includes: a dual-stream encoder that generates global semantic feature tokens and image block feature token sequences using a dual-stream coding mechanism; a cross-modal decoder that updates the global semantic feature tokens and image block feature token sequences bidirectionally using a bidirectional cross-attention mechanism; and a multi-task head that identifies the defect category, confidence level, defect center location, and defect physical quantity regression value for each of the inspected locations based on the bidirectionally updated global semantic feature tokens and image block feature token sequences. A clause-level Boolean logic rule base is used to jointly determine the inspection results of each inspection location based on the geometric measurement values, defect categories, confidence levels, defect center locations, and defect physical quantity regression values of each inspection location. The clause-level Boolean logic rule base decomposes the standard clauses of each inspection location into atomic decision items, and combines the geometric measurement value threshold, defect category confidence level threshold, and physical quantity threshold of each atomic decision item into a three-factor Boolean expression. The standard clauses are risk determination rules defined based on industry standards, safety specifications, and design requirements to assess whether the geometric measurement values, defect categories and their confidence levels, and physical quantity regression values of the inspected components meet safety requirements. The inspection results include: the inspection location, the triggering clause representing the risk determination rule triggered by the inspection location, and the inspection conclusion representing the safety control method triggered by the inspection location.
2. The tower crane inspection method as described in claim 1, characterized in that, Spatial registration of the laser point cloud set and the flight trajectory data yields a three-dimensional spatial map of the target tower crane inspection scene, including: Align the acquisition timestamp of each frame of laser point cloud data in the laser point cloud set with the UTC timestamp of each flight status data in the flight trajectory data to obtain the frame-by-frame correspondence between laser point cloud data and flight status data. For laser point cloud data and flight status data that have a frame-by-frame correspondence, the flight status data is used as the initial value to transform the laser point cloud data to a global coordinate system, forming a three-dimensional spatial map composed of all laser point clouds of the target tower crane and its surrounding scene.
3. The tower crane inspection method as described in claim 1, characterized in that, Preprocessing the RTSP stream and the laser point cloud set yields a timestamped visible light image sequence, an infrared thermal image sequence, and a laser depth image sequence, including: The RTSP stream is decoded to obtain a visible light image sequence and an infrared thermal image sequence with consistent timestamps; Based on the timestamp, extract the instantaneous laser point cloud slice corresponding to each frame of the visible light image sequence from the laser point cloud set; The laser depth image sequence is obtained by converting the instantaneous laser point cloud slices corresponding to each frame of the visible light image into laser depth images.
4. The tower crane inspection method as described in claim 1, characterized in that, A three-channel feature extraction network is used to extract visible light texture features, infrared temperature features, and laser depth features based on the visible light image sequence, the infrared thermal image sequence, and the laser depth image sequence, including: The visible light image sequence is processed by using the ResNet-50 branch to obtain visible light texture vectors and visible light texture feature maps as the visible light texture features; A lightweight 3-CNN branch is used to extract features from the infrared thermal image sequence to obtain an infrared temperature vector and an infrared thermal temperature feature map as the infrared temperature features. The laser depth image sequence is processed using a Depth-CNN branch to extract features, resulting in a laser depth vector and a laser depth feature map, which are then used as the laser depth features.
5. The tower crane inspection method as described in claim 4, characterized in that, A dual-stream Transformer-Decoder multi-task network is employed to perform cross-modal semantic recognition based on the visible light texture features, the infrared temperature features, and the laser depth features. This results in the defect category and confidence level of each inspected location, as well as the defect center location and defect physical quantity regression values. A dual-stream coding mechanism is adopted. The visible light texture vector, the infrared temperature vector, and the laser depth vector are weighted and concatenated according to the channel dimension and then fully connected to obtain a global semantic feature token. The visible light texture feature map, the infrared thermal temperature feature map, and the laser depth feature map are weighted and concatenated according to the channel dimension and then sliced to obtain each feature image block. Each feature image block is linearly mapped and then 2D positional encoded to obtain an image block feature token sequence. A bidirectional cross-attention mechanism is used to bidirectionally update the global semantic feature token and the image block feature token sequence to obtain the updated global semantic feature token and the updated image block feature token sequence; A classification head is used to predict the defect category and the confidence level of each of the inspection locations based on the updated global semantic feature tokens; Using a detection head, the defect center position of each of the inspected parts is predicted based on the updated image block feature token sequence; A regression head is used to predict the regression value of the physical quantity of defects in each of the inspection locations based on the defect center position of each inspection location.
6. The tower crane inspection method according to any one of claims 1-5, characterized in that, Also includes: Based on the inspection results of each of the inspection locations, when it is determined that there is a dangerous inspection location among the inspection locations, the target tower crane is controlled to stop immediately, and a safety hazard warning message is generated based on the inspection results of the dangerous inspection location and pushed to the responsible person's terminal. Based on the inspection results of each of the inspection locations, when it is determined that there is an abnormal inspection location among the inspection locations, a maintenance work order is generated based on the inspection results of the abnormal inspection location and pushed to the maintenance personnel's terminal.
7. A tower crane inspection system, characterized in that, include: The data acquisition module is used to acquire multimodal perception data of the target tower crane inspection scene collected by the UAV equipped with a dual-channel coaxial imaging device and a lidar, as well as the flight trajectory data of the UAV; the multimodal perception data includes RTSP streams and laser point cloud sets. The geometry determination module is used to spatially register the laser point cloud set and the flight trajectory data to obtain a three-dimensional spatial map of the target tower crane inspection scene, project the bounding boxes of each inspection part of the target tower crane onto the three-dimensional spatial map to obtain local point cloud slices of each inspection part, and calculate the geometric measurement values of each inspection part based on the local point cloud slices of each inspection part; the bounding boxes of each inspection part are obtained through the BIM model of the target tower crane; The defect identification module is used to preprocess the RTSP stream and the laser point cloud set to obtain a timestamp-synchronized visible light image sequence, infrared thermal image sequence, and laser depth image sequence; a three-channel feature extraction network is used to extract visible light texture features, infrared temperature features, and laser depth features based on the visible light image sequence, the infrared thermal image sequence, and the laser depth image sequence; the three-channel feature extraction network includes: a ResNet-50 branch for extracting visible light texture features from the visible light image sequence, a lightweight 3-CNN branch for extracting infrared temperature features from the infrared thermal image sequence, and a Depth-CNN branch for extracting laser depth features from the laser depth image sequence; A dual-stream Transformer-Decoder multi-task network is employed to perform cross-modal semantic recognition based on the visible light texture features, the infrared temperature features, and the laser depth features to obtain the defect category, confidence level, defect center location, and defect physical quantity regression value for each of the inspected locations. The dual-stream Transformer-Decoder multi-task network includes: a dual-stream encoder that generates global semantic feature tokens and image block feature token sequences using a dual-stream coding mechanism; a cross-modal decoder that updates the global semantic feature tokens and image block feature token sequences bidirectionally using a bidirectional cross-attention mechanism; and a multi-task head that identifies the defect category, confidence level, defect center location, and defect physical quantity regression value for each of the inspected locations based on the bidirectionally updated global semantic feature tokens and image block feature token sequences. The rule determination module is used to perform joint risk determination on each of the inspection locations using a clause-level Boolean logic rule base, based on the geometric measurement values, defect categories, confidence levels, defect center locations, and defect physical quantity regression values of each inspection location, to obtain the inspection results for each inspection location. The clause-level Boolean logic rule base is formed by decomposing the standard clauses for each inspection location into atomic decision items, and combining the geometric measurement value threshold, defect category confidence level threshold, and physical quantity threshold of the atomic decision items into a three-factor Boolean expression. The standard clauses are risk determination rules defined based on industry standards, safety specifications, and design requirements to evaluate whether the geometric measurement values, defect categories and their confidence levels, and physical quantity regression values of the inspection components meet safety requirements. The inspection results include: the inspection location, the triggering clause representing the risk determination rule triggered by the inspection location, and the inspection conclusion representing the safety control method triggered by the inspection location.
8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the tower crane inspection method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the tower crane inspection method as described in any one of claims 1-6.
Citation Information
Patent Citations
Automatic inspection method and device for power unmanned aerial vehicle of transformer substation
CN114092537A
Bridge defect automatic identification and inspection method based on deep learning
CN120259915A