Crawler crane real-time detection method and system for surface defects of steel wire rope during installation and disassembly based on deep learning
Patent Information
- Application Number
- CN202611031004.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]解决履带吊安拆作业中钢丝绳表面缺陷实时检测时,因钢丝绳线速度变化导致固定时间感受野的时序特征提取方式与变速运动不匹配,从而造成低速时引入冗余信息、高速时遗漏时序关联特征,影响检测准确性和稳定性的问题
本发明通过深度可分离卷积构建主干网络,并以线速度自适应的一维时间空洞卷积提取时序运动特征,使得时间感受野随钢丝绳运动速度动态调整,在变速工况下既能捕获高速运动时缺陷的大范围依赖,又能保留低速运动时的细节变化,在保持模型轻量化的前提下,有效提升了实时检测的精度与鲁棒性。
Abstract
Description
Technical Field
[0001] This invention belongs to the field of hoisting operation safety monitoring and computer vision technology, specifically involving a method for real-time detection of surface defects in wire ropes for crawler crane installation and dismantling based on deep learning. Background Technology
[0002] In crawler crane installation and dismantling operations, real-time detection of surface defects in wire ropes using video images is of practical significance for ensuring safety. Deep learning-based visual inspection methods typically utilize temporal information between consecutive video frames to capture the motion characteristics of defects, improving the ability to distinguish defects from surface textures or stains. Currently common temporal feature extraction methods often employ fixed-size 3D convolutional kernels or recursive structures with fixed time steps, whose temporal receptive field remains constant throughout the processing. However, under crawler crane installation and dismantling conditions, the linear velocity of the wire rope is not constant, exhibiting significant speed changes during lifting, lowering, speed changes, and braking. When the wire rope's speed is low, the displacement of defects between frames is minimal, and a fixed, large temporal receptive field will contain many frames with repetitive content, introducing redundant information and easily obscuring or distorting the true, subtle motion characteristics of defects. When the wire rope's speed is high, the spatial displacement of defects between frames increases, and a fixed, small temporal receptive field cannot fully cover the temporal changes of defects, resulting in insufficient extracted temporal correlation information and an increased probability of missing high-speed moving defects. This mismatch between the fixed temporal receptive field and the variable-speed motion affects the accuracy and stability of the detection method under variable-speed conditions. Meanwhile, real-time detection places high demands on lightweight models, limiting their computational capacity and complexity, making it difficult to adapt to speed changes by stacking multi-scale temporal branches or adding complex motion compensation modules. Therefore, the lack of effective means to dynamically adapt temporal feature extraction to changes in the linear velocity of the wire rope while maintaining lightweight design and real-time performance constitutes a significant technical challenge in current video detection of surface defects in wire ropes. Summary of the Invention
[0003] One object of the embodiments of the present invention is to solve at least the above-mentioned problems and / or defects, and to provide at least the advantages described below.
[0004] To address the issue of mismatch between the temporal feature extraction method of the fixed-time sensing field and the variable-speed motion caused by the change in the linear velocity of the wire rope during real-time detection of surface defects in crawler crane installation and dismantling operations, which leads to the introduction of redundant information at low speeds and the omission of temporal correlation features at high speeds, thus affecting the accuracy and stability of detection.
[0005] This addresses the problem that existing visual detection inputs fail to distinguish between the tension and relaxation states of steel wire ropes, leading to the neglect of differences in surface features of images from different stress states. This makes it difficult for the model to use mechanical state information to assist in defect identification, thus limiting the robustness of defect recognition.
[0006] This paper addresses the problem that existing wire rope surface defect detection systems do not integrate a speed-adaptive temporal feature extraction mechanism into a lightweight deep learning model, resulting in the system struggling to balance real-time detection efficiency and detection accuracy under variable speed conditions, and lacking a complete engineering implementation architecture.
[0007] Another objective of this invention is to provide a method for real-time detection of surface defects in wire ropes used for installation and dismantling of crawler cranes based on deep learning.
[0008] To achieve the above-mentioned objectives, the present invention employs the following technical solution: A deep learning-based real-time detection method for surface defects in wire ropes used in the installation and dismantling of crawler cranes includes: Step S1: During the installation and dismantling operation of the crawler crane, collect multiple consecutive video images of the wire rope surface and obtain the current linear velocity of the wire rope in real time. Step S2: Input the consecutive multi-frame video images into a preset lightweight deep learning detection model, wherein the lightweight deep learning detection model includes: The backbone feature extraction network is composed of multiple depthwise separable convolutional modules stacked together. It is used to extract spatial features independently frame by frame from each video image and output the spatial feature map corresponding to each frame. The temporal feature fusion module stacks the multi-frame spatial feature maps output by the backbone feature extraction network into a three-dimensional feature volume along the time dimension, and uses a one-dimensional temporal convolution kernel to slide along the time dimension to convolve the three-dimensional feature volume to extract the temporal motion features between frames and generate a temporally enhanced feature map. The detection output module performs defect target classification and bounding box regression on the time-enhanced feature map, and outputs the category and location information of the steel wire rope surface defects; When the temporal feature fusion module extracts temporal motion features by sliding a one-dimensional temporal convolution kernel along the time dimension, it determines the temporal convolution dilation rate based on the current linear velocity of the wire rope, and performs dilated convolution on the three-dimensional feature volume along the time dimension according to the temporal convolution dilation rate. The temporal convolution dilation rate is configured to increase with the increase of linear velocity, so that the temporal receptive field of the one-dimensional temporal convolution kernel expands with the increase of the wire rope's movement speed.
[0009] Preferably, in the deep learning-based real-time detection method for surface defects of wire ropes used in the installation and dismantling of crawler cranes, step S1, which involves acquiring multiple consecutive frames of video images of the wire rope surface, specifically includes: Multiple industrial cameras deployed in the area of wire rope drums and pulley blocks synchronously collect video image streams of the wire rope surface and obtain the operation status signal of the crawler crane installation and dismantling control system in real time. The operation status signal indicates at least that the wire rope is in a slack state or under tension. An initial image sequence is formed by extracting multiple consecutive frames from the video image stream, and the operation status signal is sampled according to the timestamp of each frame to generate a status label that is time-aligned with each frame. The status labels corresponding to each frame are expanded into single-channel matrices with the same spatial size as the image of that frame. The single-channel matrices are then stitched along the channel dimension to the pixel channels of the corresponding frame image as additional channels, forming a multi-channel image with a status channel for each frame. The continuous multi-frame video image is composed of the multi-channel images of multiple consecutive frames.
[0010] Preferably, in the deep learning-based real-time detection method for surface defects of wire ropes used in the installation and dismantling of crawler cranes, the generation of status labels time-aligned with each frame of the image further includes: The operation status signal is sampled according to the image acquisition frame rate to obtain an initial binary status identifier aligned with the timestamp of each frame image. The initial binary status identifier indicates whether the wire rope is currently in a slack state or a tension state. The current linear velocity of the wire rope is obtained in real time from the crawler crane installation and dismantling control system, and the smoothing time constant associated with the motion inertia is calculated using the current linear velocity. The initial binary state identifier sequence is subjected to a time-domain exponentially weighted moving average using the smoothing time constant to generate a state confidence value that varies continuously between 0 and 1. The state confidence value represents the probability that the wire rope is in a state of tension in the corresponding frame. Furthermore, the step of expanding the state labels corresponding to each frame into a single-channel matrix with the same spatial size as the image of that frame specifically involves expanding the state confidence value into a single-channel matrix with the same spatial size as the input image frame, and then using this single-channel matrix as the additional channel to be concatenated to the pixel channel of the corresponding frame image.
[0011] Preferably, in the deep learning-based real-time detection method for surface defects of wire ropes used in the installation and dismantling of crawler cranes, the multiple industrial camera devices are deployed at different azimuth angles around the axis of the wire rope, and the fields of view of each camera device overlap in the circumference of the wire rope. The step of extracting multiple consecutive frames from a video image stream to form an initial image sequence further includes: For each synchronous acquisition moment, acquire multiple image frames synchronously acquired by the multiple industrial camera devices; Using the pre-calibrated internal and external parameters of each camera device, distortion correction is performed on the multi-channel image frames, and perspective transformation is performed based on the cylindrical surface unfolding model with the steel wire rope axis as the reference, so that each image frame is mapped to a unified steel wire rope surface unfolding coordinate system. The mapped multi-frame images are stitched together sequentially along the circumference of the steel wire rope, and weighted fusion is performed through overlapping areas to generate a panoramic unfolded image corresponding to that moment. The initial image sequence is composed of panoramic unfolded images from multiple consecutive synchronous moments, and is used for subsequent state label expansion and the formation of multi-frame input image sequences; Furthermore, the weighted fusion through overlapping regions includes: For any pixel in the overlapping region, its fusion weight is determined linearly based on the proportion of the shortest distance from that pixel to the boundary of its original image in the overlap width; The pixel values from two adjacent images are weighted and summed according to their fusion weights, and the sum of the two weights is always 1.
[0012] Preferably, in the deep learning-based real-time detection method for surface defects of wire ropes used in the installation and dismantling of crawler cranes, each depthwise separable convolutional module in the backbone feature extraction network consists of the following structure: A channel-wise convolutional layer is used to perform spatial convolution on each input channel independently; A channel attention unit receives the output of the channel-wise convolutional layer, compresses spatial information through global average pooling and global max pooling respectively, generates channel attention weights through a shared fully connected layer, and then performs channel attention weighting with the output of the channel-wise convolutional layer to highlight channel features that are sensitive to defect textures. A pointwise convolutional layer receives the feature map weighted by the channel attention unit, performs inter-channel information fusion, and adjusts the number of output channels; In the training phase, the fully connected layer in the channel attention unit is configured to use wire rope image samples containing ultra-fine defects such as early wire breakage and micro-cracks as supervision data, and update the parameters of the fully connected layer through backpropagation to enhance the feature response of the channel attention unit to the ultra-fine defects.
[0013] Preferably, in the deep learning-based real-time detection method for surface defects of wire ropes used in the installation and dismantling of crawler cranes, step S2, which involves determining the temporal convolution dilation rate based on the current linear velocity of the wire rope and performing dilated convolution on the three-dimensional feature volume along the time dimension according to the temporal convolution dilation rate, specifically includes: In the temporal feature fusion module, when performing one-dimensional dilated convolution on the three-dimensional feature volume position by position along the time dimension, for each convolution center position in the time dimension, the current linear velocity aligned with the timestamp of that convolution center position is obtained; Based on a pre-established mapping relationship between linear velocity and expansion rate, the current linear velocity is mapped to a corresponding expansion rate value, wherein the mapping relationship is configured such that the expansion rate value monotonically and non-decreasing as the linear velocity increases. The dilation rate value is used to perform a dilated convolution operation on the current convolution center position, so that the receptive field of the one-dimensional temporal convolution kernel at different time positions is adaptively adjusted as the speed of the wire rope changes.
[0014] Preferably, in the deep learning-based real-time detection method for surface defects of wire ropes used in the installation and dismantling of crawler cranes, step S2, before the temporal feature fusion module stacks multiple frames of spatial feature maps along the time dimension into a three-dimensional feature volume, further includes: The current direction of movement of the wire rope is obtained in real time from the crawler crane installation and dismantling control system; When the current direction of motion indicates that the wire rope is moving in the opposite direction, the multi-frame spatial feature map output by the backbone feature extraction network is rearranged in reverse order in the time dimension of the frame sequence to obtain the multi-frame spatial feature map after direction compensation. The multi-frame spatial feature maps after direction compensation are stacked along the time dimension to form the three-dimensional feature volume, so that the temporal motion features extracted by the one-dimensional temporal convolution kernel are not affected by the reversal of the wire rope's motion direction.
[0015] Preferably, in the deep learning-based real-time detection method for surface defects of wire ropes used in the installation and dismantling of crawler cranes, the step of calculating the smoothing time constant related to motion inertia using the current linear velocity includes: The smoothing time constant is calculated according to the formula τ = max(τ0, k / v); Where v is the current linear velocity, k is the preset proportional coefficient, and τ0 is the preset lower limit of the smoothing time constant to avoid the smoothing time constant from tending to infinity when the wire rope is stationary or at extremely low speed.
[0016] Preferably, in the deep learning-based real-time detection method for surface defects of wire ropes used in the installation and dismantling of crawler cranes, when the current direction of motion indicates that the wire rope is in reverse motion, the multi-frame spatial feature maps output by the backbone feature extraction network are rearranged in reverse order along the time dimension of the frame sequence, specifically including: Obtain the current linear velocity aligned with the timestamp of the spatial feature map for each frame; The reverse rearrangement is performed only when the current direction of motion indicates reverse motion and the absolute value of the current linear velocity is greater than a preset direction switching speed threshold. Otherwise, the original temporal order of the multi-frame spatial feature maps is maintained.
[0017] A deep learning-based real-time detection system for surface defects in wire ropes used for installation and dismantling of crawler cranes, used to implement any of the methods described above, includes: The image acquisition device is used to acquire multiple consecutive video images of the surface of the wire rope during the installation and dismantling operations of the crawler crane; A linear velocity acquisition device is used to acquire the current linear velocity of the wire rope in real time; A data processing device is equipped with a lightweight deep learning detection model, the lightweight deep learning detection model comprising: The backbone feature extraction network is composed of multiple depthwise separable convolutional modules stacked together. It is used to extract spatial features independently frame by frame from each video image and output the spatial feature map corresponding to each frame. The temporal feature fusion module is used to stack the multi-frame spatial feature maps output by the backbone feature extraction network into a three-dimensional feature volume along the time dimension, and use a one-dimensional temporal convolution kernel to slide along the time dimension to convolve the three-dimensional feature volume to extract the temporal motion features between frames and generate a temporally enhanced feature map. The detection output module is used to classify defects and regress bounding boxes on the time-enhanced feature map, and output the category and location information of defects on the wire rope surface. When the temporal feature fusion module extracts temporal motion features by sliding the one-dimensional temporal convolution kernel along the time dimension, it determines the temporal convolution dilation rate based on the current linear velocity obtained by the linear velocity acquisition device, and performs dilated convolution on the three-dimensional feature volume along the time dimension according to the temporal convolution dilation rate. The temporal convolution dilation rate is configured to increase with the increase of linear velocity, so that the temporal receptive field of the one-dimensional temporal convolution kernel expands with the increase of the wire rope movement speed.
[0018] Compared with the prior art, the advantages and beneficial technical effects of the present invention are: This invention constructs a backbone network through depthwise separable convolutions and extracts temporal motion features using linear velocity adaptive one-dimensional temporal dilated convolutions. This allows the temporal receptive field to be dynamically adjusted according to the speed of the wire rope. Under variable speed conditions, it can capture the large-scale dependence of defects during high-speed motion while retaining the detailed changes during low-speed motion. While maintaining the lightweight nature of the model, it effectively improves the accuracy and robustness of real-time detection.
[0019] This invention expands the state of the wire rope under slack or tension into a state channel and splices it into the pixel channel of each frame image, so that the model input simultaneously contains visual texture information and mechanical state information. This helps the model learn the differentiated visual representation of defects under different stress conditions, thereby enhancing the ability to distinguish between defects and normal surface textures in complex working conditions.
[0020] This invention introduces a smooth time constant based on linear velocity to perform a time-domain exponentially weighted moving average on the binary state label, generating a continuously changing state confidence value as an additional channel. This enables the state information to reflect the gradual change process of the wire rope's stress state, avoids abrupt changes in state labels caused by instantaneous switching, and improves the temporal consistency between state features and the actual stress level of the image frame.
[0021] This invention generates a complete panoramic image of the wire rope by unfolding the images acquired by multiple cameras into a cylindrical surface, stitching them together, and linearly weighting and fusing the overlapping areas. This eliminates blind spots in a single-view perspective and smooths the discontinuities at the stitching joints, providing a complete, uniform, and boundary-free unfolded image of the wire rope surface for subsequent defect detection.
[0022] This invention enhances the sensitivity of the backbone network to weak defects such as early wire breakage and microcracks by embedding channel attention units containing global averaging and max pooling between channel-wise convolution and point-wise convolution in the depthwise separable convolution module, and by supervising the learning of its fully connected layer parameters with ultra-fine defect samples. This significantly improves the feature extraction capability of early micro-defects while controlling the amount of computation.
[0023] This invention, during temporal feature fusion, determines the dilation rate value position by position for each convolution center position in the time dimension based on the time stamp-aligned linear velocity and a preset mapping relationship. This concretizes the adaptive speed adjustment into a point-by-point dilated convolution implementation, ensuring the independence and accuracy of receptive field adjustment at each position in the time dimension.
[0024] This invention detects the direction of the steel wire rope's movement before constructing a three-dimensional feature body, and reverses the temporal dimension of multiple frames of spatial feature maps during reverse movement. This ensures that the one-dimensional temporal convolution kernel always extracts temporal features along the same equivalent direction of movement, eliminating the disruption to the temporal motion description caused by the reversal of the movement direction and guaranteeing feature consistency and direction insensitivity.
[0025] This invention avoids the calculation failure problem caused by the smoothing time constant tending to infinity when the wire rope is stationary or at extremely low speed by setting a lower limit value for the smoothing time constant and using a maximum value function constraint calculation formula. This ensures that there is a definite and usable time constant for state smoothing processing in the full speed range, including the zero speed state.
[0026] This invention filters out invalid direction switching caused by low-speed micro-motion or signal jitter by adding a linear velocity threshold condition to the direction reversal determination. The frame sequence is reversed only when the reverse motion speed substantially exceeds the preset threshold, which improves the stability of the temporal feature fusion output and the consistency of the detection results.
[0027] This invention integrates an image acquisition device, a linear velocity acquisition device, and a data processing device with a speed-adaptive temporal feature fusion detection model into a complete real-time detection system. It realizes an integrated engineering architecture from multi-source image acquisition and speed perception to variable speed adaptive defect detection, enabling the system to have both real-time performance and high precision in the installation and dismantling of crawler cranes.
[0028] Other advantages, objectives, and features of the embodiments of the present invention will be apparent in part from the following description, and in part will be understood by those skilled in the art through study and practice of the embodiments of the present invention. Detailed Implementation
[0029] To further illustrate the technical means and effects of this invention, the following embodiments are provided for further explanation. The specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention.
[0030] It should be noted that, unless otherwise specified, the experimental methods described in the following implementation plan are all conventional methods, and the reagents and materials described are all commercially available unless otherwise specified.
[0031] According to one embodiment of the present invention, a method for real-time detection of surface defects in wire ropes for installation and dismantling of crawler cranes based on deep learning includes: Step S1: During the installation and dismantling operation of the crawler crane, collect multiple consecutive video images of the wire rope surface and obtain the current linear velocity of the wire rope in real time. Step S2: Input the consecutive multi-frame video images into a preset lightweight deep learning detection model, wherein the lightweight deep learning detection model includes: The backbone feature extraction network is composed of multiple depthwise separable convolutional modules stacked together. It is used to extract spatial features independently frame by frame from each video image and output the spatial feature map corresponding to each frame. The temporal feature fusion module stacks the multi-frame spatial feature maps output by the backbone feature extraction network into a three-dimensional feature volume along the time dimension, and uses a one-dimensional temporal convolution kernel to slide along the time dimension to convolve the three-dimensional feature volume to extract the temporal motion features between frames and generate a temporally enhanced feature map. The detection output module performs defect target classification and bounding box regression on the time-enhanced feature map, and outputs the category and location information of the steel wire rope surface defects; When the temporal feature fusion module extracts temporal motion features by sliding a one-dimensional temporal convolution kernel along the time dimension, it determines the temporal convolution dilation rate based on the current linear velocity of the wire rope, and performs dilated convolution on the three-dimensional feature volume along the time dimension according to the temporal convolution dilation rate. The temporal convolution dilation rate is configured to increase with the increase of linear velocity, so that the temporal receptive field of the one-dimensional temporal convolution kernel expands with the increase of the wire rope's movement speed.
[0032] In crawler crane installation and dismantling operations, when using deep learning for real-time defect detection on wire rope surface videos, a common approach is to employ temporal feature extraction with a fixed temporal receptive field. This can be achieved by processing consecutive frames using a 3D convolutional kernel of constant size or a loop structure with a fixed time step. However, when the wire rope moves at low speeds, the displacement of defects between adjacent frames is minimal. A fixed and large temporal receptive field will cover a large number of highly repetitive frames, introducing redundant information. This makes it easy for the temporal changes of subtle defects to be obscured, resulting in the model being insensitive to early wire breaks or microcracks. Conversely, when the wire rope moves at high speeds, the spatial position of defects changes rapidly between frames. A fixed and small temporal receptive field cannot fully capture its motion trajectory, resulting in insufficient temporal correlation information and potential missed detections of high-speed moving defects. This mismatch between the temporal receptive field and the movement speed limits the accuracy and stability of the detection method under variable speed conditions.
[0033] To overcome the aforementioned problems, this embodiment presents a real-time detection method for surface defects of wire ropes used in crawler crane installation and dismantling based on deep learning. The method involves: during crawler crane installation and dismantling operations, acquiring multiple consecutive frames of video images of the wire rope surface using cameras deployed in the drum and pulley block areas, while simultaneously obtaining the current linear velocity of the wire rope from the control system in real time. These consecutive frames are then input into a lightweight deep learning detection model. This model first consists of a backbone feature extraction network composed of multiple depthwise separable convolutional modules stacked together, extracting spatial feature maps for each frame independently to obtain multi-frame spatial feature maps. Subsequently, a temporal feature fusion module stacks the multi-frame spatial feature maps along the time dimension to form a three-dimensional feature volume, and uses a one-dimensional temporal convolution kernel to slide along the time dimension to extract temporal motion features between frames. Crucially, a fixed dilation rate is not used in the convolution operation; instead, the temporal convolution dilation rate is determined based on the current linear velocity of the wire rope at each moment, and then a dilated convolution is performed on the three-dimensional feature volume along the time dimension according to this dilation rate. The dilation rate is configured to increase with increasing linear velocity, allowing the temporal receptive field of the one-dimensional temporal convolution kernel to adaptively expand with the increasing speed of the wire rope. The temporally enhanced feature map is then used by the detection output module to classify the defect category and regress the bounding box, outputting the location and type of the surface defect.
[0034] Taking a video clip of a steel wire rope surface showing both low-speed lifting and high-speed lowering processes as an example, when the steel wire rope is in the low-speed phase, the linear velocity value is small. Based on a pre-defined monotonically non-decreasing mapping relationship, a small dilation rate value is obtained. The receptive field of the one-dimensional temporally dilated convolution is correspondingly narrowed, aggregating features from only a few temporally adjacent frames. This accurately captures the minute temporal changes of the defect during slow movement, avoiding interference from irrelevant information between distant frames. When the steel wire rope enters the high-speed movement phase, the linear velocity increases, and the dilation rate automatically increases. The convolution kernel skips adjacent frames in the temporal dimension and directly acts on frames with larger intervals, effectively expanding the temporal receptive field. This allows for capturing the overall movement pattern of the defect across a longer time span, ensuring that the temporal features of high-speed defects are not lost. Throughout this process, depthwise separable convolution reduces the computational load of spatial feature extraction, while the one-dimensional temporally dilated convolution itself has a relatively low computational burden. The overall model remains lightweight, meeting the requirements of real-time detection.
[0035] This implementation method dynamically adjusts the temporal convolution dilation rate driven by linear velocity, ensuring that the same set of one-dimensional temporal convolution kernels maintains a consistent temporal receptive field that matches the actual motion scale of the defect throughout the entire variable-speed motion of the wire rope. At low speeds, it avoids redundancy and highlights details; at high speeds, it expands the field of view and prevents omissions. Thus, without significantly increasing model complexity, it systematically solves the feature extraction mismatch problem caused by variable speed, improving the adaptability and reliability of the detection method under various operating conditions of crawler crane installation and dismantling.
[0036] According to one embodiment of the present invention, in the real-time detection method for surface defects of wire ropes for installation and dismantling of crawler cranes based on deep learning, step S1, which involves acquiring multiple consecutive frames of video images of the wire rope surface, specifically includes: Multiple industrial cameras deployed in the area of wire rope drums and pulley blocks synchronously collect video image streams of the wire rope surface and obtain the operation status signal of the crawler crane installation and dismantling control system in real time. The operation status signal indicates at least that the wire rope is in a slack state or under tension. An initial image sequence is formed by extracting multiple consecutive frames from the video image stream, and the operation status signal is sampled according to the timestamp of each frame to generate a status label that is time-aligned with each frame. The status labels corresponding to each frame are expanded into single-channel matrices with the same spatial size as the image of that frame. The single-channel matrices are then stitched along the channel dimension to the pixel channels of the corresponding frame image as additional channels, forming a multi-channel image with a status channel for each frame. The continuous multi-frame video image is composed of the multi-channel images of multiple consecutive frames.
[0037] Specifically, in implementation, multiple industrial camera devices are deployed around the wire rope axis at different azimuth angles in the drum and pulley block area, with some overlap in the field of view of each device along the circumference of the wire rope. These cameras synchronously acquire video image streams of the wire rope surface, and simultaneously obtain real-time operation status signals from the crawler crane installation and dismantling control system. These signals can at least clearly indicate whether the wire rope is currently in a slack or tensile state. Multiple consecutive frames are extracted from each video image stream at a set frame rate to form an initial image sequence. For each frame, the operation status signal is sampled based on its timestamp to generate a status label aligned with the frame's time. Then, the status label corresponding to each frame is expanded into a single-channel matrix with the same spatial dimensions as the frame image. This single-channel matrix is then used as an additional channel and directly stitched along the channel dimension to the original pixel channels of the corresponding frame image, thus forming a multi-channel image where each frame has a status channel. These multiple consecutive multi-channel images ultimately constitute the multiple consecutive video images input into the subsequent detection model. This method, by labeling state as an additional channel registered with the image space, enables the backbone feature extraction network of the model to simultaneously perceive whether the image of the wire rope is acquired under tension or relaxation when extracting spatial features frame by frame. Its beneficial effect lies in the fact that the model input simultaneously contains both visual texture and mechanical state information. The network can learn the differentiated appearance of defects under different stress conditions, such as cracks widening under tension and surface texture shrinkage potentially masking microcracks under relaxation. This significantly enhances the ability to distinguish between real defects and normal surface textures or pseudo-defects in complex and variable working conditions, improving the robustness of detection.
[0038] According to one embodiment of the present invention, in the real-time detection method for surface defects of wire ropes for installation and dismantling of crawler cranes based on deep learning, the generation of status labels time-aligned with each frame of image further includes: The operation status signal is sampled according to the image acquisition frame rate to obtain an initial binary status identifier aligned with the timestamp of each frame image. The initial binary status identifier indicates whether the wire rope is currently in a slack state or a tension state. The current linear velocity of the wire rope is obtained in real time from the crawler crane installation and dismantling control system, and the smoothing time constant associated with the motion inertia is calculated using the current linear velocity. The initial binary state identifier sequence is subjected to a time-domain exponentially weighted moving average using the smoothing time constant to generate a state confidence value that varies continuously between 0 and 1. The state confidence value represents the probability that the wire rope is in a state of tension in the corresponding frame. Furthermore, the step of expanding the state labels corresponding to each frame into a single-channel matrix with the same spatial size as the image of that frame specifically involves expanding the state confidence value into a single-channel matrix with the same spatial size as the input image frame, and then using this single-channel matrix as the additional channel to be concatenated to the pixel channel of the corresponding frame image.
[0039] Specifically, after sampling the operation status signal according to the image frame rate to obtain an initial binary status identifier aligned with the timestamp of each frame, this instantaneous value, which only represents the two extreme states of relaxation or tension, is not used directly. Instead, the current linear velocity of the wire rope is further obtained in real time from the crawler crane installation and dismantling control system. Using this current linear velocity, a smoothing time constant associated with motion inertia is calculated through a preset relationship. Specifically, the formula τ is equal to the maximum value of τ0 and k divided by v, where v is the current linear velocity, k is the proportionality coefficient, and τ0 is the lower limit of the smoothing time constant to prevent numerical overflow when the wire rope is stationary or at extremely low speed. Then, a time-domain exponentially weighted moving average operation is performed on the initial binary status identifier sequence using this smoothing time constant, thereby transforming the originally abrupt binary sequence into a state confidence value that continuously varies between 0 and 1. This value represents the probability that the wire rope is actually in a tension state in the corresponding frame. Subsequently, this state confidence value is expanded into a single-channel matrix with the same spatial size as the input image frame and is stitched as an additional channel after the pixel channel. In contrast, existing processing methods often directly use the switching signals of the control system as state labels. These labels undergo abrupt changes between 0 and 1 at the moment tension is established or dissipated. However, the change in the stress state of the wire rope is actually a gradual process due to mechanical inertia. This leads to a temporal mismatch between the state labels and the actual stress level in the image frame, especially near state transition boundaries, where the model is prone to learning incorrect associations. This method introduces velocity-driven adaptive exponential smoothing, transforming the state labels into continuously changing confidence levels. Its advantages include: the state information accurately describes the gradual transition of the wire rope stress; the smoothing time constant increases as the speed decreases, consistent with the actual physical inertia characteristic of slow stress changes at low speeds; avoiding noise caused by abrupt label changes; and ensuring that the prior information provided by the state channel is highly consistent with visual features in the temporal dimension. This further enhances the stability and accuracy of the detection model during condition switching.
[0040] According to one embodiment of the present invention, in the real-time detection method for surface defects of steel wire rope for installation and dismantling of crawler crane based on deep learning, the plurality of industrial camera devices are deployed at different azimuth angles around the axis of the steel wire rope, and the fields of view of each camera device overlap in the circumference of the steel wire rope. The step of extracting multiple consecutive frames from a video image stream to form an initial image sequence further includes: For each synchronous acquisition moment, acquire multiple image frames synchronously acquired by the multiple industrial camera devices; Using the pre-calibrated internal and external parameters of each camera device, distortion correction is performed on the multi-channel image frames, and perspective transformation is performed based on the cylindrical surface unfolding model with the steel wire rope axis as the reference, so that each image frame is mapped to a unified steel wire rope surface unfolding coordinate system. The mapped multi-frame images are stitched together sequentially along the circumference of the steel wire rope, and weighted fusion is performed through overlapping areas to generate a panoramic unfolded image corresponding to that moment. The initial image sequence is composed of panoramic unfolded images from multiple consecutive synchronous moments, and is used for subsequent state label expansion and the formation of multi-frame input image sequences; Furthermore, the weighted fusion through overlapping regions includes: For any pixel in the overlapping region, its fusion weight is determined linearly based on the proportion of the shortest distance from that pixel to the boundary of its original image in the overlap width; The pixel values from two adjacent images are weighted and summed according to their fusion weights, and the sum of the two weights is always 1.
[0041] Specifically, multiple industrial cameras are deployed at different azimuth angles around the axis of the steel wire rope, and the fields of view of each camera overlap circumferentially around the steel wire rope. During implementation, for each synchronous acquisition moment, multiple image frames simultaneously acquired by these cameras are first obtained. Using the intrinsic and extrinsic parameters obtained from pre-calibration of each camera, distortion correction is performed on each image frame to eliminate geometric distortion caused by the lens. Subsequently, based on a cylindrical surface unfolding model established with the steel wire rope axis as the reference, perspective transformation is performed on the corrected images, mapping the texture information originally projected onto the cylindrical surface to a unified steel wire rope surface unfolding coordinate system, resulting in multiple rectangular unfolded images. These mapped images are arranged sequentially along the circumference of the steel wire rope, and weighted fusion is performed in the overlapping areas of adjacent images to generate a complete panoramic unfolded image corresponding to that synchronous moment. Specifically, the weighted fusion of overlapping regions involves calculating the shortest distance from any pixel in the overlapping region to the boundary of its original image. The proportion of this distance to the entire overlap width is the fusion weight of that pixel. The pixel values from two adjacent images are multiplied by their corresponding weights and then summed, with the sum of the two weights always being 1. This achieves a uniform transition from one image to the other. The initial image sequence is composed of panoramic unfolded images generated at multiple consecutive synchronous moments. Subsequently, state labels are extended on this basis to form multi-frame input. This method generates a circumferential panoramic unfolded image and provides a linear, seamless fusion strategy, providing the subsequent detection model with a continuous, uniform, and high-quality image of the entire circumferential surface of the wire rope without redundant boundaries. This allows any surface defect, regardless of its circumferential location or whether it crosses the camera's field of view, to be perceived and analyzed simultaneously on a complete and unified image. This completely eliminates blind spots and seam interference, significantly improving the completeness and accuracy of detecting circumferentially distributed defects in the wire rope.
[0042] According to one embodiment of the present invention, in the real-time detection method for surface defects of steel wire ropes for installation and dismantling of crawler cranes based on deep learning, each depthwise separable convolutional module in the backbone feature extraction network is composed of the following structure: A channel-wise convolutional layer is used to perform spatial convolution on each input channel independently; A channel attention unit receives the output of the channel-wise convolutional layer, compresses spatial information through global average pooling and global max pooling respectively, generates channel attention weights through a shared fully connected layer, and then performs channel attention weighting with the output of the channel-wise convolutional layer to highlight channel features that are sensitive to defect textures. A pointwise convolutional layer receives the feature map weighted by the channel attention unit, performs inter-channel information fusion, and adjusts the number of output channels; In the training phase, the fully connected layer in the channel attention unit is configured to use wire rope image samples containing ultra-fine defects such as early wire breakage and micro-cracks as supervision data, and update the parameters of the fully connected layer through backpropagation to enhance the feature response of the channel attention unit to the ultra-fine defects.
[0043] Specifically, in constructing the backbone feature extraction network of the lightweight deep learning detection model, the specific structure of each depthwise separable convolutional module is as follows: a channel-wise convolutional layer, a channel attention unit, and a pointwise convolutional layer. The channel-wise convolutional layer independently applies a spatial convolution kernel to each channel of the input feature map, extracting spatial texture information while maintaining channel separation. The channel attention unit receives the output of the channel-wise convolutional layer and first performs global average pooling and global max pooling in parallel, compressing the spatial information of each channel into an average descriptor value and a maximum descriptor value, respectively. These two values are then fed into a shared fully connected layer to generate a set of channel attention weights. These weights are multiplied channel-wise by the original output of the channel-wise convolutional layer, recalibrating the responses of different channels, highlighting channels sensitive to defect textures, and suppressing irrelevant background channels. Subsequently, the pointwise convolutional layer receives the weighted feature map and completes the fusion of inter-channel information and adjusts the number of output channels through a linear combination across channels. During the training phase, high-resolution wire rope image samples containing ultra-fine defects such as early wire breakage and microcracks are specifically used as supervisory data. The parameters of the fully connected layers in the channel attention units are updated by backpropagation of detection loss, forcing the attention mechanism to learn to capture weak feature patterns associated with these minute defects. This method, by inserting channel attention units trained with supervision from ultra-fine defect samples between channel-wise and point-wise convolutions, has the beneficial effect of actively enhancing the sensitivity of the backbone network to channel feature responses to ultra-fine defects such as early wire breakage and microcracks, which are easily missed, without significantly increasing the number of model parameters. This ensures that weak but critically discriminative texture signals are not drowned out during feature propagation, thereby improving the detection capability of early wire rope damage.
[0044] According to one embodiment of the present invention, in the real-time detection method for surface defects of wire ropes for installation and dismantling of crawler cranes based on deep learning, step S2, which involves determining the temporal convolution dilation rate based on the current linear velocity of the wire rope and performing dilated convolution on the three-dimensional feature body along the time dimension according to the temporal convolution dilation rate, specifically includes: In the temporal feature fusion module, when performing one-dimensional dilated convolution on the three-dimensional feature volume position by position along the time dimension, for each convolution center position in the time dimension, the current linear velocity aligned with the timestamp of that convolution center position is obtained; Based on a pre-established mapping relationship between linear velocity and expansion rate, the current linear velocity is mapped to a corresponding expansion rate value, wherein the mapping relationship is configured such that the expansion rate value monotonically and non-decreasing as the linear velocity increases; in an optional implementation, the pre-established mapping relationship between linear velocity and expansion rate can be obtained through the following calibration process: Within the typical speed range of crawler crane installation and dismantling operations, several discrete linear velocity values are selected. For each linear velocity value, several different candidate values of expansion rate are tried on a set of continuously acquired wire rope defect sample sequences, and the detection accuracy is tested. The expansion rate that achieves the highest defect detection accuracy is selected as the expansion rate value at that linear velocity. Each linear velocity and its corresponding expansion rate value are recorded in a mapping table, or a continuous mapping function is formed through curve fitting. This mapping function satisfies that the expansion rate value monotonically and non-decreasing with increasing linear velocity.
[0045] In actual deployment, the control system can dynamically determine the temporal convolution dilation rate based on the current real-time linear velocity using the mapping table or mapping function.
[0046] The dilation rate value is used to perform a dilated convolution operation on the current convolution center position, so that the receptive field of the one-dimensional temporal convolution kernel at different time positions is adaptively adjusted as the speed of the wire rope changes.
[0047] Specifically, during the temporal feature fusion module's one-dimensional dilated convolution of the 3D feature volume along the time dimension, instead of using a globally uniform dilation rate for each convolution center position in the time dimension, the current linear velocity of the wire rope, aligned with the timestamp of that convolution center position, is acquired in real time. The system pre-constructs a mapping curve from linear velocity to dilation rate, configured so that the dilation rate value monotonically and non-decreasing with increasing linear velocity. For example, when the wire rope is stationary or moving at extremely low speeds, the dilation rate is set to 1, degenerating into a standard one-dimensional tight convolution with a compact receptive field; as the linear velocity gradually increases, the dilation rate increases accordingly to 2, 3, or higher. During convolution, the mapping relationship is queried based on the linear velocity corresponding to the current convolution center position to obtain the dilation rate value for that position. Then, according to this dilation rate, several intermediate frames are skipped in the time dimension, and convolution operations are performed on frames at intervals. Thus, the receptive field of the one-dimensional temporal convolution kernel at different positions in the time dimension independently and adaptively adjusts according to the speed of the wire rope at that moment. Compared to using a single fixed expansion rate at the module input level or the overall feature volume level, this approach achieves position-by-position dynamic control of the temporal receptive field. In this embodiment, the adjustment of the temporal receptive field precisely corresponds to the actual working conditions of each frame. The expansion rate is large in high-speed frames, aggregating long-term motion cues, while the expansion rate is small in low-speed frames, focusing on short-term micro-variable details. This ensures that the extraction of temporal features throughout the entire video segment is always in an optimal state that matches the motion speed, significantly improving the consistency and accuracy of defect description under variable speed conditions.
[0048] According to one embodiment of the present invention, in the real-time detection method for surface defects of wire ropes for installation and dismantling of crawler cranes based on deep learning, in step S2, before the temporal feature fusion module stacks multiple frames of spatial feature maps along the time dimension into a three-dimensional feature volume, the method further includes: The current direction of movement of the wire rope is obtained in real time from the crawler crane installation and dismantling control system; When the current direction of motion indicates that the wire rope is moving in the opposite direction, the multi-frame spatial feature map output by the backbone feature extraction network is rearranged in reverse order in the time dimension of the frame sequence to obtain the multi-frame spatial feature map after direction compensation. The multi-frame spatial feature maps after direction compensation are stacked along the time dimension to form the three-dimensional feature volume, so that the temporal motion features extracted by the one-dimensional temporal convolution kernel are not affected by the reversal of the wire rope's motion direction.
[0049] Specifically, before performing temporal feature fusion, in addition to acquiring the current linear velocity of the wire rope, the current direction of movement of the wire rope is also acquired in real time from the crawler crane installation and dismantling control system. This direction indicates whether the wire rope is rotating forward to retract or rotating backward to release. When the current direction of movement indicates that the wire rope is moving in the opposite direction, and the absolute value of its current linear velocity is greater than a preset direction switching speed threshold, the multi-frame spatial feature maps output by the backbone feature extraction network are reversed in the time dimension of the frame sequence. That is, the original chronological order of the feature maps is reversed to obtain the direction-compensated multi-frame spatial feature maps. If the current direction of movement is not reversed, or although it is reversed but the absolute value of the linear velocity is lower than the threshold, the original temporal order of the multi-frame spatial feature maps is maintained. After the reverse reordering, the direction-compensated feature maps are stacked along the time dimension to form a three-dimensional feature volume for subsequent one-dimensional temporal convolution kernels to extract temporal motion features. This implementation method, through motion direction sensing and threshold-based reverse rearrangement, ensures that regardless of whether the wire rope moves in the forward or reverse direction, the time series observed by the one-dimensional temporal convolution kernel is essentially normalized to a motion process in the same equivalent direction. The extracted temporal motion features are insensitive to the motion direction, eliminating feature distortion or self-cancellation caused by direction reversal. This guarantees the physical consistency of defect motion description and the stability and reliability of detection results during bidirectional rope retraction and release operations. According to one embodiment of the present invention, in the real-time detection method for surface defects of wire ropes for installation and dismantling of crawler cranes based on deep learning, the step of calculating the smoothing time constant associated with motion inertia using the current linear velocity includes: The smoothing time constant is calculated according to the formula τ = max(τ0, k / v); Where v is the current linear velocity, k is the preset proportional coefficient, and τ0 is the preset lower limit of the smoothing time constant to avoid the smoothing time constant from tending to infinity when the wire rope is stationary or at extremely low speed.
[0050] Specifically, in the process of generating state labels aligned with the time of each frame of the image, the smoothing time constant associated with motion inertia is calculated using the current linear velocity. The formula τ is equal to the maximum value of τ0 and k divided by v. Here, v is the current linear velocity of the wire rope obtained in real time from the crawler crane's installation and dismantling control system, k is a preset proportional coefficient, and τ0 is a preset lower limit value for the smoothing time constant. When the wire rope is moving normally and the linear velocity v is within the normal operating range, the quotient of k divided by v is greater than τ0, and the smoothing time constant is determined by k divided by v. At this point, the higher the linear velocity, the smaller the smoothing time constant, and the faster the state confidence responds to changes in force. This aligns with the actual physical process where the force state of the wire rope changes rapidly during high-speed motion. Conversely, the lower the linear velocity, the larger the quotient of k divided by v, the larger the smoothing time constant, and the slower the change in state confidence. This conforms to the characteristic that the force state transitions smoothly due to inertia during low-speed motion. When the wire rope is stationary or its linear velocity approaches zero, the theoretical value of k divided by v will approach infinity. However, due to the formula constraint of taking the maximum value of τ0 and k divided by v, τ0 is used as the smoothing time constant, avoiding numerical overflow or smoothing failure. In contrast, if only a simple inverse proportional relationship is used to calculate the smoothing time constant without setting a lower limit, the smoothing time constant will become extremely large or even uncalculate under the condition that the wire rope is stationary or at extremely low speeds. This will cause the exponentially weighted moving average to fail to effectively update the state confidence or produce numerical instability. This method, by setting a lower limit and constraining with a maximum value function, has the advantage of obtaining a definite, usable, and physically meaningful smoothing time constant across the entire speed range, including the zero-speed state. This ensures the continued effectiveness of state inference at low speeds and when stationary, while maintaining the ability to sensitively track state changes at high speeds, thus providing a stable and reliable time smoothing foundation for subsequent state channel construction.
[0051] According to one embodiment of the present invention, in the real-time detection method for surface defects of wire ropes for installation and dismantling of crawler cranes based on deep learning, when the current direction of movement indicates that the wire rope is in reverse movement, the multi-frame spatial feature maps output by the backbone feature extraction network are rearranged in reverse order along the time dimension of the frame sequence, specifically including: Obtain the current linear velocity aligned with the timestamp of the spatial feature map for each frame; The reverse rearrangement is performed only when the current direction of motion indicates reverse motion and the absolute value of the current linear velocity is greater than a preset direction switching speed threshold. Otherwise, the original temporal order of the multi-frame spatial feature maps is maintained.
[0052] Specifically, after acquiring the current direction of the wire rope in real time from the crawler crane installation and dismantling control system, the process for reversing the direction is not an unconditional reordering of the frame sequence. Instead, it first acquires the current linear velocity aligned with the timestamp of each frame's spatial feature map and performs conditional checks. Only when two conditions are met simultaneously is the multi-frame spatial feature map output by the backbone feature extraction network reordered in the time dimension: first, the current direction of motion indicates that the wire rope is moving in the opposite direction; second, the absolute value of the current linear velocity is greater than a preset direction switching speed threshold. In this case, the original time order of the multi-frame spatial feature maps is completely reversed to obtain a direction-compensated feature map sequence. If the current direction of motion is positive, or if the current direction of motion is negative but the absolute value of the linear velocity does not reach the threshold, the original time order of the multi-frame spatial feature maps remains unchanged, and the process directly proceeds to the subsequent 3D feature volume stacking step. For example, if the preset direction switching speed threshold is 0.1 meters per second, when the wire rope moves slightly in the reverse direction at an extremely low speed of 0.05 meters per second, although the motion direction indication is reversed, the absolute value of the linear velocity is below the threshold, so no reverse reordering is triggered, and the multi-frame spatial feature map still maintains its original temporal order. When the wire rope moves in the reverse direction at a normal speed of 0.5 meters per second, the absolute value of the linear velocity exceeds the threshold, and reverse reordering is then performed. In contrast, if no speed threshold is set, and reverse reordering is triggered directly based on the motion direction signal, under conditions where the direction indication frequently jumps due to the low-speed slight movement of the wire rope, crawling, or slight jitter in the sensor signal, the temporal order of the frame sequence will repeatedly switch between forward and reverse. This causes the internal temporal structure of the three-dimensional feature body received by the temporal feature fusion module to continuously fluctuate, leading to instability in the detection output results and positional jumps of the same defect at different times. This method introduces a direction switching speed threshold as a constraint. Its beneficial effect is that it effectively filters out false direction switching caused by low-speed micro-motion and signal noise. Reverse reordering is only performed when the wire rope undergoes a truly meaningful reverse movement, ensuring the stability of the temporal structure of the time sequence feature fusion input. This allows the defect detection results to remain continuous, consistent, and traceable even under complex operating conditions with frequent direction switching.
[0053] According to one embodiment of the present invention, a deep learning-based real-time detection system for surface defects of wire ropes used in the installation and dismantling of crawler cranes is provided, which is used to implement the method described in any one of the above, comprising: The image acquisition device is used to acquire multiple consecutive video images of the surface of the wire rope during the installation and dismantling operations of the crawler crane; A linear velocity acquisition device is used to acquire the current linear velocity of the wire rope in real time; A data processing device is equipped with a lightweight deep learning detection model, the lightweight deep learning detection model comprising: The backbone feature extraction network is composed of multiple depthwise separable convolutional modules stacked together. It is used to extract spatial features independently frame by frame from each video image and output the spatial feature map corresponding to each frame. The temporal feature fusion module is used to stack the multi-frame spatial feature maps output by the backbone feature extraction network into a three-dimensional feature volume along the time dimension, and use a one-dimensional temporal convolution kernel to slide along the time dimension to convolve the three-dimensional feature volume to extract the temporal motion features between frames and generate a temporally enhanced feature map. The detection output module is used to classify defects and regress bounding boxes on the time-enhanced feature map, and output the category and location information of defects on the wire rope surface. When the temporal feature fusion module extracts temporal motion features by sliding the one-dimensional temporal convolution kernel along the time dimension, it determines the temporal convolution dilation rate based on the current linear velocity obtained by the linear velocity acquisition device, and performs dilated convolution on the three-dimensional feature volume along the time dimension according to the temporal convolution dilation rate. The temporal convolution dilation rate is configured to increase with the increase of linear velocity, so that the temporal receptive field of the one-dimensional temporal convolution kernel expands with the increase of the wire rope movement speed.
[0054] A real-time detection system for surface defects of wire ropes during the installation and dismantling of crawler cranes, based on deep learning, is constructed as follows: An image acquisition device is deployed in the area of the wire rope drum and pulley block of the crawler crane. This device continuously acquires multiple frames of video images of the wire rope surface during the installation and dismantling operation. Simultaneously, a linear velocity acquisition device is set up, which communicates with the crawler crane's installation and dismantling control system to obtain the current linear velocity of the wire rope in real time. A data processing device with computing power is deployed on-site or at the near-field edge. The data processing device loads and runs a lightweight deep learning detection model. This lightweight deep learning detection model structurally comprises three components. The backbone feature extraction network consists of multiple stacked depthwise separable convolutional modules, which independently extract spatial features from each frame of the input video image, outputting a spatial feature map corresponding to each frame. The temporal feature fusion module receives multi-frame spatial feature maps output by the backbone feature extraction network, stacks them along the time dimension to form a three-dimensional feature volume, and uses a one-dimensional temporal convolution kernel to convolve this three-dimensional feature volume along the time dimension to extract inter-frame temporal motion features, generating a temporally enhanced feature map. The detection output module performs defect target classification and bounding box regression on the temporally enhanced feature map, finally outputting the category and location information of defects on the wire rope surface. Notably, the temporal feature fusion module does not use a fixed convolution dilation rate during operation. Instead, it determines the temporal convolution dilation rate based on the current linear velocity provided in real-time by the linear velocity acquisition device, and performs dilated convolution on the three-dimensional feature volume along the time dimension according to this dilation rate. This temporal convolution dilation rate is configured to increase with increasing linear velocity, allowing the temporal receptive field of the one-dimensional temporal convolution kernel to adaptively expand with the increasing speed of the wire rope movement. In contrast, existing wire rope surface defect detection systems typically only use linear velocity data for speed marking during video playback or for post-analysis. The temporal feature extraction part of the detection model itself uses a fixed convolution kernel size or a fixed time step, and the image acquisition and algorithm inference stages are disconnected in terms of speed perception. This system integrates the image acquisition device, linear velocity acquisition device, and data processing device with a speed-adaptive temporal feature fusion model into a closed-loop whole. Its advantages lie in realizing an integrated engineering architecture from multi-source image perception, real-time speed perception to variable speed adaptive defect inference. Speed information drives the dynamic adjustment of the temporal receptive field within the detection model in real time, enabling the system to maintain a balance between efficient real-time processing and high-precision defect detection capabilities under various speed-changing conditions throughout the entire process of crawler crane installation and dismantling. This meets the stringent requirements of real-time performance and adaptability in actual hoisting operations.
[0055] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. It can be applied to various fields suitable for the embodiments of the present invention. Other modifications can be readily implemented by those skilled in the art. Therefore, without departing from the general concept defined by the claims and their equivalents, the embodiments of the present invention are not limited to the specific details.
Claims
1. A method for real-time detection of surface defects in wire ropes used for installation and dismantling of crawler cranes based on deep learning, characterized in that, include: Step S1: During the installation and dismantling operation of the crawler crane, collect multiple consecutive video images of the wire rope surface and obtain the current linear velocity of the wire rope in real time. Step S2: Input the consecutive multi-frame video images into a preset lightweight deep learning detection model, wherein the lightweight deep learning detection model includes: The backbone feature extraction network is composed of multiple depthwise separable convolutional modules stacked together. It is used to extract spatial features independently frame by frame from each video image and output the spatial feature map corresponding to each frame. The temporal feature fusion module stacks the multi-frame spatial feature maps output by the backbone feature extraction network into a three-dimensional feature volume along the time dimension, and uses a one-dimensional temporal convolution kernel to slide along the time dimension to convolve the three-dimensional feature volume to extract the temporal motion features between frames and generate a temporally enhanced feature map. The detection output module performs defect target classification and bounding box regression on the time-enhanced feature map, and outputs the category and location information of the steel wire rope surface defects; When the temporal feature fusion module extracts temporal motion features by sliding a one-dimensional temporal convolution kernel along the time dimension, it determines the temporal convolution dilation rate based on the current linear velocity of the wire rope, and performs dilated convolution on the three-dimensional feature volume along the time dimension according to the temporal convolution dilation rate. The temporal convolution dilation rate is configured to increase with the increase of linear velocity, so that the temporal receptive field of the one-dimensional temporal convolution kernel expands with the increase of the wire rope's movement speed.
2. The method for real-time detection of surface defects in wire ropes for tracked crane installation and dismantling based on deep learning as described in claim 1, characterized in that, In step S1, acquiring multiple consecutive video images of the steel wire rope surface specifically includes: Multiple industrial cameras deployed in the area of wire rope drums and pulley blocks synchronously collect video image streams of the wire rope surface and obtain the operation status signal of the crawler crane installation and dismantling control system in real time. The operation status signal indicates at least that the wire rope is in a slack state or under tension. An initial image sequence is formed by extracting multiple consecutive frames from the video image stream, and the operation status signal is sampled according to the timestamp of each frame to generate a status label that is time-aligned with each frame. The status labels corresponding to each frame are expanded into single-channel matrices with the same spatial size as the image of that frame. The single-channel matrices are then stitched along the channel dimension to the pixel channels of the corresponding frame image as additional channels, forming a multi-channel image with a status channel for each frame. The continuous multi-frame video image is composed of the multi-channel images of multiple consecutive frames.
3. The method for real-time detection of surface defects in the wire rope of a crawler crane based on deep learning as described in claim 2, characterized in that, The generation of state labels aligned with the time of each frame of image further includes: The operation status signal is sampled according to the image acquisition frame rate to obtain an initial binary status identifier aligned with the timestamp of each frame image. The initial binary status identifier indicates whether the wire rope is currently in a slack state or a tension state. The current linear velocity of the wire rope is obtained in real time from the crawler crane installation and dismantling control system, and the smoothing time constant associated with the motion inertia is calculated using the current linear velocity. The initial binary state identifier sequence is subjected to a time-domain exponentially weighted moving average using the smoothing time constant to generate a state confidence value that varies continuously between 0 and 1. The state confidence value represents the probability that the wire rope is in a state of tension in the corresponding frame. Furthermore, the step of expanding the state labels corresponding to each frame into a single-channel matrix with the same spatial size as the image of that frame specifically involves expanding the state confidence value into a single-channel matrix with the same spatial size as the input image frame, and then using this single-channel matrix as the additional channel to be concatenated to the pixel channel of the corresponding frame image.
4. The method for real-time detection of surface defects in the installation and dismantling wire rope of a crawler crane based on deep learning as described in claim 2, characterized in that, The multiple industrial camera devices are deployed at different azimuth angles around the axis of the wire rope, and the fields of view of each camera device overlap in the circumference of the wire rope. The step of extracting multiple consecutive frames from a video image stream to form an initial image sequence further includes: For each synchronous acquisition moment, acquire multiple image frames synchronously acquired by the multiple industrial camera devices; Using the pre-calibrated internal and external parameters of each camera device, distortion correction is performed on the multi-channel image frames, and perspective transformation is performed based on the cylindrical surface unfolding model with the steel wire rope axis as the reference, so that each image frame is mapped to a unified steel wire rope surface unfolding coordinate system. The mapped multi-frame images are stitched together sequentially along the circumference of the steel wire rope, and weighted fusion is performed through overlapping areas to generate a panoramic unfolded image corresponding to that moment. The initial image sequence is composed of panoramic unfolded images from multiple consecutive synchronous moments, and is used for subsequent state label expansion and the formation of multi-frame input image sequences; Furthermore, the weighted fusion through overlapping regions includes: For any pixel in the overlapping region, its fusion weight is determined linearly based on the proportion of the shortest distance from that pixel to the boundary of its original image in the overlap width; The pixel values from two adjacent images are weighted and summed according to their fusion weights, and the sum of the two weights is always 1.
5. The method for real-time detection of surface defects in wire ropes for installation and dismantling of crawler cranes based on deep learning as described in any one of claims 3 or 4, characterized in that, Each depthwise separable convolutional module in the backbone feature extraction network consists of the following structure: A channel-wise convolutional layer is used to perform spatial convolution on each input channel independently; A channel attention unit receives the output of the channel-wise convolutional layer, compresses spatial information through global average pooling and global max pooling respectively, generates channel attention weights through a shared fully connected layer, and then performs channel attention weighting with the output of the channel-wise convolutional layer to highlight channel features that are sensitive to defect textures. A pointwise convolutional layer receives the feature map weighted by the channel attention unit, performs inter-channel information fusion, and adjusts the number of output channels; In the training phase, the fully connected layer in the channel attention unit is configured to use wire rope image samples containing ultra-fine defects such as early wire breakage and micro-cracks as supervision data, and update the parameters of the fully connected layer through backpropagation to enhance the feature response of the channel attention unit to the ultra-fine defects.
6. The method for real-time detection of surface defects in wire ropes for tracked crane installation and dismantling based on deep learning as described in claim 1, characterized in that, Step S2, which involves determining the temporal convolution dilation rate based on the current linear velocity of the wire rope and performing dilated convolution on the three-dimensional feature volume along the time dimension according to this temporal convolution dilation rate, specifically includes: In the temporal feature fusion module, when performing one-dimensional dilated convolution on the three-dimensional feature volume position by position along the time dimension, for each convolution center position in the time dimension, the current linear velocity aligned with the timestamp of that convolution center position is obtained; Based on a pre-established mapping relationship between linear velocity and expansion rate, the current linear velocity is mapped to a corresponding expansion rate value, wherein the mapping relationship is configured such that the expansion rate value monotonically and non-decreasing as the linear velocity increases. The dilation rate value is used to perform a dilated convolution operation on the current convolution center position, so that the receptive field of the one-dimensional temporal convolution kernel at different time positions is adaptively adjusted as the speed of the wire rope changes.
7. The method for real-time detection of surface defects in wire ropes for tracked crane installation and dismantling based on deep learning as described in claim 1, characterized in that, In step S2, before the temporal feature fusion module stacks the multi-frame spatial feature maps along the time dimension into a three-dimensional feature volume, it further includes: The current direction of movement of the wire rope is obtained in real time from the crawler crane installation and dismantling control system; When the current direction of motion indicates that the wire rope is moving in the opposite direction, the multi-frame spatial feature map output by the backbone feature extraction network is rearranged in reverse order in the time dimension of the frame sequence to obtain the multi-frame spatial feature map after direction compensation. The multi-frame spatial feature maps after direction compensation are stacked along the time dimension to form the three-dimensional feature volume, so that the temporal motion features extracted by the one-dimensional temporal convolution kernel are not affected by the reversal of the wire rope's motion direction.
8. The method for real-time detection of surface defects in wire ropes for tracked crane installation and dismantling based on deep learning as described in claim 3, characterized in that, The calculation of the smoothed time constant associated with motion inertia using the current linear velocity includes: The smoothing time constant is calculated according to the formula τ = max(τ0, k / v); Where v is the current linear velocity, k is the preset proportional coefficient, and τ0 is the preset lower limit of the smoothing time constant to avoid the smoothing time constant from tending to infinity when the wire rope is stationary or at extremely low speed.
9. The method for real-time detection of surface defects in wire ropes for tracked crane installation and dismantling based on deep learning as described in claim 7, characterized in that, When the current direction of motion indicates that the wire rope is moving in the opposite direction, the multi-frame spatial feature maps output by the backbone feature extraction network are rearranged in reverse order along the time dimension of the frame sequence, specifically including: Obtain the current linear velocity aligned with the timestamp of the spatial feature map for each frame; The reverse rearrangement is performed only when the current direction of motion indicates reverse motion and the absolute value of the current linear velocity is greater than a preset direction switching speed threshold. Otherwise, the original temporal order of the multi-frame spatial feature maps is maintained.
10. A real-time detection system for surface defects of wire ropes used in the installation and dismantling of crawler cranes based on deep learning, characterized in that, It is used to implement the method according to any one of claims 1 to 9, comprising: The image acquisition device is used to acquire multiple consecutive video images of the surface of the wire rope during the installation and dismantling operations of the crawler crane; A linear velocity acquisition device is used to acquire the current linear velocity of the wire rope in real time; A data processing device is equipped with a lightweight deep learning detection model, the lightweight deep learning detection model comprising: The backbone feature extraction network is composed of multiple depthwise separable convolutional modules stacked together. It is used to extract spatial features independently frame by frame from each video image and output the spatial feature map corresponding to each frame. The temporal feature fusion module is used to stack the multi-frame spatial feature maps output by the backbone feature extraction network into a three-dimensional feature volume along the time dimension, and use a one-dimensional temporal convolution kernel to slide along the time dimension to convolve the three-dimensional feature volume to extract the temporal motion features between frames and generate a temporally enhanced feature map. The detection output module is used to classify defects and regress bounding boxes on the time-enhanced feature map, and output the category and location information of defects on the wire rope surface. When the temporal feature fusion module extracts temporal motion features by sliding the one-dimensional temporal convolution kernel along the time dimension, it determines the temporal convolution dilation rate based on the current linear velocity obtained by the linear velocity acquisition device, and performs dilated convolution on the three-dimensional feature volume along the time dimension according to the temporal convolution dilation rate. The temporal convolution dilation rate is configured to increase with the increase of linear velocity, so that the temporal receptive field of the one-dimensional temporal convolution kernel expands with the increase of the wire rope movement speed.