Port crane steel wire rope defect identification method and system based on machine vision
By combining the inertial measurement unit and the improved yolov8 model, the dynamic fuzzy and accuracy problems in complex environments of wire rope defect identification of port crane wire ropes are solved, and fast and accurate defect detection is achieved, improving detection efficiency and adaptability.
Patent Information
- Application Number
- CN202510315515.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The prior art is difficult to accurately identify the defects of port crane wire ropes in dynamic blur and complex environments. Traditional visual methods have poor adaptability, deep learning models have large calculation volume and poor real-time performance, and laser scanning technology is costly and susceptible to environmental interference.
Combined with the inertial measurement unit (IMU) to obtain the wire rope motion state information in real time, predict the motion state through Kalman filtering, build a fuzzy kernel function and use the inverse filtering algorithm to clear the image, use the improved yolov8 model for defect recognition, and embed the CBAM attention mechanism and angle detection head.
It realizes rapid and accurate detection of wire rope defects in dynamic fuzzy and complex environments, improves identification accuracy and detection efficiency, and reduces the demand for computing resources.
Smart Images

Figure CN120259210A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visual defect recognition, and particularly relates to a method and system for identifying wire rope defects of port cranes based on machine vision. Background Technique
[0002] Cranes are key loading and unloading equipment for logistics transfer in ports, terminals, etc. As an important load-bearing component, wire ropes are constantly under alternating loads, harsh natural environments, and frequent friction with other components, making them prone to defects such as wear, fatigue wire breakage, corrosion, deformation, and overload. The safety state of wire ropes is directly related to the stable operation of the crane system. Once a failure occurs, it may trigger serious safety accidents, resulting in damage to goods, casualties, and production stagnation.
[0003] Currently, the common methods for defect recognition are as follows:
[0004] (1) Traditional vision method: Based on threshold segmentation (such as Otsu algorithm) and morphological processing. The Otsu algorithm automatically calculates the threshold according to the image gray histogram, compares the pixels in the image with the threshold according to their gray values, and divides them into different categories to separate the target from the background, attempting to segment the defect area of the wire rope. Morphological processing then processes the image after threshold segmentation through operations such as erosion, dilation, opening, and closing, aiming to extract image features, eliminate noise, smooth boundaries, and segment the region of interest.
[0005] However, when using this method in the detection of crane wire ropes, the high-speed swinging of the wire rope causes dynamic blurring of the image, the defect features are not clear, and the gray value distribution is complex, resulting in the difficulty of accurately segmenting the defect area by the Otsu algorithm based on a fixed threshold, and prone to false segmentation or missed segmentation; the morphological processing has poor effects on noise and blurred edges in the blurred image, and is prone to destroying the true shape and features of the defects. At the same time, environmental interferences such as light changes and dust pollution will change the gray values and textures of the image, making the traditional vision method have poor adaptability and limited processing ability for defect features of different scales.
[0006] Deep learning models: such as Faster R-CNN and U-Net. Faster R-CNN is based on the Region Proposal Network (RPN). First, it extracts image features through a convolutional neural network. The RPN network generates candidate regions and performs classification and regression. Then, through the Region of Interest Pooling operation and classifiers and regressors, it realizes the accurate classification and location regression of the target. U-Net is a semantic segmentation model. Its U-shaped structure extracts high-level semantic features through the contraction path, restores the resolution and fuses features through the expansion path, and realizes the classification of image pixels. However, when Faster R-CNN is used for dynamic detection of crane wire ropes, the RPN network generates candidate regions and subsequent classification and regression operations are computationally intensive and time-consuming, resulting in slow detection speed and difficulty in meeting real-time requirements. The U-Net network structure is complex. Multiple convolution, pooling, upsampling, and deconvolution operations require a large amount of computing resources, with poor real-time performance and insufficient ability to detect tiny defects.
[0007] Laser scanning technology: It uses a laser beam to scan the surface of the wire rope. By measuring information such as the time or phase difference of the reflected laser beam, it obtains the three-dimensional information of the wire rope surface to detect whether there are defects such as wear and deformation on the wire rope.
[0008] However, when applied to the large-scale moving scenario of cranes, the equipment cost is high, including laser scanners, data acquisition systems, data processing software, etc. The economic cost of large-scale application is unbearable. The requirements for equipment installation and calibration are high. Installing and calibrating in the complex structure of cranes is time-consuming and laborious, and the vibration and displacement of the equipment during operation will reduce the detection accuracy. It is easily affected by environmental factors such as strong light, dust, rain, and fog, resulting in an increase in measurement errors and a shortening of the effective detection distance. Summary of the Invention
[0009] In view of the deficiencies in the prior art, the present invention provides a method and system for identifying wire rope defects of port cranes based on machine vision, which can achieve fast and accurate detection of wire rope surface defects in a dynamically blurred and complex environment.
[0010] The present invention provides the following technical solutions:
[0011] In the first aspect, a method for identifying wire rope defects of port cranes based on machine vision is provided, including:
[0012] Real-time acquisition of wire rope images;
[0013] Real-time acquisition of the motion state data of the wire rope through an inertial measurement unit, including the acceleration and angular velocity of the wire rope; based on the motion state data of the wire rope acquired at the previous moment, use the Kalman filter algorithm to predict the current motion state of the wire rope.
[0014] Based on the predicted current motion state of the wire rope, construct a fuzzy kernel function, and use the inverse filtering algorithm to clarify the wire rope image collected in real time;
[0015] Use the object detection model to identify defects in the clarified image.
[0016] Optionally, the step of constructing a fuzzy kernel function based on the predicted current motion state of the wire rope and using the inverse filtering algorithm to clarify the wire rope image collected in real time specifically includes:
[0017] Based on the predicted speed at the current moment, construct a fuzzy kernel function according to the speed components of the wire rope in the set X direction and the set Y direction;
[0018]
[0019] where K(x, y) is the blur degree of the image at the point (x, y), σ is the set coefficient related to the speed, v x and v y are respectively the speed components of the wire rope at the current moment in the set X direction and the set Y direction;
[0020] After normalizing the fuzzy kernel function K(x, y), perform a Fourier transform to obtain its representation H(u, v) in the frequency domain, and convert the wire rope image collected in real time into the frequency domain form F(u, v);
[0021] In the frequency domain, based on the frequency domain form F(u, v) of the wire rope image and the frequency domain form H(u, v) of the fuzzy kernel function, use the inverse filtering algorithm or the Wiener filtering algorithm to restore the frequency domain representation F(u, v) of the wire rope image to the frequency domain representation G(u, v) of a clear image;
[0022] Convert the frequency domain representation G(u, v) of the clear image back to the spatial domain through an inverse Fourier transform to achieve the clarification of the collected wire rope image.
[0023] Optionally, in the frequency domain, using the inverse filtering algorithm to restore the frequency domain representation F(u, v) of the wire rope image to the frequency domain representation G(u, v) of a clear image is specifically:
[0024]
[0025] where ∈ is the set minimum value.
[0026] Optionally, in the frequency domain, using the Wiener filtering algorithm to restore the frequency domain representation F(u, v) of the wire rope image to the frequency domain representation G(u, v) of a clear image is specifically:
[0027]
[0028] Among them, H * (u, v) is the conjugate complex number of H(u, v), and S n (u, v) is the noise power spectrum, and S f (u, v) is the power spectrum of the collected wire rope image.
[0029] Optionally, the improved yolov8 model is used to identify defects in the sharpened image, and the object detection model is the improved yolov8 model;
[0030] The improved yolov8 model uses MobileNetV3 as the backbone network, and the CBAM attention mechanism is embedded in the backbone network; the detection head of the improved yolov8 model includes an angle detection head for predicting the rotation angle of the wire rope.
[0031] Optionally, using the trained improved yolov8 model for defect identification specifically includes:
[0032] Feature extraction is performed on the sharpened image, including the contour, shape, texture features, and local direction information of the spiral texture of the wire rope;
[0033] The extracted features are input into the improved yolov8 model to predict the target bounding box and the rotation angle of the wire rope;
[0034] Based on the predicted target bounding box and rotation angle, the improved yolov8 model rotates the target bounding box in the post-processing stage of the model to obtain a rotated box;
[0035] After non-maximum suppression adjustment of the rotated box, the improved yolov8 model outputs the defect identification result of the wire rope and realizes defect localization according to the position of the rotated box.
[0036] Optionally, in the step of rotating the target bounding box by the improved yolov8 model in the post-processing stage of the model to obtain a rotated box based on the predicted target bounding box and rotation angle, the vertex coordinates (x′, y′) of the rotated box are:
[0037]
[0038] Among them, (x, y) are the vertex coordinates of the predicted target bounding box, x c and y c are the center coordinates of the predicted target bounding box respectively, and θ is the predicted rotation angle of the wire rope.
[0039] Optionally, the loss function L of the improved yolov8 model during training is:
[0040] L = L ce + L MSE + L rotation
[0041] Among them, L ce is the cross-entropy loss for defect classification, and L MSE is the mean squared error loss for position regression, and L rotation is the mean squared error loss for angle prediction;
[0042]
[0043] Among them, N is the number of samples, C is the number of classes, and y i,c is the true label (0 or 1) of the i-th sample on class c, and p i,c is the probability of the i-th sample predicted by the model on class c; y i is the true value of the i-th sample, is the value of the i-th sample predicted by the model, is the rotation angle of the i-th sample predicted by the model, is the true rotation angle of the i-th sample.
[0044] Optionally, when training the improved YOLOv8 model, an adversarial training strategy is introduced. Specifically, during the training process, image generation technology is used to add simulated noises such as haze and raindrops to the original wire rope image, and the image with noises and the original image are used as training data and input into the improved YOLOv8 model together.
[0045] In a second aspect, a port crane wire rope defect recognition system based on machine vision is provided, including:
[0046] An acquisition module for real-time acquisition of wire rope images;
[0047] A state estimation module for real-time acquisition of the motion state data of the wire rope through an inertial measurement unit, including the acceleration and angular velocity of the wire rope; based on the motion state data of the wire rope acquired at the previous moment, the Kalman filtering algorithm is used to predict the current motion state of the wire rope;
[0048] A clarification processing module for constructing a blur kernel function based on the predicted current motion state of the wire rope and using the inverse filtering algorithm to clarify the currently real-time acquired wire rope image;
[0049] A defect recognition module for using an object detection model to recognize defects in the clarified image.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] (1) The present invention innovatively integrates IMU and visual data, uses the IMU to obtain the motion state information of the wire rope in real time, combines the visual image data, and when the wire rope swings at high speed, dynamically compensates the blurred image through the motion trajectory prediction information fed back by the IMU, improving the accuracy and reliability of defect recognition.
[0052] (2) The present invention conducts a lightweight design on the YOLOv8 model, replaces the backbone network with MobileNetV3 to reduce the number of parameters and improve the inference speed; embeds the CBAM attention mechanism to enhance the feature response ability to different types of defect regions; additionally, an angle prediction head is used to achieve rotated bounding box detection, adapting to the spiral texture characteristics of the wire rope and improving the detection performance and efficiency. Description of the Drawings
[0053] Figure 1 is the flowchart of the steps of the method for identifying wire rope defects of a port crane based on machine vision according to the present invention;
[0054] Figure 2 is the flowchart of dynamically compensating for motion blur of pictures according to the present invention;
[0055] Figure 3 is the structural block diagram of the improved yolov8 model according to the present invention. Detailed Embodiments
[0056] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0057] It should be noted that the term "including" and any of its variations in the description and claims of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0058] Embodiment 1
[0059] As Figure 1 and Figure 2 shown, a method for identifying wire rope defects of a port crane based on machine vision is provided, including the following steps:
[0060] S1: Collect wire rope images in real time.
[0061] The wire rope images are captured by a high-speed industrial camera with a shooting frame rate of up to 2000fps, which can quickly capture images when the wire rope is swinging at high speed, reducing dynamic blur. For example, when the wire rope swings at a speed of 5m / s, 2 images can be captured within 1ms, which has obvious advantages compared with ordinary 30fps cameras.
[0062] S2: The motion state data of the wire rope, including the acceleration and angular velocity of the wire rope, are collected in real time through an inertial measurement unit; based on the motion state data of the wire rope collected at the previous moment, the Kalman filtering algorithm is used to predict the current motion state of the wire rope.
[0063] The inertial measurement unit IMU consists of an accelerometer and a gyroscope, and measures the acceleration and angular velocity of the wire rope in three axial directions in real time; Kalman filtering is an algorithm that uses a linear system state equation to optimally estimate the system state through system input and output observation data. Its state update equation is:
[0064] Among them, is the state prediction value at time step k based on the state prediction value at time k-1; A is the state transition matrix; is the optimal estimate value at time k-1; B is the control matrix; u k is the control input.
[0065] Calculate the position and speed of the wire rope using the processed data, and predict the motion trajectory.
[0066] S3: Based on the predicted current motion state of the wire rope, construct a fuzzy kernel function, and use the inverse filtering algorithm to clarify the currently captured real-time wire rope image.
[0067] Step S3 specifically includes:
[0068] S31: Based on the predicted speed at the current moment, construct a fuzzy kernel function according to the speed components of the wire rope in the set X direction and the set Y direction;
[0069]
[0070] Among them, K(x,y) is the blur degree of the image at point (x,y), σ is a coefficient related to the speed set, v x and v y are the speed components of the wire rope at the current moment in the set X direction and the set Y direction respectively.
[0071] The larger σ is, the larger the range of the fuzzy kernel function and the more blurred the image.
[0072] S32: Perform Fourier transform on the blurred kernel function K(x, y) after normalization to obtain its representation H(u, v) in the frequency domain, and convert the real-time acquired wire rope image into the frequency domain form F(u, v).
[0073] The method of converting the real-time acquired wire rope image into the frequency domain form F(u, v) can refer to the prior art. The blurred kernel function K(x, y) is normalized to satisfy ∑ x ∑ y K(x, y) = 1.
[0074] S33: In the frequency domain, based on the frequency domain form F(u, v) of the wire rope image and the frequency domain form H(u, v) of the blurred kernel function, use the inverse filtering algorithm or Wiener filtering algorithm to restore the frequency domain representation F(u, v) of the wire rope image to the frequency domain representation G(u, v) of a clear image.
[0075] Specifically, in the frequency domain, using the inverse filtering algorithm to restore the frequency domain representation F(u, v) of the wire rope image to the frequency domain representation G(u, v) of a clear image is specifically as follows:
[0076]
[0077] where ∈ is a set minimum value. The setting of the minimum value can avoid division-by-zero errors, and the value of the minimum value is optionally 1e - 8.
[0078] In the frequency domain, using the Wiener filtering algorithm to restore the frequency domain representation F(u, v) of the wire rope image to the frequency domain representation G(u, v) of a clear image is specifically as follows:
[0079]
[0080] where H * (u, v) is the conjugate complex number of H(u, v), S n (u, v) is the noise power spectrum, and S f (u, v) is the power spectrum of the acquired wire rope image.
[0081] Using the Wiener filtering algorithm to restore the frequency domain representation F(u, v) of the wire rope image to the frequency domain representation G(u, v) of a clear image can reduce the influence of noise.
[0082] The acquisition of the noise power spectrum and the acquisition of the power spectrum of the original image can refer to the prior art. In this application, for the acquisition of the noise power spectrum: generally, multiple image samples containing noise need to be collected, preprocessed (such as mean removal, normalization, etc.), and then Fourier transformed to convert the image from the spatial domain to the frequency domain. In the frequency domain, the power spectrum of the image is calculated, and the power spectra of multiple samples are averaged to reduce the influence of random errors, thereby obtaining the noise power spectrum. For the acquisition of the power spectrum of the original image: for the original image (i.e., the image not affected by noise pollution or blur), Fourier transform is directly performed to convert the image from the spatial domain to the frequency domain. In the frequency domain, the power spectrum of the image is calculated to obtain the power spectrum of the original image.
[0083] S34: The representation G(u, v) of the clear image in the frequency domain is converted back to the spatial domain through inverse Fourier transform to realize the clarity processing of the collected wire rope image.
[0084] The inverse Fourier transform IFT can refer to the prior art.
[0085] S4: Use the object detection model to identify defects in the image after clarity processing.
[0086] Specifically, as Figure 3 shown, use the object detection model to identify defects in the image after clarity processing. The object detection model is an improved yolov8 model; the improved yolov8 model uses MobileNetV3 as the backbone network and embeds the CBAM attention mechanism in the backbone network; the detection head of the improved yolov8 model includes an angle detection head for predicting the rotation angle of the wire rope.
[0087] The input layer of the improved yolov8 model can receive the preprocessed wire rope image (size 640×640×3). The preprocessing includes dynamic blur compensation, image enhancement after multispectral fusion, etc.
[0088] The backbone network of the improved YOLOv8 model is MobileNetV3, and its structure is composed of multiple Inverted Residual Blocks. Each block contains the following components: 1×1 expansion convolution: expand the number of channels (the expansion ratio is controlled by the expand ratio), depthwise separable convolution: 3×3 convolution to extract spatial features and reduce the computational load. SE (Squeeze-and-Excitation) module: channel attention mechanism to enhance the features of important channels. CBAM module: embedded after each inverted residual block, and channel attention and spatial attention calculations are performed in sequence. Replacing the backbone network of YOLOv8 with MobileNetV3 reduces the number of parameters by 42% and triples the inference speed. CBAM performs attention calculations on the feature map in both the channel and spatial dimensions. The ChannelAttention class obtains the global information of the channel dimension through global average pooling and global max pooling, and generates channel attention weights through two layers of convolution; the SpatialAttention class takes the average and maximum values of the feature map in the channel dimension, and generates spatial attention weights through a 7x7 convolution after concatenation. The CBAM class multiplies the feature map by the channel and spatial attention weights in sequence, enhances the feature response ability to different types of defect regions, and processes feature maps of different scales at the same time, improving the detection ability for tiny and large defects.
[0089] CBAM performs attention calculations on the feature map in both the channel and spatial dimensions. The ChannelAttention class obtains the global information of the channel dimension through global average pooling and global max pooling, and generates channel attention weights through two layers of convolution. Suppose the input feature map is F ∈ R C×H×W , the result after global average pooling The calculation formula is:
[0090]
[0091] The result after global max pooling The calculation formula is:
[0092] Then, they pass through two layers of convolution, Conv1 and Conv2 respectively, to obtain M c (F):
[0093]
[0094] where σ is the Sigmoid function.
[0095] The SpatialAttention class takes the average and maximum values of the feature map in the channel dimension, and generates spatial attention weights through a 7×7 convolution after concatenation. The result after taking the average in the channel dimension The calculation formula is: The result after taking the maximum value in the channel dimension The calculation formula is: After splicing and and passing through a 7×7 convolution to obtain M s (F): The CBAM class multiplies the feature map by the channel and spatial attention weights successively to enhance the feature response ability to different types of defect regions, and at the same time processes feature maps of different scales to improve the detection ability for tiny and large defects. The final output feature map F′ is:
[0096] F′ = F × M c (F) × M s (F)
[0097] The neck network (Neck) of the improved yolov8 model: The improved PANet can be used; it enhances the detection ability for tiny and large defects in multi-scale feature fusion. Structure: Feature Pyramid Network (FPN): Transmits high-level semantic features from top to bottom. Path Aggregation Network (PAN): Transmits low-level detailed features from bottom to top. CBAM embedding: Adds a CBAM module after the feature fusion layer to further optimize the feature response.
[0098] The detection head (Head) of the improved yolov8 model: Adds an angle detection head, which is the original branch of the yolov8 model. Bounding box prediction: Outputs the center coordinates and class probabilities. New branch: Angle prediction head, network layer: 1×1 I convolution + fully connected layer, outputs the rotation angle θ.
[0099] To ensure that the network can accurately learn the features related to the target rotation angle, a 1×1 convolutional layer can be used to adjust the number of channels to better integrate the feature information, and then the final rotation angle prediction value is output through a multi-layer perceptron (MLP) structure.
[0100] The loss function L during the training of the improved yolov8 model is:
[0101] L = L ce + L MSE + L rotation
[0102] Among them, L ce is the cross-entropy loss for defect classification, L MSE is the mean squared error loss for position regression, L rotation is the mean squared error loss for angle prediction;
[0103]
[0104] Where N is the number of samples, C is the number of classes, and y i,c is the true label (0 or 1) of the i-th sample for class c, and p i,c is the probability of the i-th sample predicted by the model for class c; y i is the true value of the i-th sample, is the value of the i-th sample predicted by the model, is the rotation angle of the i-th sample predicted by the model, is the true rotation angle of the i-th sample.
[0105] When training the improved YOLOv8 model, an adversarial training strategy is introduced. Specifically, during the training process, image generation techniques (such as the generator in GAN) are used to add simulated noises such as haze and raindrops to the original wire rope image, and the image with noise and the original image are used as training data and input into the improved YOLOv8 model together.
[0106] Adversarial training can be regarded as a game process between the generator G and the discriminator D. The goal of the generator is to generate realistic noise images to mislead the discriminator, and its loss function L G is:
[0107]
[0108] The goal of the discriminator is to accurately distinguish between real images and generated noise images, and its loss function L D is:
[0109]
[0110] Where p data (x) is the distribution of real images, and p z (z) is the distribution of noises.
[0111] In step S4, the trained improved YOLOv8 model is used for defect recognition, which specifically includes:
[0112] S41: Extract features from the clarified image, including the contour, shape, texture features of the wire rope, and the local direction information of the spiral texture.
[0113] The contour and shape of the wire rope are extracted through an edge detection algorithm, the texture features are extracted using local binary patterns, and the local direction information of the spiral texture is obtained by calculating the gradient direction of pixel points in the image.
[0114] S42: Input the extracted features into the improved YOLOv8 model to predict the target bounding box and the rotation angle of the wire rope.
[0115] S43: Based on the predicted target bounding box and rotation angle, the improved YOLOv8 model rotates the target bounding box during the post-processing stage of the model to obtain a rotated bounding box.
[0116] Specifically, perform a rotation operation on the detection bounding box according to the predicted rotation angle. This operation can be achieved using a 2D rotation matrix. The vertex coordinates (x′, y′) of the rotated bounding box are:
[0117]
[0118] where (x, y) are the vertex coordinates of the predicted target bounding box, x c and y c are the center coordinates of the predicted target bounding box respectively, and θ is the predicted rotation angle of the wire rope.
[0119] Rotate the four vertices of the detection bounding box according to the above formula to obtain a rotated detection bounding box that conforms to the actual shape. In practical applications, whether it is for on-line monitoring of wire ropes or off-line data analysis, the results based on the detection of rotated bounding boxes can more accurately locate and describe spiral texture defects, providing more reliable data support for subsequent maintenance decisions.
[0120] S44: After performing non-maximum suppression adjustment on the rotated bounding box, the improved YOLOv8 model outputs the defect recognition result of the wire rope and realizes defect localization according to the position of the rotated bounding box.
[0121] The method of non-maximum suppression NMS can refer to the existing technology. NMS is a commonly used post-processing technology in object detection, which can help the model select the most accurate one among multiple detection results.
[0122] Embodiment 2
[0123] Provide a wire rope defect recognition system based on machine vision, including:
[0124] An acquisition module for real-time acquisition of wire rope images;
[0125] A state estimation module for real-time acquisition of the motion state data of the wire rope through an inertial measurement unit, including the acceleration and angular velocity of the wire rope; based on the motion state data of the wire rope acquired at the previous moment, use the Kalman filter algorithm to predict the current motion state of the wire rope;
[0126] A clarification processing module for constructing a blur kernel function based on the predicted current motion state of the wire rope and using the inverse filtering algorithm to perform clarification processing on the currently real-time acquired wire rope image;
[0127] A defect recognition module for using an object detection model to perform defect recognition on the clarified image.
[0128] For a more specific process of the above method, reference may be made to the corresponding content disclosed in the foregoing embodiments, which will not be elaborated herein.
[0129] The embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference may be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference may be made to the description in the method part.
[0130] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0131] The above are only the preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A method for identifying defects in the steel wire ropes of port cranes based on machine vision, characterized in that, Including: Real-time acquisition of wire rope images; Real-time acquisition of the motion state data of the wire rope through an inertial measurement unit, including the acceleration and angular velocity of the wire rope; Based on the motion state data of the wire rope acquired at the previous moment, use the Kalman filter algorithm to predict the current motion state of the wire rope; Based on the predicted current motion state of the wire rope, construct a fuzzy kernel function, and use the inverse filtering algorithm to clarify the currently acquired wire rope image; Use the object detection model to identify defects in the clarified image.
2. The method for identifying defects of steel wire ropes of port cranes based on machine vision according to claim 1, wherein The step of constructing a fuzzy kernel function based on the predicted current motion state of the wire rope and using the inverse filtering algorithm to clarify the currently acquired wire rope image specifically includes: Based on the predicted velocity at the current moment, construct a fuzzy kernel function according to the velocity components of the wire rope in the set X direction and the set Y direction; Among them, K(x, y) is the degree of blurring of the image at the point (x, y), σ is a set coefficient related to the speed, and v x and v y are respectively the velocity components of the current velocity of the wire rope in the set X direction and the set Y direction; Perform Fourier transform on the normalized fuzzy kernel function K(x,y) to obtain its representation form H(u,v) in the frequency domain, and convert the currently acquired wire rope image into the frequency domain form F(u,v); In the frequency domain, based on the frequency domain form F(u,v) of the wire rope image and the frequency domain form H(u,v) of the fuzzy kernel function, use the inverse filtering algorithm or Wiener filtering algorithm to restore the frequency domain representation F(u,v) of the wire rope image to the frequency domain representation G(u,v) of a clear image; Convert the frequency domain representation G(u,v) of the clear image back to the spatial domain through inverse Fourier transform to achieve the clarification of the acquired wire rope image.
3. The method for identifying defects in the steel wire rope of a port crane based on machine vision according to claim 2, characterized in that, In the frequency domain, through the inverse filtering algorithm, restore the frequency domain representation F(u,v) of the wire rope image to the frequency domain representation G(u,v) of a clear image, specifically: where ∈ is a set minimum value.
4. The method for identifying defects of steel wire ropes of port cranes based on machine vision according to claim 2, characterized in that, In the frequency domain, through the Wiener filtering algorithm, restore the frequency domain representation F(u,v) of the wire rope image to the frequency domain representation G(u,v) of a clear image, specifically: Among them, H * (u, v) is the conjugate complex number of H(u, v), and S n (u, v) is the noise power spectrum, and S f (u, v) is the power spectrum of the collected wire rope image.
5. The method for identifying defects in the steel wire rope of a port crane based on machine vision according to claim 1, characterized in that, The object detection model used to identify defects in the clarified image is an improved yolov8 model; The improved yolov8 model uses MobileNetV3 as the backbone network, and embeds the CBAM attention mechanism in the backbone network; the detection head of the improved yolov8 model includes an angle detection head for predicting the rotation angle of the wire rope.
6. The method for identifying defects of steel ropes of port cranes based on machine vision according to claim 5, characterized in that, Using the trained improved yolov8 model for defect identification specifically includes: Extract features from the clarified image, including the contour, shape, texture features, and local direction information of the spiral texture of the wire rope; Input the extracted features into the improved yolov8 model to predict the target bounding box and the rotation angle of the wire rope; Based on the predicted target bounding box and rotation angle, the improved yolov8 model rotates the target bounding box in the post-processing stage of the model to obtain a rotated box; After non-maximum suppression adjustment of the rotated box, the improved yolov8 model outputs the defect identification result of the wire rope and realizes defect localization according to the position of the rotated box.
7. The method for identifying defects of steel wire ropes of port cranes based on machine vision according to claim 6, wherein Based on the predicted target bounding box and rotation angle, the improved YOLOv8 model rotates the target bounding box during the post-processing stage of the model to obtain the rotated box. The vertex coordinates (x ′ , y ′ ) of the rotated box are as follows: Among them, (x, y) are the vertex coordinates of the predicted target bounding box, and x c and y c are respectively the central coordinates of the predicted target bounding box, and θ is the predicted rotation angle of the wire rope.
8. The method for identifying defects in the steel wire rope of a port crane based on machine vision according to claim 5, wherein, The loss function L of the improved yolov8 model during training is: L = L ce + L MSE + L rotation Among them, L ce is the cross-entropy loss for defect classification, and L MSE is the mean squared error loss for position regression, and L rotation is the mean squared error loss for angle prediction; where N is the number of samples, C is the number of classes, and y i,c is the true label (0 or 1) of the i-th sample in class c, and p i,c is the probability of the i-th sample predicted by the model in class c; y i is the true value of the i-th sample, is the value of the i-th sample predicted by the model, is the rotation angle of the i-th sample predicted by the model, is the true rotation angle of the i-th sample.
9. The method for identifying defects in the steel wire rope of a port crane based on machine vision according to claim 5, wherein When the improved YOLOv8 model is trained, an adversarial training strategy is introduced. Specifically, during the training process, image generation technology is used to add simulated noises such as haze and raindrops to the original wire rope image, and the image with noise and the original image are used as training data and input into the improved YOLOv8 model together.
10. A wire rope defect identification system for port cranes based on machine vision, characterized in that, It includes: An acquisition module for real-time acquisition of wire rope images; A state estimation module for real-time acquisition of the motion state data of the wire rope through an inertial measurement unit, including the acceleration and angular velocity of the wire rope; Based on the motion state data of the wire rope acquired at the previous moment, the Kalman filter algorithm is used to predict the current motion state of the wire rope; A clarification processing module for constructing a blur kernel function based on the predicted current motion state of the wire rope and using the inverse filtering algorithm to clarify the currently real-time acquired wire rope image; A defect identification module for using an object detection model to identify defects in the clarified image.
Citation Information
Patent Citations
Power transmission line defect detection method based on improved YOLOv5 and blurred image enhancement
CN116416237A
Mining elevator steel wire rope damage detection method and system based on machine vision
CN118521557A