A Defect Detection Method for Traction Elevator Steel Wires Based on Deep Learning
Through the improved UCG-GAN network and EfficientNet network model, the problems of insufficient training samples and low detection accuracy in elevator wire rope defect detection are solved, efficient and automated defect detection is achieved, and the accuracy and reliability of detection are improved.
Patent Information
- Application Number
- CN202411161551.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-08-23
AI Technical Summary
The existing elevator wire rope defect detection methods have problems such as insufficient training samples, low detection accuracy and relying on a large number of equipment, making it difficult to meet the detection needs of high frequency and high accuracy.
The improved UCG-GAN network is used to generate high-quality wire rope defect images, combined with the improved EfficientNet network model for identification and detection, and the effect of image acquisition and feature extraction is improved through multi-directional acquisition and CA attention mechanism modules.
It improves the accuracy of wire rope defect image recognition and the efficiency of automated detection, reduces dependence on the equipment, and enhances the stability and reliability of the detection results.
Smart Images

Figure CN119168948B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of elevator wire rope fault detection, and in particular to a method for detecting defects of traction elevator wire ropes based on deep learning. Background Art
[0002] As one of the core components of modern urban transportation systems, the safety of traction elevators is directly related to the lives and property of the general public. The wire rope, as an important load-bearing component for elevator operation, its quality and condition are crucial for the safety and operation stability of the elevator. With the continuous increase in urban population density and the number of high-rise buildings, the usage frequency of elevators is also increasing day by day, posing higher requirements for the safety and reliability of elevator systems. However, under long-term operation and high-intensity use, wire ropes are prone to problems such as wear, corrosion, and fracture, which may lead to elevator accidents, causing serious casualties and property losses. Therefore, regular and accurate defect detection of wire ropes has become one of the urgent technical problems to be solved in the elevator industry.
[0003] According to different detection principles, the defect detection of elevator wire ropes can be roughly divided into visual inspection method and electromagnetic detection method. The visual inspection method requires inspectors to work in the dim environment of the elevator shaft. The work efficiency is low, easy to fatigue, the vision is poor, highly dependent on the experience of personnel, and the detection time is also long. The electromagnetic detection method is to ensure that the elevator is shut down and safety protection is done, and electromagnetic detection equipment, sensors and related tools are prepared; then, the electromagnetic sensor is installed on the wire rope to ensure good contact between the sensor and the wire rope; then, the electromagnetic detection equipment is started and the sensor is slowly moved along the wire rope to perform a full-length scan of the wire rope; during the scan, the equipment identifies defects inside the wire rope, such as broken wires, wear and corrosion, etc., by detecting changes in electromagnetic signals. This method requires a large amount of equipment support, the detection results are easily affected by the environment, and the accuracy is limited.
[0004] Since the elevator wire rope is exposed outside, it can be easily photographed by a camera, and the visual detection method can be used for identification. At the same time, if visual detection is carried out during the normal operation of the elevator, it can replace manual identification, greatly reducing the inspection time and labor cost. Due to the low frequency of occurrence of wire rope defects in elevators, the number of obtained wire rope defect pictures is small, resulting in the problem of insufficient diversity of training samples. Therefore, based on artificial intelligence and big data processing methods, and based on machine vision, this application combines computer science and image processing technologies to propose a method for detecting defects of traction elevator wire ropes based on deep learning, which is used to generate wire rope defect images for training, improve the image recognition accuracy of wire rope defects, and improve the efficiency of automatic detection, which is very necessary. Summary of the Invention
[0005] In view of this, the present invention proposes a method for detecting defects in traction elevator steel ropes based on deep learning, which combines an improved UCG-GAN network to generate steel rope defect images, improves the deficiency of existing defect training samples, and further identifies and detects steel rope defects through an improved EfficientNet network model.
[0006] The present invention provides a method for detecting defects in traction elevator steel ropes based on deep learning, comprising the following steps:
[0007] S1: Use a plurality of image acquisition devices to perform multi-directional acquisition to obtain real defect images of the elevator steel rope;
[0008] S2: Construct an improved UCG-GAN model, where the improved UCG-GAN model includes an independent generator G and a discriminator D;
[0009] S3: Use the real defect images and the first steel rope defect images generated by the generator G of the improved UCG-GAN model; update the parameters of the generator G and the discriminator D by backpropagation using the generated first steel rope defect images, and iteratively train the generator G and the discriminator D; use the trained generator G and discriminator D to generate second steel rope defect images, and establish a steel rope defect sample library based on the second steel rope defect images;
[0010] S4: Establish and initialize an improved EfficientNet network model, train it using the steel rope defect sample library, and update the parameters of the improved EfficientNet network model to obtain a trained improved EfficientNet network model;
[0011] S5: Use the trained improved EfficientNet network model to identify defects in the real defect images of the elevator steel rope collected in real time by a plurality of image acquisition devices, and output the defect identification results.
[0012] Based on the above technical solutions, preferably, the content of step S1 is: Configure more than two image acquisition devices, set more than two image acquisition devices on the end face of the elevator machine room floor close to the car side, and distribute them evenly in the circumferential direction; more than two image acquisition devices respectively aim at the surfaces of different positions of the same steel rope to take pictures, and obtain real defect images of the elevator steel rope.
[0013] Preferably, the generator G of the improved UCG-GAN model includes five sequentially arranged encoder blocks, an intermediate block, and five decoder blocks. Let the numbers of the five encoder blocks be 4, 3, 2, 1, 0 in sequence; the numbers of the five decoders be 1, 2, 3, 4, 5 in sequence; four skip connections are sequentially arranged between the encoder blocks and the decoder blocks. Let r ∈ {1, 2, 3, 4}. One end of each skip connection is connected to the output end of the encoder block numbered r, and the other end of each skip connection is connected to the output end of the decoder block numbered r; a CA attention mechanism module is inserted into each skip connection. The CA attention mechanism module takes the encoder numbered r and the decoder numbered r as inputs, and concatenates the output of the CA attention mechanism module with the output of the decoder numbered r as the input of the decoder numbered r + 1; the input of the encoder block at the head end is the real defect image x k , and the last decoder block outputs the first wire rope defect image
[0014] Further preferably, each encoder block and decoder block includes three convolutional layers, using leaky ReLU and batch normalization. A downsampling layer is provided at the end of each encoder block, and the output end of the downsampling layer serves as the output end of the encoder port; an upsampling layer is provided at the end of each decoder block, and the output end of the upsampling layer can serve as the output end of the decoder block.
[0015] Further preferably, the discriminator D of the improved UCG-GAN model includes a convolutional layer, five residual blocks, and a fully connected layer arranged in sequence; after the processing of the convolutional layer and each residual block, the fully connected layer outputs the discriminator score.
[0016] Even more preferably, in step S3, the parameters of the generator G and the discriminator D are updated by backpropagation using the generated first wire rope defect image, and the generator G and the discriminator D are iteratively trained. Let the average value of the discriminator output scores of k real defect images x k be D(x), and the average value of the discriminator output scores of k first wire rope defect images be Using the hinge adversarial loss, the iterative training objective of the discriminator D is to minimize the discriminator loss function λ D , and the iterative training objective of the generator G is to minimize the generator loss function λ GD : where is the expected value of the real defect image x k , is the expected value of the first wire rope defect image , max(·) is to use the ReLU activation function to ensure that the result of the loss function is non-negative; by defining the loss function λD and λ GD , during the training process, the discriminator D and the generator G continuously confront each other, and finally reach a Nash equilibrium, enabling the generator to generate the second wire rope defect image that is closest to the real defect image.
[0017] Preferably, the content of step S4 is: the constructed improved EfficientNet network model includes a first convolutional module, seven sequentially arranged graph convolutional layer modules ATConv, a second convolutional module, average global pooling, and a fully connected layer; each graph convolutional layer module ATConv includes a first ordinary convolutional layer, an asymmetric convolutional module, a Triplet attention mechanism, a second ordinary convolutional layer, and a Dropout layer arranged in sequence, and a skip connection is also arranged between the input end of the first ordinary convolutional layer and the output end of the Dropout layer; the loss function of the graph convolutional layer module ATConv is the FILOSS loss function obtained by fusing the FocalLoss loss function and the IOULOSS loss function; the improved EfficientNet network model takes the second wire rope defect image in the sample library as input, and identifies and marks the wire rope defect part for output.
[0018] Further preferably, the asymmetric convolutional module adds a horizontal asymmetric convolutional kernel and a vertical asymmetric convolutional kernel on the basis of the standard square convolution, and the input image is respectively convolved by three different convolutional kernels to extract different branch features: where M U,V,K ∈R U×V×C represents the feature map of the K-th channel with an input size of U×V, K = 1, 2,..., C, U≠V; represents the j-th convolutional kernel of the K-th channel with an input size of H×W, H≠W; O U,V,j is the output feature map corresponding to the j-th convolutional kernel, and * represents the dot product operation;
[0019] Using the additivity of convolution to fuse different branches, the fused feature output is obtained, and the dimension of the fused feature is the same as that of the input feature, satisfying: I is the input feature map matrix; K1 and K2 respectively represent different convolutional kernels; * represents the convolution operation; represents the addition of the corresponding positions of the convolutional kernels.
[0020] More preferably, the Triplet attention mechanism includes the following: the first branch is the channel attention calculation branch. The input feature passes through Z-Pool, then through a 7×7 convolution, and through the Sigmoid activation function to generate the spatial attention weight; the second branch is the interaction capture branch of channel C and spatial W dimensions. The input feature is first transformed into an H×C×W dimensional feature through the permutation function, then Z-Pool is performed on the H dimension, then processed through a 7×7 convolution and the Sigmoid activation function, and then transformed into a C×H×W dimensional feature through the permutation function; the third branch is the interaction capture branch of channel C and spatial H dimensions. The input feature is first transformed into a W×H×C dimensional feature through the permutation function, then Z-Pool is performed on the W dimension, then processed through a 7×7 convolution and the Sigmoid activation function, and then transformed into a C×H×W dimensional feature through the permutation function; the output result of the Triplet attention mechanism is obtained by adding and averaging the output features of the three branches.
[0021] More preferably, the expression of the FILOSS loss function is: FILOSS = α×FocalFoss+(1-α)×IOULOSS, where α is the weight; the FocalLoss loss function is defined as: FocalLoss = -α t (1 - p t ) γ log(p t ), where p t is the probability value predicted by the model; α t is the adjustment factor; γ is the hyperparameter; the IOULOSS loss function is defined as follows: where A is the intersection of the predicted box and the ground truth box, and B is the union of the predicted box and the ground truth box.
[0022] A traction elevator steel wire rope defect detection method based on deep learning provided by the present invention has the following beneficial effects compared with the prior art:
[0023] (1) To improve the problem that insufficient training samples affect the training of the detection model, the present invention proposes an improved UCG-GAN network and uses it for the generation of steel wire rope defect samples. Specifically, four skip connections are introduced in the generator, and the CA attention mechanism module is inserted into the skip connections. In the discriminator, to stabilize the training process, the hinge adversarial loss is introduced and incorporated into the classifier, which improves the quality and authenticity of the generated steel wire rope defect images and provides a training set basis for the subsequent model training;
[0024] (2) During the detection of wire rope defects, the present invention proposes an improved EfficientNet network model. An asymmetric convolution module, a Triplet attention mechanism, and a FILOSS loss function are introduced into each graph convolution layer module ATConv of the EfficientNet network model, which can simultaneously consider the losses in both positioning and classification, improve the performance of the target detection model for the optimization training of different samples, and enable the network model to take into account defects of both larger and smaller sizes. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0026] Figure 1 It is a flowchart of a method for detecting wire rope defects of a traction elevator based on deep learning according to the present invention;
[0027] Figure 2 It is a schematic layout diagram of several image acquisition devices for a method for detecting wire rope defects of a traction elevator based on deep learning according to the present invention;
[0028] Figure 3 It is the structure of a generative adversarial network GAN for a method for detecting wire rope defects of a traction elevator based on deep learning according to the present invention;
[0029] Figure 4 It is the structure of the generator G of an improved UCG-GAN model for a method for detecting wire rope defects of a traction elevator based on deep learning according to the present invention;
[0030] Figure 5 It is the structure of a CA attention mechanism module for a method for detecting wire rope defects of a traction elevator based on deep learning according to the present invention;
[0031] Figure 6 It is the structure of the discriminator D of a CG-GAN model for a method for detecting wire rope defects of a traction elevator based on deep learning according to the present invention;
[0032] Figure 7 It is a structural diagram of an improved EfficientNet network model for a method for detecting wire rope defects of a traction elevator based on deep learning according to the present invention;
[0033] Figure 8Schematic diagram of the graph convolutional layer module ATConv of a traction elevator wire rope defect detection method based on deep learning according to the present invention;
[0034] Figure 9 Schematic diagram of the asymmetric convolution processing process of a traction elevator wire rope defect detection method based on deep learning according to the present invention;
[0035] Figure 10 Structure diagram of the asymmetric convolution module of a traction elevator wire rope defect detection method based on deep learning according to the present invention;
[0036] Figure 11 Schematic diagram of the Triplet attention mechanism of a traction elevator wire rope defect detection method based on deep learning according to the present invention;
[0037] Figure 12 Schematic diagram of the actual detection result of a traction elevator wire rope defect detection method based on deep learning according to the present invention.
[0038] Reference numerals: 1, elevator machine room floor; 2, image acquisition device; 3, wire rope; 4, car. Detailed implementation manners
[0039] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0040] In the defect detection and identification of traction elevator wire ropes, a large number of defect pictures are required for model training. However, due to the relatively few defect faults of wire ropes, the number of defect images is small and it is difficult to collect, resulting in insufficient sample diversity of training data. In view of this, as Figure 1 shown, the present invention provides a traction elevator wire rope defect detection method based on deep learning, including the following steps:
[0041] S1: Use a number of image acquisition devices to perform multi-directional acquisition to obtain real defect images of the elevator wire rope.
[0042] Refer to Figure 2, the content of step S1 is as follows: Configure more than two image acquisition devices 2, and arrange the more than two image acquisition devices on the end face of the elevator machine room floor 1 close to the car 4 side, and they are evenly distributed circumferentially; the more than two image acquisition devices 2 respectively aim at the surfaces of different positions of the same steel wire rope 3 to take pictures, and obtain the real defect images of the elevator steel wire rope 3. As a preferred implementation method, usually the image acquisition devices 2 are arranged at equal angles in the circumferential direction of the steel wire rope 3, generally 2-4 are used, and this embodiment recommends using 3. The side of the elevator machine room floor 1 close to the car 4 can take pictures of the surface image of the steel wire rope 3 that has contacted the pulley to the greatest extent.
[0043] The image acquisition device 2 uses the IOI Victorem 4K SDI series high-speed camera, which has a 4K resolution and can clearly capture various defects of the steel wire rope. It uses a Sony Pregius CMOS chip, which is smaller in volume, higher in sensitivity and better in cost performance than the previous Flare series. According to the elevator running speed of 2m / s, the sampling frequency of the finally selected camera is set to 100fps.
[0044] S2: Construct an improved UCG-GAN model, and the improved UCG-GAN model includes an independent generator G and discriminator D.
[0045] The present invention proposes an improved UGG-GAN to generate steel wire rope defect images. The generative adversarial network GAN can generate realistic defect images through the adversarial training of the generator (Generator) and discriminator (Discriminator), thereby enriching the training data set. As Figure 3 shown, it shows the calculation process and structure of GAN, which is composed of a generator G and a discriminator D. The generator G of the GAN network generates steel wire rope defect images from random noise Z, and the discriminator D is used to distinguish the generated images, that is, the first steel wire rope defect images and real defect images. Through this adversarial mechanism, the generator G continuously improves its output and generates samples highly similar to the real steel wire rope defect images X.
[0046] The generator G adopts a combined structure of U-Net and ResNet, searches for relevant information through the CA attention mechanism module, retains more details, and improves the clarity and authenticity of the generated images. The discriminator D enhances the stability and effect of the adversarial training through the hinge adversarial loss and the classifier based on GDS Losss. These improvements enable UCG-GAN to generate high-quality images with a small number of samples and meet the actual application requirements.
[0047] Refer to Figure 4Structure of the generator G of the improved UCG - GAN model. The generator G of the improved UCG - GAN model includes 11 residual modules, specifically including five sequentially arranged encoder blocks, an intermediate block, and five decoder blocks. Let the numbers of the five encoder blocks be 4, 3, 2, 1, 0 in sequence; the numbers of the five decoders be 1, 2, 3, 4, 5 in sequence; four skip connections are sequentially arranged between the encoder blocks and the decoder blocks. Let r ∈ {1, 2, 3, 4}. One end of each skip connection is connected to the output end of the encoder block numbered r, and the other end of each skip connection is connected to the output end of the decoder block numbered r; a CA attention mechanism module is inserted into each skip connection. The CA attention mechanism module takes the encoder numbered r and the decoder numbered r as inputs, and concatenates the output of the CA attention mechanism module with the output of the decoder numbered r as the input of the decoder numbered r + 1; the input of the encoder block at the head end is the real defect image x k , and the last decoder block outputs the first wire rope defect image The structure reference of the CA attention mechanism module Figure 5 .
[0048] The CA attention mechanism enables the model to more precisely focus on important features by fusing global and local spatial information, thereby enhancing the feature representation ability. Figure 5 In [reference], the size of the input tensor is C × H × W, where C represents the number of channels, H represents the height, and W represents the width. GAP refers to Global Average Pooling. After the input tensor undergoes the Global Average Pooling (GAP) operation, the spatial dimensions (height and width) of the tensor are reduced to 1, and 1×1×C represents the size of the tensor after GAP. κ = ψ(C), which means adaptively selecting the kernel size according to the channel dimension. This is a dynamic process, and the network determines the appropriate kernel size or the number of key features to be concerned based on the input data. K = 5 output channels, which indicates that the network focuses on 5 key features or components. σ represents the sigmoid activation function.
[0049] Further preferably, each encoder block and decoder block each include three convolutional layers, using leaky ReLU and batch normalization. A downsampling layer is provided at the end of each encoder block, and the output end of the downsampling layer serves as the output end of the encoder block; an upsampling layer is provided at the end of each decoder block, and the output end of the upsampling layer can serve as the output end of the decoder block.
[0050] The discriminator D of the improved UCG - GAN model includes a convolutional layer, five residual blocks, and a fully - connected layer arranged in sequence; after the processing of the convolutional layer and each residual block, the fully - connected layer outputs the discriminator score. Refer to Figure 6 , the convolutional layer is Figure 6The light-colored part, the five residual layers are the dark blue parts stacked in sequence, and the fully connected layer is the orange part.
[0051] S3: The first wire rope defect image generated by using the real defect image and the generator G of the improved UCG-GAN model; using the generated first wire rope defect image to backpropagate and update the parameters of the generator G and the discriminator D, and generating the generator G and the discriminator D through iterative training; using the trained generator G and discriminator D to generate the second wire rope defect image, and establishing a wire rope defect sample library based on the second wire rope defect image.
[0052] Specifically, in step S3, using the generated first wire rope defect image to backpropagate and update the parameters of the generator G and the discriminator D, and generating the generator G and the discriminator D through iterative training, is to make the average value of the discriminator output scores of k real defect images x k be D(X), and the average value of the discriminator output scores of k first wire rope defect images be Adopting the hinge adversarial loss, the iterative training objective of the discriminator D is to minimize the discriminator loss function λ D and the iterative training objective of the generator G is to minimize the generator loss function λ GD : where is the expected value of the real defect image x k , is the expected value of the first wire rope defect image , max(·) is to use the ReLU activation function to ensure that the result of the loss function is non-negative; by defining the loss functions λ D and λ GD , the discriminator D and the generator G continuously confront each other during the training process, and finally reach a Nash equilibrium, so that the generator generates the second wire rope defect image closest to the real defect image.
[0053] The present invention further adds a classifier to the discriminator D, and uses GDSLoss to classify the real wire rope defect image and the generated wire rope defect image. The GDS loss adds a square to the original Generalized Dice Loss, avoiding the loss function being negative, ensuring that the category of the generated image is consistent with the conditional image, and improving the quality and authenticity of the generated wire rope defect image. The classification loss formula is: where P X and T X are the predicted region and the real region of the deterioration of category X respectively, and w X is the weight of category X.
[0054] S4: Establish and initialize the improved EfficientNet network model, train it using the wire rope defect sample library, and update the parameters of the improved EfficientNet network model to obtain the trained improved EfficientNet network model.
[0055] As Figure 7 and Figure 8 shown, the content of step S4 is as follows: The constructed improved EfficientNet network model includes a first convolutional module, seven sequentially arranged graph convolutional layer modules ATConv, a second convolutional module, average global pooling, and a fully connected layer; each graph convolutional layer module ATConv includes a first ordinary convolutional layer, an asymmetric convolutional module, a Triplet attention mechanism, a second ordinary convolutional layer, and a Dropout layer arranged in sequence, and a skip connection is also arranged between the input end of the first ordinary convolutional layer and the output end of the Dropout layer; the loss function of the graph convolutional layer module ATConv is the FILOSS loss function obtained by fusing the FocalLoss loss function and the IOULOSS loss function; the improved EfficientNet network model obtains the second wire rope defect image in the sample library as input, and identifies and marks the wire rope defect part for output.
[0056] Figure 7 In it, the 1 or 6 after ATConv represents the magnification factor n, that is, the first 1x1 convolutional layer in ATConv will expand the number of channels of the input feature matrix by n times, where k3x3 or k5x5 represents the convolutional kernel size used in the asymmetric convolution of ATConv.
[0057] The graph convolutional layer module ATConv is improved on the basis of MBConv in the EfficientDet network model. It replaces the Depwise convolution in the original MBConv module with an asymmetric convolution, replaces the original SE attention mechanism with a Triplet attention mechanism, and finally fuses the original FocalLoss loss function and the IOULOSS loss function to obtain a new FILOSS loss function, and finally forms the ATConv module:
[0058] After replacing it with an asymmetric convolution, the ATConv module can more effectively capture feature information in different directions and improve the model's ability to represent complex image structures;
[0059] After replacing it with a Triplet attention mechanism, the model can improve the attention to important regions in the feature map, thereby improving the model's ability to recognize targets;
[0060] After integrating the loss functions FocalLoss and IOULOSS, the new FILOSS can make the model perform better in dealing with imbalanced samples and target localization accuracy, and finally make the model have better robustness and accuracy in the detection task.
[0061] Refer to Figure 9 , Figure 9 Figures (a) and (b) of Figure 9 show the schematic diagrams of convolution processing using a horizontal asymmetric convolution kernel before and after horizontal flipping of the feature map, and figures (c) and (d) of Figure 9 show the schematic diagrams of convolution processing using a standard square convolution kernel before and after horizontal flipping of the feature map. It can be seen from Figure 10 that after the input feature map undergoes horizontal flipping, the features extracted by the standard square convolution change significantly, while the horizontal asymmetric convolution can still accurately extract features during the sliding convolution process. This shows that the asymmetric convolution is more robust in terms of rotational distortion characteristics. The structural diagram of the asymmetric convolution module is shown in
[0062] Figure 10 . In
[0063] , the part of ACB_N represents the input of this asymmetric convolution module; Zero Padding represents the zero-padding operation, which is used to keep the spatial dimensions of the input and output tensors consistent.
[0064] BN (Batch Normalization): Batch normalization is used to accelerate training and stabilize the training process of the neural network.
[0065] Leaky ReLU is a leaky ReLU activation function that allows a part of negative values to pass through to avoid neuron "death".
[0066] The asymmetric convolution module adds a horizontal asymmetric convolution kernel and a vertical asymmetric convolution kernel on the basis of the standard square convolution. The input image is respectively convolved by three different convolution kernels to extract different branch features: where M U,V,K ∈R U×V×C represents the feature map of the K-th channel with an input size of U×V, K = 1, 2,..., C, U≠V; represents the j-th convolution kernel of the K-th channel with an input size of H×W, H≠W; O U,V,j is the output feature map corresponding to the j-th convolution kernel, and * represents the dot product operation;
[0067] The additivity of convolution is used to fuse different branches to obtain the fused feature output. The dimension of the fused feature is consistent with that of the input feature and satisfies: I is the input feature map matrix; K1 and K2 respectively represent different convolution kernels; * represents the convolution operation; denotes the addition of the corresponding positions of the convolution kernels.
[0068] The Triplet attention mechanism includes the following: The first branch is the channel attention calculation branch. The input feature passes through Z-Pool, then through a 7×7 convolution, and through the Sigmoid activation function to generate the spatial attention weight; the second branch is the branch for capturing the interaction between the channel C and the spatial W dimensions. The input feature is first transformed into an H×C×W dimensional feature through the permutation function, then Z-Pool is performed on the H dimension, then processed through a 7×7 convolution and the Sigmoid activation function, and then transformed into a C×H×W dimensional feature through the permutation function; the third branch is the branch for capturing the interaction between the channel C and the spatial H dimensions. The input feature is first transformed into a W×H×C dimensional feature through the permutation function, then Z-Pool is performed on the W dimension, then processed through a 7×7 convolution and the Sigmoid activation function, and then transformed into a C×H×W dimensional feature through the permutation function; the output result of the Triplet attention mechanism is obtained by adding and averaging the output features of the three branches.
[0069] As Figure 11 shown, the Triplet attention mechanism consists of three parallel branches. Two of the branches are responsible for capturing the interactions between different dimensions: one focuses on the interaction between the channel dimension (C) and the spatial dimension (H or W), and the other is used to construct spatial attention. The outputs of all three branches are combined by averaging their respective weights and then aggregated. This innovative design has lateral interactions, solving the challenge of separating channel attention and spatial attention in traditional computational models. It captures the interaction between the spatial dimension and the channel dimension within the same framework. Substantially, it captures the information interactions in three dimensions simultaneously: (C, H), (C, W), and (H, W), representing the interactions between the channel, height, and width dimensions of the input data respectively.
[0070] The expression of the FILOSS loss function is: FILOSS = α × FocalLoss + (1 - α) × IOULOSS, where α is the weight; the FocalLoss loss function is defined as: FocalLoss = -α t (1 - p t )γ log(p t )), where pt is the probability value predicted by the model; α t is the adjustment factor; γ is a hyperparameter; the IOULOSS loss function is defined as follows: where A is the intersection of the predicted bounding box and the ground truth bounding box, and B is the union of the predicted bounding box and the ground truth bounding box. By combining FocalLoss and IOULOSS, the losses in both classification and localization can be comprehensively considered, improving the performance of the object detection model in dealing with easy and difficult samples and optimizing the training effect.
[0071] S5: Using the improved EfficientNet network model after training, defect recognition is performed on the real defect images of elevator wire ropes collected in real time by a number of image acquisition devices, and the defect recognition results are output.
[0072] To solve the problem that the lack of wire rope defect pictures affects the training of the detection model, the present invention proposes to improve the UCG-GAN network and use it for generating wire rope defect samples. The specific experimental results are as follows: The U-Net structure is introduced into the generator G, and 4 skip structures are added between the encoder block and the decoder block, and then the CA attention mechanism module is inserted into the skip structure; in the discriminator D, to stabilize the training process, the hinge adversarial loss is introduced and incorporated into the classifier, and a new classification loss function GDSloss is proposed to avoid the result of the loss function being negative, improving the quality and authenticity of the generated wire rope defect images. In addition, in the wire rope defect detection stage, the present invention proposes an improved EfficientNet network model, replacing the Depwise convolution in the original MBConv module with an asymmetric convolution, which has an asymmetric structure and the characteristics of scale-invariant feature transformation. The original SE attention mechanism is replaced with a Triplet attention mechanism, which can capture the interaction between the spatial dimension and the channel dimension at the same time. Finally, the original FocalLoss loss function is fused with the IOULOSS loss function to obtain a new FILOSS loss function, which simultaneously considers the losses in both classification and localization, and can improve the performance of the object detection model in dealing with easy and difficult samples and optimizing the training effect, finally forming an ATConv module. The final calculation results are as Figure 12 shown. It can be found from the figure that the improved EfficientNet model can not only accurately identify larger defects such as wear and breakage of elevator wire ropes, but also accurately identify smaller defects such as breakage defects, and the confidence levels can all reach above 0.9.
[0073] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting defects in traction elevator wire ropes based on deep learning, characterized in that: The steps include: S1: Use several image acquisition devices to perform multi-directional acquisition to obtain the real defect image of the elevator wire rope; S2: construct an improved UCG-GAN model, wherein the improved UCG-GAN model includes an independent generator G and a discriminator D; S3: Generate a first wire rope defect image using a real defect image and the generator G of the improved UCG-GAN model; Use the generated first wire rope defect image to back-propagate and update the parameters of the generator G and the discriminator D, and train the generator G and the discriminator D through iteration; Generate a second wire rope defect image using the trained generator G and the discriminator D, and establish a wire rope defect sample library based on the second wire rope defect image; S4: Establish and initialize an improved EfficientNet network model, use the wire rope defect sample library for training, update the parameters of the improved EfficientNet network model, and obtain the trained improved EfficientNet network model; the content of step S4 is: the constructed improved EfficientNet network model includes a first convolution module, seven graph convolution layer modules ATConv, a second convolution module, an average global pooling and a fully connected layer set in sequence; each graph convolution layer module ATConv includes a first ordinary convolution layer, an asymmetric convolution module, a Triplet attention mechanism, a second ordinary convolution layer and a Dropout layer set in sequence, and a jump connection is also set between the input end of the first ordinary convolution layer and the output end of the Dropout layer; the loss function of the graph convolution layer module ATConv is a FILOSS loss function obtained by fusing the FocalLoss loss function with the IOULOSS loss function; the improved EfficientNet network model obtains the second wire rope defect image in the sample library as input, identifies and marks the defective part of the wire rope for output; S5: Using the trained improved EfficientNet network model, defect recognition is performed on real defect images of elevator wire ropes collected in real time by several image acquisition devices, and the defect recognition results are output.
2. A traction elevator wire rope defect detection method based on deep learning according to claim 1, characterized in that: The content of step S1 is: configure more than two image acquisition devices, set the more than two image acquisition devices on the end surface of the elevator room floor close to the car side, and evenly distribute them in the circumferential direction; the more than two image acquisition devices are respectively aimed at the surface of different positions of the same wire rope to obtain the real defect image of the elevator wire rope.
3. A traction elevator wire rope defect detection method based on deep learning according to claim 2, characterized in that: The generator G of the improved UCG-GAN model includes five encoder blocks, one intermediate block and five decoder blocks arranged in sequence, and the five encoder blocks are numbered 4, 3, 2, 1, and 0 in sequence; the five decoders are numbered 1, 2, 3, 4, and 5 in sequence; four jump connections are arranged in sequence between the encoder block and the decoder block, and r∈{1, 2, 3, 4} is set, one end of each jump connection is connected to the output end of the encoder block numbered r, and the other end of each jump connection is connected to the output end of the decoder block numbered r; a CA attention mechanism module is inserted into each jump connection, and the CA attention mechanism module takes the encoder numbered r and the decoder numbered r as input, and the output of the CA attention mechanism module is connected in series with the output of the decoder numbered r as the input of the decoder numbered r+1; the input of the encoder block at the head end is the real defect image x k , the last decoder block outputs the first wire rope defect image 4. A traction elevator wire rope defect detection method based on deep learning according to claim 3, characterized in that: Each encoder block and decoder block includes three convolutional layers, using leaky ReLU and batch normalization. A downsampling layer is set at the end of each encoder block, and the output of the downsampling layer is used as the output of the encoder port; an upsampling layer is set at the end of each decoder block, and the output of the upsampling layer can be used as the output of the decoder block.
5. The method for detecting defects in traction elevator wire ropes based on deep learning according to claim 3, characterized in that: The discriminator D of the improved UCG-GAN model includes a convolutional layer, five residual blocks and a fully connected layer arranged in sequence; after being processed by the convolutional layer and each residual block, the fully connected layer outputs a discriminator score.
6. A method for detecting defects in traction elevator wire ropes based on deep learning according to claim 5, characterized in that: The first wire rope defect image generated in step S3 is used to back propagate and update the parameters of the generator G and the discriminator D, and the generator G and the discriminator D are trained iteratively. k The average value of the discriminator output score is D(x), k first wire rope defect images The average of the discriminator output scores is Using hinge adversarial loss, the iterative training goal of the discriminator D is to minimize the discriminator loss function λ D , the iterative training goal of the generator G is to minimize the generator loss function λ GD : in is the real defect image x k The expected value of The first wire rope defect image The expected value of max(·) is the ReLU activation function, which ensures that the loss function result is non-negative. By defining the loss function λ D and λ GD ,The discriminator D and the generator G continuously compete with each other during the training process, and finally reach a Nash equilibrium so that the generator generates the second wire rope defect image that is closest to the real defect image.
7. The method for detecting defects in traction elevator wire ropes based on deep learning according to claim 1, characterized in that: The asymmetric convolution module adds horizontal asymmetric convolution kernels and vertical asymmetric convolution kernels on the basis of standard square convolution. The input image is convolved with three different convolution kernels to extract different branch features: Among them, M U,V,K ∈R U×V×C Represents the feature map of the Kth channel with input size U×V, K=1, 2, ..., C, U≠V; represents the jth convolution kernel of the Kth channel with an input size of H×W, H≠W; O U,V,j is the output feature map corresponding to the jth convolution kernel, * represents the dot multiplication operation; The additivity of convolution is used to fuse different branches to obtain the fused feature output. The dimension of the fused feature is consistent with the dimension of the input feature, satisfying: I is the input feature map matrix; K1 and K2 represent different convolution kernels; * represents the convolution operation; Indicates the addition of the corresponding positions of the convolution kernel.
8. The method for detecting defects in traction elevator wire ropes based on deep learning according to claim 7, characterized in that: The Triplet attention mechanism includes the following contents: the first branch is the channel attention calculation branch, the input feature passes through Z-Pool, then passes through 7×7 convolution, and passes through the Sigmoid activation function to generate spatial attention weights; the second branch is the channel C and space W dimension interaction capture branch, the input feature is first transformed into H×C×W dimensional features through the permutation function, then Z-Pool is performed on the H dimension, then processed by 7×7 convolution and Sigmoid activation function, and then converted into C×H×W dimensional features through the permutation function; the third The branch is a channel C and space H dimension interactive capture branch. The input feature is first transformed into a W×H×C dimensional feature through the permutation function, then Z-Pool is performed on the W dimension, and then processed by a 7×7 convolution and a Sigmoid activation function, and then transformed into a C×H×W dimensional feature through the permutation function. The output result of the Triplet attention mechanism is obtained by adding and averaging the output features of the three branches.
9. The method for detecting defects in traction elevator wire ropes based on deep learning according to claim 1, characterized in that: The expression of FILOSS loss function is: FILOSS = α × FocalLoss + (1-α) × IOULOSS, α is the weight; The FocalLoss loss function is defined as: FocalLoss = -α t (1-p t ) γ log(p t ), where p t is the probability value predicted by the model; α t is the adjustment factor; γ is a hyperparameter; the IOULOSS loss function is defined as follows: Where A is the intersection of the predicted box and the true box, and B is the union of the predicted box and the true box.
Citation Information
Patent Citations
PCB defect data generation method based on deep learning
CN111798409A
Strip steel surface defect detection method and system
CN113450344A