Steel wire rope defect detection method with multi-dimensional interpretability

By using YOLOv8 deep learning model and multi-dimensional interpretability technology in wire rope defect detection, the problem of manual detection time-consuming and labor-intensive and deep learning model black boxing in the existing technology is solved, and efficient, accurate and transparent wire break detection effect is achieved.

CN120047424AActive Publication Date: 2025-05-27INST OF ENERGY HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ENERGY LAB)

Patent Information

Application Number
CN202510179921.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-27
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

In the detection of wire rope defects, the problem of manual detection is time-consuming and labor-intensive, the black box nature of deep learning models and the possibility of identifying wrong features, resulting in insufficient detection accuracy and low trust.

Method used

The YOLOv8 deep learning model combined with multi-dimensional interpretability technology is used to collect electromagnetic signals from the wire rope through sensors, convert them into image data, train and predict, and interpretability analysis is carried out through GradCAM++ and LIME methods to improve the transparency and accuracy of the model.

Benefits of technology

The efficiency and accuracy of wire rope defect detection are achieved, while improving the transparency and credibility of the model, ensuring the reliability of wire break detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047424A_ABST
    Figure CN120047424A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial detection maintenance and deep learning, in particular to a steel wire rope defect detection method with multi-dimensional interpretability, and the technical scheme comprises the steps of steel wire rope broken wire data set manufacturing, YOLOv8 model training and prediction and model evaluation and optimization. According to the method, the electromagnetic signal is converted into the image, the YOLOv8 deep learning model is used for training and predicting the broken wire of the steel wire rope, meanwhile, the multi-dimensional interpretability technology is used for carrying out interpretability analysis on the model from different angles, and indexes for quantifying the interpretability of the model are provided, so that the broken wire is more concrete, the identification degree is higher, and the method is suitable for popularization and application. The model detection speed is high, complex intermediate steps are not needed, model transparency can be improved, and the method plays an important role in guiding model optimization and increasing credibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of industrial inspection and maintenance and deep learning, and particularly relates to a wire rope defect detection method with multi-dimensional interpretability. Background Art

[0002] At present, wire ropes play a very important role in fields such as mining and construction. Broken wires are a key factor affecting their safety. Therefore, detecting broken wires is a very important link. Manual detection methods are time-consuming and laborious. Although the method of deep learning can avoid this problem and achieve a relatively high detection accuracy, deep learning belongs to a black box. It is difficult to believe the model and its prediction results without understanding how the model makes decisions, and it is very likely that the model makes predictions by identifying incorrect features.

[0003] For engineering materials such as wire ropes that are used frequently and under long-term stress, the electromagnetic broken wire signals generated by wire ropes in different usage environments are difficult to identify and analyze at the data level. This is because the gaps at the broken wire points are different, resulting in each broken wire signal being different in shape and size, and this phenomenon can be more clearly reflected in images.

[0004] Training a batch of images of wire rope broken wires based on a deep learning model can enable the model to learn the features of broken wires, so that the model can help us predict the broken wire situation of wire ropes. Although the speed of the model is difficult to achieve manually, there are often phenomena of insufficient accuracy or incorrect prediction. Moreover, due to the characteristics of deep learning models such as non-linearity, a huge number of parameters, and opaque internal working mechanisms, it becomes a problem to gain people's trust. And the number of broken wires in wire ropes is often an important standard for scrapping. Therefore, if the prediction is incorrect, the consequences will be relatively serious.

[0005] In summary, the present application proposes a wire rope defect detection method with multi-dimensional interpretability. Summary of the Invention

[0006] The purpose of the present invention is to propose a wire rope defect detection method with multi-dimensional interpretability for the problem that it is time-consuming and laborious to detect wire rope defects by manual detection methods in the background art.

[0007] The technical solution of the present invention: A wire rope defect detection method with multi-dimensional interpretability includes the following steps:

[0008] Production of wire rope broken wire dataset: Collect the electromagnetic signals of the wire rope through sensors to obtain data on local damage (LF) and loss of metal cross-sectional area (LMA). The English abbreviation for local damage is LF, and the English abbreviation for loss of metal cross-sectional area is LMA. Combine the local damage (LF) and loss of metal cross-sectional area data to manually label the wire rope broken wire conditions and categories. Convert the local damage (LF) and loss of metal cross-sectional area data at a fixed distance into image data and perform annotation to produce a dataset format suitable for the YOLOv8 model;

[0009] Training and prediction of the YOLOv8 model: Prepare the above dataset, select a YOLOv8 model of appropriate size, specify important hyperparameters and then train. Roughly evaluate the model based on the evaluation metrics after training. Use the trained model for prediction on a new dataset and count the prediction results;

[0010] Model evaluation and optimization: Use gradient-based visualization explanation techniques and local interpretable model-agnostic explanations to analyze the prediction results in combination. GradCAM++ is the English abbreviation for gradient-based visualization explanation techniques and LIME is the English abbreviation for local interpretable model-agnostic explanations. Judge the correctness of the model's recognition features; if correct, add an attention mechanism to judge the comprehensiveness of the features and calculate relevant interpretability metrics; Based on the accuracy and comprehensiveness of the model features, judge whether the model's accuracy and speed meet the standards. If not, take targeted measures.

[0011] Optionally, when using GradCAM++ and LIME to analyze the prediction results in combination, specifically obtain the heatmap of GradCAM++ regarding the prediction results and the pixel blocks in the image generated by LIME that have positive and negative impacts on the model, and based on the formula:

[0012] S LIME =S LIME+ +S LIME- (S LIME- ≤0,S LIME+ ≥0)

[0013] Calculate the interpretability of LIME, where S LIME represents the measure of the interpretability of LIME for the model, S LIME+ is the area of the pixel blocks in the image generated by LIME that have a positive effect on the model's decision-making, and S LIME- is the area of the pixel blocks with negative impacts.

[0014] Optionally, calculate the intersection over union of the area of the GradCAM++ heatmap in the prediction box and the area of the effective region generated by LIME in the prediction box as the interpretability metric. The calculation formula is:

[0015]

[0016] Among them, the numerator S CAM++ ∩S LIME is the intersection area of two regions, and the denominator S CAM++ ∪S LIME is the union area of two regions; then through the formula:

[0017]

[0018] the interpretability Y in the prediction box is obtained, where S box is the area of the prediction box.

[0019] Optionally, considering the influence outside the prediction box, a penalty term L is added, and the calculation formula is:

[0020]

[0021] Among them, φ is the control coefficient, C CAM++ is the area occupied by the heat map outside the prediction box, C LIME+ and C LIME- are the areas of the pixel blocks with positive and negative influences outside the prediction box respectively, S is the total area of the prediction image, S - S box represents the image area outside the prediction box, and then through the formula:

[0022] P = Y - L (L ≥ 0)

[0023] the actual interpretability P of the model is calculated, where Y is the interpretability of the prediction box and L is the penalty term.

[0024] Optionally, when adding an attention mechanism to judge the comprehensiveness of features, through the formula:

[0025] P attention = P after - P before

[0026] the contribution P attention of the interpretability brought by the attention mechanism is calculated, where P before is the interpretability P of the model before adding the attention mechanism, and P after is the interpretability of the model after adding the attention mechanism; then according to the formula:

[0027] P interpretability = P + P attention

[0028] the interpretability P interpretability of the model's prediction for a single picture is calculated, where P is the actual interpretability of the model and P attention is the contribution of the interpretability brought by the attention mechanism.

[0029] Optionally, the ratio of the broken wires detected by the model on the test set to the actual broken wires in the test set is used as the standard for judging the accuracy of the model. A model accuracy of 90% and above is considered high, a model accuracy of 70%-90% is considered medium, and a model accuracy below 70% is considered low, that is, the model accuracy does not meet the standard. After judging whether the model accuracy and speed meet the standard, when the model accuracy does not meet the standard, view the interpretability analysis, calculate the interpretability index, and observe the proportion of the heat distribution in the non-broken wire part of the image and the proportion of the continuous pixel area with a negative impact on the prediction; adopt data augmentation methods to add noise samples to the training data, and the types of noise samples include Gaussian noise and salt-and-pepper noise.

[0030] Optionally, a batch of images is generated from a complete steel wire rope, and the quantity depends on the length of the steel wire rope. All batches of images form the test set, without labeling the broken wires, and there are images without broken wires in each batch. Let the number of images on the test set be N, and use the formula:

[0031]

[0032] Calculate the actual interpretability P of the model on the test set final , which is used to comprehensively evaluate the performance of the model on the test set.

[0033] Compared with the prior art, the present application includes at least one of the following beneficial technical effects:

[0034] 1. By using images to represent the broken wires of the steel wire rope, the performance of the broken wires is made more concrete, increasing the recognition of different types of broken wires.

[0035] 2. By using the YOLOv8 deep learning model to detect broken wires, the model can capture similar features between broken wires, and the model belongs to an end-to-end method, without complex intermediate steps and with a fast detection speed.

[0036] 3. By using multi-dimensional interpretability techniques to improve the transparency of the YOLOv8 model, it can not only increase people's trust in the model, but also be used to guide the model to improve and provide directions.

[0037] 4. A method for quantifying the strength of the interpretability of the YOLOv8 model is proposed.

[0038] The present invention converts electromagnetic signals into images, uses the YOLOv8 deep learning model for training and prediction of broken wires of steel wire ropes, simultaneously adopts multi-dimensional interpretability techniques to conduct interpretability analysis on the model from different angles, and proposes an index for quantifying the strength of the model's interpretability, making the performance of broken wires more concrete and having higher recognition, the model has a fast detection speed and does not require complex intermediate steps, and can also improve the transparency of the model, which plays an important role in guiding model optimization and increasing credibility. Description of the Drawings

[0039] Figure 1 It is a flowchart of a steel wire rope defect detection method with multi-dimensional interpretability. Detailed Implementation Manner

[0040] The technical solution of the present invention will be further described below in conjunction with the drawings and specific embodiments.

[0041] Embodiment

[0042] As Figure 1 shown, a steel wire rope defect detection method with multi-dimensional interpretability proposed by the present invention consists of three parts: making a steel wire rope broken wire dataset, training a YOLOv8 model, and evaluating and optimizing the model by combining prediction results and multi-dimensional interpretability. The following is a detailed description of the three parts.

[0043] The first part is to make a steel wire rope broken wire dataset. After collecting the electromagnetic signals of the steel wire rope on the sensor, data on local damage and metal cross-sectional area loss of the steel wire rope are obtained. LMA is the English abbreviation for metal cross-sectional area loss, and LF is the English abbreviation for local damage. By combining LF data and LMA data, the broken wire condition and its broken wire category at a certain position of the steel wire rope are manually marked, and it is used as a reference for subsequent image annotation. After converting LF and LMA data at a fixed distance into image data, annotation is performed on the image, and it is made into the dataset format required by the YOLOv8 model. The first stage is to collect the electromagnetic signals of the steel wire rope to be measured through an electromagnetic sensor, convert the collected electromagnetic signal data into image data, and there is no data loss during the conversion process. Therefore, the conversion process is lossless, so as to achieve the data basis for using deep learning.

[0044] The second part is the training and prediction of the YOLOv8 model. Since the models of the YOLO series are end-to-end models, after preparing the dataset, selecting a model of appropriate size, and specifying some important hyperparameters, training can be started. At the same time, the model can be roughly evaluated according to some evaluation indicators obtained after the training ends. Finally, the trained model can be put into a new dataset for prediction, and the prediction results are counted for the analysis of the next part. The YOLOv8 model is trained and verified on the data of the previous stage to obtain a model that can both identify the type of broken wire signal and locate the position of the broken wire.

[0045] The third part evaluates and optimizes the model by combining the prediction results and multi-dimensional interpretability. The multi-dimensional interpretability includes interpretable analysis of the model's predictions from multiple perspectives, namely, pre-hoc, post-hoc, and local. When conducting interpretable analysis on the prediction results, we first analyze the prediction results by combining GradCAM++ and LIME. We can simultaneously obtain the heatmap of the prediction results from GradCAM++ and the pixel blocks in the image generated by LIME that have a positive effect and a negative impact on the model. By combining the interpretability of both, we can more accurately know the partial broken wire features recognized by the model, and then judge whether the features recognized by the model are correct, that is, whether they conform to the basis for humans to recognize broken wires in images. On this basis, if they do not conform, the dataset needs to be modified to observe whether there are biases in the dataset that cause the model to learn some features that should be ignored. If the model's judgment basis is correct, we also need to see whether the features learned by the model are comprehensive. This is to exclude the influence that the model may only focus on part of the features belonging to broken wire judgment. Therefore, an attention mechanism needs to be added to the model for pre-hoc interpretability. During this process, the interpretability degree will be calculated to judge whether the learned features are comprehensive. Finally, on the basis that the model recognizes features accurately and comprehensively, prediction data is collected. The ratio of the broken wires detected by the model on the test set to the actual broken wires in the test set is used as the standard for judging the model's accuracy. A model accuracy of 90% and above is considered high, a model accuracy of 70%-90% is considered medium, and a model accuracy of below 70% is considered low, that is, the model accuracy does not meet the standard. After judging whether the model accuracy and speed meet the standard, when the model accuracy does not meet the standard, check the interpretable analysis, calculate the interpretability index, and observe the proportion of the heat distribution in the non-broken wire part of the image and the proportion of the continuous pixel area with a negative impact on the prediction; adopt data augmentation methods to add noise samples to the training data. The types of noise samples include Gaussian noise and salt-and-pepper noise, so that the model can reduce its attention to the non-broken wire part and make the model more generalized.

[0046] For the YOLOv8 model, the model's prediction includes two parts, namely the prediction box and the category and confidence of the object within the corresponding prediction box. Let the area of the prediction box in the image be S box , and the area of the heatmap distributed in the prediction box is used as the interpretability S provided by GradCAM++ CAM++ . Let the area of the image generated by LIME that has a positive effect in the prediction box be S LIME+ , and the area of the image with a negative impact is used as S LIME- . Therefore, the interpretability of LIME is:

[0047] S LIME =S LIME+ +S LIME- (S LIME- ≤0, SLIME+ ≥ 0)

[0048] To obtain the region that the model actually focuses on and eliminate the bias caused by using only one interpretability method, we use the IOU of the two as interpretability:

[0049]

[0050] After calculating the IOU and dividing by S box we obtain the interpretability degree in the prediction box on a prediction image, which we regard as the contribution part to the interpretability degree generated by the model and denote it as Y:

[0051]

[0052] Actually, when there are heatmaps, positive effects, and negative influence pixel blocks outside the prediction box in the prediction image, it will have a negative impact on the interpretability of the model. We will add a penalty term L to reflect this impact. Let S be the area of the prediction image, the area of the heatmap outside the prediction box be C CAM++ , and the areas of the pixel blocks with positive and negative influences outside the prediction box be C LIME+ , C LIME- , and φ be the control coefficient used to control the size of this penalty term. Then this penalty term is expressed as:

[0053]

[0054] Therefore, the actual interpretability degree of this model is the difference between the interpretability degree of the prediction box that actually contributes to the model's interpretability and the penalty term outside the prediction box:

[0055] P = Y - L (L ≥ 0)

[0056] To obtain the interpretability contribution brought by the attention mechanism, we calculate P before before adding the attention mechanism and P after after adding the attention mechanism respectively, and obtain P attention by taking the difference:

[0057] P attention = P after - P before

[0058] Finally, we obtain the interpretability degree of the model's prediction for a single picture:

[0059] P interpretability = P + P attention

[0060] A batch of images is generated by a complete steel wire rope, and the quantity depends on the length of the steel wire rope. All batches of images form a test set. No broken wires are labeled, and there are images in each batch that do not contain broken wires. Let the number of pictures in the test set be N. Execute this algorithm on the entire test set, and finally use the root mean square to calculate the practical interpretability P of the model on the test set. final :

[0061]

[0062] In this embodiment, multi-dimensional interpretability is used to understand what features the model identifies as the basis for judging the presence of broken wires and the types of broken wires in the image, and to verify whether the features identified by the model are consistent with the human judgment method. Based on this, the direction of model improvement is determined, so that the model and the prediction results gain people's trust. Finally, a model that is both sufficiently interpretable and accurate and efficient is put into industrial applications. The third stage belongs to the main stage. Among them, multi-dimensional interpretability means combining multiple different interpretability methods. These interpretability methods perform interpretability analysis on the model from three different perspectives: before, after, and locally. The attention mechanism, GradCAM++, and LIME interpretability methods are used respectively, making the model have sufficient interpretability and being easily understood by people. In the used interpretability methods, the post-hoc local interpretability analysis is mainly composed of GradCAM++ and LIME, and the pre-hoc local interpretability is composed of the attention mechanism.

[0063] Among them, interpretability is used to quantify the strength of the interpretability of the model. Interpretability is based on the area of the interpretability regions generated by the interpretability methods GradCAM++ and LIME on the predicted images. Taking the prediction box as the critical point, the interpretability regions generated within the prediction box are regarded as positive contributions, while the interpretability regions generated outside the prediction box are regarded as negative impacts. The appearance of interpretability regions outside the prediction box is because the model focuses on features that are irrelevant to predicting broken wires, so it is regarded as a negative impact. Interpretability is exactly the result of the interaction between the two. The larger the result, the more features the model focuses on are generated from within the prediction box, indicating that the basis for the model to make a broken wire judgment is consistent with human judgment, thus making the model have good interpretability.

[0064] In addition, since the heatmap obtained by GradCAM++ can show some parts of the image that the model focuses on, but because this method is post-hoc local interpretability, there will be deviations in the heat distribution even for the same image. Therefore, simply using GradCAM++ to explain the model prediction results will lack persuasiveness. At this time, the pixel blocks provided by LIME that have positive and negative impacts on the model play a very good complementary effect. By innovatively using a combined method, the explanation becomes more persuasive.

[0065] Since the essence of the attention mechanism is to improve the model's ability to recognize key features and regions in images, while reducing the attention to unimportant information and enhancing the generalization ability of the model. Therefore, by using the attention mechanism, the features focused on by the model can be made more comprehensive, thereby increasing the interpretability of the model.

[0066] The above specific embodiments are only several alternative embodiments of the present invention. Based on the technical solution of the present invention and the relevant inspirations of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. A wire rope defect detection method with multi-dimensional interpretability, characterized in that: The following steps are involved: Preparation of wire rope broken wire dataset: The electromagnetic signal of the wire rope is collected by sensors to obtain local damage and metal cross-sectional area loss data. The wire rope broken wire situation and category are manually annotated based on the local damage and metal cross-sectional area loss data. The local damage and metal cross-sectional area loss data at a fixed distance are converted into image data and annotated to produce a dataset format suitable for the YOLOv8 model. YOLOv8 model training and prediction: Prepare the above dataset, select a YOLOv8 model of appropriate size, specify important hyperparameters and perform training, roughly evaluate the model based on the evaluation indicators after training, use the trained model for prediction of a new dataset and calculate the prediction results; Model evaluation and optimization: Use the gradient-based visualization interpretation technology and local interpretable model-independent interpretation to analyze the prediction results. The English abbreviation of the gradient-based visualization interpretation technology is GradCAM++, and the English abbreviation of the local interpretable model-independent interpretation is LIME to judge the correctness of the model recognition features; If correct, an attention mechanism is added to determine the comprehensiveness of the features and calculate relevant interpretability indicators. Based on the accuracy and comprehensiveness of the model features, it is determined whether the model accuracy and speed meet the standards. If not, targeted measures are taken.

2. A wire rope defect detection method with multi-dimensional interpretability according to claim 1, characterized in that: When using GradCAM++ and LIME to analyze the prediction results, we obtain the heat map of the prediction results of GradCAM++ and the pixel blocks that have positive and negative effects on the model in the image generated by LIME, and use the formula: S LIME =S LIME+ +S LIME- (S LIME- ≤0,S LIME+ ≥0) Calculate the interpretability of LIME, where S LIME Represents LIME's measure of model interpretability, S LIME+ is the area of ​​the pixel block in the image generated by LIME that has a positive effect on the model decision, S LIME- is the area of ​​the pixel block that has a negative impact.

3. A wire rope defect detection method with multi-dimensional interpretability according to claim 2, characterized in that: The intersection of the area of ​​the GradCAM++ heat map in the prediction box and the area of ​​the effective area generated by LIME in the prediction box is calculated as the interpretability metric. The calculation formula is: Among them, the molecule S CAM++ ∩S LIME is the intersection area of ​​the two regions, and the denominator S CAM++ ∪S LIME is the union area of ​​the two regions; then by the formula: Get the interpretability Y in the prediction box, where S box is the prediction box area.

4. A wire rope defect detection method with multi-dimensional interpretability according to claim 3, characterized in that: Considering the impact outside the prediction box, a penalty term L is added, and the calculation formula is: Where φ is the control coefficient, C CAM++ is the area of ​​the heat map outside the prediction box, C LIME+ and C LIME- are the areas of the pixel blocks with positive and negative effects outside the prediction box, S is the total area of ​​the predicted image, and SS box Represents the image area outside the prediction box, and then through the formula: P=YL(L≥0) The actual interpretability P of the model is calculated, where Y is the interpretability of the prediction box and L is the penalty term.

5. A wire rope defect detection method with multi-dimensional interpretability according to claim 4, characterized in that: When adding the attention mechanism to judge the comprehensiveness of features, the formula is: P attention =P after -P before Calculate the explainability contribution P brought by the attention mechanism attention , where P before is the interpretability of the model before adding the attention mechanism, P after is the interpretability of the model after adding the attention mechanism; according to the formula: P interpretability =P+P attention Calculate the interpretability P of the model's prediction for a single image interpretability , where P is the actual interpretability of the model, P attention It is the interpretability contribution brought by the attention mechanism.

6. A wire rope defect detection method with multi-dimensional interpretability according to claim 5, characterized in that: The ratio of the broken wires detected by the model on the test set to the broken wires actually existing in the test set is used as the standard for judging the accuracy of the model. Model accuracy of 90% and above is considered to be high, model accuracy of 70%-90% is considered to be medium, and model accuracy below 70% is considered to be low, that is, the model accuracy does not meet the standard. After judging whether the model accuracy and speed meet the standards, when the model accuracy does not meet the standards, check the interpretability analysis, calculate the interpretability index, observe the proportion of thermal distribution of non-broken wire parts in the image, and the proportion of continuous pixel areas that have a negative impact on the prediction; use data enhancement methods to add noise samples to the training data. The types of noise samples include Gaussian noise and salt and pepper noise.

7. A wire rope defect detection method with multi-dimensional interpretability according to claim 6, characterized in that: A batch of images is generated by a complete wire rope. The number depends on the length of the wire rope. All batches of images constitute the test set. Broken wires are not marked, and there are images without broken wires in each batch. Suppose the number of images in the test set is N, and the formula is used: Calculate the actual interpretability P of the model on the test set final , used to comprehensively evaluate the performance of the model on the test set.

Citation Information

Patent Citations

  • Interpretable kidney tumor identification method and imaging method for interpretable kidney tumor identification method

    CN113298782A

  • Information processing device, information processing method, and information processing program

    CN114586044A

  • Model integration method and device based on interpretability, equipment and medium

    CN116662888A

  • Method for detecting abnormal winding of steel wire rope reel

    CN117611544A

  • Medical abnormity violation big data risk early warning method based on unsupervised machine learning and integrated learning

    CN117764741A

Cited By

  • Alloy steel wire rope flaw detection defect model construction method, equipment and storage medium

    CN121724963A

  • An alloy steel wire rope flaw detection defect model construction method, device and storage medium

    CN121724963B