Strain clamp defect detection method based on X-ray-visible light dynamic fusion

Through the dynamic channel attention mechanism and gradient fusion optimization method, combined with the cross-modal graph neural network, the modal conflict and environmental adaptability problems in tension clamp detection are solved, and high-precision defect detection and operation and maintenance decision-making are achieved.

CN120352453APending Publication Date: 2025-07-22CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510417202.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art has problems such as single-modal limitations, poor adaptability of static fusion environments, and lack of defect semantic associations in tension clamp detection, resulting in low detection accuracy and inaccurate operation and maintenance decisions.

Method used

The dynamic channel attention mechanism (DECA) is used to adjust the modal weights in combination with environmental parameters, and the dynamic fusion of multimodal information is achieved through the gradient fusion optimization method, and a cross-modal graph neural network is built for defect causal reasoning. Combining lightweight network design and industry standard rule database, a dynamic 'fusion-semantic modeling-closed-loop decision-making' architecture is built.

Benefits of technology

It improves detection accuracy, enhances environmental robustness, realizes the accuracy of operation and maintenance decisions, and supports power personnel to quickly formulate maintenance strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005344262560000031
    Figure BDA0005344262560000031
  • Figure BDA0005344262560000041
    Figure BDA0005344262560000041
  • Figure BDA0005344262560000061
    Figure BDA0005344262560000061
Patent Text Reader

Abstract

The invention discloses a strain clamp defect detection method based on X-ray-visible light dynamic fusion, and belongs to the technical field of intelligent detection of power equipment. Aiming at the problems of single-mode detection limitation, static fusion weight solidification and internal and external defect semantic association deficiency in a traditional method, the method is innovated in the following aspects: X-ray and visible light channel weights are distributed in real time through a dynamic channel attention mechanism (DECA) in combination with environmental parameters and feature quality; dynamic fusion is realized by adopting a gradient fusion optimization method; a graph neural network (GNN) is constructed, internal cracks detected by X rays and surface corrosion identified by visible light are mapped into graph nodes, an edge connection rule is defined based on a strain clamp physical structure, causal reasoning of internal and external defects is realized by using a graph attention mechanism (GAT), detection precision and environment robustness can be effectively improved, and operation and maintenance decision precision is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a live detection device and method for strain clamps of transmission lines based on dual unmanned aerial vehicles, and belongs to the technical field of non-destructive testing of strain clamps for high-voltage transmission lines. Background Art

[0002] As a key fitting connecting conductors and towers in transmission lines, the strain clamp is used to fix conductors or lightning conductors on the strain insulator strings of non-straight towers, playing the roles of anchoring and conducting electricity. Its crimping quality directly affects the safe operation of the power grid.

[0003] With the advancement of the construction of smart grids, traditional detection methods for strain clamp defects have evolved from manual inspections to technologies such as ultrasonic flaw detection and infrared thermal imaging. However, they all have problems of detection blind spots and environmental adaptability. In recent years, intelligent detection technologies based on computer vision and deep learning have become research hotspots, especially achieving remarkable progress in the field of unmanned aerial vehicle inspections. However, existing methods still face core technical bottlenecks such as insufficient utilization of multi-modal information and inadequate defect correlation analysis.

[0004] In terms of single-modal detection, although X-ray imaging can penetrate the metal matrix to identify internal cracks and crimping voids, it is easily interfered by metal artifacts and cannot detect surface corrosion. Visible light imaging can quickly obtain appearance defects, but is significantly affected by factors such as light and weather. The miss rate of traditional methods in complex environments is as high as 23%.

[0005] In multi-modal fusion technologies, existing solutions generally adopt fixed weight strategies such as average weighting or feature splicing, and cannot dynamically adjust the modal contribution according to environmental changes. For example, when there is strong light overexposure or severe metal artifacts, the weights are still forced to be evenly divided, resulting in a decline in detection accuracy. Moreover, the independent optimization of dual-modal branches is prone to gradient conflicts, leading to unstable model training.

[0006] At the level of defect correlation analysis, existing technologies mostly stay at single-modal defect identification and do not establish a causal relationship model between internal and external defects. For example, the stress conduction relationship between surface corrosion and internal cracks has not been explored, resulting in one-sided operation and maintenance decisions. In addition, the physical relevance of the complex structures of the crimping area and bolt connection points of the strain clamp has not been fully utilized. The node connection rules of existing graph neural network methods do not combine the stress distribution characteristics of 3D CAD models, affecting the accuracy of causal reasoning.

[0007] In terms of scene adaptability, existing deep learning models have a large number of parameters, and the inference delay on the unmanned aerial vehicle side exceeds 400 ms, making it difficult to meet the requirements of real-time inspections. Moreover, the detection results are insufficiently connected with the grading disposal suggestions of industry standards such as DL / T 664-2016.

[0008] At present, there are problems in the detection of this fitting on the market, such as large limitations of single modality, poor adaptability to static fusion environment, lack of semantic association of defects, and weak pertinence of detection methods. Summary of the Invention

[0009] In view of the above defects of the prior art, the present invention proposes a dynamic channel attention mechanism (DECA) to adjust the modality weights in real time in combination with environmental parameters, a gradient balance optimization method to suppress the risk of modality dominance, cross-modal graph neural network modeling to realize defect causal reasoning, and through lightweight network design and industry standard adaptation rule library, a three-level innovation architecture of "fusion-semantic modeling-closed-loop decision-making" is constructed, breaking through the bottlenecks of traditional detection technologies in the utilization of multimodal information, defect correlation analysis and complex scene adaptability, and providing a new paradigm for the intelligent and accurate detection of strain clamps on transmission lines.

[0010] A method for detecting defects of strain clamps based on dynamic fusion of X-ray and visible light constructs a method for detecting defects of strain clamps through a three-level innovation architecture of "dynamic fusion-semantic modeling-closed-loop decision-making", which specifically includes the following steps:

[0011] S1. Dual-modal data acquisition: Use a dual-payload drone to carry an X-ray sensor and a visible light camera to synchronously acquire dual-modal image data of the strain clamp, and generate an X-ray feature map and a visible light feature map respectively.

[0012] S2. Perform metal artifact suppression and contrast enhancement processing on the X-ray image to extract internal defect features; perform illumination equalization and motion blur correction processing on the visible light image to extract surface defect features; output a new dual-modal feature map.

[0013] S3. Input the dual-modal feature map and environmental parameters into the dynamic channel attention module, dynamically allocate channel weights according to the illumination intensity and metal artifact level, and generate a preliminary fusion feature.

[0014] S4. Use the gradient fusion algorithm to balance the loss function weights of the dual-modal branches, and optimize network training in combination with the gradient truncation technique.

[0015] S5. Construct a cross-modal graph neural network, map internal defects and surface anomalies to graph nodes, define physical association edges based on the 3D CAD model of the strain clamp, and model the defect causal chain through the graph attention mechanism.

[0016] S6. Combine the defect association probability output by the graph neural network with the preset defect association rule library, match the power industry standard logic, and generate operation and maintenance suggestions and comprehensive risk assessment results.

[0017] In step S1, first, a dual-payload drone is used to carry an X-ray sensor and a visible light camera to synchronously collect dual-modal data of the strain clamp, and an environmental sensor synchronously records parameters such as light intensity and metal artifact level.

[0018] For X-ray images, wavelet transform is used to suppress metal artifacts, and adaptive histogram equalization is combined to enhance contrast; for visible light images, illumination equalization is performed through the Retinex theory, and bilateral filtering is used to eliminate motion blur. This step aims to improve the image quality and provide clear input data for subsequent processing.

[0019] Furthermore, dynamic feature fusion is carried out. The preprocessed X-ray and visible light images are respectively input into lightweight networks. The X-ray branch uses the improved MobileNetV3 to extract internal defect features, and the visible light branch uses GhostNet to extract surface defect features.

[0020] Subsequently, the features and environmental parameters are fused through a dynamic channel attention module: global average pooling is used to extract feature statistics, a multi-layer perceptron generates dynamic channel weights, and the dual-modal features are weighted and fused according to the weights.

[0021] Among them, the dynamic channel weights are allocated as follows:

[0022] W X光 = Sigmoid(f MLP (Concat(GAP(F X光 ), GAP(F 可见光 ), E env ))) (1)

[0023] Weighted fusion is performed to suppress low-quality modal noise.

[0024] F fusion = W X光 ·F X光 + (1 - W X光 )·F 可见光 (2)

[0025] The modal weights are adjusted in real time to suppress low-quality modal noise.

[0026] Among them, W X光 represents the dynamic channel weight of the X-ray modality, GAP represents the global average pooling operation; E env represents the environmental parameters, including light intensity, humidity, and metal artifact level; F X光 represents the original features of the X-ray modality, including the feature information of the X-ray image; F 可见光 represents the original features of the visible light modality, including the feature information of the visible light image; Concat represents the concatenation operation, which concatenates GAP(F X光 ), GAP(F可见光 ) and E env Concatenate according to specific dimensions and fuse multi-source information; f MLP Indicates that the multi-layer perceptron performs a non-linear transformation operation on the concatenated features to mine the complex associations between features; Sigmoid represents the activation function that maps the output of the MLP to the interval [0, 1], and finally generates the dynamic channel weight W of the X-ray modality X光 ; F fusion Represents the fused bimodal features after weighted fusion, integrating the effective information of the X-ray and visible light modalities and suppressing the noise of the low-quality modality

[0027] In step S4, gradient fusion training optimization is implemented. First, define the multi-modal loss function, Dice Loss (crack segmentation) + L1 Loss (crimping gap regression); visible light branch: Focal Loss (rust classification) + Smooth L1 Loss (bolt displacement regression), and introduce learnable parameters α and β to dynamically adjust the branch weights

[0028] Then, perform gradient fusion training optimization. Based on the validation set loss, backpropagate to adjust the modal weight parameters α (X-ray) and β (visible light) to prevent overfitting and modal bias, and then update the dynamic weights:

[0029]

[0030] When the gradient magnitude limit the update step size to ensure training stability

[0031] This optimization strategy effectively reduces cross-modal gradient conflicts and improves the generalization ability of the model

[0032] Among them, L X光 Represents the loss function of the X-ray modality, which is composed of Dice Loss and L1 Loss;

[0033] Dice Loss is used to optimize the overlap degree of the defect area; L1 Loss is used to constrain the stability of coordinate regression; L 可见光 Represents the loss function of the visible light modality, which is composed of Focal Loss and Smooth L1 Loss; Focal Loss is used to dynamically adjust the sample weights to alleviate the class imbalance problem; Smooth L1 Loss is used to optimize the regression accuracy of the bolt position in the visible light image and maintain the robustness of the positioning result after motion blur correction; α (t+1) Represents the weight parameter after the (t + 1)-th iteration update, which is calculated through backpropagation; α (t)denotes the weight parameter at the t-th iteration, which is the original parameter value before update; η represents the learning rate, which is used to control the step size of weight update and determines the amplitude of each parameter update; L val denotes the validation set loss function; denotes the validation set loss function L val the gradient of the weight α, which reflects the direction of the loss change with respect to the weight and guides the weight update direction; denotes the loss function the magnitude of the gradient, which is used to determine whether the gradient is too large.

[0034] In step S5, cross-modal semantic modeling is completed. First, node definitions are completed. The crack position (x, y), length, and depth are defined as X-ray nodes; the rust area, color gradient, and bolt displacement are defined as visible light nodes.

[0035] Based on the 3D CAD model of the strain clamp, structural dependency edges (such as the crack connection in the crimping area corresponding to the bolt node) and spatial proximity edges (Euclidean distance < 10 pixels) are defined. Message passing and feature aggregation are performed through a four-head graph attention network (GAT) to establish a causal chain of "rust → stress concentration → crack propagation". Among them, a multi-head graph attention network (4-head GAT) is used for message passing.

[0036]

[0037] Among them, h i (l+1) denotes the updated feature representation of node i at the (l + 1)-th layer in the graph neural network; σ represents the activation function; α ij h denotes the attention weight of node i to neighbor node j in the h-th attention head; W h denotes the weight matrix of the h-th attention head; h j (l) denotes the feature representation of node j at the l-th layer; N(i) represents the neighbor set of node i, which includes nodes connected by structural dependency edges (such as bolt nodes connected to cracks in the crimping area) or spatial proximity edges (Euclidean distance < 10 pixels).

[0038] Finally, in terms of risk assessment and decision-making, a defect association rule library is defined according to the DL / T 664-2016 standard: high risk (crack > 2mm and rust > 10cm 2 , replace in 72 hours), medium risk (crack exists or rust > 20cm 2 , monthly re-inspection), low risk (only a small amount of rust, quarterly observation). Combining the defect association probability output by the graph neural network, standardized operation and maintenance suggestions are generated.

[0039] The beneficial effects of the present invention are as follows: A method for detecting defects in strain clamps based on dynamic fusion of X-ray and visible light provided by the present invention designs a dynamic channel attention module (DECA) for environmental perception. Based on global average pooling, it extracts feature statistics, fuses multi-source environmental parameters such as light intensity and metal artifact level, generates channel weights through a multi-layer perceptron (MLP), and realizes adaptive weighted fusion of cross-modal features. It proposes a gradient fusion optimization algorithm, introduces learnable modal weight parameters, combines the loss of the validation set to backpropagate and dynamically adjusts the dual-modal training path, and implements gradient truncation when the gradient amplitude exceeds the threshold to suppress gradient conflicts between modalities. It constructs a cross-modal causal association graph neural network, defines structural dependence edges and spatial proximity edges based on the 3D CAD model of the strain clamp, and models the physical causal chain of "crimping area crack → bolt corrosion → stress concentration → crack propagation" through the graph attention mechanism, solving the misjudgment problem caused by isolated defect detection in traditional methods. It binds the defect association rule base with the industry standard logic, matches the preset risk level through the defect probability output by the GNN, generates interpretable operation and maintenance suggestions, greatly improves the decision traceability, and supports power personnel to quickly formulate maintenance strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a flowchart of an embodiment of the present invention;

[0041] Figure 2 is a flowchart of the dynamic channel attention module and the gradient fusion function in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The embodiments of the present invention are introduced below with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0043] This embodiment proposes a live detection device and method for strain clamps of transmission lines based on dual unmanned aerial vehicles. Aiming at the problems of single-modal detection limitations, static fusion weight curing, and lack of semantic association between internal and external defects in traditional methods, the present invention makes innovations in the following aspects: By combining the dynamic channel attention mechanism (DECA) with environmental parameters and feature quality, it real-time allocates X-ray and visible light channel weights, and uses the gradient fusion optimization method to achieve dynamic fusion; It constructs a graph neural network (GNN), maps the internal cracks detected by X-ray and the surface rust identified by visible light into graph nodes, defines edge connection rules based on the physical structure of the strain clamp, and uses the graph attention mechanism (GAT) to realize causal reasoning of internal and external defects, which can effectively improve the detection accuracy and environmental robustness, and promote accurate operation and maintenance decisions.

[0044] By establishing a three - level innovation architecture of dynamic fusion - semantic modeling - closed - loop decision - making, the detection of strain clamps has achieved a leap from "single - modality analysis" to "multi - modality causal reasoning". This can significantly improve the detection accuracy, enhance the environmental robustness, and make the operation and maintenance decision - making more accurate.

[0045] The technical solution of this application will be described in detail below in conjunction with the accompanying drawings.

[0046] As Figure 1 shown, the method of this embodiment includes the following steps:

[0047] S1. Dual - modality data acquisition: By using a dual - payload unmanned aerial vehicle (UAV) equipped with an X - ray sensor and a visible - light camera, synchronously acquire the dual - modality image data of the strain clamp, and respectively generate an X - ray feature map and a visible - light feature map.

[0048] S2. Perform metal artifact suppression and contrast enhancement processing on the X - ray image to extract internal defect features; perform illumination equalization and motion blur correction processing on the visible - light image to extract surface defect features; output a new dual - modality feature map.

[0049] S3. Input the dual - modality feature map and environmental parameters into the dynamic channel attention module, dynamically allocate channel weights according to the light intensity and metal artifact level, and generate a preliminary fusion feature.

[0050] S4. Use the gradient fusion algorithm to balance the loss function weights of the dual - modality branches, and optimize network training in combination with the gradient truncation technique.

[0051] S5. Construct a cross - modality graph neural network, map internal defects and surface anomalies to graph nodes, define physical association edges based on the 3D CAD model of the strain clamp, and model the defect causal chain through the graph attention mechanism.

[0052] S6. Combine the defect association probability output by the graph neural network with the preset defect association rule base, match the power industry standard logic, and generate operation and maintenance suggestions and comprehensive risk assessment results.

[0053] In step S1 of this embodiment, after completing data acquisition and pre - processing, first use a dual - payload UAV equipped with an X - ray sensor and a visible - light camera to synchronously acquire the dual - modality data of the strain clamp, and the environmental sensor synchronously records parameters such as light intensity and metal artifact level.

[0054] In step S2 of this embodiment, for X-ray images, wavelet transform is used to suppress metal artifacts, and adaptive histogram equalization is combined to enhance the contrast; for visible light images, illumination equalization is performed through the Retinex theory, and bilateral filtering is used to eliminate motion blur. This step aims to improve the image quality and provide clear input data for subsequent processing. The preprocessed X-ray and visible light images are respectively input into the lightweight network, where the improved MobileNetV3 is used for the X-ray branch and GhostNet is used for the visible light branch to extract features.

[0055] In step S3 of this embodiment, features and environmental parameters are fused through the Dynamic Channel Attention Module (DECA): global average pooling is used to extract feature statistics, and a Multi-Layer Perceptron (MLP) is used to generate dynamic channel weights:

[0056] W X光 = Sigmoid(f MLP (Concat(GAP(F X光 ), GAP(F 可见光 ), E env ))) (1)

[0057] The bimodal features are weighted and fused according to the weights to suppress low-quality modal noise.

[0058] F fusion = W X光 ·F X光 +(1 - W X光 )·F 可见光 (2)

[0059] Among them, W X光 represents the dynamic channel weight of the X-ray modality, GAP represents the global average pooling operation; E env represents the environmental parameters, including illumination intensity, humidity, and metal artifact level; F X光 represents the original features of the X-ray modality, including the feature information of the X-ray image; F 可见光 represents the original features of the visible light modality, including the feature information of the visible light image; Concat represents the concatenation operation, which concatenates GAP(F X光 ), GAP(F 可见光 ), and E env along a specific dimension to fuse multi-source information; f MLP represents the non-linear transformation operation of the MLP on the concatenated features, which is used to explore the complex correlations between features; Sigmoid represents the activation function, which maps the output of the MLP to the interval [0, 1], and finally generates the dynamic channel weight W X光 ; F fusionRepresents the bimodal features after weighted fusion, integrating the effective information of the X-ray and visible light modalities and suppressing the noise of the low-quality modality.

[0060] In step S4 of this embodiment, gradient fusion training optimization is performed. Based on the validation set loss, backpropagation is used to adjust the modal weight parameters α (X-ray) and β (visible light) to prevent overfitting and modal bias, and then the dynamic weights are updated. When the gradient amplitude is reached, the update step size is restricted to ensure training stability. This optimization strategy effectively reduces cross-modal gradient conflicts, improves the generalization ability of the model, and reduces the standard deviation of the detection accuracy from 3.2% to 1.5%.

[0061] Among them, the gradient fusion method includes the following steps: Define the learnable parameters α (X-ray branch weight) and β (visible light branch weight), with the initial values α = 0.5 and β = 0.5; Forward calculate the bimodal loss: LX-ray (Dice Loss + L1Loss), Lvisible light (Focal Loss + Smooth L1 Loss); Update the weights through backpropagation of the validation set loss, and the formula is:

[0062]

[0063] When the gradient amplitude is reached, truncation is performed to limit the risk of modal dominance.

[0064] Among them, L X光 represents the loss function of the X-ray modality, which is composed of Dice Loss and L1 Loss;

[0065] Dice Loss is used to optimize the overlap degree of the defect area; L1 Loss is used to constrain the stability of coordinate regression; L 可见光 represents the loss function of the visible light modality, which is composed of Focal Loss and Smooth L1 Loss; Focal Loss is used to dynamically adjust the sample weights to alleviate the class imbalance problem; Smooth L1 Loss is used to optimize the regression accuracy of the bolt position in the visible light image and maintain the robustness of the positioning result after motion blur correction; α (t+1) represents the weight parameter after the (t + 1)-th iteration update, which is obtained through backpropagation calculation; α (t) represents the weight parameter at the t-th iteration, which is the original parameter value before update; η represents the learning rate, which is used to control the update step size of the weights and determine the amplitude of each parameter update; L val represents the validation set loss function; represents the validation set loss function L val The gradient of the weight α, which reflects the direction of the loss change with respect to the weight and guides the weight update direction; represents the loss function The magnitude of the gradient, which is used to determine whether the gradient is too large.

[0066] In step S5 of this embodiment, the crack position (x, y), length, and depth are defined as X-ray nodes; the rust area, color gradient, and bolt displacement are defined as visible light nodes; the edge connection rules are defined: Structural dependency edges: Based on the stress distribution of the strain clamp 3D CAD model, connect the cracks in the crimping area to the corresponding bolt nodes; Spatial proximity edges: Automatically connect nodes with an Euclidean distance < 10 pixels.

[0067] In step S5 of this embodiment, a graph neural network (GNN) is designed according to the above node definitions and edge connection rules.

[0068] A multi-head graph attention network (4-head GAT) is used for message passing and feature aggregation to establish a causal chain of "rust → stress concentration → crack propagation".

[0069]

[0070] where h i (l+1) represents the updated feature representation of node i in the (l + 1)-th layer in the graph neural network; σ represents the activation function; α ij h represents the attention weight of node i to neighbor node j in the h-th attention head; W h represents the weight matrix of the h-th attention head; h j (l) represents the feature representation of node j in the l-th layer; N(i) represents the neighbor set of node i, including nodes connected by structural dependency edges (such as bolt nodes connected to cracks in the crimping area) or spatial proximity edges (Euclidean distance < 10 pixels).

[0071] In step S6 of this embodiment, a defect association rule library is defined according to the DL / T 664-2016 standard: High risk (crack > 2 mm and rust > 10 cm 2 , replace in 72 hours), Medium risk (crack exists or rust > 20 cm 2 , monthly re-inspection), Low risk (only a small amount of rust, quarterly observation). Combining the defect association probability output by the graph neural network, standardized operation and maintenance suggestions are generated.

[0072] In summary, in this example, the drone is used to collect X-ray and visible light images respectively, and the environmental parameters are read at the same time. Through the three-level innovation of dynamic multi-modal fusion - cross-modal semantic modeling - power scenario optimization, the pain points of modal conflict, environmental sensitivity, and semantic fragmentation in traditional detection methods are solved, and the intelligent and accurate diagnosis of tension clamp defects is realized. Its technical solution has leading detection accuracy and applicability in the industry, providing reliable technical support for the safe operation and maintenance of transmission lines.

[0073] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A method for detecting defects of strain clamps based on dynamic fusion of X-ray and visible light, characterized in that, It includes the following steps: S1. Dual-modal data acquisition: By using a dual-payload drone equipped with an X-ray sensor and a visible light camera, synchronously acquire the dual-modal image data of the strain clamp, and respectively generate an X-ray feature map and a visible light feature map. S2. Perform metal artifact suppression and contrast enhancement processing on the X-ray image, and extract internal defect features. Perform illumination equalization and motion blur correction processing on the visible light image, and extract surface defect features. Output a new dual-modal feature map. S3. Input the dual-modal feature map and environmental parameters into the dynamic channel attention module, dynamically allocate channel weights according to the illumination intensity and metal artifact level, and generate preliminary fusion features. S4. Adopt a gradient fusion algorithm to balance the loss function weights of the dual-modal branches, and optimize network training in combination with the gradient truncation technique. S5. Construct a cross-modal graph neural network, map internal defects and surface anomalies to graph nodes, define physical association edges based on the 3D CAD model of the strain clamp, and model the defect causal chain through the graph attention mechanism. S6. Combine the defect association probability output by the graph neural network with the preset defect association rule library, match the standard logic of the power industry, and generate operation and maintenance suggestions and comprehensive risk assessment results.

2. The method for detecting defects of strain clamps based on X-ray-visible light dynamic fusion according to claim 1, wherein The X-ray image processing uses a lightweight convolutional network, which includes a multi-scale feature pyramid module; the visible light image processing uses a channel separation convolutional network; the number of output feature map channels is 256.

3. The method for detecting the defects of strain clamps based on X-ray-visible light dynamic fusion according to claim 1, wherein The dynamic channel attention module fuses the X-ray features, visible light features and environmental parameter vectors through a multi-layer perceptron to generate dynamic channel weights.

4. The method for detecting the defect of the strain clamp based on X-ray-visible light dynamic fusion according to claim 1, characterized in that, The node definition of the cross-modal graph neural network includes: the position and length attributes of the X-ray crack node, and the area and color gradient attributes of the visible light rust node.

5. The method for detecting the defect of the strain clamp based on X-ray-visible light dynamic fusion according to claim 1, characterized in that, The determination logic of the defect association rule library includes: High risk: crack length > 2mm and rust area at the corresponding position > 10cm 2 ; Medium risk: crack exists or rust area > 20cm 2 ; Low risk: only surface rust and area < 5cm 2 ; Disposal suggestions: It is recommended to replace in 72 hours for high-risk situations; monthly re-inspection for medium-risk situations; quarterly observation for low-risk situations.

6. The method for detecting the defects of strain clamps based on X-ray-visible light dynamic fusion according to claim 1, characterized in that, The gradient fusion method includes the following steps: Define the learnable parameter α as the weight of the X-ray branch and β as the weight of the visible light branch, with the initial values α = 0.5 and β = 0.

5. Forward calculation of the bimodal loss: L X光 : Dice Loss + L1 Loss, L 可见光 : FocalLoss + Smooth L1 Loss; Update the weights by backpropagation of the validation set loss, and the formula is: When the gradient magnitude is truncated to limit the risk of modal dominance. Among them, L X光 represents the loss function of the X-ray modality, which is composed of Dice Loss and L1 Loss; Dice Loss is used to optimize the overlap degree of the defect area; L1 Loss is used to constrain the stability of coordinate regression; L 可见光 represents the loss function of the visible light modality, which is composed of FocalLoss and Smooth L1 Loss; Focal Loss is used to dynamically adjust the sample weights to alleviate the class imbalance problem; Smooth L1 Loss is used to optimize the regression accuracy of the bolt position in the visible light image and maintain the robustness of the positioning result after motion blur correction; α (t+1) represents the weight parameter after the (t + 1)-th iteration update, which is calculated through backpropagation; α (t) represents the weight parameter at the t-th iteration, which is the original parameter value before update; η represents the learning rate, which is used to control the step size of weight update and determine the amplitude of each parameter update; L val represents the validation set loss function; represents the validation set loss function L val the gradient of the weight α, which reflects the direction of the loss change with respect to the weight and guides the weight update direction; represents the loss function the magnitude of the gradient, which is used to determine whether the gradient is too large.

7. The method for detecting the defect of the strain clamp based on the dynamic fusion of X-ray and visible light according to claim 1, wherein In terms of risk assessment and decision-making, a defect association rule base is defined according to the DL / T 664-2016 standard: High risk: crack > 2 mm and rust > 10 cm 2 It is recommended to replace it within 72 hours. Medium risk: crack exists or rust > 20 cm 2 It is recommended to conduct monthly re-inspection. Low risk: only a small amount of rust, and it is recommended to observe quarterly; combined with the defect association probability output by the graph neural network, standardized operation and maintenance suggestions are generated.

Citation Information

Cited By

  • Automobile chassis sealant detection method and system based on multi-scale feature fusion

    CN121147218A

  • Automobile chassis sealant detection method and system based on multi-scale feature fusion

    CN121147218B

  • Nondestructive testing method for strain clamp of power transmission line based on digital ray technology

    CN121188718A