Cotter pin defect data enhancement method and system
Through the multi-source sensor synchronous acquisition and intensity conditional diffusion model, the multi-modal data alignment and generation problems in the detection of open pin defects of transmission line are solved, which improves the robustness and generalization capabilities of the detection model and reduces the calculation cost.
Patent Information
- Application Number
- CN202510546400.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the defect detection of the power transmission line opening pin faces problems such as insufficient spatial and temporal alignment accuracy of multimodal data, deviations from defect state generation and real mechanical response, overfitting of the detection model and high calculation costs.
Multimodal data is synchronously collected through multi-source sensors and hardware-triggered spatiotemporal alignment. Combined with the fusion characteristics of the gated attention mechanism, an intensity condition diffusion model is constructed, and continuous defect state data is optimized by using lightweight finite element model.
High-precision alignment and feature fusion of multimodal data are realized, defective data that conforms to the mechanical laws of materials are generated, which improves the robustness and generalization capabilities of the detection model and reduces calculation costs.
Smart Images

Figure CN120470474A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and industrial inspection technology, and more particularly to a method and system for enhancing cotter pin defect data. In particular, the present invention relates to a cotter pin defect data enhancement method that combines multimodal data fusion with a strength-controllable generation network. Background Art
[0002] Currently, there are three major technical bottlenecks in the detection of cotter pin defects in power transmission lines:
[0003] First, multimodal data (visible light, infrared, point cloud) suffers from insufficient spatiotemporal alignment accuracy due to sensor heterogeneity and acquisition asynchrony. Traditional feature fusion methods struggle to effectively extract cross-modal correlation features, limiting the robustness of defect representation.
[0004] Second, existing data augmentation methods mostly rely on geometric transformation or style transfer, which cannot generate continuous defect states driven by physical laws (such as corrosion degree and fracture stage), resulting in deviations between the generated data and the actual mechanical response.
[0005] Third, the detection model trained based on limited real defect samples is prone to overfitting, especially the generalization performance of rare defect types is significantly reduced.
[0006] In addition, although traditional finite element simulation can simulate the mechanical behavior of defects, its high computational cost and low generation efficiency make it difficult to meet the large-scale data requirements of deep learning.
[0007] Patent application document CN111815623A discloses a method for identifying missing cotter pins on power transmission lines. The above-mentioned method for identifying missing cotter pins on power transmission lines includes obtaining inspection pictures of the transmission line and generating a sample set. The sample set is trained to obtain a component detection model and a cotter pin missing detection model respectively. The inspection pictures to be detected are input into the component detection model to generate a plurality of component inspection frames. The target components in the plurality of component inspection frames are hierarchically divided to determine the component detection frames. Each component detection frame is detected using the cotter pin missing detection model to obtain the position information of the target component, the category information of the target component and the confidence level of the target component. However, this patent cannot completely solve the existing technical problems, nor can it meet the requirements of the present invention. Summary of the Invention
[0008] In view of the defects in the prior art, the purpose of the present invention is to provide a method and system for enhancing cotter pin defect data.
[0009] The method for enhancing cotter pin defect data provided by the present invention comprises:
[0010] Step S1: Synchronously collect multimodal data of the cotter pin through multi-source sensors, including visible light images, infrared image temperature distribution data, and lidar point cloud data, and achieve spatiotemporal alignment based on hardware triggering;
[0011] Step S2: Preprocess the multimodal data, including noise removal, resolution alignment, and feature extraction, and use the gated attention mechanism to dynamically fuse RGB texture, temperature gradient, and point cloud geometric features to generate multimodal fusion features;
[0012] Step S3: Construct a strength-conditional diffusion model, control the defect generation intensity through latent space parameterization, and generate continuous defect state data from slight corrosion to complete fracture in combination with strength constraint parameters;
[0013] Step S4: Input the generated defect state data into the lightweight finite element model for reverse optimization, verify the physical law consistency of the generated data, and output the multimodal defect data set for detection model training.
[0014] Preferably, the step S1 includes:
[0015] The iterative closest point algorithm is used to optimize the rigid body transformation matrix and align the lidar point cloud data with the visible light image at the sub-pixel level.
[0016] Perform bicubic interpolation on the infrared image to make its resolution consistent with that of the visible light image;
[0017] Multi-sensor hardware triggering is achieved through GPS clock synchronization to ensure acquisition time synchronization.
[0018] Preferably, step S2 includes:
[0019] Use SIFT or ORB algorithm to extract image feature key points and descriptors F img ∈R C×H×W , C is the number of channels, H is the height, and W is the width;
[0020] Calculate the two-dimensional temperature gradient field of the infrared image and quantify the abnormal heat conduction area, where x, y are the image coordinates and Δ is the pixel spacing, and obtain the feature F thermal ∈R T×H×W , T is temperature;
[0021]
[0022] Image noise is removed using median filtering or non-local mean filtering, and pixel values are scaled to the range of [0, 1]. Infrared images are subjected to histogram equalization to enhance temperature contrast and then normalized. Point clouds are de-redundant by downsampling the voxel grid or removing statistical outliers. The point cloud coordinates are centered and scaled, and the cleaned and standardized data is output for subsequent feature extraction, resulting in standardized RGB images, infrared images, and point clouds.
[0023] For any point p in the point cloud i , select the k nearest neighbor point set N(p i ), calculate the covariance matrix based on the neighborhood point set, and use principal component analysis to perform eigendecomposition on the covariance matrix to define the surface curvature feature map C(p i )=λ min / (λ1+λ2+λ3), where λ1≥λ2≥λ3 is the eigenvalue, λ min is the minimum eigenvalue, and the point cloud feature F is obtained pc ∈R D×N ;
[0024] Design a gated attention mechanism to generate attention weights α∈[0,1] by convolution of the three features H×W : Where σ1 is the sigmoid function, Point cloud features projected into image space; Dynamically adjust the contribution of each modality: Here, MLP represents a three-layer perceptron with an input dimension of 64 and an output dimension of 256, which upgrades the point cloud features to the same dimension as the image features.
[0025] Preferably, step S3 includes:
[0026] Define the defect intensity s∈[0,1], use two layers of MLP to project the intensity parameter s into the latent space, and generate the conditional vector c s ∈R 64 , c s =MLP(s)=W2·ReLU(W1·s+b1)+b2, W1 and W2 are weight matrices, b1, b2 are bias terms, and the conditional vector c s To control the defect generation intensity, as the input of the diffusion Unet encoder;
[0027] Insert the conditional normalization layer CLN into each residual block of the UNet encoder to dynamically adjust the feature distribution through the conditional vector: where γ and β are given by c s The predicted scaling and bias parameters, represents the kth layer of the decoder; δ is a constant used to ensure the stability of normalization;
[0028] The encoder output feature h enc and fusion feature F fusion Calculate cross attention: Attention(Q=h enc ,K=F fusion ,V=F fusion ), output the fused feature h fused As the initial input to the decoder;
[0029] In the DDPM inverse denoising process, the perturbation variance σ2=β is adjusted by the intensity parameter s. t ·s,β t The noise scheduling parameters are preset to cover a continuous range from slight corrosion to complete fracture.
[0030] Preferably, step S4 includes:
[0031] The low-dimensional parameter S(θ)∈R that encodes the stress field of the defect simulated by the finite element method 32 and fusion feature F fusion Combined to get F diffusion :F diffusion =Concat(F fused ,S(θ))∈R 256+32 and RGB image I rgb ∈R H×W×3 As input to the strength-controllable generative network;
[0032] Gaussian noise image sequence generated by forward diffusion The training loss is as follows:
[0033]
[0034] L phy =||φ gen -φ FEM ||
[0035] The training process minimizes the loss, where ε follows a normal distribution and ε θ is the predicted noise, φ gen and φ FEM They represent the generated images, reverse deduction of stress parameters and real simulation results through lightweight finite element model; For time step t, real image I real and the expectation of noise ε; T is the maximum time step in the diffusion process;
[0036] The total loss is:
[0037] L total =L DDPM +0.5·L phy
[0038] The trained diffusion model is used to generate multimodal data of cotter pin defects and a visible light cotter pin defect image is generated.
[0039] The cotter pin defect data enhancement system provided by the present invention includes:
[0040] Module M1: Uses multi-source sensors to synchronously collect multimodal data of the cotter pin, including visible light images, infrared image temperature distribution data, and lidar point cloud data, and achieves spatiotemporal alignment based on hardware triggering;
[0041] Module M2: Preprocesses multimodal data, including noise removal, resolution alignment, and feature extraction. It also uses a gated attention mechanism to dynamically fuse RGB texture, temperature gradient, and point cloud geometric features to generate multimodal fusion features.
[0042] Module M3: Constructs a strength-conditional diffusion model, controls the defect generation intensity through implicit space parameterization, and generates continuous defect state data from slight corrosion to complete fracture in combination with finite element simulation physical constraints;
[0043] Module M4: Input the generated defect state data into the lightweight finite element model for reverse optimization, verify the physical law consistency of the generated data, and output a multimodal defect data set for detection model training.
[0044] Preferably, the module M1 includes:
[0045] The iterative closest point algorithm is used to optimize the rigid body transformation matrix and align the lidar point cloud data with the visible light image at the sub-pixel level.
[0046] Perform bicubic interpolation on the infrared image to make its resolution consistent with that of the visible light image;
[0047] Multi-sensor hardware triggering is achieved through GPS clock synchronization to ensure acquisition time synchronization.
[0048] Preferably, the module M2 includes:
[0049] Use SIFT or ORB algorithm to extract image feature key points and descriptors F img ∈R C×H×W , C is the number of channels, H is the height, and W is the width;
[0050] Calculate the two-dimensional temperature gradient field of the infrared image and quantify the abnormal heat conduction area, where x, y are the image coordinates and Δ is the pixel spacing, and obtain the feature F thermal ∈R T×H×W , T is temperature;
[0051]
[0052] Image noise is removed using median filtering or non-local mean filtering, and pixel values are scaled to the range of [0, 1]. Infrared images are subjected to histogram equalization to enhance temperature contrast and then normalized. Point clouds are de-redundant by downsampling the voxel grid or removing statistical outliers. The point cloud coordinates are centered and scaled, and the cleaned and standardized data is output for subsequent feature extraction, resulting in standardized RGB images, infrared images, and point clouds.
[0053] For any point p in the point cloud i , select the k nearest neighbor point set N(p i ), calculate the covariance matrix based on the neighborhood point set, and use principal component analysis to perform eigendecomposition on the covariance matrix to define the surface curvature feature map C(p i )=λ min / (λ1+λ2+λ3), where λ1≥λ2≥λ3 is the eigenvalue, λ min is the minimum eigenvalue, and the point cloud feature F is obtained pc ∈R D×N ;
[0054] Design a gated attention mechanism to generate attention weights α∈[0,1] by convolution of the three features H×W : Where σ1 is the sigmoid function, Point cloud features projected into image space; Dynamically adjust the contribution of each modality: Here, MLP represents a three-layer perceptron with an input dimension of 64 and an output dimension of 256, which upgrades the point cloud features to the same dimension as the image features.
[0055] Preferably, the module M3 includes:
[0056] Define the defect intensity s∈[0,1], use two layers of MLP to project the intensity parameter s into the latent space, and generate the conditional vector c s ∈R 64 , c s =MLP(s)=W2·ReLU(W1·s+b1)+b2, W1 and W2 are weight matrices, b1, b2 are bias terms, and the conditional vector c s To control the defect generation intensity, as the input of the diffusion Unet encoder;
[0057] Insert the conditional normalization layer CLN into each residual block of the UNet encoder to dynamically adjust the feature distribution through the conditional vector: where γ and β are given by c s The predicted scaling and bias parameters, represents the kth layer of the decoder; δ is a constant used to ensure the stability of normalization;
[0058] The encoder output feature h enc and fusion feature F fusion Calculate cross attention: Attention(Q=h enc ,K=F fusion ,V=F fusion ), output the fused feature h fused As the initial input to the decoder;
[0059] In the DDPM inverse denoising process, the perturbation variance σ2=β is adjusted by the intensity parameter s. t ·s,β t The noise scheduling parameters are preset to cover a continuous range from slight corrosion to complete fracture.
[0060] Preferably, the module M4 includes:
[0061] The low-dimensional parameter S(θ)∈R that encodes the stress field of the defect simulated by the finite element method 32 and fusion feature F fusion Combined to get F diffusion :F diffusion =Concat(F fused ,S(θ))∈R 256+32 and RGB image I rgb ∈R H×W×3 As input to the strength-controllable generative network;
[0062] Gaussian noise image sequence generated by forward diffusion The training loss is as follows:
[0063]
[0064] L phy =||φ gen -φ FEM ||
[0065] The training process minimizes the loss, where ε follows a normal distribution and ε θ is the predicted noise, φ gen and φ FEM They represent the generated images, reverse deduction of stress parameters and real simulation results through lightweight finite element model; For time step t, real image I real and the expectation of noise ε; T is the maximum time step in the diffusion process;
[0066] The total loss is:
[0067] L total =L DDPM +0.5·Lphy
[0068] The trained diffusion model is used to generate multimodal data of cotter pin defects and a visible light cotter pin defect image is generated.
[0069] Compared with the prior art, the present invention has the following beneficial effects:
[0070] (1) Through hardware synchronous acquisition and iterative optimization of the rigid body transformation matrix, pixel-level multimodal alignment is achieved, and the gated attention mechanism is combined to dynamically fuse RGB texture, temperature gradient and point cloud geometric features to improve the discriminability of defect representation;
[0071] (2) Construct a strength-condition diffusion model, control the defect generation process through implicit space parameterization, and combine strength constraint parameters to generate continuous state data from slight corrosion to complete fracture, ensuring that the generated results conform to the laws of material mechanics;
[0072] (3) The physical simulation stress field parameters are embedded in the generative network, and the authenticity of the generated data is enhanced through reverse optimization of the lightweight finite element model. A multimodal defect dataset is constructed, which significantly enhances the robustness and generalization of the detection model under complex working conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0074] Figure 1 This paper presents a general scheme for cotter pin defect data enhancement method combining multimodal data fusion and strength-controllable generative network.
[0075] Figure 2 Input and output flow chart for the strength controllable network;
[0076] Figure 3 Input and output flow chart for the strength controllable network;
[0077] Figure 4 The generated visible light image. DETAILED DESCRIPTION
[0078] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0079] Example
[0080] The present invention provides a method for enhancing cotter pin defect data by combining multimodal data fusion and strength-controllable generation network. The specific implementation process is as follows:
[0081] Multimodal data fusion includes:
[0082] Step 1: Synchronously collect data through multiple source sensors, including visible light camera (RGB), infrared thermal imager (temperature distribution), and lidar (point cloud XYZ coordinates + reflection intensity). Use hardware triggering (such as GPS clock synchronization) to ensure multi-sensor time synchronization. Use the iterative closest point (ICP) algorithm to optimize the rigid body transformation matrix T and minimize the reprojection error (<0.5 pixels) to align the point cloud with the image. Then perform bicubic interpolation on the infrared image to make its resolution consistent with the visible light image, and output it as aligned multimodal data (RGB image I rgb ∈R H×W×3 , infrared image I thermal ∈R H×W , point cloud P lidar ∈R N×4 ).
[0083] Step 2: Use median filtering or non-local mean filtering to remove image noise and scale the pixel values to the range of [0,1]. The formula is I norm = (I - μ) / σ (μ is the mean, σ is the standard deviation). The infrared image is normalized after temperature contrast enhancement using histogram equalization. The point cloud is de-reduce by voxel grid downsampling or statistical outlier removal (SOR). The point cloud coordinates are centered (mean subtracted) and scaled, outputting cleaned, standardized data suitable for subsequent feature extraction. Finally, the resulting data is a standardized RGB image, infrared image, and point cloud.
[0084] Step 3: Use SIFT or ORB algorithm to extract image feature key points and descriptors F img ∈R C×H×W .
[0085] Step 4: Calculate the two-dimensional temperature gradient field of the infrared image and quantify the abnormal heat conduction area, where x, y are image coordinates and Δ is the pixel spacing (unit: mm / px). The feature F is obtained. thermal ∈R T×H×W .
[0086]
[0087] Step 5: For any point p in the point cloud i , select the k nearest neighbor point set N(p i ), calculate the covariance matrix based on the neighborhood point set, and use principal component analysis to perform eigendecomposition on the covariance matrix to define the surface curvature feature map C(pi )=λ min / (λ1+λ2+λ3), where λ1≥λ2≥λ3 is the eigenvalue, λ min is the minimum eigenvalue, and the point cloud feature F is obtained pc ∈R D ×N .
[0088] Step 6: Design a gated attention mechanism to generate attention weights α∈[0,1] by convolution of the three features. H×W : Where σ1 is the sigmoid function, is the point cloud feature projected into the image space. Dynamically adjust the contribution of each modality: Here, MLP represents a three-layer perceptron with an input dimension of 64 and an output dimension of 256. The purpose is to upgrade the point cloud features to the same dimension as the image features.
[0089] The intensity-controllable generation network includes:
[0090] Step 1: Define the defect intensity s∈[0,1], use two-layer MLP to project the intensity parameter s into the latent space, and generate the conditional vector c s ∈R 64 :c s =MLP(s)=W2·ReLU(W1·s+b1)+b2W1∈R 64×64 is the weight matrix, b1, b2 are bias terms, and the conditional vector c s To control the defect generation intensity, it is used as the input of the diffusion Unet encoder.
[0091] Step 2: Insert a conditional normalization layer CLN into each residual block of the UNet encoder to dynamically adjust the feature distribution through the conditional vector: where γ and β are given by c s The predicted scaling and bias parameters, represents the k-th layer of the decoder.
[0092] Step 3: Encoder output feature h enc and fusion feature F fusion Calculate cross attention: Attention(Q=h enc ,K=F fusion ,V=F fusion ), output the fused feature h fused As the initial input to the decoder.
[0093] Step 4: In the DDPM inverse denoising process, the perturbation variance σ2 = β is adjusted by the strength parameter s. t ·s(β tis the preset noise scheduling parameter), covering a continuous state from slight corrosion to complete fracture.
[0094] The present invention provides a method for enhancing cotter pin defect data by combining multimodal data fusion and strength controllable generation network. Figure 1 As shown, the first step is the training process:
[0095] Step 1: The drone is equipped with a multimodal preprocessing module to collect the cotter pin defect data on the transmission line and output a 256-dimensional fusion feature F. fusion , including RGB texture, infrared temperature gradient and point cloud geometric features. The visible light camera parameters in the module are 1920x1080@30fps, the infrared thermal imager is 640x480@15fps, and the laser radar is 16 beams. The process is as follows Figure 2 shown.
[0096] Step 2: Initialize the intensity controllable generation network in the original diffusion model, such as Figure 3 As shown, the low-dimensional parameter S(θ)∈R that encodes the defect stress field of the finite element method simulation is 32 and fusion feature F fusion Combined to get F diffusion :F diffusion =Concat(F fused ,S(θ))∈R 256+32 and RGB image I rgb ∈R H×W×3 As input to the strength-controllable generative network;
[0097] Step 3: Generate a Gaussian noise image sequence through forward diffusion The training loss is as follows:
[0098]
[0099] L phy =||φ gen -φ FEM ||
[0100] The training process minimizes the loss, where ε follows a normal distribution and ε θ is the predicted noise, φ gen and φ FEM They represent the generated image, the reverse deduction of stress parameters and the real simulation results through the lightweight finite element model, and the total loss is L total =L DDPM +0.5·L phy .
[0101] Step 4: Use the trained diffusion model to generate multimodal data of cotter pin defects, including RGB images (256×256×3), infrared images (256×256×1) and point clouds (10^4 points). The generated visible light cotter pin defect image (256 size) is as follows: Figure 4 shown.
[0102] This method integrates visible light, infrared, and point cloud data (with subpixel alignment) with finite element stress field modeling to achieve cross-modal, simultaneous generation of defect morphology, thermodynamic distribution, and geometric deformation, overcoming the limitations of traditional single-modal generation techniques. Through intensity-parameterized modeling (MLP+CLN) and an improved DDPM algorithm, generation efficiency surpasses traditional simulations, meeting the requirements of industrial inspection model training.
[0103] Those skilled in the art will appreciate that, in addition to implementing the system, device, and various modules provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, and the like by logically programming the method steps. Therefore, the system, device, and various modules provided by the present invention can be considered a hardware component, and the modules included therein for implementing various programs can also be considered structures within the hardware component; the modules for implementing various functions can also be considered both software programs for implementing the method and structures within the hardware component.
[0104] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A method for enhancing cotter pin defect data, characterized in that: include: Step S1: Synchronously collect multimodal data of the cotter pin through multi-source sensors, including visible light images, infrared image temperature distribution data, and lidar point cloud data, and achieve spatiotemporal alignment based on hardware triggering; Step S2: Preprocess the multimodal data, including noise removal, resolution alignment, and feature extraction, and use the gated attention mechanism to dynamically fuse RGB texture, temperature gradient, and point cloud geometric features to generate multimodal fusion features; Step S3: Construct a strength-conditional diffusion model, control the defect generation intensity through latent space parameterization, and generate continuous defect state data from slight corrosion to complete fracture by combining the strength parameters; Step S4: Input the generated defect state data into the lightweight finite element model for reverse optimization, verify the physical law consistency of the generated data, and output the multimodal defect data set for detection model training.
2. The method for enhancing the defect data of a cotter pin according to claim 1, characterized in that: The step S1 comprises: The iterative closest point algorithm is used to optimize the rigid body transformation matrix and align the lidar point cloud data with the visible light image at the sub-pixel level. Perform bicubic interpolation on the infrared image to make its resolution consistent with that of the visible light image; Multi-sensor hardware triggering is achieved through GPS clock synchronization to ensure acquisition time synchronization.
3. The method for enhancing the defect data of a cotter pin according to claim 1, characterized in that: The step S2 comprises: Use SIFT or ORB algorithm to extract image feature key points and descriptors F img ∈R C×H×W , C is the number of channels, H is the height, and W is the width; Calculate the two-dimensional temperature gradient field of the infrared image and quantify the abnormal heat conduction area, where x, y are the image coordinates and Δ is the pixel spacing, and obtain the feature F thermal ∈R T×H×W , T is temperature; Image noise is removed using median filtering or non-local mean filtering, and pixel values are scaled to the range of [0, 1]. Infrared images are subjected to histogram equalization to enhance temperature contrast and then normalized. Point clouds are de-redundant by downsampling the voxel grid or removing statistical outliers. The point cloud coordinates are centered and scaled, and the cleaned and standardized data is output for subsequent feature extraction, resulting in standardized RGB images, infrared images, and point clouds. For any point p in the point cloud i , select the k nearest neighbor point set N(p i ), calculate the covariance matrix based on the neighborhood point set, and use principal component analysis to perform eigendecomposition on the covariance matrix to define the surface curvature feature map C(p i )=λ min / (λ1+λ2+λ3), where λ1≥λ2≥λ3 is the eigenvalue, λ min is the minimum eigenvalue, and the point cloud feature F is obtained pc ∈R D×N ; Design a gated attention mechanism to generate attention weights α∈[0,1] by convolution of the three features H×W : Where σ1 is the sigmoid function, Point cloud features projected into image space; Dynamically adjust the contribution of each modality: Here, MLP represents a three-layer perceptron with an input dimension of 64 and an output dimension of 256, which upgrades the point cloud features to the same dimension as the image features.
4. The method for enhancing the defect data of a cotter pin according to claim 3, characterized in that: The step S3 comprises: Define the defect intensity s∈[0,1], use two layers of MLP to project the intensity parameter s into the latent space, and generate the conditional vector c s ∈R 64 , c s =MLP(s)=W2·ReLU(W1·s+b1)+b2, W1 and W2 are weight matrices, b1, b2 are bias terms, and the conditional vector c s To control the defect generation intensity, as the input of the diffusion Unet encoder; Insert the conditional normalization layer CLN into each residual block of the UNet encoder to dynamically adjust the feature distribution through the conditional vector: where γ and β are given by c s The predicted scaling and bias parameters, represents the kth layer of the decoder; δ is a constant used to ensure the stability of normalization; The encoder output feature h enc and fusion feature F fusion Calculate cross attention: Attention(Q=h enc ,K=F fusion ,V=F fusion ), output the fused feature h fused As the initial input to the decoder; In the DDPM inverse denoising process, the perturbation variance σ2=β is adjusted by the intensity parameter s. t ·s,β t The noise scheduling parameters are preset to cover a continuous range from slight corrosion to complete fracture.
5. The method for enhancing the defect data of a cotter pin according to claim 4, characterized in that: The step S4 comprises: The low-dimensional parameter S(θ)∈R that encodes the stress field of the defect simulated by the finite element method 32 and fusion feature F fusion Combined to get F diffusion :F diffusion =Concat(F fused ,S(θ))∈R 256+32 and RGB image I rgb ∈R H×W×3 As input to a strength-controllable generative network; Gaussian noise image sequence generated by forward diffusion The training loss is as follows: L phy =||φ gen -f FEM || The training process minimizes the loss, where ε follows a normal distribution and ε θ is the predicted noise, φ gen and φ FEM They represent the generated images, reverse deduction of stress parameters and real simulation results through lightweight finite element model; For time step t, real image I real and the expectation of noise ε; T is the maximum time step in the diffusion process; The total loss is: L total =L DDPM +0.5 L phy The trained diffusion model is used to generate multimodal data of cotter pin defects and a visible light cotter pin defect image is generated.
6. A cotter pin defect data enhancement system, characterized in that: include: Module M1: Uses multi-source sensors to synchronously collect multimodal data of the cotter pin, including visible light images, infrared image temperature distribution data, and lidar point cloud data, and achieves spatiotemporal alignment based on hardware triggering; Module M2: Preprocesses multimodal data, including noise removal, resolution alignment, and feature extraction. It also uses a gated attention mechanism to dynamically fuse RGB texture, temperature gradient, and point cloud geometric features to generate multimodal fusion features. Module M3: Constructs a strength-conditional diffusion model, controls the defect generation intensity through latent space parameterization, and generates continuous defect state data from slight corrosion to complete fracture based on the strength parameters; Module M4: Input the generated defect state data into the lightweight finite element model for reverse optimization, verify the physical law consistency of the generated data, and output a multimodal defect data set for detection model training.
7. The split pin defect data enhancement system according to claim 6, characterized in that: The module M1 includes: The iterative closest point algorithm is used to optimize the rigid body transformation matrix and align the lidar point cloud data with the visible light image at the sub-pixel level. Perform bicubic interpolation on the infrared image to make its resolution consistent with that of the visible light image; Multi-sensor hardware triggering is achieved through GPS clock synchronization to ensure acquisition time synchronization.
8. The cotter pin defect data enhancement system according to claim 6, characterized in that: The module M2 includes: Use SIFT or ORB algorithm to extract image feature key points and descriptors F img ∈R C×H×W , C is the number of channels, H is the height, and W is the width; Calculate the two-dimensional temperature gradient field of the infrared image and quantify the abnormal heat conduction area, where x, y are the image coordinates and Δ is the pixel spacing, and obtain the feature F thermal ∈R T×H×W , T is temperature; Image noise is removed using median filtering or non-local mean filtering, and pixel values are scaled to the range of [0, 1]. Infrared images are subjected to histogram equalization to enhance temperature contrast and then normalized. Point clouds are de-redundant by downsampling the voxel grid or removing statistical outliers. The point cloud coordinates are centered and scaled, and the cleaned and standardized data is output for subsequent feature extraction, resulting in standardized RGB images, infrared images, and point clouds. For any point p in the point cloud i , select the k nearest neighbor point set N(p i ), calculate the covariance matrix based on the neighborhood point set, and use principal component analysis to perform eigendecomposition on the covariance matrix to define the surface curvature feature map C(p i )=λ min / (λ1+λ2+λ3), where λ1≥λ2≥λ3 is the eigenvalue, λ min is the minimum eigenvalue, and the point cloud feature F is obtained pc ∈R D×N ; Design a gated attention mechanism to generate attention weights α∈[0,1] by convolution of the three features H×W : Where σ1 is the sigmoid function, Point cloud features projected into image space; Dynamically adjust the contribution of each modality: Here, MLP represents a three-layer perceptron with an input dimension of 64 and an output dimension of 256, which upgrades the point cloud features to the same dimension as the image features.
9. The cotter pin defect data enhancement system according to claim 8, characterized in that: The module M3 includes: Define the defect intensity s∈[0,1], use two layers of MLP to project the intensity parameter s into the latent space, and generate the conditional vector c s ∈R 64 , c s =MLP(s)=W2·ReLU(W1·s+b1)+b2, W1 and W2 are weight matrices, b1, b2 are bias terms, and the conditional vector c s To control the defect generation intensity, as the input of the diffusion Unet encoder; Insert the conditional normalization layer CLN into each residual block of the UNet encoder to dynamically adjust the feature distribution through the conditional vector: where γ and β are given by c s The predicted scaling and bias parameters, represents the kth layer of the decoder; δ is a constant used to ensure the stability of normalization; The encoder output feature h enc and fusion feature F fusion Calculate cross attention: Attention(Q=h enc ,K=F fusion ,V=F fusion ), output the fused feature h fused As the initial input to the decoder; In the DDPM inverse denoising process, the perturbation variance σ2=β is adjusted by the intensity parameter s. t ·s,β t The noise scheduling parameters are preset to cover a continuous range from slight corrosion to complete fracture.
10. The cotter pin defect data enhancement system according to claim 9, characterized in that: The module M4 includes: The low-dimensional parameter S(θ)∈R that encodes the stress field of the defect simulated by the finite element method 32 and fusion feature F fusion Combined to get F diffusion :F diffusion =Concat(F fused ,S(θ))∈R 256+32 and RGB image I rgb ∈R H×W×3 As input to the strength-controllable generative network; Gaussian noise image sequence generated by forward diffusion The training loss is as follows: L phy =||φ gen -f FEM || The training process minimizes the loss, where ε follows a normal distribution and ε θ is the predicted noise, φ gen and φ FEM They represent the generated images, reverse deduction of stress parameters and real simulation results through lightweight finite element model; For time step t, real image I real and the expectation of noise ε; T is the maximum time step in the diffusion process; The total loss is: L total =L DDPM +0.5 L phy The trained diffusion model is used to generate multimodal data of cotter pin defects and a visible light cotter pin defect image is generated.
Citation Information
Patent Citations
Power transmission line cotter pin missing identification method
CN111815623A
Cited By
Industrial part defect sample accurate generation method based on conditional diffusion model
CN121280441A