Method, device, storage medium and equipment for detecting small hardware defects of power transmission line
By improving the Grey Wolf optimization algorithm and feature fusion technology through adaptive illumination preprocessing, chaotic information sharing, and other techniques, the problems of illumination interference, angle transformation, and occlusion in small fitting inspection were solved, achieving high-precision defect detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGXI KECHEN HONGXING INFORMATION TECH CO LTD
- Filing Date
- 2026-04-22
- Publication Date
- 2026-05-29
AI Technical Summary
Existing small fitting defect detection technologies suffer from problems such as light interference, algorithms being prone to getting stuck in local optima, varying shooting angles, and occlusion leading to missed or false detections, making it difficult to meet the accurate detection needs in complex field environments.
An improved Grey Wolf optimization algorithm based on illumination adaptive preprocessing and chaotic information sharing is adopted for hyperparameter optimization. Combined with deformable convolution and spatial transformation angle adaptive feature extraction, occlusion region identification and cross-scale feature fusion processing are performed to improve the adaptability and accuracy of the detection model.
It significantly improves the accuracy and robustness of small hardware defect detection, reduces the rate of missed and false detections, and adapts to the detection needs in complex field environments.
Smart Images

Figure CN122115422A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of defect detection technology, and in particular to a method, apparatus, storage medium and equipment for detecting defects in small fittings of power transmission lines. Background Technology
[0002] Transmission line fittings are core basic components for connecting and securing power lines in a power system. The absence of pins, bolts, nuts, and other components is a critical defect, easily leading to serious safety accidents such as line detachment, equipment short circuits, and large-scale power outages, directly threatening the safety of power grid operation and the reliability of power supply. With the development of intelligent power operation and maintenance, automated defect detection based on computer vision has gradually replaced traditional manual inspections, becoming the mainstream technology for monitoring the condition of transmission line fittings. Among these, target detection algorithms, represented by the YOLO series, are the most widely used. The industry also commonly attempts to combine metaheuristic optimization algorithms to optimize the performance of detection models, thereby improving the recognition effect of small target defects.
[0003] Currently, in the field of small hardware defect detection, mainstream technical solutions mostly adopt the basic YOLO algorithm combined with the traditional Grey Wolf Optimizer (GWO) algorithm to construct the detection model. The GWO algorithm optimizes core hyperparameters of the YOLO model, such as the learning rate, number of convolutional kernels, and aspect ratio of the anchor boxes. The optimized hyperparameters are then used to configure and train the model, and finally, defect inference and recognition are performed on field-collected images based on the trained model. This type of solution relies on the combination of metaheuristic optimization and target detection algorithms to attempt to solve the problem of insufficient accuracy in identifying small targets on small hardware. However, the traditional GWO algorithm itself has significant drawbacks. Its random population initialization leads to poor population diversity, the convergence factor only decays linearly and cannot be adaptively adjusted, and the wolf pack information interaction is only one-way transmission, making it prone to getting trapped in local optima during the hyperparameter optimization process. Ultimately, this prevents the model from effectively improving its feature extraction capability for small hardware defects, making it difficult to meet the accurate detection requirements in complex field environments.
[0004] Furthermore, existing detection solutions have significant shortcomings in adaptability in image preprocessing, feature extraction, and inference deployment. Images of small hardware acquired on-site often face lighting interference such as strong light reflection, weak light vignetting, and backlight overexposure. Random and variable shooting angles can easily lead to deformation of target features, and small hardware is easily obstructed by wires, insulators, etc. Existing solutions lack targeted lighting correction, angle adaptive feature extraction, and occlusion feature compensation mechanisms, which further exacerbates the problem of high rates of missed and false detections of defects in small hardware. Summary of the Invention
[0005] In view of this, this application provides a method, apparatus, storage medium and equipment for detecting defects in small fittings of transmission lines, which can effectively solve the problems of light interference, algorithm getting trapped in local optima, variable shooting angle and occlusion leading to missed detection and false detection in small fitting detection, and can significantly improve the accuracy and robustness of small fitting defect detection.
[0006] According to a first aspect of this application, a method for detecting defects in small fittings of transmission lines is provided, comprising: The on-site images of small hardware are subjected to illumination adaptive preprocessing to obtain standardized images after eliminating illumination distortion. The gray wolf optimization algorithm is improved by adopting chaotic information sharing. The hyperparameters of the target detection model are optimized by chaotic population initialization, bidirectional information sharing search and adaptive exponential decay convergence factor adjustment, so as to obtain the target hyperparameter combination suitable for small hardware defect detection. The target hyperparameter combination is configured into the target detection model and the model is trained. Based on the trained target detection model, the standardized image is subjected to angle adaptive feature extraction by combining deformable convolution and spatial transformation to obtain a multi-scale feature map that adapts to different shooting angles. The multi-scale feature map is subjected to occlusion region identification, hierarchical feature completion and cross-scale feature fusion processing to obtain a fused feature map with enhanced defect features. Based on the fused feature map, small hardware defects are classified and located to obtain defect detection results.
[0007] According to a second aspect of this application, a device for detecting defects in small fittings of transmission lines is provided, comprising: The processing module is used to perform illumination-adaptive preprocessing on the on-site images of small hardware fittings to obtain standardized images after eliminating illumination distortion. The optimization module is used to improve the Grey Wolf optimization algorithm by adopting chaotic information sharing. Through chaotic population initialization, bidirectional information sharing search and adaptive exponential decay convergence factor adjustment, the hyperparameters of the target detection model are optimized to obtain the target hyperparameter combination suitable for small hardware defect detection. The extraction module is used to configure the target hyperparameter combination into the target detection model and train the model. Based on the trained target detection model, the module performs angle adaptive feature extraction on the standardized image by combining deformable convolution and spatial transformation to obtain a multi-scale feature map that adapts to different shooting angles. The detection module is used to perform occlusion region identification, hierarchical feature completion and cross-scale feature fusion processing on the multi-scale feature map to obtain a fused feature map with enhanced defect features, and to classify and locate small hardware defects based on the fused feature map to obtain defect detection results.
[0008] According to a third aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described method for detecting defects in small fittings of transmission lines.
[0009] According to a fourth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described method for detecting defects in small fittings of power transmission lines.
[0010] By employing the above technical solutions, this application provides a method, apparatus, storage medium, and device for detecting defects in small fittings of transmission lines. Through adaptive preprocessing of on-site images of the small fittings, it can eliminate illumination distortions such as strong light, weak light, and backlight. The use of a gray wolf optimization algorithm improved with chaotic population initialization, bidirectional information sharing search, and adaptive exponential decay convergence factor for hyperparameter optimization avoids traditional algorithms getting stuck in local optima and improves the model's feature extraction capabilities. Deformable convolution and spatial transformation enable angle-adaptive feature extraction to overcome target deformation problems caused by varying shooting angles. Furthermore, by identifying occluded areas, hierarchical feature completion, and cross-scale feature fusion to compensate and enhance occluded area features, it comprehensively solves the problems of insufficient algorithm optimization performance, illumination interference, target feature deformation, and high false negative and false positive rates of small fitting defects caused by occlusion in traditional detection schemes. This significantly improves the accuracy and reliability of small fitting defect detection in complex on-site environments.
[0011] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a method for detecting defects in small fittings of transmission lines according to an embodiment of this application is shown. Figure 2 A flowchart illustrating a method for detecting defects in small fittings of power transmission lines according to another embodiment of this application is shown. Figure 3 This paper shows a schematic diagram of the structure of a small fitting defect detection device for power transmission lines provided in an embodiment of this application. Detailed Implementation
[0013] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0014] Currently, in the field of small hardware defect detection, mainstream technical solutions mostly adopt the basic YOLO algorithm combined with the traditional Grey Wolf Optimizer (GWO) algorithm to construct the detection model. The GWO algorithm optimizes core hyperparameters of the YOLO model, such as the learning rate, number of convolutional kernels, and aspect ratio of the anchor boxes. The optimized hyperparameters are then used to configure and train the model, and finally, defect inference and recognition are performed on field-collected images based on the trained model. This type of solution relies on the combination of metaheuristic optimization and target detection algorithms to attempt to solve the problem of insufficient accuracy in identifying small targets on small hardware. However, the traditional GWO algorithm itself has significant drawbacks. Its random population initialization leads to poor population diversity, the convergence factor only decays linearly and cannot be adaptively adjusted, and the wolf pack information interaction is only one-way transmission, making it prone to getting trapped in local optima during the hyperparameter optimization process. Ultimately, this prevents the model from effectively improving its feature extraction capability for small hardware defects, making it difficult to meet the accurate detection requirements in complex field environments.
[0015] Furthermore, existing detection solutions have significant shortcomings in adaptability in image preprocessing, feature extraction, and inference deployment. Images of small hardware acquired on-site often face lighting interference such as strong light reflection, weak light vignetting, and backlight overexposure. Random and variable shooting angles can easily lead to deformation of target features, and small hardware is easily obstructed by wires, insulators, etc. Existing solutions lack targeted lighting correction, angle adaptive feature extraction, and occlusion feature compensation mechanisms, which further exacerbates the problem of high rates of missed and false detections of defects in small hardware.
[0016] Accordingly, in order to solve the above-mentioned technical problems, embodiments of the present invention provide a method for detecting defects in small fittings of transmission lines, such as... Figure 1 As shown, the method includes: Step 110: Perform illumination adaptive preprocessing on the on-site image of the small hardware to obtain a standardized image after eliminating illumination distortion.
[0017] Among them, the small fitting field image refers to the original image data containing the target of the small power fittings collected in the actual operation scenario; the illumination adaptive preprocessing refers to the image processing process that automatically selects the appropriate processing strategy according to the actual illumination conditions of the image to correct the illumination distortion problem; illumination distortion is the phenomenon of abnormal image brightness, loss of detail, and overexposure of areas caused by the complex illumination environment on site; the standardized image refers to the image data that has been processed by illumination correction and standardization, with uniform illumination conditions and standardized data format, and can be directly input into the detection model for feature extraction.
[0018] In specific application scenarios, drone terminals can be used to collect real-time images of small hardware fittings. In this embodiment of the present disclosure, the local terminal can receive the images of small hardware fittings transmitted back by the drone terminal, and perform adaptive illumination preprocessing on the images. Based on the overall illumination distribution of the image, the corresponding processing method is dynamically matched to suppress and correct various illumination anomalies, so that the image is restored to a uniform and stable illumination performance. At the same time, the image specifications are standardized, and finally a standardized image that can adapt to the processing requirements of subsequent detection models is formed.
[0019] By performing adaptive lighting preprocessing on the images of small fittings on-site and generating standardized images, various lighting interferences caused by complex on-site environments can be effectively eliminated, ensuring that the image data of the input model has stable and consistent lighting conditions, reducing feature loss and recognition deviation caused by lighting distortion, and thus reducing the probability of missed detection and false detection in the defect detection process.
[0020] Step 120: The gray wolf optimization algorithm is improved by chaotic information sharing. The hyperparameters of the target detection model are optimized by chaotic population initialization, bidirectional information sharing search and adaptive exponential decay convergence factor adjustment to obtain the target hyperparameter combination suitable for small hardware defect detection.
[0021] Among them, the chaotic information sharing improved gray wolf optimization algorithm is an optimization algorithm that introduces a chaotic mechanism for population initialization, constructs bidirectional information transmission paths between individuals, and adaptively adjusts the convergence factor by exponential decay, based on the traditional gray wolf optimization algorithm; chaotic population initialization is an initialization method that uses the traversal characteristics of chaotic sequences to construct the initial population of the algorithm, improving the uniformity and diversity of population distribution; bidirectional information sharing search is a search method that guides individuals to conduct positive transmission and negative feedback of position information between individuals and ordinary individuals, realizing information interaction of the entire population; adaptive exponential decay convergence factor adjustment is a parameter adjustment method that dynamically changes the decay method and magnitude of the convergence factor according to the algorithm iteration process and search state, rather than a fixed linear decay; the target detection model refers to a deep learning detection network used for feature extraction, classification, and localization of small hardware defects; hyperparameter optimization is a parameter optimization process that automatically searches for the optimal combination within a preset parameter space to achieve the best model detection performance; the target hyperparameter combination is a set of core model parameters adapted to the small hardware defect detection task, obtained through optimization algorithm search.
[0022] In this embodiment of the disclosure, the overall optimization process can be carried out by using the improved gray wolf optimization algorithm with chaotic information sharing. The initial population of the algorithm is constructed by means of chaotic characteristics. During the iterative search process, bidirectional information interaction between individuals within the population is realized. At the same time, the exponential decay law of the convergence factor is dynamically adjusted in combination with the algorithm running state. In this way, the relevant hyperparameters of the target detection model are globally optimized, and finally the target hyperparameter combination that can be adapted to the small hardware defect detection task is selected and determined.
[0023] By employing a chaotic information sharing-improved gray wolf optimization algorithm for hyperparameter optimization, the diversity of the initial population can be effectively enhanced. This breaks the limitation of unidirectional information transmission in traditional algorithms, enables dynamic adaptive control of the convergence process, avoids getting trapped in local optima, and obtains a hyperparameter combination that is more suitable for small hardware defect detection. This provides a better parameter basis for subsequent model training and feature extraction, thereby improving the overall detection performance of the model.
[0024] Step 130: Configure the target hyperparameter combination into the target detection model and train the model. Based on the trained target detection model, perform angle adaptive feature extraction on the standardized image by combining deformable convolution and spatial transformation to obtain multi-scale feature maps that are adapted to different shooting angles.
[0025] Among them, deformable convolution is a convolution operation that can autonomously learn offsets and adaptively capture the deformation features of the target; spatial transformation is a feature processing operation that performs viewpoint normalization and geometric correction on the feature map; angle adaptive feature extraction is a feature extraction method that automatically adjusts the feature extraction method according to the change of shooting angle and adapts to the feature extraction method of target shape from different viewpoints; multi-scale feature map is a multi-level feature expression result containing different receptive fields and corresponding to defects of different sizes.
[0026] In this embodiment of the disclosure, the optimized target hyperparameter combination can be configured into the target detection model to complete parameter initialization. Then, the model is trained using a small hardware defect sample dataset, enabling the model to have stable feature learning and detection capabilities. After inputting standardized images into the trained target detection model, angle-adaptive feature extraction is performed through deformable convolution and spatial transformation to automatically adapt to changes in target shape caused by the shooting angle, and finally outputs multi-scale feature maps containing different receptive fields.
[0027] By configuring the target hyperparameter combination into the target detection model and completing the training, the advantages of the optimization algorithm can be fully utilized to improve the model's feature learning ability for small hardware defects. Through angle-adaptive feature extraction combining deformable convolution and spatial transformation, the problem of target deformation caused by random changes in shooting angle can be effectively overcome, and the features of small hardware under different angles can be accurately extracted. The output of multi-scale feature maps can cover defect targets of different sizes, providing a rich and reliable feature foundation for subsequent occlusion processing and defect detection.
[0028] In specific application scenarios, as a preferred approach, given the large number of parameters in the trained target detection model, which makes real-time detection impossible on portable terminals at power operation and maintenance sites, the trained target detection model can be loaded onto a server for deployment to achieve real-time inference of small fitting defects. In this embodiment, the local terminal can call the trained target detection model on the server to perform angle-adaptive feature extraction on standardized images using deformable convolution and spatial transformation, obtaining multi-scale feature maps adapted to different shooting angles.
[0029] Deploying the target detection model on the server side can make full use of the server's computing resources, ensure the stable and efficient operation of the detection process, and meet the needs of real-time on-site detection and rapid processing of large batches of images.
[0030] Step 140: Perform occlusion region identification, hierarchical feature completion and cross-scale feature fusion processing on the multi-scale feature map to obtain a fused feature map with enhanced defect features. Based on the fused feature map, classify and locate defects in small hardware fittings to obtain defect detection results.
[0031] The process includes: Occlusion region identification, which involves determining the location and quantifying the proportion of occluded small hardware regions in a multi-scale feature map; Hierarchical feature completion, which involves restoring and supplementing missing features in occluded regions based on differences in occlusion degree; Cross-scale feature fusion, which involves weighted integration of feature information at different scales to achieve a unified expression of global and local features; Defect-enhanced fused feature map, which is a feature map with more prominent defect information and more complete feature expression after occlusion completion and cross-scale fusion processing; Small hardware defect classification, which involves determining the type of defect present in the small hardware based on the fused feature map; Small hardware defect localization, which involves determining the location of the defect in the image based on the fused feature map; and Defect detection result, which refers to the final detection output data containing defect type and defect location information.
[0032] In this embodiment of the disclosure, the generated multi-scale feature map can be sequentially processed by occlusion region identification, hierarchical feature completion, and cross-scale feature fusion to gradually repair the information loss caused by occlusion in the feature map and enhance the expression intensity of defect features, forming a fused feature map with enhanced defect features. The fused feature map is then input into the target detection model deployed on the server, and the model is used to complete the type determination and location determination of small hardware defects, and finally output a complete defect detection result.
[0033] By performing occlusion region identification, hierarchical feature completion, and cross-scale feature fusion on multi-scale feature maps, the problem of feature loss caused by occlusion can be effectively compensated, significantly enhancing the integrity and recognizability of defect features. Based on the enhanced fused feature maps, defect classification and localization can be carried out, which can greatly reduce the probability of missed detection and false detection in complex field environments and improve the accuracy and reliability of small hardware defect detection.
[0034] In summary, the small fitting defect detection method for transmission lines provided in this application can eliminate illumination distortions such as strong light, weak light, and backlight by performing adaptive preprocessing on the on-site images of the small fittings. The use of a gray wolf optimization algorithm improved with chaotic population initialization, bidirectional information sharing search, and adaptive exponential decay convergence factor for hyperparameter optimization avoids traditional algorithms getting stuck in local optima and improves the model's feature extraction capability. Deformable convolution and spatial transformation enable angle-adaptive feature extraction to overcome target deformation problems caused by varying shooting angles. Furthermore, by identifying occluded areas, hierarchical feature completion, and cross-scale feature fusion to compensate and enhance the features of occluded areas, the method comprehensively solves the problems of insufficient algorithm optimization performance, illumination interference, target feature deformation, and high false negative and false positive rates of small fitting defects caused by occlusion in traditional detection schemes. This significantly improves the accuracy and reliability of small fitting defect detection in complex on-site environments.
[0035] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, and to fully illustrate the implementation of this embodiment, this embodiment also provides another method for detecting defects in small fittings of transmission lines, such as... Figure 2 As shown, the method includes: Step 210: Perform illumination adaptive preprocessing on the on-site image of the small hardware to obtain a standardized image after eliminating illumination distortion.
[0036] For embodiments of this disclosure, step 210 may include the following steps: Step 210-1: Calculate the average brightness and contrast of the small hardware field image. Based on the average brightness and contrast, divide the small hardware field image into three lighting types: strong light, weak light, and backlight.
[0037] Among them, average brightness is a statistical indicator reflecting the overall brightness of the image, which is calculated by weighting the gray values of all pixels in the image; contrast is a statistical indicator reflecting the degree of gray difference in different areas of the image; and illumination type refers to the image illumination category classified according to the on-site imaging conditions, including three types: strong light, weak light, and backlight.
[0038] In this embodiment of the disclosure, pixel-level grayscale statistics can be performed on the on-site image I of the small hardware, and the average brightness L of the pixel grayscale values at all coordinate positions within the width and height range of the image can be calculated according to the following formula. avgAnd calculate the image contrast C based on the maximum and minimum pixel grayscale values: ; ; In the formula, W and H are the width and height of the small hardware scene image, respectively, I(i,j) is the pixel gray value at coordinate (i,j) in the small hardware scene image, and max(I(i,j)) and min(I(i,j)) are the maximum and minimum values of the image pixel gray value, respectively.
[0039] Furthermore, the lighting type can be automatically determined based on the set numerical thresholds. For example, an average brightness greater than 200 is determined to be a strong light scene, an average brightness less than 50 is determined to be a weak light scene, and a contrast less than 30 is determined to be a backlight scene, thereby completing the classification of lighting type in the on-site images of small hardware.
[0040] By calculating the average brightness and contrast of the small hardware field images and classifying the lighting types, we can accurately identify complex lighting conditions on site. This provides an accurate basis for subsequent targeted correction strategies to eliminate lighting distortion, ensuring that images under different lighting conditions can enter the appropriate processing flow, improving the stability and adaptability of image preprocessing, and providing a high-quality data foundation for subsequent model feature extraction and defect detection.
[0041] Step 210-2: Use differentiated combination correction strategies for different lighting types to eliminate lighting distortion and obtain lighting-corrected images.
[0042] The differentiated combination correction strategy is a processing scheme that combines different processing methods to correct illumination distortion based on the illumination type of the image. In specific application scenarios, the differentiated combination correction strategy may include: using polarization filtering and gamma compression correction for strong light illumination types, using multi-scale enhancement and local contrast enhancement correction for weak light illumination types, and using dark channel estimation and background fusion correction for backlight illumination types.
[0043] In the embodiments of this disclosure, differentiated combination correction strategies can be adopted for different lighting scenarios after determination to achieve precise elimination of lighting distortion. Specifically: (1) For strong light illumination, a combination strategy of “3×3 Gaussian blur polarization filter + adaptive gamma compression” can be adopted.
[0044] The first step is to eliminate reflections from the metal mirror by simulating polarization filtering using a 3×3 Gaussian blur. The Gaussian filtering formula is as follows: ; Wherein, G(x,y) is the value of the two-dimensional Gaussian function at coordinates (x,y), which is the weight value of the corresponding position of the Gaussian filter kernel; (x,y) is the pixel coordinate of the weight to be calculated within the Gaussian kernel, representing the relative position within the filtering window; σ is the standard deviation of the Gaussian distribution, used to control the smoothness of the Gaussian kernel, which can be 0.8 in this application. The larger σ is, the smoother the filtering effect and the stronger the de-glare capability; e is the natural constant, approximately 2.71828, which is the base of the exponential operation; (u,v) is the center coordinate of the filter kernel.
[0045] The original strong light image I is convolved with a 3×3 Gaussian kernel to obtain the de-reflected intermediate image I1.
[0046] The second step involves adaptively adjusting the gamma value γ∈[0.4,0.6] based on the image reflectivity, performing gamma compression on I1, and completing the brightness reduction and defect feature recovery of the strong light image to obtain the illumination-corrected image I under the strong light illumination type. 强光 (i,j). The formula is: ; In the formula, I 强光 (i,j) represents the output pixel value at coordinate (i,j) after gamma correction in a strong light scene, i.e., the image pixel value after strong light suppression; (i,j) represents the row and column coordinates of the pixel to be processed in the image, used to locate the specific pixel position in the image; I1(i,j) represents the input pixel value at coordinate (i,j) after Gaussian blur de-reflection processing, which is the original input data for gamma correction; γ is the gamma correction coefficient, which can take values in the range of [0.4,0.6] in this application, used to control the degree of strong light suppression. The smaller γ is, the more significant the pixel compression in the strong light area, and the stronger the strong light suppression effect.
[0047] (2) For weak light types, a combination strategy of "multi-scale Retinex enhancement + local histogram equalization" can be adopted.
[0048] The first step is to use multi-scale Retinex decomposition with scales of 3, 5, and 7 to enhance the brightness of dark areas. The core Retinex formula is: ; Where R(x,y) is the image reflection component output by the Retinex algorithm, which is the feature component containing only essential information such as defects and structure of small hardware after removing the influence of illumination; (x,y) is the spatial coordinate of the pixel to be calculated in the image, used to locate the specific pixel position in the image; I(x,y) is the pixel value of the original image to be enhanced at coordinates (x,y), that is, the pixel data of the inspection image under low light / backlight scene; F(x,y) is the Gaussian wrap function; This is the convolution operator.
[0049] The enhanced dark area image I2 is obtained by averaging the reflection components at three scales. The second step involves performing block-based local histogram equalization on I2. The image is divided into 8×8 sub-blocks, and histogram equalization is applied to each sub-block to enhance the grayscale contrast between defects and the background, resulting in the illumination-corrected image I under low-light illumination conditions. 弱光 (i,j).
[0050] (3) For backlighting, the strategy of “dark channel prior + background compensation image fusion” can be adopted.
[0051] Step 1: Extract the dark channel I of the image dark The formula is: ; Among them, I dark (x,y) represents the dark channel value of the input image at coordinates (x,y), which is the result of taking the minimum value twice for the RGB three-channel pixels within the local window, used to characterize the dark features of the image; (x,y) represents the center coordinates of the pixel in the image whose dark channel value is to be calculated, used to locate the center position of the local window; min c∈{R,G,B} This is the first minimum value operation, which means taking the minimum channel pixel value at pixel (i,j) among the three RGB color channels. This is used to eliminate color differences and extract the basic features of the dark areas of the image; c represents the RGB color channels of the image; min (i,j)∈Ω(x,y) The second minimum value operation involves taking the minimum dark channel value of all pixels within a local neighborhood window Ω(x,y) centered at (x,y) to estimate the dark reference of the local region; (i,j) represents the pixel coordinates within the local neighborhood window Ω(x,y), indicating the position of the pixel to be calculated within the window; Ω(x,y) is the local neighborhood window centered at (x,y), typically a 15×15 rectangular window in this application, used to statistically analyze the dark features of the local region; c (i,j) represents the original pixel value of color channel c at coordinates (i,j) in the input image, which is the original input data for calculating the dark channel; The second step involves generating a background compensation image I3 based on the background illumination components. The original backlight image I and the background compensation image I3 are then linearly fused at the pixel level with a weight ratio of 0.7:0.3 to improve the feature visibility of the backlight area, resulting in an illumination correction image I under the backlight illumination type. 逆光 (i,j). The formula is: ; Among them, I 逆光(i,j) represents the output pixel value at coordinate (i,j) after fusion correction in a backlit scene, i.e., the illumination-corrected image after backlight compensation; (i,j) represents the row and column coordinates of the pixel to be processed in the image, used to locate the specific pixel position in the image; I(i,j) represents the pixel value at coordinate (i,j) of the original backlit image input, which is the original input data for the fusion operation; I3(i,j) represents the pixel value at coordinate (i,j) of the background compensation image obtained after processing by algorithms such as dark channel prior, used to compensate for the loss of dark area details in backlit scenes.
[0052] Differentiated combination correction strategies for different lighting types can accurately match the correction requirements of various lighting distortions, effectively eliminate image distortion caused by strong light reflection, weak light vignetting and backlight overexposure, fully restore the defect details on the surface of small hardware, improve the overall image quality and consistency, and provide stable and reliable data support for subsequent feature extraction and defect detection.
[0053] Step 210-3: Scale the illumination-corrected image to a standard size and normalize the pixels to obtain a standardized image.
[0054] For embodiments of this disclosure, the illumination-corrected image I can be... 矫正 (i,j) (i.e. I 强光 (i,j) / I 弱光 The image (i,j) / I (backlighting) is uniformly scaled to a standard pixel size of 640×640, and the pixel values are normalized to map the pixel grayscale values to the [0,1] interval, resulting in a standardized image I. std (i,j) serves as the unified input for subsequent modules. The formula is: ; Among them, I std (i,j) is the output pixel value at coordinate (i,j) after normalization, that is, the image pixel data normalized to the interval [0,1]; (i,j) is the row and column coordinates of the pixel to be processed in the image, used to locate the specific pixel position in the image; The input pixel value at coordinates (i,j) after correction for strong light / backlight / weak light is the original input data for the normalization operation; The minimum pixel value of the entire image after correction is used to set the reference of the pixel value to zero and eliminate the overall brightness shift of the image; The maximum pixel value of the entire image after correction is used to normalize the dynamic range of pixel values and unify the pixel scale of images under different lighting conditions.
[0055] Scaling the illumination-corrected images to a standard size and performing pixel normalization ensures that all input model images maintain consistent spatial dimensions and numerical distribution, reducing detection fluctuations and computational errors caused by inconsistent image dimensions, and improving the stability and detection efficiency of model feature extraction.
[0056] Step 220: The gray wolf optimization algorithm is improved by chaotic information sharing. The hyperparameters of the target detection model are optimized by chaotic population initialization, bidirectional information sharing search and adaptive exponential decay convergence factor adjustment to obtain the target hyperparameter combination suitable for small hardware defect detection.
[0057] For embodiments of this disclosure, step 220 may include the following steps: Step 220-1: Generate a chaotic sequence using a logical mapping of a completely chaotic state. Based on the chaotic sequence, assign values to the individuals in the population of the improved gray wolf optimization algorithm for chaotic information sharing to generate the initial wolf population and complete the initialization of the chaotic population.
[0058] Among them, the fully chaotic state of the logistic mapping is a Logistic mapping method that runs with specified control parameters and can generate a uniformly traversable characteristic sequence; the chaotic sequence is a numerical sequence that is non-periodic and uncorrelated, generated by the iterative generation of the logistic mapping; the population individual is an independent search unit representing a set of hyperparameter combinations in the iterative search of the optimization algorithm; the initial wolf pack population is an initial search set composed of multiple population individuals constructed before the start of the algorithm iteration; the chaotic population initialization is a population construction method that uses the chaotic sequence to complete the assignment of population individuals, thereby improving the uniformity of the initial distribution.
[0059] For embodiments of this disclosure, a chaotic sequence can be generated using a Logistic chaotic mapping in a fully chaotic state, using the formula... Complete the iterative calculation, where x n+1 The output chaos value of the (n+1)th iteration of the Logistic mapping, with a value range of (0,1), is used to generate the initial population position for the Gray Wolf optimization algorithm, thus achieving chaotic initialization of the population; x nThe input chaotic value for the nth iteration of the Logistic mapping, ranging from (0,1), is the basic input data for the iterative calculation. μ is the chaos control parameter of the Logistic mapping, which is fixed at 4 in this application. When μ=4, the mapping is in a completely chaotic state, generating an initial population with strong ergodicity and randomness, thus preventing the Grey Wolf Optimization Algorithm from getting trapped in local optima. n is the iteration number, used to identify the iteration round of the chaotic sequence. Multiple iterations generate a chaotic initial population that meets the requirements. Furthermore, based on the generated chaotic sequence, hyperparameter interval mapping and assignment can be performed on the individuals of the population of the improved Grey Wolf Optimization Algorithm through chaotic information sharing to generate an initial wolf population of a specified size (e.g., 30), thereby completing the initialization of the chaotic population of the algorithm. Each wolf in the initial wolf population corresponds to a set of YOLOv12 hyperparameter combinations X. i =(lr i ,ratio i N convi ,w atti ), i=1,2,...,30.
[0060] Range mapping is performed on the hyperparameter values generated by chaos to match the physical value range of each hyperparameter: The learning rate lr ∈ [0.0001, 0.01]; The aspect ratio of the anchor frame is [0.5, 2.0]. Number of convolution kernels N conv ∈[16,128] (take positive integers); Attention weight w att ∈[0,1].
[0061] By using a completely chaotic state of logical mapping to construct the initial wolf population, the uniformity of the initial population distribution and global traversal can be significantly improved. This avoids the problem of insufficient population diversity caused by traditional random initialization, provides a more comprehensive search starting point for subsequent hyperparameter optimization, effectively reduces the probability of the algorithm getting stuck in local optima, and improves the stability and search accuracy of the overall optimization process.
[0062] Step 220-2: Based on the initial wolf pack population, determine the α, β, and δ-level guiding individuals and ordinary ω individuals. Forwardly transmit the position information of the α, β, and δ-level guiding individuals to the ordinary ω individuals to guide them to update their positions. Backwardly feed back the candidate positions searched by the ordinary ω individuals to the α, β, and δ-level guiding individuals to update their position information. Perform a two-way information sharing search based on the information interaction formed by forward transmission and backward feedback.
[0063] Among them, α, β, and δ-level guiding individuals are the best, second-best, and third-best individuals determined by fitness ranking in the initial wolf pack, and they play a guiding role in the search; ordinary ω individuals are the remaining individuals in the initial wolf pack other than the level guiding individuals, and they play a global exploration role; position information is a set of hyperparameters corresponding to an individual, used to characterize the individual's state in the search space; forward propagation is the process by which level guiding individuals send their own position information to ordinary ω individuals, guiding ordinary individuals to perform position updates; backward feedback is the process by which ordinary ω individuals return the high-quality candidate positions they have explored, used to update the level guiding individuals; bidirectional information sharing search is a search mechanism composed of forward propagation and backward feedback, realizing information interaction across the entire population.
[0064] In this embodiment of the disclosure, the limitation of one-way information transmission from the optimal wolves α, β, and δ to the ordinary wolf ω in the traditional GWO algorithm can be broken. A two-way information sharing search mechanism is designed to fully explore the global exploration value of the wolf ω, and improve the globality and comprehensiveness of the optimization. The core execution logic is as follows: Positive guidance: α, β, and δ wolves will determine the current optimal position (optimal hyperparameter combination) X. α X β X δ This is passed to all ω wolves, who then perform a local search according to the traditional GWO position update formula, which is: ; in, D β D δ C1 represents the Euclidean distance between the current gray wolf individual and α, β, and δ wolves, respectively, representing the positional difference between the individual and the best, second best, and third best individuals in the population; C1, C2, and C3 are random weight coefficients with values ranging from [0,2], used to simulate random interference from prey in nature, enhance the global search capability of the algorithm, and avoid getting trapped in local optima. X β X δ Let X(t) be the current position vector of α wolf, β wolf, and δ wolf in the population, representing the best, second best, and third best solutions in the current iteration, respectively; X(t) is the position vector of the current gray wolf individual in the t-th iteration, i.e., the candidate solution of the current individual; This is an absolute value operation used to calculate the distance modulus of a position vector.
[0065] ; Where X1, X2, and X3 are candidate position vectors calculated based on the positions of α wolf, β wolf, and δ wolf, respectively, representing the movement directions of an individual towards the three optimal individuals; A1, A2, and A3 are convergence factor coefficients, with values ranging from [...]. [a, a], where a decays linearly from 2 to 0 with the number of iterations, which is used to balance the algorithm's global exploration (early iteration) and local development (late iteration) capabilities; D β D δ Let be the Euclidean distances between the current gray wolf individual and α wolf, β wolf, and δ wolf, respectively.
[0066] ; Where X(t+1) is the updated position vector of the gray wolf individual in the (t+1)th iteration, which is the candidate solution for the next iteration; X1, X2, and X3 are candidate position vectors calculated based on the positions of α wolf, β wolf, and δ wolf, respectively, representing the movement direction of the individual towards the three optimal individuals; Reverse feedback: ωwolf will search for candidate optimal positions Feedback is sent to Alpha Wolf, which calculates all... The objective function value is used to filter out better candidate positions and update X. α X β X δ The set enables dynamic optimization of the wolf pack's global search direction.
[0067] By constructing a two-way information sharing search mechanism between hierarchical guided individuals and ordinary individuals, the limitations of the traditional algorithm's one-way information transmission can be broken, the global exploration ability of ordinary individuals can be fully explored, the search breadth and quality of the algorithm in the hyperparameter space can be improved, the algorithm can be prevented from converging too early, and the global optimal approximation ability and overall stability of hyperparameter optimization can be effectively improved.
[0068] Step 220-3: Construct the exponential decay base term based on the current iteration number and the total iteration number of the algorithm. Combine the change of the loss function in adjacent iteration cycles to dynamically correct the exponential decay base term, and adjust the decay rate and convergence amplitude of the convergence factor in real time to complete the adaptive exponential decay convergence factor adjustment.
[0069] Wherein, the current iteration number is the number of iterations that the optimization algorithm has completed during execution; the total number of iterations is the maximum number of iterations preset by the optimization algorithm; the exponential decay base term is a calculation term built on the natural exponent to control the basic trend of the convergence factor; and the change in the loss function between adjacent iteration cycles is the absolute difference in the value of the loss function between two consecutive iterations.
[0070] For the embodiments of this disclosure, an exponentially decaying basic term can be constructed based on the current iteration number and the total iteration number of the algorithm, and the convergence factor can be calculated using the following formula: ; in, is the adaptive convergence factor of the gray wolf optimization algorithm, used to control the search step size and convergence speed of individual gray wolves, balancing the algorithm's global exploration (early stage) and local exploitation (late stage) capabilities; e is a natural constant, approximately 2.71828, serving as the base for the exponential decay operation to achieve non-linear smooth decay of the convergence factor; t represents the current iteration number, T represents the total number of iterations (which can take the value 100), and e is a natural constant, 2e (t / T)2 For the exponentially decaying term, achieve The loss decays exponentially from 2 to 0; λ is the loss correction coefficient, ΔLoss=|Loss(t) Loss(t 1) | represents the change in the loss function between two adjacent iterations, enabling dynamic adaptive adjustment of the convergence factor. By combining the change in the loss function between adjacent iterations to dynamically correct the exponential decay base term, the decay rate and convergence amplitude of the convergence factor can be adjusted in real time, thus completing the adaptive exponential decay convergence factor adjustment.
[0071] By combining the iteration process with the change in the loss function to adaptively and exponentially adjust the convergence factor, the algorithm can maintain a large search range in the early stage to improve global exploration ability, and gradually narrow the search range in the later stage to enhance local optimization accuracy. This effectively balances the search breadth and convergence accuracy, avoids the problem of premature convergence or slow convergence, and further improves the efficiency and reliability of hyperparameter optimization.
[0072] Step 220-4: Using chaotic population initialization, bidirectional information sharing search, and adaptive exponential decay convergence factor adjustment as the iterative optimization process, and taking the detection accuracy of the target detection model as the optimization basis, the target hyperparameter combination is iteratively selected. The target hyperparameter combination includes at least the learning rate, anchor box aspect ratio, number of convolution kernels, and attention weights.
[0073] The iterative optimization process is a complete algorithm execution process that gradually approaches the optimal solution through multiple iterations; the learning rate is a core parameter that controls the step size of model parameter updates and affects the training convergence speed and stability; the aspect ratio of the anchor box is a parameter used to match the target shape and define the basic proportional relationship of the detection box; the number of convolutional kernels is the number of convolutional units used to extract features in the target detection model; and the attention weight is a weight parameter used to enhance key region features and suppress background interference.
[0074] In this embodiment of the disclosure, chaotic population initialization, bidirectional information sharing search, and adaptive exponential decay convergence factor adjustment can be integrated into a complete iterative optimization process. The mean average precision (mAP) of the target detection model on the validation set is used as the maximizing objective function F(X) for iterative optimization. During the iteration process, the learning rate, anchor box aspect ratio, number of convolutional kernels, and attention weights are interval mapped and searched. The parameter configuration is continuously optimized through multiple iterations, and finally the hyperparameter combination that maximizes the mean average precision is selected as the target hyperparameter combination.
[0075] By conducting hyperparameter iterative optimization using the complete process of the improved Grey Wolf optimization algorithm, the model parameter configuration that is most suitable for small hardware defect detection can be selected with detection accuracy as the clear guide. This fully enhances the model's feature extraction and recognition capabilities for small target defects, enabling the model to maintain higher detection accuracy and stronger generalization ability in complex field environments. It also provides a stable and efficient model foundation for subsequent angle adaptive feature extraction and occlusion compensation processing.
[0076] Step 230: Configure the target hyperparameter combination into the target detection model and train the model. Based on the trained target detection model, perform angle adaptive feature extraction on the standardized image by combining deformable convolution and spatial transformation to obtain multi-scale feature maps that are adapted to different shooting angles.
[0077] In the embodiments of this disclosure, when performing angle-adaptive feature extraction combining deformable convolution and spatial transformation on a standardized image based on a trained target detection model to obtain a multi-scale feature map adapted to different shooting angles, step 230 of the embodiment may include the following steps: Step 230-1: Input the standardized image into the backbone network of the target detection model, extract features from the standardized image through deformable convolution in the backbone network, adaptively obtain the deformation features of the small metal tool under different shooting angles, and obtain the deformable convolution feature map.
[0078] Among them, deformable convolution is a convolution operation that can autonomously learn offsets and adapt to changes in target shape; deformation features are the rotation, tilt, and offset features of small hardware caused by changes in shooting angle; deformable convolution feature map is the feature expression result containing angle deformation information after being extracted by deformable convolution.
[0079] In the embodiments of this disclosure, a standardized image can be input into the backbone network of the target detection model, and feature extraction can be performed through deformable convolutions in the backbone network, according to the formula. The calculation is complete, where y(p) is the pixel value of the output feature map; p is the target sampling position coordinate on the output feature map, representing the target pixel position of the convolution output; w(k) is the convolution kernel weight; k is the standard sampling offset within the convolution kernel; ∑ k∈K This is a summation operation for all sampling points within the convolution kernel, where K is the set of sampling points for the convolution kernel (e.g., a 3×3 convolution corresponds to 9 sampling points), used to aggregate neighborhood features; x is the pixel value of the input feature map; p0 is the center of the convolution kernel; p k Fixed offset for the convolution kernel; Δp k The deformation offset obtained by the network through autonomous learning is used to adaptively capture the deformation features of the small fitting under different shooting angles, and finally output a deformable convolutional feature map.
[0080] By inputting standardized images into the backbone network and using deformable convolution for feature extraction, the network can adaptively match the target deformation caused by different shooting angles, accurately capture the structural and defect features of small hardware, improve the feature expression capability under non-orthogonal viewpoints, and provide a rich and stable feature foundation for subsequent multi-scale fusion and spatial transformation correction.
[0081] Step 230-2: Perform channel-level splicing and fusion of deformable convolutional feature maps and multi-scale fixed convolutional feature maps to obtain an intermediate feature map that integrates multi-scale information and deformation information.
[0082] Among them, the multi-scale fixed convolutional feature map is a feature map containing multi-scale structural information extracted using fixed convolutional kernels of different sizes; the channel-level stitching fusion is a fusion method that merges multiple feature maps in the feature channel dimension to achieve complementary information; the multi-scale information is the global structural information and local detail information extracted by convolution of different receptive fields; the deformation information is the geometric deformation features of the target caused by changes in shooting angle learned by deformable convolution; the intermediate feature map is a transitional feature expression that simultaneously carries multi-scale information and deformation information after channel-level stitching fusion.
[0083] In this embodiment of the disclosure, deformable convolutional feature maps and fixed convolutional feature maps of multiple scales can be jointly input into the feature fusion unit, according to the formula... Perform channel-level splicing and fusion, where F conv This is the intermediate feature map after fusion. Concat represents the channel-level concatenation operation, F deform For deformable convolutional feature maps, F 3×3 F 5×5 F 7×7 They are three different scales ( , , The feature map is obtained by fixed convolution. By concatenating the features, global information, local details and angular deformation information are included simultaneously, resulting in an intermediate feature map that integrates multi-scale information and deformation information.
[0084] By performing channel-level concatenation and fusion of deformable convolutional feature maps with multi-scale fixed convolutional feature maps, we can effectively combine deformation adaptability and multi-scale feature representation capabilities. This enhances the ability of feature maps to describe small hardware targets of different angles and sizes, improves feature richness and robustness, and provides a higher quality feature foundation for subsequent spatial transformations and attention enhancement.
[0085] Step 230-3: Input the intermediate feature map into the region attention module of the object detection model. Use the spatial transformation network embedded in the region attention module to perform affine transformation correction on the intermediate feature map, normalize the non-orthogonal view features in the intermediate feature map into standard orthogonal view features, and output the attention-enhanced feature map after view correction.
[0086] In this embodiment, the intermediate feature map can be input into the region attention module of the object detection model. The Spatial Transformer Network (STN) embedded in the module autonomously learns the affine transformation matrix based on the intermediate feature map. It then performs rotation, translation, and scaling corrections on the intermediate feature map according to the affine transformation formula, normalizing the non-orthogonal view features to standard orthogonal view features. Specifically, the STN can adaptively transform the non-orthogonal view feature map to the orthogonal view feature space through affine transformation. The affine transformation matrix M is autonomously learned by the network, and the transformation formula is: , of which F stn (x, y) represents the feature value of the spatial transformation network at the output coordinates (x, y), i.e., the pixel value of the final feature map after pose correction and geometric normalization; (x, y) represents the target pixel coordinates on the output feature map after spatial transformation. To ensure the integrity of the coordinate mapping, x and y are usually taken as integer coordinates; ∑ i,j For the dual summation operation, it iterates through all possible pixel positions (i,j) on the input feature map to achieve global sampling and weighted aggregation of the input features; F conv (i,j) is the input feature map (i.e., the F after multi-scale fusion mentioned above). convThe original pixel value at coordinates (i,j) provides the feature data to be transformed for spatial transformation; T(x,y,i,j,M) is the spatial transformation interpolation kernel / sampling weight function, used to calculate the coordinate mapping relationship. It calculates the interpolation weight from the input position (i,j) to the output position (x,y) based on the output coordinates (x,y), the input coordinates (i,j), and the transformation matrix M; (i,j) are the original pixel coordinates on the input feature map before spatial transformation, which is the source position of the transformation operation; M is a 2×3 affine transformation matrix, which is learned autonomously by the network and used to define the geometric transformation relationship of the image (including translation, scaling, rotation, cropping, etc.). This matrix is used to achieve standardized correction of arbitrary shooting postures of small metal tools. Furthermore, the feature map F after spatial transformation is... stn An efficient attention computation method can be used to enhance the defect region features of the corrected feature map, and the final output is the attention-enhanced feature map F after viewpoint correction. att .
[0087] Accordingly, the implementation steps may include: using the spatial transformation network embedded in the region attention module to autonomously learn the affine transformation matrix based on the intermediate feature map, and performing rotation, translation and scaling correction on the intermediate feature map based on the affine transformation matrix; and using an efficient attention calculation method to enhance the defect region features of the corrected feature map to obtain the attention-enhanced feature map after viewpoint correction.
[0088] By performing affine transformation correction on the intermediate feature map through a spatial transformation network and combining it with region attention enhancement, the influence of geometric distortion caused by shooting angle shift can be effectively eliminated, and various perspective features can be uniformly standardized into standard positive perspective features. At the same time, the feature weight of defect areas can be significantly improved, enhancing the model's feature matching ability and recognition stability for small metal objects with different shooting angles.
[0089] Step 230-4: Decompose the attention enhancement feature map into multiple scales according to different receptive fields to obtain multi-scale feature maps adapted to different shooting angles.
[0090] Here, the receptive field is the range of image regions that the convolutional neural network can cover during feature extraction; multi-scale decomposition is the process of splitting the feature map into multi-level feature representations according to different receptive fields.
[0091] In this embodiment of the disclosure, the attention-enhanced feature map, after viewpoint correction and attention enhancement, can be input into the feature output structure of the target detection model. The feature map is then decomposed hierarchically according to different receptive fields to generate feature maps corresponding to small-scale, medium-scale, and large-scale targets, forming a multi-scale feature map that contains multi-level receptive field information and can adapt to various shooting angles. For example, if small-scale, medium-scale, and large-scale targets correspond to 80×80, 40×40, and 20×20 respectively, then feature maps F1, F2, and F3 at three scales of 80×80, 40×40, and 20×20 can be output.
[0092] Decomposing the attention-enhanced feature map into multiple scales according to different receptive fields enables the model to simultaneously cover small hardware defect targets of different sizes, improves the perception ability of defects of various scales, maintains stable adaptability to targets from different shooting angles, enhances the comprehensiveness and robustness of feature representation, and provides reliable multi-scale feature support for subsequent occluded area recognition and feature completion.
[0093] Step 240: Perform gradient change calculation and connected component analysis on the multi-scale feature map to identify the occlusion location and occlusion ratio of the occlusion area corresponding to the small hardware.
[0094] Among them, connected component analysis is an analysis method that aggregates consecutive pixels of the same type into a whole region to achieve region localization and quantization; occlusion region is the image region where the small fitting is covered by wires or insulators, resulting in feature loss; occlusion position is the coordinate range of the occlusion region in the feature map; occlusion ratio is the ratio of the area of the occlusion region to the area of the whole region of the small fitting.
[0095] For the embodiments of this disclosure, automatic identification and quantization of occluded regions can be achieved by calculating the pixel gradient change rate and performing connected component analysis on a multi-scale feature map F. The core steps are as follows: Calculate the pixel gradient change rate: Use the Sobel operator to calculate the horizontal and vertical gradients of the feature map, obtaining the gradient magnitude map G, as shown in the formula: ; Among them, G x The horizontal gradient feature map obtained by convolving the input feature map F with the horizontal Sobel operator is used to extract the horizontal edge and contour information of small hardware defects; G y The vertical gradient feature map obtained by convolving the input feature map F with the vertical Sobel operator is used to extract the vertical edge and contour information of small hardware defects; Sobel x Sobel y For horizontal and vertical Sobel operators; F is the input feature map to be used for edge detection (usually F after multi-scale fusion). conv or F after spatial transformation stnG is the synthesized gradient magnitude feature map, which integrates gradient information in the horizontal and vertical directions to fully characterize the edge strength and contour structure of small hardware defects, and is used for feature compensation and edge enhancement of occluded areas. , The result of squaring the horizontal and vertical gradient feature maps is used to eliminate the positive and negative differences in gradient direction, retaining only edge intensity information.
[0096] Connectivity Analysis and Quantization: Setting the Gradient Threshold G th =0.1, G <G th The pixels are identified as occlusion candidate points. Through 8-neighbor connected component analysis, consecutive occlusion candidate points are merged into a complete occlusion region. The position coordinates (x0, y0) and area S of each occlusion region are output. The occlusion ratio is calculated based on the region coordinates and area. Finally, the occlusion position and occlusion ratio η of the occlusion region corresponding to the small fitting are determined (occlusion ratio = occlusion region area / total area of small fitting region).
[0097] By calculating gradient changes and analyzing connected components, the location of occluded areas can be accurately identified and the occlusion ratio can be quantified, providing a reliable basis for subsequent hierarchical feature completion. This effectively solves the problem of feature loss caused by occlusion, improves the completeness of the model's feature perception of small fittings in complex field environments, and reduces the risk of missed detection caused by occlusion.
[0098] Step 250: Perform hierarchical feature completion on the occluded area according to the occlusion ratio to obtain the completed multi-scale feature map.
[0099] For embodiments of this disclosure, step 250 may include the following steps: Step 250-1: When the occlusion ratio is not higher than the preset ratio threshold, the small hardware structure features in the occluded area are completed by bilinear interpolation using the surrounding effective defect features to obtain the interpolation completion features of the occluded area.
[0100] Among them, the preset ratio threshold is a pre-set critical value used to distinguish between mild and severe occlusion, such as 50%; the effective surrounding defect features are the feature information of small fittings that are not occluded and have complete structures near the edge of the occluded area; the structural features of the small fittings are the inherent morphological features of the small fittings themselves, such as pins, bolts, and nuts; bilinear interpolation completion is a method of filling missing areas with features by using effective surrounding features through bilinear interpolation; the interpolation completion features are the structural feature data of the occluded area calculated by bilinear interpolation.
[0101] In this embodiment of the disclosure, under the condition of mild occlusion where the occlusion ratio is no higher than a preset threshold (e.g., η≤50%), the effective defect features around the occluded area are used as the calculation basis. A bilinear interpolation formula is employed to fill and restore the structural features of the small hardware within the occluded area. The formula uses the values of known effective feature points in the surrounding area for weighted calculation. Through continuous interpolation, a complete and smooth feature distribution is obtained, ultimately generating interpolated features capable of filling feature gaps. The bilinear interpolation formula is as follows: ; Among them, F inter (x,y) is the interpolated feature value at coordinate (x,y) obtained after bilinear interpolation calculation, which is the feature data after smoothing and high-fidelity filling of the occluded area or target position; (x,y) is the target pixel coordinate to be calculated, representing the position where the feature value needs to be obtained. For the double summation operation, it iterates through all four integer pixels (i,j) within a 2×2 neighborhood window centered at the target position (x,y) to achieve weighted aggregation of neighborhood features; F(x i ,y j ) represents the integer coordinates (x, y) within the neighborhood window. i ,y j The original feature values at () provide the basic data for interpolation calculations; These are two-way weighting coefficients, representing the target position (x, y) and the neighboring pixel (x, y), respectively. i ,y j The weights are inversely proportional to the distance in the horizontal and vertical directions. The closer the distance, the closer the weight is to 1 and the greater the contribution; the farther the distance, the smaller the weight.
[0102] In scenarios with slight occlusion, bilinear interpolation is used to complete the model by using surrounding effective defect features. This method can quickly and accurately restore the structural features of small fittings in the occluded area, maintain the continuity and integrity of the features, avoid feature breakage caused by local occlusion, and improve the model's stability and detection accuracy for targets with slight occlusion.
[0103] Step 250-2: When the occlusion ratio is higher than the preset ratio threshold, a lightweight generation network is used to restore and complete the structural features of the small fittings in the occluded area, so as to obtain the generated and completed features of the occluded area.
[0104] Among them, restoration and completion is the process of reconstructing and restoring the missing features caused by occlusion; generation and completion features are the complete structural features generated by a lightweight generative network to fill in the occluded areas.
[0105] In specific application scenarios, a lightweight generative adversarial network (Light-GAN) can be constructed as a feature compensator. The Light-GAN is trained with the structural features of unoccluded small hardware (cylinder pin, threaded bolt, hexagonal nut) as positive samples, and a differentiated feature completion strategy is adopted according to the occlusion ratio η.
[0106] In the embodiments of this disclosure, under the condition of severe occlusion where the occlusion ratio is higher than a preset ratio threshold (e.g., η>50%), a pre-trained lightweight generation network is applied to the occlusion area processing. Using the structural features of the unoccluded small hardware as a reference, the missing structural features of the small hardware are inferred and reconstructed as a whole, and the complete structural information that matches the surrounding features is directly generated to complete the feature restoration and completion of the occluded area, and finally the generated and completed features of the occluded area are obtained.
[0107] In heavily occluded scenarios, using a lightweight generative network for feature restoration and completion can effectively overcome the limitations of traditional interpolation methods in recovering complex structures, fully restore the inherent features of small hardware that are heavily occluded, ensure the authenticity and completeness of feature representation, significantly reduce the false negative rate of defect detection under high occlusion conditions, and improve the robustness of the model in complex field environments.
[0108] Step 250-3: Concatenate the interpolated or generated complete features with the non-occluded region features in the original multi-scale feature map to obtain the completed multi-scale feature map.
[0109] Among them, the original multi-scale feature map is the initial multi-scale feature data after angle adaptive extraction without occlusion completion; the non-occluded region features are the feature information in the original multi-scale feature map that is not covered and remains complete and effective; the stitching is the operation of integrating the completed features with the original effective features in the same spatial position; the completed multi-scale feature map is the result of multi-scale feature expression after occlusion region repair and feature completeness without loss.
[0110] In this embodiment of the disclosure, the corresponding interpolation completion feature or the generated completion feature can be selected according to the degree of occlusion. The completion feature and the unoccluded region feature that is preserved in the original multi-scale feature map are aligned and stitched in the same feature space, so that the missing region is filled by the completion feature, and the effective region retains the original feature information, forming a feature distribution with continuous structure and complete information, and finally obtaining the completed multi-scale feature maps F1', F2', and F3'.
[0111] By stitching together the completed features with the features of the unoccluded regions, the feature loss caused by occlusion can be completely repaired, enabling the multi-scale feature map to restore all structural information of the small fittings, improving the integrity and consistency of the features, providing a high-quality, defect-free feature foundation for subsequent cross-scale fusion and defect feature enhancement, and further reducing the risk of missed detections and false detections caused by occlusion.
[0112] Step 260: Perform cross-scale weighted fusion on the completed multi-scale feature map, and enhance the defect features by combining spatial attention mechanism to obtain the fused feature map after the defect features are enhanced.
[0113] For embodiments of this disclosure, step 260 may include the following steps: Step 260-1: Assign corresponding scale weights to the completed multi-scale feature maps respectively, and perform weighted superposition processing on the feature maps of different scales according to the scale weights to obtain cross-scale fused feature maps.
[0114] Among them, scale weight is a pre-set weighting coefficient based on the contribution of different scale features to the detection of defects in small hardware fittings; weighted superposition processing is an operation of merging feature maps of different scales after numerical weighting according to the set weights; cross-scale fusion feature map is a comprehensive feature map that integrates feature information of multiple scales and takes into account both global structure and local details.
[0115] In this embodiment of the disclosure, corresponding scale weights can be assigned to the completed multi-scale feature maps F1', F2', and F3' respectively, and upsampling / downsampling can be performed according to the scale. After unifying to the same scale, pixel-level weighted fusion is performed, as shown in the formula: ; Among them, F fusion This is a cross-scale fused feature map; (x,y) is the coordinate of any pixel on the feature map; ω1, ω2, and ω3 are the scale weights corresponding to feature maps of different scales, which need to satisfy ω1+ω2+ω3=1, and the corresponding values can be: ω1=0.5, ω2=0.3, and ω3=0.2. F1'(x,y), F2'(x,y), and F3'(x,y) are the completed multi-scale feature maps. By weighted superposition, the feature information of different scales is integrated into a unified expression, and finally the cross-scale fused feature map is obtained.
[0116] By weighting and superimposing the completed multi-scale feature maps according to scale weights, the advantages of different scales can be fully integrated, the expression of key defect information can be strengthened, the overall representation ability of the feature maps can be improved, and more stable, comprehensive and accurate feature support can be provided for subsequent defect feature enhancement and detection and identification.
[0117] Step 260-2: Input the cross-scale fused feature map into the spatial attention module, and generate a spatial attention weight map through the spatial attention module.
[0118] Among them, the spatial attention module is a feature processing module used to enhance the target region and suppress background interference in the spatial dimension; the spatial attention weight map is generated by the spatial attention module and is used to characterize the weight distribution data of different spatial locations.
[0119] In this embodiment of the disclosure, a cross-scale fusion feature map that has been weighted and superimposed on multiple scales can be input into a spatial attention module. The spatial attention module calculates the feature response of the cross-scale fusion feature map in the spatial dimension and automatically generates a spatial attention weight map that matches the size of the cross-scale fusion feature map based on the difference in feature distribution. This is used to highlight the area where the small hardware defect is located and weaken invalid background information.
[0120] By inputting the cross-scale fused feature map into the spatial attention module and generating a spatial attention weight map, it is possible to accurately focus on defect-related regions in the spatial dimension, improve the expression intensity of key features and suppress redundant background interference, further enhance the recognizability and stability of defect features, and provide a reliable guarantee for subsequent high-precision defect classification and localization.
[0121] Step 260-3: Use the spatial attention weight map to perform weighted processing on the cross-scale fusion feature map to obtain the fusion feature map after the defect features are enhanced.
[0122] In the embodiments of this disclosure, the generated spatial attention weight map can be used to perform element-wise weighting processing on the cross-scale fusion feature map, according to formula F. enhance =F fusion M space Complete the calculation, where F enhance F is the fused feature map after enhancing the defect features. fusion M is a cross-scale fusion feature map. space This is a spatial attention weight map. To perform element-wise multiplication, a weighted processing method is used to enhance the feature response of the defect region and suppress the feature of the background region, ultimately resulting in a fused feature map with enhanced defect features.
[0123] By using spatial attention weight maps to weight the cross-scale fusion feature maps, the feature representation of the small fitting defect area can be significantly highlighted and the interference of invalid background information can be suppressed. This further enhances the recognizability and saliency of the defect features and effectively improves the accuracy and reliability of subsequent defect classification and localization.
[0124] Step 270: Input the fused feature map into the classification and regression head of the target detection model to identify the defect type of the small hardware and locate the coordinates of the defect location, and output the defect detection result.
[0125] Among them, the classification regression head is the network branch structure at the end of the target detection model used to perform category determination and location regression; the defect type identification is the process of classifying and determining the types of defects existing in the small hardware; the defect location coordinate localization is the process of determining the coordinate information of the area where the defect is located in the image and generating a detection box; the defect detection result is the final detection output information containing the defect category, confidence level and detection box coordinates.
[0126] In this embodiment of the present disclosure, the fused feature map after the defect features are enhanced can be input into the classification and regression head of the target detection model. The classification branch completes the identification and determination of the defect type of the small fittings based on the fully connected operation. The regression branch completes the accurate positioning of the defect location according to the coordinate offset formula. The detection information including defect category, confidence level and bounding box coordinates is output, and finally a complete defect detection result is obtained.
[0127] By inputting the fused feature map into the classification regression head for defect identification and localization, the effective features enhanced in the early stage can be fully utilized to achieve accurate differentiation of small fitting defect types and precise location of defects, thereby improving the integrity and reliability of defect detection in complex scenarios and ensuring the overall effect of small fitting defect detection in power transmission lines.
[0128] In summary, the small fitting defect detection method for transmission lines provided in this application can eliminate various lighting distortions by performing adaptive preprocessing on the on-site images of the small fittings. The method employs a gray wolf optimization algorithm improved with chaotic initialization, bidirectional information sharing search, and adaptive exponential decay convergence factor to optimize hyperparameters, effectively improving the problems of poor population diversity and susceptibility to local optima in traditional algorithms, thus enhancing the rationality of parameter configuration and feature extraction capabilities of the detection model. Simultaneously, it utilizes deformable convolution and spatial transformation networks to achieve angle-adaptive feature extraction, adapting to target deformation caused by changes in shooting angle. Furthermore, it compensates for and enhances defect features through occlusion region identification, hierarchical feature completion, cross-scale weighted fusion, and spatial attention enhancement. Finally, it relies on the optimized model to complete defect classification and localization. This method comprehensively solves the problems of missed detection, high false detection rate, and insufficient detection accuracy of small fitting defects in complex scenes caused by lighting interference, insufficient algorithm optimization, variable target angles, and occlusion effects. It can significantly improve the stability, accuracy, and environmental adaptability of small fitting defect detection.
[0129] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides a small fitting defect detection device for power transmission lines, such as... Figure 3 As shown, the device includes: a processing module 31, an optimization module 32, an extraction module 33, and a detection module 34; Processing module 31 can be used to perform illumination adaptive preprocessing on the on-site images of small hardware to obtain standardized images after eliminating illumination distortion; Optimization module 32 can be used to improve the gray wolf optimization algorithm by adopting chaotic information sharing. Through chaotic population initialization, bidirectional information sharing search and adaptive exponential decay convergence factor adjustment, the hyperparameters of the target detection model are optimized to obtain the target hyperparameter combination suitable for small hardware defect detection. Extraction module 33 can be used to configure the target hyperparameter combination into the target detection model and train the model. Based on the trained target detection model, it performs angle adaptive feature extraction of standardized images by combining deformable convolution and spatial transformation to obtain multi-scale feature maps that adapt to different shooting angles. The detection module 34 can be used to identify occluded areas, complete hierarchical features, and fuse cross-scale features on multi-scale feature maps to obtain a fused feature map with enhanced defect features. Based on the fused feature map, small hardware defects are classified and located to obtain defect detection results.
[0130] In some embodiments of this application, the processing module 31 can be specifically used to calculate the average brightness and contrast of the small hardware scene image. Based on the average brightness and contrast, the small hardware scene image is divided into three lighting types: strong light, weak light, and backlight. Differentiated combination correction strategies are used for different lighting types to eliminate lighting distortion and obtain a lighting corrected image. The differentiated combination correction strategies include: using polarized filtering and gamma compression correction for strong light type, using multi-scale enhancement and local contrast enhancement correction for weak light type, and using dark channel estimation and background fusion correction for backlight type. The lighting corrected image is then uniformly scaled to a standard size and pixel normalized to obtain a standardized image.
[0131] In some embodiments of this application, the optimization module 32 can specifically be used to generate a chaotic sequence using a logical mapping of a completely chaotic state, assign values to individuals in the population of the improved gray wolf optimization algorithm based on the chaotic sequence to generate an initial wolf population, and complete the initialization of the chaotic population; determine α, β, and δ-level guiding individuals and ordinary ω individuals based on the initial wolf population, forward transmit the position information of the α, β, and δ-level guiding individuals to the ordinary ω individuals to guide the ordinary ω individuals to update their positions, and backfeed the candidate positions searched by the ordinary ω individuals to the α, β, and δ-level guiding individuals to update the position information of the α, β, and δ-level guiding individuals, based on... The information interaction formed by positive propagation and negative feedback enables bidirectional information sharing search. An exponential decay base term is constructed based on the current iteration number and total iteration number of the algorithm. This base term is dynamically corrected by considering the change in the loss function between adjacent iterations, and the decay rate and convergence amplitude of the convergence factor are adjusted in real time to complete the adaptive exponential decay convergence factor adjustment. The iterative optimization process uses chaotic population initialization, bidirectional information sharing search, and adaptive exponential decay convergence factor adjustment as its framework. The detection accuracy of the target detection model is used as the optimization criterion, and the target hyperparameter combination is iteratively selected. This target hyperparameter combination includes at least the learning rate, anchor box aspect ratio, number of convolutional kernels, and attention weights.
[0132] In some embodiments of this application, the extraction module 33 can be specifically used to input a standardized image into the backbone network of the target detection model, extract features from the standardized image through deformable convolution in the backbone network, adaptively acquire the deformation features of the small metal fitting under different shooting angles, and obtain a deformable convolution feature map; perform channel-level splicing and fusion of the deformable convolution feature map and the multi-scale fixed convolution feature map to obtain an intermediate feature map that integrates multi-scale information and deformation information; input the intermediate feature map into the region attention module of the target detection model, use the spatial transformation network embedded in the region attention module to perform affine transformation correction on the intermediate feature map, normalize the non-orthogonal view features in the intermediate feature map to standard orthogonal view features, and output the attention-enhanced feature map after view correction; decompose the attention-enhanced feature map into multi-scale features according to different receptive fields to obtain multi-scale feature maps adapted to different shooting angles.
[0133] In some embodiments of this application, the extraction module 33 can be specifically used to utilize the spatial transformation network embedded in the region attention module to autonomously learn the affine transformation matrix based on the intermediate feature map, and perform rotation, translation and scaling correction on the intermediate feature map based on the affine transformation matrix; and use an efficient attention calculation method to enhance the defect region features of the corrected feature map to obtain the attention-enhanced feature map after viewpoint correction.
[0134] In some embodiments of this application, the detection module 34 can be specifically used to perform gradient change calculation and connected component analysis on the multi-scale feature map to identify the occlusion position and occlusion ratio of the corresponding occlusion area of the small fitting; perform hierarchical feature completion on the occlusion area according to the occlusion ratio to obtain the completed multi-scale feature map; perform cross-scale weighted fusion on the completed multi-scale feature map, and enhance the defect features by combining spatial attention mechanism to obtain the fused feature map after the defect features are enhanced; input the fused feature map into the classification regression head of the target detection model to identify the defect type of the small fitting and locate the coordinates of the defect position, and output the defect detection result.
[0135] In some embodiments of this application, the detection module 34 can be specifically used to: when the occlusion ratio is not higher than a preset ratio threshold, use surrounding effective defect features to perform bilinear interpolation to complete the small hardware structure features in the occluded area, obtaining interpolated complete features of the occluded area; when the occlusion ratio is higher than the preset ratio threshold, use a lightweight generation network to restore and complete the small hardware structure features in the occluded area, obtaining generated complete features of the occluded area; concatenate the interpolated complete features or generated complete features with the non-occluded area features in the original multi-scale feature map to obtain a completed multi-scale feature map; assign corresponding scale weights to the completed multi-scale feature map, and perform weighted superposition processing on the feature maps of different scales according to the scale weights to obtain a cross-scale fusion feature map; input the cross-scale fusion feature map into the spatial attention module, and generate a spatial attention weight map through the spatial attention module; use the spatial attention weight map to perform weighted processing on the cross-scale fusion feature map to obtain a fusion feature map after defect feature enhancement.
[0136] It should be noted that other corresponding descriptions of the functional units involved in the small fitting defect detection device for transmission lines provided in this embodiment can be found in [reference needed]. Figure 1 and Figure 2 The corresponding descriptions in [the document] will not be repeated here.
[0137] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method for detecting defects in small fittings of the transmission line is shown.
[0138] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause an electronic device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0139] Based on the above, Figure 1 and Figure 2 The method shown, and Figure 3 To achieve the above objectives, the present application also provides an electronic device, specifically a personal computer, tablet computer, server, or other network device, as shown in the virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 and Figure 2 The method for detecting defects in small fittings of the transmission line is shown.
[0140] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0141] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0142] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0143] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.
[0144] This invention eliminates various lighting distortions by performing adaptive preprocessing on images of small hardware fittings. It employs a gray wolf optimization algorithm improved with chaotic initialization, bidirectional information sharing search, and adaptive exponential decay convergence factor to optimize hyperparameters, effectively addressing the problems of poor population diversity and susceptibility to local optima in traditional algorithms. This enhances the rationality of parameter configuration and feature extraction capabilities of the detection model. Simultaneously, it utilizes deformable convolution and spatial transformation networks to achieve angle-adaptive feature extraction, adapting to target deformation caused by changes in shooting angle. Furthermore, it compensates for and enhances defective features through occlusion region identification, hierarchical feature completion, cross-scale weighted fusion, and spatial attention enhancement. Finally, it relies on the optimized model to complete defect classification and localization. This comprehensively solves the problems of high false negative rates, low detection accuracy, and high error rates in traditional detection schemes in complex scenes due to lighting interference, insufficient algorithm optimization, variable target angles, and occlusion effects. It significantly improves the stability, accuracy, and environmental adaptability of small hardware defect detection.
[0145] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.
[0146] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. A method for detecting defects in small fittings of power transmission lines, characterized in that, include: The on-site images of small hardware are subjected to illumination adaptive preprocessing to obtain standardized images after eliminating illumination distortion. The gray wolf optimization algorithm is improved by adopting chaotic information sharing. The hyperparameters of the target detection model are optimized by chaotic population initialization, bidirectional information sharing search and adaptive exponential decay convergence factor adjustment, so as to obtain the target hyperparameter combination suitable for small hardware defect detection. The target hyperparameter combination is configured into the target detection model and the model is trained. Based on the trained target detection model, the standardized image is subjected to angle adaptive feature extraction by combining deformable convolution and spatial transformation to obtain a multi-scale feature map that adapts to different shooting angles. The multi-scale feature map is subjected to occlusion region identification, hierarchical feature completion and cross-scale feature fusion processing to obtain a fused feature map with enhanced defect features. Based on the fused feature map, small hardware defects are classified and located to obtain defect detection results.
2. The method according to claim 1, characterized in that, The images of the small hardware were subjected to illumination-adaptive preprocessing to obtain standardized images after eliminating illumination distortion, including: Calculate the average brightness and contrast of the small hardware scene image, and based on the average brightness and contrast, divide the small hardware scene image into three lighting types: strong light, weak light, and backlight. Differentiated combination correction strategies are adopted for different lighting types to eliminate lighting distortion and obtain lighting-corrected images. The differentiated combination correction strategies include: polarized filtering and gamma compression correction for strong light lighting types, multi-scale enhancement and local contrast enhancement correction for weak light lighting types, and dark channel estimation and background fusion correction for backlighting types. The illumination-corrected image is scaled to a standard size and pixel normalized to obtain a standardized image.
3. The method according to claim 1, characterized in that, The improved gray wolf optimization algorithm using chaotic information sharing optimizes the hyperparameters of the target detection model through chaotic population initialization, bidirectional information sharing search, and adaptive exponential decay convergence factor adjustment, resulting in a target hyperparameter combination suitable for small hardware defect detection, including: A chaotic sequence is generated by logical mapping of a completely chaotic state. Based on the chaotic sequence, the population individuals of the improved gray wolf optimization algorithm for chaotic information sharing are assigned values to construct an initial wolf population and complete the initialization of the chaotic population. Based on the initial wolf pack population, α, β, and δ-level guiding individuals and ordinary ω individuals are determined. The position information of the α, β, and δ-level guiding individuals is forward-transmitted to the ordinary ω individuals to guide them to update their positions. The candidate positions searched by the ordinary ω individuals are then back-fed back to the α, β, and δ-level guiding individuals to update their position information. Based on the information interaction formed by the forward transmission and the back-feedback, a bidirectional information sharing search is performed. The exponential decay base term is constructed based on the current iteration number and the total iteration number of the algorithm. The exponential decay base term is dynamically corrected by combining the change in the loss function of adjacent iteration cycles. The decay rate and convergence amplitude of the convergence factor are adjusted in real time to complete the adaptive exponential decay convergence factor adjustment. The process of iterative optimization is based on the chaotic population initialization, the bidirectional information sharing search, and the adaptive exponential decay convergence factor adjustment. The detection accuracy of the target detection model is used as the optimization criterion. The target hyperparameter combination is iteratively selected. The target hyperparameter combination includes at least the learning rate, the aspect ratio of the anchor box, the number of convolutional kernels, and the attention weights.
4. The method according to claim 1, characterized in that, The trained target detection model performs angle-adaptive feature extraction on the standardized image using deformable convolution and spatial transformation, resulting in multi-scale feature maps adapted to different shooting angles, including: The standardized image is input into the backbone network of the target detection model. The deformable convolution in the backbone network is used to extract features from the standardized image. The deformation features of the small metal tool under different shooting angles are adaptively obtained to obtain the deformable convolution feature map. The deformable convolutional feature map is fused with the multi-scale fixed convolutional feature map at the channel level to obtain an intermediate feature map that integrates multi-scale information and deformation information. The intermediate feature map is input into the region attention module of the target detection model. The spatial transformation network embedded in the region attention module is used to perform affine transformation correction on the intermediate feature map, normalize the non-frontal view features in the intermediate feature map into standard frontal view features, and output the attention-enhanced feature map after view correction. The attention-enhanced feature map is decomposed into multi-scale features based on different receptive fields to obtain multi-scale feature maps adapted to different shooting angles.
5. The method according to claim 4, characterized in that, The step involves using the spatial transformation network embedded in the region attention module to perform affine transformation correction on the intermediate feature map, normalizing the non-orthogonal viewpoint features to standard orthogonal viewpoint features, and outputting a viewpoint-corrected attention-enhanced feature map, including: Using the spatial transformation network embedded in the region attention module, the intermediate feature map is autonomously learned to learn the affine transformation matrix, and rotation, translation and scaling corrections are performed on the intermediate feature map based on the affine transformation matrix. An efficient attention calculation method is used to enhance the defect region features of the corrected feature map, resulting in an attention-enhanced feature map after viewpoint correction.
6. The method according to claim 1, characterized in that, The process involves identifying occlusion regions, performing hierarchical feature completion, and fusing cross-scale features on the multi-scale feature map to obtain a fused feature map with enhanced defect features. Based on this fused feature map, small hardware defects are classified and located to obtain defect detection results, including: Gradient change calculation and connected component analysis are performed on the multi-scale feature map to identify the occlusion position and occlusion ratio of the occlusion area corresponding to the small hardware. Based on the occlusion ratio, hierarchical feature completion is performed on the occluded region to obtain a completed multi-scale feature map. The completed multi-scale feature map is subjected to cross-scale weighted fusion, and the defect features are enhanced by combining spatial attention mechanism to obtain the fused feature map with enhanced defect features. The fused feature map is input into the classification and regression head of the target detection model to identify the defect type of the small hardware and locate the defect position by coordinates, and output the defect detection result.
7. The method according to claim 6, characterized in that, The step of performing hierarchical feature completion on the occluded region according to the occlusion ratio to obtain a completed multi-scale feature map includes: When the occlusion ratio is not higher than a preset ratio threshold, the small fitting structure features in the occluded area are completed by bilinear interpolation using the surrounding effective defect features to obtain the interpolation completion features of the occluded area. When the occlusion ratio is higher than the preset ratio threshold, a lightweight generation network is used to restore and complete the small hardware structure features of the occluded area to obtain the generated and completed features of the occluded area. The interpolated or generated complete features are concatenated with the non-occluded region features in the original multi-scale feature map to obtain the completed multi-scale feature map. The step involves performing cross-scale weighted fusion on the completed multi-scale feature map, and enhancing the defect features using a spatial attention mechanism to obtain a fused feature map with enhanced defect features, including: The completed multi-scale feature maps are assigned corresponding scale weights, and the feature maps of different scales are weighted and superimposed according to the scale weights to obtain cross-scale fused feature maps. The cross-scale fused feature map is input into the spatial attention module, and the spatial attention module generates a spatial attention weight map. The cross-scale fusion feature map is weighted using the spatial attention weight map to obtain a fusion feature map with enhanced defect features.
8. A device for detecting defects in small fittings of power transmission lines, characterized in that, include: The processing module is used to perform illumination-adaptive preprocessing on the on-site images of small hardware fittings to obtain standardized images after eliminating illumination distortion. The optimization module is used to improve the Grey Wolf optimization algorithm by adopting chaotic information sharing. Through chaotic population initialization, bidirectional information sharing search and adaptive exponential decay convergence factor adjustment, the hyperparameters of the target detection model are optimized to obtain the target hyperparameter combination suitable for small hardware defect detection. The extraction module is used to configure the target hyperparameter combination into the target detection model and train the model. Based on the trained target detection model, the module performs angle adaptive feature extraction on the standardized image by combining deformable convolution and spatial transformation to obtain a multi-scale feature map that adapts to different shooting angles. The detection module is used to perform occlusion region identification, hierarchical feature completion and cross-scale feature fusion processing on the multi-scale feature map to obtain a fused feature map with enhanced defect features, and to classify and locate small hardware defects based on the fused feature map to obtain defect detection results.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
10. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.