Iron tower screw state detection method and system based on multi-modal image fusion
By using a multimodal image fusion and cloud-edge collaborative closed-loop feedback mechanism, the problems of poor environmental adaptability and single identification dimension in tower inspection are solved. This enables multi-dimensional assessment of screw status and efficient and accurate positioning, improving the safety and efficiency of tower inspection and protecting data privacy.
Patent Information
- Application Number
- CN202511920097.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-20
AI Technical Summary
Existing tower inspection technologies suffer from poor environmental adaptability, limited identification dimensions, insufficient model generalization ability, and a lack of feedback mechanisms. These issues result in low inspection efficiency, poor accuracy, and data privacy risks, failing to meet the dual requirements of tower structural safety and functional assurance.
By employing multimodal image fusion technology and combining visible light and infrared images, a tower screw recognition model is constructed through image preprocessing and feature extraction. The model is then optimized by combining multi-source information to assess screw status and locate physical coordinates, and a closed-loop feedback mechanism for cloud-edge collaboration is established.
Ensuring clear extraction of screw features in harsh environments enables multi-dimensional status assessment, reduces false positive rates, improves detection speed and positioning accuracy, protects data privacy, adapts to inspection needs in different areas, and continuously optimizes model performance.
Smart Images

Figure CN121707982A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of power system inspection, and particularly relates to a tower screw state detection method and system based on multi-modal image fusion. BACKGROUND
[0002] Iron towers are critical infrastructure for both power transmission networks and communication networks, and their structural safety is of utmost importance. Power towers ensure the reliability of power transmission, while communication towers (including 5G / 6G base station towers, etc.) ensure stable coverage of communication signals. The safe operation of both types of towers highly depends on the integrity of the connecting components - screws. Once a screw is missing, loose, or corroded, it may cause component displacement, tower tilting, or even collapse for power towers, leading to large-scale power outages. For communication towers, it may cause structural safety hazards and equipment signal interruption, directly affecting communication service quality. With the expansion of China's power grid coverage and the extension of communication networks to the entire territory, the total number of power and communication towers has increased dramatically and the distribution scenarios have become increasingly complex: power towers are mostly located in mountainous areas, suburban areas, and across rivers, while communication towers are more widely distributed on city rooftops, densely populated residential areas, and remote border areas. Traditional tower inspection methods still rely mainly on manual climbing detection, which has significant limitations for both types of towers: first, high-altitude work is risky, especially in mountainous terrain or complex urban environments, where safety hazards are more prominent; second, detecting a single tower takes up to 1-2 hours, which cannot meet the "large-scale, high-frequency" inspection needs of millions of towers; finally, manual detection results are greatly influenced by subjective experience, and the accuracy of judging the corrosion pattern and slight looseness of screws is low, which may lead to missed detection and misjudgment, making it difficult to provide reliable protection for tower safety. To improve inspection efficiency, the industry has gradually introduced machine vision technology to assist in screw detection. Existing technologies mainly use industrial cameras to capture tower images and use image processing algorithms to identify screw status, but still have the following limitations and are not well adapted to both power and communication towers: 1. Poor environmental adaptability: existing systems mostly rely on a single visible light image, which results in increased image noise and blurred features in insufficient light scenarios such as overcast, foggy, and nighttime, or in adverse weather conditions such as rain and dust, leading to a significant decrease in recognition accuracy. In addition, communication towers are easily disturbed by direct light at night in cities, and power towers are easily obstructed by vegetation in mountainous areas, leading to uneven lighting and further increasing the difficulty of identification. 2. Single recognition dimension: most systems can only implement binary judgment of "existence" or "absence" of screws, and cannot further assess their fastening status (such as looseness) and corrosion degree. However, for power towers, loose force-bearing component screws may cause structural deformation, and for communication towers, loose screws connecting equipment may cause poor contact and signal attenuation. This single identification mode cannot meet the dual refined control needs of "structural safety + functional protection" for towers. 3. Insufficient model generalization ability and data privacy risks: existing AI models are mostly based on training samples of specific scenarios and specific screw models, and do not fully cover the differences in screws for both types of towers (such as high-strength hexagonal screws for power towers and special models such as cross and internal hexagonal screws for communication towers) and complex background interference. The performance of the model decreases dramatically when it faces new types of screws, different corrosion patterns, or unfamiliar scenarios, and it needs to be retrained with a large number of samples, which is costly and time-consuming.Meanwhile, the iron tower inspection data (such as the line layout of the power iron tower and the base station position of the communication iron tower) has high sensitivity, and direct transmission of the original image has a risk of leakage.4. Lack of feedback mechanism and low positioning accuracy: the existing system lacks an effective closed-loop feedback mechanism, and the missed detection and misjudgment cases found by manual review are difficult to efficiently feed back to model optimization, which causes the model to be unable to adapt to the dynamically changing inspection scene, and there is a common problem of "performance decay after deployment". In terms of positioning, the existing technology mainly relies on single GPS coordinate conversion and does not combine the three-dimensional structural characteristics of the iron tower, so the positioning error often exceeds 10 cm. Maintenance personnel need to spend time checking the target screw in the complex tower structure, which seriously affects the maintenance efficiency. SUMMARY
[0003] To solve the above problems, the application provides an iron tower screw state detection method and system based on multi-modal image fusion, to solve the problems of poor environmental adaptability of the prior art, the limitation of "sky detection", the inability to evaluate the fastening state (looseness) and the corrosion degree, and the difficulty in meeting the fine safety management requirements of the iron tower.
[0004] An iron tower screw state detection method based on multi-modal image fusion, comprising: obtaining multi-modal component images of a target iron tower; performing image preprocessing and feature extraction on the multi-modal component images to obtain multi-scale feature maps; constructing and training an iron tower screw recognition model based on the multi-scale feature maps; performing multi-level detection and state evaluation on the obtained component images based on the iron tower screw recognition model; locating the missing screws in physical coordinates in combination with multi-source information and detection results, and establishing a cloud-edge collaborative closed-loop feedback mechanism to update and optimize the model.
[0005] According to an embodiment of the application, obtaining multi-modal component images of a target iron tower further comprises: using an industrial camera and an infrared thermal imager to collect multi-modal component images of the target iron tower, including visible light images and infrared images, and synchronously collecting the GPS coordinates, UTC time stamp and device attitude angle of the device, the device attitude angle including the pitch angle and the yaw angle.
[0006] According to an embodiment of the application, performing image preprocessing and feature extraction on the multi-modal component images to obtain multi-scale feature maps further comprises: denoising, contrast enhancement and angle correction are performed on the obtained multi-modal component images; multi-dimensional feature extraction is performed on the processed images to obtain geometric, grayscale, texture and spectral multi-dimensional features.
[0007] According to an embodiment of the present application, the denoising, contrast enhancement and angle correction of the obtained multi-modal component image further comprises: Based on the dynamic Gaussian filtering algorithm, the obtained multi-modal component image is denoised; The CLAHE algorithm is used to perform contrast enhancement processing on the denoised image, and the RGB image is converted into a grayscale image; The SIFT algorithm is used to perform image angle correction processing on the converted grayscale image.
[0008] According to an embodiment of the present application, the image preprocessing of the obtained multi-modal component image further comprises: The blind deconvolution algorithm is used to repair the motion blur image in the multi-modal component image, wherein the motion blur image is an image with a gradient variance <100.
[0009] According to an embodiment of the present application, the image preprocessing of the obtained multi-modal component image further comprises: The U-Net rain line detection network is used to identify the rain line area in the image and remove the rain line, and the calculation formula for removing the rain line is:
[0010] wherein, is a rain line removal function, I(x, y) is an original grayscale value, is a kth rain line template, is an intensity coefficient, is a decay factor, is the Euclidean distance.
[0011] According to an embodiment of the present application, the multi-dimensional feature extraction of the processed image to obtain geometric, grayscale, texture and spectral multi-dimensional features further comprises: The improved Canny algorithm is used to perform edge detection on the processed image, and the Hough circle transformation and YOLOv8-nano model are combined to generate a screw candidate area; Based on the screw candidate area, the geometric features, grayscale features, texture features and spectral features of the image are extracted, the geometric features include area, perimeter and center coordinates; the grayscale features are the average grayscale; the texture features include LBP histogram, GLCM contrast and entropy value; the spectral feature is the spectral angle , and the calculation formula is:
[0012] wherein, is a to-be-detected spectral vector, is a standard metal screw vector, and N is the number of wavebands.
[0013] According to a specific embodiment of the present invention, constructing and training a tower screw recognition model based on multi-scale feature maps further includes: A tower screw recognition model was established based on the improved YOLOv8 architecture; Input the training set and combine a phased strategy and federated learning framework to train the model in stages, and output the probability of screw presence, the probability of missing screw, and the region confidence. The backbone network of the tower screw recognition model uses a lightweight CSPDarknet-53 and embeds an SE attention mechanism. The neck network adopts a PAN-FPN structure and fuses multi-scale feature maps. The output layer uses an improved Softmax activation function.
[0014] According to a specific embodiment of the present invention, the multi-level detection and state assessment of the acquired component images based on the tower screw recognition model further includes: The component image is input into the model and a preset sliding window is used to perform first-level detection on the component image to generate screw candidate regions; Perform secondary detection on the candidate screw region and output the probability of screw presence; The presence or absence of a screw is determined by the probability of its presence; if it exists, the condition of the screw is evaluated.
[0015] According to a specific embodiment of the present invention, determining whether a screw exists based on its probability of presence, and if it exists, further includes evaluating the screw's condition: Let P be the probability of a screw's presence. If P ≥ 0.85, the screw is considered to be present; if P < 0.6, the screw is considered to be missing. When a screw is present, a sub-pixel edge detection algorithm is used to extract the edge contours of the screw head and the nut, and the concentricity deviation δ between the two is calculated. The calculation formula is as follows:
[0016] in, , The coordinates of the center of the screw head are: , Here are the coordinates of the nut's center; Combining the temperature distribution T and concentricity deviation δ acquired by the infrared thermal imager, when δ>2 and the local temperature difference... If the temperature is above 5℃, the screw is considered loose.
[0017] According to a specific embodiment of the present invention, determining whether a screw exists based on the probability of its presence further includes: If 0.6 ≤ P < 0.85, the image is marked as a suspected image, and multiple frames of continuous verification are performed on the suspected image.
[0018] According to a specific embodiment of the present invention, evaluating the condition of the screws further includes: Fourier transform is used to extract the angular features of the torque marking line on the screw head. If the angular deviation n>5°, the loosening judgment is strengthened.
[0019] According to a specific embodiment of the present invention, the physical coordinate positioning of the missing screw by combining multi-source information and detection results further includes: Combining the detection results and the collected multi-source information data, the pixel coordinates in the component image are converted into physical coordinates through perspective transformation. The multi-source information data includes GPS coordinates, device attitude angles, and camera lens parameters. The SfM algorithm is used to construct local 3D point clouds from continuous image sequences, and the ICP algorithm is used to match and correct the physical coordinates of the screws with the local 3D point clouds to generate a location information report. The location information report includes the screw location coordinates, quantity, confidence level, and corresponding component number.
[0020] According to a specific embodiment of the present invention, establishing a closed-loop feedback mechanism for cloud-edge collaboration to update and optimize the model further includes: A cloud-edge collaborative framework is established, and detection results and manually combined data are uploaded to edge devices. When the number of images stored in the cloud reaches a preset threshold, incremental training is started, and causal reasoning and knowledge distillation techniques are used to automatically update and optimize the model. The optimized new model is then pushed to the edge device via OTA.
[0021] According to a specific embodiment of the present invention, after acquiring the multimodal component image of the target tower, the method further includes: Spatiotemporal alignment of the acquired multimodal component images is performed. The acquisition times of visible light and infrared images are synchronized by timestamps. The image offset is calculated by ORB feature point matching algorithm. The infrared image is pixel aligned by affine transformation. Dynamic weight coefficients are calculated based on a dynamic weight fusion strategy, and image fusion is performed on multimodal component images. The formula for calculating the dynamic weight coefficients is as follows:
[0022] In the formula, For dynamic weighting coefficients, Let be the entropy of the visible light image at time t. Let be the variance of the visible light image information entropy at time t. Let be the entropy of the infrared image information at time t. Let be the variance of the infrared image information entropy at time t. For time windows.
[0023] A tower screw condition detection system based on multimodal image fusion includes: The image acquisition module is used to acquire multimodal component images of the target tower; The image processing and feature extraction module is used to perform image preprocessing and feature extraction on multimodal component images to obtain multi-scale feature maps; The model building and training module is used to build and train a tower screw recognition model based on multi-scale feature maps; The detection and evaluation module is used to perform multi-level detection and status evaluation on the acquired component images based on the tower screw recognition model; The positioning module is used to locate the physical coordinates of the missing screw by combining multi-source information and detection results; The update and optimization module is used to establish a closed-loop feedback mechanism for cloud-edge collaboration to update and optimize the model.
[0024] According to a specific embodiment of the present invention, the image processing and feature extraction module further includes: The image processing module is used to denoise, enhance contrast, and correct angles on the acquired multimodal component images; The multi-dimensional feature extraction module is used to extract multi-dimensional features from the processed image, obtaining geometric, grayscale, texture, and spectral features.
[0025] According to a specific embodiment of the present invention, the image processing module further includes: The image denoising module is used to denoise the acquired multimodal component images based on the dynamic Gaussian filtering algorithm; The contrast enhancement module is used to enhance the contrast of the denoised image using the CLAHE algorithm and convert the RGB image to a grayscale image. The correction module is used to perform image angle correction processing on the converted grayscale image using the SIFT algorithm; The blurred image restoration module is used to restore motion-blurred images in multimodal component images using a blind deconvolution algorithm, wherein the motion-blurred image is an image with a gradient variance < 100; The rain line removal module is used to identify rain line regions in images and remove rain lines using the U-Net rain line detection network.
[0026] According to a specific embodiment of the present invention, the multi-dimensional feature extraction module further includes: The candidate region generation module is used to perform edge detection on the processed image using an improved Canny algorithm, and to generate screw candidate regions by combining Hough circle transform and YOLOv8-nano model. The feature extraction module is used to extract geometric features, grayscale features, texture features, and spectral features from the candidate screw region of the image. Geometric features include area, perimeter, and center coordinates; grayscale features include grayscale mean; texture features include LBP histogram, GLCM contrast, and entropy value; spectral features include spectral angle. .
[0027] According to a specific embodiment of the present invention, the model building and training module further includes: The model building module is used to build a tower screw recognition model based on the improved YOLOv8 architecture; The training module is used to input the training set and combine a phased strategy and federated learning framework to train the model in stages, and output the probability of screw presence, the probability of missing screws, and the confidence of the region. The backbone network of the tower screw recognition model uses a lightweight CSPDarknet-53 and embeds an SE attention mechanism. The neck network adopts a PAN-FPN structure and fuses multi-scale feature maps. The output layer uses an improved Softmax activation function.
[0028] According to a specific embodiment of the present invention, the detection and evaluation module further includes: The primary detection module is used to input component images into the model and perform primary detection on the component images using a preset sliding window to generate screw candidate regions. The secondary detection module is used to perform secondary detection on the screw candidate area and output the probability of the screw's presence. The evaluation module is used to determine whether a screw exists based on the probability of its presence, and if it does exist, to evaluate the screw's condition.
[0029] According to a specific embodiment of the present invention, the positioning module further includes: The coordinate transformation module is used to combine the detection results and the collected multi-source information data to convert the pixel coordinates in the component image into physical coordinates through perspective transformation. The multi-source information data includes GPS coordinates, device attitude angles, and camera lens parameters. The 3D point cloud matching module is used to construct local 3D point clouds from continuous image sequences using the SfM algorithm, and to match and correct the physical coordinates of the screws with the local 3D point clouds using the ICP algorithm, generating a location information report. The location information report includes the screw location coordinates, quantity, confidence level, and corresponding component number.
[0030] An electronic device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-described method for detecting the state of iron tower screws based on multimodal image fusion.
[0031] A computer-readable storage medium storing a computer program, which is loaded and executed by a processor to implement the above-described method for detecting the state of iron tower screws based on multimodal image fusion.
[0032] Compared with the prior art, this application has the following advantages: 1. This invention effectively overcomes the bottleneck of traditional single-modal image acquisition and fusion technology, which is limited by environmental conditions. The dynamic switching and fusion of visible light and infrared images can clearly preserve screw features under adverse conditions such as insufficient light, rain, and fog, avoiding feature loss due to environmental interference. Simultaneously, the introduction of material spectral features and three-dimensional morphological features can accurately distinguish metal screws from background interference objects such as plastic and stones, significantly reducing the misjudgment rate in complex backgrounds. Whether in sunny, cloudy, nighttime, or rainy weather, the system can stably extract key information such as screw edges and textures, ensuring the consistency and reliability of the recognition results and overcoming the limitations of existing "sky-based detection" technologies.
[0033] 2. Unlike existing technologies that can only determine whether a screw is "present" or "missing," this invention adds a screw tightness assessment submodule. It calculates the concentricity deviation between the screw head and nut using sub-pixel edge detection, combines this with temperature distribution data collected by an infrared thermal imager to determine looseness, analyzes the integrity of torque marking lines to enhance the assessment, and can also identify the degree of screw corrosion, forming a multi-dimensional status assessment system of "present-missing-loose-corroded." This refined identification capability helps inspection personnel not only discover missing screws but also provide early warnings of potential risks such as loosening and corrosion, advancing tower safety management from "post-event maintenance" to "pre-event prevention," significantly improving the structural safety level of towers.
[0034] 3. To address the privacy protection needs of inspection data from different regions, this invention employs a federated learning framework. Edge devices only upload model parameters, not raw images, and collaborative training is achieved through encrypted parameter exchange, effectively avoiding the risk of data leakage. The global model uses a weighted average to aggregate sample information from each device, improving the model's adaptability to different regions and screw types without the need for centralized storage of massive amounts of data. Compared to traditional centralized training, this mode protects data privacy while allowing the model to quickly absorb inspection experience from various regions, significantly enhancing its generalization ability. It also eliminates the need for separate model training for each region, greatly reducing deployment costs.
[0035] 4. In terms of detection efficiency, this invention designs an attention-oriented dynamic search mechanism, which performs fine-grained searching in high-confidence areas and rapid filtering in low-confidence areas, significantly improving detection speed without reducing recognition accuracy. This meets the real-time requirements of UAV inspections, and the single-frame processing efficiency is significantly better than existing technologies. In terms of positioning accuracy, by integrating multi-source data such as GPS, equipment attitude angle, and LiDAR point cloud, and through multiple rounds of calibration such as perspective transformation and ICP algorithm, the positioning error of missing screws is controlled to an extremely low range. Maintenance personnel can directly find the target location based on the positioning results without repeatedly searching on the tower, greatly shortening maintenance time.
[0036] 5. This invention establishes a closed-loop feedback framework for cloud-edge collaboration. Edge devices upload missed and misjudged cases reviewed by manual verification in real time. The cloud, through incremental training and knowledge distillation techniques, rapidly absorbs new sample experience while retaining the performance of historical models. Monthly full retraining updates the model's adaptability to new interference factors. Simultaneously, the causal reasoning module analyzes the causes of errors and generates targeted training datasets, accelerating the model's learning of weak scenarios. This allows the model to continuously evolve with changes in inspection scenarios, maintaining high recognition performance over the long term and avoiding the problem of "performance degradation after deployment" in existing systems.
[0037] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a flowchart of a method for detecting the state of iron tower screws based on multimodal image fusion according to an embodiment of the present invention.
[0040] Figure 2 This is a flowchart of a method for acquiring multimodal component images of a target iron tower according to an embodiment of the present invention.
[0041] Figure 3 This is a flowchart of a method for image preprocessing and feature extraction of multimodal component images according to an embodiment of the present invention.
[0042] Figure 4This is a flowchart of a method for denoising, contrast enhancement, and angle correction of acquired multimodal component images according to an embodiment of the present invention.
[0043] Figure 5 This is a flowchart of a method for multi-dimensional feature extraction of a processed image according to an embodiment of the present invention.
[0044] Figure 6 This is a flowchart of a method for constructing and training a tower screw recognition model based on multi-scale feature maps according to an embodiment of the present invention.
[0045] Figure 7 This is a flowchart of a method for multi-level detection and status assessment of acquired component images based on a tower screw recognition model according to an embodiment of the present invention.
[0046] Figure 8 This is a flowchart of a method for determining the presence or absence of a screw based on the probability of its presence, and evaluating the state of the screw if it exists, according to an embodiment of the present invention.
[0047] Figure 9 This is a flowchart of a method for physically locating missing screws by combining multi-source information and detection results according to an embodiment of the present invention.
[0048] Figure 10 This is a bar chart comparing recognition accuracy in different scenarios according to an embodiment of the present invention.
[0049] Figure 11 This is a line graph showing the relationship between model iteration cycle and recognition performance according to an embodiment of the present invention.
[0050] Figure 12 This is a structural diagram of a tower screw status detection system based on multimodal image fusion according to an embodiment of the present invention.
[0051] Figure 13 This is a structural diagram of an image processing and feature extraction module provided according to an embodiment of the present invention.
[0052] Figure 14 This is a structural diagram of an image processing module provided according to an embodiment of the present invention.
[0053] Figure 15 This is a structural diagram of a multi-dimensional feature extraction module provided according to an embodiment of the present invention.
[0054] Figure 16 This is a structural diagram of a model building and training module provided according to an embodiment of the present invention.
[0055] Figure 17This is a structural diagram of a detection and evaluation module provided according to an embodiment of the present invention.
[0056] Figure 18 This is a structural diagram of a positioning module provided according to an embodiment of the present invention.
[0057] Figure 19 This is a schematic diagram of a computer device structure according to an embodiment of the present invention.
[0058] Figure label: 01-Image Acquisition Module; 02-Image Processing and Feature Extraction Module; 03-Model Building and Training Module; 04-Detection and Evaluation Module; 05-Localization Module; 06-Update and Optimization Module; 021 - Image processing module; 022 - Multi-dimensional feature extraction module; 0211 - Image Denoising Module; 0212 - Contrast Enhancement Module; 0213 - Correction Module; 0214 - Blurred Image Repair Module; 0215 - Rainline Removal Module; 0221 - Candidate Region Generation Module; 0222 - Feature Extraction Module; 031 - Model building module; 032 - Training module; 041 - Level 1 Detection Module; 042 - Level 2 Detection Module; 043 - Evaluation Module; 051 - Coordinate transformation module; 052 - 3D point cloud matching module. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0060] Example 1 Additional aspects and advantages of embodiments of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of embodiments of the invention. Figures 1-11 This invention provides a method for detecting the state of iron tower screws based on multimodal image fusion, including: S1: Obtain multimodal component images of the target tower.
[0061] S2: Perform image preprocessing and feature extraction on the multimodal component images to obtain multi-scale feature maps.
[0062] S3: Construct and train a tower screw recognition model based on multi-scale feature maps.
[0063] S4: Based on the tower screw recognition model, perform multi-level detection and status assessment on the acquired component images.
[0064] S5: Combine multi-source information and detection results to locate the physical coordinates of the missing screw.
[0065] S6: Establish a closed-loop feedback mechanism for cloud-edge collaboration to update and optimize the model.
[0066] Specifically, step S1, acquiring the multimodal component images of the target tower, further includes: S11: Uses industrial cameras and infrared thermal imagers to acquire multimodal component images of the target iron tower, including visible light images and infrared images.
[0067] S12: Synchronously collect the device's GPS coordinates, UTC timestamps, and device attitude angles, including pitch and yaw angles.
[0068] In one specific embodiment of the present invention, a dual-mode acquisition system is composed of a 5-20 megapixel industrial camera and an 8-12μm wavelength infrared thermal imager. The system has a built-in light sensor. When the light intensity is ≥300 lux, visible light mode is activated for shooting, and when the light intensity is <300 lux, it automatically switches to infrared mode for shooting. During image acquisition, the device's GPS coordinates (accuracy ±0.5m), UTC timestamp (error <1ms), and device attitude angles (pitch angle, yaw angle, measurement accuracy ±0.1°) are recorded simultaneously to ensure data spatiotemporal synchronization.
[0069] In a specific embodiment of the present invention, when acquiring multimodal image data of the tower surface, a dual-mode acquisition system is composed of an industrial camera (model optional Baslerac A2500-14uc) with 5-20 megapixels and an 8-12μm infrared thermal imager (resolution 320×240-640×512, temperature measurement range -20℃ to 150℃). The camera lens focal length is 8mm-25mm (aperture F1.4-F2.8, distortion rate <1%), the thermal imager frame rate is 15-30Hz, and the temperature measurement accuracy is ±2℃. Visible light imaging is enabled under illumination intensity of 300 lux-10000 lux. When the illumination intensity is <300 lux, it automatically switches to infrared mode. The shooting angle is 30°-60° to the surface of the tower (monitored in real time by a three-axis gyroscope, with an adjustment accuracy of ±1°). The image resolution is set to 2560×1920-4096×3072 (bit depth 12bit). 10-30 frames are acquired per second. The GPS coordinates of the shooting position (accuracy ±0.5m, sampling rate 1Hz), shooting time (UTC timestamp, error <1ms), and equipment attitude angles (pitch angle -30° to 30°, yaw angle -180° to 180°, measurement accuracy ±0.1°) are recorded simultaneously.
[0070] Specifically, after obtaining the multimodal component images of the target tower in step S1, the process also includes: Spatiotemporal alignment is performed on the acquired multimodal component images. The acquisition times of visible light and infrared images are synchronized by timestamps, and the image offset is calculated using the ORB feature point matching algorithm. Pixel alignment of the infrared image is performed by affine transformation.
[0071] Dynamic weight coefficients are calculated based on a dynamic weight fusion strategy, and image fusion is performed on multimodal component images. The formula for calculating the dynamic weight coefficients is as follows: (1) In the formula, For dynamic weighting coefficients, Let be the entropy of the visible light image at time t, ranging from 0 to 8. Let be the variance of the visible light image information entropy at time t, used to reflect the stability of illumination. Let be the entropy of the infrared image information at time t, which ranges from 0 to 6. Let be the variance of the infrared image information entropy at time t. The time window is typically set to 1-3 seconds and can be dynamically adjusted according to the actual situation.
[0072] In a specific embodiment of the present invention, the acquisition time of visible light and infrared images is synchronized by timestamp (time difference ≤ 10ms) in time. In space, ORB feature point matching (extracting ≥ 200 feature points, matching threshold 30) is used to calculate the offset. Pixel-level alignment is achieved by performing 6-DOF affine transformation on the infrared image (alignment error ± 1 pixel). The alignment error is controlled within ± 1 pixel. At the same time, a dynamic weight fusion strategy is introduced. The dynamic weight coefficient is calculated by formula (1) so that the fused image maintains feature stability when the illumination changes abruptly (change rate > 50% / s), and the edge retention rate is improved by ≥ 15%.
[0073] Specifically, step S2 involves image preprocessing and feature extraction of the multimodal component image to obtain a multi-scale feature map, which further includes: S21: Denoising, contrast enhancement, and angle correction are performed on the acquired multimodal component images, further including: S211: Denoise the acquired multimodal component images based on the dynamic Gaussian filtering algorithm.
[0074] S212: The CLAHE algorithm is used to enhance the contrast of the denoised image and convert the RGB image to a grayscale image.
[0075] S213: The SIFT algorithm is used to perform image angle correction processing on the converted grayscale image.
[0076] In one specific embodiment of the present invention, a dynamic Gaussian filtering algorithm is first used to denoise the acquired multimodal component images using Gaussian filtering. This involves dynamically adjusting the filter kernel size based on the noise variance to achieve Gaussian filtering denoising. Then, the CLAHE (Contrast-Limited Adaptive Histogram Equalization) algorithm is used to enhance image contrast, and the RGB image is converted to grayscale. Finally, SIFT feature points from the SIFT algorithm are used for image angle correction.
[0077] In a specific embodiment of the present invention, when performing multi-dimensional preprocessing on the acquired image, noise is first removed by Gaussian filtering. The size of the filter kernel is dynamically adjusted to 3×3-7×7 according to the image noise intensity (3×3 when the noise variance is <30, 5×5 when the noise variance is between 30 and 50, and 3×3 median filtering is enabled when the noise variance is >50, with 2 iterations). Then, the contrast-limited adaptive histogram equalization algorithm (CLAHE) is used to enhance the image contrast, with the contrast limit threshold set to 2. The image size is 0-4.0 (2.0 for cloudy days, 4.0 for sunny days). The tiles grid size is 8×8-16×16, while retaining 1%-2% of extreme gray values to avoid loss of detail. After converting the RGB image to grayscale using a weighted formula (Gray=0.299R+0.587G+0.114B), the shooting angle deviation is corrected by an image correction algorithm based on SIFT features (extracting ≥50 feature points, matching threshold 0.75). The correction accuracy is controlled within ±0.3°.
[0078] S22: A blind deconvolution algorithm is used to repair motion-blurred images in multimodal component images, wherein the motion-blurred images are images with a gradient variance of <100.
[0079] In a specific embodiment of the present invention, for motion-blurred images (i.e., images with gradient variance < 100), a blind deconvolution algorithm is used to repair the motion-blurred image in the multimodal component image and restore its sharpness. The number of iterations is set to 10-30 (increase the number of iterations if the blurriness is high), and the regularization parameter is set to 0.001-0.01 (take a larger value when there is a lot of noise).
[0080] S23: The U-Net rain line detection network is used to identify rain line regions in the image and remove the rain lines. The formula for removing rain lines is: (2) in, Here is the rain line removal function, where I(x, y) is the original grayscale value. This is the template for the k-th type of rain line (covering straight lines, oblique lines, and curves, K=3-5 types). This is the intensity coefficient, ranging from 0.1 to 0.8, which can be dynamically adjusted according to the brightness of the rain line. This is an attenuation factor, ranging from 0.01 to 0.1, used to control the transition at the edge of the rain line. The distance from a pixel to the center of the rain line is expressed in pixels.
[0081] In a specific embodiment of the present invention, an adaptive rain removal algorithm is added to the image preprocessing. In rainy scenes, for images taken on rainy days (identified by raindrop detection operator, rain line density > 5 lines / 100×100 pixels), a rain line detection network based on deep learning is used to identify the rain line region. Preferably, the U-Net rain line detection network architecture is used, with an input of 512×512 pixels. Rain line interference is removed by formula (2), which improves the signal-to-noise ratio of the rainy day image by more than 20%, and the screw region feature retention rate is ≥ 95%.
[0082] S24: Perform multi-dimensional feature extraction on the processed image to obtain geometric, grayscale, texture, and spectral multi-dimensional features, further including: S241: An improved Canny algorithm is used to perform edge detection on the processed image, and Hough circle transform and YOLOv8-nano model are combined to generate screw candidate regions.
[0083] S242: Extract geometric features, grayscale features, texture features, and spectral features of the image based on the screw candidate region.
[0084] In a specific embodiment of the present invention, an improved Canny algorithm is first used for edge detection. Combined with the Hough circle transform and the YOLOv8-nano model, candidate screw regions are generated, and the minimum bounding rectangle is output. A multi-feature fusion set is constructed by calculating the geometric features, grayscale features, texture features, and spectral features of the candidate regions. The geometric features include the area, perimeter, and center coordinates of the screw candidate regions; the grayscale features include the mean grayscale value of the screw candidate regions; the texture features include the LBP histogram, GLCM contrast, and entropy value; and the spectral features include the spectral angle, which is calculated by acquiring the reflectance spectrum curve of the screw region using a hyperspectral camera. The calculation formula is as follows: (3) In the formula, Let be the spectral vector to be detected, i.e., the reflectance of each band. For standard metal screw vectors (pre-stored database), N is the number of bands. When When the temperature is less than 5°, it is determined to be metallic and interference is excluded.
[0085] In a specific embodiment of the present invention, during the extraction of the multi-feature fusion set of the screw candidate region, an improved Canny edge detection algorithm is first adopted. The threshold range of 50-150 is dynamically set using a gradient orientation histogram (16 orientation boxes) (the threshold is increased when edge density is high). An 8-neighborhood search strategy is used for edge connections (connection threshold ≤ 2 pixels). Combining the Hough circle transform (accumulator threshold ≥ 100) with the candidate box generation mechanism of the YOLOv8-nano model, the circle detection radius is set to a range of 5-30 pixels (step size 1 pixel). Output the minimum bounding rectangle parameters (aspect ratio 0.8-1.2, area ratio 0.7-1.0), calculate the area (50-3000 pixels), perimeter (30-200 pixels), center coordinates (subpixel accuracy ±0.1 pixels), grayscale mean (50-200), and texture features of the candidate region. Texture features include LBP histogram (16-64 bins, sampling radius 1-3 pixels), GLCM contrast (0-200, distance 1-3 pixels, angle 0° / 45° / 90° / 135°), and entropy value (0-5).
[0086] In a specific embodiment of the present invention, material spectral features are introduced into the feature extraction process. The reflectance spectral curve of the screw region is acquired using a hyperspectral camera (wavelength 400-1000nm, number of bands ≥100, spectral resolution ≤5nm), and the spectral angle is calculated. ,when It can identify metal materials and exclude interference such as plastic and stones (reducing the false recognition rate by ≥30%), improving the feature discrimination in complex backgrounds, and is especially suitable for tower detection in vegetated areas.
[0087] Specifically, step S3, which involves constructing and training a tower screw recognition model based on multi-scale feature maps, further includes: S31: Establish a tower screw recognition model based on the improved YOLOv8 architecture.
[0088] In one specific embodiment of the present invention, a tower screw recognition model is established based on an improved YOLOv8 architecture. This model supports dual-channel input of visible light and infrared images. Its backbone network uses a lightweight CSPDarknet-53, thus reducing the number of convolutional kernels by 30%, and incorporates an SE attention mechanism. The neck network adopts a PAN-FPN structure and fuses multi-scale feature maps, and the output layer uses an improved Softmax activation function (introducing a temperature coefficient τ). By inputting multi-scale feature maps into this model, the presence probability, absence probability, and region confidence of the screw can be output.
[0089] In a specific embodiment of this invention, in the process of constructing a lightweight tower screw recognition model, an improved YOLOv8 architecture is adopted. The input layer supports dual-channel fusion of visible light and infrared images (channel weights can be dynamically adjusted). The backbone network adopts a lightweight version of CSPDarknet-53, using 1×1 convolution for dimensionality reduction, reducing the number of convolution kernels by 30%, and adding an attention mechanism (SE module, compression ratio 16) to enhance the response of the screw region. The neck network adopts a PAN-FPN structure, fusing feature maps at three scales: 80×80, 40×40, and 20×20. Upsampling and convolution are used for feature fusion. The output layer adopts an improved Softmax activation function, and a temperature coefficient τ=0.5-2.0 is added to the function to output the screw presence probability, missing probability, and region confidence (0-1). The model training uses the AdamW optimizer, with the initial learning rate set to 0.001-0.01 (adaptively adjusted with batch size), the weight decay coefficient set to 0.0001-0.001, and the momentum parameter set to 0.9-0.99.
[0090] S32: Input the training set and combine the phased strategy and federated learning framework to train the model in stages, and output the probability of screw presence, probability of missing screw, and region confidence.
[0091] In a specific embodiment of this invention, the AdamW optimizer is used to train the model. First, 10,000-50,000 labeled images of iron towers are divided into a training set and a validation set in an 8:2 ratio. Then, the training set is input into the iron tower screw recognition model for training. The training is divided into two stages: the first stage trains the basic model; the second stage introduces hard example mining, assigning higher weights to samples with high error rates. During training, accuracy, recall, and F1 score are calculated every 10 iterations. If the F1 score does not improve for 5 consecutive iterations, cosine annealing learning rate scheduling is enabled. This invention utilizes multimodal image fusion technology (visible light, infrared, hyperspectral, etc.) combined with a federated learning framework to construct a lightweight AI model with strong environmental robustness and high generalization ability while protecting data privacy, thereby reducing deployment costs and improving cross-regional adaptability.
[0092] In a specific embodiment of the present invention, during the phased training and dynamic optimization of the model, 10,000-50,000 labeled images containing screws and missing screws (covering screw models M10-M30) are collected. The images are divided into four categories according to the scene: sunny, cloudy, foggy, and nighttime. Each category accounts for 20%-30% of the samples. The training set and validation set are divided in an 8:2 ratio. Stratified sampling is used to maintain category balance. When labeling, the screw area is selected by polygonal box (labeling accuracy ±1 pixel) and marked with "present", "missing", and "rusted". The first stage trains the basic model, iterating 50-100 times with a batch size of 8-32 (which can be dynamically adjusted according to GPU memory). The second stage introduces hard example mining, assigning 2-3 times the weight to samples with an error rate >30% (such as highly reflective samples or partially obscured screws), iterating 30-50 times. During training, the validation set accuracy, recall, and F1 score (accurate to 0.1%) are calculated every 10 iterations. When the F1 score does not improve for 5 consecutive iterations, cosine annealing is used for learning rate scheduling (10 iterations per cycle), with a minimum learning rate of 0.00001-0.0001, until the model achieves an accuracy ≥95% and a recall ≥90% on the validation set.
[0093] In one specific embodiment of the present invention, a federated learning framework is employed during model training. For edge devices (typically 5-20 units) distributed across different tower inspection areas, collaborative model training is achieved through encrypted parameter exchange (using homomorphic encryption algorithm, 2048-bit key length). This improves model generalization ability while protecting data privacy, resulting in a ≥10% increase in cross-regional recognition accuracy. Loss function used in local training for: (4) In the formula, M is the number of local samples. Let be the label of the i-th sample (1 / 0). Let be the predicted probability of the i-th sample. λ is the regularization coefficient, λ = 0.001 - 0.01. This is an L2 regularization term.
[0094] Global model aggregation uses weighted average for: (5) In the formula, Let N be the sample size of the k-th device, and N be the total sample size. These are the parameters for the local model.
[0095] Specifically, step S4, based on the tower screw recognition model, further includes multi-level detection and state assessment of the acquired component images, including: S41: Input the component image into the model and use a preset sliding window to perform first-level detection on the component image to generate screw candidate regions.
[0096] S42: Perform secondary detection on the candidate screw region and output the probability of screw presence.
[0097] In a specific embodiment of the present invention, during the multi-level screw detection process on the component image, the preprocessed component image is first input into the trained AI model. The first-level detection uses a sliding window (step size 10-30 pixels, window size 50×50-200×200 pixels, adaptively adjusted according to the image scale) for rapid screening, filtering out more than 70% of non-screw areas and outputting screw candidate areas. The second-level detection performs fine detection on the screw candidate areas (inputting a 224×224 pixel cropped image) and outputs the probability of screw presence in each candidate area.
[0098] In a specific embodiment of the present invention, an attention-oriented dynamic search mechanism is added to the screw detection stage. A heatmap is generated based on the region confidence of the model output, and its resolution is consistent with the input image. The gradient divergence is calculated using formula (6). In the region with divergence value < 0 (attention convergence area, confidence ≥ 0.6), a fine search with a step size of 10 pixels is used, and in the region with divergence value ≥ 0 (low confidence area), a fast search with a step size of 30 pixels is used. This improves the detection speed by more than 40% while maintaining the accuracy (accuracy fluctuation ≤ 1%), meeting the real-time requirements of UAV inspection (single frame processing time < 200ms). The formula for calculating the gradient divergence is: (6) In the formula, The gradient vector of the heatmap is calculated using the Sobel operator. This represents the image gradient components calculated using the Sobel operator in the horizontal direction (xx axis). This represents the image gradient components calculated using the Sobel operator in the vertical direction (yy axis).
[0099] S43: Determine whether a screw exists based on its probability of presence; if it exists, evaluate its condition, including: S431: Let P be the probability of the screw's existence. If P ≥ 0.85, the screw is considered to exist. If P < 0.6, the screw is considered to be missing.
[0100] S432: When a screw is present, a sub-pixel edge detection algorithm is used to extract the edge contours of the screw head and the nut, and the concentricity deviation δ between the two is calculated. The calculation formula is as follows: (7) in, , The coordinates of the center of the screw head are: , The coordinates of the nut's center are in pixels.
[0101] S433: Combining the temperature distribution T and concentricity deviation δ acquired by the infrared thermal imager, when δ>2 and the local temperature difference... If the temperature is above 5℃, the screw is considered loose.
[0102] S434: When 0.6≤P<0.85, it is marked as a suspected image, and the suspected image is continuously verified over multiple frames.
[0103] In a specific embodiment of the present invention, the screw status is determined based on the existence probability output by the model: if the probability P ≥ 0.85, it is determined to exist; if the probability P < 0.6, it is determined to be missing; if the probability 0.6 ≤ P < 0.85, it is marked as a suspected image, and the suspected image is continuously verified for multiple frames (confirmation is required for 3-5 consecutive frames, and at least 3 frames must be consistent for final determination). Based on the identification of the screw's existence, the edge contours of the screw head and nut are extracted using a sub-pixel edge detection algorithm, and the concentricity deviation δ between the two is calculated. Combined with the temperature distribution collected by the infrared thermal imager, when δ > 2 pixels and the local temperature difference is significant, the screw's status is determined. If the temperature is above 5℃, the screw is considered loose.
[0104] In a specific embodiment of the present invention, a probability graph module is introduced into the screw missing determination process. The tower components are divided into 10×10cm² sub-regions, a Markov random field is constructed to describe the correlation of screw presence, and an energy function is defined. for: (8) In the formula, The single-point potential energy of the i-th screw is negatively correlated with the model's output probability. This indicates that it exists. The potential energy of the pairs of adjacent screws is taken as the minimum value when they are in the same state, with a weighting coefficient of 0.5-1.0. The optimal state is solved by iterating the confidence propagation algorithm 5-10 times, which reduces the misclassification rate of isolated missing data caused by local occlusion (occlusion area <30%) (the misclassification rate is reduced by ≥25%).
[0105] Specifically, step S43, which assesses the condition of the screws, also includes: S435: Fourier transform is used to extract the angular features of the torque marking line on the screw head. If the angular deviation n>5°, the loosening judgment is strengthened.
[0106] In a specific embodiment of the present invention, by analyzing the integrity of the torque marking line on the screw head, Fourier transform is used to extract the angular features of the marking line. If the angular deviation n>5°, the loosening judgment is strengthened.
[0107] In a specific embodiment of the present invention, based on the identification of the presence of the screw, the edge contours of the screw head and the nut are extracted by a sub-pixel edge detection algorithm (Zernike moment method, accuracy ±0.1 pixels), and the concentricity deviation between the two is calculated according to formula (7). Combined with the temperature distribution collected by the infrared thermal imager (sampling interval 0.5℃), when Pixels and local temperature difference When the temperature is >5℃ (compared to the surrounding area), the screw is determined to be loose. At the same time, the integrity of the torque mark line on the screw head is analyzed, and the angle feature of the mark line is extracted by Fourier transform (frequency domain peak detection). When the angle deviation is >5°, the looseness determination is strengthened. In this way, the confidence can be improved by 20%, realizing the upgrade from "existence recognition" to "state assessment". The looseness recognition accuracy is ≥90%.
[0108] Specifically, step S5, which combines multi-source information and detection results to locate the physical coordinates of the missing screw, further includes: S51: Combining the detection results and the collected multi-source information data, the pixel coordinates in the component image are converted into physical coordinates through perspective transformation. The multi-source information data includes GPS coordinates, equipment attitude angles, and camera lens parameters.
[0109] S52: The SfM algorithm is used to construct a local 3D point cloud from a continuous image sequence, and the ICP algorithm is used to match and correct the physical coordinates of the screw with the local 3D point cloud to generate a location information report. The location information report includes the screw location coordinates, quantity, confidence level and corresponding component number.
[0110] In a specific embodiment of this invention, during the high-precision positioning of the missing screw, the process first combines the detection results with GPS data, equipment attitude angles, and camera lens parameters collected during image capture (the intrinsic parameter matrix is obtained using the Zhang Zhengyou calibration method, with distortion coefficients k1-k3). A perspective transformation (i.e., a 3×3 transformation matrix) is then used to convert the pixel coordinates in the image into actual physical coordinates (unit: meters). The transformation matrix is calibrated offline using a 3×3 calibration board (accuracy ±0.1mm), with the error controlled within ±3cm. For continuous images of the same component, the SfM (Structure of Motion) algorithm is used to construct a local 3D point cloud (point cloud density ≥100 points / cm²) from the continuous image sequence. 2The screw coordinates are matched with the local 3D point cloud using the ICP (Iterative Closest Point) algorithm (10-20 iterations), and the positioning deviation is further corrected so that the final positioning error is ≤ ±5cm. A location information report is generated, which includes the precise location of the missing screw (latitude and longitude accurate to 0.0001°), quantity, confidence level (0-100%), and corresponding component number (coded according to the tower design drawings).
[0111] In a specific embodiment of the present invention, when locating the position of the missing screw, laser radar point cloud data (point cloud resolution 0.5cm, ranging accuracy ±2cm) is fused, and the surface of the tower component is extracted using a plane segmentation algorithm (RANSAC iteration 50 times, threshold 0.5cm). The screw coordinates located in the image are projected onto the point cloud plane, and the projection error is calculated. ,when At that time, the iterative nearest point algorithm was used to optimize the camera's extrinsic parameters. After 3-5 iterations, the positioning error was reduced to ≤±2cm, meeting the precise location requirements of maintenance personnel. Among these, the projection error... The calculation formula is: (9) In the formula, p is any point on the point cloud plane P, and q is the point to be projected.
[0112] The formula for updating extrinsic parameters is: (10) In the formula, T k+1 Let be the transformation matrix after the (k+1)th iteration, T be a 4×4 transformation matrix, pi be the point cloud feature points, and qi be the image feature points.
[0113] The extrinsic parameter update operation is performed if and only if M ≥ 20, where M is the successfully matched feature point pair.
[0114] Specifically, step S6, establishing a closed-loop feedback mechanism for cloud-edge collaboration to update and optimize the model, further includes: A cloud-edge collaborative framework is established, and detection results and manually combined data are uploaded to edge devices. When the number of images stored in the cloud reaches a preset threshold, incremental training is started, and causal reasoning and knowledge distillation techniques are used to automatically update and optimize the model. The optimized new model is then pushed to the edge device via OTA.
[0115] In a specific embodiment of the present invention, during the establishment of a closed-loop detection result feedback mechanism, a cloud-edge collaborative framework is first built (the edge device uses NVIDIA Jetson Xavier NX, and the cloud uses a GPU cluster). The edge device uploads detection results and manually reviewed data (such as missed detections, misjudged cases, and corrected labels) in real time. When the cloud accumulates 500-1000 newly labeled images, incremental training is initiated (freezing 80% of the backbone network parameters). Knowledge distillation technology (temperature coefficient 5) is used to retain the historical model performance, and the optimized new model parameters are pushed to the edge device via OTA (transmission bandwidth ≥10Mbps). A full retraining is performed monthly (updating the learning rate to 1 / 10 of the initial value) to update the model's ability to recognize new types of screws and rusted and aged screws (rust area accounting for 30%-80%). This invention establishes a closed-loop feedback mechanism for cloud-edge collaboration. Through incremental training, causal reasoning, and knowledge distillation techniques, the model can continuously learn and optimize autonomously. At the same time, by integrating GPS, attitude angle, image features, and 3D point cloud data, it achieves centimeter-level high-precision positioning of missing screws, directly guiding maintenance operations and significantly improving the overall efficiency of inspection and maintenance.
[0116] In a specific embodiment of the present invention, a causal reasoning module is added to the detection result feedback mechanism. This module uses a Bayesian network to analyze the causes of erroneous samples (factors include 10 categories such as lighting, angle, and screw type), calculates the contribution of each factor, and automatically generates specialized training datasets (≥500 samples per category) for factors with a contribution >0.3 (e.g., backlight angle >45°, M10 small screws). A meta-learning strategy is then employed. Rapid adaptation shortens the model's adaptation cycle to new interference factors to one-third of the original time, continuously improving recognition stability in complex scenarios. Among these, contribution... The calculation formula is: (11) In the formula, f represents influencing factors, and e represents erroneous events, such as missed detections or misjudgments. For conditional probability, The frequency of factor f in real-world scenarios. Let be the overall probability of event e occurring under all conditions.
[0117] Meta-learning strategies The calculation formula is: (12) In the formula, T is the number of tasks. These are the initial parameters for the model. For learning rate, , For about gradient operator, This represents the loss for the t-th task.
[0118] Example 2 This invention provides a method for detecting the condition of tower bolts on 110kV transmission towers in plain areas using unmanned aerial vehicles (UAVs). The specific process is as follows: This embodiment targets a 110kV linear transmission tower (30m high, Q235 steel components, M16-M24 screws) in a plain area. It employs a multi-rotor drone equipped with a detection system to intelligently identify missing screws, addressing the challenges of variable lighting conditions, component glare, and rapid inspection requirements in plain areas. The specific implementation process is as follows: Step 1: Deployment and Data Acquisition of Multimodal Image Acquisition Module The DJI Matrice 350RTK drone was used as the platform, equipped with a dual-mode acquisition system: a Baslerac A2500-14uc industrial camera for visible light acquisition and a FLIRVuePro infrared thermal imager. The drone's flight altitude was set at 5-10m, the distance to the tower structure at 3-5m, and the shooting angle at 45° to the structure's surface. Light intensity was monitored in real-time by the camera's built-in light sensor: visible light mode was activated when the light intensity was between 300 lux and 10000 lux, with the exposure time automatically adjusted to 10-30ms; when the light intensity was <300 lux (e.g., dusk, cloudy days), the system switched to infrared mode, with the thermal imager gain set to medium. Synchronously recorded data included: GPS coordinates (UBLOXNEO-8M module, accuracy ±0.3m, sampling rate 1Hz), shooting time (UTC timestamp, error <1ms), and device attitude angles (pitch angle -10° to 10°, yaw angle -90° to 90°, measurement accuracy ±0.1°). Each tower can capture 200-300 frames of images, each frame is stored in 12-bit RAW format, and transmitted back to the ground edge device (NVIDIA Jetson Xavier NX) in real time via a 4G private network.
[0119] Step 2: Multi-dimensional image preprocessing Step 2-1: Noise removal, using the OpenCV library to implement dynamic filtering—first calculate the image grayscale variance, use 3×3 Gaussian filtering (standard deviation 1.0) when the variance is <30, use 5×5 Gaussian filtering (standard deviation 1.5) when the variance is between 30 and 50, and enable 3×3 median filtering (iteration 2 times) when the variance is >50, to ensure that the noise suppression rate of component edges is ≥80%.
[0120] Step 2-2: Contrast enhancement using the CLAHE algorithm. The tiles grid size is set to 12×12. The contrast limit threshold is dynamically adjusted according to the scene (4.0 for sunny days, 2.0 for cloudy days, and 1.5 for dusk). 1.5% of extreme grayscale values are retained to avoid loss of details in the reflective areas of the screw head.
[0121] Steps 2-3: Image correction. Extract SIFT feature points (≥80) from each frame of the image and match them with the pre-stored feature library of standard tower components (matching threshold 0.75). Calculate the shooting angle deviation using the homography matrix and use bilinear interpolation to achieve correction. The corrected angle deviation is ≤±0.3°.
[0122] Steps 2-4: Motion blur recovery. For blurry images caused by drone shaking (gradient variance < 100 is considered blurry), a blind deconvolution algorithm is used: the initial point spread function is set to Gaussian (standard deviation 2.0), the number of iterations is 20, and the regularization parameter is 0.005. The image sharpness is improved by ≥ 40% after recovery.
[0123] Step 3: Multi-feature fusion extraction of the screw region Step 3-1: Edge detection, using an improved Canny algorithm—first calculate the image gradient orientation histogram (16 orientation boxes), in dense edge regions (pixels with gradient values > 80 accounting for > 15%), set the high threshold to 150 and the low threshold to 50, and in sparse edge regions, set the high threshold to 100 and the low threshold to 40; edge connection uses 8-neighborhood search (maximum connection distance 2 pixels) to ensure screw edge integrity ≥ 95%.
[0124] Step 3-2: Candidate region generation, combining Hough circle transform with YOLOv8-nano candidate boxes - Hough circle accumulator threshold 120, radius search range 8-25 pixels (step 1 pixel), and outputting the minimum bounding rectangle of the screw (width-to-height ratio 0.9-1.1, area ratio 0.8-1.0), filtering out non-circular interference areas (such as nut edges, component welds).
[0125] Step 3-3: Multi-feature calculation, extract the following features from the candidate regions: Geometric features: area (100-2000 pixels²), perimeter (50-150 pixels), center coordinates (subpixel accuracy ±0.1 pixels, calculated using the Zernike moment method).
[0126] Gray-scale characteristics: mean gray-scale value (80-180), standard deviation of gray-scale value (10-30).
[0127] Texture features: LBP histogram (32 bins, sampling radius 2 pixels), GLCM contrast (50-150, distance 2 pixels, average of 4 angles), texture entropy (2-4).
[0128] Step 4: Lightweight AI Model Construction and Phased Training Step 4-1: The AI model is built using an improved YOLOv8 architecture. The input layer supports the fusion of 3-channel visible light and 1-channel infrared images (channel weights are dynamically allocated through an attention mechanism). The backbone network uses a lightweight version of CSPDarknet-53 (the number of convolutional kernels is reduced by 30%, and the number of channels is 64-256 after 1×1 convolution dimensionality reduction). An SE module (compression rate 16) is added to enhance the response of the screw region. The neck network is PAN-FPN, which fuses 80×80 (shallow features), 40×40 (middle layer), and 20×20 (deep layer) feature maps. Upsampling (bilinear interpolation) and 3×3 convolution are used for fusion. The output layer uses an improved Softmax (temperature coefficient τ=1.0) to output three probabilities of "presence / missing / corrosion" and the region confidence.
[0129] Training process: Sample preparation: Collect 15,000 labeled images (4,000 on sunny days, 3,500 on cloudy days, 3,000 at dusk, and 4,500 at night), and divide them into a training set (12,000 images) and a validation set (3,000 images) in an 8:2 ratio. Labeling was performed using the LabelMe tool, which used polygonal boxes to select the screw area (accuracy ±1 pixel) and then labeled it.
[0130] Phase 1 (Basic Training): AdamW optimizer (initial learning rate 0.005, weight decay 0.0005, momentum 0.95), batch size=16, 80 iterations, and validation set metrics (accuracy, recall, F1 score) are calculated every 10 iterations.
[0131] The second stage (difficult example mining): 1200 samples with an error rate >30% (such as highly reflective screws and partially obscured screws) were selected, assigned a weight of 2.5 times, iterated 40 times, and the learning rate was reduced to 0.001.
[0132] Learning rate scheduling: When the F1 score does not improve for 5 consecutive times (fluctuation <0.5%), cosine annealing is enabled (10 iterations per cycle), with a minimum learning rate of 0.00005, and the final validation set accuracy ≥96% and recall ≥92%.
[0133] Step 5: Multi-level screw inspection and high-precision positioning Step 5-1: Multi-level detection Level 1 rapid screening: sliding window size 100×100 pixels, step size 20 pixels, filtering 80% of non-screw areas (areas with confidence <0.3 are directly excluded).
[0134] Level 2 fine detection: The candidate region is cropped to 224×224 pixels, and the input model is used to calculate the existence probability. If the existence probability is ≥0.85, it is judged as "existing"; if the existence probability is <0.6, it is judged as "missing"; if the existence probability is between 0.6 and 0.85, it is marked as "suspected", triggering 3 consecutive frames of verification (at least 2 frames are consistent to confirm).
[0135] Step 5-2: Status Assessment For screws that "exist", calculate the concentricity deviation according to formula (7). When δ>2 and the local infrared temperature difference>5℃, it is judged as "loose". The degree of corrosion is judged by the gray standard deviation. When the gray standard deviation>30, it is marked as "corrosion".
[0136] Step 5-3: Location Implementation Perspective transformation: Camera intrinsic parameters (focal length f=12mm, principal point coordinates (1280, 960)) are calibrated offline using a 3×3 calibration plate (accuracy ±0.05mm). The transformation matrix M is solved using the four calibration points, converting pixel coordinates (x, y) to physical coordinates (X, Y). The coordinate transformation formula is as follows: X=(M[0][0]x+M[0][1]y+M[0][2]) / M[2][2] (13) Y=(M[1][0]x+M[1][1]y+M[1][2]) / M[2][2] (14) Initial error ≤ ±3cm.
[0137] Point cloud correction: For 10 consecutive frames of images of the same component, a local 3D point cloud (point cloud density 120 points / cm²) is constructed using the SfM algorithm. The screw coordinates are then matched with the point cloud using the ICP algorithm (15 iterations). The final positioning error is ≤ ±2cm.
[0138] Report generation: Outputs an inspection report, including the location of missing screws (latitude and longitude accurate to 0.0001°), quantity, condition (missing / loose / corroded), confidence level, and corresponding component number (e.g., "Tower body section 3 crossarm L1-05").
[0139] Step 6: Closed-loop feedback mechanism operation Edge devices (JetsonXavierNX) upload detection results and manually reviewed data in real time (accumulating 800 new labeled images per week). Incremental training is initiated in the cloud (GPU cluster: 4×RTX3090)—freezing 80% of the backbone network parameters, training only the neck and output layers, using knowledge distillation (temperature coefficient 5) to preserve historical model performance, and pushing new model parameters via OTA (transmission bandwidth 20Mbps, time <5 minutes). Full retraining is performed monthly (learning rate 0.0005) to update the recognition capabilities for new M24 screws and heavily corroded (60% corroded area) screws.
[0140] The effectiveness verification data is shown in Table 1: Table 1
[0141] Table 1 demonstrates the core advantages of this invention in the inspection of iron towers in plains areas. Traditional single-modal detection relies on visible light images. Accuracy drops significantly under strong light due to component reflections, and under weak / backlight conditions due to feature loss. In rainy weather, accuracy falls below 60% due to rain interference. This invention, through visible-infrared fusion, dynamic filtering, and rain removal processing, maintains an accuracy of over 92% in various scenarios, solving the "environmental sensitivity" problem. Regarding detection speed, traditional algorithms require pixel-by-pixel traversal, with single-frame processing exceeding 350ms. This invention, with its dynamic sliding window and lightweight model, improves speed by 20%-30%, adapting to the 15 frames per second acquisition requirements of drones. In terms of positioning, traditional methods rely solely on coarse GPS conversion, resulting in errors exceeding 8cm. This invention, combined with point cloud correction, controls the error within ±6cm, allowing maintenance personnel to quickly locate targets and significantly shorten maintenance time.
[0142] Example 3 This invention provides a method for detecting the condition of tower bolts in 220kV tension towers in mountainous areas using a ground robot. The specific process is as follows: This embodiment targets a 220kV tension tower in a mountainous area (tower height 45m, components include numerous corner nodes, screw sizes M18-M27, mountainous area with frequent rain and dense vegetation). A tracked ground robot (climbing angle ≤30°) equipped with a detection system is used to address issues such as vegetation obstruction, rain and fog interference, identification of multiple screw sizes, and data privacy protection. The specific implementation process is as follows: Step 1: Multimodal Image Acquisition and Spatiotemporal Alignment The robot is equipped with a dual-mode acquisition system: the visible light camera is a Hikvision MV-CA050-10GM (5 megapixels, resolution 2560×1920, lens focal length 16mm, aperture F2.0, adaptable to working conditions of -10℃-50℃), the infrared thermal imager is an Guide IR236 (resolution 384×288, 8-14μm band, temperature measurement accuracy ±1℃), and the hyperspectral camera is a Headwall Nano-Hyperspec (400-1000nm band, 128 bands, spectral resolution 4nm). Data collection strategy: The robot moves upward along the base of the iron tower (speed 0.5m / s), and collects one set of data every 0.5m. Visible light and infrared images are synchronized by timestamp (time difference ≤10ms). Spatial alignment adopts ORB feature point matching (extracting ≥200 feature points, matching distance threshold 30). A 6-DOF affine transformation is performed on the infrared image, and the alignment error is ≤±1 pixel. Multimodal fusion adopts dynamic weight strategy. The weight coefficient is calculated by formula (1), where T=2 seconds. The infrared weight in the vegetation occlusion area is increased to 0.7 to ensure that the screw features are not occluded.
[0143] Step 2: Image Preprocessing and Adaptive Rain Removal Step 2-1: Basic Preprocessing The process is the same as in Example 2, but an adaptive rain removal step is added to address the rainy characteristics of mountainous areas. This involves identifying rain line regions using a raindrop detection operator (based on gradient direction consistency) (rain line density > 8 lines / 100×100 pixels indicates a rainy day), employing a U-Net architecture rain line detection network (input 512×512 pixels, output rain line mask), and then removing the rain lines using formula (2). Type 4 rain line template, =0.2 0.7, =0.05, the signal-to-noise ratio of the image after rain removal is improved by ≥25%.
[0144] Step 2-22: Illumination Compensation Uneven sunlight distribution under trees in mountainous areas necessitates the use of formulas. Adjust the pixel grayscale so that the difference in lighting between the edge and the center is ≤10%.
[0145] Step 3: Multi-feature extraction and material spectral verification Step 3-1: Conventional Feature Extraction Same as Example 2, but for screws of various models from M18 to M27, the radius of the Hough circle is adjusted to 10-30 pixels, and the screws are divided into three categories according to their diameter: "small (10-15 pixels), medium (16-22 pixels), and large (23-30 pixels)," and features are extracted for each category.
[0146] Step 3-2: Spectral Feature Verification The reflectance spectrum of the screw area was collected by a hyperspectral camera, and the spectral angle was calculated according to formula (3). , <5° indicates a metal screw, eliminating interfering objects such as tree branches and plastic sheeting (reducing the false judgment rate by ≥35%).
[0147] Step 4: Federated Learning Model Training and Dynamic Search Detection Step 4-1: Federated Learning Training To address the data privacy needs of multiple maintenance units in mountainous areas, a federated learning framework was adopted, using five edge devices (distributed at different inspection points in the mountains). Each device had 1000 local samples, and model parameters were exchanged using homomorphic encryption (2048-bit key). The local training loss function... Global model aggregation adopts After training, the cross-regional recognition accuracy improved by ≥12%.
[0148] Step 4-2: Dynamic Search Detection Based on the gradient divergence optimization search of the model heatmap, the divergence step size is calculated by formula (6) with a step size of 10 pixels and a step size of 30 pixels for ≥0 regions, and the processing time per frame is <250ms; at the same time, the Markov random field probability graph model is introduced, and the energy function is defined according to formula (8). The false positive rate of isolated missing due to vegetation occlusion is reduced by confidence propagation (8 iterations) (false positive rate decrease ≥28%).
[0149] Step 5: LiDAR Fusion Localization and Causal Inference Feedback Step 5-1: High-precision positioning Data from a fusion robot lidar (Velodyne VLP-16, point cloud resolution 0.8cm) was used to segment the tower component plane using the RANSAC algorithm (50 iterations, threshold 0.5cm). The screw coordinates of the image were projected onto the point cloud plane. The projection error was calculated according to the formula. When the projection error was >5cm, the ICP algorithm was used to optimize the camera extrinsic parameters. The extrinsic parameters were updated according to formula (10). After 3 iterations, the positioning error was ≤±2cm.
[0150] Step 5-2: Causal Reasoning Feedback The cloud-based causal reasoning module analyzes the causes of errors through Bayesian networks, calculates the contribution based on formula (11), generates a special dataset (≥600 images per category) for factors with a contribution >0.3, and adopts a meta-learning strategy for rapid adaptation, shortening the model adaptation cycle to 1 / 3 of the original.
[0151] The effectiveness verification data is shown in Table 2: Table 2
[0152] Table 2 highlights the specific advantages of this invention in the inspection of iron towers in mountainous areas. In densely vegetated mountainous areas, traditional robots, lacking the ability to distinguish occlusion, achieve a recognition rate of only 62%. This invention, through spectral features and probabilistic graphical models, accurately distinguishes screws from vegetation, increasing the recognition rate to 93%. For multi-type screw recognition, traditional models, lacking classification training, have a low recognition rate for large M27 screws. This invention extracts features based on size classification, achieving a multi-type recognition rate of 96%. In rainy weather detection, traditional methods lack adaptive rain removal, resulting in an accuracy rate below 55%. This invention's rain removal algorithm effectively preserves screw features, achieving an accuracy rate exceeding 90%. Regarding data privacy, traditional methods require transmitting original images, posing a risk of leakage. This invention, using federated learning, only transmits parameters, protecting privacy while improving cross-regional adaptability. The model iteration cycle is shortened from 30 days to 10 days, enabling rapid adaptation to new mountainous scenarios and solving the problems of traditional models being "unsuitable" and slow updates.
[0153] Reference Figure 10 This graph highlights the core advantages of this invention in complex environments through quantitative comparison, addressing the technical shortcomings of traditional methods in terms of "environmental sensitivity." Traditional single-modal methods rely on visible light images, which suffer from inaccuracies of less than 80% under strong light due to reflection, feature loss in low light / nighttime conditions, rain interference in rainy weather, and background obfuscation due to vegetation occlusion, especially in rainy weather where the accuracy is only 58%. This invention, through multimodal fusion (visible light + infrared), adaptive rain removal, and spectral feature verification, maintains an accuracy of over 92% in various scenarios, improving by 34 percentage points in rainy weather and 31 percentage points in vegetation obfuscation. The bar chart visually presents the differences, and the numerical annotations quantify the advantages, clearly demonstrating the innovative value of "multimodal collaborative anti-interference" and "spectral feature de-obfuscation," providing data support for the system's application in complex scenarios such as mountainous and rainy areas.
[0154] Reference Figure 11 This diagram clearly illustrates the continuous optimization process of the model through closed-loop feedback, addressing the issues of poor generalization ability and long update cycles in traditional models. The initial model, without iterations, was limited by sample coverage, achieving only 85% accuracy and a 15% misclassification rate. With increasing iterations, high-value samples were supplemented through difficult example mining, and federated learning aggregated experience from multiple regions, gradually improving model accuracy and recall while continuously decreasing the misclassification rate. After 20 iterations, accuracy reached 97% and the misclassification rate was only 3%. The design combining bilinear and dotted lines simultaneously presents the dual trends of "performance improvement" and "error reduction." The numerical node quantification optimization effect confirms the innovative effectiveness of "causal reasoning difficult example mining" and "federated learning collaborative training," providing a basis for formulating model update strategies and ensuring the system's long-term adaptability to changes in inspection scenarios.
[0155] Example 4 Based on the above method, embodiments of the present invention also provide a tower screw status detection system based on multimodal image fusion, such as... Figures 12-18As shown, it includes: Image acquisition module 01 is used to acquire multimodal component images of the target iron tower.
[0156] Image processing and feature extraction module 02 is used to perform image preprocessing and feature extraction on multimodal component images to obtain multi-scale feature maps.
[0157] Model building and training module 03 is used to build and train a tower screw recognition model based on multi-scale feature maps.
[0158] The detection and evaluation module 04 is used to perform multi-level detection and status evaluation on the acquired component images based on the tower screw recognition model.
[0159] Positioning module 05 is used to locate the physical coordinates of the missing screw by combining multi-source information and detection results. The update and optimization module 06 is used to establish a closed-loop feedback mechanism for cloud-edge collaboration to update and optimize the model.
[0160] Specifically, the image acquisition module 01 is used to acquire multimodal component images of the target tower, including visible light images and infrared images, as well as synchronously acquire the GPS coordinates, UTC timestamps, and device attitude angles, including pitch and yaw angles. Specifically, this invention employs a dual-mode acquisition system consisting of a 5-20 megapixel industrial camera and an 8-12μm wavelength infrared thermal imager. The system has a built-in light sensor; when the light intensity is ≥300 lux, visible light mode is activated for shooting, and when the light intensity is <300 lux, it automatically switches to infrared mode for shooting. During image acquisition, the device's GPS coordinates (accuracy ±0.5m), UTC timestamps (error <1ms), and device attitude angles (pitch and yaw angles, measurement accuracy ±0.1°) are recorded simultaneously to ensure data spatiotemporal synchronization.
[0161] In a specific embodiment of the present invention, after the image acquisition module 01 acquires the multimodal component image of the target tower, it further includes: Spatiotemporal alignment is performed on the acquired multimodal component images. The acquisition times of visible light and infrared images are synchronized by timestamps, and the image offset is calculated using the ORB feature point matching algorithm. Pixel alignment of the infrared image is performed by affine transformation.
[0162] Dynamic weight coefficients are calculated based on a dynamic weight fusion strategy, and image fusion is performed on multimodal component images.
[0163] Specifically, in terms of time, the acquisition times of visible light and infrared images are synchronized by timestamps (time difference ≤ 10ms). In terms of space, ORB feature point matching (extracting ≥ 200 feature points, matching threshold 30) is used to calculate the offset. Pixel-level alignment is achieved by performing 6-DOF affine transformation on the infrared image (alignment error ± 1 pixel), with the alignment error controlled within ± 1 pixel. At the same time, a dynamic weight fusion strategy is introduced to maintain feature stability of the fused image when there are sudden changes in illumination (change rate > 50% / s), and the edge preservation rate is improved by ≥ 15%.
[0164] Specifically, the image processing and feature extraction module 02 also includes: Image processing module 021, used to perform noise reduction, contrast enhancement, and angle correction on the acquired multimodal component images, further includes: Image denoising module 0211 is used to denoise the acquired multimodal component images based on a dynamic Gaussian filtering algorithm.
[0165] The contrast enhancement module 0212 is used to perform contrast enhancement processing on the denoised image using the CLAHE algorithm and convert the RGB image to a grayscale image.
[0166] The correction module 0213 is used to perform image angle correction processing on the converted grayscale image using the SIFT algorithm.
[0167] The blurred image repair module 0214 is used to repair motion-blurred images in multimodal component images using a blind deconvolution algorithm, wherein the motion-blurred image is an image with a gradient variance < 100.
[0168] The rain line removal module 0215 is used to identify rain line regions in an image and remove rain lines using the U-Net rain line detection network.
[0169] In a specific embodiment of the present invention, the image denoising module 0211 first performs Gaussian filtering on the acquired multimodal component image, that is, dynamically adjusts the filter kernel size according to the noise variance to achieve Gaussian filtering denoising on the image. Then, the contrast enhancement module 0212 enhances the image contrast and converts the RGB image into a grayscale image. Finally, the correction module 0213 performs image angle correction.
[0170] In a specific embodiment of the present invention, when performing multi-dimensional preprocessing on the acquired image, noise is first removed by the image denoising module 0211. The filter kernel size is dynamically adjusted to 3×3-7×7 according to the image noise intensity (adjusted to 3×3 when the noise variance is <30, adjusted to 5×5 when the noise variance is between 30-50, and enabled with 3×3 median filtering when the noise variance is >50, with 2 iterations). Then, the contrast enhancement module 0212 is used to enhance the image contrast. The contrast limit threshold is set to 2.0-4.0 (2.0 for cloudy days and 4.0 for sunny days). The tiles grid size is 8×8-16×16. At the same time, 1%-2% of extreme gray values are retained to avoid loss of detail. After converting the RGB image to grayscale using the weighted formula (Gray=0.299R+0.587G+0.114B), the shooting angle deviation is corrected by the correction module 0213, and the correction accuracy is controlled within ±0.3°.
[0171] In a specific embodiment of the present invention, for motion-blurred images (i.e., images with gradient variance < 100), the blur image repair module 0214 is used to repair the motion-blurred images in the multimodal component images and restore clarity.
[0172] In a specific embodiment of the present invention, in a rainy scene, for images taken in the rain (identified by a raindrop detection operator, with rain line density > 5 lines / 100×100 pixels), a rain line removal module 0215 is used to identify the rain line region and remove rain line interference, thereby improving the signal-to-noise ratio of the rainy image by more than 20% and the screw region feature retention rate ≥ 95%.
[0173] The multi-dimensional feature extraction module 022 is used to extract multi-dimensional features from the processed image to obtain geometric, grayscale, texture, and spectral multi-dimensional features, and further includes: The candidate region generation module 0221 is used to perform edge detection on the processed image using the improved Canny algorithm, and to generate screw candidate regions by combining the Hough circle transform and the YOLOv8-nano model.
[0174] Feature extraction module 0222 is used to extract geometric features, grayscale features, texture features, and spectral features of an image based on the screw candidate region. Geometric features include area, perimeter, and center coordinates; grayscale features include grayscale mean; texture features include LBP histogram, GLCM contrast, and entropy value; spectral features include spectral angle. .
[0175] In a specific embodiment of the present invention, the candidate region generation module 0221 first performs edge detection, combines Hough circle transform and YOLOv8-nano model to generate screw candidate regions, and outputs the minimum bounding rectangle. The feature extraction module 0222 calculates the geometric features, grayscale features, texture features, and spectral features of the candidate regions, constructing a multi-feature fusion set. The geometric features include the area, perimeter, and center coordinates of the screw candidate regions; the grayscale features include the mean grayscale value of the screw candidate regions; the texture features include the LBP histogram, GLCM contrast, and entropy value; and the spectral features include the spectral angle, which is calculated by acquiring the reflectance spectrum curve of the screw region using a hyperspectral camera. .
[0176] In a specific embodiment of the present invention, during the extraction of the multi-feature fusion set of the screw candidate region, an improved Canny edge detection algorithm is first adopted. The threshold range of 50-150 is dynamically set using a gradient orientation histogram (16 orientation boxes) (the threshold is increased when edge density is high). An 8-neighborhood search strategy is used for edge connections (connection threshold ≤ 2 pixels). Combining the Hough circle transform (accumulator threshold ≥ 100) with the candidate box generation mechanism of the YOLOv8-nano model, the circle detection radius is set to a range of 5-30 pixels (step size 1 pixel). Output the minimum bounding rectangle parameters (aspect ratio 0.8-1.2, area ratio 0.7-1.0), calculate the area (50-3000 pixels), perimeter (30-200 pixels), center coordinates (subpixel accuracy ±0.1 pixels), grayscale mean (50-200), and texture features of the candidate region. Texture features include LBP histogram (16-64 bins, sampling radius 1-3 pixels), GLCM contrast (0-200, distance 1-3 pixels, angle 0° / 45° / 90° / 135°), and entropy value (0-5).
[0177] In a specific embodiment of the present invention, material spectral features are introduced into the feature extraction process. The reflectance spectral curve of the screw region is acquired using a hyperspectral camera (wavelength 400-1000nm, number of bands ≥100, spectral resolution ≤5nm), and the spectral angle is calculated. ,when It can identify metal materials and exclude interference such as plastic and stones (reducing the false recognition rate by ≥30%), improving the feature discrimination in complex backgrounds, and is especially suitable for tower detection in vegetated areas.
[0178] Specifically, the model building and training module 03 also includes: Model building module 031 is used to build a tower screw recognition model based on the improved YOLOv8 architecture. Training module 032 is used to input the training set and combine the phased strategy and federated learning framework to train the model in stages, and output the probability of screw presence, probability of missing screw, and region confidence.
[0179] The backbone network of the tower screw recognition model uses a lightweight CSPDarknet-53 and embeds an SE attention mechanism. The neck network adopts a PAN-FPN structure and fuses multi-scale feature maps. The output layer uses an improved Softmax activation function.
[0180] In a specific embodiment of the present invention, a tower screw recognition model is first established using model building module 031. This model supports dual-channel input of visible light and infrared images. Its backbone network uses the lightweight CSPDarknet-53, thus reducing the number of convolutional kernels by 30%, and embedding an SE attention mechanism. The neck network adopts a PAN-FPN structure and fuses multi-scale feature maps, and the output layer uses an improved Softmax activation function (introducing a temperature coefficient τ). By inputting multi-scale feature maps into this model, the presence probability, absence probability, and region confidence of the screw can be output.
[0181] Specifically, in constructing the lightweight tower screw recognition model, an improved YOLOv8 architecture is adopted. The input layer supports dual-channel fusion of visible light and infrared images (channel weights can be dynamically adjusted). The backbone network uses a lightweight version of CSPDarknet-53, employing 1×1 convolution for dimensionality reduction, reducing the number of convolution kernels by 30%, and adding an attention mechanism (SE module, compression ratio 16) to enhance the response of the screw region. The neck network adopts a PAN-FPN structure, fusing feature maps at three scales: 80×80, 40×40, and 20×20. Upsampling and convolution are used for feature fusion. The output layer uses an improved Softmax activation function, with a temperature coefficient τ = 0.5-2.0 added to the function. The output outputs the screw presence probability, missing probability, and region confidence (0-1). The model training uses the AdamW optimizer, with an initial learning rate set to 0.001-0.01 (adaptively adjusted with batch size), a weight decay coefficient set to 0.0001-0.001, and a momentum parameter set to 0.9-0.99.
[0182] In a specific embodiment of this invention, training module 032 is used to train the model. First, 10,000-50,000 labeled images of iron towers are divided into a training set and a validation set in an 8:2 ratio. Then, the training set is input into the iron tower screw recognition model for training. The training is divided into two stages: the first stage trains the basic model; the second stage introduces hard example mining, assigning higher weights to samples with high error rates. During the training process, accuracy, recall, and F1 score are calculated every 10 iterations. If the F1 score does not improve for 5 consecutive iterations, cosine annealing learning rate scheduling is enabled. This invention uses multimodal image fusion technology (visible light, infrared, hyperspectral, etc.) combined with a federated learning framework to construct a lightweight AI model with strong environmental robustness and high generalization ability while protecting data privacy, thereby reducing deployment costs and improving cross-regional adaptability.
[0183] Specifically, during the phased training and dynamic optimization of the model, 10,000-50,000 labeled images containing screws and missing screws (covering screw models M10-M30) were collected. The images were divided into four categories according to the scene: sunny, cloudy, foggy, and nighttime, with each category accounting for 20%-30% of the samples. The training set and validation set were divided in an 8:2 ratio. Stratified sampling was used to maintain category balance. When labeling, the screw area was selected by polygonal bounding box (labeling accuracy ±1 pixel) and marked with "present", "missing", and "rusted". The first stage trains the basic model, iterating 50-100 times with a batch size of 8-32 (which can be dynamically adjusted according to GPU memory). The second stage introduces hard example mining, assigning 2-3 times the weight to samples with an error rate >30% (such as highly reflective samples or partially obscured screws), iterating 30-50 times. During training, the validation set accuracy, recall, and F1 score (accurate to 0.1%) are calculated every 10 iterations. When the F1 score does not improve for 5 consecutive iterations, cosine annealing is used for learning rate scheduling (10 iterations per cycle), with a minimum learning rate of 0.00001-0.0001, until the model achieves an accuracy ≥95% and a recall ≥90% on the validation set.
[0184] In a specific embodiment of the present invention, a federated learning framework is used during model training. For edge devices (usually 5-20 units) distributed in different tower inspection areas, collaborative model training is achieved through encrypted parameter exchange (using homomorphic encryption algorithm with a key length of 2048 bits). This improves the model's generalization ability while protecting data privacy, and increases the cross-regional recognition accuracy by ≥10%.
[0185] Specifically, the detection and evaluation module 04 also includes: The first-level detection module 041 is used to input the component image into the model and perform first-level detection on the component image using a preset sliding window to generate screw candidate regions.
[0186] The secondary detection module 042 is used to perform secondary detection on the screw candidate area and output the probability of the screw's presence.
[0187] Evaluation module 043 is used to determine whether a screw exists based on the probability of its presence, and if it exists, to evaluate the screw's condition.
[0188] In a specific embodiment of the present invention, during the multi-level screw detection process on the component image, the first-level detection module 041 first inputs the preprocessed component image into the trained AI model. The first-level detection uses a sliding window (step size 10-30 pixels, window size 50×50-200×200 pixels, adaptively adjusted according to the image scale) for rapid screening, filtering out more than 70% of non-screw areas and outputting screw candidate areas. Then, the second-level detection module 042 performs fine detection on the screw candidate areas (inputting a 224×224 pixel cropped image) and outputs the screw presence probability of each candidate area. Finally, the evaluation module 043 calculates the screw presence probability, represented as P. When P ≥ 0.85, the screw is determined to exist; when P < 0.6, the screw is determined to be missing. When the screw exists, a sub-pixel edge detection algorithm is used to extract the edge contours of the screw head and nut, and the concentricity deviation δ between the two is calculated. Combined with the temperature distribution T collected by the infrared thermal imager and the concentricity deviation δ, when δ > 2 and the local temperature difference is significant, the screw is considered to be present. If the temperature is >5℃, the screw is considered loose. If 0.6≤P<0.85, the image is marked as a suspected image, and multiple frames of continuous verification are performed on the suspected image.
[0189] In a specific embodiment of the present invention, an attention-oriented dynamic search mechanism is added to the first-level detection module 041 and the second-level detection module 042. A heatmap is generated based on the region confidence of the model output, and the gradient divergence is calculated. In the region with divergence value < 0 (attention convergence area, confidence ≥ 0.6), a fine search with a step size of 10 pixels is used, and in the region with divergence value ≥ 0 (low confidence area), a fast search with a step size of 30 pixels is used. This improves the detection speed by more than 40% while maintaining the accuracy (accuracy fluctuation ≤ 1%), meeting the real-time requirements of UAV inspection (single frame processing time < 200ms).
[0190] In a specific embodiment of the present invention, the evaluation module 043 determines the screw status based on the existence probability output by the model: if the probability P ≥ 0.85, it is determined to exist; if the probability P < 0.6, it is determined to be missing; if the probability 0.6 ≤ P < 0.85, it is marked as a suspected image, and the suspected image is continuously verified for multiple frames (confirmation is required for 3-5 consecutive frames, and at least 3 frames must be consistent for final determination). Based on the identification of the screw's existence, the edge contours of the screw head and nut are extracted using a sub-pixel edge detection algorithm, and the concentricity deviation δ between the two is calculated. Combined with the temperature distribution collected by the infrared thermal imager, when δ > 2 pixels and the local temperature difference is significant, the screw's status is determined. If the temperature is above 5℃, the screw is considered loose.
[0191] In a specific embodiment of the present invention, a probability graph module is introduced into the evaluation module 043 to divide the tower components into 10×10cm² sub-regions, construct a Markov random field to describe the correlation of screw existence, define an energy function, and solve the optimal state by iterating 5-10 times through the confidence propagation algorithm, thereby reducing the misjudgment rate of isolated missing due to local occlusion (occlusion area <30%) (misjudgment rate reduction ≥25%).
[0192] In a specific embodiment of the present invention, the evaluation module 043 is also used to extract the angular features of the torque marking line on the screw head. If the angular deviation n>5°, the loosening judgment is strengthened.
[0193] In a specific embodiment of the present invention, based on the identification of the presence of a screw, the edge contours of the screw head and the nut are extracted by the evaluation module 043, and the concentricity deviation between the two is calculated. Combined with the temperature distribution collected by the infrared thermal imager (sampling interval 0.5℃), when Pixels and local temperature difference When the temperature is >5℃ (compared to the surrounding area), the screw is determined to be loose. At the same time, the integrity of the torque mark line on the screw head is analyzed, and the angle feature of the mark line is extracted by Fourier transform (frequency domain peak detection). When the angle deviation is >5°, the looseness determination is strengthened. In this way, the confidence can be improved by 20%, realizing the upgrade from "existence recognition" to "state assessment". The looseness recognition accuracy is ≥90%.
[0194] Specifically, the positioning module 05 also includes: The coordinate transformation module 051 is used to combine the detection results and the collected multi-source information data to convert the pixel coordinates in the component image into physical coordinates through perspective transformation. The multi-source information data includes GPS coordinates, device attitude angles, and camera lens parameters.
[0195] The 3D point cloud matching module 052 is used to construct a local 3D point cloud from a continuous image sequence using the SfM algorithm, and to match and correct the physical coordinates of the screw with the local 3D point cloud using the ICP algorithm, generating a location information report. The location information report includes the screw's location coordinates, quantity, confidence level, and corresponding component number.
[0196] In a specific embodiment of this invention, during the high-precision positioning of the missing screw, the coordinate transformation module 051 first uses perspective transformation (i.e., a 3×3 transformation matrix) to convert the pixel coordinates in the image into actual physical coordinates (unit: meters) based on the detection results and the GPS, equipment attitude angle, and camera lens parameters collected during shooting (the intrinsic parameter matrix is obtained through Zhang Zhengyou's calibration method, with distortion coefficients k1-k3). The transformation matrix is then calibrated offline using a 3×3 calibration plate (accuracy ±0.1mm), with the error controlled within ±3cm. For continuous images of the same component, the 3D point cloud matching module 052 constructs a local 3D point cloud (point cloud density ≥100 points / cm²) from the continuous image sequence. 2 The screw coordinates are matched with the local 3D point cloud using the ICP (Iterative Closest Point) algorithm (10-20 iterations), and the positioning deviation is further corrected so that the final positioning error is ≤ ±5cm. A location information report is generated, which includes the precise location of the missing screw (latitude and longitude accurate to 0.0001°), quantity, confidence level (0-100%), and corresponding component number (coded according to the tower design drawings).
[0197] In a specific embodiment of the present invention, when locating the position of the missing screw, laser radar point cloud data (point cloud resolution 0.5cm, ranging accuracy ±2cm) is fused, and the surface of the tower component is extracted using a plane segmentation algorithm (RANSAC iteration 50 times, threshold 0.5cm). The screw coordinates located in the image are projected onto the point cloud plane, and the projection error is calculated. ,when At that time, the iterative nearest point algorithm is used to optimize the camera's extrinsic parameters. After 3-5 iterations, the positioning error is ≤±2cm, which meets the precise location needs of maintenance personnel.
[0198] Specifically, the update and optimization module 06 is used to establish a cloud-edge collaborative framework and upload detection results and artificial composite data to edge devices. When the number of images stored in the cloud reaches a preset threshold, incremental training is started, and causal reasoning and knowledge distillation techniques are used to automatically update and optimize the model. The optimized new model is then pushed to the edge device via OTA.
[0199] In a specific embodiment of the present invention, during the establishment of a closed-loop detection result feedback mechanism, a cloud-edge collaborative framework is first built using the update and optimization module 06 (the edge device uses NVIDIA Jetson Xavier NX, and the cloud uses a GPU cluster). The edge device uploads detection results and manually reviewed data (e.g., missed detections, misjudged cases, and corrected labels) in real time. When the cloud accumulates 500-1000 newly labeled images, incremental training is initiated (freezing 80% of the backbone network parameters). Knowledge distillation technology (temperature coefficient 5) is used to retain the historical model performance, and the optimized new model parameters are pushed to the edge device via OTA (transmission bandwidth ≥10Mbps). A full retraining is performed monthly (updating the learning rate to 1 / 10 of the initial value) to update the model's ability to recognize new types of screws and rusted and aged screws (rust area accounting for 30%-80%). This invention establishes a closed-loop feedback mechanism for cloud-edge collaboration. Through incremental training, causal reasoning, and knowledge distillation techniques, the model can continuously learn and optimize autonomously. At the same time, by integrating GPS, attitude angle, image features, and 3D point cloud data, it achieves centimeter-level high-precision positioning of missing screws, directly guiding maintenance operations and significantly improving the overall efficiency of inspection and maintenance.
[0200] In a specific embodiment of the present invention, a causal reasoning module is added to the update and optimization module 06. The cause of the error samples is analyzed by Bayesian network (the factors include 10 categories such as illumination, angle, screw type, etc.). The contribution of each factor is calculated. For factors with a contribution of >0.3 (such as backlight angle >45°, M10 small screws), a special training dataset (≥500 samples per category) is automatically generated. The meta-learning strategy L_meta is used for rapid adaptation, which shortens the adaptation period of the model to new interference factors to 1 / 3 of the original, and continuously improves the recognition stability in complex scenes.
[0201] Example 5 like Figure 19 As shown, this embodiment of the invention also provides an electronic device, including a processor and a memory. The memory stores a computer program, which is loaded and executed by the processor to implement the above-described method for detecting the state of iron tower screws based on multimodal image fusion. The device in this invention can be a server, PC, PAD, mobile phone, etc.
[0202] Furthermore, this embodiment of the invention also provides a computer-readable storage medium storing a computer program, which is loaded and executed by a processor to implement the above-described method for detecting the state of iron tower screws based on multimodal image fusion.
[0203] In summary, the tower screw status detection method and system based on multimodal image fusion described in this invention has the following advantages: 1. This invention effectively overcomes the bottleneck of traditional single-modal image acquisition and fusion technology, which is limited by environmental conditions. The dynamic switching and fusion of visible light and infrared images can clearly preserve screw features under adverse conditions such as insufficient light, rain, and fog, avoiding feature loss due to environmental interference. Simultaneously, the introduction of material spectral features and three-dimensional morphological features can accurately distinguish metal screws from background interference objects such as plastic and stones, significantly reducing the misjudgment rate in complex backgrounds. Whether in sunny, cloudy, nighttime, or rainy weather, the system can stably extract key information such as screw edges and textures, ensuring the consistency and reliability of the recognition results and overcoming the limitations of existing "sky-based detection" technologies.
[0204] 2. Unlike existing technologies that can only determine whether a screw is "present" or "missing," this invention adds a screw tightness assessment submodule. It calculates the concentricity deviation between the screw head and nut using sub-pixel edge detection, combines this with temperature distribution data collected by an infrared thermal imager to determine looseness, analyzes the integrity of torque marking lines to enhance the assessment, and can also identify the degree of screw corrosion, forming a multi-dimensional status assessment system of "present-missing-loose-corroded." This refined identification capability helps inspection personnel not only discover missing screws but also provide early warnings of potential risks such as loosening and corrosion, advancing tower safety management from "post-event maintenance" to "pre-event prevention," significantly improving the structural safety level of towers.
[0205] 3. To address the privacy protection needs of inspection data from different regions, this invention employs a federated learning framework. Edge devices only upload model parameters, not raw images, and collaborative training is achieved through encrypted parameter exchange, effectively avoiding the risk of data leakage. The global model uses a weighted average to aggregate sample information from each device, improving the model's adaptability to different regions and screw types without the need for centralized storage of massive amounts of data. Compared to traditional centralized training, this mode protects data privacy while allowing the model to quickly absorb inspection experience from various regions, significantly enhancing its generalization ability. It also eliminates the need for separate model training for each region, greatly reducing deployment costs.
[0206] 4. In terms of detection efficiency, this invention designs an attention-oriented dynamic search mechanism, which performs fine-grained searching in high-confidence areas and rapid filtering in low-confidence areas, significantly improving detection speed without reducing recognition accuracy. This meets the real-time requirements of UAV inspections, and the single-frame processing efficiency is significantly better than existing technologies. In terms of positioning accuracy, by integrating multi-source data such as GPS, equipment attitude angle, and LiDAR point cloud, and through multiple rounds of calibration such as perspective transformation and ICP algorithm, the positioning error of missing screws is controlled to an extremely low range. Maintenance personnel can directly find the target location based on the positioning results without repeatedly searching on the tower, greatly shortening maintenance time.
[0207] 5. This invention establishes a closed-loop feedback framework for cloud-edge collaboration. Edge devices upload missed and misjudged cases reviewed by manual verification in real time. The cloud, through incremental training and knowledge distillation techniques, rapidly absorbs new sample experience while retaining the performance of historical models. Monthly full retraining updates the model's adaptability to new interference factors. Simultaneously, the causal reasoning module analyzes the causes of errors and generates targeted training datasets, accelerating the model's learning of weak scenarios. This allows the model to continuously evolve with changes in inspection scenarios, maintaining high recognition performance over the long term and avoiding the problem of "performance degradation after deployment" in existing systems.
[0208] Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for detecting the state of iron tower screws based on multimodal image fusion, characterized in that, include: Acquire multimodal component images of the target tower; Image preprocessing and feature extraction are performed on multimodal component images to obtain multi-scale feature maps; A tower screw recognition model was constructed and trained based on multi-scale feature maps; Multi-level detection and status assessment are performed on the acquired component images based on the tower screw recognition model; By combining multi-source information and detection results, the physical coordinates of the missing screw were located, and Establish a closed-loop feedback mechanism for cloud-edge collaboration to update and optimize the model.
2. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 1, characterized in that, The acquisition of multimodal component images of the target tower further includes: Industrial cameras and infrared thermal imagers were used to acquire multimodal component images of the target iron tower, including visible light and infrared images, as well as... The GPS coordinates, UTC timestamps, and device attitude angles of the device are collected synchronously, including pitch angle and yaw angle.
3. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 2, characterized in that, The step of image preprocessing and feature extraction of multimodal component images to obtain multi-scale feature maps further includes: The acquired multimodal component images are denoised, contrast-enhanced, and angle-corrected. Multi-dimensional feature extraction is performed on the processed image to obtain geometric, grayscale, texture, and spectral features.
4. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 3, characterized in that, The denoising, contrast enhancement, and angle correction of the acquired multimodal component images further include: The acquired multimodal component images are denoised using a dynamic Gaussian filtering algorithm. The CLAHE algorithm is used to enhance the contrast of the denoised image and convert the RGB image to a grayscale image. The SIFT algorithm is used to perform image angle correction processing on the converted grayscale image.
5. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 3, characterized in that, The image preprocessing of the acquired multimodal component images also includes: A blind deconvolution algorithm is used to repair motion-blurred images in multimodal component images, wherein the motion-blurred images are images with a gradient variance of <100.
6. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 3, characterized in that, The image preprocessing of the acquired multimodal component images also includes: The U-Net rain line detection network is used to identify rain line regions in images and remove rain lines. The calculation formula for removing rain lines is as follows: in, Here is the rain line removal function, where I(x, y) is the original grayscale value. For the k-th type of rain line template, The strength coefficient, As the attenuation factor, The distance is Euclidean.
7. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 3, characterized in that, The step of extracting multi-dimensional features from the processed image to obtain geometric, grayscale, texture, and spectral multi-dimensional features further includes: An improved Canny algorithm is used to perform edge detection on the processed image, and Hough circle transform and YOLOv8-nano model are combined to generate screw candidate regions; Based on the candidate screw region, geometric features, grayscale features, texture features, and spectral features of the image are extracted. The geometric features include area, perimeter, and center coordinates; the grayscale features include grayscale mean; the texture features include LBP histogram, GLCM contrast, and entropy value; and the spectral features include spectral angle. The calculation formula is as follows: In the formula, The spectral vector to be detected, For standard metal screw vectors, N is the number of bands.
8. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 1, characterized in that, The method for constructing and training the tower screw recognition model based on multi-scale feature maps further includes: A tower screw recognition model was established based on the improved YOLOv8 architecture; Input the training set and combine a phased strategy and federated learning framework to train the model in stages, and output the probability of screw presence, the probability of missing screw, and the region confidence. The backbone network of the tower screw recognition model adopts a lightweight CSPDarknet-53 and embeds an SE attention mechanism, the neck network adopts a PAN-FPN structure and fuses multi-scale feature maps, and the output layer adopts an improved Softmax activation function.
9. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 1, characterized in that, The multi-level detection and status assessment of the acquired component images based on the tower screw recognition model further includes: The component image is input into the model and a preset sliding window is used to perform first-level detection on the component image to generate screw candidate regions; Perform secondary detection on the candidate screw region and output the probability of screw presence; The presence or absence of a screw is determined by the probability of its presence; if it exists, the condition of the screw is evaluated.
10. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 9, characterized in that, The step of determining whether a screw exists based on its probability of presence, and then evaluating its condition if it exists, further includes: Let P be the probability of a screw's presence. If P ≥ 0.85, the screw is considered to be present; if P < 0.6, the screw is considered to be missing. When a screw is present, a sub-pixel edge detection algorithm is used to extract the edge contours of the screw head and the nut, and the concentricity deviation δ between the two is calculated. The calculation formula is as follows: in, , The coordinates of the center of the screw head are: , Here are the coordinates of the nut's center; Combining the temperature distribution T and concentricity deviation δ acquired by the infrared thermal imager, when δ > 2 and the local temperature difference... If the temperature is above 5℃, the screw is considered loose.
11. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 10, characterized in that, The method of determining whether a screw exists based on its probability of presence also includes: If 0.6 ≤ P < 0.85, the image is marked as a suspected image, and the suspected image is subjected to multi-frame continuous verification.
12. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 10, characterized in that, The assessment of screw condition also includes: Fourier transform is used to extract the angular features of the torque marking line on the screw head. If the angular deviation n > 5°, the loosening judgment is strengthened.
13. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 1, characterized in that, The step of combining multi-source information and detection results to locate the physical coordinates of the missing screw further includes: Combining the detection results and the collected multi-source information data, the pixel coordinates in the component image are converted into physical coordinates through perspective transformation. The multi-source information data includes GPS coordinates, device attitude angles, and camera lens parameters. The SfM algorithm is used to construct a local 3D point cloud from a continuous image sequence, and the ICP algorithm is used to match and correct the physical coordinates of the screws with the local 3D point cloud to generate a location information report. The location information report includes the screw location coordinates, quantity, confidence level, and corresponding component number.
14. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 1, characterized in that, The establishment of a closed-loop feedback mechanism for cloud-edge collaboration to update and optimize the model further includes: A cloud-edge collaborative framework is established, and detection results and manually combined data are uploaded to edge devices. When the number of images stored in the cloud reaches a preset threshold, incremental training is started, and causal reasoning and knowledge distillation techniques are used to automatically update and optimize the model. The optimized new model is then pushed to the edge device via OTA.
15. The method for detecting the state of iron tower screws based on multimodal image fusion according to claim 1, characterized in that, After acquiring the multimodal component images of the target tower, the process also includes: Spatiotemporal alignment of the acquired multimodal component images is performed. The acquisition times of visible light and infrared images are synchronized by timestamps. The image offset is calculated by ORB feature point matching algorithm. The infrared image is pixel aligned by affine transformation. Dynamic weight coefficients are calculated based on a dynamic weight fusion strategy, and image fusion is performed on multimodal component images. The calculation formula for the dynamic weight coefficients is as follows: In the formula, For dynamic weighting coefficients, Let be the entropy of the visible light image at time t. Let be the variance of the visible light image information entropy at time t. Let be the entropy of the infrared image information at time t. Let be the variance of the infrared image information entropy at time t. For time windows.
16. A tower screw condition detection system based on multimodal image fusion, characterized in that, include: The image acquisition module is used to acquire multimodal component images of the target tower; The image processing and feature extraction module is used to perform image preprocessing and feature extraction on multimodal component images to obtain multi-scale feature maps; The model building and training module is used to build and train a tower screw recognition model based on multi-scale feature maps; The detection and evaluation module is used to perform multi-level detection and status evaluation on the acquired component images based on the tower screw recognition model; The positioning module is used to locate the physical coordinates of the missing screw by combining multi-source information and detection results; The update and optimization module is used to establish a closed-loop feedback mechanism for cloud-edge collaboration to update and optimize the model.
17. The tower screw status detection system based on multimodal image fusion according to claim 16, characterized in that, The image processing and feature extraction module also includes: The image processing module is used to denoise, enhance contrast, and correct angles on the acquired multimodal component images; The multi-dimensional feature extraction module is used to extract multi-dimensional features from the processed image, obtaining geometric, grayscale, texture, and spectral features.
18. The tower screw status detection system based on multimodal image fusion according to claim 17, characterized in that, The image processing module further includes: The image denoising module is used to denoise the acquired multimodal component images based on the dynamic Gaussian filtering algorithm; The contrast enhancement module is used to enhance the contrast of the denoised image using the CLAHE algorithm and convert the RGB image to a grayscale image. The correction module is used to perform image angle correction processing on the converted grayscale image using the SIFT algorithm; The blurred image restoration module is used to restore motion-blurred images in multimodal component images using a blind deconvolution algorithm, wherein the motion-blurred image is an image with a gradient variance < 100; The rain line removal module is used to identify rain line regions in images and remove rain lines using the U-Net rain line detection network.
19. The tower screw status detection system based on multimodal image fusion according to claim 17, characterized in that, The multi-dimensional feature extraction module also includes: The candidate region generation module is used to perform edge detection on the processed image using an improved Canny algorithm, and to generate screw candidate regions by combining Hough circle transform and YOLOv8-nano model. The feature extraction module is used to extract geometric features, grayscale features, texture features, and spectral features from the screw candidate region of the image. The geometric features include area, perimeter, and center coordinates; the grayscale feature is the grayscale mean; the texture features include LBP histogram, GLCM contrast, and entropy value; and the spectral feature is the spectral angle. .
20. The tower screw status detection system based on multimodal image fusion according to claim 16, characterized in that, The model building and training module also includes: The model building module is used to build a tower screw recognition model based on the improved YOLOv8 architecture; The training module is used to input the training set and combine a phased strategy and federated learning framework to train the model in stages, and output the probability of screw presence, the probability of missing screws, and the confidence of the region. The backbone network of the tower screw recognition model adopts a lightweight CSPDarknet-53 and embeds an SE attention mechanism, the neck network adopts a PAN-FPN structure and fuses multi-scale feature maps, and the output layer adopts an improved Softmax activation function.
21. The tower screw status detection system based on multimodal image fusion according to claim 16, characterized in that, The detection and evaluation module also includes: The primary detection module is used to input component images into the model and perform primary detection on the component images using a preset sliding window to generate screw candidate regions. The secondary detection module is used to perform secondary detection on the screw candidate area and output the probability of the screw's presence. The evaluation module is used to determine whether a screw exists based on the probability of its presence, and if it does exist, to evaluate the screw's condition.
22. The tower screw status detection system based on multimodal image fusion according to claim 16, characterized in that, The positioning module further includes: The coordinate transformation module is used to combine the detection results and the collected multi-source information data to convert the pixel coordinates in the component image into physical coordinates through perspective transformation. The multi-source information data includes GPS coordinates, device attitude angles and camera lens parameters. The 3D point cloud matching module is used to construct a local 3D point cloud from a continuous image sequence using the SfM algorithm, and to match and correct the physical coordinates of the screw with the local 3D point cloud using the ICP algorithm, generating a location information report. The location information report includes the screw's location coordinates, quantity, confidence level, and corresponding component number.
23. An electronic device, characterized in that, include: A processor and a memory, wherein the memory stores a computer program, which is loaded and executed by the processor to implement the tower screw status detection method based on multimodal image fusion as described in any one of claims 1 to 15.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the tower screw status detection method based on multimodal image fusion as described in any one of claims 1 to 15.
Citation Information
Cited By
Full-appearance on-line visual detection method and system for automobile hexagonal head bolt
CN122089710A