Infrared visible light image data fusion method
Through cross-modal image registration and fusion technology, the problem of large workload of infrared inspection data analysis of drone is solved, intelligent identification of hidden heat hazards on transmission lines is realized, the efficiency and accuracy of grid equipment inspection are improved, and the application of big data technology in the power grid is promoted.
Patent Information
- Application Number
- CN202311699242.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-07-08
AI Technical Summary
The infrared inspection data analysis of drones is large and has low efficiency. The existing infrared-visible image registration method is insufficient in accuracy and efficiency in power grid equipment inspection, making it difficult to meet the complex environmental needs of power grid systems.
The cross-modal image registration technology is adopted to realize automatic infrared visible light registration through the MobileViT module, and the cross-modal registration image fusion method is used, combined with the Boundary Loss algorithm to optimize the edge cutting loss function to perform data fusion between infrared images and visible light images.
It realizes intelligent identification of hidden heat hazards in transmission lines, improves the efficiency and accuracy of grid equipment inspection, promotes the application of big data technology in transmission lines, and ensures the safe and stable operation of the power grid.
Smart Images

Figure CN120278890A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of infrared-visible light fusion optics, and particularly relates to a method for fusing infrared and visible light image data. Background Art
[0002] In recent years, to meet the demand for electricity in national economic and social development, new requirements have been put forward for the stability of power grid equipment operation. At the same time, along with the digital transformation and upgrading of the power grid system, artificial intelligence technology has gradually been applied in the inspection work of power grid equipment.
[0003] As the front-end technology of artificial intelligence, intelligent perception aims to detect the environmental information of the external space by using a variety of sensing devices, which can provide an information basis for the subsequent intelligent decision-making tasks and contribute to the realization of end-to-end artificial intelligence applications. Among them, vision is the most intuitive and important perception way, and the digital images obtained by imaging sensing devices are important carriers and manifestations of visual information. However, the power grid system environment is extremely complex. How to accurately capture the perception object and obtain effective information is an important prerequisite for the successful application of artificial intelligence technology. Summary of the Invention
[0004] The embodiments of this application provide a method for fusing infrared and visible light image data, which realizes the intelligent identification of typical heating hazards in transmission lines, and conducts research on infrared-visible light image calibration and fusion technology, identification technology of common heating components under limited sample conditions, research on infrared thermal hazard identification technology under the condition of dual-light fusion, research on relevant standards for the fusion analysis of unmanned aerial vehicle (UAV) infrared-visible light inspection data, development of algorithm modules for the fusion analysis of UAV infrared-visible light inspection data, etc. The research results can be applied to fields such as the defect identification and analysis of transmission lines, solve the problems of large workload and low efficiency in the data analysis of UAV infrared inspections, greatly promote the application of big data technology in transmission lines, and at the same time improve the safe and stable operation level of the power grid and ensure national economic and social development.
[0005] The embodiments of this application provide a method for fusing infrared and visible light image data, including:
[0006] Collect visible light images and infrared images captured by a dual-light camera;
[0007] Obtain the feature points in the visible light image that match the infrared image;
[0008] According to the feature points, achieve automatic registration of infrared and visible light through the registration technology of cross-modal images;
[0009] After registering the infrared image and the visible light image, use the cross-modal registration image fusion method to realize the data fusion of the mutually registered infrared image and visible light image.
[0010] In a feasible implementation, the method for obtaining the feature points in the visible light image that match the infrared image includes:
[0011] Preprocess the visible light image and the infrared image to obtain a processed visible light image and a processed infrared image;
[0012] Calculate the feature values of each pixel point in the processed visible light image and the processed infrared image;
[0013] Select the pixel points with the same area as the selected area of the infrared pixel points as the feature points according to the magnitudes of the feature values.
[0014] In a feasible implementation, the method for preprocessing the visible light image and the infrared image includes:
[0015] Perform noise reduction and smoothing on the visible light image; frame the infrared image area, extract the image information within these areas, and convert it into a vector with a fixed dimension.
[0016] In a feasible implementation, the automatic registration of infrared and visible light is achieved through the registration technology of cross-modal images by introducing the MobileViT module into this network and designing and implementing its embedding into the network model;
[0017] The MobileViT network uses Vision Transformer as a convolution to extract global feature information.
[0018] In a feasible implementation, the structure of the MobileViT module consists of three parts in total:
[0019] The first module is the local feature information encoding module, which encodes local information through two layers of convolution;
[0020] The second module is the global feature information module, which establishes global feature connections through Transformer;
[0021] The third module is the feature information fusion module, which contains skip connections to accelerate model fitting and extracts useful feature information of the image during the processing using a small number of parameters.
[0022] In a feasible implementation, a loss function situation will occur during the fusion process. Since the data area of the infrared image is larger than the area after preprocessing the visible light, the Boundary Loss algorithm is used to calculate the loss function for edge cutting. First, calculate the boundary point Dg:
[0023] Dg = ||t - zg(t)||;
[0024] Then, calculate to obtain αg(t):
[0025]
[0026] Finally, calculate to obtain the loss function L calculation:
[0027] L = δ∫ μ αg(t)qβ(t)dt;
[0028] where t is any point on μ, μ represents the spatial domain, zg(t) represents the nearest point of t to the contour g, ‖·‖ represents the L2 norm, αg is a representation of the boundary of the GT region g, qβ(t) is the output of the software, and δ represents the balance coefficient of the edge loss, taking 0.3.
[0029] An infrared-visible light image data fusion method provided by an embodiment of the present application realizes automatic registration of infrared and visible light through cross-modal image registration technology; after registering the infrared image and the visible light image, an image fusion method for cross-modal registered images is used to realize image fusion. The aim is to realize intelligent identification of typical heating hazards in transmission lines, and carry out research on infrared-visible light image calibration and fusion technology, identification technology of common heating components under limited sample conditions, research on infrared thermal hazard identification technology under the condition of dual-light fusion, research on relevant standards for unmanned aerial vehicle (UAV) infrared-visible light inspection data fusion analysis, development of algorithm modules for UAV infrared-visible light inspection data fusion analysis, etc. The research results can be applied to fields such as defect identification and analysis of transmission lines, solve the problems of large workload and low efficiency in data analysis of UAV infrared inspections, greatly promote the application of big data technology in transmission lines, and at the same time improve the safe and stable operation level of the power grid and ensure national economic and social development. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a schematic structural diagram of the infrared-visible light image data fusion method provided by the present application;
[0031] Figure 2 is a schematic diagram of a fused image under different loss functions in this method;
[0032] Figure 3 is an image quality assessment table. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] To enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0034] Unmanned aerial vehicle (UAV) infrared inspection refers to the technology of using a UAV equipped with an infrared camera to inspect a target area. The infrared camera can capture the infrared radiation emitted by an object and convert it into a visible image or video. This technology can be used to detect temperature changes in the target area, detect abnormal situations, detect equipment failures, etc.
[0035] The advantages of UAV infrared inspection include:
[0036] 1. High efficiency: The UAV can quickly inspect the target area, greatly shortening the inspection time.
[0037] 2. Safety: The UAV can avoid personnel entering dangerous areas for inspection, reducing the risk of casualties.
[0038] 3. Comprehensiveness: The UAV can inspect the target area from different angles, providing more comprehensive information.
[0039] The UAV infrared inspection technology has wide applications in the fields of electric power, petroleum, chemical industry, construction, etc. For example, in the electric power field, the UAV infrared inspection technology can be used to detect the temperature changes of high-voltage lines and timely discover potential faults; in the petroleum field, the UAV infrared inspection technology can be used to detect the temperature changes of oil wells and timely discover problems such as leaks.
[0040] Currently, there are problems of large data analysis workload and low efficiency in UAV infrared inspection.
[0041] The following will detail the specific structure of the infrared and visible light image data fusion method provided in this application in conjunction with the accompanying drawings.
[0042] Embodiment 1:
[0043] Referring to Figure 1 as shown, the embodiment of this application provides an infrared and visible light image data fusion method, including:
[0044] S100: Collect the visible light image and infrared image captured by the UAV dual-camera.
[0045] The visible light image and the infrared image are two images of the same object, at the same moment, and from the same angle.
[0046] S200: Obtain the feature points in the visible light image that match the infrared image.
[0047] S300: Based on the feature points, achieve automatic registration of infrared and visible light through cross-modal image registration technology;
[0048] Cross-modal image registration technology refers to the technology of aligning or matching images of different modalities. Here, "modality" can refer to different imaging methods, such as medical imaging modalities like CT, MRI, PET, or non-medical imaging modalities like visible light, infrared, ultrasound, etc.
[0049] The purpose of cross-modal image registration is to align images of different modalities in the same coordinate system for image fusion, image analysis, or other applications. Since images of different modalities may have different resolutions, contrasts, noise levels, etc., cross-modal image registration is a challenging problem.
[0050] The following are some common cross-modal image registration technologies:
[0051] 1. Feature-based registration: This method extracts feature points (such as key points, edges, regions, etc.) in the image and uses these feature points for matching and alignment. Common feature extraction methods include SIFT, SURF, ORB, etc.
[0052] 2. Gray-level-based registration: This method performs registration by comparing the gray-level information of the images. Common gray-level registration methods include mutual information, normalized cross-correlation, SSD, etc.
[0053] 3. Deep learning-based registration: This method uses deep learning models to learn the mapping relationship between images of different modalities and uses this mapping relationship for registration. Common deep learning registration methods include methods based on convolutional neural networks, methods based on generative adversarial networks, etc.
[0054] Cross-modal image registration technology has wide applications in the fields of medical image analysis, computer vision, robotics, etc.
[0055] S400: After registering the infrared image and the visible light image, use the cross-modal registration-oriented image fusion method to achieve image fusion.
[0056] The cross-modal registration-oriented image fusion method is a method of fusing images of different modalities into a new image after registration. Here, "modality" can refer to different imaging methods, such as medical imaging modalities like CT, MRI, PET, or non-medical imaging modalities like visible light, infrared, ultrasound, etc.
[0057] The purpose of cross-modal registration image fusion is to merge useful information from images of different modalities into one image for better image analysis or other applications. Since images of different modalities may have different resolutions, contrasts, noise levels, etc., cross-modal registration image fusion is a challenging problem.
[0058] The cross-modal registration image fusion method adopts the following method:
[0059] Pixel-based fusion method: This method fuses images of different modalities at the pixel level. Common pixel-level fusion methods include weighted average, maximum, minimum, multi-modal fusion, etc.
[0060] Or, feature-based fusion method: This method extracts feature points (such as key points, edges, regions, etc.) from images and uses these feature points for fusion. Common feature-level fusion methods include fusions based on feature extraction methods such as SIFT, SURF, ORB, etc.
[0061] Or, deep learning-based fusion method: This method uses a deep learning model to learn the mapping relationship between images of different modalities and uses this mapping relationship for fusion. Common deep learning fusion methods include methods based on convolutional neural networks, methods based on generative adversarial networks, etc.
[0062] The cross-modal registration image fusion method has extensive applications in the fields of medical image analysis, computer vision, robotics, etc.
[0063] Visible light data acquisition is to first preprocess the visible light image and the infrared image, then calculate the feature value of each pixel point in the preprocessed image, and select the pixel points with the same selection area as the infrared pixel points as feature points according to the size of the feature value;
[0064] Preprocessing is to denoise and smooth the visible light image, frame the infrared image area, extract the image information within these areas, and convert it into a vector of a fixed dimension.
[0065] To achieve automatic registration of infrared and visible light through the cross-modal image registration technology, a MobileViT module is introduced into the network, and its embedding into the network model is designed and implemented;
[0066] The MobileViT network uses the Vision Transformer as a convolution to extract global feature information.
[0067] The structure of the MobileViT module consists of three parts in total:
[0068] The first module is the local feature information encoding module, which encodes local information through two layers of convolution;
[0069] The second module is the global feature information module, which establishes global feature connections through Transformer;
[0070] The third module is the feature information fusion module, which contains skip connections to accelerate model fitting and extracts useful feature information of the image during processing using a small number of parameters.
[0071] During the fusion process, there will be a situation of the loss function. Since the infrared image data region is larger than the region after visible light preprocessing, the Boundary Loss algorithm is used to calculate the loss function for edge cutting. First, calculate the boundary point Dg:
[0072] Dg = ||t - zg(t)||;
[0073] Then calculate αg(t):
[0074]
[0075] Finally, calculate the loss function L calculation:
[0076] L = δ∫ μ αg(t)qβ(t)dt;
[0077] Among them, t is an arbitrary point on μ, μ represents the spatial domain, zg(t) represents the nearest point of t to the contour g, ‖·‖ represents the L2 norm, αg is a representation of the boundary of the GT region g, qβ(t) is the output of the software, and δ represents the balance coefficient of the edge loss, taking 0.3.
[0078] The infrared and visible light image registration technology is applied to the power equipment monitoring system, which can achieve accurate positioning while realizing the thermal image display of power equipment. The purpose of image registration is to obtain the spatial mapping relationship of different images and align the spatial positions of the same target in different images. It is a preprocessing step for image fusion. And combined with image recognition technology, the multi-modal information of the equipment can be fused into a single image to improve the visualization degree and diagnosis efficiency of equipment information. The existing infrared and visible light image registration methods can be divided into three categories: region-based, feature-based, and deep learning-based. The region-based registration method mainly uses the image correlation after Fourier transform and the mutual information and gradient information of multi-modal images to complete registration. This method depends on the linear correlation degree of image gray levels and the field of view overlap degree, and has poor adaptability to complex scenes with viewing angle, spectral differences, and distortions, and has a high computational complexity; the feature-based method uses point features, contour edges, and region features to complete registration by constructing feature descriptors.
[0079] Study the influence degree of the research angle on the registration of infrared and visible light images, including the accuracy of registration and fusion, etc., and propose the requirements and strategies for infrared and visible light data acquisition to provide effective data support for subsequent data fusion analysis.
[0080] Research on the automatic registration method of infrared-visible light images: Analyze the image information differences caused by different imaging principles during the shooting process of current dual-camera, propose a multi-modal image analysis method, study the registration technology of cross-modal images, and realize the automatic registration of infrared and visible light images.
[0081] Research on the intelligent fusion method of infrared-visible light: Analyze the reasons for the poor performance of existing cross-modal image fusion technologies, propose a new image fusion method for cross-modal registration images, and realize the precise fusion of infrared images and visible light images.
[0082] An infrared-visible light image data fusion method provided by an embodiment of this application. The purpose of the present invention is to realize the intelligent identification of typical heating hazards in transmission lines, and carry out research on infrared-visible light image calibration and fusion technology, research on the identification technology of common heating components under limited sample conditions, research on infrared thermal hazard identification technology under the condition of dual-light fusion, research on relevant standards for the fusion analysis of unmanned aerial vehicle (UAV) infrared-visible light inspection data, and development of algorithm modules for the fusion analysis of UAV infrared-visible light inspection data. The research results can be applied to fields such as defect identification and analysis of transmission lines, solve the problems of large workload and low efficiency in the data analysis of UAV infrared inspections, greatly promote the application of big data technology in transmission lines, and at the same time improve the safe and stable operation level of the power grid and ensure national economic and social development.
[0083] Embodiment 2:
[0084] This embodiment mainly introduces the construction and standards of infrared-visible light features on the target contour area. In order to verify the effectiveness and superiority of the algorithm proposed by the present invention, the following experiments are designed:
[0085] An ablation experiment was carried out on the improved DeepLabv3+ network on the dataset, and evaluations were made from subjective evaluation and objective evaluation indicators.
[0086] DeepLabv3+ is a deep learning model for semantic segmentation, and it has achieved good results on the PASCAL VOC 2012 dataset. The ablation experiment is a method for evaluating the contribution of each component in the model to the final performance.
[0087] Figure 2 It is a schematic diagram of the fused image under different loss functions in this method. Figure 3 It is an image quality evaluation table. Figure 3Among them, EN, SD, MI, Nr, SCD, and MS-SSIM refer to a method for evaluating image quality, where:
[0088] ● EN: It is the abbreviation of "Entropy", representing the information entropy of the image.
[0089] ● SD: It is the abbreviation of "Spatial Diversity", representing the spatial diversity of the image.
[0090] ● MI: It is the abbreviation of "Mutual Information", representing the mutual information of the image.
[0091] ● Nr: It is the abbreviation of "Normalized", representing normalization.
[0092] ● SCD: It is the abbreviation of "Structural Similarity Index", representing the structural similarity index.
[0093] ● MS-SSIM: It is the abbreviation of "Multi-Scale Structural Similarity Index", representing the multi-scale structural similarity index.
[0094] These metrics are usually used for image quality assessment, measuring the similarity or quality of images by calculating certain features of the images. The specific calculation methods and application scenarios may vary depending on different fields and requirements.
[0095] From Figure 2 and Figure 3 it can be seen that different δ parameters have a greater impact on the effect of blower image fusion. When δ is greater than 0.3, the characteristic information of the blower infrared has basically disappeared, which is not conducive to engineers' fault location and detection at this time, indicating that the fusion effect at this time does not meet the expectations; when δ is less than 0.3, it cannot be determined from the visual effect what the specific value of the parameter δ is when the effect reaches the best. Therefore, analyze it from objective evaluation metrics, such as Figure 3 the values shown in the table. The optimal and sub-optimal values can be determined through the values. For example: in the EN item, 6.9248 is the optimal value and 6.8952 is the sub-optimal value; in the SCD item, 1.9075 is the optimal value and 1.8927 is the sub-optimal value.
[0096] It is easy to understand that those skilled in the art can combine, split, and reorganize the embodiments of the present application based on the several embodiments provided in the present application to obtain other embodiments, and these embodiments do not exceed the protection scope of the present application.
[0097] The above specific embodiments have further elaborated in detail the objectives, technical solutions, and beneficial effects of the embodiments of the present application. It should be understood that the above are only specific embodiments of the embodiments of the present application and are not used to limit the protection scope of the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of the present application shall be included within the protection scope of the embodiments of the present application.
Claims
1. An infrared and visible light image data fusion method, characterized in that: Include; Collect visible light images and infrared images captured by a dual - light camera; Obtain the feature points in the visible light image that match the infrared image; According to the feature points, achieve automatic registration of infrared and visible light through cross - modal image registration technology; After registering the infrared image and the visible light image, use a cross - modal registration image fusion method to achieve data fusion of the mutually registered infrared image and visible light image.
2. The infrared and visible light image data fusion method according to claim 1, wherein: The method for obtaining the feature points in the visible light image that match the infrared image includes: Pre - process the visible light image and the infrared image to obtain a pre - processed visible light image and a pre - processed infrared image; Calculate the feature values of each pixel point of the pre - processed visible light image and the pre - processed infrared image; Select the pixel points with the same selection area as the infrared pixel points as feature points according to the size of the feature values.
3. The infrared and visible light image data fusion method according to claim 2, wherein: The method for pre - processing the visible light image and the infrared image includes: Perform noise reduction and smoothing on the visible light image; frame the infrared image area, extract the image information within these areas, and convert it into a vector of a fixed dimension.
4. The infrared and visible light image data fusion method according to claim 1 or 3, wherein: The realization of automatic registration of infrared and visible light through cross - modal image registration technology is to introduce a MobileViT module into the network and design and implement its embedding into the network model; The MobileViT network uses Vision Transformer as a convolution to extract global feature information.
5. The infrared and visible light image data fusion method according to claim 4, wherein: The structure of the MobileViT module consists of three parts in total: The first module is a local feature information encoding module, which encodes local information through two - layer convolution; The second module is a global feature information module, which establishes global feature connections through Transformer; The third module is a feature information fusion module, which contains skip connections to accelerate model fitting and extracts useful feature information of the image during processing with a small number of parameters.
6. The infrared and visible light image data fusion method according to claim 5, wherein: In the fusion process, there will be a situation of loss function. Since the data area of the infrared image is larger than the area after pre - processing of the visible light, the Boundary Loss algorithm is used to calculate the loss function of edge cutting. First, calculate the boundary point Dg: Dg = ||t - zg(t)||; Then calculate αg(t): Finally, calculate the loss function L calculation: L = δ∫ μ αg(t)qβ(t)dt; Where t is an arbitrary point on μ, μ represents the spatial domain, zg(t) represents the nearest point of t to the contour g, ‖·‖ represents the L2 norm, αg is a representation of the boundary of the GT region g, qβ(t) is the output of the software, and δ represents the balance coefficient of the edge loss, taking 0.3.
Citation Information
Cited By
Detection method and device for multi-light inspection fusion data of gas insulation combination switch, and medium
CN122391056A