Electrical equipment state early warning method and system fusing infrared and visible light images

By constructing an infrared-dominated dual-branch feature extraction structure and a feature space alignment mechanism for cross-modal difference representation, the problem of insufficient modal difference in the fusion of infrared and visible light images is solved, enabling high-precision identification and stable early warning of overheating defects in electrical equipment, and improving the real-time performance and reliability of detection.

CN122492623APending Publication Date: 2026-07-31STATE GRID JIANGSU ELECTRIC POWER CO LTD NANJING POWER SUPPLY COMPANY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIANGSU ELECTRIC POWER CO LTD NANJING POWER SUPPLY COMPANY
Filing Date
2026-05-09
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies do not adequately consider the modal differences between infrared and visible light images, lack targeted fusion methods, and struggle to accurately distinguish between local temperature anomalies and the overall operating status of electrical equipment. Furthermore, the algorithms are highly complex, making it difficult to balance detection accuracy and real-time performance.

Method used

A dual-branch feature extraction structure based on infrared images is constructed. Infrared features are enhanced through a feature enhancement module. A feature space alignment mechanism for cross-modal difference representation is introduced to achieve multi-level feature fusion of infrared and visible light images. Temperature anomaly determination is then performed based on the fused features.

Benefits of technology

It significantly improves the detection accuracy and stability of overheating defects in electrical equipment, reduces the risk of false alarms and missed alarms, and enhances the real-time performance and reliability of detection. It is suitable for complex backgrounds and conditions with multiple heat sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492623A_ABST
    Figure CN122492623A_ABST
Patent Text Reader

Abstract

A method and system for early warning of electrical equipment status by fusing infrared and visible light images are disclosed. Based on infrared and visible light images of the electrical equipment to be inspected, a dual-branch feature extraction network is established to extract multi-level infrared image features and visible light structural features, with additional enhancement of the infrared image features. On a level-by-level basis, candidate region features from the two modalities are fused, outputting level-by-level cross-modal fused features, which are then aggregated between layers. A temperature status determination model is invoked to analyze the aggregated features and output the anomaly identification results of the electrical equipment. This application enables more targeted deep fusion at the feature level, thereby significantly improving the detection accuracy and stability of local overheating anomalies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of electrical equipment operation status monitoring and intelligent early warning technology, specifically involving an electrical equipment status early warning method and system that integrates infrared and visible light images. Background Technology

[0002] Infrared imaging technology is widely used in electrical equipment inspection and condition monitoring because it can acquire surface temperature distribution information of equipment non-contactly. To compensate for the shortcomings of infrared imaging, some existing technologies introduce visible light images as an auxiliary information source. By overlaying or simply registering infrared and visible light images, the visualization of abnormal areas can be enhanced.

[0003] In recent years, with the development of computer vision and intelligent information processing technologies, some methods based on target detection or feature learning have been used to fuse infrared and visible light structural features for electrical equipment condition identification. These methods typically achieve the detection and localization of abnormal targets by automatically extracting and analyzing image features. For example, patent application CN115115592A simultaneously acquires visible light and infrared images of electrical equipment and fuses them according to their difference weights to obtain a fused image, enhancing imaging performance in low-light conditions and thus improving the accuracy of electrical equipment defect detection. Another example is patent application CN121147208A, which utilizes the structural information of visible light images and the thermal distribution information of infrared images. Through intelligent matching and grouping, multi-view feature fusion, and thermal distribution overlay analysis, combined with dual-stream collaborative analysis and feature pyramid extraction using a deep learning model, it achieves deep mining and accurate identification of multimodal fault features.

[0004] However, existing methods generally suffer from the following problems: First, they do not adequately consider the modal differences between infrared and visible light images, and the fusion method lacks specificity, which can easily lead to the weakening of key information; second, they have limited ability to model the correlation between local temperature anomalies and the overall operating status of the equipment, making it difficult to accurately distinguish between real faults and environmental interference; and third, in practical engineering applications, the algorithm complexity is high, making it difficult to balance detection accuracy and real-time requirements. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a multimodal intelligent detection method that can fully exploit the complementary characteristics of infrared image temperature information and visible light image structural information. This method solves the problems of insufficient modal difference modeling, single fusion level, inaccurate abnormal area localization, and difficulty in balancing detection accuracy and real-time performance in existing technologies. As a result, it enables high-precision identification, accurate localization, and stable early warning of overheating defects in electrical equipment.

[0006] The first aspect of this application discloses a method for early warning of electrical equipment status based on the fusion of infrared and visible light images, employing the following technical solution: acquiring infrared and visible light images of the electrical equipment and preprocessing them; inputting the preprocessed infrared and visible light images into a dual-branch feature extraction network with consistent structure and independent parameters for multi-level feature extraction, wherein the infrared branch enhances the features through a feature enhancement module; performing candidate region extraction and spatial deviation calibration on the multi-level infrared and visible light features to achieve cross-modal spatial alignment; employing an infrared-dominant and visible light-assisted fusion strategy at each level for weighted fusion, and finally fusing the multi-level features to obtain fused features; judging the temperature status of the electrical equipment based on the fused features, and outputting an abnormal warning result and corresponding confidence level when an abnormal temperature region is detected.

[0007] The second aspect of this application discloses an electrical equipment status early warning system based on the fusion of infrared and visible light images. This system implements the technical solution of the electrical equipment status early warning method based on the fusion of infrared and visible light images described in the first aspect of this application. The system includes: an image acquisition and preprocessing module, a dual-branch multi-level feature extraction module, an infrared feature enhancement and modeling module, a cross-modal spatial alignment module, a feature fusion module, and a status determination and early warning output module. Each module works collaboratively to achieve multi-scale, multi-level feature extraction from infrared and visible light images, infrared anomaly feature enhancement, candidate region spatial calibration, and functional complementary fusion. Based on the fused features, the system analyzes the temperature status of the electrical equipment and outputs anomaly early warning results and corresponding confidence levels.

[0008] The core innovations of this invention are as follows, and the key improvements that distinguish it from existing technologies are: (1) Construct a dual-branch feature extraction structure dominated by infrared image features. The dual branches perform feature modeling on infrared images and visible light images respectively. The infrared branch focuses on enhancing the representation capability of temperature anomaly-related features, while the visible light branch focuses on extracting equipment structure and spatial location information, realizing differentiated modeling of modal functional division rather than simple symmetrical structure extraction.

[0009] (2) In the dual-modal feature fusion stage, a feature space alignment mechanism based on cross-modal difference representation is introduced. By performing regional-level difference calculations on the multi-level enhancement features of infrared images and the corresponding visible light structural features in a unified feature representation space, a difference representation reflecting the degree of inconsistency between the two modes in spatial response positions is obtained. Based on the difference representation, the position offset, scale offset, and angle offset between cross-modal features are predicted. After geometric calibration of the center position, region size, and rotation angle of the visible light modality candidate region according to the predicted offset, it participates in the subsequent fusion process, thereby realizing explicit spatial alignment fusion with infrared thermal response distribution as a reference, improving the positioning accuracy of abnormal regions and reducing information interference caused by structural misalignment between modes.

[0010] (3) Construct an anomaly judgment mechanism for electrical equipment status early warning. Based on the fused feature information, directly output the location of the abnormal area, the degree of abnormality and the corresponding confidence level, and generate early warning results to realize closed-loop processing from feature fusion to status assessment, thereby improving the reliability and interpretability in engineering applications.

[0011] The beneficial effects of this invention lie in its construction of a dual-branch feature extraction structure oriented towards infrared and visible light modal characteristics, and the introduction of a difference-guided cross-modal fusion mechanism in the fusion stage. This allows infrared temperature information, visible light structural information, and the difference representation between the two modalities to simultaneously participate in feature expression and decision-making processes. Compared to traditional methods that only perform image overlay or shallow feature stitching, this invention can achieve more targeted deep fusion at the feature level, effectively mitigating the information attenuation problem caused by modal distribution differences, and preventing key anomalous features from being masked or diluted during the fusion process. This significantly improves the detection accuracy and stability of local overheating anomalies.

[0012] Meanwhile, by directly determining temperature anomalies based on fused features and outputting corresponding confidence information, the anomaly identification results possess quantifiable credibility evaluation capabilities. This is beneficial for hierarchical early warning and operation and maintenance decisions in practical engineering scenarios, reducing the risk of false alarms and missed alarms. Furthermore, the invention emphasizes independent branch modeling and differential constraint coordination in the fusion stage in its structural design, enabling the model to maintain high robustness and generalization ability even under complex backgrounds, multiple heat source interference, and dense equipment structures. In summary, this invention represents a substantial improvement over existing technologies in terms of fusion strategy, anomaly determination mechanism, and engineering adaptability. It not only improves the accuracy and reliability of electrical equipment overheating defect detection but also meets real-time requirements, demonstrating significant technological advancement and application value. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the network structure in this invention; Figure 2 This is a structural diagram of the VSS Block in this invention; Figure 3 This is a structural diagram of the infrared image feature enhancement module in this invention; Figure 4 This is a diagram of the cross-modal feature fusion module in this invention. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The embodiments described in this application are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.

[0015] While existing technologies have incorporated visible light images as auxiliary information sources and achieved fusion processing of infrared and visible light images through image overlay, registration, or attention fusion, most methods focus on improving visual effects or detection accuracy. They fail to establish a feature-driven mechanism centered on infrared anomaly information for scenarios involving abnormal temperatures in electrical equipment, nor do they explicitly model and finely align the spatial offset relationship between the two modal features. Existing fusion methods often lack a clear design for the functional division between infrared temperature anomaly regions and visible light structural information, making it difficult to achieve accurate localization and reliable determination of anomaly regions in complex contexts.

[0016] Therefore, there is an urgent need for an infrared and visible light collaborative processing method for electrical equipment status early warning scenarios. In the fusion process, the temperature anomaly information in the infrared image is the main factor, and the structure and position information of the visible light image are used as auxiliary constraints through explicit spatial alignment mechanism and difference guidance strategy to guide the fusion features to more accurately characterize the potential abnormal areas of the equipment. On this basis, a status judgment and early warning output mechanism is built to improve the anomaly recognition accuracy and engineering applicability.

[0017] This invention is proposed in response to the aforementioned technical requirements. By constructing an infrared-dominated dual-branch feature modeling structure, a difference-guided cross-modal alignment mechanism, and a functionally complementary fusion strategy, it achieves accurate identification and stable early warning of thermal anomaly areas in electrical equipment, effectively overcoming the problems of coarse fusion methods, unstable anomaly judgment, and insufficient system closed loop in existing technologies.

[0018] As an embodiment of this application, a specific implementation method for an electrical equipment status early warning method based on the fusion of infrared and visible light images is disclosed. The execution flow of the method embodiment is as follows: Figure 1 .

[0019] S1: As one implementation method, infrared and visible light images of the electrical equipment to be monitored are acquired. The infrared and visible light images are acquired synchronously to ensure the consistency of the dual-modal data in time and space.

[0020] In step 1, infrared and visible light images of the electrical equipment to be monitored are acquired at the same time or under similar operating conditions. The infrared images can stably reflect the temperature distribution and thermal anomaly information on the surface of the equipment, while the visible light images are used to reflect the structural shape, spatial location, and appearance of the equipment.

[0021] To ensure high consistency of dual-modal data in terms of time and spatial perspective, infrared and visible light images are acquired through synchronously triggered imaging devices or fixed installation methods to ensure the spatial correspondence of the same device location in different modal images.

[0022] Through this multimodal data acquisition, complementary images covering the thermal state and structural information of the equipment are obtained, laying a solid data foundation for subsequent multimodal feature extraction and fusion.

[0023] S2: As one implementation method, the infrared image and visible light image obtained in step 1 are preprocessed, including the following steps: 2.1: Spatial alignment processing is performed on dual-modal images based on the device structure outline or significant region features to correct the spatial offset relationship under different imaging modalities; Based on the structural outline or salient regional features of electrical equipment, a feature matching algorithm is used to spatially align infrared and visible light images, correcting spatial offsets caused by imaging equipment installation errors or differences in shooting angles, and ensuring that the same equipment part has a corresponding spatial position relationship in different modal images.

[0024] To correct infrared images With visible light images The spatial offset between them is estimated using a feature matching algorithm to estimate the transformation matrix. This involves mapping the visible light image coordinate system to the infrared image coordinate system, achieving spatial alignment. Specifically: First, extract the infrared image. With visible light images Key point set and , respectively represented as and , ; The number of feature points in each set. and The index of the feature point in the two sets.

[0025] Furthermore, corresponding point pairs are established through feature descriptor matching. , For point-to-point quantity, Point pair indexing. Estimate the affine transformation matrix using matching point pairs: ; in, and These are translation parameters; , , and The linear transformation parameters control the rotation, scaling, skewing, and cropping of the image; the translation and linear transformation parameters are solved using the least squares method to minimize the error between transformed point pairs, as follows: ; Furthermore, after obtaining the transformation matrix, the visible light image is mapped to the infrared image coordinate system through the transformation matrix, represented as follows: ;in, This is a spatially aligned visible light image. This is the original visible light image.

[0026] Bilinear interpolation is used to perform resampling and achieve spatial alignment.

[0027] 2.2: Scale the image before performing advanced noise reduction; After spatial alignment is completed, the infrared and visible light images are scaled uniformly. Based on the preset target resolution and size specifications, the dual-modal images are resized and resampled to ensure that the infrared and visible light images are consistent in pixel size, resolution and image ratio.

[0028] By unifying the scale, we can ensure that the same part of electrical equipment has a consistent spatial scale in images of different modalities. This avoids scale mismatch problems in subsequent feature extraction and cross-modal fusion due to resolution differences, and provides a stable spatial foundation for collaborative analysis of multimodal information.

[0029] In a further embodiment, to address the random thermal noise and local temperature fluctuations introduced by the thermal imaging mechanism in infrared images, a smoothing and denoising process is performed on the infrared images to suppress noise interference while preserving key temperature structure information. Specifically: The scale-unified infrared image is represented as ;in, This indicates the distribution of thermal radiation signals. This represents the random noise component caused by sensor noise and environmental interference. Considering the continuity and consistency of thermal radiation distribution within the same physical area, a weighted smoothing method based on local thermal radiation consistency is used for infrared images. (Pixel...) Centered on, in its neighborhood Perform weighted average processing within: ; in, These are the pixel values ​​of the denoised infrared image; For the neighborhood Inner pixel value; This is a weighting function used to measure the similarity between neighboring pixels and the center pixel in terms of spatial location and thermal radiation value; This is the normalization factor.

[0030] The weighting function is defined as:

[0031] in, Spatial proximity weights are used to control the influence of spatial distance on the weighting function; This is a radiation similarity weight used to control the impact of temperature differences on the weighting. The meaning of this weighting function is that for a pixel to obtain a high weight, it must simultaneously satisfy two conditions: physical proximity to the center pixel and similar temperature value. This method achieves effective smoothing in regions with gentle temperature changes and automatically reduces the cross-regional average in regions with abrupt temperature changes, thereby avoiding excessive edge blurring.

[0032] S3: As one implementation method, the infrared image and visible light image preprocessed in S2 are respectively input into a dual-branch feature extraction network to independently extract multi-level features of the two modalities. The infrared branch highlights temperature anomalies and suppresses background noise through feature enhancement and long-range dependency modeling, thereby improving the ability to represent abnormal regions.

[0033] In a further embodiment, the dual-branch feature extraction network extracts features from infrared images and visible light images respectively through a parallel dual-branch architecture; the basic feature extraction structures of the two branches are the same, but their internal parameters can be learned independently. The basic feature extraction structures of both branches employ multiple cascaded VSS Blocks (Visual State Space Blocks), with a downsampling layer set after each VSS Block.

[0034] The cascading and downsampling of VSS Blocks progressively reduce the spatial resolution and enhance semantic abstraction of features in each branch, thus forming a multi-level representation. In the infrared image branch, this process constitutes the infrared basic feature extraction unit, used to generate primary infrared features.

[0035] 3.1: In one embodiment, the processing logic of the VSS Block is briefly described. The VSS Block structure diagram is as follows: Figure 2 As shown: Each VSS Block consists of a linear normalization layer, a linear layer, a depthwise separable convolutional layer, a two-dimensional selective state space module (SS2D), and a residual connection structure. The processing flow of each VSS-Block is described below: The input features are first standardized by a linear normalization layer, and then processed by two branches: a main branch and a secondary branch. In the main branch: the standardized features are channel-mapped through a linear layer, and further processed by depthwise separable convolutions to extract local spatial features. The resulting features are then input into the SS2D module to model long-range dependencies within a two-dimensional space. The features representing these long-range dependencies are then reprocessed by a linear normalization layer and output, completing the main branch's processing.

[0036] For the sub-branch, the standard-processed features are passed through another linear layer to generate gated weight features; The feature output of the secondary branch is multiplied element-wise with the output of the main branch to achieve gating control of the feature response of the main branch. This process can be represented as follows:

[0037] in, The main branch outputs features. The gating weight features are output by the secondary branch.

[0038] After gating, the obtained features are further processed through a linear normalization layer to adjust the feature distribution, and then through a linear layer to transform the channel dimensions, thereby obtaining a more expressive feature representation. Finally, this feature is added element-wise to the input features through a residual connection structure to obtain the output feature of the VSS Block.

[0039] 3.2: Specifically, the infrared image branch includes a basic feature extraction unit and an infrared feature enhancement modeling unit. The basic feature extraction unit uses a cascaded structure of multiple VSS Blocks to extract infrared image features, obtaining primary infrared features. The infrared feature enhancement modeling unit is used to enhance the primary infrared features through multiple levels, obtaining enhanced infrared features. Specifically, the infrared feature enhancement modeling unit module is as follows: Figure 3 As shown.

[0040] See Figure 3To address the issue of insufficient feature representation capability of infrared images under complex conditions such as weak texture, uneven thermal radiation distribution, and blurred boundaries of abnormal regions, this application chooses to connect a downsampling layer after each VSS Block structure to form a base layer. That is, the preprocessed infrared image is input into the first VSS Block structure for processing, then passes through the first downsampling layer, and the features output from the first downsampling layer are input into the second VSS Block structure for processing, then through the second downsampling layer, and so on. Each base layer outputs the infrared features of that layer. For example, for the... The infrared features output at each level are denoted as follows: , , The number of basic levels (VB modules).

[0041] For each basic level's output of hierarchical infrared features, these are input into the infrared feature enhancement modeling module for enhanced modeling, thereby improving the ability to represent and identify thermal anomaly regions. For example, for the... Hierarchical infrared features output at each level The process of enhancing modeling can be described as follows: First of all, Perform dimensional transformation to obtain feature representations suitable for local window modeling. ; Furthermore, multiple local attention windows of different scales are set up to... Perform multi-scale context modeling. Assumptions are set. A local attention window, then for infrared features , will output Enhancement features at each scale, will Enhancement features at various scales are fused to improve the responsiveness of temperature anomaly regions across different spatial scales. The fusion process can be represented as follows: ; in, Indicates hierarchical infrared features Multi-scale local enhancement features, Indicates hierarchical infrared features In the Enhanced features at various scales Indicates the first The scale parameters of a local window; Indicates scale-based The local feature modeling operator is used to extract local contextual information within the corresponding receptive range.

[0042] In this embodiment, the local modeling results at each scale are integrated using an average fusion method to combine local response information at different scales and suppress feature bias that may be caused by single-scale modeling.

[0043] Furthermore, for multi-scale local enhancement features Normalization is performed to stabilize the feature distribution and improve the robustness of subsequent feature modeling; then the normalized features are restored to their dimensions to map them back to a feature representation space that is compatible with subsequent enhancement modules.

[0044] Furthermore, to compensate for the shortcomings of local feature modeling in large-scale spatial correlation modeling, multi-scale local enhancement features after dimensionality restoration (with...) are further... (This is represented as input to the Mamba state space module to model long-distance correlations between different regions, thereby enhancing the overall perception of cross-regional thermal anomaly patterns. Specifically:) The Mamba state-space module models long-range dependencies using a selective state-space model, represented as follows: ; This represents the infrared enhancement features after long-range dependency modeling. This represents a long-range dependency modeling operator based on the Mamba state-space model, used to capture long-distance relationships between different regions of a device, thereby enhancing the overall perception of cross-regional thermal anomaly patterns.

[0045] Furthermore, calling the residual join will and Fusion, output fusion features This approach aims to enhance cross-regional association representations while preserving local augmentation information. Finally, it addresses the fusion features. Global attention enhancement is performed to further strengthen the correlation between thermal anomaly regions and the overall scene, and the final infrared enhancement features are output.

[0046] Finally, the infrared image branch output Infrared enhancement features at multiple levels.

[0047] 3.3: The visible light image branch adopts the same basic feature extraction method as the infrared image branch structure, which consists of multiple VSS Blocks and downsampling layers cascaded alternately to extract multi-level structural features of the visible light image layer by layer. After each VSS Block output, it passes through the corresponding downsampling layer to obtain the input features of the next level, and finally outputs multiple levels of visible light structural features corresponding to the infrared branch levels.

[0048] The visible light structural features output from the visible light branch and the infrared enhancement output from the infrared image branch are jointly fed into the dual-modal fusion module for fusion. The visible light image branch focuses on maintaining the integrity of the device's structural outline, component shape, and spatial location information.

[0049] S4: As one implementation method of the embodiment, see Figure 4 Cross-modal fusion processing is performed on the multi-level infrared image features and visible light structural features obtained by S3.

[0050] 4.1: In a further implementation, cross-modal spatial consistency calibration and feature extraction are performed on the infrared enhancement features and visible light structural features of multiple levels output by S3. Specifically: First, calculate the feature response distribution for each level. For each level... Including infrared enhancement features and visible light structural features The response intensity at each location in the infrared enhancement feature map is calculated, and regions with response intensity higher than a set value are identified as high-response regions. These regions are the candidate regions of interest, representing the parts of the image with the strongest response to anomalies, typically corresponding to local areas of electrical equipment where thermal anomalies may exist. For example, high-response regions in infrared enhancement features usually correspond to areas of temperature anomalies.

[0051] Subsequently, using the infrared candidate region of interest as a spatial reference, corresponding candidate regions are determined in the visible light structural features of the corresponding level. The RoIAlign rotation operation is then used to extract features from both the infrared candidate region and the initial visible light corresponding region. The infrared candidate region features primarily characterize the thermal anomaly response information within the candidate region, while the initial visible light corresponding region features primarily characterize the edge, texture, connectivity, and structural layout information within the corresponding region.

[0052] Furthermore, feature mapping is performed on the infrared candidate region features and the corresponding visible light candidate regions extracted at each level using a shared feature representation space; Within a unified feature space, infrared candidate region features and their corresponding visible light candidate region features are subjected to element-wise difference, splicing, or similarity measurement operations to construct a difference representation. This difference representation includes response difference information, structural distribution difference information, and similarity association information between the infrared and visible light candidate region features within the unified feature space, expressed as the spatial position offset, scale variation, and directional variation trend of the two modes within the same candidate region.

[0053] These difference characterizations will be used to characterize the spatial inconsistencies between two modalities in the same candidate region, such as the difference between the thermal response of an infrared image and the device structure response of a visible light image.

[0054] Furthermore, using the infrared candidate region as the spatial alignment benchmark, for each level, based on the difference representation, three types of deviations are predicted through a fully connected layer: Positional deviation Translational shift caused by imaging equipment errors or differences in viewing angle; Scale bias Inconsistent region sizes are caused by variations in imaging resolution or device distance across different modalities. Angle deviation : Directional shift caused by changes in shooting angle.

[0055] These three types of deviations are used as spatial transformation parameters and applied to the center coordinates, size parameters, and rotation angle of the corresponding visible light candidate regions at the corresponding levels. This updates the center position, size, and rotation angle of the visible light candidate regions. Based on the updated region parameters, the Rotation RoIAlign is invoked to re-extract the aligned visible light candidate region features, thereby ensuring spatial correspondence between the high-response regions in the infrared mode and the device structure regions in the visible light mode. After completing the spatial consistency calibration, the infrared candidate region features and the visible light candidate region features correspond in both region position and scale, thus providing aligned input for subsequent cross-modal fusion.

[0056] 4.2: In a further implementation, cross-modal fusion is performed on the calibrated infrared and visible light corresponding candidate region features output in step 4.1.

[0057] For each level, for the corresponding infrared candidate regions and the updated visible light candidate regions within that level, their corresponding candidate region features are extracted and then paired and fused one by one; based on the level... For example, the fusion process can be represented as: ; In the formula, Indicates hierarchy The cross-modal fusion feature corresponding to the k-th candidate region and They are respectively levels The corresponding k-th infrared candidate region feature and the updated visible light corresponding candidate region feature; hierarchical The fusion adjustment coefficient is used to control the information contribution ratio of infrared candidate region features and updated visible light corresponding candidate region features in different levels of fusion processes.

[0058] Further explanation reveals that lower-level features contain richer spatial details, and visible light features provide relatively strong supplementary information on device structural boundaries, connection nodes, and local geometric relationships. Therefore, at lower levels... The possible value is 0.4; however, features at the mid-to-high level mainly provide semantic information, while the visible light mode has a weaker supplementary effect on the overall structure. The value is relatively low, and can be 0.2. By weighted integration of fusion features at different levels, the sensitivity of the infrared mode to abnormal regions and the ability of the visible light mode to characterize equipment structural information can be preserved simultaneously.

[0059] Furthermore, at each level, after cross-modal fusion is completed... The features of each candidate region are aggregated within a level. Specifically, firstly, the importance score of the fused features of each candidate region under that level is calculated, and the weight coefficient of each candidate region is obtained through normalization. Then, the fused features of multiple candidate regions within that level are weighted and summed according to the weight coefficients to obtain the aggregated fused features of that level.

[0060] Let the first Under each level The fusion features corresponding to each candidate region are: Its importance score is calculated using a fully connected layer and is expressed as: ; in, and These are used to represent the weight parameters and bias parameters for importance score calculation, respectively.

[0061] Subsequently, on the first The importance scores of all candidate regions at each level are normalized to obtain the corresponding weight parameters. , can be represented as: ; in, Indicates the first The number of candidate regions at each level, and satisfying the following conditions. After obtaining the weight coefficients of each candidate region, the weight coefficients of the first candidate region are then calculated. The aggregated fusion feature of multiple candidate regions at a given level is obtained by weighted summation of their fusion features. , represented as: ; By aggregating within the above hierarchy, the contribution of candidate regions with more significant abnormal responses or more critical structural information to the feature representation of that hierarchy can be highlighted.

[0062] 4.3: In a further implementation, inter-layer fusion and feature sufficiency are performed on the cross-modal fusion features output at each level. Specifically: First, the cross-modal fusion features from each level are adjusted to a uniform spatial resolution. Then, the features from each level are concatenated along the channel dimension. A 1x1 convolution operation is then used to compress the channels and reorganize the information in the concatenated features to obtain the final fusion features. This process can be represented as follows: ; This indicates the final inter-layer fusion characteristics; This indicates a splicing operation along the channel dimension. This indicates a fusion operator based on 1x1 convolution, used to post-process inter-layer aggregated features (including channel compression and information recombination) and output the final fused features. Through this inter-layer fusion, the complementary integration of the advantages of features from different layers can be achieved, thereby improving the comprehensive representation ability of candidate regions for potential abnormal targets.

[0063] S5: As one implementation method, based on inter-layer fusion features, the temperature status of electrical equipment is determined and anomaly warnings are issued. When an abnormal temperature region is detected, an anomaly warning result is output, along with the corresponding anomaly confidence level, to indicate potential operational risks to the equipment. This includes the following steps: First, the fusion features corresponding to each candidate region of the interlayer fusion features output by S4 are input into the temperature state determination model to analyze the temperature anomaly response information in the candidate regions and obtain the anomaly response score corresponding to each candidate region.

[0064] The temperature state determination model is used to jointly discriminate the infrared thermal response information and visible light structure information contained in the fused features, so as to improve the accuracy of identifying potential temperature anomaly regions.

[0065] Secondly, based on the abnormal response scores, anomaly determination is performed on each candidate region. When the abnormal response score corresponding to a candidate region exceeds a preset determination threshold, the candidate region is determined to be a temperature abnormality region; when the abnormal response score does not exceed the preset determination threshold, the candidate region is determined to be in a normal state or no obvious temperature abnormality is observed. The preset determination threshold can be set according to the statistical results of training samples, the performance optimization results of the validation set, or the actual application requirements.

[0066] For candidate regions identified as temperature anomaly areas, corresponding anomaly warning results are output. The anomaly warning results include at least the location of the anomaly region and the anomaly confidence level; in some embodiments, they may further include anomaly severity information. The anomaly confidence level can be calculated from the probability value, normalized score, or anomaly response score output by the temperature state determination model, and is used to characterize the degree of confidence that a temperature anomaly exists in the corresponding candidate region.

[0067] The temperature state determination model can be implemented using a determination head consisting of convolutional layers, fully connected layers, and activation functions. The convolutional layers are used to compress features and integrate local information from the fused features obtained in S4; the fully connected layers are used to discriminate abnormal responses in candidate regions; and the activation functions are used to output corresponding abnormal response scores or abnormal probability values, thereby achieving the classification and determination of the temperature state of candidate regions.

[0068] After determining the temperature status of each candidate area, the anomaly determination results can be backmapped to the original image or equipment area to form anomaly warning information for specific equipment parts, thereby providing support for equipment maintenance decisions, risk management and operational status monitoring.

[0069] As an embodiment of this application, a system for early warning of electrical equipment status that integrates infrared and visible light images is disclosed. Employing the specific implementation method described above for early warning of electrical equipment status, the system includes: A data acquisition and preprocessing module is used to simultaneously acquire infrared and visible light images of the electrical equipment to be tested and perform preprocessing. The preprocessing includes spatial alignment, scale unification, and noise reduction. In the noise reduction process, based on the continuous and consistent characteristics of thermal radiation distribution in the same physical region, the infrared image is subjected to local thermal radiation consistency weighted smoothing. A dual-branch feature extraction module is used to establish a dual-branch feature extraction network including an infrared branch and a visible light branch. The infrared branch extracts the basic infrared features of each level from the preprocessed infrared image in a cascaded manner of multiple basic feature extraction layers. After spatial modeling enhancement, the infrared enhanced features of each level are output. The visible light branch extracts the visible light structural features of each level from the preprocessed visible light image in the same cascaded manner as the infrared branch. A hierarchical cross-modal fusion module is used to perform cross-modal fusion processing on infrared enhancement features and visible light structural features. At each level, infrared candidate regions are identified by calculating the response intensity in the infrared enhancement features, and corresponding visible light candidate regions are determined to complete spatial consistency calibration. The calibrated infrared candidate regions and visible light candidate regions are fused to output hierarchical cross-modal fusion features. The inter-layer cross-modal fusion module is used to concatenate and post-process the inter-layer cross-modal fusion features output from each layer along the channel dimension to obtain inter-layer aggregated features. The inter-layer aggregated features corresponding to each candidate region are input into a pre-trained temperature determination model. The temperature determination model uses a determination head composed of convolutional layers, fully connected layers, and activation functions to calculate and output the abnormal response score corresponding to each candidate region.

[0070] As an embodiment of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it employs the specific implementation described in the above-described electrical equipment status early warning method.

[0071] As an embodiment of this application, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, it employs the specific implementation method described above for the electrical equipment status early warning method.

[0072] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0073] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.

[0074] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0075] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0076] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A method for early warning of electrical equipment status by fusing infrared and visible light images, characterized in that, include: Simultaneously acquire infrared and visible light images of the electrical equipment to be inspected and perform preprocessing; the preprocessing includes spatial alignment, scale unification, and noise reduction. In the noise reduction process, based on the continuous and consistent characteristic of thermal radiation distribution in the same physical region, the infrared image is subjected to local thermal radiation consistency weighted smoothing. A dual-branch feature extraction network including an infrared branch and a visible light branch is established. The infrared branch extracts the basic infrared features of each level from the preprocessed infrared image in a cascaded manner of multiple basic feature extraction layers. After spatial modeling enhancement, the infrared enhanced features of each level are output. The visible light branch uses the same basic hierarchical cascading method as the infrared branch to extract the visible light structural features of each level from the preprocessed visible light image; Perform cross-modal fusion processing on infrared enhancement features and visible light structural features; At each level, infrared candidate regions are identified by calculating the response intensity in the infrared enhancement features, and corresponding visible light candidate regions are determined to complete spatial consistency calibration. The calibrated infrared and visible light candidate regions are then fused to output hierarchical cross-modal fusion features. The inter-modal fusion features output from each level are spliced ​​and post-processed along the channel dimension to obtain inter-layer aggregated features. The inter-layer aggregated features corresponding to each candidate region are input into a pre-trained temperature determination model. The temperature determination model uses a determination head composed of convolutional layers, fully connected layers, and activation functions to calculate and output the abnormal response score corresponding to each candidate region.

2. The method for early warning of electrical equipment status by fusing infrared and visible light images according to claim 1, characterized in that, The infrared branch includes a basic feature extraction unit and an enhancement modeling unit; the basic feature extraction unit uses a cascaded structure of multiple basic levels to extract infrared image features, and each basic level consists of a VSS Block followed by a downsampling layer; Each base layer outputs the infrared image features of that layer, and an enhancement modeling unit is connected after each base layer to enhance the infrared image features output by the corresponding layer.

3. The method for early warning of electrical equipment status by fusing infrared and visible light images according to claim 2, characterized in that, Infrared image features are enhanced through enhanced modeling units, including: For the first The infrared image features output at each level are used to perform multi-scale contextual modeling by setting multi-scale local attention windows, with each window outputting an enhanced feature at a corresponding scale. The enhanced features of all scale windows are fused to output the hierarchy. Multi-scale local enhancement features; state-space modeling is then performed on the multi-scale local enhancement features to output infrared enhancement features; The number of infrared enhancement features corresponds to the number of layers.

4. The method for early warning of electrical equipment status by fusing infrared and visible light images according to claim 1, characterized in that, The visible light branch includes: It adopts the same cascaded structure as the basic feature extraction unit of the infrared branch, and outputs multiple levels of visible light structural features; the optical structural features and infrared enhancement features can be mapped one-to-one at the level.

5. The method for early warning of electrical equipment status by fusing infrared and visible light images according to claim 1, characterized in that, The cross-modal fusion process includes: For hierarchy The infrared enhancement features are used to calculate the response intensity at each location in the feature map, and the regions with response intensities higher than a set value are identified as candidate regions of interest. Using infrared candidate regions of interest as spatial references, visible light candidate regions of interest are determined in the visible light structural features of the corresponding level; candidate region features are extracted for the two types of candidate regions of interest respectively, and feature mapping is performed using a shared feature representation space to output the difference representation. The difference characterization includes the spatial position offset, scale variation, and directional variation trend of infrared candidate region features and visible light corresponding candidate region features in the same candidate region.

6. The method for early warning of electrical equipment status by fusing infrared and visible light images according to claim 5, characterized in that, The cross-modal fusion processing also includes: For each level, positional bias, scale bias, and angular bias are predicted based on the difference characterization, and these three types of biases are used as spatial transformation parameters and applied to the center coordinates, size parameters, and rotation angle of the visible light corresponding candidate region of the corresponding level to update the corresponding parameters of the visible light corresponding candidate region; For the corresponding infrared candidate regions of interest and the corresponding updated visible light candidate regions of interest, their corresponding candidate region features are extracted and then matched and fused one by one.

7. The method for early warning of electrical equipment status by fusing infrared and visible light images according to claim 5, characterized in that, The cross-modal fusion processing also includes: At each level, all fused candidate region features are weighted and aggregated to output the hierarchical cross-modal fusion features; the weight of each fused candidate region feature is the normalized value of its importance score.

8. An electrical equipment status early warning system integrating infrared and visible light images, executing the electrical equipment status early warning method as described in any one of claims 1-7, characterized in that, The system includes: A data acquisition and preprocessing module is used to simultaneously acquire infrared and visible light images of the electrical equipment to be tested and perform preprocessing. The preprocessing includes spatial alignment, scale unification, and noise reduction. In the noise reduction process, based on the continuous and consistent characteristics of thermal radiation distribution in the same physical region, the infrared image is subjected to local thermal radiation consistency weighted smoothing. A dual-branch feature extraction module is used to establish a dual-branch feature extraction network including an infrared branch and a visible light branch. The infrared branch extracts the basic infrared features of each level from the preprocessed infrared image in a cascaded manner of multiple basic feature extraction layers. After spatial modeling enhancement, the infrared enhanced features of each level are output. The visible light branch extracts the visible light structural features of each level from the preprocessed visible light image in the same cascaded manner as the infrared branch. A hierarchical cross-modal fusion module is used to perform cross-modal fusion processing on infrared enhancement features and visible light structural features. At each level, infrared candidate regions are identified by calculating the response intensity in the infrared enhancement features, and corresponding visible light candidate regions are determined to complete spatial consistency calibration. The calibrated infrared candidate regions and visible light candidate regions are fused to output hierarchical cross-modal fusion features. The inter-layer cross-modal fusion module is used to concatenate and post-process the inter-layer cross-modal fusion features output from each layer along the channel dimension to obtain inter-layer aggregated features. The inter-layer aggregated features corresponding to each candidate region are input into a pre-trained temperature determination model. The temperature determination model uses a determination head composed of convolutional layers, fully connected layers, and activation functions to calculate and output the abnormal response score corresponding to each candidate region.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements the electrical equipment status early warning method according to any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the electrical equipment status early warning method according to any one of claims 1-7.