Image Fusion Method and Storage Medium for AC Filter Detection
By introducing mutual information extraction module and multi-scale fusion module in the image fusion model, the problems of insufficient complementary information extraction and feature loss in infrared and visible light images are solved, significantly improving the quality of the fusion image and supporting more effective AC filter detection.
Patent Information
- Application Number
- CN202510269964.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Existing deep learning methods have insufficient complementary information extraction and feature loss problems in the fusion of infrared and visible light images, resulting in a degradation of the quality of the fusion image.
An image fusion method is designed, using a pre-trained fusion model, which includes an encoder and a decoder, which includes a first branch network, a second branch network, a mutual information extraction module and a multi-scale fusion module. The mutual information extraction module screens complementary information through the channel attention mechanism and injects it into the branch network; the multi-scale fusion module captures local and global feature information through convolution kernels of different scales.
By improving the extraction effect of complementary information and reducing feature loss, the quality of the fused image is significantly improved and the support for AC filter detection is enhanced.
Smart Images

Figure CN119784612B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an image fusion method and a storage medium for AC filter detection. Background Art
[0002] In recent years, infrared and visible light image fusion technology has received extensive attention in the fields of target detection, monitoring, remote sensing, etc. Infrared images can highlight the thermal radiation characteristics of targets in low-light or harsh environments, while visible light images provide rich detail and texture information. However, due to the different physical characteristics of the two image sources, how to effectively fuse the information of the two to improve the comprehensive expression ability of the image has become a research hotspot.
[0003] As a key device in the power system, AC filters are widely used to suppress harmonic interference in the power grid and improve power quality. In smart grids and transmission monitoring, infrared and visible light image fusion technology has gradually become an important means for equipment status monitoring. In practical applications, infrared images can intuitively reflect the thermal distribution state of AC filters and are used to detect overheating phenomena caused by overload or faults, while visible light images provide structural details of the equipment (such as wiring integrity and appearance damage) to support the evaluation of the appearance and mechanical state of the equipment. Through image fusion technology, the information of the two modalities can be combined to generate a more diagnostically valuable comprehensive image.
[0004] Image fusion technology is mainly divided into traditional algorithms and deep learning algorithms. Traditional algorithms mainly include methods such as multi-scale decomposition and sparse representation. However, these methods often rely on manually designed rules and are difficult to adapt to the non-linear relationships of multi-modal images in complex scenarios, resulting in limited fusion effects. Compared with traditional image processing methods, deep learning-based fusion models can automatically learn the features in images and complete the image fusion task without complex manual design steps. However, existing deep learning methods have problems such as insufficient extraction of complementary information between infrared and visible light images and feature loss, thus reducing the quality of the fused image. Summary of the Invention
[0005] The present invention aims to at least solve one of the technical problems in the related art to some extent. For this reason, an object of the present invention is to provide an image fusion method and a storage medium for AC filter detection to improve the fusion effect of infrared images and visible light images and contribute to subsequent AC filter detection.
[0006] In a first aspect, the present invention provides an image fusion method for AC filter detection, comprising: acquiring an infrared image and a visible light image of an AC filter; inputting the infrared image and the visible light image into a pre-trained fusion model to output a fusion image for AC filter detection; wherein, the fusion model includes an encoder and a decoder, the encoder includes a first branch network, a second branch network, a mutual information extraction module and a multi-scale fusion module, the first branch network is used to extract a first feature of the infrared image, the second branch network is used to extract a second feature of the visible light image, the mutual information extraction module is used to obtain complementary information according to the first feature and the second feature, and use a channel attention mechanism to filter out irrelevant information in the complementary information, and inject the filtered complementary information into the first branch network and the second branch network respectively, so as to splice with the first feature and the second feature respectively to obtain a first spliced feature and a second spliced feature, the multi-scale fusion module is used to perform multi-scale fusion on the first spliced feature and the second spliced feature to obtain a fusion feature, and the decoder is used to decode the fusion feature to obtain a fusion image.
[0007] According to an embodiment of the present invention, the first branch network and the second branch network have the same structure, including 4 first convolutional blocks connected in sequence, and the number of mutual information extraction modules is 3, which correspond to the latter 3 first convolutional blocks one by one; wherein, the mutual information extraction module is used to obtain complementary information according to the first feature and the second feature output by the corresponding first convolutional blocks in the two branch networks, and use a channel attention mechanism to filter out irrelevant information in the complementary information, and inject the filtered complementary information into the first branch network and the second branch network respectively, so as to perform splicing operations with the first feature and the second feature respectively.
[0008] According to an embodiment of the present invention, the mutual information extraction module obtains the feature injected into the first branch network through the following formula:
[0009]
[0010] wherein, , represents the feature containing shared information, , represents the weight, 、 represent the first feature and the second feature input by the corresponding first convolutional blocks in the two branch networks, represents the feature injected into the first branch network, and represent 1×1 and 3×3 convolution operations respectively, represents the connection operation, represents the channel attention mechanism, Represents a global average pooling operation.
[0011] According to an embodiment of the present invention, the mutual information extraction module obtains the features injected into the second branch network through the following formula:
[0012]
[0013] Wherein, , represents the feature containing shared information, , represents the weight, 、 represent the first feature and the second feature corresponding to the input of the first convolutional block in the two branch networks, represents the feature injected into the second branch network, and respectively represent 1×1 and 3×3 convolution operations, represents the concatenation operation, represents the channel attention mechanism, represents a global average pooling operation.
[0014] According to an embodiment of the present invention, the multi-scale fusion module includes a third branch network, a fourth branch network, and a local feature fusion unit, a first global feature fusion unit, and a second global feature fusion unit connected in sequence. The third branch network and the fourth branch network have the same structure, including a second convolutional block, a third convolutional block, and a fourth convolutional block; wherein, the local feature fusion module is used to perform local feature fusion on the first concatenated feature and the second concatenated feature processed by the second convolutional block to obtain a first fusion sub-feature; the first global feature fusion unit is used to perform global feature fusion on the first concatenated feature and the second concatenated feature processed by the third convolutional block, and the first fusion sub-feature to obtain a second fusion sub-feature; the second global feature fusion unit is used to perform global feature fusion on the first concatenated feature and the second concatenated feature processed by the fourth convolutional block, and the second fusion sub-feature to obtain the fusion feature.
[0015] According to an embodiment of the present invention, the local feature fusion unit obtains the first fusion sub-feature through the following formula:
[0016]
[0017] Wherein, and respectively represent the first concatenated feature and the second concatenated feature processed by the second convolutional block, represents the first fusion sub-feature, represents the concatenation operation, and respectively represent 1×1 and 3×3 convolution operations.
[0018] According to an embodiment of the present invention, the structures of the first global feature fusion unit and the second global feature fusion unit are the same. The second global feature fusion unit obtains the second fused sub-feature through the following formula:
[0019]
[0020] where, and respectively represent the first spliced feature and the second spliced feature processed by the third convolution block, represents the second fused sub-feature, , represents the intermediate fused sub-feature, represents the fused feature.
[0021] According to an embodiment of the present invention, the second convolution block includes a 1×1 convolution layer and a LeakyReLU activation layer connected in sequence. The third convolution block includes a 3×3 convolution layer and a LeakyReLU activation layer connected in sequence. The fourth convolution block includes a 5×5 convolution layer and a LeakyReLU activation layer.
[0022] According to an embodiment of the present invention, the decoder includes 4 first convolution blocks and 1 fifth convolution block connected in sequence. The first convolution block includes a 3×3 convolution layer and a LeakyReLU activation layer connected in sequence. The fifth convolution block includes a 3×3 convolution layer and a Tanh activation layer connected in sequence.
[0023] In a second aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the image fusion method for AC filter detection described in the first aspect above.
[0024] For the image fusion method and storage medium for AC filter detection in the embodiments of the present invention, in order to improve the extraction effect of complementary information, a mutual information extraction module is designed. This module combines the channel attention mechanism to fuse and screen the infrared and visible light image features, and can extract complementary information while reducing the interference of irrelevant information. At the same time, in order to reduce feature loss, a multi-scale fusion module is designed. This module captures the feature information at different scales and performs hierarchical fusion, which can improve the final fusion image effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a schematic structural diagram of a fusion model according to an embodiment of the present invention;
[0026] Figure 2Schematic diagram of the mutual information extraction module according to an embodiment of the present invention;
[0027] Figure 3 Schematic diagram of the multi-scale feature fusion module according to an embodiment of the present invention;
[0028] Figure 4 Schematic diagram of the local feature fusion unit according to an embodiment of the present invention;
[0029] Figure 5 Schematic diagram of the second global feature fusion unit according to an embodiment of the present invention;
[0030] Figure 6 Schematic diagram of obtaining a fused image based on an infrared image and a visible light image according to some examples of the present invention. Detailed implementation manners
[0031] In view of the problems of insufficient complementary information extraction and feature loss in the fusion of infrared images and visible light images of AC filters, the present invention proposes an image fusion method for AC filter detection. In this method, in order to improve the extraction effect of complementary information, a mutual information extraction module is designed. This module combines a channel attention mechanism to fuse and screen the features of infrared and visible light images, and can extract complementary information while reducing the interference of irrelevant information; at the same time, in order to reduce feature loss, a multi-scale fusion module is designed. This module captures feature information at different scales and performs hierarchical fusion, which can improve the effect of the final fused image.
[0032] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0033] The following refers to the drawings to describe an image fusion method and a storage medium for AC filter detection according to an embodiment of the present invention.
[0034] In the embodiments of the present invention, the image fusion method for AC filter detection includes the following steps S11 - S12:
[0035] S11, obtaining an infrared image and a visible light image of an AC filter;
[0036] S12, inputting the infrared image and the visible light image into a pre-trained fusion model, and outputting a fused image for AC filter detection.
[0037] Among them, as Figure 1As shown in the figure, the fusion model includes an encoder and a decoder. The encoder includes a first branch network, a second branch network, a mutual information extraction module, and a multi-scale fusion module. The first branch network is used to extract the first features of the infrared image, the second branch network is used to extract the second features of the visible light image, the mutual information extraction module is used to obtain complementary information based on the first features and the second features, and use the channel attention mechanism to filter out the irrelevant information in the complementary information, and inject the filtered complementary information into the first branch network and the second branch network respectively, so as to splice with the first features and the second features respectively to obtain the first spliced features and the second spliced features. The multi-scale fusion module is used to perform multi-scale fusion on the first spliced features and the second spliced features to obtain fusion features, and the decoder is used to decode the fusion features to obtain a fusion image.
[0038] Specifically, referring to Figure 1 , the fusion model is mainly divided into two parts: The first part is the encoder. In the encoder, the input infrared image and visible light image are respectively independently used with branch networks (such as, can include a series of consecutive convolutions) to extract the unique features of the two images respectively. Considering that the shooting scenes of the two images are the same, in order to better mine the complementary information between the two modal images, a mutual information extraction module is designed. This module first effectively extracts the complementary information of the input source images by analyzing the correlation of the two modalities, then uses the channel attention mechanism to filter out the interference of irrelevant information, and finally injects this complementary information into the two branch networks, so as to enhance the complementarity between modalities; At the same time, a multi-scale fusion module is also designed in the encoder, and by performing multi-layer fusion on multi-scale information, the quality of the fusion image is improved. The second part is the decoder. The decoder (such as, using multiple consecutive convolutions) reconstructs the fusion features of the infrared image and the visible light image, and finally obtains a fusion image, which can be used for AC filter detection (such as fault detection). The specific training process of the fusion model is as follows:
[0039] Step 1: Collect the infrared image and visible light image data of the AC filter, and randomly divide the data into a training set and a test set according to the ratio of 7:3 for samples (it can also be other ratios, which can be specifically calibrated according to needs);
[0040] Step 2: Input the training set into the designed fusion model for model training;
[0041] Step 3: After the model training is completed, input the test set and observe the fusion test results.
[0042] Among them, if the test results meet the preset conditions (such as the performance indicators meet the requirements), the trained fusion model can be used for the fusion of infrared images and visible light images, otherwise the fusion model needs to be continuously trained.
[0043] In some embodiments of the present invention, referring to Figure 1, the structures of the first branch network and the second branch network are the same, including 4 first convolutional blocks connected in sequence. The number of mutual information extraction modules is 3, which corresponds one-to-one to the latter 3 first convolutional blocks. Among them, the mutual information extraction module is used to obtain complementary information based on the first feature and the second feature output by the corresponding first convolutional blocks in the two branch networks, and use the channel attention mechanism to filter out the irrelevant information in the complementary information, and inject the filtered complementary information into the first branch network and the second branch network respectively, so as to perform splicing operations with the first feature and the second feature respectively.
[0044] Exemplarily, see Figure 1 , the first convolutional block includes a 3×3 convolutional layer and a LeakyReLU activation layer connected in sequence. Using four consecutive 3×3 convolutions and the LeakyReLU activation function independently for the input infrared image and visible light image respectively can extract the unique features of the two images.
[0045] In some examples of the present invention, the mutual information extraction module obtains the feature injected into the first branch network through the following formula (1):
[0046] (1)
[0047] Wherein, , represents the feature containing shared information, , represents the weight, 、 represent the first feature and the second feature input by the corresponding first convolutional blocks in the two branch networks, represents the feature injected into the first branch network, and respectively represent 1×1 and 3×3 convolution operations, represents the concatenation operation, represents the channel attention mechanism, represents the global average pooling operation.
[0048] In some examples of the present invention, the mutual information extraction module obtains the feature injected into the second branch network through the following formula:
[0049] (2)
[0050] Wherein, , represents the feature containing shared information, , represents the weight, 、 represent the first feature and the second feature input by the corresponding first convolutional blocks in the two branch networks, represents the feature injected into the second branch network, and respectively represent 1×1 and 3×3 convolution operations, represents a concatenation operation, represents a channel attention mechanism, represents a global average pooling operation.
[0051] Specifically, since the infrared image and the visible light image capture the same scene, they contain unique information and shared information. To better learn the complementary information between the two images, a mutual information extraction module is designed to extract the shared information of the two images, and the structure of this module is as Figure 2 shown.
[0052] See Figure 2 , and respectively represent the features corresponding to the input infrared image and visible light image. First, the features are subjected to a convolution concatenation operation and pass through two layers of 1×1 convolution to obtain features rich in shared information, as shown in the following formula (3).
[0053] (3)
[0054] Among them, and respectively represent 1×1 and 3×3 convolution operations, is a concatenation operation. Subsequently, in order to reduce the interference of irrelevant information, the obtained features containing shared information are used with a global average pooling and a channel attention mechanism to obtain the weight , as shown in formula (4).
[0055] (4)
[0056] Among them, is the channel attention mechanism, is the global average pooling operation. In addition, see Figure 2 , can also be subjected to 3×3 convolution to obtain deeper features, and multiply this feature by the weight , and then add it to the feature after 1×1 convolution to obtain features with complementary information, as shown in the above formulas (1) and (2).
[0057] Finally, the obtained features with complementary information are injected into the corresponding branch network.
[0058] In some embodiments of the present invention, such as Figure 3As shown, the multi-scale fusion module includes a third branch network, a fourth branch network, and a local feature fusion unit, a first global feature fusion unit, and a second global feature fusion unit connected in sequence. The third branch network and the fourth branch network have the same structure, including a second convolutional block, a third convolutional block, and a fourth convolutional block.
[0059] Among them, the local feature fusion module is used to perform local feature fusion on the first concatenated feature and the second concatenated feature processed by the second convolutional block to obtain a first fused sub-feature; the first global feature fusion unit is used to perform global feature fusion on the first concatenated feature and the second concatenated feature processed by the third convolutional block, and the first fused sub-feature to obtain a second fused sub-feature; the second global feature fusion unit is used to perform global feature fusion on the first concatenated feature and the second concatenated feature processed by the fourth convolutional block, and the second fused sub-feature to obtain a fused feature.
[0060] Specifically, in order to better capture features, the present invention designs a multi-scale fusion module with a structure as Figure 3 shown Figure 3 in and respectively represent the first concatenated feature (corresponding to the infrared image) and the second concatenated feature (corresponding to the visible light image) input into this module. In this module, convolution kernels with different kernels (1×1, 3×3, 5×5) are used for the first concatenated feature and the second concatenated feature respectively to capture feature information of different scales, and the activation function used is LeakyReLU (LReLU in the figure). Subsequently, these multi-scale features will be input into three fusion units (i.e., the local feature fusion unit, the first global feature fusion unit, and the second global feature fusion unit) for multi-layer feature fusion.
[0061] Exemplarily, the second convolutional block includes a 1×1 convolutional layer and a LeakyReLU activation layer connected in sequence, the third convolutional block includes a 3×3 convolutional layer and a LeakyReLU activation layer connected in sequence, and the fourth convolutional block includes a 5×5 convolutional layer and a LeakyReLU activation layer.
[0062] In some examples of the present invention, the local feature fusion unit obtains the first fused sub-feature through the following formula (5):
[0063] (5)
[0064] Among them, and respectively represent the first concatenated feature and the second concatenated feature processed by the second convolutional block, represents the first fused sub-feature, represents the concatenation operation, and respectively represent 1×1 and 3×3 convolution operations.
[0065] Specifically, the structure of the local feature fusion unit is as Figure 4 shown, Figure 4 in and respectively represent the first and second concatenated features input to the local feature fusion unit after 1×1 convolution, represents the first fusion sub-feature. Refer to Figure 4 , in the local feature fusion unit, first the two concatenated features are concatenated, and then processed in two paths. In the first path, a single convolution operation is used to generate low-dimensional features; in the second path, and operations are successively used to achieve deeper information extraction, as specifically shown in the above formula (5).
[0066] In some examples of the present invention, the structures of the first global feature fusion unit and the second global feature fusion unit are the same, and the second global feature fusion unit obtains the second fusion sub-feature through the following formula (6):
[0067] (6)
[0068] wherein, and respectively represent the first and second concatenated features after being processed by the third convolution block, represents the second fusion sub-feature, , represents the intermediate fusion sub-feature, represents the fusion feature.
[0069] Specifically, considering that smaller convolution kernels are suitable for extracting detailed features and local textures, while larger convolution kernels are more suitable for global image information. Therefore, in the global feature fusion unit, local and global features are combined simultaneously to make up for the problem of insufficient global image information extraction ability caused by the insufficient receptive field of the convolution kernel. The structure of the global feature fusion unit is as Figure 5 shown.
[0070] Taking the second global feature fusion unit as an example, Figure 5 in, and are respectively the first and second concatenated features after the 5×5 convolution operation of the fourth convolution block, is the fusion sub-feature output by the previous fusion unit, is the intermediate fusion sub-feature, is the output fusion feature. In this unit, the input features and Generate intermediate fusion sub - features using the same structure as the local feature fusion unit , and then perform a connection fusion operation on the fusion sub - features of the previous fusion unit, as shown in the above formula (6). Finally, the fusion features output by the second global feature fusion unit are input into the decoder to complete image reconstruction.
[0071] In some embodiments of the present invention, referring to Figure 1 , the decoder includes 4 first convolutional blocks and 1 fifth convolutional block connected in sequence. The first convolutional block includes a 3×3 convolutional layer and a LeakyReLU activation layer connected in sequence, and the fifth convolutional block includes a 3×3 convolutional layer and a Tanh activation layer connected in sequence.
[0072] In addition, Figure 6 shows three examples of the fused image obtained by using the image fusion method for AC filter detection according to the embodiments of the present invention. Referring to Figure 6 , compared with the infrared image and the visible light image, the fused image obtained by fusion can contain more useful information, which helps to improve the efficiency of subsequent AC filter detection (such as fault detection).
[0073] Based on the image fusion method for AC filter detection in the above - mentioned embodiments, the present invention proposes a computer - readable storage medium.
[0074] In this embodiment, a computer program is stored on the computer - readable storage medium. When the computer program is executed by a processor, the image fusion method for AC filter detection in the above - mentioned embodiments is implemented.
[0075] In summary, the image fusion method and storage medium for AC filter detection in the embodiments of the present invention utilize a fusion model of multi - scale fusion of joint mutual information, which can effectively fuse the infrared image and the visible light image of the AC filter. Among them, the mutual information extraction module can fully learn the complementary characteristics of the thermal distribution state and detail information of the input image, and the multi - scale fusion module can comprehensively capture the local and global information of the image. Specifically, it can be as follows:
[0076] 1) Use a fusion model of multi - scale fusion of joint mutual information to fuse the infrared image and the visible light image of the AC filter. This fusion model adopts an encoder - decoder structure. The encoder part designs an independent dual - encoder structure to extract the feature information of the two modalities, and at the same time enhances the complementarity between modalities by continuously injecting into the branch network through the mutual information extraction module; the features output by the two - branch network are effectively fused through the multi - scale fusion module and input into the decoder to complete the reconstruction of the fused image.
[0077] 2) The complementary information between the AC filter infrared image and the visible light image is captured by the mutual information extraction module. Based on the learning of the modal complementary characteristics, this module accurately screens the important features related to the target task by introducing a channel attention mechanism, reducing the interference of irrelevant information on the fusion effect; subsequently, the extracted complementary information is reinjected into the features of each modality, enhancing the correlation between the two modalities.
[0078] 3) The multi-scale fusion module is used to capture the local and global feature information of the image through convolutional kernels of different sizes. Subsequently, by designing a local and global feature fusion unit, the features of different scales are fused, which can fully capture the local details and global information and improve the fusion effect.
[0079] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0080] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0081] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.
[0082] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0083] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0084] In the present invention, unless otherwise clearly specified and defined, the terms such as "mounted", "connected", "connected to", "fixed", etc. should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0085] In the present invention, unless otherwise clearly specified or limited, the first feature being "on" or "under" the second feature may mean that the first and second features are in direct contact, or the first and second features are indirectly in contact via an intermediate medium. Further, the first feature being "above", "over" and "on top of" the second feature may mean that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher level than the second feature in terms of horizontal height. The first feature being "under", "beneath" and "underneath" the second feature may mean that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower level than the second feature in terms of horizontal height.
[0086] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art may make variations, modifications, substitutions and alterations to the above embodiments within the scope of the present invention.
Claims
1. An image fusion method for AC filter detection, characterized in that: include: Acquire infrared and visible light images of the AC filter; Inputting the infrared image and the visible light image into a pre-trained fusion model, and outputting a fusion image for AC filter detection; The fusion model includes an encoder and a decoder, the encoder includes a first branch network, a second branch network, a mutual information extraction module and a multi-scale fusion module, the first branch network is used to extract the first feature of the infrared image, the second branch network is used to extract the second feature of the visible light image, the mutual information extraction module is used to obtain complementary information according to the first feature and the second feature, and use the channel attention mechanism to filter out irrelevant information in the complementary information, and inject the filtered complementary information into the first branch network and the second branch network respectively, so as to respectively splice with the first feature and the second feature to obtain a first splicing feature and a second splicing feature, the multi-scale fusion module is used to perform multi-scale fusion on the first splicing feature and the second splicing feature to obtain a fused feature, and the decoder is used to decode the fused feature to obtain a fused image; The multi-scale fusion module includes a third branch network, a fourth branch network, and a local feature fusion unit, a first global feature fusion unit, and a second global feature fusion unit connected in sequence. The third branch network and the fourth branch network have the same structure, including a second convolution block, a third convolution block, and a fourth convolution block; wherein, The local feature fusion module is used to perform local feature fusion on the first splicing feature and the second splicing feature processed by the second convolution block to obtain a first fused sub-feature; The first global feature fusion unit is used to perform global feature fusion on the first splicing feature and the second splicing feature processed by the third convolution block, and the first fused sub-feature to obtain a second fused sub-feature; The second global feature fusion unit is used to perform global feature fusion on the first splicing feature and the second splicing feature processed by the fourth convolution block, and the second fused sub-feature to obtain the fused feature.
2. The image fusion method for AC filter detection according to claim 1, characterized in that: The first branch network and the second branch network have the same structure, including four first convolution blocks connected in sequence, and the number of the mutual information extraction modules is three, corresponding one to one with the three first convolution blocks that follow; Among them, the mutual information extraction module is used to obtain complementary information according to the first feature and the second feature output by the corresponding first convolutional blocks in the two branch networks, and use the channel attention mechanism to filter out irrelevant information in the complementary information, and inject the filtered complementary information into the first branch network and the second branch network respectively, so as to perform splicing operations with the first feature and the second feature respectively.
3. The image fusion method for AC filter detection according to claim 2, characterized in that: The mutual information extraction module obtains the features injected into the first branch network through the following formula: in, , represents the features containing shared information, , represents the weight, , Represents the first feature and the second feature of the first convolutional block input in the two-branch network. represents the characteristics injected into the first branch network, and Represent 1×1 and 3×3 convolution operations respectively, Indicates a connection operation. represents the channel attention mechanism, Represents a global average pooling operation.
4. The image fusion method for AC filter detection according to claim 2, characterized in that: The mutual information extraction module obtains the features injected into the second branch network through the following formula: in, , represents the features containing shared information, , represents the weight, , Represents the first feature and the second feature of the first convolutional block input in the two-branch network. represents the characteristics injected into the second branch network, and Represent 1×1 and 3×3 convolution operations respectively, Indicates a connection operation. represents the channel attention mechanism, Represents a global average pooling operation.
5. The image fusion method for AC filter detection according to claim 1, characterized in that: The local feature fusion unit obtains the first fusion sub-feature by the following formula: in, and represent the first splicing feature and the second splicing feature after being processed by the second convolution block, respectively, represents the first fusion sub-feature, Indicates a connection operation. and Represent 1×1 and 3×3 convolution operations respectively.
6. The image fusion method for AC filter detection according to claim 1, characterized in that: The first global feature fusion unit and the second global feature fusion unit have the same structure. The second global feature fusion unit obtains the second fusion sub-feature by the following formula: in, and Respectively represent the first splicing feature and the second splicing feature after being processed by the third convolution block, represents the second fusion sub-characteristic, , represents the intermediate fusion sub-feature, represents the fusion feature.
7. The image fusion method for AC filter detection according to claim 1, characterized in that: The second convolution block includes a 1×1 convolution layer and a LeakyReLU activation layer connected in sequence, the third convolution block includes a 3×3 convolution layer and a LeakyReLU activation layer connected in sequence, and the fourth convolution block includes a 5×5 convolution layer and a LeakyReLU activation layer.
8. The image fusion method for AC filter detection according to claim 1, characterized in that: The decoder includes four first convolution blocks and one fifth convolution block connected in sequence, the first convolution block includes a 3×3 convolution layer and a LeakyReLU activation layer connected in sequence, and the fifth convolution block includes a 3×3 convolution layer and a Tanh activation layer connected in sequence.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image fusion method for AC filter detection according to any one of claims 1-8 is implemented.