Diffusion model-based photovoltaic module defect image fusion method and system

The diffusion model-based method addresses the imbalance in existing solar panel image fusion by integrating spatial and channel attention to enhance texture and color details, improving defect detection accuracy.

CN120318100AActive Publication Date: 2025-07-15UNIV OF JINAN
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510395793.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing photovoltaic panel defect detection methods are difficult to effectively integrate infrared images and visible light images, resulting in excessive emphasis on thermal radiation information in infrared images and ignoring the texture details and color information of visible light images, affecting the accuracy of defect detection.

Method used

The diffusion model is used to learn the image features of the photovoltaic panels, combined with the residual network, denoising diffusion model and attention mechanism, through forward and reverse diffusion training, a multi-channel diffusion feature map is generated and spatial and channel-weighted fusion fusion is performed, and a fusion image with rich details is finally generated.

Benefits of technology

The accuracy of photovoltaic panel defect detection is improved, and by learning more visible image visual information and retaining infrared image information, it reduces feature information loss, and improves the detection effect of downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318100A_ABST
    Figure CN120318100A_ABST
Patent Text Reader

Abstract

The invention discloses a photovoltaic module defect image fusion method and system based on a diffusion model. The method comprises the following steps: respectively obtaining binary images of corresponding photovoltaic panel areas according to an infrared image and a visible light image of a photovoltaic panel; combining the binary image of the photovoltaic panel area with the corresponding infrared image and visible light image, inputting the combined image into the trained residual network, outputting a feature image, and extracting the photovoltaic panel area; the connected images are subjected to diffusion training through an improved denoising diffusion probability model to obtain a denoising network model; and obtaining a multi-channel diffusion feature map by using the de-noising network model, respectively obtaining a space weighted feature map and a channel weighted feature map by using a space attention mechanism and a channel attention mechanism, and obtaining a final fusion image. The final fusion image obtained through the method not only has rich details, but also has defect information, and the photovoltaic panel defect detection accuracy of a downstream task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and in particular to a photovoltaic module defect image fusion method and system based on a diffusion model. Background Art

[0002] As a clean energy source, solar energy is increasingly favored. With the continuous increase of the photovoltaic market, the value of the defect detection technology of photovoltaic modules has become increasingly prominent. Especially for the defect detection of photovoltaic panels, more and more researchers are committed to the research of the defect detection technology of photovoltaic panels. When performing the defect detection of photovoltaic panels, the infrared images of photovoltaic panels obtained by traditional methods can capture the thermal radiation of photovoltaic panels, but are susceptible to noise and difficult to capture the detailed information of defects. On the contrary, the visible light images of photovoltaic panels obtained usually contain rich structural and defect texture information, but are prone to problems such as occlusion or reflection. Their complementarity makes it possible for the fused image to contain the thermal object and the texture details of the defect. Therefore, fusing the infrared image and the visible light image of the photovoltaic panel is one of the problems to be solved in the defect detection of photovoltaic panels.

[0003] With the rapid development of deep learning technology, artificial intelligence technology represented by deep learning brings new possibilities for the fusion of infrared images and visible light images of photovoltaic panels. By installing an infrared camera and a visible light camera on an unmanned aerial vehicle, rapid segmentation of the photovoltaic panel can be achieved using image segmentation technology, and the images of different light sources can be adjusted to the same angle through image preprocessing. However, fusion models such as convolutional neural networks or GANs may overemphasize the thermal radiation information of the infrared image when fusing images, while ignoring the texture details and color information of the visible light image. However, in the downstream task applications of photovoltaic panel defects, not only the thermal radiation information of the infrared image can reflect the defects of the photovoltaic panel, but the texture details and color information in the visible light image are helpful for identifying the surface features of the object and providing richer information for image analysis. Moreover, there are still defect problems such as reflection and occlusion in the visible light image. Therefore, overemphasizing the thermal radiation information of the infrared image by fusion models such as convolutional neural networks or GANs will lead to a decrease in the practicality of the fused image. Moreover, existing photovoltaic image fusion methods are difficult to focus on the defects of the infrared photovoltaic image during fusion, but rather balance the global thermal radiation information and visible light information during the fusion process. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention provides a photovoltaic module defect image fusion method and system based on a diffusion model. The present invention uses a diffusion model to learn the features of the fine-tuned photovoltaic panel image, and then performs a fusion operation after learning to complete the fusion of the infrared image and the visible light image of the photovoltaic panel.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] In a first aspect, the present invention provides a method for fusing photovoltaic module defect images based on a diffusion model, including:

[0007] Obtain the infrared image and visible light image of the photovoltaic panel, and perform preprocessing; according to the preprocessed infrared image and visible light image, obtain the binary images of the corresponding photovoltaic panel regions respectively;

[0008] Merge the binary images of the photovoltaic panel regions with the corresponding infrared images and visible light images, and input them into the trained residual network to output a feature image; screen out the candidate regions of the photovoltaic panel according to the feature image, generate the corresponding masks of the candidate regions of the photovoltaic panel, and extract the photovoltaic panel regions;

[0009] Connect the infrared image and visible light image after extracting the photovoltaic panel regions, and perform forward diffusion training and reverse diffusion training through an improved denoising diffusion probability model to obtain a denoising network model;

[0010] Use the denoising network model to obtain a multi-channel diffusion feature map, based on the multi-channel diffusion feature map, and respectively obtain a spatial weighted feature map and a channel weighted feature map by using a spatial attention mechanism and a channel attention mechanism, and obtain a final fusion image according to the spatial weighted feature map and the channel weighted feature map.

[0011] As a further technical solution, the feature image output by the trained residual network contains the texture features of the photovoltaic panel region; perform forward propagation on the feature image, and screen out the candidate regions of the photovoltaic panel with a probability greater than 0.7 through the non-maximum suppression method.

[0012] As a further technical solution, align the features of all screened candidate regions of the photovoltaic panel to a fixed size, input them into the trained mask head network, output the corresponding masks of the candidate regions of the photovoltaic panel, and then combine them with the corresponding infrared images and visible light images to extract the photovoltaic panel regions.

[0013] As a further technical solution, the forward diffusion training has T steps, and a subset S{x 1t ,x 2t ,...,x st} is selected from 1 to T, and Gaussian noise is gradually added until it approaches pure noise. The specific formula is as follows:

[0014] Among them, and respectively represent adding the x is th and the k is -1th Gaussian noise, Z represents the standard normal distribution, Denoted as controlling x is is the variance of the Gaussian noise added in is , and γ represents the Gaussian noise.

[0015] As a further technical solution, the reverse diffusion training is a process of denoising the noise added in the forward diffusion training process. The specific formula is as follows:

[0016]

[0017]

[0018] where ε θ represents the noise prediction network, σ represents the variance of represents the mean of

[0019] As a further technical solution, the infrared image and visible light image of the photovoltaic panel are used to generate an attention map by using the cross-attention mechanism and used as a constraint condition in the reverse diffusion training process. The specific formula is as follows: Q = I vis , K = I ir , V = I ir ;

[0020] As a further technical solution, the specific method for obtaining the spatial weighted feature map by using the spatial attention mechanism is as follows: First, perform two parallel convolution operations on the input multi-channel diffusion feature map and add the two feature maps; then, through a sigmoid activation function, obtain the spatial attention weight; finally, multiply the input multi-channel diffusion feature map by the spatial attention weight M s to obtain the spatial weighted feature map;

[0021] The specific method for obtaining the channel weighted feature map by using the channel attention mechanism is as follows: First, perform global average pooling and global maximum pooling on the input multi-channel diffusion feature map to obtain two one-dimensional vectors respectively; then, splice the two vectors, pass through a fully connected layer, then through a ReLU activation function, then through another fully connected layer, and finally through a sigmoid activation function to obtain the channel attention weight; multiply the input multi-channel diffusion feature map by the channel attention weight to obtain the channel weighted feature map.

[0022] In a second aspect, the present invention provides a photovoltaic module defect image fusion system based on a diffusion model, including the following modules:

[0023] An image acquisition module, configured to: acquire the infrared image and visible light image of the photovoltaic panel and perform preprocessing; respectively obtain the binary images of the corresponding photovoltaic panel regions according to the preprocessed infrared image and visible light image;

[0024] The photovoltaic panel area extraction module is configured to: merge the binary image of the photovoltaic panel area with the corresponding infrared image and visible light image, and input them into the trained residual network to output a feature image; screen out the candidate photovoltaic panel areas according to the feature image, generate the corresponding masks of the candidate photovoltaic panel areas, and extract the photovoltaic panel areas;

[0025] The diffusion training module is configured to: connect the infrared image and the visible light image after extracting the photovoltaic panel area, and perform forward diffusion training and reverse diffusion training through an improved denoising diffusion probability model to obtain a denoising network model;

[0026] The image fusion module is configured to: obtain a multi-channel diffusion feature map by using the denoising network model, and based on the multi-channel diffusion feature map, respectively obtain a spatially weighted feature map and a channel weighted feature map by using a spatial attention mechanism and a channel attention mechanism, and obtain a final fused image according to the spatially weighted feature map and the channel weighted feature map.

[0027] One or more technical solutions of the present invention have the following beneficial effects:

[0028] 1. By performing diffusion training (forward diffusion training and reverse diffusion training) on the obtained infrared image and visible light image of the photovoltaic panel, the present invention enables the model to learn as much visual information in the visible light image as possible, such as color, structural details, etc., making up for the loss of too much visual information in the visible light in the prior art methods, while retaining more visual information of the infrared image.

[0029] 2. In the process of fusing the feature maps, the present invention performs fusion based on spatial attention and channel attention. The advantage of this is to reduce the loss of feature information in the fused multi-scale feature maps. The final fused image obtained by the fusion method of the present invention not only has rich details but also has defect information, improving the accuracy of photovoltaic panel defect detection in downstream tasks.

[0030] 3. The present invention generates an attention map by using a cross-attention mechanism for the infrared image and visible light image of the photovoltaic panel, and uses it as a constraint condition in the reverse diffusion training process. It can focus on the defect areas in the infrared image, learn more features of the infrared image in the defect areas, and learn more features of the visible light image in the non-defect areas, so that the learned features can more accurately represent the information of the infrared and visible light images. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0032] Figure 1 It is a schematic flowchart of the photovoltaic module defect image fusion method based on the diffusion model in the first embodiment of the present invention;

[0033] Figure 2 It is a schematic framework diagram of the photovoltaic module defect image fusion method based on the diffusion model in the first embodiment of the present invention;

[0034] Figure 3 It is a schematic diagram of the diffusion model training process in the first embodiment of the present invention;

[0035] Figure 4 It is a schematic diagram of the image fusion process using diffusion features in the first embodiment of the present invention. Detailed implementation manners

[0036] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0037] First embodiment

[0038] In this embodiment, a photovoltaic module defect image fusion method based on the diffusion model is provided. As Figure 1 shown, the specific method includes the following steps:

[0039] S1: Obtain the infrared image and visible light image of the photovoltaic panel and perform preprocessing; according to the preprocessed infrared image and visible light image, obtain the corresponding binary images of the photovoltaic panel area respectively;

[0040] S2: Merge the binary images of the photovoltaic panel area with the corresponding infrared image and visible light image, and input them into the trained residual network to output a feature image; screen out the candidate areas of the photovoltaic panel according to the feature image, generate the corresponding masks of the candidate areas of the photovoltaic panel, and extract the photovoltaic panel area;

[0041] S3: Connect the infrared image and visible light image after extracting the photovoltaic panel area, and perform forward diffusion training and reverse diffusion training through an improved denoising diffusion probabilistic model to obtain a denoising network model;

[0042] S4: Use the denoising network model to obtain a multi-channel diffusion feature map. Based on the multi-channel diffusion feature map, use the spatial attention mechanism and channel attention mechanism to obtain a spatially weighted feature map and a channel-weighted feature map respectively, and obtain the final fused image according to the spatially weighted feature map and the channel-weighted feature map.

[0043] In step S1, an infrared image and a visible light image of a photovoltaic panel at the same angle obtained by an infrared camera and a visible light camera at the same moment are respectively acquired. The method for preprocessing is as follows: the infrared image and the visible light image of the photovoltaic panel are respectively processed by a Gaussian filter to enhance the edges and contours of the scenes in the images. The preprocessed infrared image and visible light image are input into a trained U-Net image segmentation model, and after output, a binary image corresponding to the photovoltaic panel area is obtained.

[0044] In step S2, the binary image of the photovoltaic panel area is merged with the corresponding infrared image and visible light image, and then they are respectively input into a trained residual network, such as ResNet50, to extract multi-scale features. conv2x in ResNet50 is used to extract shallow features, enhance edge detection, and adapt to the rectangular boundary of the photovoltaic panel. conv5x in ResNet50 is used to extract deep features and capture global context. Finally, a feature image containing the texture features of the photovoltaic panel area is output. The feature map is input into a region proposal network, such as RPN, for forward propagation. RPN outputs candidate boxes and their probabilities of being photovoltaic panels, and the top 1000 candidate boxes with the highest probabilities are retained. The photovoltaic panel candidate regions with probabilities greater than 0.7 are selected by non-maximum suppression method. The features of all photovoltaic panel candidate regions are aligned to a fixed size through ROIAlign and then input into a mask head network. In this embodiment, the mask head network selects a fully convolutional neural network, including convolutional layers, batch normalization layers, activation functions, upsampling, and sigmoid activation functions. The mask head network outputs the corresponding masks for the photovoltaic panel candidate regions, which are then combined with the corresponding infrared images and visible light images to extract the photovoltaic panel areas.

[0045] In step S3, as Figure 3 shown, the 1-channel infrared image and the 3-channel visible light image are connected to form a 4-channel image, denoted as I. To adapt to the characteristic that the high pixel height of the photovoltaic panel image may lead to too long training time, this embodiment improves the original DDPM algorithm (Denoising Diffusion Probability Model), which is called the Accelerated Denoising Diffusion Probability Model. In the original DDPM algorithm, the forward process has T steps and the backward process also has T steps, but the process of obtaining the final result actually does not depend on the forward process. Therefore, to accelerate sampling, in this embodiment, the T steps of sampling in the original DDPM model are changed to select a subset S = {x 1t , x 2t ,..., x st} from 1 to T, and Gaussian noise is gradually added until it approaches pure noise. The specific formula is as follows:

[0046]

[0047] Where and respectively represent adding Gaussian noise at the x is -th and the x is -1-th time, Z represents the standard normal distribution, represents the variance of the Gaussian noise added in controlling x is , and γ represents the Gaussian noise.

[0048] Reverse diffusion training is a process of denoising the noise added in the forward diffusion training process. The specific formula is as follows:

[0049]

[0050]

[0051] Among them, ε θ represents the noise prediction network, σ represents the variance, represents the mean of.

[0052] In this embodiment, in order to focus on the defective areas in the infrared image and learn more features of the infrared image in the defective areas, and learn more visible light image features in the non-defective areas, the present invention generates an attention map by using the cross-attention mechanism for the photovoltaic panel infrared image I ir and the visible light image I vis as a constraint condition in the reverse diffusion process. The formula involved in this step is as follows:

[0053] Q = I vis , K = I ir , V = I ir ;

[0054]

[0055] Among them, Q, K, V represent Query, Key, Value in the attention mechanism; CA iv represents the cross-attention map.

[0056] In step S4, a multi-channel diffusion feature map is obtained by using the denoising network model, and a multi-channel fusion module is used to fuse the multi-channel diffusion feature maps of multiple stages in the denoising network. Specifically: in this embodiment, the multi-channel diffusion feature maps of multiple stages in the denoising network are fused through the spatial attention mechanism and the channel attention mechanism.

[0057] The spatial attention mechanism is: first, two parallel convolutional operations are performed on the input multi-channel diffusion feature map and the two feature maps are added together; then, through a sigmoid activation function, the spatial attention weight is obtained; finally, the input multi-channel diffusion feature map F and the spatial attention weight M sMultiply to obtain the spatially weighted feature map. The formula is as follows:

[0058] M s = σ(conv1(F) + conv2(F));

[0059] F' = F × M s ;

[0060] where F represents the multi-channel diffusion feature map; M s represents the spatial attention weight, and conv represents the convolution operation.

[0061] The channel attention mechanism is as follows: First, perform global average pooling and global max pooling on the input multi-channel diffusion feature map to obtain two one-dimensional vectors respectively; then concatenate these two vectors, pass through a fully connected layer, then through the ReLU activation function, and then through another fully connected layer (or two parallel fully connected layers), and finally through the sigmoid activation function to obtain the channel attention weight M c ; Multiply the input multi-channel diffusion feature map F by the channel attention weight to obtain the channel-weighted feature map. The formula is as follows:

[0062]

[0063] M c = σ(FC(ReLU(FC(concat(F avg , F max ))));

[0064] F' = F × M c ;

[0065] where F represents the multi-channel diffusion feature map; M c represents the channel attention weight; H represents the height of the feature map; W represents the width of the feature map; σ represents the sigmoid activation function.

[0066] To enhance the important regions jointly attended to by the two attentions, the channel-weighted feature map and the spatially weighted feature map are fused by multiplying the attention weights. The fused feature map is upsampled to restore it to the original image size and then passed through the output layer to output the final fused image.

[0067] In this embodiment, by using the multi-channel gradient loss function, a 3-channel image can be directly fused and generated, eliminating the step of color space transformation. During the process of fusing images, a 3-channel image is directly generated. However, the existing gradient loss is designed for single-channel fused images. To directly generate a 3-channel fused image while keeping the gradient unchanged, this embodiment applies a brand-new 3-channel gradient loss function. The loss function formula involved in this step is as follows:

[0068]

[0069] Among them, L MCG represents the multi-channel gradient loss, and L MCI represents the multi-channel intensity loss, represents the channels (red, green, blue) of the fused image, represents the gradient operator, represents the channels of the input visible light image. When the multi-channel intensity loss and the multi-channel gradient loss are calculated, they are added together to obtain the final loss, making full use of the diffusion characteristics and directly generating a 3-channel fused image.

[0070] Example Two

[0071] In this example, a photovoltaic module defect image fusion system based on a diffusion model is provided, including the following modules:

[0072] An image acquisition module, configured to: acquire the infrared image and visible light image of the photovoltaic panel and perform preprocessing; obtain the binary images of the corresponding photovoltaic panel areas according to the preprocessed infrared image and visible light image;

[0073] A photovoltaic panel area extraction module, configured to: merge the binary image of the photovoltaic panel area with the corresponding infrared image and visible light image, and input them into the trained first convolutional neural network model to output a feature image; screen out the candidate photovoltaic panel areas according to the feature image, generate the corresponding masks of the candidate photovoltaic panel areas, and extract the photovoltaic panel areas;

[0074] A diffusion training module, configured to: register the infrared image and visible light image after extracting the photovoltaic panel area, and perform forward diffusion training and reverse diffusion training through an improved denoising diffusion probability model to obtain a denoising network model;

[0075] An image fusion module, configured to: obtain a multi-channel diffusion feature map by using the denoising network model, and based on the multi-channel diffusion feature map, respectively obtain a spatially weighted feature map and a channel-weighted feature map by using a spatial attention mechanism and a channel attention mechanism, and obtain the final fused image according to the spatially weighted feature map and the channel-weighted feature map.

[0076] Example Three

[0077] The purpose of this example is to provide a computer-readable storage medium for storing a computer program to complete the method described in Example One.

[0078] The method in the first embodiment can be directly embodied as being executed by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can be located in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0079] Embodiment Four

[0080] The purpose of this embodiment is to provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, which can be used to complete the method described in the first embodiment. For the sake of brevity of description, it will not be elaborated here.

[0081] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0082] The memory can include read-only memory and random access memory, and provide instructions and data to the processor. A part of the memory can also include non-volatile random access memory. For example, the memory can also store information about the device type.

[0083] For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for fusing defect images of photovoltaic modules based on a diffusion model, characterized in that, Including: Obtain the infrared image and visible light image of the photovoltaic panel, and perform preprocessing; According to the preprocessed infrared image and visible light image, obtain the binary images of the corresponding photovoltaic panel regions respectively; Merge the binary images of the photovoltaic panel regions with the corresponding infrared images and visible light images, and input them into the trained residual network to output a feature image; Screen out the candidate regions of the photovoltaic panel according to the feature image, generate the corresponding masks of the candidate regions of the photovoltaic panel, and extract the photovoltaic panel regions; Connect the infrared image and visible light image after extracting the photovoltaic panel regions, and perform forward diffusion training and reverse diffusion training through an improved denoising diffusion probabilistic model to obtain a denoising network model; Use the denoising network model to obtain a multi-channel diffusion feature map. Based on the multi-channel diffusion feature map, use the spatial attention mechanism and channel attention mechanism to obtain the spatial weighted feature map and channel weighted feature map respectively, and obtain the final fusion image according to the spatial weighted feature map and channel weighted feature map.

2. The method for fusing defective images of photovoltaic modules based on a diffusion model according to claim 1, characterized in that, The feature image output by the trained residual network contains the texture features of the photovoltaic panel region; Perform forward propagation on the feature image, and screen out the candidate regions of the photovoltaic panel with a probability greater than 0.7 through the non-maximum suppression method.

3. The method for fusing defect images of photovoltaic modules based on a diffusion model according to claim 1, wherein Align the features of all the screened candidate regions of the photovoltaic panel to a fixed size, input them into the trained mask head network, output the corresponding masks of the candidate regions of the photovoltaic panel, and then combine them with the corresponding infrared images and visible light images to extract the photovoltaic panel regions.

4. The method for fusing defect images of photovoltaic modules based on a diffusion model according to claim 1, characterized in that The forward diffusion is trained for T steps, and a subset S{x 1t ,x 2t ,...,x st} is selected from 1 to T, and Gaussian noise is gradually added until it approaches pure noise. The specific formula is as follows: Among them, and respectively represent adding Gaussian noise for the x is -th and the x is - 1-th times, Z represents the standard normal distribution, is expressed as the variance of the Gaussian noise added in controlling x is , and γ represents the Gaussian noise.

5. The method for fusing defect images of photovoltaic modules based on a diffusion model according to claim 1, wherein, The reverse diffusion training is a process of denoising the noise added in the forward diffusion training process. The specific formula is as follows: Among them, ε θ represents the noise prediction network, σ represents the variance of, represents the mean of.

6. The method for fusing defect images of photovoltaic modules based on a diffusion model according to claim 1, wherein Generate an attention map by using the cross-attention mechanism for the infrared image and visible light image of the photovoltaic panel, and use it as a constraint condition in the reverse diffusion training process. The specific formula is as follows: Q = I vis , K = I ir , V = I ir ; 7. The method for fusing defect images of photovoltaic modules based on a diffusion model according to claim 1, wherein The specific method for obtaining the spatially weighted feature map using the spatial attention mechanism is as follows: First, two parallel convolutional operations are performed on the input multi-channel diffusion feature map and the two feature maps are added together; then, a sigmoid activation function is used to obtain the spatial attention weights; finally, the input multi-channel diffusion feature map is multiplied by the spatial attention weight M s to obtain the spatially weighted feature map; The specific method of obtaining the channel weighted feature map by using the channel attention mechanism is as follows: First, perform global average pooling and global maximum pooling on the input multi-channel diffusion feature map to obtain two one-dimensional vectors respectively; then concatenate these two vectors, pass through a fully connected layer, then pass through the ReLU activation function, then pass through another fully connected layer, and finally pass through the sigmoid activation function to obtain the channel attention weight; multiply the input multi-channel diffusion feature map by the channel attention weight to obtain the channel weighted feature map.

8. A photovoltaic module defect image fusion system based on a diffusion model, characterized in that Including the following modules: Image acquisition module, configured to: Obtain the infrared image and visible light image of the photovoltaic panel, and perform preprocessing; According to the preprocessed infrared image and visible light image, obtain the binary images of the corresponding photovoltaic panel regions respectively; Photovoltaic panel region extraction module, configured to: Merge the binary images of the photovoltaic panel regions with the corresponding infrared images and visible light images, and input them into the trained residual network to output a feature image; Screen out the candidate regions of the photovoltaic panel according to the feature image, generate the corresponding masks of the candidate regions of the photovoltaic panel, and extract the photovoltaic panel regions; Diffusion training module, configured to: Connect the infrared image and visible light image after extracting the photovoltaic panel regions, and perform forward diffusion training and reverse diffusion training through an improved denoising diffusion probabilistic model to obtain a denoising network model; The image fusion module is configured to: obtain a multi-channel diffusion feature map by using the denoising network model, and based on the multi-channel diffusion feature map, respectively obtain a spatially weighted feature map and a channel weighted feature map by adopting a spatial attention mechanism and a channel attention mechanism, and obtain a final fused image according to the spatially weighted feature map and the channel weighted feature map.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the diffusion model-based photovoltaic module defect image fusion method according to any one of claims 1-7.

10. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the diffusion model-based photovoltaic module defect image fusion method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Photovoltaic module defect intelligent detection method fusing visible light and infrared images

    CN116091472A

  • Photovoltaic hot spot defect detection method and device and storage medium

    CN118735890A

  • Amodal instance segmentation using diffusion models

    US20240169541A1