Double-stage photovoltaic panel defect detection method and system based on multispectral spatial model
By adopting a two-stage detection method of multispectral spatial model in photovoltaic panel defect detection, and using variable polymerization convolution and timing dynamic image registration modules, the problems of limited detection accuracy and high cost in the prior art are solved, and efficient and robust photovoltaic panel defect detection are achieved.
Patent Information
- Application Number
- CN202510111131.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-23
AI Technical Summary
When processing high-resolution aerial images, existing photovoltaic panel defect detection technology has problems such as limited detection accuracy, high cost and poor robustness, especially in multispectral images, spectral information cannot be fully mined.
The two-stage photovoltaic panel defect detection method based on multispectral spatial model is adopted. The photovoltaic panel area is first extracted through a lightweight photovoltaic panel positioning detection network based on variable polymerization convolution, and then the multispectral image calibration is performed through the timing dynamic image registration module. Finally, the image spectrum and spatial dimension are linearly modeled using the multispectral spatial model to enhance the characterization ability of defect characteristics.
High-precision detection of photovoltaic panel defects is achieved, irrelevant background interference is reduced, detection accuracy and robustness are improved, and detection costs are reduced.
Smart Images

Figure CN120147225A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of photovoltaic panel defect detection, and particularly to a two-stage photovoltaic panel defect detection method and system based on a multi-spectral state space model. Background Art
[0002] As a clean and renewable energy solution, solar photovoltaic technology has received extensive attention and application. With the large-scale deployment of photovoltaic power generation systems, the daily maintenance problems of photovoltaic panels have become increasingly prominent. During actual operation, photovoltaic panels often expose various potential faults and defects such as hot spots and debris, which may lead to energy loss, system efficiency decline, and even system failures in extreme cases. Therefore, developing accurate and efficient photovoltaic panel defect detection technology is crucial for ensuring the reliability and stability of photovoltaic systems.
[0003] Currently, photovoltaic panel defect detection technologies are mainly divided into two categories. The first category is the detection method based on electrical characteristics, which mainly collects electrical operation data through various sensors and uses mathematical models or machine learning to analyze whether there are defects. However, a large number of sensors result in too high detection costs for photovoltaic panels. The second category is the detection method based on aerial images, which mainly relies on machine learning or deep learning technologies and has gradually become the mainstream solution for photovoltaic panel defect detection. Algorithms based on traditional machine learning, such as support vector machines and random forest algorithms, usually rely on manually designed extractors and need to manually construct different and complex recognition relationships in different scenarios, with poor generalization ability and robustness. With the rapid development of deep learning, detection methods based on deep learning, such as models like Faster R-CNN, YOLO, and DETR, have significantly improved in feature representation ability. However, these methods usually directly identify the collected images. However, aerial images have large scales, high resolutions, and a large number of irrelevant background pixels, and are limited by the model's structure size when processing high-resolution images, which seriously affects the model's detection ability. In addition, defects will cause uneven heating of the photovoltaic panel, resulting in significant differences in the spectral characteristics of the photovoltaic panel in bands such as infrared, near-infrared, etc. compared with normal photovoltaic panels. Therefore, using multi-spectral images including infrared, near-infrared, far-infrared, and RGB bands for photovoltaic panel defect detection has great advantages. However, existing algorithms have not fully exploited and utilized the spectral information in multi-spectral images under resource-constrained conditions, resulting in limited accuracy of photovoltaic panel defect detection. Summary of the Invention
[0004] In order to solve the above technical problems existing in the prior art, the present invention provides a two-stage photovoltaic panel defect detection method based on a multi-spectral state space model, and the technical solution is as follows:
[0005] On the one hand, a two-stage photovoltaic panel defect detection method based on a multi-spectrum state space model is provided, and the method includes:
[0006] S1. Obtain a high-resolution aerial RGB photovoltaic panel image to be detected and a multi-spectrum state aerial photovoltaic panel image;
[0007] S2. Input the RGB aerial photovoltaic panel image into a lightweight photovoltaic panel positioning detection network based on variable aggregation convolution in the first stage. The photovoltaic panel positioning detection network locates and detects the position of the photovoltaic panel in the RGB aerial photovoltaic panel image and outputs the position of the photovoltaic panel;
[0008] S3. Use the RGB aerial photovoltaic panel image as a reference image and the multi-spectrum state aerial photovoltaic panel image as a to-be-calibrated image, and input them into a temporal dynamic image registration module. The temporal dynamic image registration module uses a deep learning registration method based on a skip network to accurately register a single-frame RGB image and a multi-spectrum state image, and obtains a multi-spectrum state calibrated image through temporal dynamic accumulation operation of multiple frames of images;
[0009] S4. Crop the multi-spectrum state calibrated image according to the position of the photovoltaic panel output by the photovoltaic panel positioning detection network to obtain a multi-spectrum state photovoltaic panel defect image of the corresponding area;
[0010] S5. Input the multi-spectrum state photovoltaic panel defect image into a photovoltaic panel defect detection network in the second stage. The photovoltaic panel defect detection network uses a multi-spectrum state space model to linearly model the image spectrum and spatial dimensions, effectively adapts to the characteristics of spectral order changes, enhances the characterization ability of defect features, and outputs the defect category and defect detection frame of the photovoltaic panel to be detected, completing the photovoltaic panel defect detection.
[0011] Optionally, the photovoltaic panel positioning detection network includes a backbone network, a neck network, and a detection output end;
[0012] The backbone network is divided into four stages, which are connected in sequence. In each of the first three stages, the network is composed of a convolution module and a cross-stage feature extraction module connected in series. In the last stage, an SPPF module is added on this basis; the convolution module includes a convolutional layer, a batch normalization layer, and a Leaky ReLU activation function. Among them, the Leaky ReLU activation function is an improved version of ReLU, and the formula is as follows: In the formula, α is a small constant. When the input is negative, the output remains negative but is not exactly zero, and it provides a non-zero gradient α for negative inputs, which helps the gradient propagate more smoothly throughout the network and effectively solves the problem of dead neurons that may occur during the training of the standard ReLU. The cross-stage feature extraction module splits the channels of the input feature map into two parts on average and inputs them into two branches respectively. One branch passes through variable aggregation convolution, and the other branch is processed through variable aggregation convolution and a bottleneck module. After the outputs of the two branches are concatenated, the final result is obtained through variable aggregation convolution. The bottleneck module consists of two convolutional modules and uses residual connection to add the input and output. The SPPF module includes three max pooling layers in series. After the output of each max pooling layer is concatenated with the original input, the final output is obtained through a convolutional module. The last three stages of the backbone network output feature maps of small, medium, and large scales respectively, and there is a two-fold downsampling relationship between the sizes. These feature maps will be passed as inputs to the neck network.
[0013] The neck network inputs the small, medium, and large scale feature maps from the backbone network and processes them through cross-stage feature extraction modules, convolution, and upsampling operations: First, the small scale feature map is upsampled to the same spatial resolution as the medium scale feature map and concatenated with the medium scale feature map, then input into the cross-stage feature extraction module L1. The output of L1 is upsampled to the same resolution as the large scale feature map, concatenated with the large scale feature map, and then input into the cross-stage feature extraction module L2 to generate a large scale fusion feature. The output of L2 is downsampled to the medium scale resolution through convolution and concatenated with the output of L1, then input into the cross-stage feature extraction module L3 to obtain a medium scale fusion feature. The output of L3 is restored to the small scale resolution through convolution, concatenated with the small scale feature map, and then input into the cross-stage feature extraction module L4 to obtain a small scale fusion feature map. Through the top-down feature fusion and bottom-up feature enhancement paths, the high-level semantic information and low-level detail features are effectively fused to achieve more accurate object detection in different scales and complex scenes.
[0014] The detection output end concatenates the three scale fusion features from the neck network, upsamples them to the original image size, and outputs the position of the photovoltaic panel through convolution operations.
[0015] Optionally, in the variable aggregation convolution, k parallel convolutional kernels are embedded in one convolutional layer, and through an attention mechanism composed of an average pooling layer, a fully connected layer, a Leaky ReLU activation function, a fully connected layer, and softmax stacked in sequence, the weights of all convolutional kernels are dynamically adjusted to identify photovoltaic panels of different scales and distinguish irrelevant backgrounds on high-resolution images. Each convolutional kernel performs a convolution operation on the input image tensor, multiplies it by its respective weight, and finally aggregates and adds them to obtain the final convolution result, which is then output through batch normalization and the Leaky ReLU activation function.
[0016] Optionally, in the temporal dynamic image registration module, during the single-frame image registration process, the reference RGB image I F and the multi-spectral image I to be calibrated M are respectively input into the skip network to extract their respective feature maps. After splicing the two obtained feature maps, they are downsampled and flattened through a convolutional layer and then input into a fully connected layer to obtain the geometric transformation parameters P{t x ,t y ,θ,s x ,s y ,β x ,β y}, where t x and t y are translation parameters, θ is the rotation angle, s x and s y are scaling factors, and β x and β y are shear factors. The multi-spectral image to be calibrated is geometrically transformed and adjusted using these parameters to achieve registration, obtaining the calibrated image IR;
[0017] The skip network takes an image as input, first extracts preliminary features through convolution and max-pooling layers; then further extracts deep features through three double-branch residual structures and cross-stage feature extraction modules. The double-branch residual structure includes a single-layer convolutional branch and a three-layer convolutional branch. After fusion, it is processed using the Leaky ReLU activation function, and then the deep features of the image are further extracted through two cross-stage feature extraction modules, and finally the feature map is output.
[0018] Optionally, in the temporal dynamic image registration module, during the temporal dynamic accumulation operation, n RGB images {I F1 ,I F2 ,I F3 ...I Fn} in the image video segment and the corresponding n multi-spectral images {I M1 ,I M2 ,I M3 ...I Mn}Perform image registration to obtain the registered image {IR 1 , IR 2 , IR 3 ... IR n}, and then IR 1 is cumulated frame by frame with IR 2 to obtain IR 12 , and IR 12 is registered with I F2 to obtain IRR 12 ; IRR 12 is cumulated frame by frame with IR 3 to obtain IRR 123 , and then IRR 123 is registered with I F3 to obtain IRRR 123 ; By repeating the above steps, the final multi-spectral calibration image IRR..R 12..n is obtained, where the frame accumulation operation refers to adding and averaging each pixel value corresponding to two images.
[0019] Optionally, the photovoltaic panel defect detection network consists of a photovoltaic panel defect detection backbone network and a multi-scale fusion detection output end;
[0020] The photovoltaic panel defect detection backbone network first extracts photovoltaic panel defect features through a 3D convolution module, which is composed of 3D convolution, batch normalization, and Leaky ReLU activation function stacked in sequence, and the input and output are connected through residual connections; Next, the feature map is sent to the subsequent stages, which are divided into four stages, and each stage is connected in sequence. The first stage consists of a state space feature extraction module; The second and third stages are composed of a downsampling module and a state space feature extraction module in series; The fourth stage adds an SPPF module on the basis of the first two stages to enhance the extraction of key features and multi-scale feature fusion. Through the above four stages, feature maps of different scales are generated, including large, medium, and small scales, and the feature map sizes decrease according to a factor of two downsampling relationship. The three generated scale features will be used as the input of the multi-scale fusion detection output end;
[0021] The multi-scale fusion detection output end receives small, medium, and large-scale feature maps from the photovoltaic panel defect detection backbone network. First, multi-scale fusion processing is performed through a state space feature extraction module, convolution, and upsampling operations. Among them, the small-scale feature map is upsampled to the same spatial resolution as the medium-scale feature map and concatenated with the medium-scale feature map, and then input into the state space feature extraction module F1. The output of F1 is upsampled to the same resolution as the large-scale feature map, concatenated with the large-scale feature map, and then input into the state space feature extraction module F2 to generate a large-scale fusion feature. The output of F2 is downsampled to the medium-scale resolution through convolution and concatenated with the output of F1, and then input into the state space feature extraction module F3 to obtain a medium-scale fusion feature. The output of F3 is restored to the small-scale resolution through convolution, concatenated with the small-scale feature map, and then input into the state space feature extraction module F4 to obtain a small-scale fusion feature map. Then, the fusion feature maps of these three scales are concatenated and upsampled to the original image size. Then, 1×1 convolution is used to adjust the dimension of the feature map, and then processed through two 3×3 convolutions and one 1×1 convolution. Finally, the defect target category and location are output to complete the photovoltaic panel defect detection.
[0022] Optionally, for the state space feature extraction module, the input feature map is first preliminarily extracted through convolution, batch normalization, and the LeakyReLU activation function. The preliminarily processed feature is input into the local space module, which is sequentially stacked by depthwise separable convolution, batch normalization, convolution, the LeakyReLU activation function, and convolution, and the output of the module is added to the input through a residual structure.
[0023] Then, the output of the local space module is input into the multi-spectral state space model after layer normalization processing to extract and combine the information of the spectral dimension and the spatial dimension to enhance the expression of the photovoltaic panel defect features.
[0024] Finally, the obtained feature map is input into a linear layer for further processing, and a residual connection is adopted. The output of the LeakyReLU activation function in the preliminary processing is added and connected to the outputs of the multi-spectral state space model and the linear layer respectively to obtain the output feature map.
[0025] Optionally, for the multi-spectral state space model, the input first undergoes image serialization processing, cutting the feature map into 9 spatial slices and converting it into an image sequence.
[0026] Then these image sequences are respectively input into the spectral dimension branch and the spatial dimension branch for processing. In the spatial dimension branch, the multi-spectral state space model linearly models the 9 slices of the image sequence along the spatial dimension. Considering that the image sequence has context relationships and spatial dependencies in four directions, the image sequence is first scanned forward and backward row by row and column by column, and block matrix rearrangement is performed. The four groups of rearranged image sequences are input into the SSM module. After the four output feature sequences are processed by summation and averaging, they are decoded back into an image tensor and passed through a convolution layer, finally obtaining the spatial dimension features;
[0027] In the spectral dimension branch, the multi-spectral state space model linearly models the c spectral channels of the image sequence. First, the spectral dimension sequence is adjusted, and square matrix rearrangement is performed through forward and backward scans. The two groups of rearranged image sequences are input into the SSM module. The two output image feature sequences are also processed by summation and averaging, then adjusted to the original dimension, decoded back into an image tensor, and after passing through a convolution layer, the final spectral dimension features are obtained;
[0028] Finally, the spatial dimension features and the spectral dimension features are fused to learn the joint features of the learning target in space and spectrum, generating the output of the final feature map.
[0029] Optionally, the SSM module is a discrete state space model obtained by discretizing a continuous state space model. By discretizing the state transition process in continuous time, the original continuous state transition matrix and projection matrix are converted into forms suitable for discrete time, realizing effective modeling of time series data.
[0030] On the other hand, a two-stage photovoltaic panel defect detection system based on a multi-spectral state space model is provided. The system includes:
[0031] An acquisition module for acquiring high-resolution RGB aerial photovoltaic panel images and multi-spectral aerial photovoltaic panel images to be detected;
[0032] A positioning and detection module for inputting the RGB aerial photovoltaic panel image into a lightweight photovoltaic panel positioning and detection network based on variable aggregation convolution in the first stage. The photovoltaic panel positioning and detection network performs positioning and detection on the position of the photovoltaic panel in the RGB aerial photovoltaic panel image and outputs the position of the photovoltaic panel;
[0033] A registration module for using the RGB aerial photovoltaic panel image as a reference image and the multi-spectral aerial photovoltaic panel image as a to-be-calibrated image, inputting them into a temporal dynamic image registration module. The temporal dynamic image registration module uses a deep learning registration method based on a skip network to perform precise registration of a single-frame RGB image and a multi-spectral image, and obtains a multi-spectral calibrated image through temporal dynamic accumulation operations of multiple frames of images;
[0034] A cutting module, configured to cut the multi-spectral calibration image according to the position of the photovoltaic panel output by the photovoltaic panel positioning detection network, so as to obtain a multi-spectral photovoltaic panel defect image of the corresponding area;
[0035] A defect detection module, configured to input the multi-spectral photovoltaic panel defect image into a photovoltaic panel defect detection network in a second stage. The photovoltaic panel defect detection network uses a multi-spectral space model to perform linear modeling on the image spectrum and spatial dimension, effectively adapts to the characteristics of spectral order changes, enhances the characterization ability of defect features, and outputs the defect category and defect detection frame of the photovoltaic panel to be detected, thereby completing the photovoltaic panel defect detection.
[0036] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned two-stage photovoltaic panel defect detection method based on the multi-spectral space model.
[0037] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned two-stage photovoltaic panel defect detection method based on the multi-spectral space model.
[0038] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0039] 1. The present invention proposes a two-stage photovoltaic panel defect detection algorithm. Different from common deep learning algorithms that directly perform single-stage defect detection on high-resolution aerial images, first, through a lightweight photovoltaic panel positioning detection network based on variable aggregation convolution in the first stage, the photovoltaic panel area in the RGB image is extracted, and the corresponding multi-spectral calibration image is cut according to this area, thereby minimizing the interference of irrelevant backgrounds to the greatest extent.
[0040] 2. The temporal dynamic image registration module of the present invention, compared with traditional image registration methods, in a complex photovoltaic panel aerial photography scenario, this module realizes the accurate registration of a single-frame RGB image and a multi-spectral image through a registration network based on a skip network, and effectively improves the signal-to-noise ratio to improve the calibration image quality through the temporal dynamic accumulation operation of multiple frames of images, can achieve high-precision alignment of RGB and multi-spectral high-resolution images in an aerial photography scenario, and has strong robustness.
[0041] 3. The photovoltaic panel defect detection network in the second stage of the present invention, combined with the multi-spectral space model, can perform linear modeling on the spatial and spectral dimensions, enhances the characterization ability of photovoltaic panel defect features, improves the detection accuracy. At the same time, the multi-spectral space model is designed based on the SSM architecture and has a linear computational complexity, ensuring the feasibility of this method in actual deployment. Brief Description of the Drawings
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 It is a flowchart of a two-stage photovoltaic panel defect detection method based on a multi-spectrum state space model provided by an embodiment of the present invention;
[0044] Figure 2 It is an overall block diagram of a two-stage photovoltaic panel defect detection method based on a multi-spectrum state space model provided by an embodiment of the present invention;
[0045] Figure 3 It is a block diagram of a photovoltaic panel positioning detection network structure provided by an embodiment of the present invention;
[0046] Figure 4 It is a block diagram of a variable aggregation convolution structure provided by an embodiment of the present invention;
[0047] Figure 5 It is a block diagram of a timing dynamic image registration module structure provided by an embodiment of the present invention;
[0048] Figure 6 It is a block diagram of a photovoltaic panel defect detection network structure provided by an embodiment of the present invention;
[0049] Figure 7 It is a block diagram of a state space feature extraction module structure provided by an embodiment of the present invention;
[0050] Figure 8 It is a block diagram of a multi-spectrum state space model structure provided by an embodiment of the present invention;
[0051] Figure 9 It is a block diagram of a two-stage photovoltaic panel defect detection system based on a multi-spectrum state space model provided by an embodiment of the present invention;
[0052] Figure 10 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0053] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the drawings and specific embodiments.
[0054] The embodiments of the present invention provide a two-stage photovoltaic panel defect detection method based on a multi-spectrum state space model. This method can be implemented by an electronic device, which can be a terminal or a server.Figure 1 The flowchart of the method is shown as follows, Figure 2 The overall block diagram of the method is shown as follows. The processing flow may include the following steps:
[0055] S1. Obtain a high-resolution RGB aerial photovoltaic panel image to be detected and a multispectral aerial photovoltaic panel image;
[0056] In the embodiment of the present invention, a drone equipped with a high-resolution RGB camera and a multispectral camera is used to take an aerial photo of the photovoltaic power station area to obtain the RGB image and multispectral image data of the photovoltaic panel to be detected. The RGB image is used for the first-stage photovoltaic panel position detection. The channel data dimension is relatively low, only including three channels (red, green, and blue), and the calculation and storage requirements are less, which can ensure the lightweight of the system; the multispectral image covers bands such as infrared, near-infrared, and far-infrared. These bands are more sensitive to the uneven heat sources of the defective parts of the photovoltaic panel, so they are used to identify potential defects of the photovoltaic panel in the second stage.
[0057] S2. Input the RGB aerial photovoltaic panel image into the lightweight photovoltaic panel positioning detection network based on variable aggregation convolution in the first stage. The photovoltaic panel positioning detection network locates and detects the position of the photovoltaic panel in the RGB aerial photovoltaic panel image and outputs the position of the photovoltaic panel;
[0058] Optionally, as Figure 3 shown, the photovoltaic panel positioning detection network includes a backbone network, a neck network, and a detection output end;
[0059] The backbone network is divided into four stages, which are connected in sequence. In each of the first three stages, the network is composed of a convolution module and a cross-stage feature extraction module connected in series. In the last stage, an SPPF module is added on this basis; the convolution module includes a convolution layer, a batch normalization layer, and a Leaky ReLU activation function. Among them, the Leaky ReLU activation function is an improved version of ReLU (the formula of the ReLU activation function is as follows: During the training process of the ReLU activation function, the output for negative inputs is 0, resulting in some neurons being unable to be activated all the time, having no impact on the final output of the network. In addition, the derivative of ReLU for negative inputs is always 0. During backpropagation, the gradient in the negative input area is 0, so that the weights of these neurons cannot be updated. This phenomenon is called the "dead neuron" phenomenon, which directly affects the learning ability of the model), and the formula is as follows: In the formula, α is a small constant (usually taken as 0.01). When the input is negative, the output still remains negative, but it will not be exactly zero, and it provides a non-zero gradient α for the negative input, which helps the gradient to propagate more smoothly throughout the network and effectively solves the problem of dead neurons that may occur in the standard ReLU during the training process. The cross-stage feature extraction module splits the channels of the input feature map into two parts on average and inputs them into two branches respectively. One branch goes through variable aggregation convolution, and the other branch is processed through variable aggregation convolution and a bottleneck module. After the outputs of the two branches are concatenated, the final result is obtained through variable aggregation convolution. The bottleneck module consists of two convolution modules and uses residual connection to add the input and output. The SPPF module includes three max pooling layers in series. After the output of each max pooling layer is concatenated with the original input, the final output is obtained through a convolution module. The last three stages of the backbone network output feature maps of small, medium, and large scales respectively, and there is a two-fold downsampling relationship between the sizes. These feature maps will be passed as inputs to the neck network.
[0060] The neck network inputs the small, medium, and large scale feature maps from the backbone network and processes them through cross-stage feature extraction modules, convolution, and upsampling operations: First, the small scale feature map is upsampled to the same spatial resolution as the medium scale feature map and concatenated with the medium scale feature map, and then input into the cross-stage feature extraction module L1. The output of L1 is upsampled to the same resolution as the large scale feature map, concatenated with the large scale feature map, and then input into the cross-stage feature extraction module L2 to generate a large scale fusion feature. The output of L2 is downsampled to the medium scale resolution through convolution and concatenated with the output of L1, and then input into the cross-stage feature extraction module L3 to obtain a medium scale fusion feature. The output of L3 is restored to the small scale resolution through convolution, concatenated with the small scale feature map, and then input into the cross-stage feature extraction module L4 to obtain a small scale fusion feature map. Through top-down feature fusion (transferring semantic information from high-level to low-level to enhance the expression ability of low-level features) and bottom-up feature enhancement path (starting from low-level features and gradually adding semantic information to enhance the network's comprehensive understanding of details and semantics), the high-level semantic information and low-level detail features are effectively fused to achieve more accurate object detection in different scales and complex scenes.
[0061] The detection output end concatenates the three-scale fusion features from the neck network, upscales them to the original image size through upsampling, and outputs the position of the photovoltaic panel through convolution operations.
[0062] Optionally, as Figure 4As shown, the variable aggregation convolution embeds k parallel convolutional kernels in a convolutional layer, and dynamically adjusts the weights of all convolutional kernels through an attention mechanism composed of an average pooling layer, a fully connected layer, a Leaky ReLU activation function, a fully connected layer, and softmax in sequence, so as to identify photovoltaic panels of different scales and distinguish irrelevant backgrounds on high-resolution images. Each convolutional kernel performs a convolution operation on the input image tensor, multiplies it by its respective weight, and finally aggregates and adds them to obtain the final convolution result, and then outputs the result through batch normalization and the Leaky ReLU activation function.
[0063] S3. Use the RGB aerial photovoltaic panel image as the reference image and the multi-spectral aerial photovoltaic panel image as the image to be calibrated, and input them into the temporal dynamic image registration module. The temporal dynamic image registration module uses a deep learning registration method based on a skip network to accurately register a single-frame RGB image and a multi-spectral image, and obtains a multi-spectral calibration image through the temporal dynamic accumulation operation of multiple frames of images;
[0064] Optionally, as Figure 5 shown, in the process of single-frame image registration, the temporal dynamic image registration module inputs the reference RGB image I F and the image I M to be calibrated, a multi-spectral image, into the skip network respectively to extract their respective feature maps. After splicing the two obtained feature maps, downsample and flatten them through a convolutional layer, and input them into a fully connected layer to obtain geometric transformation parameters P{t x ,t y ,θ,s x ,s y ,β x ,β y}, where t x and t y are translation parameters, θ is the rotation angle, s x and s y are scaling factors, β x and β y are shear factors. Geometric image transformation adjustment (including translation, rotation, scaling, and shearing) is performed on the image to be calibrated with these parameters to achieve registration and obtain a calibrated image IR;
[0065] The skip network takes an image as input. First, it extracts preliminary features through convolutional and max pooling layers; then, it further extracts deep features through three double-branch residual structures and cross-stage feature extraction modules. The double-branch residual structure includes a single-layer convolutional branch and a three-layer convolutional branch. After fusion, it is processed by the Leaky ReLU activation function, and then further extracts the deep features of the image through two layers of cross-stage feature extraction modules, and finally outputs the feature map.
[0066] Optionally, in the temporal dynamic image registration module, during the temporal dynamic accumulation operation, n RGB images {I F1 , I F2 , I F3 ... I Fn} in the image video segment are registered with the corresponding n multi-spectral images {I M1 , I M2 , I M3 ... I Mn} to obtain registered images {IR 1 , IR 2 , IR 3 ... IR n}. Then, IR 1 and IR 2 are frame-accumulated to obtain IR 12 , and IR 12 and I F2 are image-registered to obtain IRR 12 ; IRR 12 and IR 3 are frame-accumulated to obtain IRR 123 , and then IRR 123 and I F3 are image-registered to obtain IRRR 123 ; by repeating the above steps, the final multi-spectral calibration image IRR..R 12..n is obtained, where the frame-accumulation operation refers to taking the average after adding the corresponding pixel values of two images.
[0067] In the embodiment of the present invention, through the temporal dynamic image registration module, precise alignment between high-resolution aerial RGB images of photovoltaic panels and corresponding multi-spectral images can be achieved, and it has strong robustness.
[0068] S4. Crop the multi-spectral calibration image according to the position of the photovoltaic panel output by the photovoltaic panel positioning detection network to obtain a multi-spectral photovoltaic panel defect image of the corresponding area;
[0069] S5. Input the multi-spectral photovoltaic panel defect image into the photovoltaic panel defect detection network in the second stage. The photovoltaic panel defect detection network uses a multi-spectral space model to linearly model the image spectrum and spatial dimensions, effectively adapts to the characteristics of spectral order changes, enhances the characterization ability of defect features, and outputs the defect category and defect detection frame of the photovoltaic panel to be detected, completing the photovoltaic panel defect detection.
[0070] Optionally, as Figure 6 shown, the photovoltaic panel defect detection network consists of a photovoltaic panel defect detection backbone network and a multi-scale fusion detection output end;
[0071] The backbone network for photovoltaic panel defect detection first extracts the defect features of the photovoltaic panel through a 3D convolution module, which is composed of 3D convolution, batch normalization, and the Leaky ReLU activation function stacked in sequence, and the input and output are connected through a residual connection. Next, the feature map is fed into the subsequent stages. The subsequent stages are divided into four stages, which are connected in sequence. The first stage consists of a state space feature extraction module. The second and third stages are composed of a downsampling module and a state space feature extraction module connected in series. The fourth stage adds an SPPF module based on the first two stages to enhance the extraction of key features and multi-scale feature fusion. Through the above four stages, feature maps of different scales are generated, including large, medium, and small scales. The sizes of the feature maps decrease according to a two-fold downsampling relationship. The three generated scales of features will be used as the input of the multi-scale fusion detection output end.
[0072] The multi-scale fusion detection output end, with the input being the small, medium, and large scale feature maps from the backbone network for photovoltaic panel defect detection, first performs multi-scale fusion processing through a state space feature extraction module, convolution, and upsampling operations. Among them, the small scale feature map is upsampled to the same spatial resolution as the medium scale feature map and concatenated with the medium scale feature map, and then input into the state space feature extraction module F1. The output of F1 is upsampled to the same resolution as the large scale feature map, concatenated with the large scale feature map, and then input into the state space feature extraction module F2 to generate a large scale fusion feature. The output of F2 is downsampled to the medium scale resolution through convolution and concatenated with the output of F1, and then input into the state space feature extraction module F3 to obtain a medium scale fusion feature. The output of F3 is restored to the small scale resolution through convolution and concatenated with the small scale feature map, and then input into the state space feature extraction module F4 to obtain a small scale fusion feature map. Then, the fusion feature maps of these three scales are concatenated and upsampled to the original image size. Then, 1×1 convolution is used to adjust the dimension of the feature map, and then processed through two 3×3 convolutions and one 1×1 convolution to finally output the defect target category and location, completing the photovoltaic panel defect detection.
[0073] Optionally, as Figure 7 shown, for the state space feature extraction module, the input feature map is first preliminarily processed through convolution, batch normalization, and the Leaky ReLU activation function. The preliminarily processed feature is input into the local space module, which is composed of depthwise separable convolution, batch normalization, convolution, the Leaky ReLU activation function, and convolution stacked in sequence, and the output of the module is added to the input through a residual structure.
[0074] After the output of the local space module is processed by layer normalization, it is input into the multi-spectral state space model to extract and combine the information of the spectral dimension and the spatial dimension, enhancing the expression of photovoltaic panel defect features;
[0075] Finally, the obtained feature map is input into a linear layer for further processing, and residual connection is adopted. The output of the Leaky ReLU activation function in the preliminary processing is added and connected to the outputs of the multi-spectral state space model and the linear layer respectively to obtain the output feature map.
[0076] Optionally, as Figure 8 shown, for the multi-spectral state space model, the input first undergoes image serialization processing, cutting the feature map into 9 spatial slices and converting it into an image sequence;
[0077] Then these image sequences are respectively input into the spectral dimension branch and the spatial dimension branch for processing. In the spatial dimension branch, the multi-spectral state space model performs linear modeling on the 9 slices of the image sequence along the spatial dimension. Considering that the image sequence has context relationships and spatial dependencies in four directions, the image sequence is first scanned forward and backward row by row and column by column, and block matrix rearrangement is performed. The four groups of rearranged image sequences are input into the SSM module. After the four output feature sequences are processed by summation and averaging, they are decoded back into an image tensor and passed through a layer of convolution to finally obtain the spatial dimension features;
[0078] In the spectral dimension branch, the multi-spectral state space model performs linear modeling on the c spectral channels of the image sequence. First, the spectral dimension sequence is adjusted, and square matrix rearrangement is performed through forward and backward scanning. The two groups of rearranged image sequences are input into the SSM module. After the two output image feature sequences are also processed by summation and averaging, they are adjusted to the original dimension, decoded back into an image tensor, and passed through a layer of convolutional layer to obtain the final spectral dimension features;
[0079] Finally, the spatial dimension features and the spectral dimension features are fused to learn the joint features of the target in space and spectrum, generating the final feature map output.
[0080] Optionally, the SSM module is a discrete state space model, obtained by discretizing the continuous state space model. By discretizing the state transition process in continuous time, the original continuous state transition matrix and projection matrix are converted into forms suitable for discrete time, realizing effective modeling of time series data.
[0081] In the implementation of the multi-spectral state space model in the embodiments of the present invention, the SSM module can better capture the spatial features and spectral information of the input data by performing sequence modeling on the input data in the spatial dimension and spectral dimension. Especially in the spectral dimension, the SSM with sequential modeling can adapt to the sequential changes of the spectral bands, improve the modeling accuracy of the model for spectral information, and enhance the performance of the overall feature learning.
[0082] In the embodiments of the present invention, the training process of the overall model composed of the lightweight photovoltaic panel positioning and detection network based on variable aggregation convolution in the first stage, the time-series dynamic image registration module, and the photovoltaic panel defect detection network in the second stage is as follows:
[0083] 1) Data collection.
[0084] In the context of a photovoltaic power station, use a drone equipped with a high-resolution RGB camera and a multi-spectral state camera to take aerial photos of the photovoltaic power station area to obtain multiple segments of RGB images and multi-spectral state image data at different times and perspectives.
[0085] Construction of the photovoltaic panel positioning dataset and the photovoltaic panel defect detection dataset.
[0086] The photovoltaic panel positioning dataset is labeled by annotating the high-resolution RGB aerial photos, marking the photovoltaic panel area as the detection and positioning target, and the remaining areas as the background. The photovoltaic panel defect detection dataset is labeled based on the cropped multi-spectral state photovoltaic panel defect images, marking defect types such as hot spots and fragments as the detection targets, and other areas as the background. After annotation, the photovoltaic panel positioning dataset and the defect detection dataset are respectively divided into a training set and a test set according to the ratio of 8:2.
[0087] 3) Model training.
[0088] Compare the outputs of photovoltaic panel position detection and photovoltaic panel defect detection with the corresponding ground truth labels, calculate their respective detection loss functions respectively, and then update the corresponding network parameters by backpropagation according to the detection loss functions to optimize the model. After multiple rounds of iterative training, finally obtain a photovoltaic panel positioning and photovoltaic panel defect detection model with the optimal detection accuracy.
[0089] During the training process of the time-series dynamic image registration module, by calculating the MSD loss function between the reference image I F and the calibrated image IR, use an optimization algorithm for backpropagation, update the network parameters and complete multiple rounds of training optimization. The formula of the MSD loss function is as follows:
[0090]
[0091] where I F and I Mrespectively represent the reference image and the corresponding multi-spectral image, n represents the total number of pixels in the image, and I F(i) represents the reference image I F the value of the i-th pixel, I M (T (i,p) ) represents the corresponding multi-spectral image I after the geometric image transformation T (i,p) the value of the i-th pixel in M .
[0092] As Figure 9 shown, an embodiment of the present invention also provides a two-stage photovoltaic panel defect detection system based on a multi-spectral space model. The system includes:
[0093] An acquisition module 910, configured to acquire a high-resolution RGB aerial photovoltaic panel image to be detected and a multi-spectral aerial photovoltaic panel image;
[0094] A positioning and detection module 920, configured to input the RGB aerial photovoltaic panel image into a lightweight photovoltaic panel positioning and detection network based on variable aggregation convolution in the first stage. The photovoltaic panel positioning and detection network performs positioning and detection on the position of the photovoltaic panel in the RGB aerial photovoltaic panel image and outputs the position of the photovoltaic panel;
[0095] A registration module 930, configured to use the RGB aerial photovoltaic panel image as a reference image and the multi-spectral aerial photovoltaic panel image as an image to be calibrated, and input them into a temporal dynamic image registration module. The temporal dynamic image registration module uses a deep learning registration method based on a skip network to perform precise registration of a single-frame RGB image and a multi-spectral image, and obtains a multi-spectral calibrated image through temporal dynamic accumulation operations of multiple frames of images;
[0096] A cropping module 940, configured to crop the multi-spectral calibrated image according to the position of the photovoltaic panel output by the photovoltaic panel positioning and detection network to obtain a multi-spectral photovoltaic panel defect image of the corresponding area;
[0097] A defect detection module 950, configured to input the multi-spectral photovoltaic panel defect image into a photovoltaic panel defect detection network in the second stage. The photovoltaic panel defect detection network uses a multi-spectral space model to perform linear modeling on the spectral and spatial dimensions of the image, effectively adapts to the characteristics of spectral order changes, enhances the characterization ability of defect features, and outputs the defect category and defect detection frame of the photovoltaic panel to be detected, completing the photovoltaic panel defect detection.
[0098] The function structure of a two-stage photovoltaic panel defect detection system based on a multi-spectral space model provided by an embodiment of the present invention corresponds to a two-stage photovoltaic panel defect detection method provided by an embodiment of the present invention, and will not be elaborated here.
[0099] Figure 10 FIG. Figure 10 is a schematic structural diagram of an electronic device 1000 provided by an embodiment of the present invention. The electronic device 1000 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1001 and one or more memories 1002. Among them, at least one instruction is stored in the memory 1002, and the at least one instruction is loaded and executed by the processor 1001 to implement the steps of the above-mentioned dual-stage photovoltaic panel defect detection method based on the multi-spectral state space model.
[0100] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions. The above instructions can be executed by a processor in a terminal to complete the above-mentioned dual-stage photovoltaic panel defect detection method based on the multi-spectral state space model. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0101] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a magnetic disk, or an optical disc, etc.
[0102] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A two-stage photovoltaic panel defect detection method based on a multi-spectral space model, characterized in that: The method comprises: S1. Acquire high-resolution RGB aerial photovoltaic panel images and multi-spectral aerial photovoltaic panel images to be detected; S2, inputting the RGB aerial photovoltaic panel image into the first stage lightweight photovoltaic panel positioning detection network based on variable aggregation convolution, the photovoltaic panel positioning detection network performs positioning detection on the position of the photovoltaic panel in the RGB aerial photovoltaic panel image, and outputs the position of the photovoltaic panel; S3, taking the RGB aerial photovoltaic panel image as a reference image and the multi-spectral aerial photovoltaic panel image as an image to be calibrated, and inputting them into a time series dynamic image registration module, wherein the time series dynamic image registration module uses a deep learning registration method based on a skip network to perform accurate registration of a single-frame RGB image and a multi-spectral image, and obtains a multi-spectral calibration image through a time series dynamic accumulation operation of multiple frames of images; S4, cropping the multi-spectral calibration image according to the position of the photovoltaic panel output by the photovoltaic panel positioning detection network to obtain a multi-spectral photovoltaic panel defect image of the corresponding area; S5. Input the multi-spectral photovoltaic panel defect image into the second-stage photovoltaic panel defect detection network. The photovoltaic panel defect detection network uses a multi-spectral space model to perform linear modeling on the image spectrum and spatial dimensions, effectively adapt to the characteristics of spectral sequence changes, enhance the characterization capability of defect features, output the defect category and defect detection frame of the photovoltaic panel to be detected, and complete the photovoltaic panel defect detection.
2. The method according to claim 1, characterized in that The photovoltaic panel positioning detection network includes a backbone network, a neck network, and a detection output terminal; The backbone network is divided into four stages, each of which is connected in sequence. The first three stages are each composed of a convolution module and a cross-stage feature extraction module in series, and the last stage adds an SPPF module on this basis; the convolution module includes a convolution layer, a batch normalization layer and a Leaky ReLU activation function, wherein the Leaky ReLU activation function is an improved version of ReLU, and the formula is as follows: In the formula, α is a small constant. When the input is negative, the output remains negative, but it will not be completely zero, and a non-zero gradient α is provided for the negative input, which helps the gradient to propagate more smoothly throughout the network, and effectively solves the problem of dead neurons that may occur in the standard ReLU during training. The cross-stage feature extraction module evenly splits the channel of the input feature map into two parts, inputs two branches respectively, one branch is processed by variable aggregation convolution, and the other branch is processed by variable aggregation convolution and bottleneck module. The outputs of the two branches are spliced and then processed by variable aggregation convolution to obtain the final result. The bottleneck module consists of two convolution modules, and the input and output are added by residual connection. The SPPF module includes three maximum pooling layers in series. The output of each maximum pooling layer is spliced with the original input, and then the final output is obtained by the convolution module. The last three stages of the backbone network output small, medium and large scale feature maps respectively, and there is a two-fold downsampling relationship between the sizes. These feature maps will be passed as input to the neck network. The neck network inputs the small, medium and large scale feature maps from the backbone network, and processes them through the cross-stage feature extraction module, convolution and upsampling operations: first, the small scale feature map is restored to the same spatial resolution as the medium scale feature map through upsampling, and is spliced with the medium scale feature map, and is input to the cross-stage feature extraction module L1. The output of L1 is restored to the same resolution as the large scale feature map through upsampling, and is spliced with the large scale feature map, and is input to the cross-stage feature extraction module L2 to generate large scale fusion features. The output of L2 is downsampled to the medium scale resolution through convolution, and is spliced with the output of L1, and is input to the cross-stage feature extraction module L3 to obtain medium scale fusion features. The output of L3 is restored to the small scale resolution through convolution, and is spliced with the small scale feature map, and is input to the cross-stage feature extraction module L4 to obtain a small scale fusion feature map. Through the top-down feature fusion and bottom-up feature enhancement path, high-level semantic information is effectively fused with low-level detail features, so as to achieve more accurate target detection in different scales and complex scenes; The detection output end splices the three scale fusion features from the neck network, restores them to the original image size through upsampling, and outputs the position of the photovoltaic panel through convolution operation.
3. The method according to claim 1, characterized in that The variable aggregate convolution embeds k parallel convolution kernels in a convolution layer, and dynamically adjusts the weights of each convolution kernel through an attention mechanism composed of an average pooling layer, a fully connected layer, a Leaky ReLU activation function, a fully connected layer, and a softmax stacked in sequence, so as to identify photovoltaic panels of different sizes and distinguish irrelevant backgrounds on high-resolution images. Each convolution kernel performs a convolution operation on the input image tensor and then multiplies it by its own weight. Finally, the final convolution result is obtained by aggregation and addition, and then the result is output through batch normalization and the Leaky ReLU activation function.
4. The method according to claim 1, characterized in that The temporal dynamic image registration module, in the single-frame image registration process, references the RGB image I F and the multi-spectral image I to be calibrated M They are input into the skip network to extract their respective feature maps. After the two feature maps are spliced, they are downsampled and flattened through the convolution layer and input into the fully connected layer to obtain the geometric transformation parameters P{t x ,t y ,θ,s x ,s y ,β x ,β y }, where t x ,t y is the translation parameter, θ is the rotation angle, s x 、s y is the scaling factor, β x , β y is the shear factor, through which the geometric image transformation adjustment is performed on the multi-spectral image to be calibrated to achieve registration, and the calibration image IR is obtained; The skip network takes an image as input, first extracts preliminary features through convolution and maximum pooling layers; then further extracts deep features through a three-fold dual-branch residual structure and a cross-stage feature extraction module. The dual-branch residual structure includes a single-layer convolution branch and a three-layer convolution branch. After fusion, it is processed using a Leaky ReLU activation function, and then further extracted through a two-layer cross-stage feature extraction module. The deep-level features of the image are finally output. The feature map is output.
5. The method according to claim 1, characterized in that The temporal dynamic image registration module, in the temporal dynamic accumulation operation, converts n frames of RGB images {I F1 ,I F2 ,I F3 ...I Fn } and the corresponding n-frame multi-spectral image {I M1 ,I M2 ,I M3 ...I Mn } to perform image registration and obtain the registered images {IR1, IR2, IR3...IR n }, then perform frame accumulation of IR1 and IR2 to obtain IR 12 , and IR 12 with I F2 Perform image registration and obtain IRR 12 ; IRR 12 Perform frame accumulation with IR3 to obtain IRR 123 , and then IRR 123 with I F3 Perform image registration and obtain IRRR 123 ; By repeating the above steps, the final multi-spectral calibration image IRR..R is obtained 12..n , where the frame accumulation operation refers to adding each pixel value corresponding to the two images and taking the average.
6. The method according to claim 1, characterized in that The photovoltaic panel defect detection network is composed of a photovoltaic panel defect detection backbone network and a multi-scale fusion detection output terminal; The photovoltaic panel defect detection backbone network first extracts photovoltaic panel defect features through a 3D convolution module, and the 3D convolution module is composed of 3D convolution, batch normalization and Leaky ReLU activation functions stacked in sequence, and the input and output are connected through residuals; next, the feature map is sent to the subsequent stage, and the subsequent stage is divided into four stages, each stage is connected in sequence, the first stage is composed of a state space feature extraction module; the second and third stages are composed of a downsampling module and a state space feature extraction module in series; the fourth stage adds an SPPF module on the basis of the first two stages to enhance the extraction of key features and multi-scale feature fusion. Through the above four stages, feature maps of different scales are generated, including large, medium and small scales, and the feature map size decreases according to a two-fold downsampling relationship. The generated three scale features will be used as the input of the multi-scale fusion detection output end; The multi-scale fusion detection output end inputs the small, medium and large scale feature maps from the photovoltaic panel defect detection backbone network, and first performs multi-scale fusion processing through the state space feature extraction module, convolution and upsampling operations, wherein the small scale feature map is restored to the same spatial resolution as the medium scale feature map by upsampling, and is spliced with the medium scale feature map, and input to the state space feature extraction module F1, the output of F1 is restored to the same resolution as the large scale feature map by upsampling, and after splicing with the large scale feature map, is input to the state space feature extraction module F2 to generate a large scale fusion feature, and the output of F2 The image is downsampled to a medium-scale resolution through convolution, and spliced with the output of F1, and input into the state-space feature extraction module F3 to obtain a medium-scale fusion feature. The output of F3 is restored to a small-scale resolution through convolution, spliced with the small-scale feature map, and input into the state-space feature extraction module F4 to obtain a small-scale fusion feature map. The fusion feature maps of these three scales are then spliced and restored to the original image size through upsampling. A 1×1 convolution is then used to adjust the dimension of the feature map, and then two 3×3 convolutions and one 1×1 convolution are performed to finally output the defect target category and position to complete the photovoltaic panel defect detection.
7. The method according to claim 6, characterized in that The state space feature extraction module, the input feature map is firstly extracted and processed by convolution, batch normalization and Leaky ReLU activation function, and the features after preliminary processing are input to the local space module, and the local space module is sequentially stacked by depthwise separable convolution, batch normalization, convolution, Leaky ReLU activation function and convolution, and the output of the module is added to the input through the residual structure; Then, the output of the local space module is input into the multi-spectral space model after layer normalization, and the spectral dimension and spatial dimension information are extracted to enhance the defect feature expression of the photovoltaic panel; Finally, the obtained feature map is input into the linear layer for further processing, and a residual connection is adopted to add and connect the output of the Leaky ReLU activation function in the preliminary processing with the output of the multi-spectral space model and the linear layer respectively to obtain the output feature map.
8. The method according to claim 7, characterized in that The multi-spectral spatial model, the input is first processed by image serialization, the feature map is cut into 9 spatial slices, and then converted into an image sequence; Then these image sequences are input into the spectral dimension branch and the spatial dimension branch for processing. In the spatial dimension branch, the multi-spectral space model linearly models the 9 slices of the image sequence along the spatial dimension. Considering that the image sequence has contextual relationships and spatial dependencies in four directions, the image sequence is firstly forwardly and backwardly scanned by rows and columns, and the block matrix is rearranged. The four groups of rearranged image sequences are input into the SSM module. The four feature sequences output are decoded back to the image tensor after addition and averaging, and finally the spatial dimension features are obtained after a layer of convolution. In the spectral dimension branch, the multi-spectral space model linearly models the c spectral channels of the image sequence. First, the spectral dimension sequence is adjusted, and the block matrix is rearranged through forward and reverse scanning. The two rearranged image sequences are input into the SSM module. The two output image feature sequences are also adjusted to the original dimension after addition and averaging, and decoded back to the image tensor. After a convolution layer, the final spectral dimension feature is obtained. Finally, the spatial dimension features are fused with the spectral dimension features to learn the joint features of the target in space and spectrum and generate the final feature map output.
9. The method according to claim 8, characterized in that The SSM module is a discrete state space model, which is obtained by discretizing the continuous state space model. By discretizing the continuous time state transition process, the original continuous state transfer matrix and projection matrix are converted into a form suitable for discrete time, thereby realizing effective modeling of time series data.
10. A dual-stage photovoltaic panel defect detection system based on a multi-spectral space model, characterized in that: The system comprises: An acquisition module, used to acquire high-resolution RGB aerial photovoltaic panel images and multi-spectral aerial photovoltaic panel images to be detected; A positioning detection module is used to input the RGB aerial photovoltaic panel image into a lightweight photovoltaic panel positioning detection network based on variable aggregation convolution in the first stage, and the photovoltaic panel positioning detection network performs positioning detection on the position of the photovoltaic panel in the RGB aerial photovoltaic panel image and outputs the position of the photovoltaic panel; A registration module is used to use the RGB aerial photovoltaic panel image as a reference image and the multi-spectral aerial photovoltaic panel image as an image to be calibrated, and input them into a time-series dynamic image registration module. The time-series dynamic image registration module uses a deep learning registration method based on a skip network to perform accurate registration of a single-frame RGB image and a multi-spectral image, and obtains a multi-spectral calibration image through a time-series dynamic accumulation operation of multiple frames of images; A cropping module, used to crop the multi-spectral calibration image according to the position of the photovoltaic panel output by the photovoltaic panel positioning detection network to obtain a multi-spectral photovoltaic panel defect image of the corresponding area; The defect detection module is used to input the multi-spectral photovoltaic panel defect image into the second-stage photovoltaic panel defect detection network. The photovoltaic panel defect detection network uses a multi-spectral space model to perform linear modeling on the image spectrum and spatial dimensions, effectively adapt to the characteristics of spectral sequence changes, enhance the characterization capability of defect features, output the defect category and defect detection frame of the photovoltaic panel to be detected, and complete the photovoltaic panel defect detection.
Citation Information
Patent Citations
Photovoltaic module EL defect detection method based on image processing and deep learning fusion
CN113989241A
Photovoltaic panel detection method, unmanned aerial vehicle and computer readable storage medium
CN116385421A
Photovoltaic cell defect detection method
CN118735874A
Photovoltaic panel defect detection system and method, computer equipment and storage medium
CN118941493A