An ultra-lightweight photovoltaic module image segmentation method based on SegMamba
By adopting SegMamba's ultra-lightweight method in photovoltaic module image segmentation, using PMM Layer and multi-scale cross feature fusion technology, the problems of high computational complexity and low efficiency of the existing photovoltaic module image segmentation methods are solved, and high precision, real-time performance and low-cost photovoltaic module image segmentation are achieved.
Patent Information
- Application Number
- CN202411438524.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-10-15
AI Technical Summary
The existing photovoltaic module image segmentation methods have problems such as high computational complexity, low efficiency, insufficient robustness, insufficient model generalization capabilities, and poor details and edge processing.
Using the ultra-lightweight photovoltaic module image segmentation method based on SegMamba, image features are extracted through PatchEmbed2D, feature extraction and fusion are performed using Parallel Multi-Mamba Layer (PMM Layer) blocks, and high-resolution segmentation maps are generated by combining multi-scale cross-feature fusion decoder and image post-processing technology.
It significantly improves image segmentation accuracy, achieves efficient processing and real-time performance, reduces deployment costs, and improves the robustness and generalization capabilities of the model.
Smart Images

Figure CN119006498B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an ultra-lightweight photovoltaic module image segmentation method based on SegMamba. Background Art
[0002] Traditional photovoltaic module image segmentation methods are mostly based on manual annotation segmentation and some traditional deep learning methods. Manual methods use the quadrilateral features of photovoltaic arrays to segment photovoltaic arrays, but the segmentation accuracy is limited and time-consuming. In recent years, the Transformer model based on the self-attention mechanism has been applied to the field of computer vision. It realizes the interaction of global information through the self-attention mechanism, which can reveal the dependencies between different positions and scales of input data. It has certain advantages over traditional manual methods and traditional models (such as FCN, UNet, Segformer, etc.) in IRT image segmentation, but it still has high computational complexity and low efficiency.
[0003] Dependence on specific features and lack of robustness: Traditional segmentation methods based on quadrilateral features are highly dependent on specific geometric features and require frequent parameter adjustments in the face of changing environments, which limits the segmentation accuracy and robustness.
[0004] High computational cost: Although the self-attention mechanism Transformer model (such as Segformer) can capture global information and dependencies between different scales, it has high computational complexity when processing high-resolution images and is not suitable for scenarios with high real-time requirements.
[0005] Insufficient model generalization ability: Although existing deep learning methods perform well, they are greatly affected by the diversity and quality of training data. When faced with new scenarios, they lack generalization ability, which affects segmentation accuracy.
[0006] Details and edges are not processed precisely: Although models such as Segformer can capture global information and dependencies between different scales, they may not be precise enough when processing image details and edges, especially in key areas such as the boundaries of photovoltaic arrays and shadow junctions, where segmentation discontinuities or blurred edges may occur. For this reason, the present invention proposes an ultra-lightweight photovoltaic module image segmentation method based on SegMamba. Summary of the invention
[0007] The object of the present invention is to provide an ultra-lightweight photovoltaic module image segmentation method based on SegMamba to solve the problems raised in the above background technology.
[0008] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: an ultra-lightweight photovoltaic module image segmentation method based on SegMamba, the specific steps comprising: S1: data preprocessing and optimization: S11: using a drone equipped with a dual-spectrum camera to perform low-altitude high-resolution photography, and simultaneously capturing infrared and visible light images of the photovoltaic array; after multiple shots, a total of several infrared and visible light images are obtained, each image having a resolution of 640×512 pixels; using infrared thermal imaging for analysis; after screening, finally retaining high-quality infrared thermal images;
[0009] S12: Use the Labelme annotation tool to manually annotate the infrared image to obtain accurate image label information, divide the image into two categories: background and photovoltaic array Pv, and annotate it in PASCAL VOC format;
[0010] S13: To eliminate the noise in the image, a two-step filtering strategy is adopted. First, bilateral filtering technology is applied to reduce the edge blur and detail loss in the image. Then, Gaussian filtering is used to further smooth the image and remove the remaining random noise.
[0011] S2: Feature extraction: This paper proposes an ultra-lightweight encoder to solve the problem of application overhead caused by large segmentation model calculation and parameter quantity. The image is extracted through PatchEmbed2D, and then the number of channels of the image is changed through a block composed of two PMM Layer blocks in series, and then a normalization operation is performed to generate the first stage feature map. Similarly, the feature map generated in the previous stage is subjected to the above operation to generate the feature map of the next stage, and finally four images with different resolutions are generated;
[0012] S3: Feature Fusion: The decoder uses multi-scale cross-feature fusion to efficiently restore the resolution of the image from multiple scale levels through horizontal cross connections. First, the features of different resolutions generated by the encoder in four different stages and the post-processed features are used as input. A recursive and cross-level strategy is adopted. Through multiple iterations of upsampling and skip connection mechanisms, the rich semantic information of the high level is directly passed to the lower level, deeply and carefully fused with the features of the subsequent stages and converted into the final segmentation mask. A high-resolution segmentation map is generated through the classification and segmentation module cls_seg.
[0013] S4: Image post-processing: Morphological processing is used to eliminate tiny noises and fill defects in the target area; edge smoothing technology is used to smooth segmentation boundaries and reduce jagged edges; regional optimization strategies are used to merge adjacent or subdivide overly large areas to accurately depict object contours; and threshold fine-tuning is used to further improve segmentation accuracy.
[0014] Preferably, in each PMM Layer block, the feature X with the number of channels C is first normalized and then divided into Four features, each with C / 4 channels; each feature is then simultaneously input into a PM Block consisting of two serial Mamba modules for convolution; in each channel, 1 / 4 of the original input feature is added to the output feature after passing through a Sigmoid function, which ensures that the total number of channels remains unchanged and maintains high accuracy while minimizing the number of parameters to make the model lightweight
[0015] Preferably, the process formula of step S2 is expressed as:
[0016]
[0017] i=1,2,3,4
[0018] i = 1,2,3,4
[0019] i=1,2,3,4
[0020]
[0021]
[0022] in, is the input feature with C channels; The number of channels output after being evenly divided is C / 4 feature; represent The result after passing the sigmoid activation function; for The output characteristics after passing through the PMM Layer module; For the general and The result after addition; It is the result of connecting the four parallel output results, and the number of channels is C; is the final mapped feature; LN is LayerNorm, which represents the layer normalization function; Sp is the Split operation used to evenly divide the input features; Sigmoid is the activation function; PM is the PMMLayer operation, and Cat is the Concat operation used to fuse the extracted features output by four independent channels to generate features with a channel number of C; Pro is the linear projection operation, which helps to integrate information from different sub-channels and Mamba modules to ensure the consistency and integrity of the features.
[0023] Preferably, the encoder is composed of PatchEmbed2D, PMM Layer blocks and Groupnorm normalization blocks which are alternately connected in series.
[0024] Preferably, the step S4 can also effectively distinguish the foreground and background of the image by adaptive threshold processing Otsu method to generate a clear binary image.
[0025] Compared with the prior art, the present invention has the following beneficial effects: the image segmentation accuracy is significantly improved:
[0026] By introducing the Parallel Multi-Mamba Layer (PMM Layer), feature diversity and deep mining are achieved in the feature extraction stage. This parallel complementary mechanism accurately captures the subtle features and key differences of photovoltaic module images, greatly improving the segmentation accuracy and providing a solid guarantee for applications such as photovoltaic module defect detection and crack identification.
[0027] Efficient processing and real-time performance:
[0028] While pursuing high precision, the ultra-lightweight design greatly reduces the number of model parameters, achieving efficient processing and real-time performance. The number of parameters in the data set used by the present invention is as low as 0.85M, which is about 78% lower than the traditional segmentation method (such as segformer). This feature enables the invention to operate efficiently in an environment with limited computing resources, shorten processing time, and improve processing speed. This is especially important for photovoltaic monitoring systems that require real-time feedback, which can ensure timely detection and response to potential problems and improve overall operation and maintenance efficiency.
[0029] Deployment costs are significantly reduced:
[0030] The lightweight design of the model not only reduces the demand for computing resources, but also significantly reduces the cost of hardware deployment. Photovoltaic power stations and distributed photovoltaic systems can flexibly select hardware equipment according to actual needs, without relying on high-cost high-performance computing equipment, thereby reducing the overall investment threshold and enhancing the deployability and flexibility of the system.
[0031] Robustness and generalization ability:
[0032] Mamba's structure may have better noise resistance than convolution and Transformer for noise and abnormal data in photovoltaic array infrared images. Since Mamba's state space model contains memory capabilities and can process continuous input data more smoothly, it can provide more robust predictions when there are abnormal points or lighting changes in the image, and better adapt to complex and changeable practical application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0034] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0035] See also Figure 1 , which is the first embodiment of the present invention, provides a technical solution: a method for image segmentation of ultra-lightweight photovoltaic modules based on SegMamba, the specific steps of which include: S1: data preprocessing and optimization: S11: using a drone equipped with a dual-spectrum camera to perform low-altitude high-resolution photography, and simultaneously capture infrared and visible light images of the photovoltaic array; after multiple shots, a total of several infrared and visible light images are obtained, each with a resolution of 640×512 pixels; infrared thermal imaging is used for analysis; after screening, high-quality infrared thermal images are finally retained;
[0036] S12: Then, the infrared image is manually labeled using the Labelme labeling tool to obtain accurate image label information, and the image is divided into two categories: background and photovoltaic array Pv, and labeled using the PASCAL VOC format;
[0037] S13: In view of the fact that infrared thermal imaging technology is susceptible to the dual influence of external factors such as reflection interference, medium absorption and reflection, background radiation, and internal factors of the instrument, the image is often accompanied by significant noise; a two-step filtering strategy is adopted, firstly, bilateral filtering technology is applied to reduce edge blur and detail loss in the image, and then Gaussian filtering is used to further smooth the image and remove the remaining random noise;
[0038] S2: Feature extraction: First input The image is extracted through PatchEmbed2D, and then the number of channels of the image is changed through a block composed of two PMM Layer blocks in series, and then a normalization operation is performed to generate the first stage feature map. Similarly, the feature map generated in the previous stage is subjected to the above operation to generate the feature map of the next stage, and finally four images with different resolutions are generated;
[0039] S3: Feature fusion: The decoder adopts multi-scale cross-feature fusion to efficiently restore the resolution of the image from multiple scale levels through horizontal cross-connections. First, the features of different resolutions generated by the encoder in four different stages and the post-processed features are taken as input. Taking the features generated in the fourth stage as an example, the resolution is first adjusted to the same as stage3 by upsampling, that is, Then, the PMM Layer or ConvModule (1x1 convolution) is used to halve the number of feature channels of the adjusted stage4 and the original stage3 to C / 2 for fusion; a recursive and cross-level strategy is adopted, through multiple iterations of upsampling and skip connection mechanism, the rich semantic information of the high level is directly transmitted to the lower level, deeply and carefully fused with the features of the subsequent stages and converted into the final segmentation mask; a high-resolution segmentation map is generated through the classification and segmentation module cls_seg;
[0040] S4: Image post-processing: Morphological processing is used to eliminate tiny noises and fill defects in the target area; edge smoothing technology is used to smooth segmentation boundaries and reduce jagged edges; regional optimization strategies are used to merge adjacent or subdivide overly large areas to accurately depict object contours; and threshold fine-tuning is used to further improve segmentation accuracy.
[0041] In this embodiment, preferably, in each PMM Layer block, the feature X with the number of channels C is first normalized and then divided into Four features, each with C / 4 channels; each feature is then simultaneously and parallelly input into a PM Block consisting of two serial Mamba modules for convolution; in each channel, 1 / 4 of the input original features are added to the output features after passing through a Sigmoid function, which ensures that the total number of channels remains unchanged and maintains high accuracy while minimizing the number of parameters to make the model lightweight.
[0042] In this embodiment, preferably, the process formula of step S2 is expressed as:
[0043]
[0044] i=1,2,3,4
[0045] i = 1,2,3,4
[0046] i=1,2,3,4
[0047]
[0048]
[0049] in, is the input feature with C channels; The number of channels output after being evenly divided is C / 4 feature; represent The result after passing the sigmoid activation function; for The output characteristics after passing through the PMM Layer module; For the general and The result after addition; It is the result of connecting the four parallel output results, and the number of channels is C; is the final mapped feature; LN is LayerNorm, which represents the layer normalization function; Sp is the Split operation used to evenly divide the input features; Sigmoid is the activation function; PM is the PMMLayer operation, and Cat is the Concat operation used to fuse the extracted features output by four independent channels to generate features with a channel number of C; Pro is the linear projection operation, which helps to integrate information from different sub-channels and Mamba modules to ensure the consistency and integrity of the features.
[0050] In this embodiment, preferably, the encoder is composed of PatchEmbed2D, PMM Layer blocks and Groupnorm normalization blocks which are alternately connected in series.
[0051] In this embodiment, preferably, the step S4 can also effectively distinguish the foreground and background of the image through the adaptive threshold processing Otsu method to generate a clear binary image.
[0052] Although embodiments of the present invention have been shown and described (see the detailed description above for details), it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for ultra-lightweight photovoltaic module image segmentation based on SegMamba, characterized by: The specific steps include: S1: data preprocessing and optimization: S11: using a drone equipped with a dual-spectrum camera to take low-altitude high-resolution photos, and simultaneously capture infrared and visible light images of the photovoltaic array; after multiple shots, a total of several infrared and visible light images were obtained, each with a resolution of 640×512 pixels; infrared thermal imaging was used for analysis; after screening, high-quality infrared thermal images were finally retained; S12: Use the Labelme annotation tool to manually annotate the infrared image to obtain accurate image label information, divide the image into two categories: background and photovoltaic array Pv, and annotate it in PASCAL VOC format; S13: To eliminate the noise in the image, a two-step filtering strategy is adopted. First, bilateral filtering technology is applied to reduce the edge blur and detail loss in the image. Then, Gaussian filtering is used to further smooth the image and remove the remaining random noise. S2: Feature extraction: To address the application overhead caused by the large number of segmentation model calculations and parameters, an ultra-lightweight encoder is proposed. In this encoder, an H×W×3 image is first input, and image features are extracted through PatchEmbed2D. Then, a block composed of two ParallelMulti-Mamba Layer blocks is used to change the number of channels of the image and perform a normalization operation to generate the first-stage feature map. Similarly, the feature map generated in the previous stage is subjected to the above operations to generate the feature map of the next stage. Finally, four images with different resolutions are generated. S3: Feature Fusion: The decoder uses multi-scale cross-feature fusion to efficiently restore the resolution of the image from multiple scale levels through horizontal cross connections. First, the features of different resolutions generated by the encoder in four different stages and the post-processed features are used as input. A recursive and cross-level strategy is adopted. Through multiple iterations of upsampling and skip connection mechanisms, the rich semantic information of the high level is directly passed to the lower level, deeply and carefully fused with the features of the subsequent stages and converted into the final segmentation mask. A high-resolution segmentation map is generated through the classification and segmentation module cls_seg. S4: Image post-processing: Morphological processing is used to eliminate tiny noises and fill defects in the target area; edge smoothing technology is used to smooth segmentation boundaries and reduce jagged edges; regional optimization strategies are used to merge adjacent or subdivide overly large areas to accurately depict object contours; and threshold fine-tuning is used to further improve segmentation accuracy.
2. The ultra-lightweight photovoltaic module image segmentation method based on SegMamba according to claim 1, characterized in that: In each ParallelMulti-Mamba Layer block, the feature X with the number of channels C is first normalized and then divided into Four features, each with C / 4 channels; then each feature is simultaneously and parallelly input into the PMM Block composed of two serial Mamba modules for convolution operation; in each channel, 1 / 4 of the original input features are added to the output features after passing through the Sigmoid function, and finally the four features are added. This parallel convolution strategy can extract the spatial dimension features of the image at different scales at the same time, and realize the comprehensive and accurate capture of local details and global context information, ensuring the diversity and depth of feature extraction, and minimizing the number of parameters while ensuring that the total number of channels remains unchanged and maintains high precision, making the model lightweight.
3. The ultra-lightweight photovoltaic module image segmentation method based on SegMamba according to claim 2, characterized in that: The process formula of step S2 is expressed as: Out=Pro(X out ) in, is the input feature with C channels; The number of channels output after being evenly divided is C / 4 feature; represent The result after passing the sigmoid activation function; for The result features after passing through the ParallelMulti-Mamba Layer module; Y i C / 4 For the general and The result after addition; X out It is the result of connecting the four parallel output results, and its number of channels is C; Out is the feature after the final mapping; LN is the LayerNorm representation layer normalization function; Sp is the Split operation used to evenly divide the input features; Sigmoid is the activation function; PM is the Parallel Multi-Mamba Layer operation, and Cat is the Concat operation used to fuse the extracted features output by four independent channels to generate features with the number of channels C; Pro is the linear projection operation Linear Projection, which helps to integrate information from different sub-channels and Mamba modules to ensure the consistency and integrity of the features.
4. The ultra-lightweight photovoltaic module image segmentation method based on SegMamba according to claim 1, characterized in that: The encoder is composed of PatchEmbed2D, Parallel Multi-Mamba Layer blocks and Groupnorm normalization blocks, which are alternately connected in series.
5. The ultra-lightweight photovoltaic module image segmentation method based on SegMamba according to claim 1, characterized in that: The step S4 can also effectively distinguish the foreground and background of the image by using the adaptive threshold processing Otsu method to generate a clear binary image.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method based on double-branch feature fusion
CN115797931A
Battery pack SOC and SOE estimation method driven by Mama architecture
CN118444159A