A ground cloud map fine-grained segmentation method based on an improved encoder-decoder structure

By using a U-shaped network with an improved encoder-decoder structure, combined with selectable channel attention modules, parallel dilated convolution, and CARAFE upsampling modules, the problem of binary segmentation of clouds and sky in ground-based cloud image segmentation is solved, realizing the foundation for fine-grained segmentation and photovoltaic power prediction.

CN117911693BActive Publication Date: 2026-05-26NORTH CHINA ELECTRIC POWER UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTH CHINA ELECTRIC POWER UNIV
Filing Date
2023-12-29
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing ground-based cloud image segmentation methods only distinguish between clouds and sky as binary values, which cannot meet the needs of photovoltaic power prediction. Furthermore, they do not consider the impact of clouds on solar irradiance and lack fine-grained segmentation datasets.

Method used

A U-shaped network with an improved encoder-decoder structure is adopted, and selectable channel attention modules, parallel dilated convolution modules, and CARAFE upsampling modules are introduced to construct a fine-grained segmentation dataset of ground cloud images, thereby improving cloud recognition accuracy and fine-grained feature extraction.

Benefits of technology

It achieves high-precision fine-grained segmentation of ground-based cloud maps, meeting the basic requirements for photovoltaic power prediction and improving the accuracy of cloud map segmentation and fine-grained feature extraction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117911693B_ABST
    Figure CN117911693B_ABST
Patent Text Reader

Abstract

This invention discloses a fine-grained segmentation method for ground cloud images based on an improved encoder-decoder structure, comprising the following steps: constructing a fine-grained dataset of ground cloud images, classifying clouds into five categories and manually labeling them to obtain fine-grained ground cloud images of different types; selecting a U-shaped network as the base model; introducing selectable channel attention before encoder downsampling to better capture the multi-scale features of fine-grained ground cloud images; proposing a parallel dilated convolution module to replace the convolutional blocks of the original bottleneck layer structure, increasing the receptive field range and further learning fine-grained features; introducing a CARAFE upsampling module to avoid the inability to recover semantic information during upsampling; training the model and comparing it with other segmentation models. This invention applies a U-shaped network to ground cloud image segmentation, achieving fine-grained segmentation of ground cloud images, outputting multi-category color images, and achieving the highest segmentation accuracy compared to other semantic segmentation networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image analysis technology, and in particular to a fine-grained segmentation method for ground cloud maps based on an improved encoder-decoder structure. Background Technology

[0002] Clouds are a common and important natural phenomenon in daily life. They are aggregates of various shapes floating in the sky, composed of water droplets, ice crystals, or a mixture of both, and cover more than 60% of the global land area. Cloud observation is an important and challenging task. With the continuous invention and improvement of various hardware facilities, cloud observation has gradually shifted from manual to automated observation. Currently, automated observation is divided into satellite cloud observation and ground-based cloud observation. Satellite cloud observation is more suitable for describing large-scale cloud information and its changes, while ground-based cloud observation is an important way to accurately obtain local sky cloud cover information. Ground-based cloud observation is convenient to operate, easy to observe, and has low equipment costs. It mainly uses ground-based all-sky cloud imagers to acquire ground-based cloud images, which mainly reflect the cloud base information, cloud distribution, changes, and movement in local areas of the sky. It can further effectively and accurately obtain the three elements of cloud shape, cloud coverage, and cloud base height, which is of great significance for further accurate prediction of photovoltaic power. The ground-based cloud observation system mainly consists of two parts: ground-based cloud image segmentation and classification. Ground-based cloud image segmentation is the basis of ground-based cloud image classification and an important component of photovoltaic power prediction. High-precision ground cloud map segmentation is one of the key factors for accurate prediction of photovoltaic power. Therefore, this paper focuses on the automatic segmentation of ground cloud maps.

[0003] Ground-based cloud image segmentation methods are mainly divided into traditional thresholding methods and deep learning methods. Thresholding methods utilize the grayscale difference between cloud targets and the sky background extracted from cloud images, treating the image as a combination of two different grayscale regions. A suitable threshold is selected to determine whether each pixel in the image belongs to the cloud or the sky, thus generating a binary segmented image. However, thresholding segmentation methods have low segmentation accuracy and cannot perform further fine-grained segmentation, failing to meet the needs of photovoltaic power prediction.

[0004] In recent years, with the development of artificial intelligence technology, using deep learning to process ground-based cloud images and then using computer vision and image processing technology to segment the ground-based cloud images has become the main method.

[0005] However, there are two main challenges in using deep learning methods to segment ground-based cloud images:

[0006] 1. Solar energy, with its advantages of safety, high efficiency, abundant resources, and a strong industrial foundation, has been widely used globally. Currently, solar power generation is mainly achieved through photovoltaics (PV). PV power generation is affected by multiple factors, including local cloud cover variations, solar irradiance, and solar cell performance. Among these, local cloud cover variations are a significant cause of unstable and intermittent PV power generation. Accurate cloud cover information obtained through cloud observation is crucial for the precise prediction of PV power generation. However, current ground-based cloud image segmentation methods, which only separate clouds from the sky, cannot accurately predict future PV power, and there is a lack of suitable fine-grained ground-based cloud image segmentation datasets for both cloud image segmentation and subsequent PV power prediction.

[0007] 2. Ground-based cloud images contain a wealth of semantic information. Different types of clouds have varying irradiance, thus affecting photovoltaic power prediction differently. However, currently available ground-based cloud image segmentation datasets are only suitable for binary segmentation of clouds and sky background. When classifying clouds based on images, cloud height information should be ignored, focusing primarily on cloud texture features, brightness information, and location distribution. Ideal cloud classification should consider the physical processes of cloud formation, development, and evolution. Internationally, cloud classification is mainly based on cloud brightness, color, size, and height characteristics. The World Meteorological Organization's cloud classification also does not consider the impact of clouds on solar irradiance, and therefore is not entirely applicable to photovoltaic power prediction. Summary of the Invention

[0008] Against the backdrop of the aforementioned technology, this invention ignores the influence of cloud height. Under the guidance of meteorological experts, it reclassifies clouds into five categories based on the radiation characteristics, cloud shape, and meteorological characteristics of different types of clouds. This leads to the construction of a fine-grained segmentation dataset for ground-based cloud images. The fine-grained images are then applied to ground-based cloud image segmentation to solve the problems existing in ground-based cloud segmentation. This allows the cloud images to meet the requirements of fine-grained segmentation while laying a foundation for subsequent photovoltaic power prediction.

[0009] The purpose of this invention is to provide a fine-grained segmentation method for ground-based cloud images based on an improved encoder-decoder structure, addressing the problem that current ground-based cloud segmentation methods merely distinguish between clouds and sky, often resulting in binary classification. This invention utilizes a U-shaped network with an improved encoder-decoder structure, introduces selectable channel attention modules, and designs a parallel dilated convolution module and a CARAFE upsampling method to further mine fine-grained features of ground-based cloud images, thereby improving segmentation accuracy.

[0010] To achieve the above objectives, the present invention provides the following solution:

[0011] A ground-based cloud map fine-grained segmentation method based on an improved encoder-decoder structure includes the following steps:

[0012] Ignoring the influence of cloud height, clouds are reclassified into five categories based on solar irradiance characteristics, cloud shape, and meteorological characteristics of different cloud types. Ground-based cloud images are labeled to obtain color-labeled images of different categories, and a fine-grained segmentation dataset of ground-based cloud images is constructed. The five cloud types are: cumulonimbus and nimbostratus, cumulus, stratus, cirrus, and clear sky.

[0013] The fine-grained segmentation dataset of the ground cloud map is constructed and input into the preset deep learning model to obtain fine-grained color segmentation images.

[0014] The preset deep learning model is as follows: an encoder-decoder structure is selected as the basic architecture, and an improved U-shaped network is used as the backbone network. In the process of improving the U-shaped network, an optional channel attention module is introduced to improve cloud recognition accuracy and fully reflect the fine-grained attributes of the cloud map. A parallel dilated convolution module is proposed to increase the receptive field while keeping the feature map size unchanged, and to further mine the fine-grained semantic information of the feature map. A CARAFE upsampling module is introduced to avoid more information loss and better recover fine-grained semantic features.

[0015] In the process of improving the U-shaped network, an optional channel attention module is introduced to improve cloud recognition accuracy and fully reflect the fine-grained attributes of cloud images. Specifically, this includes:

[0016] SK attention is a channel attention mechanism with dynamic selectivity. It has three operations: Split, Fuse, and Select. It can adaptively adjust the size of the receptive field to capture target objects of different scales, thereby improving the recognition accuracy of blocky clouds and cloud edges in ground-based cloud maps and fully reflecting the fine-grained attributes of cloud maps. For a given feature map, two transformations are first performed to obtain two feature maps, U1 and U2. Then, they are fused by pixel-wise summation to obtain U, as shown in equation (1).

[0017] U = U1 + U2 (1)

[0018] Then, channel statistics are generated by simply embedding global information using global average pooling. Furthermore, a compact feature is created to guide precise and adaptive selection, implemented through a simple fully connected (fc) layer, which reduces dimensionality for efficiency.

[0019] The weight score corresponding to each channel is calculated and applied to the feature map. Finally, the feature map is fused to obtain the final output image. Compared with the input image, the output image has been refined by the information channels and incorporated more key information, thus enhancing the key information of the image.

[0020] In the process of improving the U-shaped network, a parallel dilated convolution module was proposed to increase the receptive field while keeping the feature map size unchanged, thereby further mining the fine-grained semantic information of the feature map, specifically including:

[0021] First, the input fine-grained feature map is fed into a 3×3 convolutional block for dimensionality reduction, and then passed through a normalization and ReLU activation function layer.

[0022] Secondly, there are three parallel branches. In order to better mine the fine-grained semantic features of the ground cloud map, dilated convolution is introduced. Multiple dilated convolutions are stacked. However, dilated convolution has a grid effect. When multiple dilated convolutions are stacked, some pixels are not utilized at all, which will lose some of the continuity and integrity of information, thus affecting the effect of fine-grained segmentation of the ground cloud map. To address this problem, a hybrid dilated convolution method is designed. When using multiple dilated convolutions, the common divisor of the stacked convolution dilation kernels is no greater than 1.

[0023] The first branch has one convolutional block with an inflation factor of 1; the second branch has two convolutional blocks with inflation factors of 1 and 3 respectively; the third branch has three convolutional blocks with inflation factors of 1, 3 and 5 respectively. The kernel size of all convolutional blocks is 3×3. Finally, the three branches are merged. Through the parallel dilated convolution module, fine-grained semantic information of the feature map can be further mined, and high-precision segmentation results can be achieved.

[0024] The formula for dilated convolution is:

[0025] K = (rate-1) × (k-1) + k (2)

[0026] In the formula, K is the dilated receptive field, rate is the dilated convolution rate, and k is the kernel size.

[0027] In the process of improving the U-shaped network, CARAFE upsampling is used to replace the linear upsampling of the original decoder, avoiding the loss of fine-grained information and better recovering semantic features. Specifically, this includes:

[0028] CARAFE consists of two steps: a kernel prediction module and a content-aware reassembly module. The kernel prediction module comprises three sub-modules: a channel compressor, a content encoder, and a kernel normalizer. It first reduces the number of channels from C to C using a 1×1 convolution. mReducing the number of channels in the input feature map leads to a reduction in parameters and computational costs in subsequent steps, making CARAFE more efficient. Within the same budget, a larger kernel size can also be used for the content encoder. Reducing the number of feature channels within an acceptable range does not harm performance. The content encoder then has a size k. encoder ×k encoder The convolution reduces the number of channels from C. m Change to C UP The purpose of this step is to generate a reconstructed kernel based on the content of the input features, and then process the N(χ) region in the kernel obtained in the previous step. l ,k up Normalization is performed using the Softmax normalization function to make the convolution weight kernel 1, as shown in equation (3) to obtain the size W. l' Kernel:

[0029] W l′ =φ(N(χ) l ,k encoder (3)

[0030] For the content-aware reorganization module, the features in the local area are reorganized by the function φ, as shown in equation (4). Under the action of the reorganization kernel, N(χ) l ,k up Each pixel within the region contributes differently to the upsampled pixels based on the feature content, allowing for greater focus on information about relevant points within the local region. Therefore, the reconstructed feature map possesses richer semantic information than the original feature map. The reconstructed kernel χ... l’ The calculation is as follows:

[0031]

[0032] This invention discloses the following technical effects: It provides a fine-grained segmentation method for ground-based cloud images based on an encoder-decoder architecture. The method constructs a fine-grained segmentation model for ground-based cloud images, using an encoder-decoder structure as the basic segmentation model and an improved U-shaped network as the backbone network. This reduces network parameters while extracting multi-scale features at different levels. To address the difficulty in distinguishing between different cloud types and the inability to capture small target clouds and blurry clouds, an optional channel attention module is introduced before the encoder downsampling operation to better extract fine-grained features from the ground-based cloud image. To address the low segmentation accuracy, a parallel dilated convolution module is designed to expand the receptive field without increasing computational load, fully mining the semantic information of the ground-based cloud image. To address the information loss during sampling on the original U-shaped network, a CARAFE upsampling module is introduced. The model is trained, and the decoder branch generates the final color prediction result, achieving the best segmentation accuracy compared to other semantic segmentation methods. This method applies fine-grained attributes to the field of ground-based cloud image segmentation, achieving fine-grained segmentation of ground-based cloud images while laying the foundation for subsequent photovoltaic power prediction. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a flowchart of a ground cloud map fine-grained segmentation method based on an improved encoder-decoder structure according to an embodiment of the present invention;

[0035] Figure 2 This is a schematic diagram of the selective channel attention module according to an embodiment of the present invention;

[0036] Figure 3 This is a schematic diagram of the structure of the parallel dilated convolution module in an embodiment of the present invention;

[0037] Figure 4 This is a schematic diagram of the overall structure of the CARAFE upsampling module according to an embodiment of the present invention;

[0038] Figure 5 This is a schematic diagram of the overall structure of an embodiment of the present invention;

[0039] Figure 6 This is a diagram illustrating the effect of fine-grained segmentation of the ground-based cloud in an embodiment of the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] The purpose of this invention is to provide a ground-based cloud image fine-grained segmentation method based on an improved encoder-decoder structure. This method addresses the shortcomings of existing ground-based cloud segmentation methods, which only distinguish between clouds and the sky using binary segmentation, do not consider the influence of solar irradiance in cloud classification, and cannot meet the requirements for subsequent photovoltaic power prediction. This invention lays the foundation for subsequent photovoltaic power prediction while satisfying the requirements of ground-based cloud fine-grained segmentation.

[0042] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0043] like Figure 1 As shown, the flowchart of a ground cloud map fine-grained segmentation method based on an improved encoder-decoder structure provided by the present invention includes the following steps:

[0044] Ignoring the influence of cloud height, clouds are reclassified into five categories based on solar irradiance characteristics, cloud shape, and meteorological characteristics of different cloud types. Ground-based cloud images are labeled to obtain color-labeled images of different categories, and a fine-grained segmentation dataset of ground-based cloud images is constructed. The five cloud types are: cumulonimbus and nimbostratus, cumulus, stratus, cirrus, and clear sky.

[0045] The fine-grained segmentation dataset of the ground cloud map is constructed and input into the preset deep learning model to obtain fine-grained color segmentation images;

[0046] The pre-defined deep learning model uses an encoder-decoder structure as its basic architecture and an improved U-shaped network as its backbone. The improved U-shaped network incorporates selectable channel attention modules to improve cloud recognition accuracy and fully reflect the fine-grained attributes of the cloud image. A parallel dilated convolution module is proposed to increase the receptive field while maintaining the feature map size, further mining the fine-grained semantic information of the feature map. A CARAFE upsampling module is introduced to avoid further information loss and better recover fine-grained semantic features.

[0047] Deep learning models require a large number of image samples for training. Ground-based cloud image segmentation is a crucial step in photovoltaic power generation prediction. However, currently available ground-based cloud image segmentation datasets are only suitable for binary segmentation of clouds and sky background. When classifying clouds based on images, cloud height information should be ignored, and the focus should be on cloud texture features, brightness information, and location distribution. For example, Buch and Sun et al. used texture analysis to classify cloud images taken by the All-Sky Imager (WSI) into five different categories: altocumulus, cirrus, stratus, cumulus, and clear sky. Heinle et al. used the K-nearest neighbor classification method to classify clouds into seven categories: cumulus, cirrus and cirrostratus, cirrocumulus and altocumulus, clear sky, stratocumulus, stratus and altostratus, cumulus rain and nimbostratus. Zhu et al. classified ground-based clouds into five categories based on factors such as irradiance and meteorological characteristics, but this only classified the entire image, not down to the pixel level. However, none of the above classification methods consider the impact of clouds on solar irradiance, and therefore are not applicable to photovoltaic power prediction.

[0048] Table 1. Fine-grained cloud classification criteria for photovoltaic power generation prediction

[0049]

[0050] The varying rates of attenuation of solar irradiance are a significant factor contributing to the low accuracy of photovoltaic (PV) power generation prediction. Ideal cloud classification should consider the physical processes of cloud formation, development, and evolution. Internationally, cloud classification is primarily based on brightness, color, size, and altitude characteristics. The World Meteorological Organization's cloud classification also fails to consider the impact of clouds on solar irradiance, thus it is not entirely applicable to PV power generation prediction. This invention ignores the influence of cloud height and, by comprehensively considering the solar irradiance characteristics, cloud shape, and meteorological characteristics of different cloud types, reclassifies clouds into five categories, as shown in Table 1: cumulonimbus and nimbostratus, cumulus, stratus, cirrus, and clear sky. Among these, cumulonimbus and nimbostratus are dark gray, the thickest, and have the largest area, typically occupying the entire sky. The solar irradiance attenuation rate of cumulonimbus and nimbostratus is greater than 90%. Cumulus clouds, appearing as clumps, are the most common type of cloud in clear skies. Their solar irradiance attenuation is greater than 90% when obstructing the sun, and close to zero when not obstructing it. Stratus clouds are thinner than cumulus clouds, with an attenuation rate of approximately 50%. Cirrus clouds are fibrous or elongated, white in color, and generally appear as small, isolated patches or on the outermost layer of cumulus clouds. Their solar irradiance attenuation is less than 10%. Clear skies exhibit a solar irradiance attenuation rate close to zero. The radiative characteristics of each cloud type were obtained by statistically analyzing the clear sky coefficient of multiple images for that cloud type. Referring to the cloud classification criteria in Table 1, and under the guidance of authoritative meteorological experts, 300 ground-based cloud images with rich, fine-grained characteristics were selected from the Kiel University ground-based cloud image dataset in Germany. These images were uniformly cropped to 1704×1704 pixels, and LabelMe software was used to label different cloud categories with different colors. Purple represents cumulonimbus and nimbostratus clouds, green represents cumulonimbus clouds, yellow represents stratus clouds, red represents cirrus clouds, and the rest represents clear skies in blue, thus constructing a fine-grained segmentation dataset of ground-based cloud maps for photovoltaic power generation prediction.

[0051] The encoder-decoder architecture is selected, and the U-shaped network is improved as the backbone network, specifically including:

[0052] A selective channel attention module is introduced before encoder downsampling to better capture clouds of different sizes; a parallel dilated convolution module is proposed to replace the convolution block of the original bottleneck layer, which can perform more refined segmentation of small target clouds and cloud genus boundaries; a CARAFE upsampling module is introduced to replace the original linear upsampling, which can better recover semantic information and prevent the loss of fine-grained features.

[0053] In improving the U-shaped network, an optional channel attention module is introduced to enhance cloud recognition accuracy and fully reflect the fine-grained attributes of cloud images. Specifically, this includes:

[0054] SK attention is a channel attention mechanism with dynamic selectivity, such as... Figure 2As shown, it has three operations: Split, Fuse, and Select, which can adaptively adjust the receptive field size to capture target objects of different scales, further improving the recognition accuracy of blocky clouds and cloud edges in ground-based cloud maps, and fully reflecting the fine-grained attributes of cloud maps. For a given feature map, two transformations are first performed to obtain two feature maps, U1 and U2, and then they are fused by pixel-wise summation to obtain U, as shown in equation (1):

[0055] U = U1 + U2 (1)

[0056] Then, channel statistics are generated by simply embedding global information using global average pooling. Furthermore, a compact feature is created to guide precise and adaptive selection. This is implemented through a simple fully connected (fc) layer, reducing dimensionality for improved efficiency.

[0057] The weight score for each channel is calculated and applied to the feature map. Finally, the feature maps are fused to obtain the final output image. Compared with the input image, the output image, after information channel refinement, incorporates more key information, enhancing the key information of the image.

[0058] In improving the U-shaped network, a parallel dilated convolution module was proposed to increase the receptive field while maintaining the feature map size, thereby further mining the fine-grained semantic information of the feature maps. Specifically, this includes:

[0059] First, the input fine-grained feature map is fed into a 3×3 convolutional block for dimensionality reduction, and then passed through a normalization and ReLU activation function layer.

[0060] Secondly, there are three parallel branches. To better extract fine-grained semantic features from the ground-based cloud image, dilated convolution is introduced, using multiple dilated convolutions stacked together. However, dilated convolution has a raster effect; when multiple dilated convolutions are stacked, some pixels are completely unutilized, resulting in a loss of information continuity and integrity, thus affecting the fine-grained segmentation effect of the ground-based cloud image. To address this issue, a hybrid dilated convolution method is designed, where the common divisor of the stacked convolution kernels is no greater than 1.

[0061] The first branch has one convolutional block with an expansion coefficient of 1; the second branch has two convolutional blocks with expansion coefficients of 1 and 3 respectively; the third branch has three convolutional blocks with expansion coefficients of 1, 3, and 5 respectively. All convolutional blocks have a kernel size of 3×3, and the last three branches are merged. Figure 3The process of the parallel dilated convolution module is demonstrated. The parallel dilated convolution module can further mine fine-grained semantic information of the feature map, and finally achieve high-precision segmentation results.

[0062] The formula for dilated convolution is:

[0063] K = (rate-1) × (k-1) + k (2)

[0064] In the formula, K is the receptive field of dilated convolution, rate is the dilated convolution rate, and k is the kernel size.

[0065] In improving the U-shaped network, CARAFE upsampling is used to replace the linear upsampling of the original decoder, avoiding the loss of fine-grained information and better recovering semantic features. Specifically, this includes:

[0066] CARAFE consists of two steps: a kernel prediction module and a content-aware reassembly module. Figure 4 As shown, the Kernel Prediction Module consists of three sub-modules: a channel compressor, a content encoder, and a kernel normalizer. It first reduces the number of channels from C to C0 using a 1×1 convolution. m Reducing the number of channels in the input feature map leads to a reduction in parameters and computational costs in subsequent steps, making CARAFE more efficient. Within the same budget, a larger kernel size can also be used for the content encoder, and reducing feature channels within an acceptable range does not harm performance. The content encoder then has a size k. encoder ×k encoder The convolution reduces the number of channels from C. m Change to C UP The purpose of this step is to generate a reconstructed kernel based on the content of the input features. Then, the kernel obtained in the previous step is processed within the region N(χ)... l ,k up Normalization is performed using the Softmax normalization function to make the convolution weight kernel 1, as shown in equation (3) to obtain the size W. l' The kernel.

[0067] W l′ =φ(N(χ) l ,k encoder (3)

[0068] For the content-aware reorganization module, features within a local region are reorganized using a function φ, as shown in equation (4). Under the action of the reorganization kernel, N(χ) l ,k upEach pixel within the region contributes differently to the upsampled pixels based on the feature content. This allows for greater focus on information about relevant points within the local region, resulting in a reconstructed feature map with richer semantic information than the original. The reconstructed kernel χ... l’ The calculation is as follows:

[0069]

[0070] The overall structure of the method of the present invention is as follows: Figure 5 As shown in the figure. This invention uses an encoder-decoder structure as the basic segmentation model and an improved U-shaped network as the backbone network to extract multi-scale features at different levels while reducing network parameters. To address the difficulty in distinguishing between different cloud types and the inability to capture small target clouds and blurred clouds, an optional channel attention module is introduced before the encoder downsampling operation to better extract fine-grained features of the ground-based cloud image. To address the low segmentation accuracy, a parallel dilated convolution module is designed to expand the receptive field without increasing computational load, fully mining the semantic information of the ground-based cloud image. To address the information loss during sampling on the original U-shaped network, a CARAFE upsampling module is introduced. After training the model, the decoder branch generates the final color prediction result. The segmentation effect of this invention is shown in the figure. Figure 6 As shown, this method achieves the best segmentation accuracy compared to other semantic segmentation methods. This method incorporates fine-grained attributes into the ground-based cloud map segmentation domain, enabling fine-grained segmentation of ground-based cloud maps while laying the foundation for subsequent photovoltaic power prediction.

[0071] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A fine-grained segmentation method for ground cloud maps based on an improved encoder-decoder structure, characterized in that, Includes the following steps: Ignoring the influence of cloud height, clouds are reclassified into five categories based on solar irradiance characteristics, cloud shape, and meteorological characteristics of different cloud types. Ground-based cloud images are labeled to obtain color-labeled images of different categories, and a fine-grained segmentation dataset of ground-based cloud images is constructed. The five cloud types are: cumulonimbus and nimbostratus, cumulus, stratus, cirrus, and clear sky. The fine-grained segmentation dataset of the ground cloud map is constructed and input into the preset deep learning model to obtain fine-grained color segmentation images; The preset deep learning model is as follows: an encoder-decoder structure is selected as the basic architecture, and an improved U-shaped network is used as the backbone network. In the process of improving the U-shaped network, an optional channel attention module is introduced to improve cloud recognition accuracy and fully reflect the fine-grained attributes of cloud images. A parallel dilated convolution module is proposed to increase the receptive field while keeping the feature map size unchanged, and to further mine the fine-grained semantic information of the feature map. A CARAFE upsampling module is introduced to avoid more information loss and better recover fine-grained semantic features. In the process of improving the U-shaped network, a parallel dilated convolution module was proposed to increase the receptive field while keeping the feature map size unchanged, thereby further mining the fine-grained semantic information of the feature map, specifically including: First, the input fine-grained feature map is fed into a 3×3 convolutional block for dimensionality reduction, and then passed through a normalization and ReLU activation function layer; Secondly, there are three parallel branches. In order to better mine the fine-grained semantic features of the ground cloud map, dilated convolution is introduced. When multiple dilated convolutions are stacked, some pixels are not utilized at all, which will lose some of the continuity and integrity of information, thus affecting the effect of fine-grained segmentation of the ground cloud map. To address this issue, a hybrid dilated convolution method is designed. When using multiple dilated convolutions, the common divisor of the stacked convolution dilation kernels is no greater than 1. The first branch has one convolutional block with an inflation factor of 1; the second branch has two convolutional blocks with inflation factors of 1 and 3 respectively; the third branch has three convolutional blocks with inflation factors of 1, 3 and 5 respectively. The kernel size of all convolutional blocks is 3×3. Finally, the three branches are merged. Through the parallel dilated convolution module, fine-grained semantic information of the feature map can be further mined, and high-precision segmentation results can be achieved. The formula for dilated convolution is: (2) In the formula, For the receptive field of void convolution, The dilated convolution rate, The kernel size; In the process of improving the U-shaped network, CARAFE upsampling is used to replace the linear upsampling of the original decoder, avoiding the loss of fine-grained information and better recovering semantic features. Specifically, this includes: CARAFE consists of two steps: a kernel prediction module and a content-aware reassembly module. The kernel prediction module comprises three sub-modules: a channel compressor, a content encoder, and a kernel normalizer. It first reduces the number of channels from C to C using a 1×1 convolution. m Reducing the number of channels in the input feature map leads to a reduction in parameters and computational costs in subsequent steps, making CARAFE more efficient. Within the same budget, a larger kernel can also be used for the content encoder, which then has a size k. encoder ×k encoder The convolution reduces the number of channels from C. m Change to C UP The purpose of this step is to generate a reconstructed kernel based on the content of the input features, and then process the N(χ) region in the kernel obtained in the previous step. l , k up Normalization is performed using the Softmax normalization function to make the convolution weight kernel 1, as shown in equation (3) to obtain the size W. l' Kernel: (3) For the content-aware reorganization module, the features in the local area are reorganized by the function φ, as shown in equation (4). Under the action of the reorganization kernel, N(χ) l , k up Each pixel within the region contributes differently to the upsampled pixels based on the feature content, resulting in a reconstructed kernel χ. l’ The calculation is as follows: (4)。 2. The ground cloud map fine-grained segmentation method based on an improved encoder-decoder structure according to claim 1, characterized in that, In improving the U-shaped network, an optional channel attention module is introduced to enhance cloud recognition accuracy and reflect the fine-grained attributes of the cloud image. Specifically, this includes: SK attention is a channel attention mechanism with dynamic selectivity. It has three operations: Split, Fuse, and Select. It can adaptively adjust the receptive field size, capture target objects of different scales, improve the recognition accuracy of blocky clouds and cloud edges in ground-based cloud maps, and reflect the fine-grained attributes of cloud maps. For a given feature map, it first performs two transformations to obtain two feature maps, U1 and U2, and then fuses them by pixel-wise summation to obtain U, as shown in equation (1): (1) Then, global information is embedded by using global average pooling to generate channel statistics. Furthermore, a feature is created to guide precise and adaptive selection, which is implemented through a simple fully connected FC layer. To improve efficiency, the dimensionality is reduced. The weight score corresponding to each channel is calculated and applied to the feature map. Finally, the feature map is fused to obtain the final output image. The output image is compared with the input image. After the information channels are refined, key information is fused and the key information of the image is enhanced.