Fine-grained foundation cloud segmentation identification method based on lightweight neural network
By constructing a ground-based cloud segmentation model based on a lightweight neural network, and combining attention enhancement and edge perception modules, the problems of boundary blurring and detail loss in ground-based cloud image segmentation are solved, achieving high-precision cloud body recognition and adaptive enhancement, which is suitable for edge computing devices and mobile observation platforms.
Patent Information
- Application Number
- CN202511297721.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-11
AI Technical Summary
Existing ground-based cloud image segmentation methods are insufficient in fine-grained structure recognition, making it difficult to accurately depict the contour shape, edge transitions, and local texture differences of cloud bodies. Especially in cloud scenarios where multiple scales and discontinuous structures coexist, the models are prone to overly smoothed boundaries and loss of details, failing to meet the needs of refined meteorological observation and automatic cloud cover recognition.
A fine-grained ground-based cloud segmentation and recognition method based on lightweight neural networks is adopted. The ground-based cloud segmentation model is constructed using the DeepLabV3+ framework, and the backbone feature is extracted by combining the MobileNetV2 network. An attention enhancement module and an edge perception fusion module are added. Strong light interference is handled by solar suppression channel attention and multi-scale spatial attention. The solar region is estimated by combining the image acquisition time and device location. An edge perception fusion module is introduced to enhance the response of edge regions. The model training is optimized using the total loss function.
It significantly improves the ability of ground-based cloud segmentation models to model cloud edge details and multi-scale structures, enhances the segmentation and recognition accuracy and adaptability in ground-based cloud image scenarios, is suitable for edge computing devices and mobile observation platforms, and strengthens the robustness and accuracy of boundary blurring and texture details.
Smart Images

Figure CN120877124A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of meteorological technology, and in particular to a fine-grained ground-based cloud segmentation and recognition method based on a lightweight neural network. Background Technology
[0002] Ground-based cloud image segmentation is a crucial data source for meteorological observations, and its results directly impact the accuracy of subsequent tasks such as cloud cover estimation and cloud shape identification. Compared to natural image segmentation, ground-based cloud images are characterized by blurred boundaries, complex texture variations, and variable lighting conditions, placing higher demands on the structural modeling capabilities of semantic segmentation models. While current mainstream segmentation methods possess strong semantic modeling capabilities, they still have significant shortcomings in fine-grained structure recognition, struggling to accurately depict the contours, edge transitions, and local texture differences of cloud bodies. Especially in cloud field scenarios with multiple scales and discontinuous structures, models are prone to issues such as overly smoothed boundaries and loss of detail, failing to meet the requirements of refined meteorological observations and automatic cloud cover identification. Summary of the Invention
[0003] To address the aforementioned problems and technical requirements, this application proposes a fine-grained ground-based cloud segmentation and recognition method based on a lightweight neural network. The technical solution of this application is as follows: A fine-grained ground-based cloud segmentation and recognition method based on a lightweight neural network, comprising: Collect visible cloud images of the entire sky, preprocess the images, and input them into the ground-based cloud segmentation model to obtain the ground-based cloud segmentation and recognition results. The ground-based cloud segmentation model is built on DeepLabV3+ and includes a backbone feature extraction module, an attention enhancement module, a context modeling module, and a decoder module. The backbone feature extraction module uses the MobileNetV2 network. The backbone feature extraction module extracts features from the input all-sky visible light cloud image and outputs shallow feature maps and deep feature maps. The shallow feature map is input into the attention enhancement module, which performs dual attention guidance of strong light suppression and multi-scale structure on the shallow feature map before inputting it into the decoder module. The deep feature map is input into the context modeling module and then into the decoder module. The decoder module uses a 1*1 convolution to adjust the number of channels in the input shallow feature map before inputting it into the edge perception fusion module. The edge perception fusion module performs edge enhancement perception on the input shallow feature map to obtain a joint feature map. The joint feature map is then concatenated and fused with the upsampled deep feature map, and then passed through a convolutional layer and an upsampling layer in sequence to obtain the ground-based cloud segmentation and recognition result.
[0004] Its further technical solution is that the attention enhancement module includes solar suppression channel attention and multi-scale spatial attention; The solar suppression channel attention is based on the solar projection region in the all-sky visible light cloud image. After performing strong light suppression attention guidance on the input shallow feature map, it is multiplied element-wise with the input shallow feature map to obtain the intermediate feature map. After multi-scale spatial attention is applied to the intermediate feature map, it is multiplied element-wise with the intermediate feature map and then output to the decoder module.
[0005] Further technical solutions include: the fine-grained ground-based cloud segmentation and identification method also includes: A mask is applied to the solar projection area in the visible light cloud image of the entire sky to obtain a solar region mask image; The solar region mask map is embedded into the solar suppression channel attention of the ground-based cloud segmentation model. During the attention calculation of the shallow feature map, the solar suppression channel attention uses the solar region mask map for mask modulation to achieve strong light suppression attention guidance.
[0006] A further technical solution involves determining the solar projection area in the visible light cloud image of the entire sky, including: Based on the image acquisition time and geographical location of the all-sky visible light cloud image, the real-time position of the sun is determined using a solar position prediction model, and the solar projection area in the all-sky visible light cloud image is mapped.
[0007] Its further technical solution is that the edge-aware fusion module includes parallel edge branches and context branches; The edge branch extracts edge information from the input shallow feature map to obtain an edge-aware feature map; The context branch extracts multi-scale local contextual information from the input shallow feature map to obtain a context-enhanced feature map; The edge-aware feature map and the context-enhanced feature map are fused along the channel dimension and then passed through a convolutional layer to generate a joint feature map with edge-enhanced awareness capabilities.
[0008] The further technical solution is that the edge branch in the edge-aware fusion module includes an edge extraction layer, a normalization layer and an activation function cascaded in sequence. The edge extraction layer performs adaptive edge information extraction based on the image perspective distortion perception of the shallow feature map to obtain an edge response map. The edge response map is then processed by the normalization layer and the activation function to obtain an edge-aware feature map.
[0009] A further technical solution involves the edge extraction layer calculating the value of each pixel in the input shallow feature map. Normalized radial distance between the image center point and the image center point And using normalized radial distance For standard level Sobel kernels and standard vertical Sobel core The corrected horizontal Sobel kernels are obtained by modifying them separately. and the corrected vertical Sobel kernel :
[0010] Using the modified horizontal Sobel kernel and the corrected vertical Sobel kernel For pixels Edge information extraction is performed; among which, It is the adjustment coefficient.
[0011] The further technical solution is that the context branch includes multiple convolutional kernels of different sizes arranged in parallel. The multiple convolutional kernels extract multi-scale local contextual information from the input shallow feature map in parallel. The local contextual information output by the multiple convolutional kernels is convolved and fused in the channel dimension to form a context-enhanced feature map.
[0012] Further technical solutions include: the fine-grained ground-based cloud segmentation and identification method also includes: The total loss function used when training the ground-based cloud segmentation model. Among them, the pixel-level cross-entropy loss term The edge loss term is used to measure the semantic segmentation accuracy of the ground-based cloud segmentation model. The structural loss term is used to measure the edge detection quality of the ground-based cloud segmentation model. Used to measure the quality of local structure detection in ground-based cloud segmentation models. and This is the loss weight hyperparameter.
[0013] Its further technical solution is, edge loss term , This is a prediction map of the ground-based cloud segmentation model based on the training samples. The edge image obtained after edge extraction operator. These are the true label images of the training samples. The edge image obtained after edge extraction operator. It is a binary cross-entropy loss function; Structural loss term , It is a prediction graph of the training samples. and its actual label image The structural similarity index.
[0014] The beneficial technical effects of this application are: This application discloses a fine-grained ground-based cloud segmentation and recognition method based on a lightweight neural network. This method utilizes a ground-based cloud segmentation model to perform pixel-level ground-based cloud segmentation and recognition on all-sky visible cloud images. The ground-based cloud segmentation model used is based on the DeepLabV3+ framework and is specifically optimized to address the problems of edge blurring, insufficient detail response, and high computational resource requirements of traditional semantic segmentation models when processing ground-based cloud images. While maintaining its advantages of multi-scale context-aware structure, the backbone feature extraction module is replaced with a MobileNetV2 network, and an attention enhancement module is added to provide dual attention guidance for shallow features through strong light suppression and multi-scale structure. An edge-aware fusion module is added to the decoder to further enhance edge perception of shallow feature maps. Thus, while maintaining the lightweight overall network architecture, the modeling ability of the ground-based cloud segmentation model for cloud edge details and multi-scale structure is significantly improved. It also has high ground-based cloud segmentation and recognition accuracy in complex ground-based cloud image scenarios and has strong adaptability and deployment efficiency, which can flexibly adapt to the resource conditions of different platforms.
[0015] The attention enhancement module integrates solar suppression channel attention and multi-scale spatial attention to adapt to the unique strong light interference and multi-scale structural features of ground-based cloud images. It estimates the solar region by combining image acquisition time and device location, and uses solar suppression channel attention for explicit suppression, effectively mitigating misjudgments caused by bright solar interference areas in ground-based cloud images. Simultaneously, multi-scale aggregation in the spatial dimension enhances the model's ability to identify multi-scale cloud structures. The feature maps guided by dual attention are used for downstream decoding and fusion, effectively strengthening the model's ability to model boundary ambiguity and texture details, and improving the accuracy and fine-grained consistency of semantic segmentation of ground-based cloud images.
[0016] To address the edge distortion issues caused by fisheye lens imaging in ground-based cloud images and the inherent boundary blurring of cloud structures, an edge-aware fusion module added to the decoder employs a distortion-aware Sobel edge adaptive extraction mechanism. This mechanism moderately enhances the response of edge regions, improves the ability to model the boundaries of stretched and distorted areas, and fuses with contextual information to generate a joint feature map with enhanced edge perception capabilities. This feature map will participate in the resolution restoration process in subsequent decoding stages, improving the model's robustness and accuracy in boundary blurring and detail recognition.
[0017] In addition to optimizing the network structure, a total loss function containing pixel-level cross-entropy loss, edge loss, and structure loss was introduced during the model training phase. By jointly optimizing the above three metrics, the total loss function ensures the global semantic segmentation accuracy while focusing on optimizing the model's prediction performance in the boundary region. This effectively improves the quality of the segmentation map in terms of edge coherence, local texture alignment, and detail preservation, making it particularly suitable for semantic segmentation tasks of ground-based cloud images with characteristics such as blurred edges and complex structures. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a fine-grained ground-based cloud segmentation and identification method according to an embodiment of this application. Detailed Implementation
[0019] The specific embodiments of this application will be further described below with reference to the accompanying drawings.
[0020] This application discloses a fine-grained ground-based cloud segmentation and recognition method based on a lightweight neural network, which includes the following steps: Step S1 involves acquiring full-sky visible cloud images using an observation platform. This platform is deployed at a fixed ground observation point and can continuously acquire full-sky visible cloud images with high temporal resolution. Full-sky visible cloud images typically have a 360° field of view coverage, recording cloud field changes under different times and meteorological conditions. In practical applications, the observation platform used is usually a full-sky imager.
[0021] Step S2: Perform image preprocessing on the visible light cloud image of the entire sky.
[0022] Image preprocessing includes scaling the all-sky visible cloud image to an image size supported by the ground-based cloud segmentation model, and then performing brightness standardization and normalization on the scaled all-sky visible cloud image to mitigate light interference caused by different time periods or weather conditions.
[0023] Step S3: Input the preprocessed all-sky visible light cloud image into the ground-based cloud segmentation model to obtain the ground-based cloud segmentation and recognition results.
[0024] The key to improving the performance of ground-based cloud segmentation and recognition in this application lies in the optimized design of a ground-based cloud segmentation model for line-of-sight semantic segmentation. This model is built based on DeepLabV3+. Please refer to [link / reference needed]. Figure 1The network structure diagram shown illustrates that this ground-based cloud segmentation model includes a backbone feature extraction module, an attention enhancement module, a context modeling module, and a decoder module. Specifically, the backbone feature extraction module extracts features from the input all-sky visible cloud image and outputs shallow and deep feature maps. The shallow feature map is then passed through the attention enhancement module and input into the decoder module, while the deep feature map is input into the context modeling module and then into the decoder module. The decoder module uses a 1x1 convolution to adjust the channel count of the input shallow feature map before inputting it into the edge-aware fusion module. The edge-aware fusion module enhances the edge perception of the input shallow feature map to obtain a joint feature map. This joint feature map is then concatenated and fused with the upsampled deep feature map, and subsequently passed through convolutional layers and upsampling layers to gradually restore the spatial resolution, resulting in a pixel-level ground-based cloud segmentation and recognition result consistent with the size of the all-sky visible cloud image. The context modeling module directly adopts the module structure in DeepLabV3+, such as the common ASPP module, utilizing multi-scale dilated convolution and image pooling to enhance context modeling capabilities. This application will not elaborate on this part further.
[0025] While maintaining the advantages of DeepLabV3+'s multi-scale context-aware structure, this ground-based cloud segmentation model is adapted and optimized according to the characteristics and requirements of ground-based cloud segmentation, mainly including the following aspects: (1) The backbone feature extraction module in the traditional DeepLabV3+ uses the Xception backbone network. The Xception backbone network has a large number of model parameters and high computational complexity. However, this fine-grained ground cloud segmentation and recognition method is often required for use on edge computing devices or mobile observation platforms in practical applications. Therefore, in order to ensure good engineering adaptability and promotion value, the backbone feature extraction module in this ground cloud segmentation model adopts the MobileNetV2 network to replace the original Xception backbone network. The MobileNetV2 network adopts a depth-separable convolutional design, which can retain the multi-scale semantic information modeling capability while greatly reducing model parameters and computational load. This makes the overall segmentation network more suitable for embedded devices or edge computing scenarios, improves operating efficiency and deployment flexibility, and is suitable for edge computing devices or mobile observation platforms.
[0026] (2) In the traditional DeepLabV3+, the shallow feature map is directly input into the decoder module. In this application, an attention enhancement module is added to perform strong light suppression and multi-scale structure dual attention guidance on the shallow feature map output by the backbone feature extraction module before inputting it into the decoder module.
[0027] In ground-based cloud segmentation and recognition scenarios using all-sky visible light cloud images, bright areas of the sun in the images are easily misidentified as clouds, and the scale of clouds varies considerably. Therefore, adding an attention enhancement module to perform dual attention guidance on shallow feature maps can effectively adapt to the strong light interference and multi-scale structural features unique to all-sky visible light cloud images. The shallow feature maps guided by dual attention are used for downstream decoding and fusion, effectively enhancing the ground-based cloud segmentation model's ability to model boundary ambiguity and texture details, and improving the accuracy and fine-grained consistency of semantic segmentation of ground-based cloud images.
[0028] In one embodiment, the attention enhancement module includes solar suppression channel attention and multi-scale spatial attention. Solar suppression channel attention performs strong light suppression attention guidance on the input shallow feature map based on the solar projection region in the all-sky visible light cloud image, and then performs element-wise multiplication with the input shallow feature map to obtain an intermediate feature map. Multi-scale spatial attention performs multi-scale structure attention guidance on the intermediate feature map, and then performs element-wise multiplication with the intermediate feature map before outputting it to the decoder module.
[0029] To address the common problem of bright solar areas in all-sky visible light cloud images being easily misidentified as clouds, the first step is to determine the solar projection area within the all-sky visible light cloud image. For determining the real-time position of the sun, since the sun's position and movement are predictable, a solar position prediction model can be used. This model records the sun's position at different geographical locations at different times. After determining the image acquisition time and geographical location of the all-sky visible light cloud image (i.e., the location of the observation platform), inputting these two locations into the solar position prediction model will determine the sun's real-time position, which can then be mapped to the solar projection area in the all-sky visible light cloud image.
[0030] Then, a mask is applied to the solar projection region in the visible cloud image of the entire sky to obtain a solar region mask map. This solar region mask map is embedded into the solar suppression channel attention of the ground-based cloud segmentation model. During the attention calculation of the shallow feature map, the solar suppression channel attention utilizes this solar region mask map for mask modulation to achieve strong light suppression attention guidance. Mask modulation during attention calculation can force the attention weight of the mask location (i.e., the solar projection region) to be set to a very small value, thereby preventing the model from noticing information at these locations. This effectively suppresses false activation of the solar region, thus improving the model's robustness in segmenting the region near the sun. Specific mask modulation methods can be found in existing attention calculation implementations, and will not be elaborated here.
[0031] Multi-scale spatial attention enhances the ability to identify key regions such as cloud structures at different scales by introducing pooling operations with different receptive field sizes. In one example, 3×3 and 5×5 receptive fields are used to accommodate cloud structures at scales such as cloud edges, isolated small clouds, and cirrus clouds.
[0032] (3) In traditional DeepLabV3+, after the shallow feature map is input into the decoder module, it is directly concatenated and fused with the upsampled deep feature map after undergoing a 1*1 convolution to adjust the number of channels. In this application, an edge-aware fusion module is added to the decoder module. The shallow feature map input into the decoder module is first subjected to a 1*1 convolution to adjust the number of channels, and then input into the edge-aware fusion module. The edge-aware fusion module performs edge enhancement perception on the input shallow feature map to obtain a joint feature map, and then concatenates and fuses it with the upsampled deep feature map.
[0033] The edge-aware fusion module can further enhance the ability of the ground-based cloud segmentation model to model blurred edge regions and local details. The resulting joint feature map will participate in the resolution restoration process in the subsequent decoding stage, thereby improving the robustness and accuracy of the ground-based cloud segmentation model in terms of boundary blurring and detail identification.
[0034] In one embodiment, the edge-aware fusion module includes parallel edge branches and context branches: (a) Edge branches extract edge information from the input shallow feature map to obtain an edge-aware feature map. For example... Figure 1 As shown, the edge branch includes an edge extraction layer, a normalization layer, and an activation function cascaded in sequence. The edge extraction layer performs adaptive edge information extraction based on the image perspective distortion perception of the shallow feature map to obtain an edge response map. After the edge response map is processed by the normalization layer and the activation function, an edge-aware feature map is obtained.
[0035] To acquire full-sky visible cloud images, observation platforms typically use fisheye lenses. However, while fisheye lenses expand the field of view, they also introduce edge distortion. Therefore, to adapt to the aforementioned imaging characteristics of full-sky visible cloud images, the edge extraction layer employs a distortion-aware Sobel edge adaptive extraction mechanism, implemented as follows: Calculate each pixel in the input shallow feature map Normalized radial distance between the image center point and the image center point And using normalized radial distance For standard level Sobel kernels and standard vertical Sobel core The corrected horizontal Sobel kernels are obtained by modifying them separately. and the corrected vertical Sobel kernel :
[0036] Then, the modified horizontal Sobel kernel is used. and the corrected vertical Sobel kernel For pixels Edge information extraction is performed. It is the adjustment coefficient.
[0037] This distortion-aware Sobel edge adaptive extraction mechanism can moderately enhance the response of edge regions and improve the ability to model the boundaries of stretched and distorted regions.
[0038] (b) The context branch extracts multi-scale local context information from the input shallow feature map to obtain a context-enhanced feature map. The context branch includes multiple convolutional kernels of different sizes arranged in parallel. Each kernel extracts multi-scale local context information from the input shallow feature map in parallel. The local context information output by each kernel is concatenated along the channel dimension and then fused by convolution to form the context-enhanced feature map. In one example, the context branch includes three convolutional kernels arranged in parallel, with sizes of 3×3, 5×5, and 7×7, to enhance the modeling capability of cloud structures at different scales.
[0039] Edge branches are used for edge information extraction, while context branches are used for context information completion. Finally, the edge-aware feature map and the context-enhanced feature map are fused along the channel dimension and then passed through a convolutional layer to generate a joint feature map with edge-enhanced awareness, thereby enhancing the model's ability to respond to structural details.
[0040] The ground-based cloud segmentation model also includes a model training process before use, including: 1. Construct a sample dataset.
[0041] Full-sky visible cloud images were acquired through observation platforms such as all-sky imagers, and sample images were obtained through image preprocessing. To improve the training and generalization ability of the model, data augmentation was performed through processes such as horizontal flipping and brightness perturbation. Then, pixel-level segmentation labels were added to the sample images to obtain the ground truth label images of the sample images. The resulting sample dataset was constructed, and the sample dataset was divided into training and test sets.
[0042] 2. Set up as follows Figure 1The network structure shown inputs sample images from the training set into the ground-based cloud segmentation model to obtain predicted images. The total loss function is calculated based on the predicted and ground truth labels of the sample images, and the network parameters of the ground-based cloud segmentation model are optimized according to the total loss function. In one example, the Adam optimizer is used for gradient updates during training, with 300 iterations and 16 images input to the network model each time. The initial learning rate is set to 0.001 to accelerate convergence and avoid overfitting.
[0043] In addition to optimizing the network structure of the ground-based cloud segmentation model compared to DeepLabV3+ in the three aspects mentioned above, this embodiment also optimizes the total loss function used by the ground-based cloud segmentation model during training to further improve the segmentation accuracy in the cloud boundary region: The total loss function used when training the ground-based cloud segmentation model. In addition to including the regular pixel-level cross-entropy loss In addition, an edge loss term was added. and structural loss term Total loss function The calculation formula is:
[0044] Pixel-level cross-entropy loss term The semantic segmentation accuracy of the ground-based cloud segmentation model is measured by calculating the difference between the predicted label of each pixel in the ground truth label map and the predicted label of that pixel in the prediction map.
[0045] Marginal loss term This is used to measure the edge detection quality of a ground-based cloud segmentation model. The edge loss term can be expressed as: , This is a prediction map of the ground-based cloud segmentation model based on the training samples. The edge image obtained after edge extraction operator. These are the true label images of the training samples. The edge image obtained after edge extraction operator. It is a binary cross-entropy loss function. The edge extraction operator used is either the Canny operator or the Sobel operator.
[0046] Structural loss term This is used to measure the quality of local structure detection in a ground-based cloud segmentation model. The structure loss term can be expressed as: , It is a prediction graph of the training samples. and its actual label image The structural similarity index.
[0047] and This is the loss weight hyperparameter, used to control the influence strength of marginal and structural terms.
[0048] By jointly optimizing the above three indicators, the total loss function While ensuring the accuracy of global semantic segmentation, the model focuses on optimizing the prediction performance in the boundary region, effectively improving the quality of the segmentation map in terms of edge coherence, local texture alignment and detail preservation. It is particularly suitable for semantic segmentation tasks of ground-based cloud images with characteristics such as blurred edges and complex structures.
[0049] 3. Test the trained ground-based cloud segmentation model using the test set to evaluate its accuracy using metrics such as Accuracy, Precision, and Recall.
[0050] The above descriptions are merely preferred embodiments of this application, and this application is not limited to the above embodiments. It is understood that other improvements and variations that can be directly derived or conceived by those skilled in the art without departing from the spirit and concept of this application should be considered to be included within the protection scope of this application.
Claims
1. A fine-grained ground-based cloud segmentation and recognition method based on a lightweight neural network, characterized in that, The fine-grained ground-based cloud segmentation and identification method includes: A full-sky visible light cloud image is acquired, and after image preprocessing, it is input into a ground-based cloud segmentation model to obtain the ground-based cloud segmentation and recognition result. The ground-based cloud segmentation model is built on DeepLabV3+ and includes a backbone feature extraction module, an attention enhancement module, a context modeling module, and a decoder module. The backbone feature extraction module uses the MobileNetV2 network. The backbone feature extraction module extracts features from the input all-sky visible light cloud image and outputs shallow feature maps and deep feature maps. The shallow feature map is input to the attention enhancement module, which performs dual attention guidance of strong light suppression and multi-scale structure on the shallow feature map before inputting it to the decoder module. The deep feature map is input to the context modeling module and then to the decoder module. The decoder module uses a 1*1 convolution to adjust the number of channels in the input shallow feature map before inputting it into the edge perception fusion module. The edge perception fusion module performs edge enhancement perception on the input shallow feature map to obtain a joint feature map. The joint feature map is then concatenated and fused with the upsampled deep feature map, and then passed through a convolutional layer and an upsampling layer in sequence to obtain the ground-based cloud segmentation and recognition result.
2. The fine-grained ground-based cloud segmentation and identification method according to claim 1, characterized in that, The attention enhancement module includes solar suppression channel attention and multi-scale spatial attention; The solar suppression channel attention is based on the solar projection region in the all-sky visible light cloud image. After performing strong light suppression attention guidance on the input shallow feature map, it is multiplied element-wise with the input shallow feature map to obtain the intermediate feature map. After multi-scale spatial attention is applied to the intermediate feature map, it is multiplied element-wise with the intermediate feature map and then output to the decoder module.
3. The fine-grained ground-based cloud segmentation and identification method according to claim 2, characterized in that, The fine-grained ground-based cloud segmentation and identification method further includes: A mask is applied to the solar projection area in the visible light cloud image of the entire sky to obtain a solar region mask image; The solar region mask map is embedded into the solar suppression channel attention of the ground-based cloud segmentation model. During the attention calculation of the shallow feature map, the solar suppression channel attention uses the solar region mask map for mask modulation to achieve strong light suppression attention guidance.
4. The fine-grained ground-based cloud segmentation and identification method according to claim 2, characterized in that, Determining the solar projection area in a full-sky visible cloud image includes: Based on the image acquisition time and geographical location of the all-sky visible light cloud image, the real-time position of the sun is determined using a solar position prediction model, and the solar projection area in the all-sky visible light cloud image is mapped.
5. The fine-grained ground-based cloud segmentation and identification method according to claim 1, characterized in that, The edge-aware fusion module includes parallel edge branches and context branches; Edge branches extract edge information from the input shallow feature map to obtain an edge-aware feature map; The context branch extracts multi-scale local contextual information from the input shallow feature map to obtain a context-enhanced feature map; The edge-aware feature map and the context-enhanced feature map are fused along the channel dimension and then passed through a convolutional layer to generate a joint feature map with edge-enhanced awareness capabilities.
6. The fine-grained ground-based cloud segmentation and identification method according to claim 5, characterized in that, The edge branch in the edge-aware fusion module includes a cascaded edge extraction layer, a normalization layer, and an activation function. The edge extraction layer performs adaptive edge information extraction based on the perception of image viewpoint distortion in the shallow feature map to obtain an edge response map. The edge response map is then processed by the normalization layer and the activation function to obtain an edge-aware feature map.
7. The fine-grained ground-based cloud segmentation and identification method according to claim 6, characterized in that, The edge extraction layer calculates the value of each pixel in the input shallow feature map. Normalized radial distance between the image center point and the image center point And using normalized radial distance For standard level Sobel kernels and standard vertical Sobel core The corrected horizontal Sobel kernels are obtained by modifying them separately. and the corrected vertical Sobel kernel : Using the modified horizontal Sobel kernel and the corrected vertical Sobel kernel For pixels Edge information extraction is performed; among which, It is the adjustment coefficient.
8. The fine-grained ground-based cloud segmentation and identification method according to claim 5, characterized in that, The context branch consists of multiple convolutional kernels of different sizes arranged in parallel. Each kernel extracts multi-scale local contextual information from the input shallow feature map in parallel. The local contextual information output by each kernel is concatenated along the channel dimension and then fused by convolution to form a context-enhanced feature map.
9. The fine-grained ground-based cloud segmentation and identification method according to claim 1, characterized in that, The fine-grained ground-based cloud segmentation and identification method further includes: The total loss function used when training the ground-based cloud segmentation model. Among them, the pixel-level cross-entropy loss term The edge loss term is used to measure the semantic segmentation accuracy of the ground-based cloud segmentation model. The structural loss term is used to measure the edge detection quality of the ground-based cloud segmentation model. Used to measure the quality of local structure detection in ground-based cloud segmentation models. and This is the loss weight hyperparameter.
10. The fine-grained ground-based cloud segmentation and identification method according to claim 9, characterized in that, Marginal loss term , This is a prediction map of the ground-based cloud segmentation model based on the training samples. The edge image obtained after edge extraction operator. These are the true label images of the training samples. The edge image obtained after edge extraction operator. It is a binary cross-entropy loss function; Structural loss term , It is a prediction graph of the training samples. and its actual label image The structural similarity index.
Citation Information
Patent Citations
Ground-based cloud picture fine-grained segmentation method based on improved encoder-decoder structure
CN117911693A
All-sky nephogram detection and classification method based on attention pyramid network
CN119360120A
Three-dimensional lidar point cloud semantic segmentation method and apparatus based on deep learning
WO2024130776A1