Satellite cloud detection method and device based on transfer learning and ground-based physical enhancement

CN122435476BActive Publication Date: 2026-09-15齐鲁空天信息研究院
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610556303.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-24
Publication Date
2026-09-15
Estimated Expiration
2046-04-24

AI Technical Summary

Technical Problem

首先,现有方法多直接采用静止卫星官方二级产品作为标签进行全监督训练,这导致模型继承了静止卫星传感器低空间分辨率及传统物理阈值算法的固有误差

Benefits of technology

[0017] This invention proposes a cascaded cloud mask inversion architecture based on cross-sensor transfer learning and ground-based physical constraints, and adopts a hierarchical and progressive accuracy improvement strategy: the first stage constructs a global basic inversion model based on multi-scale deep neural networks to output a large field of view initial probability field; the second stage constructs a lightweight nonlinear residual prediction model with multi-source auxiliary environmental parameters to implement local physical correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435476B_ABST
    Figure CN122435476B_ABST
Patent Text Reader

Abstract

The application discloses a satellite cloud detection method and device based on transfer learning and ground physical enhancement, and belongs to the technical field of satellite parameter inversion and deep learning image segmentation. Firstly, the data of stationary satellites, polar orbit satellites, spaceborne laser radars and ground stations are geometrically strictly registered and parallax corrected to construct a precision cascaded dataset; then a deep neural network is constructed, after pre-training by using large sample coarse labels, the encoder is frozen, the decoder is fine-tuned by using high-precision polar orbit satellite labels, and a global initial cloud mask probability field is output; finally, the track truth value of the spaceborne laser radar is taken as a residual error learning target, a light-weight regression model is trained in combination with physical constraints such as ground atmospheric precipitable water, the rectification ability of sparse points is extrapolated to the whole domain, and the final cloud mask is generated through residual error superposition and thresholding. The application effectively improves the continuity of the cloud boundary and the precision of the broken cloud discrimination, reduces the false alarm rate of complex underlying surfaces, and realizes real-time driving correction of sparse high-value data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of satellite parameter inversion and deep learning image segmentation technology, and particularly relates to a satellite cloud detection method and device based on transfer learning and ground-based physical enhancement. Background Technology

[0002] Clouds play a crucial role in weather monitoring and climate change detection, and cloud masking is a primary step in retrieving all cloud parameters (such as cloud phase, cloud optical thickness, and cloud top height). Existing cloud detection technologies can be broadly categorized into two types: traditional methods based on physical mechanisms and data-driven machine learning / deep learning methods. Traditional physical methods primarily rely on the spectral differences, texture features, and geometric morphology of clouds in the visible to thermal infrared bands, utilizing techniques such as channel thresholds, spectral indices, and multi-criteria joint decision-making to achieve cloud detection. These methods offer advantages such as strong interpretability, fast computation speed, and mature operational application, and are therefore widely used in the operational production of mainstream satellite products such as MODIS and LandSat. However, due to interference from the surface environment (such as high-albedo ice and snow, deserts) and the complex and variable atmospheric conditions, fixed threshold rules often fail to adequately account for cloud scenarios across different time phases and regions, leading to missed detections or false positives. In contrast, machine learning or deep learning methods utilize high-quality satellite-annotated datasets to construct machine learning models or neural network models to achieve cloud / clear-sky pixel classification. While traditional machine learning models have addressed the difficulty of threshold selection to some extent through nonlinear mapping, they still heavily rely on manually designed features and ignore the spatial texture and contextual information of clouds, thus limiting their generalization ability. In contrast, deep learning models, through end-to-end hierarchical feature extraction, can automatically capture the deep semantic information and contextual features of clouds, achieving significantly higher detection accuracy than traditional methods in complex backgrounds.

[0003] While deep learning has demonstrated powerful nonlinear feature extraction capabilities in cloud parameter inversion, it still faces dual constraints in practical applications: data quality and physical mechanisms. First, existing methods often directly use official geostationary satellite Level 2 data as labels for fully supervised training. This results in models inheriting the inherent errors of low spatial resolution from geostationary satellite sensors and traditional physical thresholding algorithms. Although some studies have attempted to fine-tune using high-precision polar-orbiting satellite data (such as MODIS), the lack of precise spatiotemporal registration and parallax correction leads to pixel-level alignment errors between heterogeneous data sources. This causes models to exhibit smoothing of segmentation boundaries and loss of high-frequency textures when processing cloud edges and fragmented cloud regions. Second, single passive optical remote sensing methods have fundamental flaws. Existing models primarily rely on end-to-end mapping of spectral and texture features, lacking atmospheric thermodynamic constraints independent of the optical imaging link. This makes it difficult to achieve "cloud-ground" decoupling against complex underlying surface backgrounds, resulting in a high false alarm rate in specific surface areas. Finally, the high-value information from multi-source heterogeneous data has not yet been effectively integrated and utilized. Although existing research involves active detection data from spaceborne lidar (such as CALIPSO), it is usually only used for post-result accuracy verification, lacking an effective multi-modal fusion mechanism. Existing inversion architectures struggle to transform the high spatial resolution advantage of MODIS, the true vertical structure of CALIPSO, and the continuous physical constraints of ground-based observations into real-time correction and joint driving capabilities for inversion models.

[0004] In view of the shortcomings of the prior art, the technical problem to be solved by the present invention is how to comprehensively utilize multi-source heterogeneous observation data, design a data learning system with strict geometric registration and cascading accuracy, and provide a cloud mask inversion method based on cross-sensor transfer learning and ground-based physical enhancement for the target observation area by means of a deep learning architecture. Summary of the Invention

[0005] To address the aforementioned problems, this invention proposes a satellite cloud detection method and device based on transfer learning and ground-based physical enhancement. It employs a two-stage cloud mask inversion technology architecture of "global basic inversion + regional residual compensation." The core idea of ​​this architecture is: first, to extract global general cloud features using deep networks based on geostationary satellite (e.g., FY-4B) observation data and secondary inversion products; second, to perform transfer fine-tuning of the basic model based on high-precision polar-orbiting satellite observation products (e.g., MODIS products) with strict spatiotemporal matching; and finally, to learn the nonlinear error mapping relationship between "primary basic inversion probability - ground-based physical constraint parameters - onboard active detection reference benchmark" using a lightweight regional enhancement model, thereby achieving an accuracy improvement from a global "surface" to a local "point" and back to a "surface." The specific technical solution is as follows:

[0006] A satellite cloud detection method based on transfer learning and ground-based physical augmentation includes the following steps:

[0007] First, geometrically rigorous registration and spatiotemporal alignment are performed on multi-source heterogeneous observation data from geostationary satellites, polar-orbiting satellites, spaceborne lidar, and ground-based stations. Parallax correction is used to eliminate projection deviations caused by cloud height and observation angle, and a large-sample basic pre-training set and a small-sample high-precision fine-tuning set are constructed.

[0008] Then, a deep neural network consisting of an encoder and a decoder is constructed. Fully supervised pre-training is performed using the large sample basic pre-training set. The encoder parameters are then frozen and the decoder is transferred and fine-tuned using the small sample high-precision fine-tuning set to obtain the basic model and output the global initial cloud mask probability field.

[0009] Finally, using the sparse trajectory cloud detection ground truth provided by the spaceborne lidar as the residual learning objective, the predicted probability of the basic model on the corresponding trajectory is extracted and the residual is calculated. With the ground-based atmospheric precipitable water, the predicted probability of the basic model and the auxiliary environmental features as inputs, a lightweight regression model is trained to learn the nonlinear mapping between the residual and the multi-source features. The regression model is applied to the entire domain, and the residual correction field is predicted pixel by pixel and superimposed on the initial probability field. After thresholding, the final cloud mask is generated.

[0010] A satellite cloud detection device based on transfer learning and ground-based physical augmentation includes the following modules:

[0011] The multi-source air-ground data spatiotemporal matching module performs geometrically rigorous registration and spatiotemporal alignment of multi-source heterogeneous observation data from geostationary satellites, polar-orbiting satellites, spaceborne lidar, and ground-based stations. It eliminates projection deviations caused by cloud height and observation angle through parallax correction and constructs a large-sample basic pre-training set and a small-sample high-precision fine-tuning set.

[0012] The cloud parameter inversion module based on transfer learning constructs a deep neural network consisting of an encoder and a decoder. It performs fully supervised pre-training using the large sample basic pre-training set, freezes the encoder parameters, and performs transfer fine-tuning of the decoder using the small sample high-precision fine-tuning set to obtain the basic model and output the global initial cloud mask probability field.

[0013] The physical constraint-based regional residual enhancement module uses the sparse trajectory cloud detection ground truth provided by the spaceborne lidar as the residual learning target, extracts the prediction probability of the basic model on the corresponding trajectory and calculates the residual, and uses the ground-based atmospheric precipitable water, the prediction probability of the basic model and auxiliary environmental features as inputs to train a lightweight regression model to learn the nonlinear mapping between the residual and multi-source features. The regression model is applied to the whole domain, predicts the residual correction field pixel by pixel and superimposes it on the initial probability field, and generates the final cloud mask after thresholding.

[0014] An electronic device includes: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method.

[0015] A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to implement the method described thereon.

[0016] The present invention has the following beneficial effects:

[0017] This invention proposes a cascaded cloud mask inversion architecture based on cross-sensor transfer learning and ground-based physical constraints, and adopts a hierarchical and progressive accuracy improvement strategy: the first stage constructs a global basic inversion model based on multi-scale deep neural networks to output a large field of view initial probability field; the second stage constructs a lightweight nonlinear residual prediction model with multi-source auxiliary environmental parameters to implement local physical correction.

[0018] This invention employs a cross-sensor transfer fine-tuning method based on heterogeneous observation geometric correction and multi-source label cascade optimization. Spatial registration is used as a prerequisite constraint for cross-sensor transfer learning. After completing the global pre-training of the network using large-sample coarse-resolution labels from geostationary satellites, parameters such as cloud top height are introduced to perform reverse parallax correction on the high-precision labels of polar-orbiting satellites to eliminate projection bias. Then, under the premise of freezing the weight parameters of the backbone encoder, the decoder network is fine-tuned in a directional manner.

[0019] This invention employs a regional physical residual enhancement mechanism based on the dimensional transformation of "sparse points to continuous surfaces". It constructs a feature space by explicitly introducing physical parameters (such as ground-based atmospheric precipitable water) independent of the passive optical imaging link and auxiliary geographic information, and uses the high-precision one-dimensional trajectory profile provided by spaceborne lidar (such as CALIPSO) as the benchmark truth to train a lightweight nonlinear regression model, establish the mapping relationship between the physical background and the inversion residual, and then extrapolate the sparse physical correction capability along the track to the global continuous surface.

[0020] This invention overcomes the noise limitations of traditional pixel-level inversion, significantly improving the spatial continuity of cloud boundaries and the accuracy of fragmented cloud discrimination. Existing official L2 products from geostationary satellites are mostly based on traditional pixel-level physical thresholding algorithms, processing single pixels in isolation while ignoring the local spatial correlation of cloud clusters. This results in inversion results often containing severe salt-and-pepper noise and jagged edges. This invention, by introducing high-quality MODIS labels optimized for disparity correction and downsampling for transfer fine-tuning, effectively corrects the overfitting of pre-trained models to the inherent noise of L2 labels.

[0021] This invention effectively solves the problem of "physical obfuscation" in passive remote sensing under complex underlying surfaces and atmospheric backgrounds. Addressing the issue of existing pure visual models easily obfuscating spectral features under complex atmospheric backgrounds, this invention introduces ground-based PWV data as supplementary physical information. The aim is to use total atmospheric water vapor information to correct the inversion results, providing the model with a background reference independent of optical images. Based on the cloud mask probability output by the basic model, this invention further introduces a region enhancement model for residual correction. This correction model uses the previously constructed physical-environment features as input and learns the nonlinear relationship between the prediction residuals of the basic model and multi-source auxiliary features to achieve local adaptive optimization of the initial cloud detection results.

[0022] This invention achieves "real-time driven" rather than "post-event verification" of sparse, high-value data. Existing technologies typically only use CALIPSO and ground-based observation data for post-event accuracy verification, resulting in a waste of data value. This invention proposes a "sparse point-continuous surface" residual enhancement architecture, successfully transforming sparse vertical probe ground truth and ground-based observation data into real-time, pixel-level correction capabilities for inversion results, maximizing the synergistic benefits of multi-source heterogeneous data.

[0023] This invention features a flexible and efficient architecture that meets the real-time processing requirements of operational applications. The "deep learning foundation + lightweight enhancement" architecture employed demonstrates high engineering feasibility. Once the base model is trained, it can be fixed; regional adaptive adjustments only require retraining the lightweight LightGBM model (taking only sub-minutes), with extremely fast inference speed. This avoids the high cost of fully retraining a massive deep network to adapt to different regions, making it extremely easy to deploy and promote in meteorological operational systems. Attached Figure Description

[0024] Figure 1 Overall method architecture diagram;

[0025] Figure 2 Flowchart of the spatiotemporal matching module for multi-source air and ground data;

[0026] Figure 3 InceptionResnetV2+UNet++ cloud parameter inversion network architecture diagram;

[0027] Figure 4 Flowchart of cloud parameter inversion for transfer learning;

[0028] Figure 5 Flowchart of residual compensation for regional cloud parameter inversion;

[0029] Figure 6 Segmentation results of the transfer learning model;

[0030] Figure 7 The segmentation results of the region enhancement model. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above objectives, this invention adopts the following technical solution.

[0032] like Figure 1 As shown, the method of the present invention mainly includes three steps: spatiotemporal matching of multi-source data from air and ground, cloud parameter inversion based on transfer learning, and regional residual enhancement based on physical constraints.

[0033] Step 1. Spatiotemporal matching of multi-source air and ground data:

[0034] This step aims to establish a geometrically rigorous, hierarchically increasing multi-source heterogeneous data system. Considering the differences in observation geometry, temporal resolution, and spatial resolution among geostationary satellites, polar-orbiting satellites, spaceborne lidar, and ground-based stations, this invention implements the following data processing steps, such as... Figure 2 As shown:

[0035] Acquire full-year daytime observation data from the Fengyun-4B geostationary meteorological satellite (FY-4B). Establish a 15-minute time-matching framework for the FY-4B's two-dimensional flat-scan observation mode. Utilizing the north-to-south scanning characteristics of FY-4B, read the scan time file and accurately record the actual scan time of each row of pixels across the entire disk. Simultaneously, acquire secondary cloud product data (MOD06) from the Moderate Resolution Imaging Spectroradiometer (MODIS). Taking into account the 5-minute scan swath data characteristics of MODIS, use every 15-minute interval of FY-4B as a reference to match three adjacent MOD06 files (e.g., for FY-4B at 00:00, match MOD06 data at 00:00, 05:00, and 10:00). These three swath data are then stitched together to serve as the polar-orbiting satellite data to be matched for the current moment.

[0036] A standard geographic latitude and longitude grid was established with a spatial resolution of 0.04 degrees to match the nominal resolution (4km) of FY-4B. The geographic range was preferably set to 50°S-50°N and 80°E-180°E to eliminate areas with excessive distortion at the edges of the FY-4B full-disk observations. The first-level observation data (L1), geolocation data (GEO), second-level official product data (L2), and MODIS mosaic data of FY-4B were all reprojected onto this standard grid. Specifically, for the L1, GEO, L2, and scan time data of FY-4B, the GDAL tool and nearest neighbor interpolation method were used to map them to the nearest neighbor geographic grid nodes.

[0037] To address the discrete data characteristics of MOD06 cloud mask products with a resolution of 1km, this invention employs a mode-based spatial constraint downsampling strategy: Using the center of a standard grid as the center, the 16 nearest MODIS pixels within a preset threshold distance (e.g., 4500 meters) are selected; the mode strategy is used to determine the classification result of the central grid; simultaneously, an interpolation distance matrix is ​​established, and grid points with interpolation distances exceeding 3000 meters are removed from the mask to prevent excessive edge interpolation expansion. Furthermore, the MODIS cloud mask labels undergo binarization and merging processing: "uncertain cloud" is classified as "cloud," and "uncertain clear sky" is classified as "clear sky," generating high-precision binary classification mask labels containing only clouds and clear sky.

[0038] The FY-4B and MODIS scan times were iterated over at each grid point, and the absolute value of the time difference was calculated. A time threshold of 5 minutes was set, retaining only regions with observation intervals within 5 minutes, while masking was used to remove regions exceeding this threshold. Simultaneously, spatial continuity filtering was performed to remove moments with an effective spatial matching pixel range of less than 300×300, ensuring the spatial integrity of the samples.

[0039] To address the parallax issue arising from the different observation angles between geostationary and polar-orbiting satellites, this invention implements rigorous parallax correction: First, the cloud top height (CTH) data is read from the MOD06 product. Given that CTH is a continuous variable, an inverse distance-squared weighted interpolation method is used to map it to the grid. The calculation formula is as follows:

[0040] ;

[0041] in, As weight, The distance of each pixel from the center point. To minimize the value, division by zero is avoided. To prevent interpolation expansion in non-cloud areas, cloud masking results are used to mask the interpolated CTH data.

[0042] Subsequently, based on the cloud top height data for each pixel, the nadir position (133°E) of the FY-4B satellite, and the observation geometry, a parallax correction model was established to reverse-correct the MODIS cloud mask product from its "true geographical location" to the "apparent position" of FY-4B. This step eliminates projection bias caused by cloud height and observation angle, ultimately generating FY4B-MODIS sample pairs that are strictly registered in time and space. The parallax correction process is as follows:

[0043] (1) For each pixel, let its latitude and longitude coordinates under FY-4B geometric positioning be... The height of the cloud top is The unit is meters, and the observed zenith angle is... azimuth angle is .

[0044] (2) Based on the "section approximation method", the horizontal displacement of the ground projection caused by cloud parallax is approximated as a function of cloud top height and observed zenith angle. Convert to radians :

[0045] ;

[0046] And calculate the horizontal displacement distance (displacement modulus along the observation azimuth direction):

[0047] ;

[0048] in, This represents the horizontal displacement distance due to parallax. Increase or As the value increases, the displacement also increases, which is consistent with the physical law that projection shift is more significant under high clouds and oblique observation.

[0049] (3) Consider the above displacement as occurring on the WGS-84 ellipsoid from the original pixel position. Along azimuth forward distance The geodesic forward problem yields corrected coordinates. :

[0050] ;

[0051] in, This represents the forward geodesic operator for the ellipsoid (given the starting point, azimuth, and distance, it outputs the latitude and longitude of the ending point). This operator guarantees higher geographical consistency over a large area compared to planar approximations.

[0052] Subsequently, cloud-aerosol lidar and CALIPSO (California Infrared Pathfinder Satellite) level 2 cloud product data for the target observation area were acquired. Cloud top height data were read, and using the same reverse parallax correction method, the one-dimensional cloud mask trajectory of CALIPSO was projected onto the 0.04-degree grid of the FY-4B "apparent position" as a high-precision reference label. Simultaneously, ground-based atmospheric precipitable water (PWV) data from several ground-based meteorological stations within the target observation area were acquired. Kriging interpolation was used to spatially interpolate the discrete station data, generating a continuous two-dimensional water vapor field strictly aligned with satellite pixels, serving as a physical constraint feature. Digital elevation data and land surface type data for the target observation area were acquired and resampled to the same spatial resolution as topographic feature elements.

[0053] Based on the above matching data, a slice dataset suitable for deep learning models was constructed: the L1 data, GEO data, and L2 data of FY-4B were cropped into 256×256 pixel blocks with a step size of 128 pixels. The 15 channels of the L1 data of FY-4B and the satellite zenith angle data in the GEO data were selected as input features. The input data were standardized. The L1 data was Z-score standardized based on the mean and standard deviation of the entire dataset, and the cosine value of the satellite zenith angle data was calculated to ensure that all input features were distributed on the same order of magnitude. Thus, two types of training sets were constructed: (1) a large sample basic pre-training set: labeled with the official cloud mask product of the L2 data of FY-4B; (2) a small sample high-precision fine-tuning set: labeled with the MODIS cloud mask product after parallax correction and downsampling.

[0054] Step 2. Cloud parameter inversion based on transfer learning:

[0055] This step aims to construct a high-precision deep learning-based inversion network to provide a high-confidence global initial probability field for subsequent regional residual enhancement. Addressing the feature fusion challenge between deep semantic feature extraction and high-frequency spatial detail preservation in cloud parameter inversion tasks, this invention constructs a multi-scale network topology architecture based on a "multi-branch residual encoder-dense skip connection decoder" (e.g., an improved combination of Inception-ResNetV2 and U-Net++ architecture) to achieve deep mining and reconstruction of multi-scale features. Simultaneously, this step employs a two-stage transfer learning strategy across sensors: "large-sample coarse-label global feature pre-training" and "small-sample high-precision label fine-tuning," thereby achieving efficient nonlinear mapping of complex feature spaces.

[0056] A schematic diagram of the overall network structure is shown below. Figure 3 As shown. The network input is a multidimensional feature tensor. , denoted as:

[0057] ;

[0058] in, For feature dimension, The input image spatial dimensions are specified. The network output is a binary cloud mask (or a cloud / clear sky logits probability map). , denoted as:

[0059] ;

[0060] In the encoder section, the Inception-ResNetV2 network is used as the backbone feature extractor, responsible for extracting multi-level features from the FY-4B multispectral image, ranging from low-level texture to high-level semantics. The encoder consists of a Stem module and three sets of Inception-ResNet modules (A / B / C), progressively downsampling and expanding the receptive field to output 5 levels of multi-scale features. .

[0061] The Stem module consists of multiple stacked convolutional and pooling layers, used to quickly reduce the spatial resolution of the input image and expand the number of channels. Input features First, a set of convolutional feature extraction units (including 3×3 and 1×1 convolutions) are used. After each convolution, batch normalization and non-linear activation are performed sequentially. In the Stem stage, a 3×3 convolution with a stride of 2 is used to complete the first downsampling, reducing the feature map resolution from... Down to Then, a second downsampling was performed using 3×3 max pooling (MaxPool with a stride of 2), further reducing the feature map resolution to [the desired value]. Finally, shallow features are obtained. Features of the middle and shallow layers :

[0062] ;

[0063] ;

[0064] Before entering the Inception-ResNet-A group, this embodiment first performs a 3×3 max pooling downsampling (with a stride of 2) on the output features of the previous stage, reducing the spatial resolution of the feature map from... Down to This expands the receptive field and reduces the computational cost of subsequent multi-branch convolutions. The Inception-ResNet-A module is used to... Multi-scale texture and local semantic extraction are performed at various scales. The main module employs a three-branch parallel convolutional structure combined with a residual scaling mechanism. Its specific structure is as follows:

[0065] (1) First branch: 1×1 convolution, mapping the input channel 320 to 32;

[0066] (2) Second branch: 1×1→3×3 concatenated convolution, which maps the input channel 320 to 32→32 in sequence;

[0067] (3) Third branch: 1×1→3×3→3×3 concatenated convolution, which maps the input channel 320 to 32→48→64 in sequence;

[0068] The outputs of the three branches are concatenated in the channel dimension to obtain 128-channel features, and then the channels are aligned back to 320 through a 1×1 convolution. Residual scaling and activation are then applied.

[0069] ;

[0070] in, Input feature map to the module, To output the feature map, The activation function is ReLU. This represents the residual components after 1×1 alignment following the fusion of the three branches. This is the residual scaling factor. Finally, after cascading 10 A modules, the third-level feature is output:

[0071] ;

[0072] The Reduction-A module is used to reduce features from downsampling to Furthermore, a multi-branch downsampling fusion strategy is employed during the downsampling process to reduce information loss. Its specific structure is a three-way parallel downsampling:

[0073] (1) First branch: 3×3 convolution with stride of 2, mapping the channels from 320 to 384, and realizing spatial downsampling to ;

[0074] (2) Second branch: 1×1→3×3→3×3 (stride 2) concatenated convolution, mapping the channels from 320 to 256→256→384 in sequence, and performing downsampling in the last layer. ;

[0075] (3) Third branch: 3×3 max pooling (step size 2) to achieve spatial downsampling to The channel remains at 320;

[0076] The three branch outputs are concatenated along the channel dimension to obtain the Reduction-A output with 384 + 384 + 320 = 1088 channels. This output is then fed into the Inception-ResNet-B group for further deep semantic extraction. The Inception-ResNet-B module employs a two-branch structure to expand the receptive field while maintaining computational control, thereby enhancing its ability to model large-scale cloud clusters and cloud band structures.

[0077] (1) First branch: 1×1 convolutional branch, which maps the input feature channels from 1088 to 192;

[0078] (2) Second branch: Factorization convolution branch, which uses a convolution combination of 1×1→1×7→7×1 in sequence (corresponding channel mapping is 1088→128→160→192) to obtain a larger equivalent receptive field;

[0079] The outputs of the two branches are concatenated along the channel dimension to obtain 384-channel features, and then the channels are aligned back to 1088 through a 1×1 convolution. The residual scaling residual structure is also used. Finally, after being concatenated through 20 B modules, the output is obtained as the fourth-level features:

[0080] ;

[0081] Completed via Reduction-B arrive The downsampling also employs a parallel fusion of multi-branch stride convolution and pooling to improve information preservation during the downsampling stage; the Inception-ResNet-C module uses a deeper semantic extraction structure:

[0082] (1) First branch: 1×1 convolutional branch, which maps the input feature channels from 2080 to 192;

[0083] (2) Second branch: Factorization convolution branch, which uses 1×1→1×3→3×1 convolution combination in sequence (corresponding channel mapping is 2080→192→224→256).

[0084] The outputs of the two branches are concatenated to obtain 448-channel features, which are then aligned back to 2080 through a 1×1 convolution, followed by residual scaling and activation. At the end of the encoder, the Inception-ResNet-C module is first cascaded nine times, and then a residual module without activation is used once at the end as a bottleneck to reduce the non-linear loss of high-level semantics in the final stage. Subsequently, a 1×1 convolution compresses the channels from 2080 to 1536, forming the final deep feature output.

[0085] ;

[0086] In the decoder section, a nested dense skip connection structure based on U-Net++ is adopted. The decoder contains multiple levels of upsampling modules and feature fusion modules. In this implementation, the decoder has a total of 11 decoding node blocks, distributed by scale as follows:

[0087] ;

[0088] Each decode node block contains the following processing steps:

[0089] (1) Perform 2× upsampling on the main input features to restore spatial resolution;

[0090] (2) Concatenate with the SkipConnection feature in the channel dimension;

[0091] (3) The output of the node is obtained by performing feature reshaping and channel compression after two 3×3 convolutions (each convolution is combined with normalization and nonlinear activation).

[0092] The processing procedure can be represented as follows:

[0093] ;

[0094] in, For a 2× upsampling operator, This is the set of skip / dense fusion features for this node. This represents the "two-layer 3×3 convolution reconstruction" operator.

[0095] Final node It only receives the features from the output of the previous scale (H / 2) after 2× upsampling, and outputs them through two layers of 3×3 convolution. This node serves as the "final feature reconstruction block", integrating high-resolution semantics and edge information into the input representation of the segmentation head.

[0096] This invention implements a training method based on cross-sensor knowledge distillation, such as... Figure 4 As shown. First, a large-scale pre-training dataset is constructed. The input is the FY4B Level 1 multi-channel data and angle data mentioned above, and the label is the official Level 2 cloud mask product (dataset A). The above deep network is pre-trained in a fully supervised manner, so that the model learns the basic texture, shape and atmospheric radiative transfer characteristics of clouds, and obtains pre-trained weights with basic generalization ability.

[0097] Subsequently, a higher-precision fine-tuning dataset was constructed. The input remained the same FY4B data, but the labels were replaced with MODIS secondary cloud mask products (dataset B) that had undergone spatiotemporal matching and downsampling processing. The aforementioned pre-trained weights were loaded, and fine-tuning training was performed using a strategy of freezing encoder parameters and only opening decoder parameters. This process utilizes high-quality MODIS labels as prior knowledge, and while maintaining the encoder's ability to extract spectral features from geostationary satellites, it reshapes the decoder's feature mapping logic. This corrects the label noise inherited from the pre-trained model, significantly improving the accuracy of cloud edge and fine cloud discrimination, and effectively improving the problems of cloud omissions and misclassifications in complex surface environments such as snow cover, deserts, and water bodies.

[0098] Step 3. Region residual enhancement based on physical constraints:

[0099] This step aims to use physical observation data independent of the passive optical imaging link to perform nonlinear residual correction on the systematic bias of the basic model in a specific region. The overall process is as follows: Figure 5 As shown. This invention innovatively constructs a cross-dimensional enhancement architecture of "sparse points-continuous surfaces", and the specific implementation steps are as follows.

[0100] To incorporate ground-based physical constraints and obtain high-precision ground-value labels, a spatiotemporally aligned heterogeneous dataset (dataset C) first needs to be constructed. Then, a two-dimensional water vapor field is obtained by interpolating ground-based atmospheric precipitable water data from the target observation area. The cloud classification products observed by the CALIPSO satellite over the target observation area (Shandong Province in this embodiment) are used as high-precision reference benchmarks for cloud detection. The classification results are then probabilistically mapped: pixels identified as "clouds" are assigned a value of 1.0, and pixels identified as "clear sky" are assigned a value of 0.0. These probabilities are recorded as the benchmark true value probabilities. .

[0101] Based on the latitude and longitude of the CALIPSO satellite's nadir trajectory and the observation time, the spatiotemporal corresponding basic model prediction results are extracted. Specifically, the extracted basic model results are the raw confidence probability values ​​without binarization processing, denoted as... The probability residual for each matching point is calculated using the following formula:

[0102] ;

[0103] in, This refers to the regression objective residual that the region enhancement model needs to learn. This indicates that the basic model has a tendency to miss detections. This indicates that the basic model has a tendency to produce false alarms.

[0104] This invention employs the LightGBM algorithm to construct a lightweight regional augmentation model, which aims to establish a nonlinear mapping relationship between ground physical quantities and foundation model errors. The model input feature vector is a hybrid feature vector of physical constraints and observation geometry. Specifically, it includes:

[0105] (1) Physical constraint: Ground water vapor value after spatiotemporal matching , used to indicate the current thermodynamic state of the atmosphere;

[0106] (2) Basic information item: the predicted probability of the basic model ;

[0107] (3) Auxiliary environment items: solar zenith angle, satellite zenith angle, ground elevation, and surface type data, which are used to help the model understand the influence of observation geometry and underlying surface environment on radiative transfer and cloud-ground spectral confusion effect, and enhance the model context information;

[0108] The training objective of the model is to enable the LightGBM regression model to learn the mapping function. To make its output as close as possible to the true residual. :

[0109] ;

[0110] in, This is the probability correction amount for the model prediction. For ground-based water vapor data, Based on the probability, These are auxiliary environmental factor variables.

[0111] Using the trained LightGBM model, the physical correction capability is extended from sparse CALIPSO trajectories to the entire target region. First, the ground-based water vapor field of the current target region is analyzed. Probability graphs obtained from basic model inference The corresponding geometric auxiliary data is input into the pre-trained LightGBM model. The model performs point-by-point inference for each pixel in the image and outputs a probability residual correction field covering the entire image. A linear superposition strategy is used to apply the predicted residual field to the base probability map. The final cloud mask probability is... The calculation formula is as follows:

[0112] ;

[0113] in, This indicates that the value is limited to... Within the range. Finally, the corrected probability map is adjusted using a threshold of 0.5. Binarization is performed to generate the final high-precision region-enhanced cloud mask product.

[0114] To verify the effectiveness and generalization ability of the cloud mask inversion model based on cross-sensor transfer learning described in this invention, this embodiment constructs an experimental environment based on the PyTorch deep learning framework, with the computing platform configured as a single NVIDIA GeForce RTX4090 graphics processing unit (GPU). The experimental dataset is randomly divided according to a preset ratio, with the training set, validation set, and test set ratio set at 70%:15%:15%. The model input is 16-channel FY-4B multispectral feature data including satellite zenith angle, and the output is a binary classification cloud mask, where the label category is defined as cloud and clear sky.

[0115] During the model training phase, the Adam optimizer was used in conjunction with a cosine annealing learning rate scheduling algorithm for iterative parameter updates. The Focal Loss function was adopted, which introduces a focusing parameter. The weights of easily classifiable samples are dynamically reduced, forcing the model to focus its training on sparse, difficult-to-classify samples (such as thin clouds, fragmented clouds, or label edge regions), ensuring the robustness of the base model. This invention selects Intersection over Union (IoU), Precision, and Recall as core evaluation metrics to monitor the model's performance on the training and validation sets in real time, comprehensively assessing the model's convergence speed and overfitting risk. Monitoring data during training shows that the model reached a stable convergence state at the 80th epoch, with IoU ratios of 0.8898 and 0.8809 for the training and test sets, respectively; recall rates of 0.9439 and 0.9392; and precision rates of 0.9395 and 0.9342. The two sets of data are highly similar, indicating that the model did not exhibit overfitting. To further verify the model's generalization performance, inference tests were conducted on an independent test set using the optimal model parameters. Statistical results show that the intersection-union ratio of the test set reached 0.8848, the recall rate was 0.9403, and the precision rate was 0.9374. The test set metrics are highly consistent with those of the training and validation sets, confirming that the method described in this invention has excellent stability and generalization ability, and can maintain high recall (reducing false negatives) while also achieving high precision (reducing false positives).

[0116] Figure 6The visualization results of cloud segmentation using the model of this invention on the test set are shown. From left to right, the columns display the true-color composite image (3 / 2 / 1 channels), standard ground truth labels, and model prediction results. Visual comparison results show that the model of this invention achieves accurate segmentation in various land and cloud formations, preserving the detailed texture of the cloud system well. It also exhibits robust performance in complex conditions such as land-sea boundaries and fragmented cloud areas. The mutual corroboration between subjective visual effects and objective quantitative indicators fully demonstrates that the cloud detection method described in this invention can achieve consistent, reliable, and high-precision segmentation results on the FY-4B and MOD06 fused samples.

[0117] Figure 7 This paper presents an example of cloud mask results for a target observation area (such as Shandong Province) generated after residual correction, and compares it with the official cloud mask product from the same period's FY-4B. As can be seen from the figure, the corrected model effectively suppresses false cloud pixels caused by abrupt changes in surface type and spectral mixing effects at the land-sea interface, significantly improving detection accuracy near the coastline. Simultaneously, the model demonstrates stronger detection capabilities for thin clouds, cirrus clouds, and other clouds with low optical thickness, better distinguishing clouds from clear-sky atmosphere and improving the capture of subtle cloud structures. These results validate the dual role of the region enhancement model in reducing false alarms and enhancing the identification of weak targets.

[0118] Another aspect of the present invention provides a satellite cloud detection device based on transfer learning and ground-based physical augmentation, comprising the following modules:

[0119] The multi-source air-ground data spatiotemporal matching module performs geometrically rigorous registration and spatiotemporal alignment of multi-source heterogeneous observation data from geostationary satellites, polar-orbiting satellites, spaceborne lidar, and ground-based stations. It eliminates projection deviations caused by cloud height and observation angle through parallax correction and constructs a large-sample basic pre-training set and a small-sample high-precision fine-tuning set.

[0120] The cloud parameter inversion module based on transfer learning constructs a deep neural network consisting of an encoder and a decoder. It performs fully supervised pre-training using the large sample basic pre-training set, freezes the encoder parameters, and performs transfer fine-tuning of the decoder using the small sample high-precision fine-tuning set to obtain the basic model and output the global initial cloud mask probability field.

[0121] The physical constraint-based regional residual enhancement module uses the sparse trajectory cloud detection ground truth provided by the spaceborne lidar as the residual learning target, extracts the prediction probability of the basic model on the corresponding trajectory and calculates the residual, and uses the ground-based atmospheric precipitable water, the prediction probability of the basic model and auxiliary environmental features as inputs to train a lightweight regression model to learn the nonlinear mapping between the residual and multi-source features. The regression model is applied to the whole domain, predicts the residual correction field pixel by pixel and superimposes it on the initial probability field, and generates the final cloud mask after thresholding.

[0122] Another aspect of the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method.

[0123] Another aspect of the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to implement the method described thereon.

[0124] The present invention also has the following alternative technical solutions:

[0125] 1. Alternatives to the basic inversion module

[0126] Backbone network: Although Inception-ResNetV2 is preferred as the encoder in this embodiment, it can also be replaced with other feature extraction networks such as ResNet-50 / 101, EfficientNet, and SwinTransformer, which have multi-scale feature extraction capabilities.

[0127] Segmentation architecture: Although U-Net++ is preferred, it can also be replaced by semantic segmentation architectures with multi-scale feature fusion capabilities such as DeepLabV3+, SegFormer, LinkNet or FPN.

[0128] Tag source: In addition to MODIS cloud mask products, high-precision cloud products from polar-orbiting satellites such as VIIRS (NPP), MERSI (FY3F) or SLSTR (Sentinel-3) can also be used, such as cloud phase and cloud top height, and after the same parallax correction and resampling processing, they can be used as tags.

[0129] 2. Alternatives to the Region Enhancement Module

[0130] Regression Algorithm: Although LightGBM is preferred in this invention, this lightweight regression model can also be replaced by other tree-based ensemble learning algorithms or feedforward neural networks (such as XGBoost, CatBoost, Random Forest or Multilayer Perceptron MLP, etc.).

[0131] Physical constraint variables: In addition to ground-based PWV, other multidimensional continuous parameters that can characterize the atmospheric thermodynamic state and dynamic characteristics can be introduced, such as temperature and humidity profiles and atmospheric boundary layer height provided by numerical weather prediction models, as auxiliary physical constraint features to be included in the regression model.

[0132] 3. Alternative solutions for data processing and matching

[0133] Interpolation methods: In addition to Kriging interpolation, spatialization of the ground water vapor field can also be achieved using methods such as the inverse distance weighting method and spline function method to generate a continuous two-dimensional background field.

[0134] Parallax correction: In addition to geometric projection correction based on cloud top height, in specific scenarios, data-driven image-level feature registration methods (such as optical flow or stereo matching networks based on deep learning) can also be used to achieve automatic registration of heterogeneous images.

Claims

1. A satellite cloud detection method based on transfer learning and ground-based physical augmentation, characterized in that, Includes the following steps: First, geometrically rigorous registration and spatiotemporal alignment are performed on multi-source heterogeneous observation data from geostationary satellites, polar-orbiting satellites, spaceborne lidar, and ground-based stations. Parallax correction is used to eliminate projection deviations caused by cloud height and observation angle, and a large-sample basic pre-training set and a small-sample high-precision fine-tuning set are constructed. Then, a deep neural network consisting of an encoder and a decoder is constructed. Fully supervised pre-training is performed using the large sample basic pre-training set. The encoder parameters are then frozen and the decoder is transferred and fine-tuned using the small sample high-precision fine-tuning set to obtain the basic model and output the global initial cloud mask probability field. Finally, using the sparse trajectory cloud detection ground truth provided by the spaceborne lidar as the residual learning objective, the predicted probability of the basic model on the corresponding trajectory is extracted and the residual is calculated. With the ground-based atmospheric precipitable water, the predicted probability of the basic model and the auxiliary environmental features as inputs, a lightweight regression model is trained to learn the nonlinear mapping between the residual and the multi-source features. The regression model is applied to the entire domain, and the residual correction field is predicted pixel by pixel and superimposed on the initial probability field. After thresholding, the final cloud mask is generated.

2. The method according to claim 1, characterized in that, The geometrically rigorous registration and spatiotemporal alignment include: Geostationary satellite and polar-orbiting satellite data are reprojected onto a unified geographic grid. Polar-orbiting satellite swath data with time differences within a threshold are matched based on the observation time of geostationary satellites. A mode downsampling strategy is used to map high-resolution polar-orbiting satellite cloud masks onto the grid. Inverse parallax correction is performed on polar-orbiting satellite labels based on cloud top height and observation geometry to generate pixel-level aligned sample pairs. The one-dimensional cloud mask trajectory of the spaceborne lidar is projected onto the same grid after parallax correction as a high-precision reference label; Spatial interpolation is performed on discrete ground-based atmospheric precipitable water data to generate a continuous two-dimensional water vapor field aligned with the grid.

3. The method according to claim 1, characterized in that, The encoder of the deep neural network uses a multi-scale residual network to extract multi-level features from shallow texture to deep semantics, and the decoder uses a nested dense jump connection structure to upsample and fuse the features of each layer of the encoder, and finally outputs a binary classification cloud mask probability map.

4. The method according to claim 1, characterized in that, During the transfer fine-tuning process, all encoder weights of the pre-trained model are frozen, and only the decoder parameters are opened. The model is then retrained using a high-precision polar-orbiting satellite cloud mask with parallax correction as a supervision signal, so that the decoder can reshape the feature mapping logic to correct the label noise inherited by the pre-trained model.

5. The method according to claim 1, characterized in that, In the residual calculation, the true value of the spaceborne lidar is probabilistically mapped to a cloud probability of 1 and a clear sky probability of 0. The output of the basic model is the original confidence probability without binarization. The difference between the two is used as the learning target of the regression model. A positive value indicates the tendency of the basic model to miss detections, and a negative value indicates the tendency of false alarms.

6. The method according to claim 1, characterized in that, The input feature vector of the lightweight regression model includes: ground-based atmospheric precipitable water, basic model prediction probability, solar zenith angle, satellite zenith angle, ground elevation, and surface type data; the regression model uses a tree-based ensemble learning algorithm or a feedforward neural network.

7. The method according to claim 1, characterized in that, The residual correction field and the initial probability field are superimposed using a linear superposition strategy. After superposition, the superimposed field is clipped to the [0,1] interval and then binarized with a threshold of 0.5 to generate the final cloud mask, thereby realizing the physical correction capability extrapolation from sparse trajectories to a global continuous surface.

8. A satellite cloud detection device based on transfer learning and ground-based physical augmentation, characterized in that, Includes the following modules: The multi-source air-ground data spatiotemporal matching module performs geometrically rigorous registration and spatiotemporal alignment of multi-source heterogeneous observation data from geostationary satellites, polar-orbiting satellites, spaceborne lidar, and ground-based stations. It eliminates projection deviations caused by cloud height and observation angle through parallax correction and constructs a large-sample basic pre-training set and a small-sample high-precision fine-tuning set. The cloud parameter inversion module based on transfer learning constructs a deep neural network consisting of an encoder and a decoder. It performs fully supervised pre-training using the large sample basic pre-training set, freezes the encoder parameters, and performs transfer fine-tuning of the decoder using the small sample high-precision fine-tuning set to obtain the basic model and output the global initial cloud mask probability field. The physical constraint-based regional residual enhancement module uses the sparse trajectory cloud detection ground truth provided by the spaceborne lidar as the residual learning target, extracts the prediction probability of the basic model on the corresponding trajectory and calculates the residual, and uses the ground-based atmospheric precipitable water, the prediction probability of the basic model and auxiliary environmental features as inputs to train a lightweight regression model to learn the nonlinear mapping between the residual and multi-source features. The regression model is applied to the whole domain, predicts the residual correction field pixel by pixel and superimposes it on the initial probability field, and generates the final cloud mask after thresholding.

9. An electronic device, characterized in that, include: One or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, cause the processor to perform the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for inverting water quality lacking remote sensing image by combining migration and space-time deep learning

    CN119625551A

  • Atmospheric correction method and device based on deep learning inversion AOD (Argon Oxygen Decarburization) and medium

    CN121837962A