Satellite cloud cover retrieval method and device based on joint driving of space-time physics
Patent Information
- Application Number
- CN202611134107.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-08-28
AI Technical Summary
[0005]有鉴于此,本申请实施例提供了一种基于时空物理联合驱动的卫星云量反演方法及装置,以解决静止气象卫星连续多光谱时序数据与极轨卫星云量产品之间的像元偏差的技术问题
[0019] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.
Smart Images

Figure CN122657748A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of satellite cloud parameter inversion and deep learning technologies, and in particular to a satellite cloud amount inversion method and apparatus based on spatiotemporal physics joint driving. Background Technology
[0002] Cloud cover, as a core meteorological element, directly governs the radiation and heat balance of the Earth-atmosphere system and is an indispensable key input parameter for numerical weather prediction, climate change monitoring, and atmospheric radiation transfer simulation. Compared to traditional products that only provide binary cloud masks, continuous cloud cover fields can more accurately characterize sub-pixel cloud coverage, especially in the transition zone between "likely cloudy" and "likely clear," providing richer information for downstream applications. Acquiring near real-time, high-precision cloud cover data with spatiotemporal continuity has become an important prerequisite for ensuring the safety of solar photovoltaic grid connection, optimizing regional energy dispatch, and improving the ability to provide refined meteorological services.
[0003] Existing cloud cover retrieval techniques mainly fall into two categories: traditional physical methods and deep learning methods. Traditional physical methods typically rely on multispectral thresholding or statistical testing to first generate a binary cloud mask, and then calculate cloud cover based on this mask. Deep learning methods, on the other hand, usually utilize fully convolutional network architectures to automatically extract deep spatial features from multi-channel remote sensing images from meteorological satellites, thereby establishing a complex nonlinear mapping relationship between satellite-observed radiance values and continuous cloud cover. Compared to traditional physical methods, deep learning models are better at mining the local texture and global contextual information of images, effectively eliminating the over-reliance on complex physical parameters and prior fixed thresholds, and significantly improving retrieval efficiency and overall accuracy.
[0004] However, existing deep learning methods still have drawbacks when performing transport inversion, such as large bias in cloud cover estimation in transition zones, significant decrease in inversion accuracy and high-frequency detail fidelity at cloud edges or in areas of rapid advection and severe convection, and further increase inversion error due to non-rigid offsets. Summary of the Invention
[0005] In view of this, embodiments of this application provide a satellite cloud cover inversion method and apparatus based on spatiotemporal physics joint driving, in order to solve the technical problem of pixel deviation between continuous multispectral time-series data of geostationary meteorological satellites and cloud cover products of polar-orbiting satellites.
[0006] A first aspect of this application provides a satellite cloud cover inversion method based on spatiotemporal physics joint driving, comprising:
[0007] Preprocessing polar-orbiting satellite cloud product data yields historical cloud cover data;
[0008] Based on the observation time, historical observation data of meteorological satellites and cloud product data of polar-orbiting satellites are registered. Based on the registered historical observation data of meteorological satellites and historical cloud cover data of polar-orbiting satellites, a spatiotemporal slice sample set is constructed. The spatiotemporal slice sample set uses time-series spectral feature sequences constructed based on historical observation data of meteorological satellites as samples, and uses the historical cloud cover data at the corresponding time of each feature as sample labels.
[0009] A pre-trained spatiotemporal hybrid feature extraction network is constructed, and the pre-trained spatiotemporal hybrid feature extraction network is trained using a spatiotemporal slice sample set to obtain the trained spatiotemporal hybrid feature extraction network.
[0010] Real-time meteorological satellite observation data is acquired, and satellite cloud cover is obtained by inverting the real-time observation data using a trained spatiotemporal hybrid feature extraction network.
[0011] The pre-trained spatiotemporal hybrid feature extraction network includes a 3D spatiotemporal encoder, a multi-scale temporal aggregation bottleneck module, a skip-connected temporal fusion dimensionality reduction module, a fusion deformable convolution-aligned 2D decoder, and a dual-branch multi-task output head.
[0012] A second aspect of this application provides a satellite cloud cover inversion device based on spatiotemporal physics joint driving, comprising:
[0013] The cloud product data preprocessing module is configured to preprocess polar-orbiting satellite cloud product data to obtain historical cloud volume data;
[0014] The dataset construction module is configured to register historical meteorological satellite observation data and polar-orbiting satellite cloud product data based on observation time, and construct a spatiotemporal slice sample set based on the registered historical meteorological satellite observation data and historical cloud cover data of polar-orbiting satellites. The spatiotemporal slice sample set uses time-series spectral feature sequences constructed based on historical meteorological satellite observation data as samples, and uses the historical cloud cover data at the corresponding time of each feature as sample labels.
[0015] The model training module is configured to build a pre-trained spatiotemporal hybrid feature extraction network, train the pre-trained spatiotemporal hybrid feature extraction network using a spatiotemporal slice sample set, and obtain the trained spatiotemporal hybrid feature extraction network.
[0016] The inversion module is configured to acquire real-time observation data from meteorological satellites and use a trained spatiotemporal hybrid feature extraction network to invert satellite cloud cover based on the real-time observation data.
[0017] The pre-trained spatiotemporal hybrid feature extraction network includes a 3D spatiotemporal encoder, a multi-scale temporal aggregation bottleneck module, a skip-connected temporal fusion dimensionality reduction module, a fusion deformable convolution-aligned 2D decoder, and a dual-branch multi-task output head.
[0018] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0019] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.
[0020] The beneficial effects of this application embodiment compared with the prior art are as follows: This application embodiment utilizes high-frequency continuous observations from geostationary meteorological satellites within a short timescale to construct a spatiotemporal sequence input containing multiple time points and multiple spectral channels to explicitly characterize the advection movement and deformation evolution of cloud systems; in terms of model structure, through three-dimensional spatiotemporal feature encoding and multi-scale temporal aggregation mechanisms, effective compression and cross-scale fusion of temporal information are achieved, thereby reducing computational complexity while retaining key evolutionary clues; in the decoding and reconstruction stage, a deformable alignment feature fusion strategy is introduced to adaptively align and fuse multi-scale jump connection features, improving the expressive ability of boundary details and non-rigid cloud structures; simultaneously, by constructing auxiliary constraints or auxiliary tasks related to physical radiation consistency, physical consistency guidance is applied to the prediction results during training to suppress the accumulation of errors caused by label noise and spatiotemporal mismatch, ultimately achieving high-precision inversion of cloud parameters. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating a satellite cloud cover inversion method based on spatiotemporal physics joint driving provided in this application embodiment.
[0023] Figure 2 This is a schematic diagram of the spatiotemporal hybrid feature extraction network structure provided in the embodiments of this application.
[0024] Figure 3 This is a schematic diagram of the structure of the three-dimensional spatiotemporal encoder provided in the embodiments of this application.
[0025] Figure 4 This is a schematic diagram of the spatiotemporal attention gating module provided in the embodiments of this application.
[0026] Figure 5This is a schematic diagram of the structure of the depth 3D separable convolution module provided in the embodiments of this application.
[0027] Figure 6 This is a schematic diagram of the structure of the 3D max pooling module provided in the embodiments of this application.
[0028] Figure 7 This is a schematic diagram of the structure of the multi-scale time-series aggregation bottleneck module provided in the embodiments of this application.
[0029] Figure 8 This is a schematic diagram of the structure of the time-series fusion dimensionality reduction module with skip connections provided in the embodiments of this application.
[0030] Figure 9 This is a schematic diagram of the structure of the two-dimensional decoder provided in the embodiment of this application.
[0031] Figure 10 This is a comparison chart of the inversion results for typical difficult scenarios in the test set.
[0032] Figure 11 This is a schematic diagram of a satellite cloud inversion device based on spatiotemporal physics joint drive provided in an embodiment of this application. Detailed Implementation
[0033] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0034] The following will describe in detail, with reference to the accompanying drawings, a satellite cloud cover inversion method and apparatus based on spatiotemporal physics joint drive according to embodiments of this application.
[0035] As mentioned above, existing deep learning methods for cloud cover inversion suffer from several drawbacks, including significant biases in cloud cover estimation during transition zones, a marked decrease in accuracy and high-frequency detail fidelity at cloud edges or in areas of rapid advection and intense convection, and further amplification of inversion errors due to non-rigid migration. In other words, existing cloud cover inversion techniques are limited by the disruption of temporal continuity and the inherent physical observation parallax across sensors, often resulting in overly smoothed transition zones, severe ghosting at cloud edges, and high absolute radiometric errors.
[0036] Specifically, traditional physical methods suffer from limited resolution, jagged boundaries, and complete neglect of sub-pixel cloud information, resulting in a severe lack of spatial representation capabilities and leading to large biases in cloud amount estimation in transition zones.
[0037] Deep learning methods often employ single-time-point two-dimensional convolutions or brute-force channel stacking of multiple time-series images, lacking explicit modeling of cloud advection, deformation, and formation / dissipation trajectories. When faced with complex weather systems, the models cannot effectively utilize the temporal spectral features of high-frequency observations from geostationary satellites, resulting in a significant decrease in inversion accuracy and high-frequency detail fidelity at cloud edges, in areas of rapid advection, and in regions of intense convection.
[0038] Meanwhile, there is an unavoidable spatial matching error between geostationary satellite data and polar-orbiting satellite cloud cover labels—parallax and time difference problems caused by differences in observation perspective, transit time difference, and resampling mechanism, which cause complex non-rigid offsets of pixels at the same geographical location in different satellite observations, resulting in background confusion and feature misalignment, further amplifying the inversion error.
[0039] In view of this, this application provides a satellite cloud cover inversion method based on spatiotemporal physics joint driving. This method makes full use of the high temporal resolution advantage of geostationary satellites, explicitly extracts cloud movement trajectories and performs non-rigid spatial alignment, so as to effectively integrate spatiotemporal feature information of multiple time, multiple channels and multiple scales, and realize high-precision spatiotemporal resolution spatiotemporal inversion of continuous cloud cover fields. This provides more accurate and reliable continuous cloud cover field data support for downstream applications such as refined weather forecasting, climate monitoring, atmospheric radiation transfer simulation and solar energy resource assessment.
[0040] Figure 1 This is a flowchart illustrating a satellite cloud cover retrieval method based on spatiotemporal physics joint driving, as provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps:
[0041] In step S101, the polar-orbiting satellite cloud product data is preprocessed to obtain historical cloud cover data.
[0042] In step S102, historical observation data of meteorological satellites and cloud product data of polar-orbiting satellites are registered based on the observation time, and a spatiotemporal slice sample set is constructed based on the registered historical observation data of meteorological satellites and historical cloud cover data of polar-orbiting satellites.
[0043] Among them, the spatiotemporal slice sample set uses time-series spectral feature sequences constructed based on historical meteorological satellite observation data as samples, and uses historical cloud cover data corresponding to each feature as sample labels.
[0044] In step S103, a pre-trained spatiotemporal hybrid feature extraction network is constructed, and the pre-trained spatiotemporal hybrid feature extraction network is trained using a spatiotemporal slice sample set to obtain the trained spatiotemporal hybrid feature extraction network.
[0045] In step S104, real-time observation data from meteorological satellites are acquired, and satellite cloud cover is obtained by inverting the real-time observation data using a trained spatiotemporal hybrid feature extraction network.
[0046] The pre-trained spatiotemporal hybrid feature extraction network includes a 3D spatiotemporal encoder, a multi-scale temporal aggregation bottleneck module, a skip-connected temporal fusion dimensionality reduction module, a fusion deformable convolution-aligned 2D decoder, and a dual-branch multi-task output head.
[0047] In some embodiments of this application, the method may be executed by a server or by a terminal device with certain processing capabilities.
[0048] In some embodiments of this application, the polar-orbiting satellite cloud product data can be preprocessed first to obtain historical cloud cover data. Then, the historical meteorological satellite observation data and the polar-orbiting satellite cloud product data are registered based on the observation time, and a spatiotemporal slice sample set is constructed based on the registered historical meteorological satellite observation data and historical polar-orbiting satellite cloud cover data.
[0049] Furthermore, a pre-trained spatiotemporal hybrid feature extraction network can be constructed, and the pre-trained spatiotemporal hybrid feature extraction network can be trained using the constructed spatiotemporal slice sample set to obtain the trained spatiotemporal hybrid feature extraction network.
[0050] Finally, real-time meteorological satellite observation data can be obtained, and the satellite cloud cover can be retrieved based on the real-time observation data using a trained spatiotemporal hybrid feature extraction network.
[0051] According to the technical solution provided in the embodiments of this application, high-frequency continuous observations of geostationary meteorological satellites within a short time scale are used to construct a spatiotemporal sequence input containing multiple time points and multiple spectral channels to explicitly characterize the advection movement and deformation evolution of cloud systems. In terms of model structure, through three-dimensional spatiotemporal feature encoding and multi-scale temporal aggregation mechanism, effective compression and cross-scale fusion of time dimension information are achieved, thereby reducing computational complexity while retaining key evolutionary clues. In the decoding and reconstruction stage, a deformable alignment feature fusion strategy is introduced to adaptively align and fuse multi-scale jump connection features, improving the expressive ability of boundary details and non-rigid cloud structures, and achieving high-precision inversion of cloud parameters.
[0052] In some embodiments of this application, preprocessing of polar-orbiting satellite cloud product data may include:
[0053] Polar-orbiting satellite cloud product data is converted into a continuous numerical field based on preset weights;
[0054] Parallax correction is performed on the continuous numerical field by combining synchronously observed cloud top height data;
[0055] A continuous field resampling strategy based on spatial interpolation or kernel weighting is used to reconstruct the disparity-corrected continuous numerical field;
[0056] Multi-scale aggregation of the reconstructed continuous numerical field yields historical cloud cover data.
[0057] In other words, after reading cloud product data represented by discrete cloud detection masks, this embodiment of the application abandons the traditional binary label classification method and instead converts it into a continuous numerical field. This numerical field can be, for example, a floating-point numerical field.
[0058] During the transformation, considering the typically "conservative" nature of cloud product data, the weight of certain cloud pixels is set to 1.0, the weight of possible cloud pixels to 0.3, the weight of possible clear-sky pixels to 0.05, and the weight of certain clear-sky pixels to 0.0. This constructs an initial cloud cover probability distribution field, preserving sub-pixel-level cloud cover gradient information for subsequent processing and avoiding the jagged edges caused by traditional binarization. The specific weights of different pixel types can be adjusted according to actual conditions; no restrictions are imposed here.
[0059] Based on the obtained continuous cloud cover field, parallax correction can be further performed by combining synchronously observed cloud top height data. In this embodiment, the geometric offset vector from the actual geographical location to the apparent satellite position is calculated pixel by pixel according to the satellite observation geometry, the Earth's curvature model, and the cloud top height. Based on this, the latitude and longitude coordinates of the cloud pixels are remapped, transforming the cloud cover field from the geographic coordinate system to the satellite apparent coordinate system. This effectively eliminates cloud position projection errors caused by differences in cloud height and satellite observation perspective, achieving spatial alignment between polar-orbiting satellite data and geostationary satellite perspective, thereby improving the registration accuracy of multi-source satellite data fusion.
[0060] To address the spatial stretching and data holes caused by non-uniform pixel displacement during parallax correction, embodiments of this application employ a continuous field resampling strategy based on spatial interpolation or kernel weighting for reconstruction. Spatial interpolation can be, for example, bilinear interpolation, bicubic interpolation, inverse distance weighted interpolation, and radial basis function interpolation, while kernel weighting can be, for example, Gaussian weighting.
[0061] Taking Gaussian weighting as an example, a high-resolution intermediate grid can be defined, and the parallax-corrected cloud pixels can be projected onto this intermediate grid using the Gaussian kernel function as a spatial weighting factor. During this process, the smooth diffusion characteristics of the Gaussian kernel are utilized to allow cloud cover values to naturally cover the pixel gaps caused by parallax stretching. For areas with severe stretching, a transitional value between 0.0 and 1.0 is automatically generated, thus achieving adaptive hole filling while ensuring the continuity of the physical field.
[0062] Finally, the reconstructed high-resolution cloud cover field can be subjected to multi-scale aggregation processing. In this embodiment, the high-resolution intermediate grid data is downsampled to the target resolution by performing average pooling on the cloud cover values within a local window, generating preprocessed historical cloud cover data.
[0063] In some embodiments of this application, the spatiotemporal slice sample set can be constructed in the following manner:
[0064] Spatiotemporal registration and quality screening are performed based on the observation time difference between meteorological satellites and polar-orbiting satellites; among them, the observation time difference between historical observation data of meteorological satellites and cloud product data of polar-orbiting satellites after spatiotemporal registration and quality screening is less than a preset time threshold.
[0065] In time Based on, backtracking A time-series spectral feature sequence was constructed using historical meteorological satellite observation data at consecutive moments. It is a positive integer greater than 1;
[0066] The initial sample is obtained by spatially aligning each feature in the time-series spectral feature sequence with historical cloud cover data.
[0067] The initial samples are cropped and sampled using a sliding window to obtain a spatiotemporal slice sample set that balances sample coverage with local spatial context information.
[0068] In other words, spatiotemporal registration and quality screening of two types of observation data can be performed based on the full-disk horizontal scan observation time of geostationary meteorological satellites and the orbital transit scan time of polar-orbiting satellites. Specifically, the spatial regions where the time difference between the two observations exceeds a preset time threshold are eliminated, using the time difference as a criterion to reduce label noise introduced by rapid cloud evolution and perspective differences, thereby improving the consistency and reliability of the samples. The specific value of the preset time threshold can be set according to actual needs; in some examples, it can be set to 5 minutes.
[0069] Considering meteorological satellites at a preset cycle To characterize the advection and deformation features of cloud systems over short timescales through continuous observation, this application's embodiments use the current time as the starting point. Based on the baseline, backtracking is used to select... , , , A temporal spectral feature sequence is constructed using observation data from five consecutive time points. Each time point contains 15 multispectral channel features, thus forming a spatiotemporal spectral tensor sequence for network input. In practical applications, the number of observation time points can be adjusted as needed; no limitation is imposed here.
[0070] During the sample construction phase, the 15-channel multispectral features from the five consecutive time points mentioned above can be spatially aligned with the corresponding historical cloud cover data labels from polar-orbiting satellites. These are then cropped into 256×256 pixel image patches using a fixed window, and a sliding window sampling method with a step size of 128 pixels is employed to balance sample coverage and local spatial context information. This results in a spatiotemporal slice dataset suitable for deep learning training and inference. The specific values of the image patch size and the sliding window sampling step size can be adjusted according to actual needs and are not limited here.
[0071] In some embodiments of this application, a pre-trained spatiotemporal hybrid feature extraction network incorporating physical radiation consistency constraints can be constructed. The input is a multispectral spatiotemporal tensor of a geostationary meteorological satellite. Taking the sample constructed above as an example, its dimensions are B×15×5×H×W (representing batch size, 15 spectral channels, 5 consecutive time steps, spatial height, and width, respectively), and the output is a two-dimensional continuous cloud cover field B×1×H×W. During the training phase, a physical radiation reconstruction field can be output simultaneously for physical consistency constraints. This network includes a three-dimensional spatiotemporal encoder, a multi-scale temporal aggregation bottleneck module, a skip-connected temporal fusion dimensionality reduction module, a two-dimensional decoder with fused deformable convolution alignment, and a dual-branch multi-task output head, as shown in the following structure. Figure 2 As shown.
[0072] refer to Figure 2 The pre-trained spatiotemporal hybrid feature extraction network may include an input module for inputting samples containing time steps T; a three-level encoder and a three-level decoder, wherein each encoder and its corresponding decoder are connected by a skip connection to a temporal fusion dimensionality reduction module, and the output of the last encoder is connected to the decoder through a multi-scale temporal aggregation bottleneck module. The output of the decoder is connected to a dual-branch multi-task output head, which includes a cloud cover output head and a radiation output head.
[0073] In some embodiments of this application, the three-dimensional spatiotemporal encoder samples samples from the spatiotemporal slice sample set through a three-level encoder; wherein, the sampling strategy of each level encoder is that the time dimension is not downsampled and the spatial dimension is downsampled.
[0074] Each encoder provided in this application includes a feature extraction module, a spatiotemporal attention gating module, and a pooling module.
[0075] The feature extraction module is used to extract features.
[0076] The spatiotemporal attention gating module generates channel attention weights and spatial attention weights based on the statistical responses of the current-level features in the channel and spatial dimensions, respectively. These weights are then used to element-wise weight the output features of the current-level convolutional module, resulting in adaptively weighted features. The channel and spatial attention weights range from 0 to 1.
[0077] The pooling module is used to perform spatial dimension downsampling on the adaptively weighted features at this level.
[0078] In some examples, the number of feature channels extracted by the first-level encoder is less than the number of feature channels extracted by the second-level encoder, and the number of feature channels extracted by the second-level encoder is less than the number of feature channels extracted by the third-level encoder; furthermore, the feature size after spatial dimension downsampling by the first-level encoder is greater than the feature size after spatial dimension downsampling by the second-level encoder, and the feature size after spatial dimension downsampling by the second-level encoder is greater than the feature size after spatial dimension downsampling by the third-level encoder.
[0079] Figure 3 This is a schematic diagram of the structure of the three-dimensional spatiotemporal encoder provided in an embodiment of this application. Figure 3 As shown, taking the sample set constructed above as an example, the 3D spatiotemporal encoder can include three encoding stages. Each stage adopts a strategy of no downsampling in the temporal dimension and downsampling in the spatial dimension to maintain a temporal resolution of 5 while reducing the spatial resolution, thereby extracting the spatiotemporal coupling features of cloud system evolution over time. At the same time, a spatiotemporal attention gating module is set after the output of each encoding stage to perform joint adaptive weighting of the channel dimension and the spatial dimension to highlight significant responses related to cloud system structure.
[0080] In the first-level encoding stage: the input B×15×5×H×W is first processed by a 3D convolutional feature extraction unit (3×3×3 convolution, 3D normalization, and nonlinear activation) to map the number of channels to 32; then, it is weighted by a spatiotemporal attention gating module to obtain the first-level spatiotemporal feature E1 (B×32×5×H×W). Subsequently, the spatial dimension is downsampled by a 3D max pooling module with a pooling kernel size of 1×2×2 and a stride of 1×2×2, and the feature size becomes H / 2×W / 2.
[0081] The second-level encoding stage: a deep 3D separable convolution module (depth convolution + pointwise convolution) is used to expand the number of channels from 32 to 64. The second-level feature E2 (B×64×5×H / 2×W / 2) is obtained through the spatiotemporal attention gating module. Then, the spatial resolution is reduced to H / 4×W / 4 by the same 3D max pooling module.
[0082] The third-level encoding stage: Similarly, the number of channels is expanded from 64 to 128, and the third-level feature E3 (B×128×5×H / 4×W / 4) is obtained through the spatiotemporal attention gating module; then the spatial resolution is reduced to H / 8×W / 8 through the 3D max pooling module.
[0083] Figure 4 This is a schematic diagram of the spatiotemporal attention gating module provided in an embodiment of this application. Figure 4 As shown, the module includes a channel attention branch and a spatial attention branch. It calculates channel dimension weights and spatial dimension weights for the input features respectively. The weights are constrained to 0 to 1 by the Sigmoid normalization function and then multiplied element by element with the input features, thereby improving the response weights of highly correlated regions or channels such as cloud edges, cloud textures, and dynamically changing regions, while suppressing background or redundant channel responses.
[0084] In this system, the three-dimensional global adaptive average pooling is Global AvgPool3D, the 1×1×1 convolutional layer is 1×1×1 Conv, the 3×3×3 convolutional layer is 3×3×3 Conv, and ReLU and Sigmoid are both activation functions.
[0085] Figure 5 This is a schematic diagram of the structure of the depth 3D separable convolution module provided in the embodiments of this application. It includes a 3D channel-wise convolution (Dev3D) with a size of 3×3×3; a 3D pointwise convolution (Dev3D) with a size of 1×1×1; and a Norm normalization function and a ReLU activation function.
[0086] Figure 6 This is a schematic diagram of the structure of the 3D max pooling module provided in an embodiment of this application. The 3D max pooling module is referred to as MaxPool3D.
[0087] In some embodiments of this application, the multi-scale temporal aggregation bottleneck module includes a first temporal aggregation branch and a second temporal aggregation branch;
[0088] The first temporal aggregation branch uses a preset time kernel and a first preset spatial kernel to perform temporal aggregation and channel compression on the output signal of the three-dimensional spatiotemporal encoder, and removes the time dimension of the compressed features to obtain the two-dimensional features of the first branch.
[0089] The second temporal aggregation branch uses a preset temporal kernel and a second preset spatial kernel to perform temporal aggregation and channel compression on the output signal of the three-dimensional spatiotemporal encoder, and removes the temporal dimension of the compressed features to obtain the two-dimensional features of the second branch.
[0090] The preset time kernel covers all time steps of the time series spectral feature sequence, and the scale of the first preset spatial kernel is different from that of the second preset spatial kernel;
[0091] The multi-scale temporal aggregation bottleneck module also includes a splicing module, which is used to fuse the two-dimensional features of the first branch and the two-dimensional features of the second branch to obtain two-dimensional bottleneck fused features.
[0092] In other words, the deepest output of the encoder undergoes semantic enhancement via 3D depthwise separable convolution, increasing the number of channels from 128 to 256, resulting in the bottleneck spatiotemporal feature B3D (B×256×5×H / 8×W / 8). To achieve effective 3D-to-2D conversion and enhance the ability to characterize spatiotemporal relationships at different scales, this embodiment employs two parallel temporal aggregation branches to perform full-coverage aggregation of the time dimension. The structure of this multi-scale temporal aggregation bottleneck module is as follows: Figure 7 As shown.
[0093] Branch 1 uses a 3D convolution with a temporal kernel covering 5 time steps and a spatial kernel of 1×1 to perform temporal aggregation and channel compression, outputting B×128×1×H / 8×W / 8. After removing the time dimension of length 1, a 2D feature B1 (B×128×H / 8×W / 8) is obtained.
[0094] Branch 2 uses a 3D convolution with a temporal kernel covering 5 time steps and a spatial kernel of 3×3 (with boundary filling in the spatial direction) to perform spatiotemporal joint aggregation. After removing the temporal dimension, a 2D feature B2 (B×128×H / 8×W / 8) is obtained.
[0095] By concatenating B1 and B2 in the channel dimension, we can obtain the two-dimensional bottleneck fusion feature B (B×256×H / 8×W / 8), which is then fused and shaped by two-dimensional normalization and nonlinear activation.
[0096] In some embodiments of this application, the skip-connected temporal fusion dimensionality reduction module is skip-connected to a three-level encoder to convert the three-dimensional features output by each level encoder into two-dimensional skip-connected features.
[0097] The process of converting the target's 3D features output by the target encoder into 2D skip connection features includes:
[0098] Jointly expand the temporal and channel dimensions of the target's three-dimensional features;
[0099] The target 3D features after joint unfolding are sequentially subjected to 2D convolutional feature extraction, 2D normalization and nonlinear activation to obtain the 2D skip connection features corresponding to the target encoder.
[0100] The target encoder can be any level encoder, and the target 3D feature is the 3D feature output by the target encoder.
[0101] Furthermore, the fused deformable convolution-aligned two-dimensional decoder performs feature decoding through a three-level decoder to obtain high-resolution decoded features;
[0102] The first-level decoder performs feature decoding in the following manner:
[0103] The two-dimensional bottleneck fusion features are upsampled, and the upsampled features are fused with the two-dimensional skip connection features corresponding to the third-level encoder to obtain the first concatenation feature.
[0104] The first splicing feature is extracted by deformable convolution alignment feature extraction to obtain the first alignment feature;
[0105] The first alignment feature is input into the feature refining module to obtain the first-level decoding output feature;
[0106] The second-level decoder performs feature decoding in the following manner:
[0107] The first-level decoded output features are upsampled, and the upsampled features are fused with the two-dimensional skip connection features corresponding to the second-level encoder to obtain the second concatenation features.
[0108] The second splicing feature is extracted by deformable convolution alignment feature extraction to obtain the second alignment feature;
[0109] The second alignment feature is input into the feature refining module to obtain the second-level decoding output feature;
[0110] The third-level decoder performs feature decoding in the following manner:
[0111] The second-level decoded output features are upsampled, and the upsampled features are fused with the two-dimensional skip connection features corresponding to the first-level encoder to obtain the third concatenation features.
[0112] The third splicing feature is extracted by deformable convolution alignment feature extraction to obtain the third alignment feature;
[0113] The third alignment feature is input into the feature refining module to obtain the decoded output feature;
[0114] The feature refinement module is used to reshape and compensate for details of the input features.
[0115] Simultaneously, deformable convolutional alignment feature extraction of the target splicing features can include:
[0116] Offset prediction is performed using N concatenated feature extraction layers to obtain the target two-dimensional offset and modulation mask corresponding to the target splicing features; where the target splicing features are any one of the first splicing features, the second splicing features, and the third splicing features; the objective function for offset prediction is to maximize the effective receptive field, where N is a positive integer greater than 1;
[0117] The target 2D offset and modulation mask are used to drive deformable convolutional layers to align and extract target stitching features, resulting in initial target alignment features.
[0118] The initial alignment features of the target are subjected to two-dimensional normalization and nonlinear activation processing, and the processed initial alignment features are added to the target splicing features to obtain the target alignment features.
[0119] In other words, since the encoder outputs three-dimensional spatiotemporal features while the decoder adopts a two-dimensional structure, this embodiment sets a temporal fusion dimensionality reduction module at each skip connection to convert the three-dimensional features of the corresponding scale into two-dimensional features while preserving cloud movement and deformation clues as much as possible. The structure of this temporal fusion dimensionality reduction module is shown in Figure 8.
[0120] refer to Figure 8 For any three-dimensional feature (B×C×5×h×w) at any scale, the time dimension and channel dimension can be jointly expanded to obtain B×(5·C)×h×w, and then adaptively fused by 3×3 two-dimensional convolution, two-dimensional normalization and nonlinear activation to output two-dimensional skip connection feature B×C×h×w.
[0121] Therefore, after temporal fusion and dimensionality reduction, E1 becomes E1_2D (B×32×H×W), E2 becomes E2_2D (B×64×H / 2×W / 2), and E3 becomes E3_2D (B×128×H / 4×W / 4).
[0122] The 2D decoder comprises three decoding stages, each executed sequentially: transposed convolutional upsampling (restoring spatial resolution and reducing the number of channels), channel concatenation with 2D skip connection features of the corresponding scale, deformable convolutional alignment feature extraction, and two-layer convolutional refinement and fusion, to progressively restore a high-resolution cloud cover representation. Its structure is as follows: Figure 9 As shown.
[0123] refer to Figure 9 In the first-level decoding stage, the two-dimensional bottleneck feature B (B×256×H / 8×W / 8) is first upsampled by a 2×2 transposed convolution, increasing the spatial resolution to H / 4×W / 4, while reducing the number of channels from 256 to 128, resulting in U3 (B×128×H / 4×W / 4). U3 is then concatenated with E3_2D in the channel dimension to obtain the concatenated feature C3 (B×256×H / 4×W / 4).
[0124] The module then proceeds to the deformable convolution alignment feature extraction module. This module first predicts the 2D offset and modulation mask required for deformable sampling using two concatenated 3×3 convolutions, giving the offset prediction a larger effective receptive field. The modulation mask is normalized using Sigmoid to constrain its value range. The offset and modulation mask are then used to drive a 3×3 deformable convolution to align and extract features from C3, outputting 128-channel features, which are then normalized and non-linearly activated. Simultaneously, the input features are matched with channels through a 1×1 convolution via the residual path and added to the main branch output to enhance gradient propagation and training stability. This module outputs A3 (B×128×H / 4×W / 4).
[0125] Finally, A3 enters the two-layer convolutional refinement and fusion module (two consecutive 3×3 two-dimensional convolutions combined with normalization and non-linear activation) to further shape and compensate for details, resulting in the first-level decoding output D3 (B×128×H / 4×W / 4).
[0126] In the second-level decoding stage, D3 is first upsampled to H / 2×W / 2 via a 2×2 transposed convolution, reducing the number of channels from 128 to 64, resulting in U2 (B×64×H / 2×W / 2). Then, U2 is concatenated with E2_2D to obtain C2 (B×128×H / 2×W / 2). C2 is then processed by a deformable convolution alignment feature extraction module to output a 64-channel aligned feature A2 (B×64×H / 2×W / 2), which is then refined and fused by a two-layer convolution to obtain the second-level decoding output D2 (B×64×H / 2×W / 2).
[0127] In the third-level decoding stage, D2 is first upsampled to H×W via a 2×2 transposed convolution, reducing the number of channels from 64 to 32, resulting in U1 (B×32×H×W). Then, U1 is concatenated with E1_2D to obtain C1 (B×64×H×W). C1 is then processed by a deformable convolutional alignment feature extraction module to output a 32-channel aligned feature A1 (B×32×H×W), which is then refined and fused by a two-layer convolutional module to obtain the third-level decoding output D1 (B×32×H×W).
[0128] In some embodiments of this application, the dual-branch multi-task output head includes a main task output branch and an auxiliary task output branch. The main task output branch is used to determine the two-dimensional continuous cloud cover prediction field based on the decoded output features; the auxiliary task output branch is used to determine the physical radiation reconstruction field based on the decoded output features during the model training phase.
[0129] In other words, the high-resolution feature D1 output by the decoder is fed into two output branches in parallel:
[0130] The main task output branch uses a 1×1 two-dimensional convolution to map the number of channels from 32 to 1, and outputs a two-dimensional continuous cloud cover prediction field B×1×H×W.
[0131] The auxiliary task output branch uses a 1×1 two-dimensional convolution to map the number of channels from 32 to 15, outputting a physical radiation reconstruction field B×15×H×W. During the training phase, the consistency constraint between the auxiliary branch and the input observations is used as an auxiliary supervision signal to guide the network to learn a spatiotemporal representation that satisfies physical consistency. During the inference phase, only the main task branch can be output to obtain the cloud cover prediction result.
[0132] This application presents an embodiment of a spatiotemporal feature extraction and adaptive gating architecture based on continuous meteorological sequences. It constructs a multi-time, multispectral spatiotemporal input tensor for high-frequency continuous observation sequences from geostationary meteorological satellites. A three-dimensional spatiotemporal encoder is used to extract the dynamic evolution features of cloud clusters during advection and deformation processes at each level. A spatiotemporal joint attention gating module is embedded at the output of each encoding level to jointly and adaptively weight the channel dimension and spatial dimension information, thereby enhancing the network's ability to represent cloud cluster edges and dynamically changing regions and suppressing redundant background interference.
[0133] This application establishes a multi-scale temporal aggregation dimensionality reduction mechanism for 3D feature-to-2D conversion. Parallel temporal aggregation branches are set in the deepest layer of the encoder, containing two aggregation paths that fully cover the temporal dimension but have different convolutional kernel sizes in the spatial dimension. These paths are used to capture the absolute numerical changes of pixels over time and the spatial translation relationships between adjacent pixels, respectively. After channel-level fusion, the outputs of each branch reduce the 3D spatiotemporal features to 2D spatial features, thereby preserving spatiotemporal correlation clues at different scales while completing temporal compression. This provides an informative and computationally controllable intermediate representation for subsequent 2D decoding and reconstruction.
[0134] This application proposes a temporally guided deformable decoding and reconstruction strategy to eliminate cross-source parallax. After multi-scale skip connections and fusion at each level in the network decoding stage, deformable alignment feature extraction units are connected in series. This method utilizes the cloud motion prior features retained after feature dimensionality reduction and aggregation to drive deformable convolution operators to calculate spatial offsets. It performs adaptive spatial resampling and feature compensation for non-rigid deformation of cloud targets and pixel-level misalignment caused by differences in cross-sensor observation perspectives, effectively eliminating edge ghosting and feature mismatch under large-scale advection, and achieving high-fidelity continuous field reconstruction of cloud parameters.
[0135] To verify the technical effects of the technical solutions provided in the embodiments of this application, the following experiment was designed:
[0136] Existing cloud cover inversion techniques are limited by the fragmentation of temporal continuity and the inherent physical observation parallax across sensors. Inversion results often exhibit excessive smoothing in transition zones, severe ghosting at cloud edges, and high absolute radiometric errors. This patent, by introducing explicit three-dimensional spatiotemporal evolution feature extraction and a deformable spatial alignment mechanism guided by motion priors, effectively endows geostationary satellite models with non-rigid parallax compensation and hydrodynamic sensing capabilities. To verify the effectiveness and generalization ability of the spatiotemporal joint inversion model in this application, an experimental environment based on the PyTorch deep learning framework was constructed, with the computing platform configured as a single NVIDIA GeForce RTX 4090 graphics processing unit (GPU). The experimental dataset was randomly divided according to a preset ratio, with the training set, validation set, and test set set at 70%:15%:15%. The model input consisted of five consecutive frames of multispectral temporal feature data from geostationary meteorological satellites, including the time dimension, and the output was a continuous cloud cover field (0~100%), labeled with high-precision polar-orbiting satellite cloud products.
[0137] During the model training phase, the AdamW optimizer is used in conjunction with a cosine annealing learning rate scheduling algorithm for iterative parameter updates. To address the edge blurring and pixel-level misalignment issues that are prone to occur in continuous variable regression tasks, this application proposes a multi-task joint loss function optimization mechanism that integrates physical constraints and high-frequency sensing. First, in the backbone cloud inversion task, a joint constraint of Smooth L1 Loss and Multi-Scale SSIM (MS-SSIM) is employed to simultaneously minimize pixel-level absolute radiometric errors and enhance the consistency between cloud texture and high-frequency edge structures at multiple scales. Simultaneously, a multi-task collaborative output architecture is introduced at the decoding end to construct a physical radiometric consistency constraint auxiliary task based on satellite multi-channel observation data. The loss function is designed as follows: ;in, , , These are the smoothing L1 loss, MS-SSIM loss, and physical radiation uniformity loss, respectively. , , These are the dynamic weighting coefficients for each loss.
[0138] During the parameter iteration process of model training, this embodiment adopts a dynamic weight scheduling strategy for the main task and auxiliary task: the initial weight coefficient of the physical radiation consistency auxiliary task is set to be greater than the MS-SSIM weight coefficient in the main task. As the training iteration cycle increases, the weight coefficient of the physical auxiliary task is calculated and reduced cycle by cycle using a preset decay function, while the weight coefficient of MS-SSIM loss is increased simultaneously. Thus, in the early stage of training, the radiation reconstruction loss dominates the initial spatial mapping of network parameters, and in the later stage of training, it smoothly transitions to the high-frequency structural similarity loss dominating model optimization.
[0139] For evaluation metrics, Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Structural Similarity Index (SSIM) were selected as core metrics. Training monitoring results show that the model tends to stabilize and converge after the 108th epoch. On the independent test set, compared with the baseline model with channel stacking, the MAE and RMSE of the model in this application embodiment are further reduced, while the SSIM is significantly improved. Specifically, the quantitative comparison shows that the baseline MAE is 0.069, RMSE is 0.147, and SSIM is 0.727; while the MAE of the model in this application embodiment decreases to 0.058 (a decrease of 15.9%), the RMSE decreases to 0.124 (a decrease of 15.6%), and the SSIM increases to 0.778 (an increase of 7.0%).
[0140] To further examine the performance in complex regions, this application introduces segmented hierarchical cloud cover assessment and high-frequency mask assessment at cloud edges. Results show that in regions significantly affected by cross-sensor spatiotemporal mismatches, such as cloud edges, the error level of the model in this application is further reduced. Specifically, the RMSE in the edge region significantly decreases from 0.234 in the baseline model to 0.201 (a decrease of 14.1%), indicating that the method provided in this application can improve the inversion robustness in complex boundary regions while maintaining global accuracy.
[0141] To further verify the inversion robustness of the embodiments of this application under various complex meteorological and observation scenarios, Figure 10The comparison of inversion results for typical difficult scenarios in the test set is shown. Rows and columns from left to right are: multispectral near-true color composite image RGB (3-2-1), polar-orbiting satellite real cloud cover label, prediction results of the temporal-physical 3D-UNet model of this application embodiment, prediction results of the existing channel stacked baseline model Stacked 2D U-Net, and error difference map between the model of this application embodiment and the baseline model (black boxes in the figures indicate areas of significant error). Here, RGB refers to red, green, and blue; (3-2-1) refers to the near-color composite based on the mapping of the 3rd, 2nd, and 1st channels of the satellite payload to the red, green, and blue bands respectively; TempPhy 3DU-N indicates that the temporal-physical 3D-UNet model provided in this application embodiment can be represented by TempPhy 3D U-Net, the channel stacked baseline model can be represented by Stacked 2D U-Net, and the error difference map between the model of this application embodiment and the baseline model can be represented by Diff(3D-2D).
[0142] Figure 10 Five representative complex scenarios were selected from top to bottom: small-scale sparse cloud areas, thin cloud transition areas at the land-sea interface, rapidly advection storm areas, complex surface areas, and cross-sensor observation areas with large parallax. Figure 10 The third and fourth columns mark the MAE, RMSE, and SSIM values for TempPhy 3D U-Net and Stacked 2D U-Net, respectively. The color bars on the right of the second to fourth columns represent cloud fraction, the proportion of the current pixel covered by clouds, ranging from 0 to 1, where 0 represents clear skies and 1 represents full cloud cover. The color bar on the right of the fifth column represents the difference in cloud fraction inversion results, which is the value in the third column minus the value in the fourth column, used to highlight the improvement in inversion performance of the model designed in this embodiment compared to the baseline model. As can be seen from the figure, the temporal-physical 3D-UNet architecture of this embodiment exhibits lower cloud fraction inversion errors compared to the baseline model in the aforementioned scenarios. This architecture effectively captures spatiotemporal and parallax features and successfully suppresses the obfuscation of static surface features, achieving high-precision reconstruction of continuous cloud fraction in both spatial location and physical values.
[0143] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0144] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0145] Figure 11This is a schematic diagram of a satellite cloud cover inversion device based on spatiotemporal physics joint driving, provided in an embodiment of this application. Figure 11 As shown, the device includes:
[0146] The cloud product data preprocessing module 1101 is configured to preprocess polar-orbiting satellite cloud product data to obtain historical cloud volume data;
[0147] The dataset construction module 1102 is configured to register historical meteorological satellite observation data and polar-orbiting satellite cloud product data based on observation time, and construct a spatiotemporal slice sample set based on the registered historical meteorological satellite observation data and historical cloud cover data of polar-orbiting satellites. The spatiotemporal slice sample set uses time-series spectral feature sequences constructed based on historical meteorological satellite observation data as samples, and uses the historical cloud cover data at the corresponding time of each feature as sample labels.
[0148] The model training module 1103 is configured to construct a pre-trained spatiotemporal hybrid feature extraction network, train the pre-trained spatiotemporal hybrid feature extraction network using a spatiotemporal slice sample set, and obtain the trained spatiotemporal hybrid feature extraction network.
[0149] Inversion module 1104 is configured to acquire real-time observation data from meteorological satellites and use a trained spatiotemporal hybrid feature extraction network to invert satellite cloud cover based on the real-time observation data.
[0150] The pre-trained spatiotemporal hybrid feature extraction network includes a 3D spatiotemporal encoder, a multi-scale temporal aggregation bottleneck module, a skip-connected temporal fusion dimensionality reduction module, a fusion deformable convolution-aligned 2D decoder, and a dual-branch multi-task output head.
[0151] According to the technical solution provided in the embodiments of this application, high-frequency continuous observations of geostationary meteorological satellites within a short time scale are used to construct a spatiotemporal sequence input containing multiple time points and multiple spectral channels to explicitly characterize the advection movement and deformation evolution of cloud systems. In terms of model structure, through three-dimensional spatiotemporal feature encoding and multi-scale temporal aggregation mechanism, effective compression and cross-scale fusion of time dimension information are achieved, thereby reducing computational complexity while retaining key evolutionary clues. In the decoding and reconstruction stage, a deformable alignment feature fusion strategy is introduced to adaptively align and fuse multi-scale jump connection features, improving the expressive ability of boundary details and non-rigid cloud structures, and achieving high-precision inversion of cloud parameters.
[0152] In some implementations, the polar-orbiting satellite cloud product data is preprocessed, including: converting the polar-orbiting satellite cloud product data into a continuous numerical field based on preset weights; performing parallax correction on the continuous numerical field by combining synchronously observed cloud top height data; reconstructing the parallax-corrected continuous numerical field using a continuous field resampling strategy based on spatial interpolation or kernel weighting; and performing multi-scale aggregation on the reconstructed continuous numerical field to obtain historical cloud cover data.
[0153] In some implementations, the spatiotemporal slice sample set is constructed as follows: spatiotemporal registration and quality screening are performed based on the observation time difference between meteorological satellites and polar-orbiting satellites; wherein, the observation time difference between the historical observation data of meteorological satellites and the cloud product data of polar-orbiting satellites after spatiotemporal registration and quality screening is less than a preset time threshold; based on time... Based on, backtracking A time-series spectral feature sequence was constructed using historical meteorological satellite observation data from consecutive moments. The value is a positive integer greater than 1; each feature in the time-series spectral feature sequence is spatially aligned with historical cloud cover data to obtain an initial sample; the initial sample is cropped and subjected to sliding window sampling to obtain a spatiotemporal slice sample set that takes into account both sample coverage and local spatial context information.
[0154] In some implementations, the 3D spatiotemporal encoder samples samples from a spatiotemporal slice sample set using a three-stage encoder. Each stage of the encoder employs a sampling strategy of not downsampling in the temporal dimension but downsampling in the spatial dimension. Each stage includes a feature extraction module, a spatiotemporal attention gating module, and a pooling module. The feature extraction module performs feature extraction. The spatiotemporal attention gating module generates channel attention weights and spatial attention weights based on the statistical responses of the features at this stage in the channel and spatial dimensions, respectively. It then uses these channel and spatial attention weights to element-wise weight the output features of the convolutional module at this stage, obtaining an adaptive weighting. The weighted features are defined as follows: the channel attention weight and spatial attention weight are defined as follows, with values ranging from 0 to 1; the pooling module is used to perform spatial dimension downsampling on the adaptively weighted features of this level; the number of feature channels extracted by the first-level encoder is less than the number of feature channels extracted by the second-level encoder, and the number of feature channels extracted by the second-level encoder is less than the number of feature channels extracted by the third-level encoder; furthermore, the feature size after spatial dimension downsampling by the first-level encoder is greater than the feature size after spatial dimension downsampling by the second-level encoder, and the feature size after spatial dimension downsampling by the second-level encoder is greater than the feature size after spatial dimension downsampling by the third-level encoder.
[0155] In some implementations, the multi-scale temporal aggregation bottleneck module includes a first temporal aggregation branch and a second temporal aggregation branch. The first temporal aggregation branch uses a preset time kernel and a first preset spatial kernel to perform temporal aggregation and channel compression on the output signal of the 3D spatiotemporal encoder, and removes the temporal dimension of the compressed features to obtain the first branch's two-dimensional features. The second temporal aggregation branch uses a preset time kernel and a second preset spatial kernel to perform temporal aggregation and channel compression on the output signal of the 3D spatiotemporal encoder, and removes the temporal dimension of the compressed features to obtain the second branch's two-dimensional features. The preset time kernel covers all time steps of the temporal spectral feature sequence, and the scale of the first preset spatial kernel is different from that of the second preset spatial kernel. The multi-scale temporal aggregation bottleneck module also includes a splicing module for fusing the first branch's two-dimensional features and the second branch's two-dimensional features to obtain two-dimensional bottleneck fused features.
[0156] In some implementations, the skip-connected temporal fusion dimensionality reduction module is skip-connected to a three-level encoder to convert the three-dimensional features output by each encoder into two-dimensional skip-connected features. The conversion of the target three-dimensional features output by the target encoder into two-dimensional skip-connected features includes: jointly unfolding the temporal and channel dimensions of the target three-dimensional features; sequentially performing two-dimensional convolutional feature extraction, two-dimensional normalization, and nonlinear activation processing on the jointly unfolded target three-dimensional features to obtain the two-dimensional skip-connected features corresponding to the target encoder; the target encoder is any level encoder, and the target three-dimensional features are the three-dimensional features output by the target encoder.
[0157] In some implementations, the 2D decoder with deformable convolutional alignment is used for feature decoding through a three-level decoder to obtain high-resolution decoded features. The first-level decoder performs feature decoding as follows: upsampling the 2D bottleneck fusion features, and fusing the upsampled features with the 2D skip connection features corresponding to the third-level encoder using channel-dimensional splicing to obtain the first spliced features; extracting deformable convolutional alignment features from the first spliced features to obtain the first aligned features; and inputting the first aligned features into a convolutional feature refinement module to obtain the first-level decoder output features. The second-level decoder performs feature decoding as follows: upsampling the first-level decoder output features, and fusing the upsampled features with the 2D skip connection features corresponding to the second-level encoder... The first-level encoder performs channel-dimensional concatenation feature fusion to obtain the second concatenated feature; the second concatenated feature is then subjected to deformable convolutional alignment feature extraction to obtain the second aligned feature; the second aligned feature is input into the convolutional feature refinement module to obtain the second-level decoded output feature; the third-level decoder performs feature decoding as follows: the second-level decoded output feature is upsampled, and the upsampled feature is fused with the corresponding two-dimensional skip connection feature of the first-level encoder through channel-dimensional concatenation feature fusion to obtain the third concatenated feature; the third concatenated feature is then subjected to deformable convolutional alignment feature extraction to obtain the third aligned feature; the third aligned feature is input into the convolutional feature refinement module to obtain the decoded output feature; wherein, the convolutional feature refinement module is used to shape and compensate for details of the input feature.
[0158] In some implementations, deformable convolutional alignment feature extraction is performed on the target splicing features, including: using N cascaded feature extraction layers to predict the offset, obtaining the target two-dimensional offset and modulation mask corresponding to the target splicing features; wherein, the target splicing features are any one of the first splicing features, the second splicing features, and the third splicing features; the objective function for offset prediction is to maximize the effective receptive field, and N is a positive integer greater than 1; using the target two-dimensional offset and modulation mask to drive deformable convolutional layers to align and extract features from the target splicing features, obtaining the target initial alignment features; performing two-dimensional normalization and nonlinear activation processing on the target initial alignment features, and adding the processed target initial alignment features to the target splicing features to obtain the target alignment features.
[0159] In some implementations, the dual-branch multi-task output head includes a main task output branch and an auxiliary task output branch; wherein, the main task output branch is used to determine the two-dimensional continuous cloud cover prediction field based on the decoded output features; and the auxiliary task output branch is used to determine the physical radiation reconstruction field based on the decoded output features during the model training phase.
[0160] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0161] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A satellite cloud cover inversion method based on spatiotemporal physics joint driving force, characterized in that, include: Preprocessing polar-orbiting satellite cloud product data yields historical cloud cover data; Based on the observation time, historical observation data of meteorological satellites and cloud product data of polar-orbiting satellites are registered. Based on the registered historical observation data of meteorological satellites and historical cloud cover data of polar-orbiting satellites, a spatiotemporal slice sample set is constructed. The spatiotemporal slice sample set uses time-series spectral feature sequences constructed based on historical observation data of meteorological satellites as samples, and uses the historical cloud cover data at the time corresponding to each feature as sample labels. A pre-trained spatiotemporal hybrid feature extraction network is constructed, and the pre-trained spatiotemporal hybrid feature extraction network is trained using the spatiotemporal slice sample set to obtain the trained spatiotemporal hybrid feature extraction network. Real-time meteorological satellite observation data is acquired, and satellite cloud cover is obtained by inverting the real-time observation data using the trained spatiotemporal hybrid feature extraction network. The pre-trained spatiotemporal hybrid feature extraction network includes a three-dimensional spatiotemporal encoder, a multi-scale temporal aggregation bottleneck module, a skip-connected temporal fusion dimensionality reduction module, a two-dimensional decoder with fused deformable convolution alignment, and a dual-branch multi-task output head.
2. The satellite cloud cover retrieval method based on spatiotemporal physics joint driving as described in claim 1, characterized in that, Preprocessing of polar-orbiting satellite cloud product data includes: Polar-orbiting satellite cloud product data is converted into a continuous numerical field based on preset weights; Parallax correction is performed on the continuous numerical field by combining synchronously observed cloud top height data; A continuous field resampling strategy based on spatial interpolation or kernel weighting is used to reconstruct the disparity-corrected continuous numerical field; Multi-scale aggregation of the reconstructed continuous numerical field yields historical cloud cover data.
3. The satellite cloud cover retrieval method based on spatiotemporal physics joint driving as described in claim 1, characterized in that, The spatiotemporal slice sample set is constructed in the following manner: Spatiotemporal registration and quality screening are performed based on the observation time difference between meteorological satellites and polar-orbiting satellites; among them, the observation time difference between historical observation data of meteorological satellites and cloud product data of polar-orbiting satellites after spatiotemporal registration and quality screening is less than a preset time threshold. In time Based on, backtracking A time-series spectral feature sequence was constructed using historical meteorological satellite observation data at consecutive moments. It is a positive integer greater than 1; The features in the time-series spectral feature sequence are spatially aligned with the historical cloud cover data to obtain the initial sample. The initial samples are cropped and sampled using a sliding window to obtain a spatiotemporal slice sample set that balances sample coverage and local spatial context information.
4. The satellite cloud cover retrieval method based on spatiotemporal physics joint driving as described in claim 1, characterized in that, The three-dimensional spatiotemporal encoder samples samples from the spatiotemporal slice sample set through a three-level encoder; wherein, the sampling strategy of each level encoder is that the time dimension is not downsampled and the spatial dimension is downsampled. Each level encoder includes a feature extraction module, a spatiotemporal attention gating module, and a pooling module; The feature extraction module is used to extract features; the spatiotemporal attention gating module is used to generate channel attention weights and spatial attention weights based on the statistical responses of the features at this level in the channel dimension and spatial dimension, respectively, and to use the channel attention weights and spatial attention weights to perform element-wise weighting on the output features of the convolution module at this level to obtain adaptively weighted features; wherein, the values of the channel attention weights and spatial attention weights are in the range of 0 to 1; the pooling module is used to perform spatial dimension downsampling on the adaptively weighted features at this level; The number of feature channels extracted by the first-stage encoder is less than the number of feature channels extracted by the second-stage encoder, and the number of feature channels extracted by the second-stage encoder is less than the number of feature channels extracted by the third-stage encoder. Furthermore, the feature size after spatial dimension downsampling by the first-stage encoder is greater than the feature size after spatial dimension downsampling by the second-stage encoder, and the feature size after spatial dimension downsampling by the second-stage encoder is greater than the feature size after spatial dimension downsampling by the third-stage encoder.
5. The satellite cloud cover inversion method based on spatiotemporal physics joint driving as described in claim 4, characterized in that, The multi-scale temporal aggregation bottleneck module includes a first temporal aggregation branch and a second temporal aggregation branch; The first temporal aggregation branch uses a preset time kernel and a first preset spatial kernel to perform temporal aggregation and channel compression on the output signal of the three-dimensional spatiotemporal encoder, and removes the time dimension of the compressed features to obtain the two-dimensional features of the first branch. The second temporal aggregation branch uses the preset time kernel and the second preset spatial kernel to perform temporal aggregation and channel compression on the output signal of the three-dimensional spatiotemporal encoder, and removes the time dimension of the compressed features to obtain the second branch two-dimensional features; The preset time kernel covers all time steps of the time-series spectral feature sequence, and the scale of the first preset spatial kernel is different from that of the second preset spatial kernel; The multi-scale temporal aggregation bottleneck module also includes a splicing module, which is used to fuse the two-dimensional features of the first branch and the two-dimensional features of the second branch to obtain two-dimensional bottleneck fusion features.
6. The satellite cloud cover retrieval method based on spatiotemporal physics joint driving as described in claim 5, characterized in that, The time-series fusion dimensionality reduction module with skip connections is connected to the three-level encoder and is used to convert the three-dimensional features output by each level encoder into two-dimensional skip connection features. The process of converting the target's 3D features output by the target encoder into 2D skip connection features includes: Jointly expand the temporal and channel dimensions of the target's three-dimensional features; The target 3D features after joint unfolding are sequentially subjected to 2D convolutional feature extraction, 2D normalization and nonlinear activation to obtain the 2D skip connection features corresponding to the target encoder. The target encoder can be any level encoder, and the target three-dimensional feature is the three-dimensional feature output by the target encoder.
7. The satellite cloud cover retrieval method based on spatiotemporal physics joint driving as described in claim 6, characterized in that, The fused deformable convolution-aligned two-dimensional decoder performs feature decoding through a three-level decoder to obtain high-resolution decoded features. The first-level decoder performs feature decoding in the following manner: The two-dimensional bottleneck fusion feature is upsampled, and the upsampled feature is fused with the two-dimensional skip connection feature corresponding to the third level encoder to obtain the first splicing feature. The first splicing feature is subjected to deformable convolutional alignment feature extraction to obtain the first alignment feature; The first alignment feature is input into the feature refining module to obtain the first-level decoding output feature; The second-level decoder performs feature decoding in the following manner: The first-level decoded output features are upsampled, and the upsampled features are fused with the two-dimensional skip connection features corresponding to the second-level encoder to obtain the second concatenation features. The second splicing feature is subjected to deformable convolutional alignment feature extraction to obtain the second alignment feature; The second alignment feature is input into the feature refining module to obtain the second-level decoding output feature; The third-level decoder performs feature decoding in the following manner: The second-level decoded output features are upsampled, and the upsampled features are fused with the two-dimensional skip connection features corresponding to the first-level encoder to obtain the third concatenation features. The third splicing feature is subjected to deformable convolutional alignment feature extraction to obtain the third alignment feature; The third alignment feature is input into the feature refining module to obtain the decoded output feature; The feature refinement module is used to reshape and compensate for details of the input features.
8. The satellite cloud cover retrieval method based on spatiotemporal physics joint driving as described in claim 7, characterized in that, Deformable convolutional alignment feature extraction is performed on the target splicing features, including: Offset prediction is performed using N cascaded feature extraction layers to obtain the target two-dimensional offset and modulation mask corresponding to the target splicing features; wherein, the target splicing features are any one of the first splicing features, the second splicing features, and the third splicing features; the objective function of the offset prediction is to maximize the effective receptive field, and N is a positive integer greater than 1; The target two-dimensional offset and modulation mask are used to drive deformable convolutional layers to align and extract target splicing features, thereby obtaining initial target alignment features. The initial alignment features of the target are subjected to two-dimensional normalization and nonlinear activation processing, and the processed initial alignment features are added to the target splicing features to obtain the target alignment features.
9. The satellite cloud cover retrieval method based on spatiotemporal physics joint driving as described in claim 1, characterized in that, The dual-branch multi-task output head includes a main task output branch and an auxiliary task output branch; The main task output branch is used to determine the two-dimensional continuous cloud cover prediction field based on the decoded output features. The auxiliary task output branch is used to determine the physical radiation reconstruction field based on the decoded output features during the model training phase.
10. A satellite cloud cover inversion device based on spatiotemporal physics joint driving, characterized in that, include: The cloud product data preprocessing module is configured to preprocess polar-orbiting satellite cloud product data to obtain historical cloud volume data; The dataset construction module is configured to register historical meteorological satellite observation data and polar-orbiting satellite cloud product data based on observation time, and construct a spatiotemporal slice sample set based on the registered historical meteorological satellite observation data and historical cloud cover data of polar-orbiting satellites; the spatiotemporal slice sample set uses time-series spectral feature sequences constructed based on historical meteorological satellite observation data as samples, and uses the historical cloud cover data at the corresponding time of each feature as sample labels. The model training module is configured to construct a pre-trained spatiotemporal hybrid feature extraction network, train the pre-trained spatiotemporal hybrid feature extraction network using the spatiotemporal slice sample set, and obtain the trained spatiotemporal hybrid feature extraction network. The inversion module is configured to acquire real-time observation data from meteorological satellites and use the trained spatiotemporal hybrid feature extraction network to invert satellite cloud cover based on the real-time observation data. The pre-trained spatiotemporal hybrid feature extraction network includes a three-dimensional spatiotemporal encoder, a multi-scale temporal aggregation bottleneck module, a skip-connected temporal fusion dimensionality reduction module, a two-dimensional decoder with fused deformable convolution alignment, and a dual-branch multi-task output head.